← Back to list
Job · Senior

Senior Site Reliability Engineer, DGX Cloud

DevOps / SRE • Senior • Remote • Full-time • Switzerland Switzerland

NVIDIA's DGX Cloud team is looking for an experienced Senior Site Reliability Engineer to keep high-performance AI clusters reliable for researchers and enterprise clients worldwide.

Responsibilities

  • ▹Support operational and reliability aspects of large-scale Kubernetes clusters
  • ▹Define SLOs/SLIs and monitor error budgets
  • ▹Support services through launch reviews and capacity management
  • ▹Operate GPU workloads across AWS, GCP, Azure, OCI and private clouds
  • ▹Scale and evolve systems through automation
  • ▹Lead triage and root-cause analysis of high-severity incidents
  • ▹Participate in on-call rotation

Requirements

  • ▹BS in Computer Science or related field, or equivalent experience
  • ▹10+ years operating production services
  • ▹Expert-level Kubernetes administration and containerization
  • ▹Infrastructure automation tools (Terraform, Ansible, Chef, Puppet)
  • ▹Proficiency in a high-level language (e.g. Python, Go)
  • ▹In-depth Linux, networking (TCP/IP) and cloud security knowledge
  • ▹SRE principles: SLOs, SLIs, error budgets, incident handling
  • ▹Building observability stacks (OpenTelemetry, Prometheus, Grafana, ELK, Splunk)

Nice to have

  • ▹Operating GPU-accelerated clusters with KubeVirt in production
  • ▹Applying generative AI to reduce operational toil
  • ▹Workflow orchestration platforms (Temporal, Cadence, Airflow, Argo Workflows)

Soft skills

Balanced incident response, blameless postmortemsCross-functional collaboration

About the company

NVIDIA has been transforming computer graphics and accelerated computing for 25+ years; DGX Cloud delivers a fully managed AI platform on major cloud providers.

Similar jobs