← Back to list
Job · Senior

Senior Backend Engineer, Managed Kubernetes & Slurm

Backend Developer • Senior • On-site • Full-time United States New York, USA

Lightning AI, the company behind PyTorch Lightning, is hiring a Senior Backend Engineer for its Managed Services team, building the control planes, backend services, and automation that provision, operate, and scale Kubernetes and Slurm clusters across its GPU fleet. Based in San Francisco, NYC, or Seattle with a hybrid schedule (minimum 2 office days/week).

Responsibilities

  • Design, build, and operate backend services in Go or Python for Lightning AI's managed infrastructure platform
  • Develop control plane services that provision, orchestrate, and manage Kubernetes and Slurm clusters
  • Build distributed systems automating cluster lifecycle management, workload scheduling, and infrastructure provisioning
  • Develop platform capabilities using Kubernetes APIs, controllers, and operators
  • Improve reliability, scalability, security, and observability of the managed platform through automation
  • Diagnose and resolve complex production issues across Kubernetes, distributed systems, and networking
  • Collaborate with infrastructure, AI, and platform engineering teams
  • Contribute to technical design, architecture, mentoring, and on-call operations

Requirements

  • Significant professional experience designing, building, and operating production backend systems in Go or Python
  • Deep hands-on experience with Kubernetes or Slurm, including large-scale production environments
  • Strong understanding of distributed systems and cloud-native architectures
  • Experience designing scalable backend services, APIs, and automation for infrastructure or platform operations
  • Strong understanding of networking, storage, and cloud infrastructure fundamentals
  • Familiarity with observability, CI/CD, testing, and incident response
  • Ability to own complex technical projects while collaborating across engineering teams

Nice to have

  • Kubernetes platform development experience (operators, controllers, CRDs, control plane components)
  • Slurm administration, scheduling, or HPC environment experience
  • Infrastructure-as-Code or GitOps tooling (Terraform, Crossplane, Helm, Argo CD, Flux, Kustomize)
  • GPU infrastructure, AI platforms, or large-scale compute environments
  • Kubernetes networking, storage, or multi-tenant platform architecture
  • Event-driven systems, gRPC, or distributed messaging
  • Building CI/CD systems for production environments
  • Contributions to cloud-native or open source infrastructure projects

Soft skills

Moving with urgency and owning outcomesOpen, direct communicationBuilding and empowering strong teamsContinuous self-improvement and openness to feedbackLong-term, scalable thinking

What we offer

  • Comprehensive health coverage (medical, dental, vision)
  • Meaningful equity (RSUs)
  • 401(k) matching (US) / pension contributions (UK)
  • Unlimited PTO, company and floating holidays
  • Two-week company-wide winter break
  • Paid parental and family leave
  • Annual learning and development allowance
  • Wellness and work-from-home stipends
  • Four weeks of paid sabbatical after four years of service
  • Flexible schedules and hybrid work model
  • Complimentary meals at office hubs

About the company

Lightning AI is the company behind PyTorch Lightning, building an end-to-end platform for developing, training, and deploying AI systems. Following its merger with Voltage Park, it combines developer-first tooling with cost-efficient, large-scale compute, running an AI-native cloud powered by 25,000+ H100, B200, and GB300 GPUs.

Similar jobs