← Back to list
Job · Senior

Senior Software Engineer - Core Cloud Platform

Software Engineer • Senior • Remote • Full-time United States San Francisco Office (Fremont St), USA

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure. As a Senior Software Engineer on the Core Cloud Platform team, you'll build the control-plane systems powering Lambda's GPU cloud — APIs, workflows, orchestration layers, and schedulers turning physical GPU infrastructure into reliable, customer-facing cloud capacity.

Responsibilities

  • Build and operate core cloud platform services for compute lifecycle, bare metal hosts, capacity, placement, and maintenance workflows
  • Design reliable APIs, backend services, state machines, and orchestration systems
  • Work on bare metal lifecycle systems: launch, terminate, restart/reboot, host reclaim, validation, quarantine, return-to-pool
  • Improve deployment, observability, testing, alerting, runbooks, and operational readiness
  • Debug complex production issues across distributed services, infrastructure, and networking
  • Partner with infrastructure, networking, fleet, security, support, and product teams
  • Contribute to architecture, design docs, code reviews, incident follow-through, and mentoring

Requirements

  • Bachelor's degree or equivalent working experience
  • 6+ years professional software engineering experience in production backend/distributed systems
  • Strong in Python, Go, or similar backend/systems language
  • Experience designing and operating APIs, workflow engines, schedulers, or orchestration services
  • Understanding of reliability fundamentals: fault tolerance, idempotency, retries, state machines, failure handling
  • Experience with cloud or cloud-like infrastructure primitives (compute, networking, storage, capacity, identity, fleet ops)
  • Comfortable with Linux, containers, Kubernetes, infrastructure automation, service deployment
  • Owned production services, on-call experience, operational improvements
  • Care about testability, CI/CD, observability, metrics, logging, alerting
  • Ability to drive ambiguous infrastructure problems to clear designs and implementation
  • Clear communication across engineering, product, support, infrastructure, and leadership

Nice to have

  • Experience building cloud control planes, compute platforms, schedulers, or orchestration systems
  • Experience with bare metal, GPU infrastructure, HPC, Kubernetes, Slurm, or large-scale AI/ML infrastructure
  • Host lifecycle, provisioning, validation, firmware, BMC/Redfish, fleet management experience
  • Familiarity with Temporal, Airflow, event-driven systems, or durable workflow orchestration
  • Networking experience with VPCs, firewalls, SDN, InfiniBand, routing
  • Security-minded engineering: identity, authorization, audit logging, attestation, tenant isolation

Soft skills

Passion for distributed systems and cloud infrastructureCommitment to operational excellenceClear, proactive communication about risks and tradeoffsHigh standards while moving fast

What we offer

  • Generous cash & equity compensation
  • Health, dental, vision coverage for employees and dependents
  • Wellness and commuter stipends for select roles
  • 401(k) with 2% company match (USA)
  • Flexible paid time off

About the company

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers from AI researchers to enterprises and hyperscalers. Founded in 2012, with 500+ employees and growing fast.

Education: Alapdiploma vagy azzal egyenértékű szakmai tapasztalat

Similar jobs