← Back to list
Job · Senior

Senior Site Reliability Engineer - SDN

DevOps / SRE • Senior • Remote • Full-time United States USA

An SRE engineer scaling Lambda's cloud platform, operating and improving the multi-tenant SDN networking infrastructure and Kubernetes-based control-plane services while reducing operational toil through automation.

Responsibilities

  • Operate and scale Lambda's multi-tenant cloud networking platform and SDN infrastructure
  • Operate and improve Kubernetes-based control-plane services and dataplane software on SmartNICs
  • Develop tooling and automation to reduce operational toil and improve reliability
  • Collaborate with software, platform and networking teams to improve reliability and deployment workflows
  • Deploy and maintain network monitoring, observability and management tools
  • Improve deployment safety through CI/CD pipelines, GitOps workflows, testing and progressive rollouts
  • Drive operational excellence through observability, incident management, capacity planning, postmortems and on-call

Requirements

  • 5+ years in Site Reliability Engineering, Production Engineering or similar
  • Experience operating large-scale distributed systems in production
  • Experience with Kubernetes application lifecycle management, upgrades and troubleshooting
  • Experience with on-call rotations and incident response
  • Strong troubleshooting across Linux, Kubernetes, distributed systems and networking
  • Experience with observability platforms, monitoring, alerting and metrics
  • Comfortable on the Linux command line with solid understanding of the Linux networking stack
  • Experience with multi-datacenter and hybrid cloud environments
  • Experience automating infrastructure with Python, Ansible or similar
  • Experience designing and operating CI/CD and GitOps deployment workflows

Nice to have

  • Experience building/operating Software Defined Networks (OpenStack Neutron, OVN, OVS)
  • Experience operating production-scale SDNs in cloud environments (e.g. AWS VPC-like networking)
  • Software development in Go and/or Python (C a plus)
  • Automating infrastructure/network configuration with Kubernetes, Helm, Terraform, Ansible
  • Deep understanding of the Linux networking stack, SR-IOV, DPDK
  • Understanding of the SDN ecosystem and modern cloud networking architectures
  • Diagnosing complex production issues across infrastructure, networking and application layers

Soft skills

Cross-team collaborationProblem-solving and troubleshooting under pressureOperational discipline, on-call ownership

What we offer

  • Generous cash and equity compensation
  • Health, dental and vision coverage
  • Wellness and commuter stipends
  • 401k plan with 2% company match
  • Flexible paid time off

About the company

Engineering at Lambda builds and scales its cloud offering, including the Lambda website, cloud APIs and systems, and internal tooling for deployment, management and maintenance.

Similar jobs