An SRE role on Runpod's Reliability team, responsible for the availability, performance, and operational excellence of the global platform through SLO design, observability, and automation.
Responsibilities
- ▹Define and implement SLIs/SLOs for critical services
- ▹Lead incident response and coordinate cross-team mitigation efforts
- ▹Conduct blameless postmortems and ensure corrective actions are completed
- ▹Perform production readiness reviews for new services and features
- ▹Identify systemic risks and drive preventative improvements
- ▹Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana)
- ▹Improve signal-to-noise ratio in alerts and reduce alert fatigue
- ▹Automate recurring operational workflows using Python, Go, and Bash
- ▹Improve deployment safety and strengthen CI/CD reliability
- ▹Partner with engineering teams to improve system resilience and fault tolerance
Requirements
- ▹5+ years of experience in SRE, Reliability Engineering, or Production Engineering
- ▹Strong Linux systems and networking expertise
- ▹Experience managing containerized production systems
- ▹Strong understanding of distributed systems and failure modes
- ▹Experience defining and managing SLIs/SLOs
- ▹Proven incident response and postmortem leadership experience
- ▹Strong scripting or programming skills
- ▹Experience with monitoring and alerting systems
- ▹Excellent written communication skills
- ▹Successful completion of a background check
Nice to have
- ▹Experience with GPU infrastructure or AI/ML platforms
- ▹Experience improving reliability in high-growth or large-scale environments
- ▹Familiarity with GPU observability tooling
- ▹Experience with Infrastructure as Code
- ▹Experience working in startup environments
- ▹Experience building internal reliability platforms or frameworks
Soft skills
Proactive problem solvingAutomation-first thinkingStrong ownership of production systems
What we offer
- ▹Competitive base pay of $150,000–$200,000
- ▹Meaningful equity in a fast-growing company
- ▹Generous medical, dental & vision plans
- ▹Flexible PTO
- ▹Remote-first, Slack-driven collaborative culture
About the company
Runpod is the AI Developer Cloud, used by more than one million developers to experiment, train, and scale AI, with over 20 billion inference requests processed. The company closed a $100M Series A in June 2026.
Similar jobs

Job
Site Reliability Engineer 2
Kong
ArgocdClickhouse
+12
💰 Salary: not specified
🔀 Hybrid
Washington
🗣️ EN

Job
DevOps Director
Air
CloudformationCypress
+15
💰 Salary: not specified
🏢 On-site
Pittsburgh
🗣️ EN

Job
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
D Matrix
AI/ML
+9
💰 Salary: not specified
🔀 Hybrid
Santa Clara
🗣️ EN

Job
Site Reliability Engineer
Axle
AI/ML
+18
$140,000–$155,000/yr
gross
🏢 On-site
Frederick
🗣️ EN
Job
Site Reliability Engineer (SRE)
Bright Vision Technologies
Datadog
+8
$100,000–$150,000/yr
gross
🌍 Remote
New Jersey 🇺🇸 United States of America
🗣️ EN

Job
DevOps Engineer Backend
Airapps
Argocd
+16
$109,000–$194,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN
