An SRE role on Runpod's Reliability team, responsible for the availability, performance, and operational excellence of the global platform through SLO design, observability, and automation.
Responsibilities
- ▹Define and implement SLIs/SLOs for critical services
- ▹Lead incident response and coordinate cross-team mitigation efforts
- ▹Conduct blameless postmortems and ensure corrective actions are completed
- ▹Perform production readiness reviews for new services and features
- ▹Identify systemic risks and drive preventative improvements
- ▹Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana)
- ▹Improve signal-to-noise ratio in alerts and reduce alert fatigue
- ▹Automate recurring operational workflows using Python, Go, and Bash
- ▹Improve deployment safety and strengthen CI/CD reliability
- ▹Partner with engineering teams to improve system resilience and fault tolerance
Requirements
- ▹5+ years of experience in SRE, Reliability Engineering, or Production Engineering
- ▹Strong Linux systems and networking expertise
- ▹Experience managing containerized production systems
- ▹Strong understanding of distributed systems and failure modes
- ▹Experience defining and managing SLIs/SLOs
- ▹Proven incident response and postmortem leadership experience
- ▹Strong scripting or programming skills
- ▹Experience with monitoring and alerting systems
- ▹Excellent written communication skills
- ▹Successful completion of a background check
Nice to have
- ▹Experience with GPU infrastructure or AI/ML platforms
- ▹Experience improving reliability in high-growth or large-scale environments
- ▹Familiarity with GPU observability tooling
- ▹Experience with Infrastructure as Code
- ▹Experience working in startup environments
- ▹Experience building internal reliability platforms or frameworks
Soft skills
Proactive problem solvingAutomation-first thinkingStrong ownership of production systems
What we offer
- ▹Competitive base pay of $150,000–$200,000
- ▹Meaningful equity in a fast-growing company
- ▹Generous medical, dental & vision plans
- ▹Flexible PTO
- ▹Remote-first, Slack-driven collaborative culture
About the company
Runpod is the AI Developer Cloud, used by more than one million developers to experiment, train, and scale AI, with over 20 billion inference requests processed. The company closed a $100M Series A in June 2026.
Ähnliche Stellen

Stelle
Site Reliability Engineer 2
Kong
ArgocdClickhouse
+12
💰 Gehalt: keine Angabe
🔀 Hybrid
Washington
🗣️ EN

Stelle
DevOps Director
Air
CloudformationCypress
+15
💰 Gehalt: keine Angabe
🏢 Vor Ort
Pittsburgh
🗣️ EN

Stelle
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
D Matrix
AI/ML
+9
💰 Gehalt: keine Angabe
🔀 Hybrid
Santa Clara
🗣️ EN

Stelle
Site Reliability Engineer
Axle
AI/ML
+18
120 207–133 086 €/Jahr
brutto
🏢 Vor Ort
Frederick
🗣️ EN
Stelle
Site Reliability Engineer (SRE)
Bright Vision Technologies
Datadog
+8
85 862–128 793 €/Jahr
brutto
🌍 Remote
New Jersey 🇺🇸 United States of America
🗣️ EN

Stelle
DevOps Engineer Backend
Airapps
Argocd
+16
93 590–166 572 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN
