SRE role owning the reliability of Baseten's multi-cloud Kubernetes-based ML infrastructure, building observability tooling and automated remediation workflows.
Responsibilities
- ▹Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking
- ▹Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code
- ▹Author and improve runbooks for recurring failure patterns, structured for low-context, safe execution
- ▹Identify high-frequency failure patterns and convert them into automated or self-healing mitigations
- ▹Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management
- ▹Define and instrument SLOs and SLIs across customer workloads and internal services
Requirements
- ▹Extensive hands-on experience with Kubernetes (multi-cloud experience across EKS, GKE, or similar is a strong plus)
- ▹Experience building and maintaining scalable infrastructure
- ▹Strong foundation in observability tooling: metrics (VictoriaMetrics, Prometheus), logging (Loki, ELK), dashboards (Grafana), alerting pipelines
- ▹Experience with infrastructure-as-code (Terraform, Helm) and GitOps workflows (Flux CD, ArgoCD)
- ▹Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis
Nice to have
- ▹Observability-as-code experience
- ▹Familiarity with incident management platforms (incident.io or similar)
- ▹Curiosity about how ML models are deployed and served at scale
Soft skills
Comfort navigating ambiguity and making principled tradeoffsProcess- and escalation-minded alongside hands-on engineeringCross-functional collaboration with engineering, forward-deployed and product teams
What we offer
- ▹Competitive compensation including meaningful equity
- ▹100% coverage of medical, dental and vision insurance for employee and dependents
- ▹Flexible PTO including a company-wide winter break (closed between Christmas Eve and New Year's Day)
- ▹Paid parental leave
- ▹Fertility and family-building stipend through Carrot
- ▹Company-facilitated 401(k)
- ▹Exposure to a variety of ML startups, offering great learning and networking opportunities
About the company
Baseten powers mission-critical inference for some of the world's most dynamic AI companies, including Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By combining applied AI research, flexible infrastructure and seamless developer tooling, they help companies bring cutting-edge models into production. The company is growing fast and recently closed a $1.5B Series F round led by Altimeter Capital, Conviction Partners and Spark Capital.
Ähnliche Stellen

Stelle
Site Reliability Engineer 2
Kong
ArgocdClickhouse
+12
💰 Gehalt: keine Angabe
🔀 Hybrid
Washington
🗣️ EN

Stelle
Site Reliability Engineer
TeamViewer Germany GmbH
ArgocdDatadog
+9
85 862–107 327 €/Jahr
brutto
🏢 Vor Ort
Austin
🗣️ EN

Stelle
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
D Matrix
AI/ML
+9
💰 Gehalt: keine Angabe
🔀 Hybrid
Santa Clara
🗣️ EN

Stelle
Site Reliability Engineer
Axle
AI/ML
+18
120 207–133 086 €/Jahr
brutto
🏢 Vor Ort
Frederick
🗣️ EN

Stelle
DevOps Engineer Backend
Airapps
Argocd
+16
93 590–166 572 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
Site Reliability Engineer (SRE)
Airapps
CloudformationDatadog
+10
99 600–171 724 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN
