SRE role owning the reliability of Baseten's multi-cloud Kubernetes-based ML infrastructure, building observability tooling and automated remediation workflows.
Responsibilities
- ▹Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking
- ▹Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code
- ▹Author and improve runbooks for recurring failure patterns, structured for low-context, safe execution
- ▹Identify high-frequency failure patterns and convert them into automated or self-healing mitigations
- ▹Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management
- ▹Define and instrument SLOs and SLIs across customer workloads and internal services
Requirements
- ▹Extensive hands-on experience with Kubernetes (multi-cloud experience across EKS, GKE, or similar is a strong plus)
- ▹Experience building and maintaining scalable infrastructure
- ▹Strong foundation in observability tooling: metrics (VictoriaMetrics, Prometheus), logging (Loki, ELK), dashboards (Grafana), alerting pipelines
- ▹Experience with infrastructure-as-code (Terraform, Helm) and GitOps workflows (Flux CD, ArgoCD)
- ▹Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis
Nice to have
- ▹Observability-as-code experience
- ▹Familiarity with incident management platforms (incident.io or similar)
- ▹Curiosity about how ML models are deployed and served at scale
Soft skills
Comfort navigating ambiguity and making principled tradeoffsProcess- and escalation-minded alongside hands-on engineeringCross-functional collaboration with engineering, forward-deployed and product teams
What we offer
- ▹Competitive compensation including meaningful equity
- ▹100% coverage of medical, dental and vision insurance for employee and dependents
- ▹Flexible PTO including a company-wide winter break (closed between Christmas Eve and New Year's Day)
- ▹Paid parental leave
- ▹Fertility and family-building stipend through Carrot
- ▹Company-facilitated 401(k)
- ▹Exposure to a variety of ML startups, offering great learning and networking opportunities
About the company
Baseten powers mission-critical inference for some of the world's most dynamic AI companies, including Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By combining applied AI research, flexible infrastructure and seamless developer tooling, they help companies bring cutting-edge models into production. The company is growing fast and recently closed a $1.5B Series F round led by Altimeter Capital, Conviction Partners and Spark Capital.
Similar jobs

Job
Site Reliability Engineer 2
Kong
ArgocdClickhouse
+12
💰 Salary: not specified
🔀 Hybrid
Washington
🗣️ EN

Job
Site Reliability Engineer
TeamViewer Germany GmbH
ArgocdDatadog
+9
$100,000–$125,000/yr
gross
🏢 On-site
Austin
🗣️ EN

Job
Site Reliability Engineer - AI Accelerator Infrastructure - Contract
D Matrix
AI/ML
+9
💰 Salary: not specified
🔀 Hybrid
Santa Clara
🗣️ EN

Job
Site Reliability Engineer
Axle
AI/ML
+18
$140,000–$155,000/yr
gross
🏢 On-site
Frederick
🗣️ EN

Job
DevOps Engineer Backend
Airapps
Argocd
+16
$109,000–$194,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN

Job
Site Reliability Engineer (SRE)
Airapps
CloudformationDatadog
+10
$116,000–$200,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN
