← Zurück zur Liste
Stelle · Principal

Staff Site Reliability Engineer

DevOps / SRE • Principal • Remote • Vollzeit Vereinigte Staaten USA

As a Staff Site Reliability Engineer, you'll proactively find and analyze reliability problems across Replit's stack, then design and implement software and systems that create step-function improvements — while mentoring the broader engineering team.

Responsibilities

  • Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions
  • Define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) with product and engineering teams
  • Act as a senior leader during high-impact incidents, guiding the team to rapid resolution and driving blameless post-mortems
  • Architect, build, and improve automation and infrastructure as code (Terraform, Pulumi), maintain CI/CD pipelines
  • Performance-tune and optimize large-scale cloud deployments with a deep focus on Kubernetes, Docker, and GCP
  • Dive deep into debugging extremely difficult technical problems across the stack and implement long-term fixes
  • Review feature and system designs from across the company, owning reliability, scalability, security, and operational integrity
  • Educate and mentor the broader engineering team, making reliability a core value
  • Write high-quality, well-tested code in Python or Go for internal tools and third-party integrations

Requirements

  • 8-10 years of experience in Site Reliability Engineering or similar roles
  • Strong programming skills in Python or Go, writing high-quality, well-tested code
  • Deep understanding of distributed systems, with experience designing, building, scaling, and maintaining production services
  • Deep experience with container orchestration platforms, specifically Kubernetes, and cloud-native technologies
  • Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions
  • Strong incident management skills with extensive experience leading incident response for complex systems
  • Experience with infrastructure as code (e.g. Terraform, Pulumi) and configuration management tools
  • Excellent written and verbal communication, with mentoring experience from junior to principal levels

Nice to have

  • Deep experience with Google Cloud Platform (GCP) services and tools
  • Expert-level knowledge of modern observability platforms (Prometheus, Grafana, Datadog, OpenTelemetry)
  • Experience designing and building reliable systems capable of handling high throughput and low latency
  • Significant experience with Go and Terraform
  • Familiarity with rapid-growth, startup environments
  • Experience writing company-facing blog posts and training materials

Soft skills

Staff-level leadership presence in mentoring and guidanceExcellent communication and a bias toward open, transparent cultural practicesCritical thinking under pressurePassion for making software creation accessible

What we offer

  • Competitive salary and equity
  • 401(k) program with a 4% match (US only)
  • Health, dental, vision, and life insurance
  • Short-term and long-term disability
  • Paid parental, medical, and caregiver leave
  • Flexible Time Off (FTO) plus holidays
  • Commuter benefits (in-office only)
  • Monthly wellness stipend
  • Autonomous work environment
  • In-office set-up reimbursement (in-office only)
  • Quarterly team gatherings
  • In-office amenities (in-office only)

About the company

Replit is the agentic software creation platform that enables anyone to build applications using natural language, serving millions of developers worldwide and democratizing software development.

Ähnliche Stellen