← Back to list
Job · Senior

Senior Site Reliability Engineer

DevOps / SRE • Senior • Remote • Full-time United States USA

Join Replit's Site Reliability Engineering team to ensure the reliability, scalability, and performance of infrastructure serving millions of developers worldwide, bridging development and operations through automation and best practices.

Responsibilities

  • Design and implement comprehensive monitoring and alerting systems using modern observability tools
  • Create dashboards and metrics that provide real-time visibility into system health and performance
  • Implement logging strategies that enable quick problem identification and resolution
  • Architect and implement infrastructure automation using tools like Terraform, Ansible, or Pulumi
  • Design and maintain CI/CD pipelines that enable reliable and consistent deployments
  • Create self-healing systems that can automatically respond to common failure scenarios
  • Define and implement SLOs and SLIs together with product and engineering teams
  • Lead incident response efforts, conduct thorough post-mortems, and implement preventative improvements
  • Develop and maintain runbooks for critical services
  • Build tools and processes that reduce Mean Time To Recovery (MTTR)
  • Identify and resolve performance bottlenecks, implement capacity planning, reduce latency across global regions

Requirements

  • 4-8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering)
  • Strong programming skills in languages commonly used for automation (Python, Go, or similar)
  • Deep understanding of distributed systems
  • Experience with container orchestration platforms (Kubernetes) and cloud-native technologies
  • Proven track record of implementing and maintaining monitoring/observability solutions
  • Strong incident management skills with experience leading incident response
  • Experience with infrastructure as code and configuration management tools

Nice to have

  • Experience with Google Cloud Platform (GCP) services and tools
  • Knowledge of modern observability platforms (Prometheus, Grafana, Datadog, etc.)

Soft skills

Problem-solving mindset for complex operational challengesSelf-directed and autonomous, while collaborating effectively cross-functionallyStrong communication skills for both technical and non-technical audiencesContinuous learning, staying current with industry best practicesStrong belief in automation

What we offer

  • Competitive salary and equity
  • 401(k) program with a 4% match (US only)
  • Health, dental, vision, and life insurance
  • Short-term and long-term disability
  • Paid parental, medical, and caregiver leave
  • Flexible Time Off (FTO) plus holidays
  • Commuter benefits (in-office only)
  • Monthly wellness stipend
  • Autonomous work environment
  • In-office set-up reimbursement (in-office only)
  • Quarterly team gatherings
  • In-office amenities (in-office only)

About the company

Replit is the agentic software creation platform that enables anyone to build applications using natural language, serving millions of developers worldwide and democratizing software development.

Similar jobs