← Zurück zur Liste
Stelle · Senior

Senior Site Reliability Engineer

DevOps / SRE • Senior • Vor Ort • Vollzeit Bulgarien Bulgarien

Remote Senior SRE role on GoDaddy's Global Sustaining Engineering team, owning production services end-to-end and mentoring junior SREs while applying LLM-driven tooling to incident response.

Responsibilities

  • Design, implement, and operate scalable, highly available production services while diagnosing and resolving complex infrastructure, network, and application issues
  • Build and maintain alerting pipelines, dashboards, and SLO-driven monitoring strategies using Icinga, Prometheus, and Grafana
  • Lead incident response end-to-end, performing root-cause analysis, authoring blameless post-mortems, and driving corrective actions to closure
  • Develop and extend Infrastructure as Code coverage and build internal tooling to eliminate manual, repetitive operational work
  • Mentor SRE I and SRE II engineers through code reviews, debugging sessions, and knowledge-sharing talks
  • Apply LLM-driven log analysis, anomaly detection, and generative AI tools to accelerate incident response and runbook creation, validating all outputs before use

Requirements

  • 6+ years of professional experience in Site Reliability Engineering or Platform Engineering with demonstrated success leading organization-wide reliability programs
  • Deep hands-on expertise with Kubernetes (deployments, operators, custom resources) and Docker in production environments
  • Advanced Linux experience solving problems involving kernel internals, TCP/IP, DNS, and load balancers
  • Proficiency in Python for production-grade automation and scripting, with working knowledge of Bash
  • Expertise in Ansible and at least one additional Infrastructure as Code tool such as Terraform or Pulumi
  • Hands-on mastery of Icinga, Prometheus, and Grafana
  • Understanding of large language models, embeddings, and basic machine learning pipelines, with the ability to evaluate and integrate AI-ops tools

Nice to have

  • Experience building and maintaining CI/CD pipelines using Jenkins, GitLab CI, or GitHub Actions
  • Demonstrated experience defining and managing Service Level Objectives, Service Level Indicators, and Service Level Agreements across production

What we offer

  • Paid time off
  • Retirement savings (e.g. 401k, pension schemes)
  • Bonus/incentive eligibility
  • Equity grants
  • Participation in the employee stock purchase plan
  • Competitive health benefits
  • Other family-friendly benefits, including parental leave

About the company

At GoDaddy the future of work looks different for each team: some in-office, some hybrid, some fully remote. This is a remote position on the Global Sustaining Engineering team, which sits at the intersection of software engineering and infrastructure, ensuring the services customers depend on are fast, resilient, and always available.

Ähnliche Stellen