← Zurück zur Liste
Stelle · Senior

Sr. Site Reliability Engineer - SRE

DevOps / SRE • Senior • Remote • Vollzeit Spanien Spanien

A growing SRE team is looking for an experienced engineer to ensure the reliability, scalability and performance of mission-critical services, taking a leading role in automation and reducing operational toil.

Responsibilities

  • Design, implement and maintain highly available, scalable, resilient systems
  • Act as SME for observability: monitoring, alerting, logging, tracing, dashboards, synthetic testing
  • Develop robust, maintainable software and self-service tooling to automate operational tasks
  • Identify and eliminate operational toil through automation and process improvement
  • Lead incident response, participate in on-call rotations, drive blameless post-mortems
  • Define, implement and track SLIs, SLOs and error budgets
  • Leverage infrastructure as code, GitOps and CI/CD automation with Terraform, Flux and GitHub Actions
  • Provide reliability expertise during system design reviews
  • Document processes, build runbooks, mentor engineers
  • Leverage AI responsibly to accelerate investigations, documentation and toil reduction

Requirements

  • Demonstrated experience operating and improving production systems at scale in an SRE, Production Engineering or Platform Engineering role
  • Ability to rapidly build accurate mental models of complex distributed systems across infrastructure, applications, networking, identity and observability
  • Strong troubleshooting skills with a methodical, evidence-driven approach to incident response and root cause analysis
  • Experience defining and using SLIs, SLOs and error budgets to guide reliability decisions
  • Excellent written and verbal communication skills
  • Kubernetes platforms including Amazon EKS, and service mesh technologies such as Istio
  • Cloud infrastructure and services within AWS
  • Identity and access management systems including Auth0 and AWS IAM
  • Networking fundamentals including DNS, load balancing, routing, TLS, connectivity troubleshooting
  • GitOps workflows and infrastructure automation using Flux and Terraform
  • Observability platforms and practices
  • CI/CD systems and engineering workflows

Soft skills

Prioritizes service stability and customer impact during incidentsStays calm under pressure, gathers facts, communicates clearlySystems-thinking approach to problem-solvingBalances short-term remediation with long-term reliability improvements

Ähnliche Stellen