← Back to list
Job · Senior

Site Reliability Engineer

DevOps / SRE • Senior • Remote • Full-time Spain Spain

HostPapa's CloudBlue business unit is hiring a Site Reliability Engineer to ensure the reliability, scalability and observability of its multi-tenant SaaS platforms used by service providers worldwide.

Responsibilities

  • Define SLIs, SLOs and error budgets for critical services
  • Influence system architecture for reliability, scalability and operability
  • Reduce operational toil through automation
  • Design and operate the observability stack (Datadog, Grafana, Elastic Stack)
  • Develop alerting strategies and dashboards
  • Design high-availability architectures: redundancy, failover, disaster recovery
  • Conduct capacity planning, load testing and performance optimization
  • Act as senior responder during production incidents
  • Own blameless post-mortems and drive improvements
  • Improve reliability of Kubernetes-based platforms
  • Partner on deployment safety and rollback strategies
  • Maintain runbooks and operational documentation

Requirements

  • 3+ years as an SRE, DevOps Engineer or Production Engineer
  • Proven experience operating highly available, multi-tenant SaaS platforms
  • Hands-on experience with observability and monitoring tools (Datadog, Grafana, Elasticsearch/Kibana)
  • Solid understanding of Linux, networking and distributed systems
  • Experience with containerized environments

Soft skills

Incident leadershipCross-team communicationProactive problem-solving

About the company

HostPapa is a fast-growing web hosting company operating in 39 countries. Its CloudBlue business powers cloud commerce for major Telcos, distributors and MSPs.

Similar jobs