← Zurück zur Liste

Stelle
· Senior
Senior SRE Engineer
DevOps / SRE
• Senior
• Hybrid
• Vollzeit
•
Barcelona, Spanien
Senior Site Reliability Engineer for Codeway's growing platform, owning reliability, performance and security: defining SLIs, SLOs and error budgets, operating a multi-cluster Kubernetes environment and leading incident response. The SRE practice is still early in its maturity, so you will help build its foundations.
Responsibilities
- ▹Define, instrument and report on SLIs, SLOs and error budgets across critical services, so reliability decisions are driven by data rather than opinion
- ▹Own observability end-to-end (metrics, logs, traces, dashboards and alerting) and drive measurable reductions in detection and resolution times
- ▹Reduce alert noise and false positives so on-call engineers can trust what wakes them up
- ▹Run reliability reviews and an error-budget policy that shapes how teams prioritize between shipping and stability
- ▹Operate, scale and upgrade the multi-cluster Kubernetes (GKE) environment: cluster lifecycle, autoscaling, networking, ingress and resource management
- ▹Act as the deep-expertise escalation point for cluster and platform issues across dozens of services
- ▹Own capacity planning, performance and cloud cost efficiency, balancing spend against reliability targets
- ▹Build self-service platform tooling that lets product teams move quickly without needing to become infrastructure experts
- ▹Embed security into the platform through RBAC and least-privilege, secrets management, image and dependency scanning, network policies and a disciplined patching cadence
- ▹Partner with the security function on vulnerability remediation, audit readiness and secure-by-default infrastructure
- ▹Own disaster recovery: define and regularly validate RTO/RPO targets through DR drills and failure testing
- ▹Contribute to architecture and production-readiness reviews so reliability and security are designed in
- ▹Lead the on-call rotation and act as incident commander during production incidents
- ▹Run blameless postmortems, quantify impact and track corrective actions through to closure
- ▹Build and maintain Infrastructure as Code (Terraform) and CI/CD pipelines, enforcing GitOps and progressive delivery with automated rollbacks
- ▹Systematically identify, measure and eliminate operational toil through automation
- ▹Within the first 12 months, establish shared measurement with the Security team (vulnerability remediation times, patching cadence and coverage, access review completion, security incident detection and response times) and a single agreed view of operational health across engineering and security
- ▹Connect platform performance to the business it supports: the reliability of the marketing and growth stack, the speed and safety of the path from commit to production, and the infrastructure cost behind serving and acquiring users
Requirements
- ▹Experience operating high-traffic, always-on production systems at meaningful scale, typically gained over 5-8 years in SRE, Platform or DevOps roles
- ▹Hands-on production Kubernetes experience: you have run clusters day to day, through upgrades, autoscaling and real troubleshooting under load
- ▹A strong cloud engineering background, along with solid Linux and networking fundamentals
- ▹A track record of defining and operating with SLOs and error budgets, and comfort being measured on reliability outcomes
- ▹Experience with Infrastructure as Code and CI/CD pipeline design
- ▹Depth in observability tooling: instrumentation, dashboarding and alert design
- ▹A genuine security-first mindset, where least-privilege, secrets hygiene and vulnerability management are habits
- ▹Scripting and automation fluency in at least one language, used to build tooling and remove toil
- ▹Incident-command experience: owning on-call, running blameless postmortems and driving resolution times down over time
- ▹Ability to communicate clearly with both engineers and leadership, especially under pressure
Nice to have
- ▹Experience with high-scale consumer or mobile app backends, or with AI/ML inference workloads and their scaling characteristics
- ▹Experience with GitOps and progressive-delivery patterns such as canary and blue-green rollouts
- ▹Familiarity with service mesh, API gateways, or multi-region and multi-cluster topologies
- ▹Cloud cost optimization and FinOps discipline at scale
- ▹Exposure to compliance initiatives (SOC 2, ISO 27001, GDPR) and broader DevSecOps practice
- ▹Chaos engineering or resilience testing experience
- ▹Relevant certifications in Kubernetes, cloud or DevOps disciplines
- ▹Experience supporting many independent services and teams concurrently in a fast-shipping, product-led environment
Soft skills
Clear communication with engineers and leadershipCommunication under pressureInitiative in building standards and processesOwnershipProactive, security-first mindset
What we offer
- ▹Competitive compensation package
- ▹Meal compensation
- ▹Full health benefits: unlimited private health insurance and coverage of the HPV vaccine
- ▹Pet adoption support: primary healthcare expenses, parasite vaccinations and microchip costs within the first year after adoption
- ▹Tech stack: MacBook, iPhone 15 Pro, mouse, keyboard, adjustable desk with a 4K screen and any other gadget needed for the job
- ▹Sport activities support (gym membership)
- ▹Flexible schedule: productivity matters, not tracking every minute on site
- ▹English course support
- ▹Office in the heart of Barcelona in the Edifici Estel
- ▹In-office coffee shop (Codebrew) with healthy snacks at all hours
- ▹Free breakfast and lunch at Codebrew
- ▹No dress code
- ▹Gaming area with a PS5 corner
- ▹Software support: subscription to any software needed for the job
- ▹Public transportation support through an additional monthly compensation
About the company
Codeway is a global consumer tech company with more than 400M users worldwide. Since 2020 it has built and scaled 60+ mobile apps across creativity, productivity, wellness, language learning and entertainment, including Retake AI, Cleanup, Learna and DramaPops. In 2024 it became the most downloaded app publisher on iOS. The team of 300+ people works across Istanbul and Barcelona.
Ähnliche Stellen

Stelle
· Senior
Senior DevOps Engineer
Lemon.io
AI/ML
+38
💰 Gehalt: keine Angabe
🌍 Remote

Stelle
· Senior
DevOps Engineer - Senior- EY GDS Spain - Hybrid
EY
+9
💰 Gehalt: keine Angabe
🏢 Vor Ort
Malaga
🗣️ EN

Stelle
· Senior
Site Reliability Engineer - DevSecOps Engineer
Kyndryl
+7
💰 Gehalt: keine Angabe
🔀 Hybrid
Madrid
🗣️ EN

Stelle
· Senior
Senior DevOps Engineer for Industrial AI Cloud (m/f/d)
T-Systems Iberia
Github Actions
+8
💰 Gehalt: keine Angabe
🏢 Vor Ort
A Coruña
🗣️ EN

Stelle
· Senior
Senior DevOps Engineer for Industrial AI Cloud (m/f/d)
T-Systems Iberia
Github Actions
+8
💰 Gehalt: keine Angabe
🏢 Vor Ort
A Coruña
🗣️ EN

Stelle
· Senior
Senior DevOps Engineer for Industrial AI Cloud (m/f/d)
T-Systems Iberia
Github Actions
+8
💰 Gehalt: keine Angabe
🏢 Vor Ort
A Coruña
🗣️ EN