← Zurück zur Liste

Stelle
· Senior
Senior Site Reliability Engineer
DevOps / SRE
• Senior
• Remote
• Vollzeit
•
EU/EMEA
As a Site Reliability Engineer at Runware you ensure the platform stays reliable, performant and resilient as it scales, in a hands-on technical role across software, infrastructure and production operations.
Responsibilities
- ▹Own and improve the reliability, availability and performance of critical production services
- ▹Define and evolve reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
- ▹Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in the on-call rotation
- ▹Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting improvements
- ▹Reduce operational toil through automation, automated remediation and improvements to deployment safety and resilience
- ▹Work with Engineering and DevOps on capacity planning, performance, scaling and architectural improvements
Requirements
- ▹Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
- ▹Strong understanding of distributed systems, comfortable debugging across applications, databases, queues, containers, networking and infrastructure
- ▹Experience designing and operating observability systems using metrics, logs and distributed tracing
- ▹Understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing toil
- ▹Experience with Kubernetes, containers, IaC and automated deployment, plus the ability to write software and automation in Python, Go or PHP
- ▹Strong ownership of production problems and willingness to join the on-call rotation
Nice to have
- ▹Experience operating high-throughput or low-latency APIs and distributed systems
- ▹Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
- ▹Experience with RabbitMQ or other messaging and queueing systems
- ▹Experience operating MySQL, Redis, ClickHouse or similar production data systems
- ▹Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure
- ▹Experience building automated scaling, capacity management or self-healing systems
Soft skills
Strong ownershipCollaboration across teams
What we offer
- ▹Remote-first collective, meeting in person twice a year
- ▹Core hours for collaboration, otherwise a free schedule
- ▹Real downtime after fast, intense release cycles
About the company
Runware builds high-performance infrastructure and products for fast, scalable inference across image, video and emerging modalities. Its Serverless platform lets customers deploy and scale their own AI models on production-grade GPU infrastructure.
Ähnliche Stellen

Stelle
· Senior
Senior AI-assisted DevOps Engineer
Improvado
Clickhouse
+13
💰 Gehalt: keine Angabe
🌍 Remote
🗣️ EN
Himalayas

Stelle
· Senior
Senior SRE (Self-hosted)
GitGuardian
ArgocdClickhouse
+23
💰 Gehalt: keine Angabe
🏢 Vor Ort
Paris
🗣️ EN

Stelle
· Senior
Senior Site Reliability Engineer (remote within EMEA)
FYUL
Argocd
+22
💰 Gehalt: keine Angabe
🌍 Remote
🗣️ EN
Himalayas

Stelle
· Senior
Senior DevOps Engineer
Trading 212
ClickhouseCloudformation
+18
💰 Gehalt: keine Angabe
🔀 Hybrid
Sofia
🗣️ EN

Stelle
· Senior
Senior Site Reliability Engineer
Okta
AI/ML
+7
💰 Gehalt: keine Angabe
🏢 Vor Ort
Bengaluru
🗣️ EN

Stelle
· Senior
Senior DevOps Engineer (Cloud-Native, AI-Driven Platform)
septeo
AI/MLBitbucket
+16
💰 Gehalt: keine Angabe
🌍 Remote
España la Vieja
🗣️ EN