← Zurück zur Liste
Stelle · Senior

Senior Site Reliability Engineer

DevOps / SRE • Senior • Remote • Vollzeit • Europäische Union EU/EMEA

As a Site Reliability Engineer at Runware you ensure the platform stays reliable, performant and resilient as it scales, in a hands-on technical role across software, infrastructure and production operations.

Responsibilities

  • ▹Own and improve the reliability, availability and performance of critical production services
  • ▹Define and evolve reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
  • ▹Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in the on-call rotation
  • ▹Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting improvements
  • ▹Reduce operational toil through automation, automated remediation and improvements to deployment safety and resilience
  • ▹Work with Engineering and DevOps on capacity planning, performance, scaling and architectural improvements

Requirements

  • ▹Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
  • ▹Strong understanding of distributed systems, comfortable debugging across applications, databases, queues, containers, networking and infrastructure
  • ▹Experience designing and operating observability systems using metrics, logs and distributed tracing
  • ▹Understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing toil
  • ▹Experience with Kubernetes, containers, IaC and automated deployment, plus the ability to write software and automation in Python, Go or PHP
  • ▹Strong ownership of production problems and willingness to join the on-call rotation

Nice to have

  • ▹Experience operating high-throughput or low-latency APIs and distributed systems
  • ▹Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
  • ▹Experience with RabbitMQ or other messaging and queueing systems
  • ▹Experience operating MySQL, Redis, ClickHouse or similar production data systems
  • ▹Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure
  • ▹Experience building automated scaling, capacity management or self-healing systems

Soft skills

Strong ownershipCollaboration across teams

What we offer

  • ▹Remote-first collective, meeting in person twice a year
  • ▹Core hours for collaboration, otherwise a free schedule
  • ▹Real downtime after fast, intense release cycles

About the company

Runware builds high-performance infrastructure and products for fast, scalable inference across image, video and emerging modalities. Its Serverless platform lets customers deploy and scale their own AI models on production-grade GPU infrastructure.

Ähnliche Stellen