← Zurück zur Liste
Stelle · Senior

Senior Site Reliability Engineer (Hiring Globally)

DevOps / SRE • Senior • Remote • Vollzeit Serbien Serbien

Operating a distributed, camera-based video monitoring and AI alerting platform. This is a reliability-first role: owning SLOs, error budgets, trustworthy alerting, capacity planning, and disaster recovery across a mixed cloud-edge infrastructure.

Responsibilities

  • Define SLIs/SLOs and manage error budgets
  • Build trustworthy observability (dashboards, custom metrics, OpenTelemetry)
  • Lead incident response, write blameless postmortems
  • Own and improve the on-call rotation
  • Plan capacity and performance (GPU, Kafka/MSK, RDS, Redis)
  • Own business continuity and disaster recovery (backups, replication, failover)

Requirements

  • Experience operating large-scale, mixed (cloud + edge) infrastructure
  • SLO/error-budget-driven mindset
  • Familiarity with Java microservices, Python services, AWS
  • Experience with Kafka (MSK), PostgreSQL/RDS, Redis, DynamoDB
  • Docker, Terraform, observability tooling (Datadog, Grafana, Prometheus)

Soft skills

Calm, composed incident responseTreating reliability as a measured number, not a feelingProactive in taming noisy systems

About the company

The company runs a distributed, camera-based video monitoring and AI alerting platform connecting cloud infrastructure with thousands of on-premise edge devices.

Ähnliche Stellen