← Zurück zur Liste
Stelle · Senior

Senior Site Reliability Engineer (SRE) – Kubernetes

DevOps / SRE • Senior • Hybrid • Vollzeit • Polen Kraków, Polen

Software Mind is looking for a Senior SRE for the team building the platform behind ServiceNow's AI-first user interfaces. The role owns production reliability on Kubernetes and troubleshooting of both the Node.js and JVM sides of the system.

Responsibilities

  • ▹Support the deployment, operation and reliability of production services running on Kubernetes
  • ▹Monitor service health and investigate production incidents across distributed applications
  • ▹Participate in on-call support, incident response, root cause analysis, postmortems and reliability improvements
  • ▹Troubleshoot application runtime, networking and service-to-service issues in collaboration with engineering teams
  • ▹Support CI/CD, GitOps-based deployments, observability and production monitoring
  • ▹Work within a client-directed backlog and established priorities

Requirements

  • ▹5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services
  • ▹3+ years of hands-on production Kubernetes experience strongly preferred
  • ▹Kubernetes production operations: deployment, scaling, rollout/rollback, resource tuning and service-to-service troubleshooting
  • ▹Strong production incident response experience: on-call, runbooks, postmortems and paging hygiene
  • ▹Splunk experience for log aggregation, search and production troubleshooting
  • ▹Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing ones
  • ▹CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
  • ▹Strong Linux and networking fundamentals, including DNS, load balancing, TCP/HTTP, HTTP/2 and Kubernetes networking
  • ▹Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime; may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, JVM GC log analysis, thread dumps, JVM tuning and Java service latency investigation
  • ▹Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion and JWT-based service authentication

Nice to have

  • ▹Web Components / Lit experience, for first-level debugging of UI-related issues
  • ▹Server-side rendering or isomorphic runtime experience
  • ▹Canary rollout / multi-version production operations
  • ▹Distributed tracing and request-context correlation
  • ▹KEDA or event-driven autoscaling
  • ▹Experience with enterprise platform integration layers

Soft skills

Openness and ownershipCollaboration with engineering teams

What we offer

  • ▹Flexible employment and remote work
  • ▹International projects with leading global clients
  • ▹International business trips
  • ▹Non-corporate atmosphere
  • ▹Language classes
  • ▹Internal and external training
  • ▹Private healthcare and insurance
  • ▹Multisport card
  • ▹Well-being initiatives

About the company

Software Mind develops solutions for companies around the globe, working with tech giants and unicorns. The position is in the AI Experience Framework team, which builds the platform powering ServiceNow's AI-first user interfaces.

Languages: Angol: nagyon jó írásban és szóban

Ähnliche Stellen