← Zurück zur Liste
Stelle · Senior

Senior SRE / Cloud Engineer

Cloud Engineer • Senior • Remote • Vollzeit Ukraine Ukraine

An experienced SRE / Cloud Engineer is sought with strong hands-on experience supporting production cloud infrastructure, focused on reliability, incident response, automation and observability.

Responsibilities

  • Operate, maintain and improve production cloud infrastructure across AWS, Azure, GCP, Windows or hybrid environments
  • Build and maintain monitoring, logging, metrics, tracing, dashboards and alerting
  • Improve observability coverage across infrastructure, applications, databases and network dependencies
  • Tune alerts to reduce noise and improve signal quality
  • Participate in incident response, troubleshooting, root cause analysis and post-incident remediation
  • Build automation for infrastructure operations, deployments, health checks and runbooks
  • Partner with engineering teams to define SLOs, SLIs, error budgets and operational readiness standards
  • Support cloud networking, DNS, TLS, load balancing, IAM, storage and compute operations
  • Maintain Infrastructure as Code and configuration management practices
  • Create and maintain runbooks, operational documentation and escalation procedures
  • Identify and remediate production risks through automation and architecture improvements

Requirements

  • Hands-on production cloud infrastructure experience
  • Strong experience with AWS, GCP, Windows and/or hybrid cloud environments
  • Experience building and maintaining observability, monitoring, logging, dashboards and alerting
  • Strong troubleshooting skills across infrastructure, networking, application and cloud service layers
  • Experience with Linux systems, networking fundamentals, DNS, TLS, IAM, load balancers, storage and compute
  • Infrastructure as Code experience using Terraform, CloudFormation or Pulumi
  • Scripting and automation skills in Bash, Python or Go
  • Experience participating in production incident response and postmortem processes
  • Familiarity with SRE practices (SLOs, SLIs, error budgets, toil reduction, operational readiness)
  • Ability to work closely with engineering teams to improve reliability

Nice to have

  • Kubernetes and Cloud-Native Platforms
  • GitOps experience (Flux or Argo CD)
  • Jenkins and Ansible
  • Service Mesh, Ingress and API Gateway experience
  • High Availability architectures
  • Multi-region environments
  • Disaster Recovery solutions
  • Secrets Management (Vault, AWS Secrets Manager, External Secrets)
  • Security, Compliance and Vulnerability Management
  • On-call operations and runbook design

Soft skills

Troubleshooting live systemsClose collaboration with engineering teamsProactive risk identification and remediation

What we offer

  • Culture of relentless performance: 99% project success rate and 30%+ year-over-year revenue growth
  • Competitive pay

Ähnliche Stellen