← Back to list
Job · Senior

Senior Site Reliability Engineer (remote within EMEA)

DevOps / SRE • Senior • Remote • Full-time Romania Romania

As a senior individual contributor on FYUL's Platform Infrastructure team, own the reliability of the container platform and cloud foundation while mentoring Associate and mid-level SREs.

Responsibilities

  • Architect and manage highly available, secure, scalable infrastructure across multiple AWS accounts using infrastructure as code
  • Design and operate Amazon EKS clusters: networking policies, persistent storage, scaling strategies
  • Own and evolve core platform services: cloud networking, Kubernetes, databases, and messaging systems
  • Drive large-scale automation projects and set standards for Terraform/Terragrunt and GitOps (ArgoCD)
  • Drive observability and incident response initiatives (Grafana, Prometheus, Loki, Tempo, Mimir), participate in on-call, write runbooks, ADRs, and postmortems
  • Lead security efforts (IAM, encryption, secure logging) and drive cost optimization across the platform
  • Mentor mid-level SREs, provide feedback, support onboarding of new team members

Requirements

  • Solid Linux systems administration background and comfort scripting in Python
  • Strong AWS knowledge: EKS, IAM (roles, policies, IRSA), VPC networking, RDS, S3, SQS
  • Hands-on experience operating and troubleshooting Kubernetes (EKS) at production scale

Nice to have

  • Experience in multi-account AWS environments

Soft skills

Mentoring and giving feedback to mid-level engineersCommunicating complex technical concepts clearly to engineers and non-technical stakeholdersPartnering with product engineering squads

About the company

FYUL's Platform Infrastructure team builds and operates the company's container platform and cloud foundation, fostering a DevOps culture through self-service tooling for product engineering teams.

Similar jobs