← Back to list
Job · Senior

Senior Site Reliability Engineer (Guardicore AI Platform) - Remote

DevOps / SRE • Senior • Remote • Full-time Spain Spain

A Senior SRE role on Akamai's Guardicore Data and AI Platform Team, owning the reliability, availability, performance, and operational readiness of the security-focused AI platform.

Responsibilities

  • Operate secure, highly available Kubernetes infrastructure for core microservices, data pipelines, observability, and internal tooling
  • Enhance platform reliability, observability, security, performance, and cost efficiency
  • Provide guidance to engineers and developers on service performance
  • Lead complex production investigations and drive long-term improvements
  • Leverage LLMs and AI-driven automation to auto-remediate incidents and streamline operations
  • Partner across DevOps, Software, Data, AI, and Security engineering teams
  • Participate in on-call rotations, guiding restoration of service-impacting issues

Requirements

  • 5+ years of experience in SRE, DevOps, or Platform Engineering
  • Ability to design and implement monitoring and observability using Prometheus and Grafana
  • Extensive production expertise with Kubernetes, Docker, Helm, and clouds (GCP, Azure, Linode, AWS) on Linux
  • Exceptional troubleshooting across network, system, application, and database layers
  • Experience with GitOps, CI/CD, and Infrastructure as Code
  • Scripting and programming proficiency in Python, Go, and Bash
  • Daily use of AI tools in operational tasks

Soft skills

Technical leadershipCross-team collaborationOwnership

About the company

Akamai provides the world's most distributed platform from cloud to edge, helping the giants of the digital world work faster and stay more secure across cloud/edge, security, content delivery, and AI.

Similar jobs