← Back to list
Job

Site Reliability Engineer - IDP

DevOps / SRE • Remote • Full-time • Portugal Portugal

Intermedia is hiring a Site Reliability Engineer responsible for the reliability, availability and performance of its most important applications and services. The role is primarily remote, with occasional visits to the office in Coimbra.

Responsibilities

  • ▹Ensure the availability, performance and reliability of critical applications and services through robust monitoring, alerting and optimization strategies
  • ▹Define, measure and maintain SLIs, SLOs and error budgets
  • ▹Partner with development teams to improve performance, reduce latency and increase application resilience
  • ▹Work with platform and DevOps teams to align infrastructure and application reliability
  • ▹Define reliability standards and operational guardrails for platform capabilities and golden paths
  • ▹Partner with platform engineering teams to design resilient self-service capabilities
  • ▹Automate operational tasks such as deployments, rollbacks, scaling, failover and recovery
  • ▹Continuously improve CI/CD pipelines to reduce manual intervention and support safe, progressive delivery
  • ▹Integrate automated validation, reliability checks and operational guardrails into development and deployment workflows
  • ▹Implement and maintain observability across production systems, including metrics, logs, traces and dashboards
  • ▹Develop dashboards, alerts and operational views for real-time visibility into system health
  • ▹Act as a key responder during incidents, troubleshooting, mitigating and resolving production issues
  • ▹Conduct root cause analysis and drive long-term corrective actions to prevent recurrence
  • ▹Run fire drills, game days and chaos engineering exercises to validate resilience under failure
  • ▹Monitor resource usage, capacity trends and scaling behavior to support future growth
  • ▹Partner with security teams on secure communication, access controls and data protection
  • ▹Lead or contribute to production readiness and operational review meetings
  • ▹Promote reliability engineering best practices and strengthen the organization's operational maturity

Requirements

  • ▹Bachelor's degree in Computer Science, Engineering or a related field, or equivalent practical experience
  • ▹Proven experience in Site Reliability Engineering, Platform Engineering or Infrastructure/DevOps roles with strong operational ownership
  • ▹Strong expertise in application monitoring, observability platforms, incident response and production troubleshooting
  • ▹Strong understanding of reliability concepts such as SLIs, SLOs, error budgets, alerting quality and incident management
  • ▹Scripting and automation skills with Python, Bash, Terraform, Ansible or similar
  • ▹Experience with cloud platforms such as AWS, Azure or Google Cloud
  • ▹Strong knowledge of CI/CD pipelines, deployment automation and progressive delivery
  • ▹Strong knowledge of infrastructure as code and configuration management
  • ▹Experience with containerization and orchestration, such as Docker and Kubernetes
  • ▹Strong problem-solving skills, operational judgment and attention to detail
  • ▹Excellent communication and collaboration skills across engineering, platform and security teams

Nice to have

  • ▹Experience with chaos engineering practices and tools
  • ▹Experience supporting internal platforms or platform engineering teams
  • ▹Familiarity with developer portals, golden paths, service catalogs or self-service platform patterns
  • ▹Understanding of developer experience metrics and operational maturity for internal platforms
  • ▹Familiarity with microservices architectures and multi-tenant environments
  • ▹Experience with modern observability stacks and telemetry standards
  • ▹Understanding of UCaaS and CCaaS platforms, especially voice and communication service flows
  • ▹Experience leading reliability initiatives, incident reviews or production improvement programs
  • ▹Familiarity with capacity planning, resilience testing and operational readiness

Soft skills

Strong problem-solvingExcellent communication and collaborationAnalytical thinkingHands-on, proactive attitude

What we offer

  • ▹Primarily remote work with occasional visits to the Coimbra office
  • ▹Offices planned in Aveiro and Porto in the future

About the company

Intermedia is a provider of cloud communications and collaboration technology with a record of growth and profitability and a culture of promoting from within. Many employees have been with the company for 10 to 20+ years.

Education: Informatikai, mérnöki vagy kapcsolódó területen szerzett alapképzés, vagy ezzel egyenértékű gyakorlati tapasztalat

Similar jobs