← Zurück zur Liste
Stelle · Principal

Staff Software Engineer (L4) – SRE, Platform Engineering

Softwareentwickler • Principal • Remote • Vollzeit • Irland Irland

As Staff Site Reliability Engineer in Twilio's Platform Engineering organization you are accountable for production health and resiliency. A hands-on engineering role where you write code that improves reliability and have impact across multiple teams. The role is remote, based in Ireland.

Responsibilities

  • ▹Own the reliability posture of production services in your area: availability, latency, capacity, efficiency, performance, and the monitoring and alerting that makes them visible.
  • ▹Define, instrument and operate against SLIs and SLOs, and use error budgets to drive engineering priorities.
  • ▹Identify trends and problem areas that threaten stability, and provide a path forward to mitigate risk before it reaches customers.
  • ▹Drive down repair items and prevent classes of incidents rather than resolving them one at a time.
  • ▹Improve detection, response and recovery, reducing time to acknowledge, engage, mitigate and restore, with fewer people pulled in.
  • ▹Design for failure: strengthen failure domains, validate recovery paths, and make production changes safer to ship and safer to roll back.
  • ▹Participate in on-call for the services you support, and lead the response when production is degraded.
  • ▹Write post-mortems that identify true root causes, and drive the follow-up work to completion.
  • ▹Oversee efforts to identify, diagnose, report and document production problems across all reliability dimensions.
  • ▹Write, configure and deploy code that measurably improves service reliability: maintainable, reviewed, documented and well tested.
  • ▹Orchestrate complex changes across systems and services, documenting design changes, technical decisions, migration plans and upgrades.
  • ▹Lead debugging, troubleshooting and analysis of service architecture and design.
  • ▹Use code review to drive up the quality of your coworkers' code.
  • ▹Reduce the operational overhead required to run infrastructure and services.
  • ▹Drive projects from conception to completion for efforts spanning the concerns of your team; break projects into milestones and tasks, track progress and communicate updates to stakeholders.
  • ▹Identify and communicate changes that may impact stability.

Requirements

  • ▹8+ years of related engineering experience, with a substantial portion focused on reliability, infrastructure or platform engineering.
  • ▹Demonstrated accountability for production systems: you have carried a pager for services that mattered and owned the outcome when they failed.
  • ▹Strong software engineering fundamentals: you build and ship production code, not only configure tooling.
  • ▹Experience defining and operating against SLIs and SLOs, and using error budgets to inform engineering priorities.
  • ▹Depth in production operations: incident command, post-mortem analysis, capacity planning and observability.
  • ▹A record of preventing recurrence: reducing incident classes and operational toil, not just closing tickets.
  • ▹Experience driving changes that span multiple teams, and the communication skills to build alignment without formal authority.
  • ▹A track record of improving the engineers around you through code review, design feedback and mentorship.
  • ▹Experience with large-scale distributed systems in a cloud environment.

Nice to have

  • ▹Familiarity with infrastructure-as-code, container orchestration and GitOps-style delivery.
  • ▹Experience with multi-region architecture, failure-domain design or regional expansion work.
  • ▹Background in chaos engineering, game days or other proactive resilience validation.

Soft skills

Communication and alignment without formal authorityMentoringWorking without day-to-day guidanceCollaboration across teamsAccountability

What we offer

  • ▹Competitive pay
  • ▹Generous time off
  • ▹Ample parental and wellness leave
  • ▹Healthcare
  • ▹Retirement savings program
  • ▹Support for volunteering and donation efforts
  • ▹Remote-first work (offerings vary by location)

About the company

Twilio delivers communications solutions to hundreds of thousands of businesses and empowers millions of developers worldwide. It is a remote-first company; in-person attendance may occasionally be needed for team gatherings, off-sites or customer meetings.

Ähnliche Stellen