← Back to list
Job · Senior

Senior Site Reliability Engineer

DevOps / SRE • Senior • Remote • Full-time European Union EU/EMEA

Alpaca, a US-headquartered global agent-first brokerage infrastructure company, is hiring a Senior Site Reliability Engineer to help keep its brokerage platform reliable, observable, and operable across cloud infrastructure, Kubernetes, and the data layer.

Responsibilities

  • Operate production day-to-day: oncall, incident response, postmortems, and follow-ups
  • Own reliability practice: define and refine SLIs/SLOs and error budgets
  • Strengthen observability across metrics, logs, traces, and alerting
  • Ship infrastructure through code in a GitOps workflow, covering both cloud resources and Kubernetes workloads
  • Look after PostgreSQL: performance tuning, schema and migration review, online migrations on large tables, HA/DR, and CDC pipelines
  • Mentor engineers on reliability and database fundamentals through code review, design review, and pairing

Requirements

  • 4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership
  • Hands-on experience operating production services on Kubernetes, shipping infrastructure as code in a GitOps workflow
  • Solid working knowledge of PostgreSQL in production: query plans, pg_stat_*, indexing and schema trade-offs, safe online migrations on non-trivial tables
  • Cloud networking fundamentals (VPCs, routing, L4/L7 load balancing, DNS, TLS) and cross-service connectivity debugging

Soft skills

Ownership of on-call responsibilitiesMentoringCross-functional collaboration

About the company

Alpaca is a US-headquartered, global agent-first brokerage infrastructure company for stocks, ETFs, options, crypto, fixed income, and 24/5 trading, serving over 10 million brokerage accounts across 40 countries.

Similar jobs