← Back to list

Job
· Principal
Staff Infrastructure Engineer
Platform Engineer
• Principal
• Remote
• Full-time
•
EU/EMEA
You architect and own the cloud platform every Headway engineer deploys on: deploys should be boring, scaling automatic, infrastructure self-serve and cost attributable. A Staff-level technical leadership role in the Infrastructure Engineering team.
Responsibilities
- ▹Deployment architecture and blast-radius containment: redesign deployment so that a mistake in one part of the service cannot block or take down others, continuing the shift toward per-service deploy isolation and functional-area slices
- ▹Container footprint and networking: own the ECS and EKS footprint, evaluate broader EKS adoption for AI workloads and design the next iteration of inter-service network connectivity
- ▹Capacity and scaling: own capacity for a spiky workload, with floors computed ahead of demand rather than chased by reactive scaling, self-deriving from data, with drift caught early
- ▹Self-serve infrastructure platform: build the Terraform self-serve platform with guardrails so engineering teams own their standard infrastructure changes and Reliability Engineering reviews only the non-standard ones
- ▹Cloud cost attribution and controls: stand up per-team cost attribution across AWS, Datadog and LLM spend, making infrastructure costs visible and attributable
- ▹Python runtime and dependency health: own how the monolith behaves under load (garbage collection, event loop contention, runtime limits) and lead the framework and package upgrades most teams defer
Requirements
- ▹8 or more years in platform, infrastructure or SRE roles at companies running significant production traffic
- ▹Deep AWS expertise and production ownership of compute and networking at scale (ECS, EKS, RDS, networking, IAM)
- ▹Strong infrastructure-as-code experience, particularly Terraform, including designing self-serve platforms for other engineering teams
- ▹Hands-on autoscaling and capacity engineering, and container orchestration with ECS and/or EKS
- ▹Track record making deploys safe and self-serve for other teams, not just your own
- ▹Staff-level influence: you drive decisions across team boundaries and raise the infrastructure bar org-wide without requiring management authority
Nice to have
- ▹FinOps and cloud cost optimization experience
- ▹Kubernetes and EKS depth
- ▹Observability tooling at scale (Datadog)
- ▹Experience in healthcare or other regulated environments
- ▹Experience with event driven systems
Soft skills
Technical leadership and cross-team alignmentHigh degree of ownershipThriving in ambiguityMaking other engineers better through architecture reviews and runbooks
About the company
Headway is building a more accessible mental healthcare system in the US: it automates insurance administration, and over 75,000 providers run their practice on its software, serving over 1 million patients. A Series D company with $325M+ in funding.
Similar jobs
Job
· Principal
Staff Platform Engineer
iFIT
DatadogGithub ActionsLlm
+3
$155,000–$190,000/yr
gross
🌍 Remote
🗣️ EN

Job
· Principal
Senior Staff Machine Learning Platform Engineer
Faire
AI/MLDatabricksDatadog
+16
💰 Salary: not specified
🌍 Remote
Anywhere in the World
🗣️ EN
Job
· Principal
Principal / Staff / Senior Infrastructure Engineer
Allspice
Datadog
+15
💰 Salary: not specified
🔀 Hybrid
Boston
🗣️ EN

Job
· Principal
Staff Platform Engineer
Horizon3ai
Datadog
+4
$199,750–$270,000/yr
gross
🌍 Remote
US
🗣️ EN

Job
· Principal
Staff AI Platform Engineer: Agent & Retrieval Infrastructure
Bedrockocean
CloudformationLlm
+1
💰 Salary: not specified
🌍 Remote
🗣️ EN

Job
· Principal
Staff Platform Engineer
Prefect
DagsterDatadog
+5
$214,000–$309,000/yr
gross
🌍 Remote
🗣️ EN