← Back to list
Job

Senior Engineering Manager, Compute

Backend Developer • On-site • Full-time • 📍 United States

As the leader of the Compute team, this role is responsible for making Temporal's compute layer invisible to customers while scaling, operating, and keeping it cost-efficient for the world's most demanding AI workloads.

Responsibilities

  • Set strategic direction for Compute: own the strategy and standards of excellence across design, delivery, and operations
  • Provide technical leadership: build and grow a high-ownership team, stay close to design docs and code rather than managing from a distance
  • Drive roadmap & trajectory: chart the path from today's compute toward next-generation compute platforms, grounded in customer feedback
  • Own operational excellence: run on-call and incident response, drive blameless postmortems and systemic fixes
  • Guide technical depth: make hard architectural decisions for large-scale, multi-tenant compute (workload isolation, security, scheduling, fleet efficiency, performance)
  • Own capacity, supply & economics: utilization, capacity and supply planning, and cost-per-unit-of-compute across CPU and accelerated compute
  • Drive cross-team & customer execution: partner with leadership, Product, SDK, UX/DX, Security, and design-partner customers to align priorities and unblock delivery

Requirements

  • Proven experience leading software engineering teams that build and operate large-scale compute platforms or fleets, with strong operational practices
  • 12+ years in software and/or infrastructure engineering, including 7+ years of people management and demonstrated ownership of delivery and live-site outcomes
  • Deep distributed-systems and compute infrastructure depth, with hands-on judgment to guide architecture and execution
  • Experience operating multi-tenant compute that other people's production workloads depend on
  • Bachelor's degree in Computer Science or related field, or equivalent practical experience; advanced degree a plus
  • Excellent communication skills, with the ability to partner across engineering, product, and leadership and fold customer feedback into the roadmap
  • Strong leadership, coaching, and performance management; ability to grow engineers and build a healthy, accountable, high-ownership team
  • Excellence in execution: planning, prioritization, and delivering iterative milestones in an ambiguous, fast-moving environment while managing unplanned work
  • Fleet thinking: utilization, goodput, capacity and supply planning, and cost discipline as first-class engineering concerns
  • Live-site reliability craft: on-call, incident management & response, and postmortem-driven continuous improvement
  • Strong command of the building blocks of a compute platform: multi-tenant isolation and security, scheduling, and resource management
  • Ability to review and raise the bar on technical artifacts (design docs, code reviews) across a distributed-systems codebase

Nice to have

  • MicroVMs and virtualization (Firecracker, gVisor, Edera) or managed-compute primitives (AWS Fargate, GCP Cloud Run, AWS Lambda), and/or Kubernetes internals
  • Building serverless or hosted-compute products from 0 to 1, including the rapid-delivery-vs-durable-platform tradeoffs
  • Multi-cloud delivery across AWS and GCP
  • Cold-start, warm-pool, and scheduling/latency optimization for on-demand compute
  • Agent sandboxes, secure execution of untrusted code, or other AI-agent infrastructure
  • GPU / accelerated compute: fractional GPUs (MIG, MPS, time-slicing), GPU scheduling, training vs. inference fleets, and multi-tenant GPU isolation

Soft skills

Strategic thinker with a hands-on approachComfortable shaping a space that doesn't fully exist yetObsessed with reliability when customers bet production on itDefaults to working backwards from customersBalances speed to ship against durable, planet-scale design

What we offer

  • Eligible to participate in Temporal's equity plan

About the company

Companies at the frontier of the AI revolution run on Temporal. OpenAI runs on Temporal, handling millions of requests. Cursor runs its cloud coding agents on Temporal at over 50 million actions a day across 7M+ workflows, and more than a third of the pull requests its users merge now come from those agents. Replit, Lovable, Abridge, and Hebbia build their agents on it too. In the last year alone, AI-native companies executed 1.86 trillion actions on Temporal Cloud, and the curve is still bending upwards. Backed by a recent $300M Series D at a $5B valuation, the company is building the durable execution layer the agentic era depends on.

Languages: Angol: Felsőfok
Education: Számítástechnika vagy kapcsolódó terület, BSc (vagy azzal egyenértékű gyakorlati tapasztalat); mesterfokozat előny

Similar jobs