← Zurück zur Liste
Stelle · Senior

Senior Engineering Manager, Production Engineering

Engineering Manager • Senior • Hybrid • Vollzeit Vereinigte Staaten San Francisco, USA

A leadership role over Harvey's Infrastructure Production Engineering organization, owning the compute/networking infrastructure, Kubernetes platform, and workflow orchestration that let engineering teams move quickly with confidence.

Responsibilities

  • Lead, mentor, and grow a team of infrastructure engineers
  • Foster a culture of operational excellence, quality, and continuous improvement
  • Partner with Engineering, Security, Product, and AI Infrastructure leaders to define long-term infrastructure strategy
  • Drive technical direction for compute, networking, Kubernetes, workflow orchestration, and production operations
  • Lead cross-functional initiatives to improve reliability, scalability, security, and cost optimization
  • Own and operate global compute and network infrastructure
  • Lead capacity planning, demand forecasting, and fleet lifecycle management
  • Operate and improve the Kubernetes platform (cluster provisioning, upgrades, monitoring, automation)
  • Own the Temporal-based workflow orchestration platform
  • Drive infrastructure cost optimization
  • Build and maintain secure infrastructure foundations (IAM, network isolation, secrets management, auditing, compliance)
  • Develop scalable Infrastructure-as-Code and automation frameworks (e.g. Terraform, Pulumi)
  • Establish comprehensive observability, monitoring, alerting, and incident response practices

Requirements

  • 7+ years of software or infrastructure engineering experience, including 5+ years leading teams
  • Deep expertise operating large-scale cloud infrastructure on AWS, Azure, or GCP
  • Strong hands-on experience operating Kubernetes in production (cluster lifecycle, networking, reliability)
  • Experience building and operating large-scale distributed systems
  • Experience with infrastructure automation and Infrastructure-as-Code (e.g. Terraform, Pulumi)
  • Strong understanding of compute infrastructure, networking, capacity planning, and fleet management
  • Experience designing and operating observability platforms (monitoring, logging, alerting, incident response)
  • Strong understanding of infrastructure security (IAM, network security, secrets management, compliance)
  • Demonstrated success leading complex cross-functional technical initiatives
  • Excellent communication skills for both engineering and executive audiences
  • A systems-thinking mindset and passion for simple, reliable, scalable infrastructure

Nice to have

  • Experience operating workflow orchestration platforms such as Temporal
  • Experience supporting AI/ML or LLM infrastructure at scale
  • Experience managing GPU fleets, high-performance compute infrastructure, or large-scale capacity planning
  • Experience with multi-cloud or hybrid cloud environments

Soft skills

team leadership and mentoringcross-functional communicationsystems thinkingculture building, operational excellence

About the company

Harvey is the AI platform trusted by the world's leading law firms and enterprises. Its infrastructure is the foundation that powers every customer interaction, model inference, and production workload.

Ähnliche Stellen