← Back to list
Job · Principal

Staff Software Engineer, Production Engineering

Software Engineer • Principal • Hybrid • Full-time United States San Francisco, USA

You'll build and operate Harvey's production infrastructure — compute, networking, and the Kubernetes platform — that powers its legal and enterprise AI products.

Responsibilities

  • Design, build, and operate the production infrastructure that powers Harvey's products and AI workloads
  • Drive technical direction across compute infrastructure, networking, Kubernetes, and workflow orchestration
  • Lead complex, cross-functional technical initiatives that improve reliability, scalability, and security
  • Build and operate Harvey's global compute and network infrastructure with high availability
  • Develop capacity models, demand forecasts, and fleet lifecycle automation
  • Operate and continuously improve the Kubernetes platform (cluster provisioning, upgrades, monitoring)
  • Develop scalable Infrastructure-as-Code and automation frameworks using Terraform and Pulumi
  • Improve observability, monitoring, alerting, and incident response
  • Participate in the on-call rotation and lead incident response when needed

Requirements

  • 10+ years of software, infrastructure, SRE, or production engineering experience
  • Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or GCP
  • Strong hands-on experience operating Kubernetes in production
  • Experience building and operating distributed systems
  • Experience with infrastructure automation and Infrastructure-as-Code (Terraform, Pulumi)
  • Strong understanding of compute infrastructure, networking, capacity planning, and fleet management
  • Experience designing observability systems (monitoring, logging, alerting, incident response)
  • Strong understanding of infrastructure security (IAM, network security, secrets management)
  • Track record of driving complex, cross-functional technical initiatives without formal authority
  • Excellent communication skills

Nice to have

  • Experience supporting AI/ML or LLM infrastructure at scale
  • Experience operating GPU fleets or high-performance compute infrastructure
  • Experience with multi-cloud or hybrid cloud environments
  • Experience building internal platforms or developer tooling that improves engineering velocity

Soft skills

Systems-thinking mindset with a passion for simple, reliable solutionsExcellent communication with technical and non-technical partnersIndependent, persuasive technical leadership without formal authority

About the company

Harvey is building the AI platform trusted by the world's leading law firms and enterprises, reshaping how legal and professional services work with frontier agentic AI.

Similar jobs