← Zurück zur Liste

Stelle
· Lead
Lead Hardware Deployment Engineer
Sonstige
• Lead
• Hybrid
• Vollzeit
•
Memphis, USA
This role owns end-to-end bring-up of GPU compute hardware for the world's largest AI training clusters, based in Memphis, TN. You'll build and lead a dedicated hardware deployment team responsible for L11 rack integration, bring-up, and post-L11 repair across multiple data halls concurrently. Your team's throughput is one of the most critical-path activities in the company, directly determining how fast compute comes online.
Responsibilities
- ▹Lead, hire, and develop a dedicated hardware deployment team (deployment engineers, technicians, repair technicians) with full ownership of team structure and staffing
- ▹Own L11 rack integration and compute hardware bring-up across multiple data halls concurrently, from delivery dock to healthy production handoff
- ▹Drive aggressive bring-up timelines: 95%+ node availability within days of rack delivery, 100% closure within one week per data hall
- ▹Own post-L11 hardware health, running systematic health pushes to sustain greater than 98% node availability before turnover to operations
- ▹Internalize non-RMA hardware repairs to maximize recovery, minimize backlogs, and reduce dependence on OEM turnaround times
- ▹Develop and enforce vendor SLAs for OEM and supplier responsibilities, preventing accumulation of unrepaired hardware
- ▹Perform root cause analysis of hardware failures found during L11 and drive corrective actions with vendors and internal engineering
- ▹Partner with site operations on hardware debugging and repair, and train site operations teams for future deployments
- ▹Build, document, and continuously improve deployment processes, tooling, and training so bring-up capability scales across sites
Requirements
- ▹5+ years of hands-on experience deploying, integrating, or repairing compute/server hardware at data center scale
- ▹Direct experience with L11 (rack-level) integration and bring-up of GPU or accelerator-based systems
- ▹Demonstrated experience leading technician or engineering teams in a fast-paced deployment, manufacturing, or data center environment
- ▹Deep troubleshooting skills across servers, GPUs, NVLink/fabric interconnects, high-speed networking, and liquid cooling systems
- ▹Willingness to work on-site in Memphis, TN, including extended hours and weekends during critical bring-up phases
Nice to have
- ▹Experience with NVIDIA GB200/GB300 NVL72 or similar rack-scale liquid-cooled GPU systems
- ▹Experience standing up a new team or function from scratch, including hiring, training, and process development
- ▹Experience managing OEM/ODM vendor relationships (e.g., Dell, Supermicro), including SLA definition and enforcement
- ▹Experience with hardware failure analysis, RMA processes, and component-level repair strategies at fleet scale
- ▹Experience with data center automation, burn-in/validation tooling, and hardware health telemetry
- ▹Track record of driving step-change improvements in deployment velocity or cost
Soft skills
Strong communication skillsStrong work ethic and prioritizationHands-on initiative and ownershipAbility to lead teams under high pressure
About the company
The company's mission is to build AI systems that accurately understand the universe and help humanity in its pursuit of knowledge. The team is small, highly motivated, and operates with a flat structure focused on engineering excellence, expecting every employee to be hands-on and take initiative.
Ähnliche Stellen

Stelle
· Lead
Lead Mechatronics Engineer
Anduril
124 721–165 726 €/Jahr
brutto
🏢 Vor Ort
Costa Mesa
🗣️ EN
Stelle
· Lead
Lead Engineer, Computational Knowledge Graph
Avathon
AI/MLData Science
+9
128 138–187 936 €/Jahr
brutto
🏢 Vor Ort
Pleasanton
🗣️ EN

Stelle
· Lead
Analytics Lead, Safety Reporting
Lyft
Data ScienceSQL
💰 Gehalt: keine Angabe
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
· Lead
Data Acquisition Lead, Frontier Environments
Scale AI
155 816–194 770 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
· Lead
Data Center Operations Lead - Partner Site Operations
Anthropic
273 362–345 974 €/Jahr
brutto
🔀 Hybrid
Austin
🗣️ EN

Stelle
· Lead
E-7A Systems Engineer (Experienced or Lead) | Requirements and Verification
Boeing
Data Science
102 383–138 518 €/Jahr
brutto
🏢 Vor Ort
Tukwila
🗣️ EN