As a Research Engineer at Lightning AI, you'll optimize training and inference workloads running on Lightning AI infrastructure, sitting at the intersection of ML systems, AI infrastructure, and performance engineering.
Responsibilities
- ▹Optimize large-scale training and inference workloads across GPUs, accelerators, and distributed systems
- ▹Work directly with customers to analyze workloads and improve performance, scalability, and reliability
- ▹Develop inference pipelines, model serving systems, and performance-oriented tooling for production AI
- ▹Design and implement profiling, debugging, and observability tools to guide optimization
- ▹Work across the software stack so performance improvements are accessible via clean APIs and automation
- ▹Partner with hardware vendors and ecosystem partners for efficient execution (NVIDIA, TPU, emerging accelerators)
- ▹Contribute to open-source projects: features, tooling, documentation, community engagement
Requirements
- ▹Strong expertise with deep learning frameworks, especially PyTorch
- ▹Experience with large-scale training or inference workloads
- ▹Familiarity with distributed systems and parallelism strategies (data/model/pipeline parallelism, checkpointing, elastic scaling, distributed inference)
- ▹Strong software engineering fundamentals: API design, tooling, debugging complex systems, production-quality code
- ▹Experience analyzing and improving performance bottlenecks in ML systems or distributed workloads
- ▹Excellent collaboration and communication skills, including with customers and external contributors
- ▹Ability to work in ambiguous, fast-moving environments across multiple layers of the stack
- ▹Bachelor's degree in Computer Science, Engineering, or a related field
Nice to have
- ▹Experience with inference optimization: quantization, speculative decoding, mixed precision, memory-efficient training, throughput/latency optimization
- ▹Experience with CUDA, Triton, TensorRT, vLLM, SGLang, Dynamo, or related ML systems/inference tooling
- ▹Experience contributing to open-source ML, infrastructure, or AI systems projects
- ▹Startup experience or highly cross-functional environments
- ▹Advanced degree (Master's or PhD) in AI, machine learning, systems, or related fields
Soft skills
Excellent collaboration and communication, including with customers and external contributorsComfortable in ambiguous, fast-moving environmentsPractical, systems-minded problem solvingOpenness to open-source community work
What we offer
- ▹Anticipated annual base salary $120,000-$250,000, with discretionary bonus and meaningful equity
- ▹Comprehensive medical, dental, and vision coverage (U.S.); private medical and dental insurance (U.K.)
- ▹Retirement and financial wellness support (U.S.); pension contribution (U.K.)
- ▹Generous paid time off, plus holidays
- ▹Paid parental leave
- ▹Professional development support
- ▹Wellness and work-from-home stipends
- ▹Flexible work environment
About the company
Lightning AI is the company behind PyTorch Lightning, building an end-to-end platform for developing, training, and deploying AI systems since 2019. Following its merger with Voltage Park, it combines developer-first software with cost-efficient, large-scale compute. With hubs in New York City, San Francisco, Seattle, and London, it is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.
Languages: Angol: Felsőfok
Education: Alapképzés (BSc) informatika, mérnöki vagy kapcsolódó területen
Similar jobs

Job
AI Research Scientist- Multimodal Foundational Models
Bosch
AI/MLData ScienceMlflow
+1
$165,000–$195,000/yr
gross
🏢 On-site
Sunnyvale
🗣️ EN

Job
Applied Research Scientist, AI Research
Descript
AI/ML
💰 Salary: not specified
🔀 Hybrid
San Francisco
🗣️ EN

Job
Research Scientist- Robotics AI
Bosch
AI/MLCppData Science
+2
💰 Salary: not specified
🏢 On-site
Sunnyvale
🗣️ EN

Job
AI Research Resident
Normal Computing
AI/MLLlm
$150,000/yr
gross
🏢 On-site
New York City
🗣️ EN

Job
Research Engineer, Mid-Training
Cognition
AI/MLLlm
💰 Salary: not specified
🏢 On-site
San Francisco
🗣️ EN

Job
AI Research Engineer
Normal Computing
AI/MLLlm
$200,000–$400,000/yr
gross
🔀 Hybrid
New York City
🗣️ EN
