← Zurück zur Liste
Stelle

Research Engineer

AI Research Scientist • Vor Ort • Vollzeit Vereinigte Staaten New York, USA

As a Research Engineer at Lightning AI, you'll optimize training and inference workloads running on Lightning AI infrastructure, sitting at the intersection of ML systems, AI infrastructure, and performance engineering.

Stack

Responsibilities

  • Optimize large-scale training and inference workloads across GPUs, accelerators, and distributed systems
  • Work directly with customers to analyze workloads and improve performance, scalability, and reliability
  • Develop inference pipelines, model serving systems, and performance-oriented tooling for production AI
  • Design and implement profiling, debugging, and observability tools to guide optimization
  • Work across the software stack so performance improvements are accessible via clean APIs and automation
  • Partner with hardware vendors and ecosystem partners for efficient execution (NVIDIA, TPU, emerging accelerators)
  • Contribute to open-source projects: features, tooling, documentation, community engagement

Requirements

  • Strong expertise with deep learning frameworks, especially PyTorch
  • Experience with large-scale training or inference workloads
  • Familiarity with distributed systems and parallelism strategies (data/model/pipeline parallelism, checkpointing, elastic scaling, distributed inference)
  • Strong software engineering fundamentals: API design, tooling, debugging complex systems, production-quality code
  • Experience analyzing and improving performance bottlenecks in ML systems or distributed workloads
  • Excellent collaboration and communication skills, including with customers and external contributors
  • Ability to work in ambiguous, fast-moving environments across multiple layers of the stack
  • Bachelor's degree in Computer Science, Engineering, or a related field

Nice to have

  • Experience with inference optimization: quantization, speculative decoding, mixed precision, memory-efficient training, throughput/latency optimization
  • Experience with CUDA, Triton, TensorRT, vLLM, SGLang, Dynamo, or related ML systems/inference tooling
  • Experience contributing to open-source ML, infrastructure, or AI systems projects
  • Startup experience or highly cross-functional environments
  • Advanced degree (Master's or PhD) in AI, machine learning, systems, or related fields

Soft skills

Excellent collaboration and communication, including with customers and external contributorsComfortable in ambiguous, fast-moving environmentsPractical, systems-minded problem solvingOpenness to open-source community work

What we offer

  • Anticipated annual base salary $120,000-$250,000, with discretionary bonus and meaningful equity
  • Comprehensive medical, dental, and vision coverage (U.S.); private medical and dental insurance (U.K.)
  • Retirement and financial wellness support (U.S.); pension contribution (U.K.)
  • Generous paid time off, plus holidays
  • Paid parental leave
  • Professional development support
  • Wellness and work-from-home stipends
  • Flexible work environment

About the company

Lightning AI is the company behind PyTorch Lightning, building an end-to-end platform for developing, training, and deploying AI systems since 2019. Following its merger with Voltage Park, it combines developer-first software with cost-efficient, large-scale compute. With hubs in New York City, San Francisco, Seattle, and London, it is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.

Languages: Angol: Felsőfok
Education: Alapképzés (BSc) informatika, mérnöki vagy kapcsolódó területen

Ähnliche Stellen