← Zurück zur Liste

Stelle
· Principal
Member of Technical Staff – Research Engineer
AI Research Scientist
• Principal
• Vor Ort
• Vollzeit
•
San Francisco (USA), USA
In this large-scale training-systems research engineering role at Black Forest Labs, you'll sit at the intersection of research and engineering, tackling GPU performance, low-precision training, and distributed-systems debugging.
Responsibilities
- ▹Improve the performance, reliability, and numerical stability of production training runs for large multimodal generative models
- ▹Profile full training steps across model code, attention, kernels, data loading, encoders, communication, optimizer steps, checkpointing, and memory pressure
- ▹Implement and validate GPU-level optimizations: fused kernels, attention paths, low-precision matmuls, quantization kernels, CUDA/Triton/CuTe/CUTLASS experiments
- ▹Push lower-precision training forward, including FP8/MXFP8/FP4-style paths, weight and activation quantization, and convergence/quality tradeoffs
- ▹Work with researchers to translate architecture changes into efficient training implementations
- ▹Debug distributed training failures: NaNs, loss spikes, silent numerical drift, memory leaks, stragglers, bad nodes, NCCL issues
- ▹Build benchmarking and profiling harnesses that make performance claims trustworthy across hardware and configurations
- ▹Help the training team move quickly on urgent bottlenecks, turning repeated failures into better abstractions and tools
Requirements
- ▹Experience working deeply on large-scale training systems, ideally as part of a training group working closely with researchers
- ▹Strong PyTorch fluency, including comfort reading and modifying low-level training code rather than only using high-level APIs
- ▹Experience with distributed training concepts such as FSDP, tensor/model/context/sequence parallelism, activation checkpointing, and NCCL
- ▹Hands-on experience improving training throughput, memory footprint, or stability in real training runs
- ▹Experience profiling GPU workloads with tools like Nsight Systems, Nsight Compute, torch profiler, or custom telemetry
- ▹Practical GPU performance judgment: the ability to verify correctness, numerical behavior, and performance
- ▹Understanding of low-precision training and quantization tradeoffs: FP8, MXFP8, FP4/NVFP4-style formats, scaling, accumulation, numerical validation
- ▹Good research judgment: partnering with researchers on ablations and correctly interpreting what measurements do and do not prove
- ▹Comfortable operating in ambiguity, whether the task is a clean implementation, a production fire, or a diagnostic puzzle
Nice to have
- ▹Supported or co-owned training for a frontier foundation model that shipped or reached a major release
- ▹Written or substantially improved forward/backward GPU kernels, with strong measurement and validation discipline
- ▹Worked on attention performance, variable sequence length training, or non-standard attention patterns
- ▹Experience on Hopper or Blackwell-class GPUs
- ▹Experience with low-precision training
- ▹Experience with diffusion, flow matching, DiT, and multimodal generative model training; LLM/autoregressive background is fine if eager to learn the diffusion stack quickly
- ▹Can move naturally between profiler traces, kernel code, distributed systems failures, and research discussions
Soft skills
research judgmentcomfort with ambiguitycollaboration with researcherscommunicating resultsrigorous validation
What we offer
- ▹Competitive base salary: EU €130,000–€240,000 + equity; US $180,000–$290,000 + equity
- ▹Flexible working model: Freiburg or SF at least 2 days a week (or one full week every other week), or remote with a monthly in-person week, with travel costs covered
About the company
Black Forest Labs is the team behind Latent Diffusion, Stable Diffusion, and FLUX — foundational generative AI technologies used by millions of creators, developers, and businesses to create images and video. Headquartered in Freiburg, Germany, with a growing presence in San Francisco.
Ähnliche Stellen

Stelle
· Principal
Staff Gen AI Research Scientist
Anduril
AI/MLHuggingfaceLangchainLlm
+1
188 896–250 717 €/Jahr
brutto
🏢 Vor Ort
Costa Mesa
🗣️ EN

Stelle
· Principal
Member of Technical Staff (AI Researcher)
Perplexity
AI/MLCppLlm
188 896–416 431 €/Jahr
brutto
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
· Principal
Senior / Staff Machine Learning Research Scientist, Agents
Scale AI
AI/MLLlm
+1
21 637–27 047 €/Mon.
brutto
🏢 Vor Ort
San Francisco
🗣️ EN

Stelle
· Principal
Principal Research Scientist/TLM - AI Scaling
Databricks
AI/MLDatabricksLlmMlflow
+3
231 827–291 931 €/Jahr
brutto
🏢 Vor Ort
Mountain View
🗣️ EN

Stelle
· Principal
Member of Technical Staff, AI Research
Psi
AI/ML
💰 Gehalt: keine Angabe
🔀 Hybrid
Boston
🗣️ EN

Stelle
· Principal
Staff AI Research Engineer
Agility Robotics
AI/MLCpp
185 462–290 213 €/Jahr
brutto
🏢 Vor Ort
Any Office (Fremont
🗣️ EN