← Back to list

Job
· Principal
Member of Technical Staff – Research Engineer
AI Research Scientist
• Principal
• On-site
• Full-time
•
San Francisco (USA), USA
In this large-scale training-systems research engineering role at Black Forest Labs, you'll sit at the intersection of research and engineering, tackling GPU performance, low-precision training, and distributed-systems debugging.
Responsibilities
- ▹Improve the performance, reliability, and numerical stability of production training runs for large multimodal generative models
- ▹Profile full training steps across model code, attention, kernels, data loading, encoders, communication, optimizer steps, checkpointing, and memory pressure
- ▹Implement and validate GPU-level optimizations: fused kernels, attention paths, low-precision matmuls, quantization kernels, CUDA/Triton/CuTe/CUTLASS experiments
- ▹Push lower-precision training forward, including FP8/MXFP8/FP4-style paths, weight and activation quantization, and convergence/quality tradeoffs
- ▹Work with researchers to translate architecture changes into efficient training implementations
- ▹Debug distributed training failures: NaNs, loss spikes, silent numerical drift, memory leaks, stragglers, bad nodes, NCCL issues
- ▹Build benchmarking and profiling harnesses that make performance claims trustworthy across hardware and configurations
- ▹Help the training team move quickly on urgent bottlenecks, turning repeated failures into better abstractions and tools
Requirements
- ▹Experience working deeply on large-scale training systems, ideally as part of a training group working closely with researchers
- ▹Strong PyTorch fluency, including comfort reading and modifying low-level training code rather than only using high-level APIs
- ▹Experience with distributed training concepts such as FSDP, tensor/model/context/sequence parallelism, activation checkpointing, and NCCL
- ▹Hands-on experience improving training throughput, memory footprint, or stability in real training runs
- ▹Experience profiling GPU workloads with tools like Nsight Systems, Nsight Compute, torch profiler, or custom telemetry
- ▹Practical GPU performance judgment: the ability to verify correctness, numerical behavior, and performance
- ▹Understanding of low-precision training and quantization tradeoffs: FP8, MXFP8, FP4/NVFP4-style formats, scaling, accumulation, numerical validation
- ▹Good research judgment: partnering with researchers on ablations and correctly interpreting what measurements do and do not prove
- ▹Comfortable operating in ambiguity, whether the task is a clean implementation, a production fire, or a diagnostic puzzle
Nice to have
- ▹Supported or co-owned training for a frontier foundation model that shipped or reached a major release
- ▹Written or substantially improved forward/backward GPU kernels, with strong measurement and validation discipline
- ▹Worked on attention performance, variable sequence length training, or non-standard attention patterns
- ▹Experience on Hopper or Blackwell-class GPUs
- ▹Experience with low-precision training
- ▹Experience with diffusion, flow matching, DiT, and multimodal generative model training; LLM/autoregressive background is fine if eager to learn the diffusion stack quickly
- ▹Can move naturally between profiler traces, kernel code, distributed systems failures, and research discussions
Soft skills
research judgmentcomfort with ambiguitycollaboration with researcherscommunicating resultsrigorous validation
What we offer
- ▹Competitive base salary: EU €130,000–€240,000 + equity; US $180,000–$290,000 + equity
- ▹Flexible working model: Freiburg or SF at least 2 days a week (or one full week every other week), or remote with a monthly in-person week, with travel costs covered
About the company
Black Forest Labs is the team behind Latent Diffusion, Stable Diffusion, and FLUX — foundational generative AI technologies used by millions of creators, developers, and businesses to create images and video. Headquartered in Freiburg, Germany, with a growing presence in San Francisco.
Similar jobs

Job
· Principal
Staff Gen AI Research Scientist
Anduril
AI/MLHuggingfaceLangchainLlm
+1
$220,000–$292,000/yr
gross
🏢 On-site
Costa Mesa
🗣️ EN

Job
· Principal
Member of Technical Staff (AI Researcher)
Perplexity
AI/MLCppLlm
$220,000–$485,000/yr
gross
🏢 On-site
San Francisco
🗣️ EN

Job
· Principal
Senior / Staff Machine Learning Research Scientist, Agents
Scale AI
AI/MLLlm
+1
$25,200–$31,500/mo
gross
🏢 On-site
San Francisco
🗣️ EN

Job
· Principal
Principal Research Scientist/TLM - AI Scaling
Databricks
AI/MLDatabricksLlmMlflow
+3
$270,000–$340,000/yr
gross
🏢 On-site
Mountain View
🗣️ EN

Job
· Principal
Member of Technical Staff, AI Research
Psi
AI/ML
💰 Salary: not specified
🔀 Hybrid
Boston
🗣️ EN

Job
· Principal
Staff AI Research Engineer
Agility Robotics
AI/MLCpp
$216,000–$338,000/yr
gross
🏢 On-site
Any Office (Fremont
🗣️ EN