← Zurück zur Liste
Stelle · Principal

Member of Technical Staff – Research Engineer

AI Research Scientist • Principal • Vor Ort • Vollzeit Vereinigte Staaten San Francisco (USA), USA

In this large-scale training-systems research engineering role at Black Forest Labs, you'll sit at the intersection of research and engineering, tackling GPU performance, low-precision training, and distributed-systems debugging.

Responsibilities

  • Improve the performance, reliability, and numerical stability of production training runs for large multimodal generative models
  • Profile full training steps across model code, attention, kernels, data loading, encoders, communication, optimizer steps, checkpointing, and memory pressure
  • Implement and validate GPU-level optimizations: fused kernels, attention paths, low-precision matmuls, quantization kernels, CUDA/Triton/CuTe/CUTLASS experiments
  • Push lower-precision training forward, including FP8/MXFP8/FP4-style paths, weight and activation quantization, and convergence/quality tradeoffs
  • Work with researchers to translate architecture changes into efficient training implementations
  • Debug distributed training failures: NaNs, loss spikes, silent numerical drift, memory leaks, stragglers, bad nodes, NCCL issues
  • Build benchmarking and profiling harnesses that make performance claims trustworthy across hardware and configurations
  • Help the training team move quickly on urgent bottlenecks, turning repeated failures into better abstractions and tools

Requirements

  • Experience working deeply on large-scale training systems, ideally as part of a training group working closely with researchers
  • Strong PyTorch fluency, including comfort reading and modifying low-level training code rather than only using high-level APIs
  • Experience with distributed training concepts such as FSDP, tensor/model/context/sequence parallelism, activation checkpointing, and NCCL
  • Hands-on experience improving training throughput, memory footprint, or stability in real training runs
  • Experience profiling GPU workloads with tools like Nsight Systems, Nsight Compute, torch profiler, or custom telemetry
  • Practical GPU performance judgment: the ability to verify correctness, numerical behavior, and performance
  • Understanding of low-precision training and quantization tradeoffs: FP8, MXFP8, FP4/NVFP4-style formats, scaling, accumulation, numerical validation
  • Good research judgment: partnering with researchers on ablations and correctly interpreting what measurements do and do not prove
  • Comfortable operating in ambiguity, whether the task is a clean implementation, a production fire, or a diagnostic puzzle

Nice to have

  • Supported or co-owned training for a frontier foundation model that shipped or reached a major release
  • Written or substantially improved forward/backward GPU kernels, with strong measurement and validation discipline
  • Worked on attention performance, variable sequence length training, or non-standard attention patterns
  • Experience on Hopper or Blackwell-class GPUs
  • Experience with low-precision training
  • Experience with diffusion, flow matching, DiT, and multimodal generative model training; LLM/autoregressive background is fine if eager to learn the diffusion stack quickly
  • Can move naturally between profiler traces, kernel code, distributed systems failures, and research discussions

Soft skills

research judgmentcomfort with ambiguitycollaboration with researcherscommunicating resultsrigorous validation

What we offer

  • Competitive base salary: EU €130,000–€240,000 + equity; US $180,000–$290,000 + equity
  • Flexible working model: Freiburg or SF at least 2 days a week (or one full week every other week), or remote with a monthly in-person week, with travel costs covered

About the company

Black Forest Labs is the team behind Latent Diffusion, Stable Diffusion, and FLUX — foundational generative AI technologies used by millions of creators, developers, and businesses to create images and video. Headquartered in Freiburg, Germany, with a growing presence in San Francisco.

Ähnliche Stellen