← Back to list
Job · Mid-level

Machine Learning Engineer — Training Optimization

AI / ML Engineer • Mid-level • Remote • Full-time • European Union EU/EMEA

Featherless AI is looking for an ML Engineer focused on training optimization to scale and improve large-scale model training for speed, stability, and cost, working closely with researchers.

Responsibilities

  • ▹Optimize large-scale model training pipelines (throughput, convergence, stability, cost)
  • ▹Improve distributed training strategies (data, model, and pipeline parallelism)
  • ▹Tune optimizers, schedulers, batch sizing, and precision (bf16 / fp16 / fp8)
  • ▹Reduce training time and compute cost via profiling, bottleneck analysis, and systems-level improvements
  • ▹Collaborate with researchers on architecture-aware training strategies
  • ▹Build and maintain robust training infrastructure (checkpointing, fault tolerance, reproducibility)
  • ▹Evaluate and integrate new training techniques (gradient checkpointing, ZeRO, FSDP, custom kernels)
  • ▹Own training performance metrics

Requirements

  • ▹Strong experience training large neural networks (LLMs or similarly large models)
  • ▹Hands-on experience with training optimization, not just model usage
  • ▹Understanding of backpropagation, optimization algorithms, and training dynamics
  • ▹Understanding of distributed systems for ML training
  • ▹Experience with PyTorch (required)
  • ▹Comfort working close to hardware (GPUs, memory, networking constraints)

Nice to have

  • ▹Large-scale distributed training (multi-node, multi-GPU)
  • ▹Familiarity with DeepSpeed, FSDP, Megatron, or custom training stacks
  • ▹Experience optimizing training on AMD or NVIDIA GPUs
  • ▹Contributions to open-source ML infrastructure or research codebases
  • ▹Exposure to non-Transformer architectures (RNNs, hybrid models)

Soft skills

Moving fluidly between research ideas and production-ready code

What we offer

  • ▹Real ownership at Series-A stage
  • ▹Work on cutting-edge models and training systems at scale
  • ▹Small, highly technical team with fast feedback loops
  • ▹Competitive compensation and meaningful equity

Similar jobs