← Back to list
Job · Principal

Staff Machine Learning Systems Engineer

Other • Principal • Remote • Full-time United States USA

Reddit's ML Platform team owns the infrastructure powering recommendations and content discovery; as a Staff ML Infrastructure Engineer you'll lead a platform for large-scale ML models.

Responsibilities

  • Design end-to-end MLOps model lifecycle patterns to boost ML engineer velocity
  • Zero-to-one development and support of a graph ML codebase and platform
  • Collaborate on performance tuning, training time, efficiency and GPU costs
  • Optimize batch data processing (Apache Beam, Spark, Ray Data)
  • Architect pipelines for graph data structures with billions of nodes and tens of billions of edges

Requirements

  • 8+ years of ML infrastructure experience, including model training and deployment
  • Hands-on ML optimization, including memory and GPU profiling
  • Deep experience with cloud technologies (GCP BigQuery, GCS, Terraform)
  • Hands-on experience administering MLOps tools (MLflow, Wandb)
  • Proficiency in Python, PyTorch, TensorFlow
  • Deep experience with distributed training frameworks (Ray, Kubernetes)

Nice to have

  • Experience with graph databases (Neo4j, JanusGraph, TigerGraph)
  • Experience with graph neural networks and frameworks (PyTorch Geometric, Deep Graph Library)

Soft skills

strong organizational and communication skillsadvocate for platform usersintuition for the ML development lifecycle

What we offer

  • Medical, dental and vision insurance
  • 401k program with employer match
  • Generous time off and parental leave
  • Competitive base salary ($230,000–$322,000) plus equity

About the company

Reddit is a community of communities, built on shared interests, passion, and trust, home to some of the internet's most open and authentic conversations. With over 100,000 active communities and approximately 126 million daily active unique visitors, Reddit is one of the internet's largest sources of information.

Similar jobs