← Back to list

Job
· Principal
Member of Technical Staff - Inference
Other
• Principal
• Remote
• Full-time
•
EU/EMEA
Prime Intellect is hiring an inference engineer spanning cloud LLM serving, LLM inference optimization, and RL systems, building scalable model-serving infrastructure and integrating it into the training stack.
Stack
Responsibilities
- ▹Build a multi-tenant LLM serving platform that operates across cloud GPU fleets
- ▹Design placement and scheduling algorithms for heterogeneous accelerators
- ▹Implement multi-region/zone failover and traffic shifting for resilience and cost control
- ▹Build autoscaling, routing, and load balancing to meet throughput/latency SLOs
- ▹Optimize model distribution and cold-start times across clusters
- ▹Integrate and contribute to LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM
- ▹Optimize configurations for tensor/pipeline/expert parallelism, prefix caching, and memory management
- ▹Profile kernels, memory bandwidth, and transport; apply quantization and speculative decoding
- ▹Develop reproducible performance suites for latency, throughput, context length, batch size, and precision
- ▹Embed and optimize distributed inference within the RL stack
- ▹Establish CI/CD with artifact promotion, performance gates, and reproducible builds
- ▹Build metrics, logs, and tracing; structured incident response and SLO management
- ▹Document architectures, playbooks, and API contracts; mentor and collaborate cross-functionally
Requirements
- ▹3+ years building and running large-scale ML/LLM services with clear latency/availability SLOs
- ▹Hands-on experience with at least one inference backend (vLLM, SGLang, TensorRT-LLM)
- ▹Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo
- ▹Deep understanding of prefill vs. decode, KV-cache behavior, batching, sampling, speculative decoding, and parallelism strategies
- ▹Comfortable debugging CUDA/NCCL, drivers/kernels, containers, service mesh/networking, and storage end to end
- ▹Python: systems tooling and backend services
- ▹PyTorch: LLM inference engine development, integration, and deployment readiness
- ▹Cloud & automation: AWS/GCP service experience, cloud deployment patterns
- ▹Kubernetes: running infrastructure at scale with containers
- ▹GPU & networking: CUDA runtime, NCCL, InfiniBand, GPU-aware bin-packing and scheduling across heterogeneous fleets
Nice to have
- ▹CUDA/Triton kernel development, Nsight Systems/Compute profiling
- ▹Systems performance languages: Rust, C++
- ▹Data & observability: Kafka/PubSub, Redis, gRPC/Protobuf, Prometheus/Grafana, OpenTelemetry
- ▹Infrastructure automation: Terraform/Ansible, infrastructure-as-code
- ▹Open source contributions to serving, inference, or RL infrastructure projects
Soft skills
Openness to open development and community contributionCross-functional collaboration with researchers and engineers
What we offer
- ▹Cash compensation range of $150-300k with significant equity incentives
- ▹Flexible work arrangement (remote or San Francisco office)
- ▹Full visa sponsorship and relocation support
- ▹Professional development budget
- ▹Regular team off-sites and conference attendance
About the company
Prime Intellect is building the open superintelligence stack: the infrastructure frontier AI labs build internally, made available to every ambitious AI team. Its platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment into one full-stack system. The company is backed by leading investors including Founders Fund, Radical Ventures, and NVIDIA.
Similar jobs

Job
· Principal
Principal Engineer (L5)
Twilio
AI/MLDatadog
+4
💰 Salary: not specified
🌍 Remote
Anywhere in the World
🗣️ EN

Job
· Principal
(Canada) Principal ML System Engineer
PointClickCare
AI/MLDatabricks
+7
$176,000–$195,000/yr
gross
🌍 Remote
or Mississauga
🗣️ EN

Job
· Principal
Staff Engineer - Heartbeat AI (m/f/d)
1Komma5Grad
AI/MLBigqueryDatadog
+12
💰 Salary: not specified
🌍 Remote
🗣️ EN

Job
· Principal
Staff Software Development Engineer - SDM
Delinea
Grpc
+3
$150,000–$185,000/yr
gross
🌍 Remote
U.S. Remote
🗣️ EN

Job
· Principal
Staff/Lead Python Engineer (FastAPI, Orchestration)
Mimica
Llm
+1
💰 Salary: not specified
🌍 Remote
🗣️ EN

Job
· Principal
Senior/Staff/Principal Engineer
Canonical
CppData Science
+9
💰 Salary: not specified
🌍 Remote
🗣️ EN