← Zurück zur Liste

Stelle
· Principal
LLM Serving Engineer (Cloud AI Engineering), Senior / Staff
AI / ML Engineer
• Principal
• Vor Ort
• Vollzeit
•
San Diego, USA
Qualcomm's Cloud AI team develops hardware and software solutions for inference acceleration. The LLM Serving Engineer builds a scalable LLM inference platform, spanning the full product lifecycle from research and development to commercial deployment.
Responsibilities
- ▹Build a scalable LLM inference platform (disaggregated serving, KV-cache management, advanced parallelism, speculative algorithms, model optimization, specialized kernels).
- ▹Contribute to LLM serving packages (vLLM, SGLang, TGI, Triton Inference Server, Dynamo, LLM-d).
- ▹Work closely with customers and internal compiler, firmware and platform teams.
- ▹Understand advanced algorithms (attention mechanisms, MoEs) and numerics to identify optimization opportunities.
- ▹Drive efficient serving through smart autoscaling, load balancing and routing.
- ▹Engage with open-source serving communities to evolve the framework.
Requirements
- ▹Bachelor's degree in Computer Science, Engineering, Information Systems or related field and 4+ years of experience (or Master's + 3 years, or PhD + 2 years).
- ▹MS in Computer Science, Machine Learning, Computer Engineering or Electrical Engineering.
Nice to have
- ▹Hands-on experience with LLM serving/orchestration packages (Triton Inference Server, vLLM, SGLang, Ollama, llm-d, KServe, LMCache, MoonCake).
- ▹Deep understanding of foundational LLMs, VLMs, SLMs, and transformer-based architectures.
- ▹Strong experience developing language models with PyTorch.
- ▹Strong computer science fundamentals: algorithms, data structures, parallel and distributed programming.
- ▹Understanding of computer architecture, ML accelerators, in-memory processing and distributed systems.
- ▹Strong Python development skills for large-scale projects.
- ▹Experience analyzing, profiling and optimizing deep learning workloads.
- ▹Open-source contributions to any GenAI package.
- ▹Experience architecting and developing large-scale distributed systems.
- ▹High-level kernel design experience (PyTorch, CUDA, Triton).
- ▹Knowledge of torch.compile or torchDynamo.
- ▹PhD in Computer Science, Computer Engineering or Machine Learning.
Soft skills
Excellent communication and problem-solving skills.Thriving in a fast-paced, collaborative environment.Proactive learning about the latest inference optimization techniques.
What we offer
- ▹Annual discretionary bonus program.
- ▹Annual RSU grants.
- ▹Competitive benefits package.
About the company
Qualcomm is applying its traditional strengths in digital wireless technology to play a central role in the evolution of Cloud AI, developing hardware and software solutions for inference acceleration.
Education: Számítástechnika, informatika, mérnöki vagy villamosmérnöki alapszak (MSc előny).
Ähnliche Stellen

Stelle
· Principal
Principal Applied Scientist- AI
UiPath
AI/MLHuggingfaceLlm
+2
188 476–229 111 €/Jahr
brutto
🏢 Vor Ort
Bellevue
🗣️ EN

Stelle
· Principal
Principal Engineer, AI/ML Architecture - 11310
Coupa
AI/MLLlm
208 361–291 705 €/Jahr
brutto
🏢 Vor Ort
Foster City
🗣️ EN

Stelle
· Principal
Principal AI/ML Engineer
Pragmatike
AI/MLData Science
+5
💰 Gehalt: keine Angabe
🔀 Hybrid
San Francisco
🗣️ EN

Stelle
· Principal
Staff Machine Learning Engineer, Applied Research
Pinterest
AI/MLCppLlmMlflow
+3
163 670–336 968 €/Jahr
brutto
🔀 Hybrid
San Francisco
🗣️ EN

Stelle
· Principal
Staff Machine Learning Engineer, Personalization
Spotify
AI/MLBigqueryLlm
+1
💰 Gehalt: keine Angabe
🏢 Vor Ort
New York
🗣️ EN

Stelle
· Principal
AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff
Qualcomm
AI/MLLlm
154 239–231 358 €/Jahr
brutto
🏢 Vor Ort
San Diego
🗣️ EN