← Back to list
Job · Principal

LLM Serving Engineer (Cloud AI Engineering), Senior / Staff

AI / ML Engineer • Principal • On-site • Full-time United States San Diego, USA

Qualcomm's Cloud AI team develops hardware and software solutions for inference acceleration. The LLM Serving Engineer builds a scalable LLM inference platform, spanning the full product lifecycle from research and development to commercial deployment.

Responsibilities

  • Build a scalable LLM inference platform (disaggregated serving, KV-cache management, advanced parallelism, speculative algorithms, model optimization, specialized kernels).
  • Contribute to LLM serving packages (vLLM, SGLang, TGI, Triton Inference Server, Dynamo, LLM-d).
  • Work closely with customers and internal compiler, firmware and platform teams.
  • Understand advanced algorithms (attention mechanisms, MoEs) and numerics to identify optimization opportunities.
  • Drive efficient serving through smart autoscaling, load balancing and routing.
  • Engage with open-source serving communities to evolve the framework.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Information Systems or related field and 4+ years of experience (or Master's + 3 years, or PhD + 2 years).
  • MS in Computer Science, Machine Learning, Computer Engineering or Electrical Engineering.

Nice to have

  • Hands-on experience with LLM serving/orchestration packages (Triton Inference Server, vLLM, SGLang, Ollama, llm-d, KServe, LMCache, MoonCake).
  • Deep understanding of foundational LLMs, VLMs, SLMs, and transformer-based architectures.
  • Strong experience developing language models with PyTorch.
  • Strong computer science fundamentals: algorithms, data structures, parallel and distributed programming.
  • Understanding of computer architecture, ML accelerators, in-memory processing and distributed systems.
  • Strong Python development skills for large-scale projects.
  • Experience analyzing, profiling and optimizing deep learning workloads.
  • Open-source contributions to any GenAI package.
  • Experience architecting and developing large-scale distributed systems.
  • High-level kernel design experience (PyTorch, CUDA, Triton).
  • Knowledge of torch.compile or torchDynamo.
  • PhD in Computer Science, Computer Engineering or Machine Learning.

Soft skills

Excellent communication and problem-solving skills.Thriving in a fast-paced, collaborative environment.Proactive learning about the latest inference optimization techniques.

What we offer

  • Annual discretionary bonus program.
  • Annual RSU grants.
  • Competitive benefits package.

About the company

Qualcomm is applying its traditional strengths in digital wireless technology to play a central role in the evolution of Cloud AI, developing hardware and software solutions for inference acceleration.

Education: Számítástechnika, informatika, mérnöki vagy villamosmérnöki alapszak (MSc előny).

Similar jobs