← Back to list
Job · Principal

AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff

AI / ML Engineer • Principal • On-site • Full-time United States San Diego, USA

Qualcomm's Cloud AI team develops hardware and software solutions for inference acceleration. The AI Performance Engineer analyzes and optimizes the performance of state-of-the-art generative AI models (LLM, VLM, diffusion), spanning research through commercial deployment.

Responsibilities

  • Convert, optimize and deploy models for efficient inference using PyTorch and ONNX.
  • Understand advanced algorithms (attention mechanisms, MoEs) and numerics to identify optimization opportunities.
  • Analyze and optimize performance of LLM, VLM and diffusion models for inference, scaling for throughput and latency.
  • Map next-generation AI workloads onto current and future hardware designs.
  • Work closely with customers and internal compiler, firmware and platform teams.
  • Analyze complex performance or stability issues to root cause.
  • Build engineering solutions for continuous insight into AI workload performance.
  • Design and implement high-level kernels (e.g. in Triton) for efficient, low-level code.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Information Systems or related field and 6+ years of experience (or Master's + 5 years, or PhD + 4 years).
  • MS in Computer Science, Machine Learning, Computer Engineering or Electrical Engineering.

Nice to have

  • Hands-on experience building and optimizing language models, notably in PyTorch and ONNX, preferably in production.
  • Deep understanding of transformer architectures, attention mechanisms and performance tradeoffs.
  • Experience with workload mapping strategies such as sharding or parallelism.
  • Strong Python programming skills.
  • Understanding of computer architecture, ML accelerators, in-memory processing and distributed systems.
  • Background in neural network operators and math fundamentals (linear algebra, math libraries).
  • Understanding of machine learning compilers.
  • Experience with accuracy convergence and its evaluation methods.
  • Knowledge of torch.compile or torchDynamo.
  • PhD in Computer Science, Computer Engineering or Machine Learning.

Soft skills

Strong communication and problem-solving skills.Ability to learn and work effectively in a fast-paced, collaborative environment.Proactive learning about the latest inference optimization techniques.

What we offer

  • Annual discretionary bonus program.
  • Annual RSU grants.
  • Competitive benefits package.

About the company

Qualcomm is applying its traditional strengths in digital wireless technology to play a central role in the evolution of Cloud AI, developing hardware and software solutions for inference acceleration.

Education: Számítástechnika, informatika, mérnöki vagy villamosmérnöki alapszak (MSc előny).

Similar jobs