← Back to list
Job

AI Platforms Leader (Enterprise AI Platforms)

Other • On-site • Full-time United States San Diego, USA

Qualcomm is seeking an experienced AI Platforms Leader to own the strategy, architecture, and operation of its end-to-end AI platform, spanning on-prem GPU clusters and cloud services (AWS/GCP/Azure). The role leads a ~10-engineer global team delivering reliable, secure, and cost-efficient infrastructure for training, fine-tuning, inference, retrieval, and agentic orchestration. Full-time onsite role in San Diego, CA (5 days per week).

Responsibilities

  • Define the multi-year strategy and roadmap for a multi-tenant, hybrid (on-prem + cloud) AI platform
  • Establish platform SLAs/SLOs, reliability goals, and security/compliance guardrails
  • Operate and optimize on-prem GPU clusters (Kubernetes + GPU operator and/or Slurm), capacity planning and scheduling
  • Deliver MLOps/LLMOps as a product: data prep, training/fine-tuning, model registry, governance, evaluation, safe deployment
  • Implement CI/CD for models, prompts, and agents with automated evaluation and canary/A-B rollouts
  • Lead agentic AI orchestration (A2A patterns) and MCP server integration
  • Leverage cloud AI/ML platforms (AWS/Azure/GCP) for training and inference
  • Own core platform services: identity/RBAC, secrets, observability, vector stores, model gateways
  • Lead a ~10-engineer global platform/SRE/MLOps team with 24x7 on-call readiness and incident response
  • Manage strategic vendor relationships and licensing

Requirements

  • 15+ years overall engineering/technology experience, including ~10 years building and operating large-scale platforms (AI/ML, data, or HPC)
  • 5+ years leading a ~10-engineer team across platform/SRE/MLOps/LLMOps
  • Hands-on experience operating on-prem GPU clusters (Kubernetes + GPU operator and/or Slurm)
  • Strong MLOps/LLMOps experience across the full model lifecycle (data → training → registry → deployment)
  • Deep AWS/GCP/Azure experience with managed Kubernetes (EKS/AKS/GKE)
  • DevOps/platform engineering: CI/CD, GitOps, IaC (Terraform/Bicep/Helm), Docker, Kubernetes
  • Understanding of agent orchestration, A2A patterns, and operating MCP servers in production
  • Bachelor's degree in Engineering, Computer Science, or related field

Nice to have

  • Master's or PhD in CS/EE/Math or related field
  • Experience with PyTorch, CUDA/cuDNN, Triton Inference Server, vLLM, KServe, Ray, Slurm
  • High-throughput storage (Lustre, BeeGFS, Ceph) and vector databases (FAISS, Milvus, Pinecone)
  • MLOps toolchain: MLflow, Airflow/Argo, Weights & Biases, LangSmith
  • Security/governance: OIDC/RBAC, OPA policy as code, AWS Secrets Manager/Azure Key Vault
  • Agentic frameworks: Semantic Kernel, LangChain, CrewAI, AutoGen

Soft skills

Strong communication with executives and technical leaders, with clear metrics and business value storytellingGlobal, cross-time-zone team leadershipCoaching, hiring, and performance management

What we offer

  • Competitive annual discretionary bonus program
  • Opportunity for annual RSU grants
  • Highly competitive benefits package

About the company

Qualcomm Incorporated is a global technology company. This role sits within the Engineering Group's Software Engineering organization, building and operating the company's enterprise-wide AI platform infrastructure.

Education: Alapdiploma mérnöki, informatikai vagy rokon területen (mester/PhD előny)

Similar jobs