← Zurück zur Liste
Stelle

Applied AI Inference Engineer

AI / ML Engineer • Vor Ort • Vollzeit Vereinigte Staaten San Francisco, USA

You'll own the LLM inference stack end to end: profiling where time and cost go, bringing modern optimization techniques into real deployments, and getting deep into the serving code when defaults aren't good enough - while working directly with customer engineering teams to take workloads from proof of concept to fully monitored production.

Responsibilities

  • Bring current inference techniques into production and refine them
  • Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches
  • Work down into the serving stack, from frameworks like vLLM and SGLang to the CUDA kernels underneath, profiling and running in-depth analysis to find and fix performance problems
  • Adapt and scale optimization methods across many kinds of ML models, with an emphasis on large language models
  • Profile and tune deployments against clear targets for latency, throughput, and cost, keeping them dependable under real traffic
  • Tailor deployments to each customer's models and constraints, partnering with their engineering teams from proof of concept through to a live, well-monitored production service
  • Build and support the software and product features around the inference stack in production, using one or more general-purpose languages, primarily Python
  • Experiment quickly: shape fuzzy goals into clear specs and focused proofs of concept, and ship well-tested results without delay
  • Own delivery end to end, from the first experiment through to production optimization, drafting features and product requirement documents with other engineering and product teams

Requirements

  • A Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field
  • Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python
  • Familiarity with methods for optimizing LLMs for high throughput / low latency inference
  • Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level
  • A firm grasp of how GPUs are built and how they behave
  • Clear interest and hands-on experience with large language models
  • A working knowledge of AI/ML pipelines and the full path of developing and deploying ML models
  • Strong communication skills, particularly when explaining hard technical topics to customers and teammates

Nice to have

  • A track record of making software systems run faster, especially for large language models
  • Experience with CUDA or comparable technologies
  • A strong command of software engineering fundamentals, with a record of building and shipping AI/ML inference systems
  • Experience with Docker and Kubernetes
  • Prior work building or tuning AI/ML projects, particularly in a customer-facing setting

Soft skills

Customer orientationPragmatic problem-solvingOwnership and accountabilityClear communication to technical and non-technical audiences

What we offer

  • Competitive compensation and equity packages (RSUs)
  • Paid time off, paid holidays & leave of absence programs
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance & emergency assistance
  • Daily meals allowance

About the company

Crusoe is the only vertically integrated AI infrastructure company built from the ground up, owning and operating every layer of the stack - from electrons to tokens - to power the world's most ambitious AI workloads.

Languages: Angol: Felsőfok

Ähnliche Stellen