← Zurück zur Liste

Stelle
Applied AI Inference Engineer
AI / ML Engineer
• Vor Ort
• Vollzeit
•
San Francisco, USA
You'll own the LLM inference stack end to end: profiling where time and cost go, bringing modern optimization techniques into real deployments, and getting deep into the serving code when defaults aren't good enough - while working directly with customer engineering teams to take workloads from proof of concept to fully monitored production.
Responsibilities
- ▹Bring current inference techniques into production and refine them
- ▹Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches
- ▹Work down into the serving stack, from frameworks like vLLM and SGLang to the CUDA kernels underneath, profiling and running in-depth analysis to find and fix performance problems
- ▹Adapt and scale optimization methods across many kinds of ML models, with an emphasis on large language models
- ▹Profile and tune deployments against clear targets for latency, throughput, and cost, keeping them dependable under real traffic
- ▹Tailor deployments to each customer's models and constraints, partnering with their engineering teams from proof of concept through to a live, well-monitored production service
- ▹Build and support the software and product features around the inference stack in production, using one or more general-purpose languages, primarily Python
- ▹Experiment quickly: shape fuzzy goals into clear specs and focused proofs of concept, and ship well-tested results without delay
- ▹Own delivery end to end, from the first experiment through to production optimization, drafting features and product requirement documents with other engineering and product teams
Requirements
- ▹A Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field
- ▹Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python
- ▹Familiarity with methods for optimizing LLMs for high throughput / low latency inference
- ▹Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level
- ▹A firm grasp of how GPUs are built and how they behave
- ▹Clear interest and hands-on experience with large language models
- ▹A working knowledge of AI/ML pipelines and the full path of developing and deploying ML models
- ▹Strong communication skills, particularly when explaining hard technical topics to customers and teammates
Nice to have
- ▹A track record of making software systems run faster, especially for large language models
- ▹Experience with CUDA or comparable technologies
- ▹A strong command of software engineering fundamentals, with a record of building and shipping AI/ML inference systems
- ▹Experience with Docker and Kubernetes
- ▹Prior work building or tuning AI/ML projects, particularly in a customer-facing setting
Soft skills
Customer orientationPragmatic problem-solvingOwnership and accountabilityClear communication to technical and non-technical audiences
What we offer
- ▹Competitive compensation and equity packages (RSUs)
- ▹Paid time off, paid holidays & leave of absence programs
- ▹Comprehensive health, dental & vision insurance
- ▹Employer contributions to HSA account
- ▹Paid parental leave
- ▹Paid life insurance, short-term and long-term disability
- ▹Professional development & tuition reimbursement
- ▹Mental health & wellness support
- ▹Commuter benefits (parking & transit)
- ▹Cell phone stipend
- ▹401(k) Retirement plan with company match up to 4% of salary
- ▹Volunteer time off
- ▹Global travel insurance & emergency assistance
- ▹Daily meals allowance
About the company
Crusoe is the only vertically integrated AI infrastructure company built from the ground up, owning and operating every layer of the stack - from electrons to tokens - to power the world's most ambitious AI workloads.
Languages: Angol: Felsőfok
Ähnliche Stellen

Stelle
AI / ML Engineer
Accenture Federal Services
AI/MLCpp
+7
💰 Gehalt: keine Angabe
🏢 Vor Ort
Tampa
🗣️ EN
Stelle
AI/ML Engineer
540
AI/MLData Science
+7
💰 Gehalt: keine Angabe
🏢 Vor Ort
Arlington
🗣️ EN

Stelle
Director, AI Engineering (Data Science)
Blend360
AI/MLFlinkBigquery
+16
154 998–206 664 €/Jahr
brutto
🔀 Hybrid
Columbia
🗣️ EN

Stelle
Applied AI Product Engineer (Remote Opportunity)
Vetsez
AI/MLData Science
+10
💰 Gehalt: keine Angabe
🌍 Remote
Philadelphia
🗣️ EN

Stelle
Applied AI Engineer
Twenty
AI/ML
+6
136 915–226 470 €/Jahr
brutto
🏢 Vor Ort
Arlington
🗣️ EN

Stelle
AI/ML Engineers and Analysts - Future Consideration
Analytic Services Inc
AI/MLData ScienceHuggingface
+5
💰 Gehalt: keine Angabe
🏢 Vor Ort
US - Various
🗣️ EN