A technical owner within Fireworks' Strategic Partnerships team, making Fireworks the default inference and fine-tuning layer across partner (primarily Microsoft Azure) AI architectures.
Responsibilities
- ▹Act as technical lead on co-sell motions with Strategic Partners, building joint reference architectures and shared POCs
- ▹Build end-to-end POCs and MVPs with partner engineering teams inside their own codebases and infrastructure
- ▹Run load tests and establish latency, throughput, and cost baselines against realistic customer traffic profiles
- ▹Deploy and validate new model families on inference frameworks (vLLM, SGLang), tuning shapes and serving patterns
- ▹Guide customers on model selection and fine-tuning strategy (SFT, DPO, RFT)
- ▹Build and run fine-tuning pipelines directly with customers
- ▹Design and implement evaluation frameworks that measure production-quality metrics, not just benchmark scores
- ▹Own the feedback loop between partner ecosystems and Fireworks engineering
- ▹Ship external technical content such as reference architectures, integration guides, and benchmark posts
- ▹Track pipeline health weekly and flag risks and opportunities to Field leadership
Requirements
- ▹3+ years in pre-sales, partner engineering, forward-deployed, or technical consulting roles
- ▹Demonstrated ability to build production software with customers, not just advise on it
- ▹Strong Python skills, comfortable reading, writing, and debugging production code
- ▹Familiarity with Kubernetes and infrastructure engineering
- ▹Hands-on fluency with LLM inference: latency/throughput tradeoffs, batching, quantization, structured outputs, function calling
- ▹Real fine-tuning experience (LoRA at minimum, RFT a strong plus)
- ▹Deep familiarity with the Azure AI stack (Azure Foundry, Azure OpenAI Service, Azure ML, AKS, Entra/RBAC)
- ▹Exceptional communication across executive and engineering levels
Nice to have
- ▹5+ years owning a technical relationship with a hyperscaler or major SI, not just supporting one
- ▹Experience with inference serving frameworks (vLLM, SGLang, TensorRT-LLM)
- ▹Prior role at a hyperscaler, AI-native cloud, or inference provider
- ▹Deep familiarity with other strategic partner stacks
- ▹Experience with agentic frameworks (LangChain, LlamaIndex, custom tool-use pipelines)
- ▹Background in model evaluation and awareness of benchmark gaming risks
- ▹Published technical blog posts or reference architectures
- ▹Track record taking GenAI POCs from prototype to production-scale deployments
Soft skills
Exceptional communication across executive and engineering audiencesStakeholder management in large, multi-party organizationsIndependent problem-solving and prioritizationCross-team and cross-partner collaboration
What we offer
- ▹$200,000-$260,000 on-target earnings plus equity
- ▹Meaningful equity in a fast-growing startup
- ▹Competitive salary and comprehensive benefits package
About the company
Fireworks is building the future of generative AI infrastructure, delivering one of the industry's fastest and most scalable inference platforms. A Series C company valued at $4 billion, backed by investors including Benchmark, Sequoia, Lightspeed, and Index, and founded by veterans of Meta PyTorch and Google Vertex AI.
Similar jobs

Job
Deployed Engineer (Federal)
LangChain
Langchain
+2
💰 Salary: not specified
🔀 Hybrid
Washington DC
🗣️ EN

Job
Technology Evangelist
Diagrid
IstioLangchain
+6
💰 Salary: not specified
🔀 Hybrid
San Francisco
🗣️ EN

Job
Deployed Engineer (Chicago)
LangChain
+3
💰 Salary: not specified
🔀 Hybrid
Chicago
🗣️ EN

Job
Deployed Engineer (Early Career-NYC)
LangChain
+3
💰 Salary: not specified
🏢 On-site
New York
🗣️ EN

Job
Deployed Engineer (Houston)
LangChain
+3
💰 Salary: not specified
🔀 Hybrid
Houston
🗣️ EN

Job
Deployed Engineer (Dallas)
LangChain
+3
💰 Salary: not specified
🔀 Hybrid
Dallas
🗣️ EN
