← Zurück zur Liste
Stelle

Technical Support Engineer (Inference) - US Weekends

IT Support / Helpdesk • Remote • Vollzeit Europäische Union EU/EMEA

Provide technical support to Together AI customers building training, fine-tuning, and inference solutions on GPU clusters running on Kubernetes.

Responsibilities

  • Engage directly with customers to resolve complex technical challenges involving GPU clusters and inference/fine-tuning services
  • Act as a customer-facing SRE ensuring inference endpoints running on Kubernetes remain healthy, stable, and performant
  • Become a product expert in Gen AI solutions, serving as the last line of technical defense before Engineering and Product escalation
  • Assist with hardware and platform migrations by validating system health and traffic routing
  • Monitor dashboards to detect anomalies and escalate with data-backed analysis
  • Manage customer-facing communications during incidents, translating technical findings into clear, evidence-backed updates
  • Execute infrastructure changes via pull requests (infra-as-code) for endpoint configuration, model bringup/bringdown, and capacity scaling
  • Flag engine-level bugs with logs and reproduction steps for engineering
  • Collaborate across Engineering, Research, and Product teams to address customer concerns
  • Identify patterns in support cases to help drive the product roadmap
  • Maintain detailed documentation of system configurations, procedures, and troubleshooting guides
  • Provide flexible support coverage during holidays, nights, and weekends

Requirements

  • 6+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering, with at least 1 year supporting an AI service
  • Experience as an SRE or DevOps engineer working with Kubernetes
  • Strong technical background with knowledge of AI, ML, and GPU technology

Soft skills

Customer-facing communicationCross-team collaborationFlexibility around weekend/on-call scheduling

About the company

Together AI is a pioneering AI company providing training, fine-tuning, and inference solutions to customers on GPU clusters.

Ähnliche Stellen