← Back to list

Job
Machine Learning Engineer - Voice Conversion
AI / ML Engineer
• Remote
• Full-time
•
U.S. or Europe, EU/EMEA
Cantina, a social platform founded by Sean Parker with an advanced AI character creator, is hiring a Research/ML Engineer for its Speech team to build state-of-the-art speech systems end to end, focused on voice conversion and related tasks.
Responsibilities
- ▹Architect, implement, pre-train, fine-tune and post-train/align (e.g. GRPO/DPO) large-scale speech models
- ▹Design, run and analyze scientific experiments to advance model understanding
- ▹Develop and improve dev tooling to enhance team productivity
- ▹Contribute across the full stack, from low-level optimizations to high-level model design
- ▹Define data requirements and collaborate on acquisition, curation, augmentation, labeling quality and synthetic data strategies
- ▹Design automated objective/subjective evaluations: listening tests, SV/WER/ASR-based metrics, robustness and bias checks, red-team studies
- ▹Harden the training-evaluation-inference pipeline; profile latency, memory and cost; meet production SLAs with monitoring and rollback
- ▹Contribute to safety/consent guardrails and misuse mitigation
Requirements
- ▹Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data)
- ▹Deep hands-on experience with diffusion and/or flow-matching transformers, samplers, schedules, conditioning mechanisms, and distillation
- ▹Deep hands-on experience training audio VAEs, neural audio codecs and vocoders, latent/tokenizer design, reconstruction and perceptual objectives, adversarial training
- ▹Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent)
- ▹Strong software engineering skills with a track record of building complex systems
- ▹Strong PyTorch skills and performance work (profiling, CUDA/Triton/C++ as needed), writing reliable production-quality code
- ▹Experience shipping large-scale speech/audio or multimodal generative models to production
Soft skills
Results-oriented, flexible, willing to pick up whatever moves the needleEnjoys collaborating closely with infra, data and product teamsEnjoys designing experiments, listening tests and metrics correlated with user-perceived quality
About the company
Cantina is a social platform founded by Sean Parker with the most advanced AI character creator, building lifelike social bots that interact across voice, video and text.
Similar jobs
Job
Freelance AI Engineer
hived
AI/MLAzure DevopsCpp
+14
💰 Salary: not specified
🏢 On-site
Vlaanderen

Job
ML Engineer – Robotics
Clera
AI/MLCppRos
+1
$220,000–$300,000/yr
gross
🌍 Remote
🗣️ EN

Job
AI Engineer
Armada
AI/MLCpp
+2
$154,560–$193,200/yr
gross
🏢 On-site
Bellevue Office
🗣️ EN

Job
Machine Learning Engineer, Speech - Joint Audio-Video Modeling
Cantina
AI/MLCpp
$200,000–$220,000/yr
gross
🌍 Remote
U.S. or Europe
🗣️ EN

Job
Machine Learning Engineer, TTS
Cantina
AI/MLCpp
$200,000–$220,000/yr
gross
🏢 On-site
🗣️ EN

Job
Machine Learning Engineer
Cuesta Partners
AI/MLData Science
+4
💰 Salary: not specified
🌍 Remote
🗣️ EN