Machine Learning Engineer - Voice Conversion
What engineering roles in crypto pay
218 salaries · our own dataThis role pays $200k-$220k, above the $202k median for engineering roles in crypto on this board.
As a Research/ML Engineer on the Speech Team at Cantina, you'll build state-of-the-art speech systems end-to-end, from data specifications through production inference. You'll drive the model-data-eval flywheel for voice conversion and adjacent tasks like controllable TTS and voice design, partnering closely with research, data, and infrastructure teams to ship fast, reliable, and cost-aware models.
What you'll do
- Architect, implement, pre-train, fine-tune, and post-train/align large-scale speech models using methods like GRPO and DPO.
- Design, run, and analyze scientific experiments to advance your understanding of model behaviour and capabilities.
- Develop and improve developer tooling to enhance team productivity.
- Contribute to the entire stack, from low-level optimizations to high-level model design.
- Define data requirements and collaborate on acquisition, curation, augmentation, labeling quality, and synthetic data strategies.
- Design automated objective and subjective evaluations, including listening tests, WER/ASR-based metrics, robustness and bias checks, and red-team studies.
- Harden the training, evaluation, and inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback.
- Contribute to safety and consent guardrails and to misuse/abuse mitigation for responsible speech technology.
What you bring
- Exceptional research and development experience with large-scale audio models (more than 8B parameters and more than 500k hours of data).
- Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.
- Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders; experience with latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training.
- Strong experience with multi-node, multi-GPU distributed training using FSDP, DeepSpeed, or equivalent.
- Strong software engineering skills and a proven track record of building complex systems.
- Strong proficiency with PyTorch and performance work including profiling, CUDA, Triton, and C++ as needed; ability to write reliable production-quality code.
- Shipped large-scale speech, audio, or multimodal generative models to production.
- Background working with large-scale ML data and the ability to iterate on data quality using both subjective and objective signals.
- Experience with voice cloning, speech control/steerability, or expressive speech generation.
- Notable publications and/or open-source contributions in speech, audio, or ML.
What we offer
- Annual base salary of $200,000-$220,000 (€170,000-€190,000).
- Generous company equity.
- Medical, dental, and vision insurance with 99.99% of premiums covered by Cantina.
- 42 days of paid time off, including 15 PTO days, 10 sick days, 15 company holidays, and 2 floating holidays.
- Generous parental leave and fertility support.
- 401(k) retirement savings plan.
- Lifestyle spending account of $500 per month.
- Complimentary lunch and snacks for in-office employees.
- One Medical membership.
About Cantina
Cantina is a social platform founded by Sean Parker that combines advanced AI character creation with seamless interaction across voice, video, and text. The platform enables users to create personalized AI characters, generate scalable content, and participate in group conversations powered by lifelike, social AI bots.
