Machine Learning Engineer - Voice Conversion
Cantina · Remote (U.S. or Europe) · $200k–$220k
- Posted
- 18 days ago
- Last confirmed live
- 1 day ago
- Published range
- $200k–$220k
What this role involves
Cantina is seeking a Research/ML Engineer to join their Speech Team, focusing on building state-of-the-art speech systems end-to-end, including voice conversion and controllable TTS. The role involves model building, experimental design, tool development, and full-stack contribution, with a strong emphasis on large-scale audio models and production deployment. The position offers a salary range of $200,000-$220,000.
Skills this posting asks for
- pytorch
- diffusion transformers
- flow-matching transformers
- audio vae
- neural audio codecs
- vocoders
- distributed training
- fsdp
- deepspeed
- cuda
- triton
- c++
- voice cloning
- speech control
- expressive speech generation
- grpo
- dpo
- asr
- wer
- sv
Requirements
- Level: senior
From the employer’s posting
About Cantina: Cantina is a new social platform founded by Sean Parker with the most advanced AI character creator. Our bots are lifelike, social creatures that can interact wherever people are online—across voice, video, and text. Create yourself, imagine someone new, or choose from thousands of…
Read the full description on Cantina’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Cantina
- Media Software Engineer, Speech (Senior-Staff Levels)Remote
- Staff Software Engineer, PlatformRemote (U.S. or Canada)
- Machine Learning Engineer, Speech - Joint Audio-Video ModelingRemote (U.S. or Europe)
- Machine Learning Engineer, OpsRemote (U.S. or Europe)
- Staff Software Engineer, BackendLos Angeles or San Francisco
- Product Manager, WebCalifornia