Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

plaud · Remote

Posted
107 days ago
Last confirmed live
1 day ago

What this role involves

Plaud is seeking a Machine Learning Engineer focused on inference and serving for speech LLMs. The role involves building and deploying high-throughput, low-latency inference engines, optimizing GPU performance, and implementing real-time audio streaming. The position sits between the ML training and backend infrastructure teams.

Skills this posting asks for

  • inference engines
  • large language models
  • speech models
  • continuous batching
  • kv cache management
  • pagedattention
  • gpu architectures
  • nvidia ampere
  • nvidia hopper
  • memory hierarchy
  • vllm
  • tensorrt-llm
  • sglang
  • nvidia triton inference server
  • websockets
  • webrtc
  • neural audio codecs
  • speculative decoding
  • lookahead decoding
  • chunked prefill
  • post-training quantization
  • fp8
  • int8
  • awq

From the employer’s posting

ABOUT PLAUD INC. Plaud is building the world's most trusted AI work companion for professionals to elevate productivity and performance through note-taking solutions, loved by over 1,500,000 users worldwide since 2023. With a mission to amplify human intelligence, Plaud is building the next-genera…

Read the full description on plaud’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at plaud

All 15 roles at plaud