Inference Optimization ML Engineer

Rhoda AI · Mountain View

Posted
103 days ago
Last confirmed live
1 day ago

What this role involves

Rhoda AI is hiring an Inference Optimization MLE to optimize large multimodal models for production deployment. The role focuses on end-to-end inference performance, including latency, throughput, and efficiency improvements. The candidate will work closely with research and robotics teams to bridge the gap between training and real-world deployment.

Skills this posting asks for

  • inference optimization
  • ml systems
  • pytorch
  • jax
  • quantization
  • distillation
  • pruning
  • model compilation
  • tensorrt
  • torch.compile
  • xla
  • attention mechanisms
  • kv caching
  • cuda
  • triton
  • inference serving
  • triton inference server
  • vllm
  • torchserve
  • multimodal models
  • video model inference
  • edge deployment
  • cloud deployment
  • speculative decoding

Requirements

  • 3 years of experience
  • Level: mid

From the employer’s posting

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists cap…

Read the full description on Rhoda AI’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Rhoda AI

All 7 roles at Rhoda AI