Staff ML Engineer, Generative Model Performance & Efficiency

waymo · Mountain View, California, United States, New York City, New York, United States

Posted
4 days ago
Last confirmed live
1 day ago

What this role involves

Waymo is hiring a Staff ML Engineer to optimize the performance and efficiency of generative models used for simulation. The role involves analyzing model architectures, applying compression and parallelism techniques, and building profiling tools. Candidates need an MS/PhD in CS/ML/Robotics, 5+ years of deep learning experience, and proficiency in JAX/Flax with expertise in model optimization.

Skills this posting asks for

  • deep learning
  • transformers
  • diffusion models
  • mixture of experts
  • quantization
  • fp8
  • int4
  • pruning
  • knowledge distillation
  • efficient attention mechanisms
  • tpus
  • gpus
  • xla
  • model parallelism
  • data parallelism
  • tensor parallelism
  • pipeline parallelism
  • expert parallelism
  • low-latency serving
  • high-throughput serving
  • training pipeline optimization
  • profiling
  • xprof
  • perfetto

Requirements

  • 5 years of experience
  • Level: staff

From the employer’s posting

Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve acce…

Read the full description on waymo’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at waymo

All 128 roles at waymo