Senior Machine Learning Engineer, Model Training and Reinforcement Learning

nebius · Palo Alto, California, United States

Posted
3 days ago
Last confirmed live
2 days ago

What this role involves

Nebius is hiring a Senior ML Systems Engineer to build and maintain large-scale training and RL infrastructure for frontier model improvement. The role involves integrating distributed training frameworks, debugging GPU performance, and partnering with research scientists to turn algorithmic recipes into scalable systems.

Skills this posting asks for

  • megatron-lm
  • deepspeed
  • pytorch
  • fsdp
  • dtensor
  • ray
  • verl
  • slime
  • areal
  • openrlhf
  • nccl
  • cuda
  • distributed training
  • reinforcement learning
  • gpu profiling
  • parallelism strategies
  • tensor parallelism
  • pipeline parallelism
  • sequence parallelism
  • context parallelism
  • expert parallelism
  • data parallelism
  • checkpointing
  • experiment orchestration

Requirements

  • Level: senior

From the employer’s posting

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, w…

Read the full description on nebius’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at nebius

All 30 roles at nebius