Senior Machine Learning Engineer, Model Training and Reinforcement Learning
nebius · Palo Alto, California, United States
- Posted
- 3 days ago
- Last confirmed live
- 2 days ago
What this role involves
Nebius is hiring a Senior ML Systems Engineer to build and maintain large-scale training and RL infrastructure for frontier model improvement. The role involves integrating distributed training frameworks, debugging GPU performance, and partnering with research scientists to turn algorithmic recipes into scalable systems.
Skills this posting asks for
- megatron-lm
- deepspeed
- pytorch
- fsdp
- dtensor
- ray
- verl
- slime
- areal
- openrlhf
- nccl
- cuda
- distributed training
- reinforcement learning
- gpu profiling
- parallelism strategies
- tensor parallelism
- pipeline parallelism
- sequence parallelism
- context parallelism
- expert parallelism
- data parallelism
- checkpointing
- experiment orchestration
Requirements
- Level: senior
From the employer’s posting
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, w…
Read the full description on nebius’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at nebius
- IT Infrastructure Engineer – RMA & Hardware DiagnosticsMinnesota, United States
- AI/ML Specialist Solutions ArchitectRemote
- Frontend Engineer - User InterfaceRemote
- Infrastructure Site Reliability EngineerUnited States
- Partner Solutions ArchitectRemote
- Principal ML Solutions Architect - Token FactoryUnited States