ML Infrastructure Engineer

nebius · Remote

Posted
3 days ago
Last confirmed live
2 days ago

What this role involves

Seeking an ML Infrastructure Engineer to lead GPU platform benchmarking for ML/AI workloads. Responsibilities include profiling GPU performance, debugging and optimizing workloads, and developing tools for performance metrics. Requires deep understanding of deep learning frameworks and GPU stack.

Skills this posting asks for

  • pytorch
  • jax
  • megatron-lm
  • tensorrt-llm
  • cuda
  • nccl
  • docker
  • kubernetes
  • python
  • vllm
  • sglang
  • tensorrt

From the employer’s posting

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, w…

Read the full description on nebius’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at nebius

All 30 roles at nebius