Research Engineer - Distributed Training

Prime Intellect · San Francisco

Posted
49 days ago
Last confirmed live
2 days ago

What this role involves

Prime Intellect is hiring a Research Engineer focused on distributed training. The role involves building and optimizing distributed training infrastructure for large-scale pre-training and RL workloads, contributing to open-source frameworks like prime-rl. The ideal candidate has strong systems engineering experience in AI/ML infrastructure, deep familiarity with PyTorch and distributed training frameworks, and experience optimizing training performance.

Skills this posting asks for

  • pytorch
  • pytorch distributed
  • deepspeed
  • fsdp
  • megatron
  • vllm
  • ray
  • cuda
  • triton
  • distributed training
  • data parallelism
  • tensor parallelism
  • pipeline parallelism
  • gpu architecture
  • performance profiling
  • kernel optimization
  • communication optimization
  • memory optimization
  • rl training
  • async rollout
  • compiler optimization
  • runtime optimization
  • multi-node gpu clusters
  • high-performance networking

From the employer’s posting

OWN YOUR INTELLIGENCE Prime Intellect is building the open superintelligence stack: the infrastructure frontier AI labs build internally, made available to every ambitious AI team. Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and…

Read the full description on Prime Intellect’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Prime Intellect

All 4 roles at Prime Intellect