Research Engineer - RL Infrastructure

Prime Intellect · San Francisco

Posted
49 days ago
Last confirmed live
2 days ago

What this role involves

Prime Intellect is hiring a Research Engineer to build and optimize systems infrastructure for large-scale RL and distributed training. The role involves contributing to the prime-rl framework, improving training efficiency, and implementing low-level performance optimizations. Ideal candidates have strong systems engineering experience in AI/ML infrastructure and familiarity with distributed training frameworks.

Skills this posting asks for

  • pytorch
  • pytorch distributed
  • deepspeed
  • fsdp
  • megatron
  • vllm
  • ray
  • cuda
  • triton
  • distributed training
  • data parallelism
  • tensor parallelism
  • pipeline parallelism
  • gpu architecture
  • performance profiling
  • kernel optimization
  • memory optimization
  • communication optimization
  • compiler optimization
  • runtime optimization
  • rl infrastructure
  • async rollout
  • multi-node gpu clusters
  • high-performance networking

From the employer’s posting

OWN YOUR INTELLIGENCE Prime Intellect is building the open superintelligence stack: the infrastructure frontier AI labs build internally, made available to every ambitious AI team. Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and…

Read the full description on Prime Intellect’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Prime Intellect

All 4 roles at Prime Intellect