Research Engineer - RL Infrastructure
Prime Intellect · San Francisco
- Posted
- 49 days ago
- Last confirmed live
- 2 days ago
What this role involves
Prime Intellect is hiring a Research Engineer to build and optimize systems infrastructure for large-scale RL and distributed training. The role involves contributing to the prime-rl framework, improving training efficiency, and implementing low-level performance optimizations. Ideal candidates have strong systems engineering experience in AI/ML infrastructure and familiarity with distributed training frameworks.
Skills this posting asks for
- pytorch
- pytorch distributed
- deepspeed
- fsdp
- megatron
- vllm
- ray
- cuda
- triton
- distributed training
- data parallelism
- tensor parallelism
- pipeline parallelism
- gpu architecture
- performance profiling
- kernel optimization
- memory optimization
- communication optimization
- compiler optimization
- runtime optimization
- rl infrastructure
- async rollout
- multi-node gpu clusters
- high-performance networking
From the employer’s posting
OWN YOUR INTELLIGENCE Prime Intellect is building the open superintelligence stack: the infrastructure frontier AI labs build internally, made available to every ambitious AI team. Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and…
Read the full description on Prime Intellect’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Prime Intellect
- Member of Technical Staff - Full Stack Software EngineerSan Francisco
- Research Engineer - Distributed TrainingSan Francisco
- Solutions Architect - AI InfrastructureSan Francisco