Research Engineer - Distributed Training
Prime Intellect · San Francisco
- Posted
- 49 days ago
- Last confirmed live
- 2 days ago
What this role involves
Prime Intellect is hiring a Research Engineer focused on distributed training. The role involves building and optimizing distributed training infrastructure for large-scale pre-training and RL workloads, contributing to open-source frameworks like prime-rl. The ideal candidate has strong systems engineering experience in AI/ML infrastructure, deep familiarity with PyTorch and distributed training frameworks, and experience optimizing training performance.
Skills this posting asks for
- pytorch
- pytorch distributed
- deepspeed
- fsdp
- megatron
- vllm
- ray
- cuda
- triton
- distributed training
- data parallelism
- tensor parallelism
- pipeline parallelism
- gpu architecture
- performance profiling
- kernel optimization
- communication optimization
- memory optimization
- rl training
- async rollout
- compiler optimization
- runtime optimization
- multi-node gpu clusters
- high-performance networking
From the employer’s posting
OWN YOUR INTELLIGENCE Prime Intellect is building the open superintelligence stack: the infrastructure frontier AI labs build internally, made available to every ambitious AI team. Our platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and…
Read the full description on Prime Intellect’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Prime Intellect
- Member of Technical Staff - Full Stack Software EngineerSan Francisco
- Research Engineer - RL Infrastructure San Francisco
- Solutions Architect - AI InfrastructureSan Francisco