AI Training Infrastructure Engineer
Designworkstalent · Remote
- Posted
- 33 days ago
- Last confirmed live
- 1 day ago
What this role involves
This role involves building and scaling distributed training infrastructure for large AI models, focusing on reliability, efficiency, and operational excellence across GPU clusters. The position is hybrid in Bellevue, WA, and targets senior and staff levels. The company is a well-funded AI infrastructure startup backed by a parent organization.
Skills this posting asks for
- pytorch distributed
- deepspeed
- megatron-lm
- ray
- distributed training
- gpu clusters
- machine learning infrastructure
- distributed systems
- model training
- fault tolerance
- checkpointing
- automation
- python
Requirements
- Level: senior
- Remote policy: hybrid
From the employer’s posting
AI TRAINING INFRASTRUCTURE ENGINEER Location: Hybrid | Bellevue, WA Area Titles: Senior and Staff (multiple roles available) BUILD THE TRAINING INFRASTRUCTURE POWERING NEXT-GENERATION AI MODELS ABOUT THE OPPORTUNITY A well-funded, rapidly growing AI infrastructure company is building a next…
Read the full description on Designworkstalent’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.