AI Training Infrastructure Engineer

Designworkstalent · Remote

Posted
33 days ago
Last confirmed live
1 day ago

What this role involves

This role involves building and scaling distributed training infrastructure for large AI models, focusing on reliability, efficiency, and operational excellence across GPU clusters. The position is hybrid in Bellevue, WA, and targets senior and staff levels. The company is a well-funded AI infrastructure startup backed by a parent organization.

Skills this posting asks for

  • pytorch distributed
  • deepspeed
  • megatron-lm
  • ray
  • distributed training
  • gpu clusters
  • machine learning infrastructure
  • distributed systems
  • model training
  • fault tolerance
  • checkpointing
  • automation
  • python

Requirements

  • Level: senior
  • Remote policy: hybrid

From the employer’s posting

AI TRAINING INFRASTRUCTURE ENGINEER Location: Hybrid | Bellevue, WA Area Titles: Senior and Staff (multiple roles available) BUILD THE TRAINING INFRASTRUCTURE POWERING NEXT-GENERATION AI MODELS ABOUT THE OPPORTUNITY A well-funded, rapidly growing AI infrastructure company is building a next…

Read the full description on Designworkstalent’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Designworkstalent

All 3 roles at Designworkstalent