Staff Cluster Infrastructure Engineer

atoms · San Francisco, CA

Posted
16 days ago
Last confirmed live
1 day ago

What this role involves

Atoms is seeking a Staff Cluster Infrastructure Engineer to own and automate the GPU compute fabric for training foundation models. The role involves managing GPU clusters, automating bare-metal bring-up, building software abstractions, and ensuring reliability and scalability. Requires 6+ years experience with GPU compute on Kubernetes and strong programming skills in Python or Go.

Skills this posting asks for

  • gpu compute
  • kubernetes
  • python
  • go
  • terraform
  • cloudformation
  • linux
  • networking
  • infrastructure-as-code
  • automation
  • bare-metal

Requirements

  • 6 years of experience
  • Level: staff
  • Remote policy: onsite

From the employer’s posting

Who we are Atoms is building the machines that power the next era of progress. Over the last decade, software has transformed the digital world. But the physical world, where food is made, minerals are mined, goods are moved, and industries are run, remains far les…

Read the full description on atoms’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at atoms

All 19 roles at atoms