Staff Cluster Infrastructure Engineer
atoms · San Francisco, CA
- Posted
- 16 days ago
- Last confirmed live
- 1 day ago
What this role involves
Atoms is seeking a Staff Cluster Infrastructure Engineer to own and automate the GPU compute fabric for training foundation models. The role involves managing GPU clusters, automating bare-metal bring-up, building software abstractions, and ensuring reliability and scalability. Requires 6+ years experience with GPU compute on Kubernetes and strong programming skills in Python or Go.
Skills this posting asks for
- gpu compute
- kubernetes
- python
- go
- terraform
- cloudformation
- linux
- networking
- infrastructure-as-code
- automation
- bare-metal
Requirements
- 6 years of experience
- Level: staff
- Remote policy: onsite
From the employer’s posting
Who we are Atoms is building the machines that power the next era of progress. Over the last decade, software has transformed the digital world. But the physical world, where food is made, minerals are mined, goods are moved, and industries are run, remains far les…
Read the full description on atoms’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at atoms
- Recruiting Technical Program ManagerSan Francisco, CA
- Software Engineer, OnboardSan Francisco, CA
- Product Manager, Integration & EcosystemMountain View, CA
- Robotics Test EngineerSan Francisco, CA
- Senior Full Stack EngineerSan Francisco, CA
- Senior Machine Learning EngineerSan Francisco, CA