Staff Engineer, Distributed Storage and HPC & AI Infrastructure
togetherai · San Francisco
- Posted
- 50 days ago
- Last confirmed live
- Today
What this role involves
Staff engineer responsible for architecting and scaling multi-petabyte distributed storage systems for AI training and inference. Manages high-performance parallel filesystems and object stores, develops Kubernetes-native storage operators, and optimizes data paths for GPU clusters. Requires deep expertise in distributed storage, Kubernetes, and programming in Go and Python.
Skills this posting asks for
- vast
- weka
- ceph
- lustre
- kubernetes
- go
- python
- terraform
- ansible
- helm
- argocd
- prometheus
- grafana
- thanos
- csi
- s3
- minio
- r2
- rdma
- infiniband
- ext4
- xfs
- lvm
- nvme
Requirements
- 8 years of experience
- Level: staff
From the employer’s posting
About the Role In this role, you will operate, scale, and optimize multi-petabyte storage systems purpose-built for the world’s largest AI training and inference workloads. You’ll manage and scale high-performance parallel filesystems and object stores, evaluate and int…
Read the full description on togetherai’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at togetherai
- GTM Data Analytics EngineerSan Francisco
- Senior Product Manager, Model APIs & Developer ExperienceSan Francisco
- Senior Software Engineer - Together Cloud InfrastructureSan Francisco
- AI Infrastructure Systems EngineerSan Francisco
- Software Engineer, Customer InsightsSan Francisco
- Senior Product Engineer, FullstackSan Francisco