Staff Engineer, Distributed Storage and HPC & AI Infrastructure

togetherai · San Francisco

Posted
50 days ago
Last confirmed live
Today

What this role involves

Staff engineer responsible for architecting and scaling multi-petabyte distributed storage systems for AI training and inference. Manages high-performance parallel filesystems and object stores, develops Kubernetes-native storage operators, and optimizes data paths for GPU clusters. Requires deep expertise in distributed storage, Kubernetes, and programming in Go and Python.

Skills this posting asks for

  • vast
  • weka
  • ceph
  • lustre
  • kubernetes
  • go
  • python
  • terraform
  • ansible
  • helm
  • argocd
  • prometheus
  • grafana
  • thanos
  • csi
  • s3
  • minio
  • r2
  • rdma
  • infiniband
  • ext4
  • xfs
  • lvm
  • nvme

Requirements

  • 8 years of experience
  • Level: staff

From the employer’s posting

About the Role In this role, you will operate, scale, and optimize multi-petabyte storage systems purpose-built for the world’s largest AI training and inference workloads. You’ll manage and scale high-performance parallel filesystems and object stores, evaluate and int…

Read the full description on togetherai’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at togetherai

All 21 roles at togetherai