Staff Software Engineer, Inference / Compute Infrastructure Engineering
togetherai · San Francisco
- Posted
- 31 days ago
- Last confirmed live
- 2 days ago
What this role involves
This role involves building software systems that manage the lifecycle of GPU infrastructure, treating hardware provisioning as a software problem. The engineer will design state machines, declarative APIs, and self-healing mechanisms to automate cluster management for inference and training. The position requires strong software engineering skills and experience with workflow orchestration and control planes.
Skills this posting asks for
- go
- python
- rust
- temporal
- cadence
- kubernetes
- kafka
- nats
- sqs
- cuda
- gpu
- ci/cd
- infrastructure as code
- state machines
- workflow orchestration
- event-driven systems
- control planes
- reconciliation loops
Requirements
- Level: staff
From the employer’s posting
About the Role We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its full lifecycle — turning racks of GPUs into r…
Read the full description on togetherai’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at togetherai
- Senior Product Manager, Model APIs & Developer ExperienceSan Francisco
- Senior Software Engineer - Together Cloud InfrastructureSan Francisco
- Software Engineer, Customer InsightsSan Francisco
- Senior Product Engineer, FullstackSan Francisco
- Senior Network EngineerSan Francisco
- Product Manager, AI InfrastructureSan Francisco