Senior Backend Engineer, Inference Platform
togetherai · San Francisco
- Posted
- 43 days ago
- Last confirmed live
- 2 days ago
What this role involves
Senior Backend Engineer responsible for building and optimizing the inference platform for generative AI models. Works with large-scale GPU infrastructure, contributes to open-source inference projects, and requires deep experience in distributed systems and low-level performance optimization.
Skills this posting asks for
- rust
- go
- python
- typescript
- kubernetes
- container orchestration
- cuda
- triton
- nccl
- infiniband
- nvlink
- mpi
- sglang
- vllm
- nvidia dynamo
- load balancing
- auto-scaling
- multi-tenant
- prefix caching
- distributed systems
- os concepts
- llms
- generative models
Requirements
- 5 years of experience
- Level: senior
From the employer’s posting
About the Role Together AI is building the Inference Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling developers, enterprises, and researchers to harness the…
Read the full description on togetherai’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at togetherai
- Senior Product Manager, Model APIs & Developer ExperienceSan Francisco
- Senior Software Engineer - Together Cloud InfrastructureSan Francisco
- Software Engineer, Customer InsightsSan Francisco
- Senior Product Engineer, FullstackSan Francisco
- Staff Software Engineer, Inference / Compute Infrastructure EngineeringSan Francisco
- Senior Network EngineerSan Francisco