Site Reliability Engineer, Production
Thinking Machines · San Francisco · $350k–$475k
The role
This role focuses on driving end-to-end reliability for Tinker, a fine-tuning API. The engineer will define service level objectives, design monitoring, and lead incident response. They will work with engineering and research teams to harden multi-tenant isolation and resource scheduling.
- Pay
- $350k–$475k
- Location
- San Francisco
- Work mode
- Onsite
- Level
- Mid
- Education
- Bachelors
- Sponsorship
- Offered
What we know that the posting doesn’t say
- Seen 1 day agostill listed on the employer’s careers page
- Posted 8 days agothe first time we saw it
About Thinking Machines
Thinking Machines builds AI to extend human will and judgment, training frontier models and developing interfaces for human-AI communication.
What you would do
- Define and own end-to-end reliability from CI/CD to production observability.
- Develop service level objectives for distributed training systems.
- Design and implement monitoring across the full training path.
- Drive incident response for platform issues and ensure systematic improvements.
- Harden multi-tenant isolation and resource scheduling for workload co-scheduling.
- Collaborate with security teams to address production vulnerabilities.
Must have
- Bachelor's degree or equivalent in computer science or engineering.
- Experience in distributed systems, cloud infrastructure, or site reliability engineering.
- Proficiency writing software to solve reliability problems.
- Experience with production incident response and postmortems.
- Strong communication and coordination skills across teams.
Nice to have
- Deep experience operating production cloud services at scale.
- Background in distributed training frameworks and infrastructure failures.
- Track record building checkpoint and recovery systems for long-running jobs.
- Expertise in Kubernetes at scale with heterogeneous GPU workloads.
What you get
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
Experience and education
- Bachelors degree or equivalent experience
Key skills
- distributed systems
- cloud infrastructure
- site reliability engineering
- kubernetes
- gpu workloads
- incident response
- monitoring
- observability
- automation
- tooling
- distributed training
- checkpoint systems
- recovery systems
- multi-tenant isolation
- resource scheduling
ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future wort…
Extracted from the employer’s posting. Read it in full on Thinking Machines’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Thinking Machines
- Product Manager - Post TrainingRemote
- Software Engineer, ProductRemote
- Site Reliability Engineer, Post TrainingSan Francisco
- Software Engineer, Evaluation Platform / InfraSan Francisco
- Software Engineer, Research ToolsRemote
- Software Engineer, SandboxingRemote