Site Reliability Engineer, Production

Thinking Machines · San Francisco · $350k–$475k

The role

This role focuses on driving end-to-end reliability for Tinker, a fine-tuning API. The engineer will define service level objectives, design monitoring, and lead incident response. They will work with engineering and research teams to harden multi-tenant isolation and resource scheduling.

Pay
$350k–$475k
Location
San Francisco
Work mode
Onsite
Level
Mid
Education
Bachelors
Sponsorship
Offered

What we know that the posting doesn’t say

  • Seen 1 day agostill listed on the employer’s careers page
  • Posted 8 days agothe first time we saw it

About Thinking Machines

Thinking Machines builds AI to extend human will and judgment, training frontier models and developing interfaces for human-AI communication.

What you would do

  • Define and own end-to-end reliability from CI/CD to production observability.
  • Develop service level objectives for distributed training systems.
  • Design and implement monitoring across the full training path.
  • Drive incident response for platform issues and ensure systematic improvements.
  • Harden multi-tenant isolation and resource scheduling for workload co-scheduling.
  • Collaborate with security teams to address production vulnerabilities.

Must have

  • Bachelor's degree or equivalent in computer science or engineering.
  • Experience in distributed systems, cloud infrastructure, or site reliability engineering.
  • Proficiency writing software to solve reliability problems.
  • Experience with production incident response and postmortems.
  • Strong communication and coordination skills across teams.

Nice to have

  • Deep experience operating production cloud services at scale.
  • Background in distributed training frameworks and infrastructure failures.
  • Track record building checkpoint and recovery systems for long-running jobs.
  • Expertise in Kubernetes at scale with heterogeneous GPU workloads.

What you get

  • Generous health, dental, and vision benefits.
  • Unlimited paid time off.
  • Paid parental leave.
  • Relocation support as needed.

Experience and education

  • Bachelors degree or equivalent experience

Key skills

  • distributed systems
  • cloud infrastructure
  • site reliability engineering
  • kubernetes
  • gpu workloads
  • incident response
  • monitoring
  • observability
  • automation
  • tooling
  • distributed training
  • checkpoint systems
  • recovery systems
  • multi-tenant isolation
  • resource scheduling

ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future wort…

Extracted from the employer’s posting. Read it in full on Thinking Machines’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Thinking Machines

All 17 roles at Thinking Machines