Software Engineer, Evaluation Platform / Infra

Thinking Machines · San Francisco · $300k–$475k

The role

This role involves designing and building a self-serve platform for authoring, running, and analyzing model evaluations. The engineer will work across Python frameworks, data pipelines, APIs, and user interfaces, collaborating with research teams to improve evaluation workflows.

Pay
$300k–$475k
Location
San Francisco
Work mode
Onsite
Employment
Intern
Level
Mid
Experience
2+ yrs
Education
Bachelors
Sponsorship
Offered

What we know that the posting doesn’t say

  • Seen 1 day agostill listed on the employer’s careers page
  • Posted 9 days agothe first time we saw it

About Thinking Machines

Thinking Machines builds AI that extends human will and judgment, training frontier models and developing tools for model customization and human-AI communication.

What you would do

  • Design and build the evaluation platform for authoring, running, and analyzing model evaluations
  • Work across evaluation libraries, distributed systems, data pipelines, APIs, and user-facing applications
  • Build flexible abstractions for evaluation tasks, environments, graders, datasets, and model outputs
  • Make evaluation results reproducible through versioning, provenance, observability, and quality controls
  • Partner with researchers to turn bespoke workflows into self-serve systems
  • Work with the Research Tooling engineering team

Must have

  • Bachelor's degree or equivalent in CS, engineering, ML, or related field
  • Two years of post-grad software or ML engineering experience
  • Experience building evaluations, benchmarks, or model-quality systems for LLMs or multimodal models
  • Strong software engineering fundamentals and system reliability experience
  • Proficiency in at least one backend language (Python or Rust)
  • Experience with databases, data pipelines, or distributed systems
  • Comfort working across the stack from problem discovery to deployment
  • Experience collaborating with cross-functional partners and subject-matter experts

Nice to have

  • Track record of building frameworks, SDKs, or developer tools with thoughtful abstractions
  • Experience with distributed job execution, workflow orchestration, or sandboxed environments
  • Experience building interfaces for inspecting complex data or debugging model behavior
  • Familiarity with LLM or multimodal model evaluation including model-based grading
  • Experience working closely with researchers to turn evolving needs into durable systems
  • Experience at a startup building technically complex products end to end

What you get

  • Generous health, dental, and vision benefits
  • Unlimited paid time off
  • Paid parental leave
  • Relocation support as needed

Experience and education

  • 2+ years of relevant experience
  • Bachelors degree or equivalent experience

Key skills

  • python
  • rust
  • react
  • typescript
  • databases
  • data pipelines
  • distributed systems
  • evaluation frameworks
  • benchmarks
  • graders
  • model-quality systems
  • large language models
  • multimodal models
  • versioning
  • provenance
  • observability
  • failure recovery
  • quality controls
  • workflow orchestration
  • sandboxed environments
  • large-scale data processing
  • model-based grading
  • human evaluation
  • synthetic data

ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future wort…

Extracted from the employer’s posting. Read it in full on Thinking Machines’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Thinking Machines

All 17 roles at Thinking Machines