Research Engineer, Infrastructure, Training Systems

Thinking Machines · Remote · $350k–$475k

The role

This role involves designing and building the core systems that enable scalable, efficient training of large models. The goal is to make experimentation and training fast and reliable, allowing research teams to focus on science. The position is ideal for someone with deep systems and performance expertise and a curiosity for machine learning at scale.

Pay
$350k–$475k
Location
Remote
Work mode
Onsite
Level
Mid
Education
Bachelors
Sponsorship
Offered

What we know that the posting doesn’t say

  • Seen 1 day agostill listed on the employer’s careers page
  • Posted 35 days agothe first time we saw it

About Thinking Machines

Thinking Machines builds AI that extends human will and judgment, training frontier models and developing interfaces to broaden human-AI communication.

What you would do

  • Design and implement distributed training systems scaling across thousands of GPUs.
  • Develop high-performance optimizations to maximize throughput and efficiency.
  • Build reusable frameworks for training reproducibility, reliability, and scalability.
  • Establish standards for reliability, maintainability, and security.
  • Collaborate with researchers and engineers on scalable infrastructure.
  • Publish learnings through documentation, open-source libraries, or technical reports.

Must have

  • Bachelor's degree or equivalent in computer science, electrical engineering, or similar.
  • Strong engineering skills for performant, maintainable code and debugging.
  • Understanding of deep learning frameworks like PyTorch or JAX.
  • Ability to thrive in a highly collaborative environment.
  • Bias for action and initiative across different stacks and teams.

Nice to have

  • Experience with distributed training for the world's largest models.
  • Track record of improving research productivity through infrastructure design.
  • Contributions to open-source ML infrastructure like PyTorch, XLA, Megatron-LM, or DeepSpeed.

What you get

  • Generous health, dental, and vision benefits.
  • Unlimited paid time off.
  • Paid parental leave.
  • Relocation support as needed.

Experience and education

  • Bachelors degree or equivalent experience

Key skills

  • pytorch
  • jax
  • xla
  • megatron-lm
  • deepspeed
  • distributed training
  • high-performance computing
  • infrastructure design
  • debugging

ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future wort…

Extracted from the employer’s posting. Read it in full on Thinking Machines’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Thinking Machines

All 17 roles at Thinking Machines