Research Engineer, Infrastructure, Training Systems
Thinking Machines · Remote · $350k–$475k
The role
This role involves designing and building the core systems that enable scalable, efficient training of large models. The goal is to make experimentation and training fast and reliable, allowing research teams to focus on science. The position is ideal for someone with deep systems and performance expertise and a curiosity for machine learning at scale.
- Pay
- $350k–$475k
- Location
- Remote
- Work mode
- Onsite
- Level
- Mid
- Education
- Bachelors
- Sponsorship
- Offered
What we know that the posting doesn’t say
- Seen 1 day agostill listed on the employer’s careers page
- Posted 35 days agothe first time we saw it
About Thinking Machines
Thinking Machines builds AI that extends human will and judgment, training frontier models and developing interfaces to broaden human-AI communication.
What you would do
- Design and implement distributed training systems scaling across thousands of GPUs.
- Develop high-performance optimizations to maximize throughput and efficiency.
- Build reusable frameworks for training reproducibility, reliability, and scalability.
- Establish standards for reliability, maintainability, and security.
- Collaborate with researchers and engineers on scalable infrastructure.
- Publish learnings through documentation, open-source libraries, or technical reports.
Must have
- Bachelor's degree or equivalent in computer science, electrical engineering, or similar.
- Strong engineering skills for performant, maintainable code and debugging.
- Understanding of deep learning frameworks like PyTorch or JAX.
- Ability to thrive in a highly collaborative environment.
- Bias for action and initiative across different stacks and teams.
Nice to have
- Experience with distributed training for the world's largest models.
- Track record of improving research productivity through infrastructure design.
- Contributions to open-source ML infrastructure like PyTorch, XLA, Megatron-LM, or DeepSpeed.
What you get
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
Experience and education
- Bachelors degree or equivalent experience
Key skills
- pytorch
- jax
- xla
- megatron-lm
- deepspeed
- distributed training
- high-performance computing
- infrastructure design
- debugging
ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future wort…
Extracted from the employer’s posting. Read it in full on Thinking Machines’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Thinking Machines
- Product Manager - Post TrainingRemote
- Software Engineer, ProductRemote
- Site Reliability Engineer, Post TrainingSan Francisco
- Site Reliability Engineer, ProductionSan Francisco
- Software Engineer, Evaluation Platform / InfraSan Francisco
- Software Engineer, Research ToolsRemote