Research Engineer, Infrastructure, Numerics
Thinking Machines · Remote · $350k–$475k
The role
This role focuses on improving the numerical foundations of distributed training stacks for large-scale LLMs. The engineer will design and optimize infrastructure, implement low-precision numerics, and collaborate with research teams. It is an evergreen role for expressing interest.
- Pay
- $350k–$475k
- Location
- Remote
- Work mode
- Onsite
- Level
- Mid
- Education
- Bachelors
- Sponsorship
- Offered
What we know that the posting doesn’t say
- Seen 1 day agostill listed on the employer’s careers page
- Posted 35 days agothe first time we saw it
About Thinking Machines
Thinking Machines builds AI that extends human will and judgment, training frontier models and developing interfaces for human-AI communication.
What you would do
- Design and optimize distributed training infrastructure for large-scale LLMs.
- Implement and evaluate low-precision numerics like BF16, MXFP8, NVFP4.
- Develop kernels and communication primitives for mixed and low-precision arithmetic.
- Collaborate with research teams on model architectures and training recipes.
- Prototype and benchmark scaling strategies like data, tensor, and pipeline parallelism.
- Contribute to internal orchestration and monitoring systems for distributed experiments.
Must have
- Bachelor's degree or equivalent in relevant field.
- Understanding of deep learning frameworks like PyTorch or JAX.
- Thrive in collaborative cross-functional environment.
- Bias for action and initiative across different stacks and teams.
- Strong engineering skills in floating-point numerics and distributed systems.
Nice to have
- Familiarity with distributed frameworks like PyTorch/XLA, DeepSpeed, Megatron-LM.
- Experience implementing FP8, INT8, or block-floating point formats.
- Prior contributions to open-source deep learning infrastructure.
- Publications related to numerical optimization or systems for large models.
- Experience training and supporting large-scale AI models.
- Track record of improving research productivity through infrastructure design.
What you get
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
Experience and education
- Bachelors degree or equivalent experience
Key skills
- pytorch
- jax
- pytorch/xla
- deepspeed
- megatron-lm
- bf16
- mxfp8
- nvfp4
- fp8
- int8
- block-floating point
- distributed systems
- low-precision arithmetic
- floating-point numerics
- kernel optimization
- communication frameworks
- data parallelism
- tensor parallelism
- pipeline parallelism
ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future wort…
Extracted from the employer’s posting. Read it in full on Thinking Machines’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Thinking Machines
- Product Manager - Post TrainingRemote
- Software Engineer, ProductRemote
- Site Reliability Engineer, Post TrainingSan Francisco
- Site Reliability Engineer, ProductionSan Francisco
- Software Engineer, Evaluation Platform / InfraSan Francisco
- Software Engineer, Research ToolsRemote