Research Engineer, Infrastructure, Numerics

Thinking Machines · Remote · $350k–$475k

The role

This role focuses on improving the numerical foundations of distributed training stacks for large-scale LLMs. The engineer will design and optimize infrastructure, implement low-precision numerics, and collaborate with research teams. It is an evergreen role for expressing interest.

Pay
$350k–$475k
Location
Remote
Work mode
Onsite
Level
Mid
Education
Bachelors
Sponsorship
Offered

What we know that the posting doesn’t say

  • Seen 1 day agostill listed on the employer’s careers page
  • Posted 35 days agothe first time we saw it

About Thinking Machines

Thinking Machines builds AI that extends human will and judgment, training frontier models and developing interfaces for human-AI communication.

What you would do

  • Design and optimize distributed training infrastructure for large-scale LLMs.
  • Implement and evaluate low-precision numerics like BF16, MXFP8, NVFP4.
  • Develop kernels and communication primitives for mixed and low-precision arithmetic.
  • Collaborate with research teams on model architectures and training recipes.
  • Prototype and benchmark scaling strategies like data, tensor, and pipeline parallelism.
  • Contribute to internal orchestration and monitoring systems for distributed experiments.

Must have

  • Bachelor's degree or equivalent in relevant field.
  • Understanding of deep learning frameworks like PyTorch or JAX.
  • Thrive in collaborative cross-functional environment.
  • Bias for action and initiative across different stacks and teams.
  • Strong engineering skills in floating-point numerics and distributed systems.

Nice to have

  • Familiarity with distributed frameworks like PyTorch/XLA, DeepSpeed, Megatron-LM.
  • Experience implementing FP8, INT8, or block-floating point formats.
  • Prior contributions to open-source deep learning infrastructure.
  • Publications related to numerical optimization or systems for large models.
  • Experience training and supporting large-scale AI models.
  • Track record of improving research productivity through infrastructure design.

What you get

  • Generous health, dental, and vision benefits.
  • Unlimited paid time off.
  • Paid parental leave.
  • Relocation support as needed.

Experience and education

  • Bachelors degree or equivalent experience

Key skills

  • pytorch
  • jax
  • pytorch/xla
  • deepspeed
  • megatron-lm
  • bf16
  • mxfp8
  • nvfp4
  • fp8
  • int8
  • block-floating point
  • distributed systems
  • low-precision arithmetic
  • floating-point numerics
  • kernel optimization
  • communication frameworks
  • data parallelism
  • tensor parallelism
  • pipeline parallelism

ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future wort…

Extracted from the employer’s posting. Read it in full on Thinking Machines’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Thinking Machines

All 17 roles at Thinking Machines