Research Engineer, Infrastructure, Kernels
Thinking Machines · Remote · $350k–$475k
The role
This role involves designing, optimizing, and maintaining the compute foundations for large-scale language model training. The engineer will develop high-performance ML kernels, enable efficient low-precision arithmetic, and improve the distributed compute stack. Collaboration with researchers and systems architects is key to bridging algorithmic design with hardware efficiency.
- Pay
- $350k–$475k
- Location
- Remote
- Work mode
- Onsite
- Level
- Mid
- Education
- Bachelors
- Sponsorship
- Offered
What we know that the posting doesn’t say
- Seen 1 day agostill listed on the employer’s careers page
- Posted 35 days agothe first time we saw it
About Thinking Machines
Thinking Machines builds AI that extends human will and judgment, training frontier models and developing interfaces to broaden human-AI communication.
What you would do
- Design and implement custom ML kernels for core LLM operations.
- Reduce memory bandwidth bottlenecks and improve kernel compute efficiency.
- Collaborate with research teams to align kernel optimizations with model architecture.
- Develop and maintain a library of reusable kernels and performance benchmarks.
- Contribute to infrastructure stability, scalability, and reproducibility.
- Document and share insights through talks, papers, or open-source contributions.
Must have
- Bachelor's degree or equivalent in relevant field.
- Strong engineering skills for performant, maintainable code.
- Understanding of deep learning frameworks and their system architectures.
- Thrive in a highly collaborative environment.
- Bias for action and initiative across different stacks and teams.
- Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.
- Ability to analyze, profile, and optimize compute-intensive workloads.
Nice to have
- Experience training or supporting large-scale language models.
- Track record of improving research productivity through infrastructure design.
- Experience developing or tuning kernels for deep learning frameworks.
- Familiarity with tensor parallelism, pipeline parallelism, or distributed frameworks.
- Experience implementing low-precision formats or contributing to compiler stacks.
- Contributions to open-source GPU, ML systems, or compiler optimization projects.
- Prior research or engineering experience in numerical optimization or scalable AI infrastructure.
What you get
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
Experience and education
- Bachelors degree or equivalent experience
Key skills
- cuda
- cute
- triton
- pytorch
- jax
- gpu programming
- deep learning
- kernel optimization
- low-precision arithmetic
- distributed computing
- tensor parallelism
- pipeline parallelism
- xla
- tvm
- fp8
- int8
- block floating point
ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future wort…
Extracted from the employer’s posting. Read it in full on Thinking Machines’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Thinking Machines
- Product Manager - Post TrainingRemote
- Software Engineer, ProductRemote
- Site Reliability Engineer, Post TrainingSan Francisco
- Site Reliability Engineer, ProductionSan Francisco
- Software Engineer, Evaluation Platform / InfraSan Francisco
- Software Engineer, Research ToolsRemote