Software Engineer, Supercomputing
Thinking Machines · Remote · $350k–$475k
The role
Thinking Machines seeks a software engineer to design, build, and operate GPU supercomputing environments for large-scale AI training and inference. The role involves automating clusters, writing management software, and partnering with researchers to optimize performance.
- Pay
- $350k–$475k
- Location
- Remote
- Work mode
- Onsite
- Level
- Mid
- Education
- Bachelors
- Sponsorship
- Offered
What we know that the posting doesn’t say
- Seen 1 day agostill listed on the employer’s careers page
- Posted 35 days agothe first time we saw it
About Thinking Machines
Thinking Machines builds AI that extends human will and judgment, training frontier models and developing interfaces for human-AI communication.
What you would do
- Operate and automate large GPU clusters including provisioning and capacity planning.
- Write software that abstracts cluster management for training and inference.
- Extend scheduling and orchestration for topology-aware placement and multi-tenancy.
- Monitor and improve operational metrics of speed, reliability, and error recovery.
- Build reliable storage and artifact paths for datasets, checkpoints, and logs.
- Partner with researchers to unblock scale runs and advise on performance trade-offs.
Must have
- Bachelor's degree or equivalent in computer science or engineering.
- Proficiency in Python or Rust.
- Experience operating large-scale clusters and container orchestration systems.
- Comfort operating across the stack and owning projects end-to-end.
- Thrive in collaborative environments with cross-functional partners.
- Bias for action and initiative to work across different stacks and teams.
Nice to have
- Strong systems background in Linux, networking, and infrastructure-as-code.
- Familiarity with CUDA/NCCL and performance profiling for distributed training.
- Prior work supporting large-scale model training or inference environments.
- Understanding of deep learning frameworks like PyTorch, TensorFlow, or JAX.
- Track record of working in fast-paced environments balancing care with urgency.
What you get
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
Experience and education
- Bachelors degree or equivalent experience
Key skills
- python
- rust
- kubernetes
- slurm
- cuda
- nccl
- pytorch
- tensorflow
- jax
- linux
- networking
- infrastructure-as-code
ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future wort…
Extracted from the employer’s posting. Read it in full on Thinking Machines’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Thinking Machines
- Product Manager - Post TrainingRemote
- Software Engineer, ProductRemote
- Site Reliability Engineer, Post TrainingSan Francisco
- Site Reliability Engineer, ProductionSan Francisco
- Software Engineer, Evaluation Platform / InfraSan Francisco
- Software Engineer, Research ToolsRemote