Software Engineer, Inference Runtime
LM Studio · Remote
- Posted
- 15 days ago
- Last confirmed live
- Today
What this role involves
LM Studio is seeking an Inference Runtime Software Engineer to advance its inference stack for on-device and cloud AI. The role involves integrating new inference engines, optimizing model execution across CPU and GPU targets, and contributing to open-source projects. Candidates should have strong Python and C++ skills, deep understanding of transformer architectures, and experience with profiling and inference systems.
Skills this posting asks for
- python
- c++
- pytorch
- llama.cpp
- mlx
- executorch
- vllm
- sglang
- tensorrt-llm
- cuda
- metal
- vulkan
- rocm
- profiling
- debugging
- transformer architectures
- model inference
- distributed execution
- batching
- scheduling
- caching
Requirements
- Level: senior
- Remote policy: hybrid
From the employer’s posting
LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family. As a team, we work w…
Read the full description on LM Studio’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at LM Studio
- Software Engineer, Agent HarnessNew York City
- Software Engineer, ApplicationNew York City