Software Engineer, Inference Runtime

LM Studio · Remote

Posted
15 days ago
Last confirmed live
Today

What this role involves

LM Studio is seeking an Inference Runtime Software Engineer to advance its inference stack for on-device and cloud AI. The role involves integrating new inference engines, optimizing model execution across CPU and GPU targets, and contributing to open-source projects. Candidates should have strong Python and C++ skills, deep understanding of transformer architectures, and experience with profiling and inference systems.

Skills this posting asks for

  • python
  • c++
  • pytorch
  • llama.cpp
  • mlx
  • executorch
  • vllm
  • sglang
  • tensorrt-llm
  • cuda
  • metal
  • vulkan
  • rocm
  • profiling
  • debugging
  • transformer architectures
  • model inference
  • distributed execution
  • batching
  • scheduling
  • caching

Requirements

  • Level: senior
  • Remote policy: hybrid

From the employer’s posting

LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family. As a team, we work w…

Read the full description on LM Studio’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at LM Studio

All 3 roles at LM Studio