MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling)
Weekday AI · United States; Canada; United Kingdom
The role
This role is for a client of Weekday AI, a leading AI lab, and involves building foundational AI models. The MLOps Engineer will design tasks and solutions across GPU kernels, performance profiling, debugging, and inference serving to generate training data. The position requires hands-on experience in ML systems and infrastructure, with a focus on writing and evaluating technical content.
- Location
- United States; Canada; United Kingdom
- Level
- Mid
- Experience
- 2+ yrs
What we know that the posting doesn’t say
- Seen todaystill listed on the employer’s careers page
- Posted 2 days agothe first time we saw it
About Weekday AI
The role is for a client of Weekday AI, described as a leading AI lab with a cutting-edge GenAI team.
What you would do
- Design challenging tasks across GPU kernels, profiling, debugging, and serving.
- Write accurate, well-structured solutions for MLOps and ML systems tasks.
- Guide research and engineering teams to close knowledge gaps.
- Evaluate MLOps tasks and provide clear, written technical feedback.
- Develop guidelines and rubrics for kernel optimization and profiling.
- Collaborate with subject matter experts to ensure data consistency.
Must have
- 2+ years professional experience in ML systems or infrastructure.
- Hands-on experience in at least one of the four specified areas.
- Production experience with JAX and/or PyTorch.
- Familiarity with modern accelerators like A100, H100, B200, or TPU.
- Ability to reason about throughput, latency, and memory trade-offs.
- Demonstrable career progression.
- Ability to engage reliably for at least 40 hours/week on weekdays.
- Strong written communication skills.
- Hands-on systems role, not applied modeling or data science.
- Experience writing or optimizing custom GPU kernels.
- Experience with performance profiling and trace analysis.
- Experience debugging distributed or accelerator-bound workloads.
Nice to have
- Experience in more than one of the four specified areas.
- Framework-level depth with custom operators or distributed training.
- Experience with compiler and graph-level work.
- Experience with vLLM, SGLang, TensorRT-LLM, or Ray Serve.
- Experience with KV cache, paged attention, or continuous batching.
- Experience with Kineto, torch.profiler, Nsight, or XLA profiler.
- Experience with FSDP, DDP, DeepSpeed, or Megatron.
Experience and education
- 2+ years of relevant experience
Key skills
- cuda
- triton
- pallas
- kineto
- torch.profiler
- nsight
- xla
- jax
- vllm
- sglang
- tensorrt-llm
- ray serve
- kv cache
- paged attention
- continuous batching
- pytorch
- fsdp
- ddp
- deepspeed
- megatron
- a100
- h100
- b200
- tpu
This role is for one of our clients Compensation: $90-$120 per hour Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking MLOps Engineers with hands-on experience in large language model infrastructure across any of four areas: GPU k…
Extracted from the employer’s posting. Read it in full on Weekday AI’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Weekday AI
- Data Scientist Talent NetworkRemote
- Software Engineer, Go - Codebase Q&ARemote
- Software Engineer, C - Codebase Q&ARemote
- Software Engineer, Python - Codebase Q&ARemote
- Incident management / reliability / SRE EvaluatorRemote
- Cloud / DevOps Engineer (Infra & IaC)Remote