AI Platform Engineer, Training and Inference
Saviynt · Milpitas, California
- Posted
- 97 days ago
- Last confirmed live
- 1 day ago
What this role involves
Owns the Ray ecosystem for distributed training on H100 clusters and builds the LLM inference mesh with vLLM, SGLang, and NVIDIA Triton. Operates the full model promotion lifecycle from shadow mode to canary rollout. Integrates RAG retrieval and builds RL training infrastructure using Flyte and Ray RLlib.
Skills this posting asks for
- ray
- kubegray
- gke
- distributed training
- ddp
- nccl
- vllm
- sglang
- nvidia triton
- pytorch
- mlflow
- flyte
- python
- tensorrt
- onnx
- ray train
- ray serve
- ppo
- rl
- rag
- pgvector
- qdrant
- continuous batching
- checkpointing
From the employer’s posting
AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and red…
Read the full description on Saviynt’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Saviynt
- Distinguished Software Engineer, Data PlatformMilpitas, California
- Principal Site Reliability Engineer, Google CloudAtlanta
- Senior Principal Software Engineer, AI OnboardingMilpitas, California
- Associate Principal Software Engineer, AI OnboardingSan Francisco
- Principal Software Engineer, AI OnboardingMilpitas, California