AI Platform Engineer, Training and Inference

Saviynt · Milpitas, California

Posted
97 days ago
Last confirmed live
1 day ago

What this role involves

Owns the Ray ecosystem for distributed training on H100 clusters and builds the LLM inference mesh with vLLM, SGLang, and NVIDIA Triton. Operates the full model promotion lifecycle from shadow mode to canary rollout. Integrates RAG retrieval and builds RL training infrastructure using Flyte and Ray RLlib.

Skills this posting asks for

  • ray
  • kubegray
  • gke
  • distributed training
  • ddp
  • nccl
  • vllm
  • sglang
  • nvidia triton
  • pytorch
  • mlflow
  • flyte
  • python
  • tensorrt
  • onnx
  • ray train
  • ray serve
  • ppo
  • rl
  • rag
  • pgvector
  • qdrant
  • continuous batching
  • checkpointing

From the employer’s posting

AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and red…

Read the full description on Saviynt’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Saviynt

All 6 roles at Saviynt