Senior Principal AI Engineer

cerence · Remote

Posted
17 days ago
Last confirmed live
3 days ago

What this role involves

This role focuses on designing and operating distributed training systems for large neural networks across GPU clusters. Responsibilities include optimizing multi-node, multi-GPU execution, diagnosing bottlenecks, and improving training stability. The ideal candidate has deep experience with distributed systems, GPU orchestration tools, and training frameworks like PyTorch Distributed, Megatron-LM, and DeepSpeed.

Skills this posting asks for

  • slurm
  • kubernetes
  • ray
  • runai
  • nccl
  • rdma
  • infiniband
  • nvlink
  • pytorch distributed
  • megatron-lm
  • deepspeed
  • activation checkpointing
  • zero offload
  • distributed systems
  • ml systems
  • gpu clusters
  • data parallelism
  • tensor parallelism
  • pipeline parallelism
  • hpc

Requirements

  • Level: senior
  • Remote policy: remote

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at cerence

All 4 roles at cerence