Staff Software Engineer, AI Runtime

Databricks · Mountain View, California; San Francisco, California

Posted
5 days ago
Last confirmed live
Today

What this role involves

Databricks is hiring a Staff Software Engineer for AI Runtime to build and scale the managed GPU training platform. The role involves architecting distributed training systems, improving GPU efficiency, and leading engineering initiatives for large-scale AI training. The ideal candidate has 10+ years of distributed systems experience with deep knowledge of GPU training infrastructure and frameworks like PyTorch.

Skills this posting asks for

  • pytorch
  • fsdp
  • deepspeed
  • megatron
  • distributed systems
  • gpu training
  • high-performance computing
  • ml systems
  • checkpointing
  • distributed parallelism
  • gpu scheduling
  • fault tolerance
  • observability
  • api design
  • cli

Requirements

  • 10 years of experience
  • Level: staff

From the employer’s posting

P-1930At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infras…

Read the full description on Databricks’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Databricks

All 216 roles at Databricks