Senior Site Reliability Engineer - Fleet

Lambda · Remote

Posted
8 days ago
Last confirmed live
1 day ago

What this role involves

Lambda is hiring a Senior Site Reliability Engineer to build and operate monitoring, automate cluster lifecycle, and troubleshoot large-scale HPC clusters for AI workloads. The role requires 7+ years of experience and presence in San Francisco or Bellevue office 4 days per week. Responsibilities include on-call rotations, incident response, and collaboration with engineering teams.

Skills this posting asks for

  • site reliability engineering
  • hpc engineering
  • devops
  • monitoring
  • alerting
  • prometheus
  • grafana
  • clickhouse
  • automation
  • configuration management
  • ansible
  • terraform
  • linux
  • infiniband
  • roce
  • clos fabrics
  • 100gbe
  • ethernet
  • switching
  • gpu-direct
  • nccl
  • python
  • go
  • pytorch

Requirements

  • 7 years of experience
  • Level: senior
  • Remote policy: onsite

From the employer’s posting

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintellige…

Read the full description on Lambda’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Lambda

All 25 roles at Lambda