Senior Site Reliability Engineer -AI Infrastructure Operations
nscaleoperationsukltd · Houston; San Francisco; Seattle · $170k–$265k
- Posted
- 4 days ago
- Last confirmed live
- 2 days ago
- Published range
- $170k–$265k
What this role involves
Nscale is hiring a Senior Site Reliability Engineer to own reliability for critical production services in AI infrastructure. The role involves setting reliability standards, mentoring other SREs, and building automation to reduce toil. Requires 6-10 years of SRE experience with strong software engineering skills and deep knowledge of Linux, networking, and Kubernetes.
Skills this posting asks for
- python
- go
- linux
- networking
- distributed systems
- kubernetes
- virtualized environments
- bare-metal environments
- ai workloads
- gpu workloads
- high-performance computing
- slo
- observability
- alerting
- incident process
- on-call
- infiniband
- rdma
Requirements
- 6 years of experience
- Level: senior
From the employer’s posting
About NscaleNscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-nativestartups and global enterprises, from bare metal up through the platform services teams actually buildon. Our culture runs on ownership, accountabili…
Read the full description on nscaleoperationsukltd’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at nscaleoperationsukltd
- Principal Technical Program Manager (TPM) - AI Infrastructure OperationsHouston; New York; San Francisco; Seattle
- Senior Technical Product Manager, Observability New York
- Principal Front-End Network EngineerHouston; New York; San Francisco; Seattle
- Senior Cloud Native Platform Engineer New York
- Physical Security Engineer (US)Houston; New York; San Francisco; Seattle
- Senior Network EngineerUS