Infrastructure Site Reliability Engineer
nebius · United States
- Posted
- 3 days ago
- Last confirmed live
- 2 days ago
What this role involves
Nebius is seeking a Staff Network Site Reliability Engineer to build and operate the network infrastructure for its AI cloud platform. The role involves defining reliability goals, driving improvements, owning incident response, and building observability and automation. Required skills include production Linux, networking fundamentals, and software development in Go or Python.
Skills this posting asks for
- linux
- networking
- go
- python
- iac
- cicd
- container platforms
- high-availability systems
- incident response
- sli/slo
- observability
- alerting
- automation
- load balancers
- tunneling
- nat64
- ebpf
- xdp
- dpdk
- ftrace
- kernel networking
- testing labs
- staged rollouts
- automated verification
Requirements
- Level: staff
From the employer’s posting
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, w…
Read the full description on nebius’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at nebius
- IT Infrastructure Engineer – RMA & Hardware DiagnosticsMinnesota, United States
- AI/ML Specialist Solutions ArchitectRemote
- Frontend Engineer - User InterfaceRemote
- Partner Solutions ArchitectRemote
- Principal ML Solutions Architect - Token FactoryUnited States
- Senior Site Reliability Engineer (In-Office Required)New York City, New York, United States