Principal SRE - AI Inference
Cerebras Systems · Headquarters/Sunnyvale Office
- Posted
- 47 days ago
- Last confirmed live
- 1 day ago
What this role involves
Cerebras Systems is hiring a Principal SRE for AI Inference to architect and scale reliability for their high-performance inference infrastructure. The role involves defining technical architecture, building self-service platforms, and driving reliability practices across multi-datacenter environments. Requires 15+ years of experience in SRE or platform engineering at large scale.
Skills this posting asks for
- sre
- infrastructure engineering
- platform engineering
- control planes
- schedulers
- orchestration systems
- capacity management
- reliability automation
- slo
- sli
- error budgets
- blameless postmortems
- chaos testing
- capacity forecasting
- self-service platforms
- internal tooling
- observability
- automation
- mentoring
- incident management
- mtbf
- mttr
Requirements
- 15 years of experience
- Level: principal
From the employer’s posting
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transf…
Read the full description on Cerebras Systems’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Cerebras Systems
- Cluster Operations Software EngineerRemote
- Hardware Analytics EngineerRemote
- Senior Quality Assurance EngineerRemote
- AI Inference Core - Senior SW Engineer for Platform & DevOpsRemote
- Senior Front End Design Engineer (Microarchitecture) - Sunnyvale HQHeadquarters/Sunnyvale Office
- Software Engineer, Cluster Deployment Headquarters/Sunnyvale Office