Site Reliability Engineer
runpod · Remote
- Posted
- 43 days ago
- Last confirmed live
- Today
What this role involves
Runpod, an AI developer cloud, is hiring a Site Reliability Engineer for its Reliability team. The role focuses on ensuring platform stability, defining SLOs, improving observability, and automating operational workflows. The position is remote-first and involves cross-functional collaboration.
Skills this posting asks for
- prometheus
- grafana
- python
- go
- bash
- ci/cd
- slo
- sli
- observability
- monitoring
- alerting
- automation
- incident response
- postmortems
- production readiness
- distributed systems
- gpu
Requirements
- Remote policy: remote
From the employer’s posting
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed…
Read the full description on runpod’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at runpod
- Senior Data Engineer Remote
- Senior Developer Advocate Remote
- Developer Relations Engineer Remote
- Software Engineer (Full-Stack) Remote
- Senior Product Manager Remote
- Director of Infrastructure EngineeringRemote