Site Reliability Engineer

runpod · Remote

Posted
43 days ago
Last confirmed live
Today

What this role involves

Runpod, an AI developer cloud, is hiring a Site Reliability Engineer for its Reliability team. The role focuses on ensuring platform stability, defining SLOs, improving observability, and automating operational workflows. The position is remote-first and involves cross-functional collaboration.

Skills this posting asks for

  • prometheus
  • grafana
  • python
  • go
  • bash
  • ci/cd
  • slo
  • sli
  • observability
  • monitoring
  • alerting
  • automation
  • incident response
  • postmortems
  • production readiness
  • distributed systems
  • gpu

Requirements

  • Remote policy: remote

From the employer’s posting

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed…

Read the full description on runpod’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at runpod

All 8 roles at runpod