SRE

Baseten · Remote

Posted
103 days ago
Last confirmed live
1 day ago

What this role involves

Baseten is hiring a Site Reliability Engineer to define and codify day 2 operations for their ML infrastructure platform. The role involves building robust systems, automations, and observability tooling to ensure platform reliability at scale. The SRE will work closely with engineering, forward-deployed, and product teams to learn from failure patterns and raise operational standards.

Skills this posting asks for

  • kubernetes
  • eks
  • gke
  • victoriametrics
  • prometheus
  • loki
  • elk
  • grafana
  • terraform
  • helm
  • fluxcd
  • argocd
  • incident.io
  • observability-as-code
  • gitops
  • multi-cloud
  • sre
  • incident response
  • post-mortem
  • runbooks
  • automation
  • ml infrastructure

Requirements

  • Level: senior

From the employer’s posting

ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the front…

Read the full description on Baseten’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Baseten

All 14 roles at Baseten