SRE
Baseten · Remote
- Posted
- 103 days ago
- Last confirmed live
- 1 day ago
What this role involves
Baseten is hiring a Site Reliability Engineer to define and codify day 2 operations for their ML infrastructure platform. The role involves building robust systems, automations, and observability tooling to ensure platform reliability at scale. The SRE will work closely with engineering, forward-deployed, and product teams to learn from failure patterns and raise operational standards.
Skills this posting asks for
- kubernetes
- eks
- gke
- victoriametrics
- prometheus
- loki
- elk
- grafana
- terraform
- helm
- fluxcd
- argocd
- incident.io
- observability-as-code
- gitops
- multi-cloud
- sre
- incident response
- post-mortem
- runbooks
- automation
- ml infrastructure
Requirements
- Level: senior
From the employer’s posting
ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the front…
Read the full description on Baseten’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Baseten
- AI EngineerRemote
- Software Engineer - AI Developer ProductivityRemote
- Software Engineer - Testing FrameworksRemote
- Software Engineer - Continuous DeliveryRemote
- Software Engineer - ObservabilityRemote
- Product Manager, EnterpriseRemote