Principal Site Reliability Engineer, Machine Learning

Cambridge Mobile Telematics · Cambridge, MA

Posted
19 days ago
Last confirmed live
1 day ago

What this role involves

Cambridge Mobile Telematics is seeking a Principal Site Reliability Engineer for Machine Learning to ensure the operational health of Ray clusters on AWS EKS and Databricks workloads. The role involves owning SLOs, maintaining observability, and automating infrastructure using Terraform and CI/CD. Candidates need 7+ years of SRE experience and expertise in AWS, Kubernetes, and Python.

Skills this posting asks for

  • aws
  • ec2
  • ecs
  • eks
  • sqs
  • lambda
  • dynamodb
  • rds
  • aurora
  • s3
  • iam
  • kubernetes
  • docker
  • terraform
  • python
  • cloudwatch
  • datadog
  • ray
  • databricks
  • linux
  • ci/cd

Requirements

  • 7 years of experience
  • Level: principal

From the employer’s posting

CMT is looking for a Principal Site Reliability Engineer I, Machine Learning  to help us change the world. CMT has helped protect over 65 million drivers and prevent over 126,000 crashes worldwide. We build AI to solve some of the most difficult challenges in mobility — und…

Read the full description on Cambridge Mobile Telematics’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Cambridge Mobile Telematics

All 10 roles at Cambridge Mobile Telematics