Site Reliability Engineer - AI Accelerator Infrastructure - Contract

d-Matrix · Remote

Posted
40 days ago
Last confirmed live
1 day ago
Type
contractor

What this role involves

d-Matrix is hiring a Site Reliability Engineer for a 6-month contract with potential conversion to full-time. The role involves owning reliability, automation, and observability of infrastructure including colocation, on-premises GPU clusters, and cloud environments. Responsibilities include hands-on infrastructure work, IaC with Terraform/Ansible, monitoring with Prometheus/Grafana/DataDog, and supporting customer-facing platform services.

Skills this posting asks for

  • terraform
  • ansible
  • prometheus
  • grafana
  • datadog
  • aws
  • azure
  • gcp
  • infiniband
  • roce
  • high-speed ethernet
  • linux
  • networking
  • storage
  • ci/cd
  • infrastructure as code
  • configuration management
  • monitoring
  • alerting
  • incident response
  • capacity planning
  • hardware troubleshooting
  • server provisioning
  • os configuration

Requirements

  • Level: senior
  • Remote policy: onsite

From the employer’s posting

At d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration. We value humility and believ…

Read the full description on d-Matrix’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at d-Matrix

All 11 roles at d-Matrix