Site Reliability Engineer - AI Accelerator Infrastructure - Contract
d-Matrix · Remote
- Posted
- 40 days ago
- Last confirmed live
- 1 day ago
- Type
- contractor
What this role involves
d-Matrix is hiring a Site Reliability Engineer for a 6-month contract with potential conversion to full-time. The role involves owning reliability, automation, and observability of infrastructure including colocation, on-premises GPU clusters, and cloud environments. Responsibilities include hands-on infrastructure work, IaC with Terraform/Ansible, monitoring with Prometheus/Grafana/DataDog, and supporting customer-facing platform services.
Skills this posting asks for
- terraform
- ansible
- prometheus
- grafana
- datadog
- aws
- azure
- gcp
- infiniband
- roce
- high-speed ethernet
- linux
- networking
- storage
- ci/cd
- infrastructure as code
- configuration management
- monitoring
- alerting
- incident response
- capacity planning
- hardware troubleshooting
- server provisioning
- os configuration
Requirements
- Level: senior
- Remote policy: onsite
From the employer’s posting
At d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration. We value humility and believ…
Read the full description on d-Matrix’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at d-Matrix
- Technical Program Manager - Hardware/Systems, PrincipalRemote
- Principal Software Engineer, SDK & Lowering StackRemote
- Staff Software Engineer, SIMD KernelsRemote
- Principal System Software Engineer, AI Inference ExecutionRemote
- Principal Software Engineer in Test (SDET)Remote
- Senior Staff Software Engineer, Developer and Qualification ToolsRemote