Software Engineer, Hardware Health

OpenAI · San Francisco

Posted
103 days ago
Last confirmed live
1 day ago

What this role involves

This role involves building infrastructure for monitoring and maintaining the health of OpenAI's global compute fleet. Responsibilities include defining health signals, building automated remediation systems, and debugging hardware failures at scale. The candidate will own node lifecycle workflows and partner with reliability and provider teams.

Skills this posting asks for

  • python
  • shell_scripting
  • sql
  • promql
  • distributed_systems
  • debugging
  • operational_tooling
  • linux
  • pcie
  • infiniband
  • roce
  • networking
  • gpu_clusters
  • observability
  • telemetry
  • automated_remediation
  • fleet_lifecycle_management
  • node_lifecycle
  • health_signals
  • health_checks

Requirements

  • 7 years of experience
  • Level: senior

From the employer’s posting

ABOUT THE TEAM The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated reme…

Read the full description on OpenAI’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at OpenAI

All 131 roles at OpenAI