Software Engineer, Hardware Health
OpenAI · San Francisco
- Posted
- 103 days ago
- Last confirmed live
- 1 day ago
What this role involves
This role involves building infrastructure for monitoring and maintaining the health of OpenAI's global compute fleet. Responsibilities include defining health signals, building automated remediation systems, and debugging hardware failures at scale. The candidate will own node lifecycle workflows and partner with reliability and provider teams.
Skills this posting asks for
- python
- shell_scripting
- sql
- promql
- distributed_systems
- debugging
- operational_tooling
- linux
- pcie
- infiniband
- roce
- networking
- gpu_clusters
- observability
- telemetry
- automated_remediation
- fleet_lifecycle_management
- node_lifecycle
- health_signals
- health_checks
Requirements
- 7 years of experience
- Level: senior
From the employer’s posting
ABOUT THE TEAM The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet. Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated reme…
Read the full description on OpenAI’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at OpenAI
- Product Marketing Manager, CybersecurityRemote
- Product Manager, YouthRemote
- Software Engineer, API SafetyRemote
- Machine Learning Engineer, Multimodal Perception and AuthenticationRemote
- Social Marketing Manager, Developers Remote
- Data Engineer, Monetization Data PlatformMountain View