Infrastructure Engineer (GPU & Compute)

lightningai · London, England, United Kingdom; New York, New York, United States; San Francisco, California, United States; Seattle, Washington, United States

Posted
12 days ago
Last confirmed live
1 day ago

What this role involves

Lightning AI is seeking an Infrastructure Engineer to own image management, system diagnostics, and validation across large-scale bare-metal compute infrastructure with a focus on GPUs. The role involves developing automation, improving reliability, and enabling cluster bring-up for AI/ML and HPC workloads. The engineer will work with Linux, GPU diagnostics tools, and hardware management systems.

Skills this posting asks for

  • gpu diagnostics
  • nvidia dcgm
  • linux
  • python
  • ipmi
  • redfish
  • pxe
  • bare-metal provisioning
  • image management
  • system validation
  • hardware qualification
  • virtualization
  • automation
  • firmware validation
  • driver validation
  • performance analysis
  • hpc
  • ai/ml

Requirements

  • Remote policy: remote
  • Visa sponsorship: no

From the employer’s posting

Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction. …

Read the full description on lightningai’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at lightningai

All 16 roles at lightningai