Principal Observability Platform Engineer

nscaleoperationsukltd · US

Posted
4 days ago
Last confirmed live
2 days ago

What this role involves

This role involves owning the technical direction of Nscale's observability platform for GPU clusters and AI workloads. The candidate will define architecture, drive platform decisions, and mentor the team. Required experience includes 8+ years in SRE or platform engineering with deep knowledge of observability tools like Prometheus, Grafana, and Kubernetes.

Skills this posting asks for

  • prometheus
  • thanos
  • victoriametrics
  • grafana
  • loki
  • tempo
  • opentelemetry
  • clickhouse
  • elastic
  • python
  • go
  • kubernetes
  • terraform
  • ansible
  • slurm
  • infrastructure-as-code
  • observability
  • platform engineering
  • sre
  • alerting
  • metrics
  • logs
  • tracing
  • incident management

Requirements

  • 8 years of experience
  • Level: principal

From the employer’s posting

Principal Observability Platform Engineer – Nscale About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior result…

Read the full description on nscaleoperationsukltd’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at nscaleoperationsukltd

All 27 roles at nscaleoperationsukltd