Senior Software Engineer, Observability

togetherai · San Francisco · $200k–$280k

Posted
43 days ago
Last confirmed live
2 days ago
Published range
$200k–$280k

What this role involves

Together AI is building an AI Acceleration Cloud. The observability team designs and implements scalable observability platforms using tools like Prometheus, Grafana, and OpenTelemetry. The role involves developing automated monitoring, alerting, and anomaly detection systems, and collaborating on distributed tracing and incident response.

Skills this posting asks for

  • prometheus
  • grafana
  • clickhouse
  • clickstack
  • opentelemetry
  • aws
  • gcp
  • azure
  • go
  • python
  • terraform
  • ansible
  • helm
  • docker
  • kubernetes
  • postgresql
  • mongodb
  • redis
  • ai/ml infrastructure
  • gpu clusters
  • chaos engineering
  • open-source observability
  • security monitoring
  • compliance frameworks

Requirements

  • Level: senior
  • Remote policy: remote

From the employer’s posting

About the Role Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. The AI Infrastructure team at Together AI is at…

Read the full description on togetherai’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at togetherai

All 19 roles at togetherai