Staff Site Reliability Engineer (Production Engineer)- Federal

Zscaler · Bellevue, Washington, USA; Boston, Massachusetts, USA; Crystal City, Virginia, USA; Dallas, Texas, USA; Denver, Colorado, USA; McLean, Virginia, USA; New York City, New York, USA; San Jose, California, USA; Short Hills, New Jersey, USA · $119k–$170k

The role

This is a staff-level Site Reliability Engineering role focused on the reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure. The role involves software-first SRE practices, including writing production-grade code, automation, and leading incident response. The position is hybrid, requiring three days onsite in San Jose, CA, and requires US Citizenship.

Pay
$119k–$170k
Location
Bellevue, Washington, USA; Boston, Massachusetts, USA; Crystal City, Virginia, USA; Dallas, Texas, USA; Denver, Colorado, USA; McLean, Virginia, USA; New York City, New York, USA; San Jose, California, USA; Short Hills, New Jersey, USA
Work mode
Hybrid
Employment
Full time
Level
Staff
Experience
5+ yrs
Sponsorship
Not offered

What we know that the posting doesn’t say

  • Seen todaystill listed on the employer’s careers page
  • Posted 7 days agothe first time we saw it

About Zscaler

Zscaler provides a cloud security platform that securely connects users, devices, and applications. Their Zero Trust Exchange platform is designed to protect customers from cyberattacks and data loss.

What you would do

  • Maintain high availability across large-scale bare-metal Linux/BSD fleets and Kubernetes clusters.
  • Lead full-cycle incident response with cross-stack troubleshooting using low-level OS and network tools.
  • Automate infrastructure lifecycle, provisioning, configuration, and releases using Ansible, Python, and Go.
  • Own end-to-end telemetry using Prometheus and OpenTelemetry, defining SLOs and error budgets.
  • Perform architectural reviews, kernel upgrades, capacity tuning, and strict CI/CD validation.
  • Embed operability standards like telemetry and rollback safety into service design.

Must have

  • US Citizenship required due to customer nature.
  • 5+ years in SRE, Production Engineering, or Systems Engineering.
  • Experience operating high-scale, low-latency production platforms.
  • Proven ability to write and debug code in Python, Go, or Bash.
  • Hands-on experience writing Ansible playbooks for automation.
  • Deep knowledge of Linux OS internals and kernel troubleshooting.
  • Understanding of networking protocols and packet-level analysis.
  • Experience with DNS resolution workflows and TLS handshakes.
  • Proficiency with TCP/IP mechanics and tcpdump packet captures.
  • Foundational understanding of AI/ML technologies and solutions.
  • Experience leveraging or securing AI-driven solutions in domain.
  • Ability to analyze storage discrepancies like df vs du.

Nice to have

  • Hands-on experience operating FreeBSD or BSD systems in production.
  • Proven expertise running and scaling Kubernetes in high-traffic environments.
  • Experience troubleshooting Kubernetes in low-latency environments.
  • Experience with workflow orchestration platforms like Temporal.
  • Deep experience with Prometheus or OpenTelemetry ecosystems.
  • Experience using AI/ML frameworks or AIOps for root-cause analysis.

What you get

  • Various health plans
  • Time off for vacation and sick leave
  • Parental leave options
  • Retirement options
  • Education reimbursement
  • In-office perks

Experience and education

  • 5+ years of relevant experience

Key skills

  • python
  • go
  • bash
  • ansible
  • linux
  • kubernetes
  • prometheus
  • opentelemetry
  • tcpdump
  • strace
  • lsof
  • iostat
  • vmstat
  • gdb
  • dns
  • tls
  • tcp/ip
  • freebsd
  • temporal
  • ai/ml
  • aiops
  • ci/cd
  • slo
  • error budgets

Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any locatio…

Extracted from the employer’s posting. Read it in full on Zscaler’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Zscaler

All 45 roles at Zscaler →