Staff Site Reliability Engineer (Production Engineer)- Federal
Zscaler · Bellevue, Washington, USA; Boston, Massachusetts, USA; Crystal City, Virginia, USA; Dallas, Texas, USA; Denver, Colorado, USA; McLean, Virginia, USA; New York City, New York, USA; San Jose, California, USA; Short Hills, New Jersey, USA · $119k–$170k
The role
This is a staff-level Site Reliability Engineering role focused on the reliability and performance of Zscaler's high-throughput bare-metal and cloud infrastructure. The role involves software-first SRE practices, including writing production-grade code, automation, and leading incident response. The position is hybrid, requiring three days onsite in San Jose, CA, and requires US Citizenship.
- Pay
- $119k–$170k
- Location
- Bellevue, Washington, USA; Boston, Massachusetts, USA; Crystal City, Virginia, USA; Dallas, Texas, USA; Denver, Colorado, USA; McLean, Virginia, USA; New York City, New York, USA; San Jose, California, USA; Short Hills, New Jersey, USA
- Work mode
- Hybrid
- Employment
- Full time
- Level
- Staff
- Experience
- 5+ yrs
- Sponsorship
- Not offered
What we know that the posting doesn’t say
- Seen todaystill listed on the employer’s careers page
- Posted 7 days agothe first time we saw it
About Zscaler
Zscaler provides a cloud security platform that securely connects users, devices, and applications. Their Zero Trust Exchange platform is designed to protect customers from cyberattacks and data loss.
What you would do
- Maintain high availability across large-scale bare-metal Linux/BSD fleets and Kubernetes clusters.
- Lead full-cycle incident response with cross-stack troubleshooting using low-level OS and network tools.
- Automate infrastructure lifecycle, provisioning, configuration, and releases using Ansible, Python, and Go.
- Own end-to-end telemetry using Prometheus and OpenTelemetry, defining SLOs and error budgets.
- Perform architectural reviews, kernel upgrades, capacity tuning, and strict CI/CD validation.
- Embed operability standards like telemetry and rollback safety into service design.
Must have
- US Citizenship required due to customer nature.
- 5+ years in SRE, Production Engineering, or Systems Engineering.
- Experience operating high-scale, low-latency production platforms.
- Proven ability to write and debug code in Python, Go, or Bash.
- Hands-on experience writing Ansible playbooks for automation.
- Deep knowledge of Linux OS internals and kernel troubleshooting.
- Understanding of networking protocols and packet-level analysis.
- Experience with DNS resolution workflows and TLS handshakes.
- Proficiency with TCP/IP mechanics and tcpdump packet captures.
- Foundational understanding of AI/ML technologies and solutions.
- Experience leveraging or securing AI-driven solutions in domain.
- Ability to analyze storage discrepancies like df vs du.
Nice to have
- Hands-on experience operating FreeBSD or BSD systems in production.
- Proven expertise running and scaling Kubernetes in high-traffic environments.
- Experience troubleshooting Kubernetes in low-latency environments.
- Experience with workflow orchestration platforms like Temporal.
- Deep experience with Prometheus or OpenTelemetry ecosystems.
- Experience using AI/ML frameworks or AIOps for root-cause analysis.
What you get
- Various health plans
- Time off for vacation and sick leave
- Parental leave options
- Retirement options
- Education reimbursement
- In-office perks
Experience and education
- 5+ years of relevant experience
Key skills
- python
- go
- bash
- ansible
- linux
- kubernetes
- prometheus
- opentelemetry
- tcpdump
- strace
- lsof
- iostat
- vmstat
- gdb
- dns
- tls
- tcp/ip
- freebsd
- temporal
- ai/ml
- aiops
- ci/cd
- slo
- error budgets
Zscaler (NASDAQ: ZS) accelerates digital transformation so customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange™️ platform protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any locatio…
Extracted from the employer’s posting. Read it in full on Zscaler’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at Zscaler
- Senior Staff Rust DeveloperBellevue, Washington, USA; San Jose, California, USA
- Staff Software Engineer, Go (Service Platform & Orchestration)San Jose, California, USA
- Staff Software Development Engineer (Microservices)San Jose, California, USA
- Staff Software Engineer - React.jsSan Jose, California, USA
- Sr. Staff Software Development Engineer-AI Security (Network, Security, LLM)Bellevue, Washington, USA; San Jose, California, USA
- Senior Staff Software Engineer (C / Networking / Dataplane)San Jose, California, USA