Staff+ Software Engineer, Safeguards Evals

Anthropic · San Francisco, CA | New York City, NY

Posted
2 days ago
Last confirmed live
Today

What this role involves

This role involves building evaluation infrastructure for an agentic investigation system that monitors misuse of Claude. Responsibilities include constructing high-quality eval datasets, measuring agent performance, and productionizing evaluations into release pipelines. The work sits at the intersection of applied ML research and engineering, focusing on trust and safety.

Skills this posting asks for

  • python
  • data pipelines
  • llms
  • data analysis
  • research prototyping
  • production code
  • agent evaluation frameworks
  • benchmarks
  • automated grading systems
  • trust and safety
  • content moderation
  • abuse detection
  • red teaming
  • adversarial testing
  • jailbreak research
  • synthetic data generation
  • data augmentation
  • distributed systems
  • large-scale data processing
  • reinforcement learning

Requirements

  • Level: staff

From the employer’s posting

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, en…

Read the full description on Anthropic’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Anthropic

All 142 roles at Anthropic