Cyber Evaluations Engineer

Anthropic
San Francisco, Washington
Workplace: OnsiteFull timeUSD 300,000 - 405,000 annuallyFunction: CybersecurityEducation: bachelorsSkills: ["Communication","Collaboration","Analytical thinking"]

Design and run cyber-relevant capability and safety evaluations to measure safeguards robustness in new models. Execute per-release testing ahead of major launches, analyze results on jailbreaks and prompt bypasses, and communicate findings to engineering and policy stakeholders. Build and tune detection probes, help translate policy into layered abuse-detection architecture, and develop internal tooling to run and score evaluations over time.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
3 days ago

Cyber Evaluations Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Design and run cyber-relevant capability and safety evaluations to measure safeguards robustness in new models. Execute per-release testing ahead of major launches, analyze results on jailbreaks and prompt bypasses, and communicate findings to engineering and policy stakeholders. Build and tune detection probes, help translate policy into layered abuse-detection architecture, and develop internal tooling to run and score evaluations over time.
Location: San Francisco, Washington
Workplace: Onsite
Employment Type: Full time
Job Function: Cybersecurity

Key Responsibilities

  • •Design and run capability, uplift, and safety evaluations to assess cyber-relevant risk in new models.
  • •Execute per-release safeguard-robustness testing ahead of major model launches.
  • •Analyze evaluation results and clearly communicate findings to the team and stakeholders.
  • •Design, prototype, and tune detection probes for cyber misuse.
  • •Build and maintain internal tooling to run and score evaluations, collaborating with policy and engineering to improve safeguards.

Pay and Benefits

Salary: USD 300,000 - 405,000 annually
Perks:Paid LeaveParental Leave

Key Requirements

  • •Experience building or running evaluations, benchmarks, or test suites for software or ML systems and delivering results on short timelines.
  • •Hands-on cybersecurity experience (e.g., CTF participation, vulnerability research, exploit development, or security research).
  • •Proficiency in Python.
  • •Strong ability to communicate evaluation results to cross-functional stakeholders or policy stakeholders.
  • •Minimum education: Bachelor’s degree or equivalent combination of education, training, and/or experience.
Experience:CybersecuritySoftware securityML evaluation
Education:Bachelor's
Skills:CommunicationCollaborationAnalytical thinking
Languages:English
Tech Stack:PythonSigmaYARASuricataSIEM

Eligibility

Security Clearance:Secret

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn