ML/Research Engineer, Safeguards

Anthropic
San Francisco
Workplace: OnsiteFull timeUSD 350,000 - 500,000 annuallyFunction: Research & Scientific (R&D)Experience: 4+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","Problem-solving"]

Join the Safeguards ML team to build scalable systems that detect and mitigate AI misuse, develop defenses for agentic risks, and monitor harms across contexts. Work across the research-to-deployment pipeline, contribute to red-teaming and adversarial robustness efforts, and help ensure model safety and user wellbeing through scalable classifiers and evaluation environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
10 months ago

ML/Research Engineer, Safeguards

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Join the Safeguards ML team to build scalable systems that detect and mitigate AI misuse, develop defenses for agentic risks, and monitor harms across contexts. Work across the research-to-deployment pipeline, contribute to red-teaming and adversarial robustness efforts, and help ensure model safety and user wellbeing through scalable classifiers and evaluation environments.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Develop classifiers to detect misuse and anomalous behavior at scale, including synthetic data pipelines for training classifiers and methods to automatically source representative evaluations
  • •Build systems to monitor harms that span multiple exchanges and coordinate signals across contexts
  • •Evaluate and improve the safety of agentic products by developing threat models and testing environments, and deploying mitigations for prompt injections
  • •Conduct research on automated red-teaming, adversarial robustness, and other research that helps test for or find misuse
  • •Collaborate across teams to translate research findings into production-ready ML systems and safeguards

Pay and Benefits

Salary: USD 350,000 - 500,000 annually

Key Requirements

  • •Have 4+ years of experience in ML engineering, research engineering, or applied research, in academia or industry
  • •Have proficiency in Python and experience building ML systems
  • •Are comfortable working across the research-to-deployment pipeline, from exploratory experiments to production systems
  • •Are worried about misuse risks of AI systems, and want to work to mitigate them
  • •Have strong communication skills and ability to explain complex technical concepts to non-technical stakeholders
Experience:4+ yearsAI safetyMachine learningResearch
Education:Bachelor's
Skills:CommunicationCollaborationProblem-solving
Languages:English
Tech Stack:PythonTransformersMachine learning

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn