Research Engineer (Senior Staff+), Safeguards Labs

Anthropic
San Francisco, New York
Workplace: OnsiteFull timeUSD 350,000 - 850,000 annuallyFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Communication","Problem-solving","Collaboration","Independence","Rigor"]

Research-engineering role at Anthropic Safeguards Labs to define and execute safety-focused research, run end-to-end experiments, and prototype signals for real-time safeguards. You’ll work on detecting misuse of Claude, building classifiers, and transferring successful prototypes to production with a small, high-leverage team blending research and engineering.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Anthropic
Anthropic
3 months ago

Research Engineer (Senior Staff+), Safeguards Labs

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Research-engineering role at Anthropic Safeguards Labs to define and execute safety-focused research, run end-to-end experiments, and prototype signals for real-time safeguards. You’ll work on detecting misuse of Claude, building classifiers, and transferring successful prototypes to production with a small, high-leverage team blending research and engineering.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Lead and contribute to research projects investigating new methods for detecting misuse of Claude, identifying malicious organizations and accounts, strengthening model safeguards, and other safety needs.
  • •Design and run offline analyses over model usage data to surface abuse patterns, build classifiers and detection systems, and evaluate their effectiveness.
  • •Develop and iterate on prototypes that could eventually feed signals into the real-time safeguards path, partnering with engineers on tech transfer.
  • •Contribute to a broader research portfolio investigating methods for detecting abusive behavior in chat-based or agentive workflows, and for training the model to robustly refrain from dangerous responses or behaviors without over-refusing.
  • •Build evaluations and methodologies for measuring whether safeguards actually work, including in agentic settings.

Pay and Benefits

Salary: USD 350,000 - 850,000 annually
Equity and Bonus:Equity
Perks:Paid LeaveParental LeaveEquity Donation

Key Requirements

  • •Have a track record of independently driving research projects from ambiguous problem statements to concrete results, ideally in AI, ML, security, integrity, or a related technical field.
  • •Are comfortable scoping your own work and switching between research, engineering, and analysis as a project demands.
  • •Have working familiarity with how large language models operate — sampling, prompting, training — even if LLMs aren’t your primary background.
  • •Are proficient in Python and comfortable working with large datasets.
  • •Care about the societal impacts of AI and want your work to directly reduce real-world harm.
Experience:AIMLSecuritySafetyResearch
Education:Bachelor's
Skills:CommunicationProblem-solvingCollaborationIndependenceRigor
Languages:English
Tech Stack:PythonLLMsMachine learningData analysis

Company Brief

Anthropic
Develops large-scale AI systems and safety research to create reliable, steerable, and interpretable AI assistants and models for commercial and research applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn