Research Engineer, Safety

Decagon
San Francisco, New York
Workplace: OnsiteFull timeUSD 200,000 - 400,000 annuallyFunction: Research & Scientific (R&D)Experience: 2+ yearsSkills: ["Experimental judgment","Risk tradeoffs","Ownership","Handling ambiguity"]

Build and evaluate safety mechanisms for conversational AI agents end-to-end, from spotting real-world failure modes to shipping safeguards in production. You’ll design adversarial evaluations and red-team datasets, develop classifiers/judges and post-training or runtime safeguards, and analyze production traces to prevent prompt injection, unsafe tool use, sensitive-data disclosure, policy violations, and hallucinated commitments. Work cross-functionally with Security, Product, Infrastructure, Legal, and enterprise stakeholders.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Decagon
Decagon
1 day ago

Research Engineer, Safety

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Build and evaluate safety mechanisms for conversational AI agents end-to-end, from spotting real-world failure modes to shipping safeguards in production. You’ll design adversarial evaluations and red-team datasets, develop classifiers/judges and post-training or runtime safeguards, and analyze production traces to prevent prompt injection, unsafe tool use, sensitive-data disclosure, policy violations, and hallucinated commitments. Work cross-functionally with Security, Product, Infrastructure, Legal, and enterprise stakeholders.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Research and build safeguards against prompt injection, unsafe tool use, sensitive-data disclosure, policy violations, and hallucinated commitments.
  • •Build adversarial evaluations, simulations, red-team datasets, and regression suites informed by production failures.
  • •Develop and deploy classifiers, judges, reward signals, post-training methods, and runtime safeguards for safer agent behavior.
  • •Analyze production traces and incidents to identify root causes, test mitigations, and measure their impact.
  • •Partner with Security, Product, Infrastructure, Legal, and customer-facing teams to translate enterprise requirements into scalable safeguards and rollout practices.

Pay and Benefits

Salary: USD 200,000 - 400,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionLife InsuranceDisabilityRetirement401kParental LeaveMonthly Stipend

Key Requirements

  • •2+ years of experience in AI/ML engineering, research, or AI safety.
  • •Hands-on experience evaluating, post-training, or deploying language models or agentic systems.
  • •Experience with modern post-training techniques such as reinforcement learning, preference optimization, distillation, model routing, and synthetic-data generation.
  • •Experience with adversarial testing, model red teaming, prompt injection, policy enforcement, privacy, or safe tool use.
  • •Fluency in Python and modern ML tooling, with strong experimental judgment to ship production systems.
Experience:2+ yearsAI/MLAI safetyLanguage modelsAgentic systems
Skills:Experimental judgmentRisk tradeoffsOwnershipHandling ambiguity
Tech Stack:PythonReinforcement learningPreference optimizationDistillationModel routingSynthetic-data generationPrompt injectionAdversarial testingRed teaming

Company Brief

Decagon
Builds conversational AI agents and a platform that automates customer support across chat, email, and voice, enabling brands to deliver concierge-level customer experiences at scale.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2023
Glassdoor
Glassdoor: 3.9
WebsiteLinkedInGlassdoor