Researcher, Agent Safety, Oversight and System Mitigations

OpenAI
San Francisco
Workplace: HybridFull timeUSD 380,000 - 500,000 annuallyFunction: Research & Scientific (R&D)Skills: ["Rigorous reasoning","Threat modeling","Experimentation","Evaluation design","Evidence-based iteration"]

Work on oversight and system-level mitigations that help increasingly capable AI agents act safely and autonomously in real environments. Build practical controls for agent actions, including sandboxing and permission boundaries, and partner with a Codex harness team to productionize them. Red-team agentic systems to measure prevention of data exfiltration and unsafe tool use, while improving the safety–productivity tradeoff through rigorous evaluations and experimentation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
2 days ago

Researcher, Agent Safety, Oversight and System Mitigations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Work on oversight and system-level mitigations that help increasingly capable AI agents act safely and autonomously in real environments. Build practical controls for agent actions, including sandboxing and permission boundaries, and partner with a Codex harness team to productionize them. Red-team agentic systems to measure prevention of data exfiltration and unsafe tool use, while improving the safety–productivity tradeoff through rigorous evaluations and experimentation.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Design, build, and evaluate system-level controls for agent actions, including how they fit into broader systems like sandboxing, process isolation, and permission boundaries.
  • •Partner with a Codex harness engineering team to productionize the AI controls.
  • •Red-team end-to-end agentic systems to measure whether controls prevent data exfiltration, unsafe tool use, and other harmful outcomes.
  • •Improve the safety–productivity tradeoff by measuring and reducing missed harmful actions, unnecessary blocks, approval burden, and latency.

Pay and Benefits

Salary: USD 380,000 - 500,000 annually
Equity and Bonus:Equity
Perks:Relocation

Key Requirements

  • •Safety- or security-minded approach to agent behavior, with the ability to reason about security boundaries.
  • •Strong systems or security instincts with concrete understanding of isolation boundaries, permissions, attack surfaces, and failure modes.
  • •Ability to turn ambiguous safety questions into concrete threat models, reproducible experiments, and practical mitigations.
  • •Experience building experimental infrastructure and designing evaluations that distinguish robust mitigations from brittle ones.
  • •Deep interest in frontier AI alignment, safety, and control.
Experience:AI safetyAgentic systemsSecurityFrontier AI alignmentThreat modeling
Skills:Rigorous reasoningThreat modelingExperimentationEvaluation designEvidence-based iteration

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor