Model Policy Manager, Agentic Safety

OpenAI
San Francisco
Workplace: HybridFull timeUSD 207,000 - 335,000 annuallyFunction: Government Affairs & Public PolicySkills: ["Communication","Collaboration","Analytical thinking"]

Shape how OpenAI identifies and mitigates real-world risks from model misalignment as models become more autonomous. Investigate harmful behavior across long trajectories and translate findings into behavioral policies, evaluations, monitoring, and safeguards. Partner with research, engineering, security, and product teams to balance safety, utility, and business risk, and inform deployment decisions, system cards, and safeguards reports.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
2 days ago

Model Policy Manager, Agentic Safety

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Shape how OpenAI identifies and mitigates real-world risks from model misalignment as models become more autonomous. Investigate harmful behavior across long trajectories and translate findings into behavioral policies, evaluations, monitoring, and safeguards. Partner with research, engineering, security, and product teams to balance safety, utility, and business risk, and inform deployment decisions, system cards, and safeguards reports.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Government Affairs & Public Policy
Seniority: Manager level

Key Responsibilities

  • •Identify vulnerabilities that emerge as models interact with tools, data, and external systems, and translate them into model- and system-level safeguards.
  • •Develop threat models and empirical frameworks for harmful outcomes from misaligned behavior, including building frameworks for understanding misalignment-driven harms.
  • •Identify underlying behaviors and system conditions driving harmful outcomes, and turn findings into policy frameworks, evaluation criteria, online measurement, and safeguards.
  • •Develop human data campaigns and gold sets to ground measurement and evaluation of emerging behaviors and risks.
  • •Partner across research, engineering, security, and product teams, inform deployment decisions and safeguards reports, and build monitoring approaches to detect regressions and emerging risks.

Pay and Benefits

Salary: USD 207,000 - 335,000 annually
Equity and Bonus:Equity
Perks:Relocation

Key Requirements

  • •Strong background in AI agent safety, privacy, security, cybersecurity, or adjacent fields with an adversarial mindset to investigate harmful outcomes.
  • •Demonstrated interest in AI alignment and a strong understanding of technical drivers of misaligned model behavior.
  • •Technical fluency to work directly with evaluation and training data, interpret results, and identify limitations, patterns, and investigation opportunities.
  • •Comfort working hands-on with model data and evaluation results, including inspecting examples and analyzing failure patterns and data quality.
  • •Can use empirical evidence to develop and refine safety policies and safeguards, translating ambiguous risks into measurable evaluation criteria.
Experience:AI agent safetyAI alignmentCybersecurityPrivacyModel evaluation
Skills:CommunicationCollaborationAnalytical thinking

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor