Research, Safety

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Communication"]

Work on AI safety research to make models safe and trustworthy, exploring how training and data shape refusal and engagement with harmful or dual-use requests. Contribute across the stack—from pre-training data curation and safety-focused fine-tuning to evaluations and red-teaming—designing experiments, building safety evaluations (including long-horizon/agentic tasks), and developing mitigations based on failure modes you discover.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
1 day ago

Research, Safety

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Work on AI safety research to make models safe and trustworthy, exploring how training and data shape refusal and engagement with harmful or dual-use requests. Contribute across the stack—from pre-training data curation and safety-focused fine-tuning to evaluations and red-teaming—designing experiments, building safety evaluations (including long-horizon/agentic tasks), and developing mitigations based on failure modes you discover.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Conduct AI safety research focused on how models learn to handle harmful or dual-use requests.
  • •Work across the development stack, including safety-focused fine-tuning, evaluations, and red-teaming.
  • •Build data filtering pipelines and quality classifiers to shape what models learn and study downstream safety effects.
  • •Design and maintain safety evaluations, including long-horizon and agentic-task behavior measurement.
  • •Red-team models and products to surface failure modes, jailbreaks, and emergent risks, then design mitigations for what you find.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeaveParental LeaveRelocation

Key Requirements

  • •Bachelor’s degree (or equivalent) in Computer Science, Machine Learning, Physics, Mathematics, or a related field with strong theoretical and empirical grounding.
  • •Background in AI safety research with hands-on experience in at least one area such as RLHF/RLAIF, alignment and preference modeling, deliberative alignment, safety evaluations, or red-teaming.
  • •Proficiency in Python and familiarity with deep learning frameworks such as PyTorch, TensorFlow, or JAX, with comfort debugging distributed training and scaling code.
  • •Ability to clearly communicate by explaining complex technical concepts in writing.
  • •Experience with long-horizon, multi-step, or agentic-task evaluations, synthetic data generation, and/or modern red-teaming/jailbreaking techniques (preferred).
Experience:AI safetyMachine learningDeep learning
Education:Bachelor's
Skills:Communication
Tech Stack:PythonPyTorchTensorFlowJAX

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website