Staff Software Engineer, RL Environments

Scale AI
San Francisco, New York
Workplace: OnsiteFull timeUSD 252,000 - 315,000 annuallyFunction: Software EngineeringExperience: 8+ yearsSkills: ["Communication"]

Own the technical foundation for building, running, verifying, and delivering reinforcement learning (RL) environments at scale. Design the end-to-end platform—sandboxed execution, environment packaging/versioning, rollout orchestration, trajectory capture, and verifier frameworks—while also instrumenting real applications and building graders that hold up under adversarial optimization. This is a hands-on staff engineering role where you set technical direction and write the hard parts.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Scale AI
Scale AI
5 days ago

Staff Software Engineer, RL Environments

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Own the technical foundation for building, running, verifying, and delivering reinforcement learning (RL) environments at scale. Design the end-to-end platform—sandboxed execution, environment packaging/versioning, rollout orchestration, trajectory capture, and verifier frameworks—while also instrumenting real applications and building graders that hold up under adversarial optimization. This is a hands-on staff engineering role where you set technical direction and write the hard parts.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Own how Scale builds, runs, verifies, and delivers RL environments at scale.
  • •Design the RL environment platform including sandboxed execution, environment packaging/versioning, rollout orchestration, trajectory capture, verifier frameworks, and authoring surfaces.
  • •Instrument real applications and design task suites to expose capability gaps.
  • •Build graders that withstand adversarial optimization and help produce trustworthy reward signals.
  • •Set technical direction across multiple teams while still writing the hard parts.

Pay and Benefits

Salary: USD 252,000 - 315,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionRetirement BenefitsLearning BudgetPaid LeaveCommuter Benefits

Key Requirements

  • •8+ years of software engineering experience with strong fundamentals in distributed systems, system design, data structures, and algorithms.
  • •Strong Python skills and experience shipping production software; comfort with at least one other stack area (TypeScript/React, Go, Rust, or similar).
  • •Deep experience with containerization and sandboxed execution, including Docker, VMs, gVisor/Firecracker, Kubernetes, or equivalent.
  • •Experience building or operating high-throughput backend systems including orchestration, job scheduling, queuing, and large-scale data pipelines.
  • •Hands-on experience with LLMs and agent loops (tool calling, MCP, or eval harnesses) and intuition for what a training signal teaches.
Experience:8+ yearsReinforcement learningLLMsDistributed systems
Skills:Communication
Languages:English
Tech Stack:PythonTypeScriptReactGoRustDockerVMsGVisorFirecrackerKubernetesLLMsAgent loopsTool callingMCPEval harnessesRLHFRLAIFRLVRGRPOPPO

Company Brief

Scale AI
Provides data labeling, annotation, and infrastructure services to accelerate machine learning and AI development. Supplies high-quality training data, tooling, and APIs for customers in autonomous vehicles, mapping, robotics, and enterprise AI applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn