Software Engineer, Data Infrastructure - Research

OpenAI
San Francisco
Workplace: OnsiteFull timeUSD 250,000 - 380,000 annuallyFunction: Software EngineeringSkills: ["Collaboration","Humility","Ownership","Communication"]

Design and build dataset infrastructure for OpenAI’s next-gen training stack. Create standardized dataset APIs, scale pipelines across thousands of GPUs, and proactively test performance bottlenecks, collaborating with multimodal researchers and infra teams to ensure unified, efficient, and easy-to-consume data for training and inference.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
11 months ago

Software Engineer, Data Infrastructure - Research

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 58 minutes agoStatus: Live

Job Summary

Design and build dataset infrastructure for OpenAI’s next-gen training stack. Create standardized dataset APIs, scale pipelines across thousands of GPUs, and proactively test performance bottlenecks, collaborating with multimodal researchers and infra teams to ensure unified, efficient, and easy-to-consume data for training and inference.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and maintain standardized dataset APIs for multimodal data that cannot fit in memory.
  • •Build proactive testing and scale validation pipelines for dataset loading at GPU scale.
  • •Collaborate with teammates to integrate datasets into training and inference pipelines with a great user experience.
  • •Document and maintain dataset interfaces so they are discoverable and easy for other teams to adopt.
  • •Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged.

Pay and Benefits

Salary: USD 250,000 - 380,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Design and maintain standardized dataset APIs, including for multimodal data that cannot fit in memory.
  • •Build proactive testing and scale validation pipelines for dataset loading at GPU scale.
  • •Collaborate with teammates to integrate datasets into training and inference pipelines with a great user experience.
  • •Document and maintain dataset interfaces to be discoverable and easy for other teams to adopt.
  • •Establish safeguards and validation systems to keep datasets reproducible and unchanged once standardized.
Experience:Distributed systemsData pipelinesInfrastructure
Skills:CollaborationHumilityOwnershipCommunication
Tech Stack:APIsDistributed systemsGPUMultimodal datasets

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor