Software Engineer, Infrastructure - Analytics Platform

OpenAI
San Francisco
Workplace: HybridFull timeUSD 230,000 - 385,000 annuallyFunction: Software EngineeringSkills: ["Rust","C++","Kubernetes","Distributed systems","Async","Profiling","Latency","Throughput","Memory behavior","I/O","Networking","Debugging"]

Staff-level software engineer to own production-critical infrastructure end to end for the Analytics Platform, focusing on backend/system engineering, low-level performance, and distributed services. Build and operate Rust or C++ backends at scale, design data/serving systems, debug bottlenecks, and ensure reliability through Kubernetes and on-call practices while collaborating with researchers to accelerate OpenAI's research workflows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
5 months ago

Software Engineer, Infrastructure - Analytics Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 55 minutes agoStatus: Live

Job Summary

Staff-level software engineer to own production-critical infrastructure end to end for the Analytics Platform, focusing on backend/system engineering, low-level performance, and distributed services. Build and operate Rust or C++ backends at scale, design data/serving systems, debug bottlenecks, and ensure reliability through Kubernetes and on-call practices while collaborating with researchers to accelerate OpenAI's research workflows.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Own critical infrastructure across design, implementation, rollout, operation, and iteration.
  • •Build and operate performant backend systems in Rust or C++ that support core research workflows.
  • •Design and improve distributed data and serving systems, including tradeoffs around partitioning, replication, consistency, retries, backpressure, and failure isolation.
  • •Debug real production bottlenecks across latency, throughput, contention, hot spots, and overload behavior.
  • •Operate business-critical services through on-call, incidents, postmortems, observability, rollout safety, and zero-downtime migrations.

Pay and Benefits

Salary: USD 230,000 - 385,000 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Strong systems experience in Rust or C++, especially in performance-sensitive infrastructure.
  • •Hands-on experience building performance-sensitive backend systems and operating production-critical systems.
  • •Experience designing, building, or operating distributed systems or distributed databases at meaningful scale.
  • •Comfort working below typical service abstractions including concurrency, async execution, memory behavior, serialization, I/O, networking, profiling, and failure analysis.
  • •Strong technical judgment, ownership, and ability to ship practical first versions and improve them through production feedback.
Experience:AnalyticsResearch
Skills:RustC++KubernetesDistributed systemsAsyncProfilingLatencyThroughputMemory behaviorI/ONetworkingDebugging
Languages:English
Tech Stack:RustC++KubernetesClickHouseDistributed databasesSQL

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor