Software Engineer, Observability

OpenAI
San Francisco
Workplace: OnsiteFull timeUSD 255,000 - 405,000 annuallyFunction: Software EngineeringSkills: ["Communications","Problem-solving","Team collaboration","Systems design","Debugging"]

Join OpenAI's Applied AI Infrastructure team to build the observability platform for large-scale AI systems. You’ll own core observability infrastructure, contribute AI-native tooling, and help create dashboards and notebook-like debugging experiences. Collaborate with engineers, researchers, and operations across the company to ship reliable, scalable, observable production systems powering GPT-based products.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
6 months ago

Software Engineer, Observability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join OpenAI's Applied AI Infrastructure team to build the observability platform for large-scale AI systems. You’ll own core observability infrastructure, contribute AI-native tooling, and help create dashboards and notebook-like debugging experiences. Collaborate with engineers, researchers, and operations across the company to ship reliable, scalable, observable production systems powering GPT-based products.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Own core observability infrastructure, including distributed logging, time series, and trace storage.
  • •Build AI-native tools that help engineers detect, understand, and resolve issues autonomously.
  • •Contribute to UI experiences like dashboards, notebooking, or interactive debugging.
  • •Collaborate closely with engineers, researchers, user ops, and other teams across the company to build the next generation observability product.
  • •Spend time solving real production reliability problems to improve system performance at scale.

Pay and Benefits

Salary: USD 255,000 - 405,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Experience operating large-scale distributed systems in production, especially logging systems or time-series databases.
  • •Strong fundamentals in systems, networking, and cloud infrastructure (Kubernetes, AWS, etc.).
  • •Full-stack chops or product sensibilities and ability to build real tools used by engineers.
  • •Ability to thrive in ambiguous environments and solve unscoped problems.
  • •Familiarity with observability tooling (e.g., Prometheus, OpenTelemetry) is a plus.
Experience:AIObservabilityInfrastructureCloudDistributed systems
Skills:CommunicationsProblem-solvingTeam collaborationSystems designDebugging
Tech Stack:KubernetesAWSPrometheusOpenTelemetryDistributed loggingTime seriesTrace storageNotebook-like UIUI dashboards

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor