Senior Support Engineer - Toronto

OpenAI
Toronto
Workplace: RemoteFull timeFunction: Solutions Engineering & Sales EngineeringEducation: bachelorsSkills: ["Communication","Cross-functional collaboration","Troubleshooting","Incident coordination"]

Collaborate with strategic enterprise accounts and product teams to troubleshoot complex issues on the OpenAI API platform. You’ll design and run operational processes for monitoring top customers and a 24x7 response team, partner with Infrastructure and Engineering on reliability readiness, and lead incident response with incident RCAs and post-mortems. Use automation and AI, advanced observability (metrics, logging, tracing, SLIs/SLOs), and scripting (e.g., Python) to improve support workflows and dashboards.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
3 days ago

Senior Support Engineer - Toronto

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Collaborate with strategic enterprise accounts and product teams to troubleshoot complex issues on the OpenAI API platform. You’ll design and run operational processes for monitoring top customers and a 24x7 response team, partner with Infrastructure and Engineering on reliability readiness, and lead incident response with incident RCAs and post-mortems. Use automation and AI, advanced observability (metrics, logging, tracing, SLIs/SLOs), and scripting (e.g., Python) to improve support workflows and dashboards.
Location: Toronto
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Mid level

Key Responsibilities

  • •Act as a foremost technical troubleshooting expert for the OpenAI API platform, serving as the last line of defense before core Engineering.
  • •Proactively scale support operations using automation and AI advancements to improve response and detection of customer-impacting issues.
  • •Configure and use advanced monitoring and alerting workflows to detect issues in real time.
  • •Partner with engineering on reliability reviews and operational readiness for new features, launches, and customer requirement updates.
  • •Design and refine incident response processes and documentation, and analyze operational metrics and RCAs to improve monitoring dashboards and support workflows.

Key Requirements

  • •Bachelor’s degree in Computer Science or related field, with a strong software engineering foundation.
  • •8+ years in technical operations roles (e.g., SRE/NOC), including designing monitoring systems and resolving production issues in mission-critical environments.
  • •Deep familiarity with monitoring, alerting, and observability practices, including metrics, logging, and tracing for distributed systems (SLIs/SLOs, alert tuning, dashboards).
  • •Proven experience leading incident response for high-severity outages, performing real-time incident coordination, root cause analysis, and driving post-mortems and action items.
  • •Strong skills in scripting or software engineering (e.g., Python) to automate repetitive tasks and integrate tools.
Education:Bachelor's
Skills:CommunicationCross-functional collaborationTroubleshootingIncident coordination
Tech Stack:PythonMonitoringAlertingObservabilityMetricsLoggingTracingSLIsSLOsDashboardsCloud infrastructureLoad balancersDatabasesContainerized applicationsDistributed systems

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor