Software Engineer, Inference - Multi Modal

OpenAI
San Francisco
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["CPU/GPU optimization","Distributed systems","Scaling","Inference","Model deployment"]

Join OpenAI’s Inference team to design and build high-performance inference infrastructure for multimodal models, delivering real-time audio, image, and other modalities at scale. Collaborate with researchers and product engineers to deploy state-of-the-art capabilities, optimize GPU-accelerated pipelines, and improve system-level performance like tensor parallelism and hardware abstraction layers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
1 year ago

Software Engineer, Inference - Multi Modal

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 minutes agoStatus: Live

Job Summary

Join OpenAI’s Inference team to design and build high-performance inference infrastructure for multimodal models, delivering real-time audio, image, and other modalities at scale. Collaborate with researchers and product engineers to deploy state-of-the-art capabilities, optimize GPU-accelerated pipelines, and improve system-level performance like tensor parallelism and hardware abstraction layers.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and implement inference infrastructure for large-scale multimodal models.
  • •Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs.
  • •Enable experimental research workflows to transition into reliable production services.
  • •Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities.
  • •Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers.

Key Requirements

  • •Experience building and scaling inference systems for LLMs or multimodal models
  • •Worked with GPU-based ML workloads and understood performance dynamics of large models with images or audio
  • •Design and implement inference infrastructure for large-scale multimodal models
  • •Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs
  • •Collaborate with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities
Experience:MultimodalLLMInference
Skills:CPU/GPU optimizationDistributed systemsScalingInferenceModel deployment
Tech Stack:VLLMTensorRT-LLMGPUTensor parallelism

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor