Software Engineer, Infrastructure

Exa
San Francisco
Workplace: OnsiteFull timeUSD 150,000 - 300,000 annuallyFunction: Software EngineeringSkills: ["Rust","Kubernetes","GPU","Batch processing","Observability","Distributed systems"]

Design and operate large-scale infrastructure for Exa’s AI-powered search engine, focusing on GPU clusters, Kubernetes-based orchestration, and cloud batch processing. You’ll improve reliability, observability, and performance across the stack while building tooling to scale a massive GPU-driven cluster and distributed workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Exa
Exa
1 year ago

Software Engineer, Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Design and operate large-scale infrastructure for Exa’s AI-powered search engine, focusing on GPU clusters, Kubernetes-based orchestration, and cloud batch processing. You’ll improve reliability, observability, and performance across the stack while building tooling to scale a massive GPU-driven cluster and distributed workloads.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Build Kubernetes orchestration on a multi-million-dollar GPU cluster to support scalable workloads
  • •Scale AWS batchjob system to handle map-reduce jobs across tens of thousands of machines
  • •Design GPU scheduling software to maximize cluster utilization and efficiency
  • •Develop and improve observability across production systems for reliability and performance
  • •Contribute to tooling and infrastructure powering Exa’s systems to move fast as an engineering org

Pay and Benefits

Salary: USD 150,000 - 300,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Experience designing and operating large-scale infrastructure (GPU clusters or large Kubernetes clusters or cloud batchjob systems)
  • •Proven focus on reliability, observability, and optimization across the stack
  • •Ability to design and operate Kubernetes orchestration for GPU clusters
  • •Experience scaling batchjob systems (e.g., AWS batch) to thousands of machines
  • •Familiarity with GPU scheduling, observability tooling, and distributed systems
Experience:AI infrastructureGPU clusteringKubernetesDistributed systemsCloud batch processing
Skills:RustKubernetesGPUBatch processingObservabilityDistributed systems
Tech Stack:RustKubernetesRayAWSGPUHPCObservability

Eligibility

Visa:H1BOPTSTEM OPTO1E3
Work Authorization:Sponsorship available.

Company Brief

Exa
Builds an embeddings-based neural search engine and web search API for AI applications, offering crawling, retrieval, and deep research tools to serve developers and AI agents with up-to-date web knowledge.
Industry: API Platforms
Company Size: Small (11 to 50 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn