Software Engineer, Inference (AI Data Engineering)

SpaceX
Palo Alto
Workplace: OnsiteFull timeUSD 135,000 - 210,000 annuallyFunction: Software EngineeringExperience: 2+ yearsEducation: bachelorsSkills: ["Problem-solving","Collaboration","Ownership","Respect","Fairness"]

Build and optimize SpaceX’s high-performance AI inference platform that serves internal models powering mission-critical applications. Own end-to-end design for scalable distributed model serving, from request routing and SDKs to global KV cache and continuous batching. Improve latency and throughput with low-level GPU kernel optimizations, quantization, and speculative decoding, while delivering high-concurrency, 100% uptime systems with strong observability and CI/CD for reliable endpoint deployments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
SpaceX
SpaceX
2 days ago

Software Engineer, Inference (AI Data Engineering)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 minutes agoStatus: Live

Job Summary

Build and optimize SpaceX’s high-performance AI inference platform that serves internal models powering mission-critical applications. Own end-to-end design for scalable distributed model serving, from request routing and SDKs to global KV cache and continuous batching. Improve latency and throughput with low-level GPU kernel optimizations, quantization, and speculative decoding, while delivering high-concurrency, 100% uptime systems with strong observability and CI/CD for reliable endpoint deployments.
Location: Palo Alto
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Develop highly reliable, high-throughput inference systems to serve internal AI models across SpaceX
  • •Architect and implement scalable distributed infrastructure for model serving (e.g., load balancing, auto-scaling, batch scheduling, global KV cache, continuous batching)
  • •Optimize inference latency and throughput under production workloads, including GPU kernel work, quantization, and speculative decoding
  • •Build high-concurrency serving systems with 100% uptime, low tail latency, and strong observability
  • •Own end-to-end components and tooling, including request routing, SDK development, rate limiting, benchmarking/acceleration, and CI/CD for endpoint deployment

Pay and Benefits

Salary: USD 135,000 - 210,000 annually
Perks:Health InsuranceVisionDental401kPaid ParentalLife InsurancePaid LeavePaid Holidays

Key Requirements

  • •Bachelor’s degree in computer science, engineering, math, or a scientific discipline; or 2+ years of professional software development experience in lieu of a degree
  • •Experience designing, implementing, and maintaining reliable, horizontally scalable distributed systems
  • •1+ years building full-stack or backend production systems
  • •1+ years of experience with Rust or C++
  • •Ability to work onsite in Palo Alto (remote/hybrid not considered)
Experience:2+ years
Education:Bachelor's
Skills:Problem-solvingCollaborationOwnershipRespectFairness
Tech Stack:RustC++PythonGoSGLangVLLMTritonTensorRT-LLMDockerKubernetesPostgreSQLClickHouseMongoDBGRPCRESTCI/CDContinuous integrationContinuous deliveryMonitoringGPU kernels

Eligibility

Nationality:US National

Company Brief

SpaceX
Designs, manufactures, and launches advanced rockets and spacecraft for commercial and government customers, aiming to reduce space transportation costs and enable human life on Mars through reusable launch vehicles and integrated space systems.
Industry: Aerospace Manufacturing
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: Hawthorne, United States
Founded: 2002
Glassdoor
Glassdoor: 4.2
WebsiteLinkedIn