Software Engineer - Training/Inference (C++)

X AI
Palo Alto
Workplace: OnsiteFull timeUSD 180,000 - 440,000Function: Software EngineeringSkills: ["Communication","Prioritization","Curiosity"]

Build and optimize large-scale model serving systems for Grok, end-to-end. Own distributed infrastructure such as global KV cache, continuous batching, load balancing, and auto-scaling, then drive low-level GPU and inference optimizations including quantization and speculative decoding. Improve latency, throughput, and reliability under production workloads, develop debugging tools across the full stack, and create CI/CD infrastructure for seamless endpoint and engine updates.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
X AI
X AI
1 year ago

Software Engineer - Training/Inference (C++)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build and optimize large-scale model serving systems for Grok, end-to-end. Own distributed infrastructure such as global KV cache, continuous batching, load balancing, and auto-scaling, then drive low-level GPU and inference optimizations including quantization and speculative decoding. Improve latency, throughput, and reliability under production workloads, develop debugging tools across the full stack, and create CI/CD infrastructure for seamless endpoint and engine updates.
Location: Palo Alto
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Architect and implement scalable distributed infrastructure for model serving, including load balancing, auto-scaling, batch scheduling, and global KV cache.
  • •Optimize latency and throughput of model inference under real production workloads.
  • •Build reliable, high-concurrency serving systems with strong uptime and tail latency performance.
  • •Benchmark, fine-tune, and accelerate inference engines, including low-level GPU kernel work and code generation.
  • •Develop tools to trace, replay, and fix issues across the full stack, and create CI/CD infrastructure for endpoint deployment and inference engine updates.

Pay and Benefits

Salary: USD 180,000 - 440,000
Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kDisability InsuranceLife InsuranceEquity

Key Requirements

  • •Deep low-level systems programming experience in C/C++ or Rust.
  • •Experience with large-scale, high-concurrent production serving.
  • •Experience with GPU inference engines such as vLLM, SGLang, Triton, or TensorRT-LLM.
  • •Strong background in systems optimization including batching, caching, load balancing, and parallelism.
  • •Experience with inference reliability and benchmarking, plus designing CI/CD infrastructure for inference.
Experience:AI inferenceLarge-scale serving
Skills:CommunicationPrioritizationCuriosity
Languages:English
Tech Stack:C++CRustGPU kernelsQuantizationSpeculative decodingVLLMSGLangTritonTensorRT-LLMCI/CDKV cache

Company Brief

X AI
Develops advanced artificial intelligence models and research aimed at building safe, general AI and understanding the fundamental nature of the universe. Focuses on large-scale AI systems, research publications, and building foundational AI capabilities.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Headquarters: San Francisco, United States
Founded: 2023
Website