Research Engineer, Infrastructure, Inference

Thinking Machines Lab
San Francisco
Workplace: HybridFull timeUSD 350,000 - 475,000 annuallyFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Collaboration","Initiative"]

Design and scale the infrastructure that powers large AI model inference and deployment at research and production scale. Collaborate with researchers and engineers to improve performance, latency, throughput, and efficiency through new techniques and architectures. Optimize compute and codebase utilization on GPUs, extend orchestration frameworks for distributed inference and evaluation, and set standards for reliability, observability, and reproducibility. Share learnings via documentation, open-source libraries, or technical reports.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thinking Machines Lab
Thinking Machines Lab
23 hours ago

Research Engineer, Infrastructure, Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Design and scale the infrastructure that powers large AI model inference and deployment at research and production scale. Collaborate with researchers and engineers to improve performance, latency, throughput, and efficiency through new techniques and architectures. Optimize compute and codebase utilization on GPUs, extend orchestration frameworks for distributed inference and evaluation, and set standards for reliability, observability, and reproducibility. Share learnings via documentation, open-source libraries, or technical reports.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Work alongside researchers and engineers to bring cutting-edge AI models into production.
  • •Collaborate with research teams to enable high-performance inference for novel architectures.
  • •Design and implement techniques, tools, and architectures to improve performance, latency, throughput, and efficiency.
  • •Optimize the codebase and compute fleet (e.g., GPUs) to fully utilize hardware compute, bandwidth, and memory.
  • •Extend orchestration frameworks (e.g., Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving.

Pay and Benefits

Salary: USD 350,000 - 475,000 annually
Perks:Health InsuranceDentalVisionPaid LeavePaid ParentalRelocation

Key Requirements

  • •Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • •Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
  • •Experience with inference serving systems optimized for throughput and latency (e.g., SGLang, vLLM).
  • •Thrive in a highly collaborative, cross-functional environment with subject matter experts.
  • •Strong engineering skills, contributing performant, maintainable code and debugging in complex codebases.
Experience:Open-sourceML infrastructure
Education:Bachelor's
Skills:CollaborationInitiative
Tech Stack:PyTorchJAXSGLangVLLMKubernetesRaySLURMGPUsTritonDeepSpeedXLA

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Thinking Machines Lab
Develops enterprise AI solutions, custom large language models, and ML platforms to help organizations deploy intelligent applications. Services include data engineering, model development, and AI consulting for scale and production readiness.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Mumbai, India
Website