LLM/ML Engineer (Inference)

Reducto
San Francisco
Workplace: OnsiteFull timeUSD 200,000 - 300,000 annuallyFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Python","PyTorch","CUDA","Triton","TensorRT","VLLM","Optimum"]

Join a hands-on ML/AI core infra team to design and optimize scalable inference systems for state-of-the-art models. You’ll implement robust serving architectures, reduce latency, integrate advanced inference techniques, collaborate with research, and build tooling to enable rapid experimentation in a fast-paced, in-person SF environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reducto
Reducto
1 year ago

LLM/ML Engineer (Inference)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Join a hands-on ML/AI core infra team to design and optimize scalable inference systems for state-of-the-art models. You’ll implement robust serving architectures, reduce latency, integrate advanced inference techniques, collaborate with research, and build tooling to enable rapid experimentation in a fast-paced, in-person SF environment.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Architect and implement robust, scalable inference systems for serving state-of-the-art AI models
  • •Optimize model serving infrastructure for high throughput and low latency at scale
  • •Develop and integrate advanced inference optimization techniques
  • •Work closely with the research team to bring cutting-edge capabilities into production
  • •Build developer tools and infrastructure to support rapid experimentation and deployment

Pay and Benefits

Salary: USD 200,000 - 300,000 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Deep expertise in Python and PyTorch
  • •Experience with modern inference systems (TGI, vLLM, TensorRT-LLM, Optimum)
  • •Strong foundation in low-level operating systems concepts (multithreading, memory management, networking, storage, performance, scale)
  • •Ability to architect robust, scalable inference systems for state-of-the-art AI models and optimize model serving infrastructure
  • •Experience creating custom tooling for testing and optimization
Experience:3+ yearsMachine learningAIInferenceLow-level systems
Skills:PythonPyTorchCUDATritonTensorRTVLLMOptimum
Tech Stack:PythonPyTorchCUDATritonTensorRTVLLMOptimum

Eligibility

Visa:H1B
Work Authorization:Sponsorship available.

Company Brief

Reducto
Builds AI-driven tools for automated transcription, summarization, and insights extraction from audio and video to help teams analyze conversations, surface key moments, and accelerate decision-making workflows.
Industry: SaaS
Website