Software Engineer - Model Performance

Baseten
San Francisco, New York, Toronto, Canada, Montreal
Workplace: OnsiteFull timeUSD 150,000 - 250,000 annuallyFunction: Software EngineeringEducation: bachelorsSkills: ["Problem-solving","Communication","Collaboration"]

Software Engineer focused on ML performance to join Baseten’s Model Performance team. You’ll productionize techniques for ML model inference, debug performance issues in PyTorch/TensorRT-based stacks, and optimize large language models across diverse hardware. Collaborate with a cross-functional team, own projects from idea to production, and contribute to open-source ML model tooling in a fast-paced startup environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Baseten
Baseten
2 years ago

Software Engineer - Model Performance

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Software Engineer focused on ML performance to join Baseten’s Model Performance team. You’ll productionize techniques for ML model inference, debug performance issues in PyTorch/TensorRT-based stacks, and optimize large language models across diverse hardware. Collaborate with a cross-functional team, own projects from idea to production, and contribute to open-source ML model tooling in a fast-paced startup environment.
Location: San Francisco, New York, Toronto, Canada, Montreal
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure.
  • •Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues.
  • •Apply and scale optimization techniques across a wide range of ML models, particularly large language models.
  • •Collaborate with a diverse team to design and implement innovative solutions.
  • •Own projects from idea to production.

Pay and Benefits

Salary: USD 150,000 - 250,000 annually
Equity and Bonus:Equity
Perks:EquityMedicalDentalVision401kParantal LeavePaid Leave

Key Requirements

  • •Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.
  • •Experience with one or more general-purpose programming languages, such as Python or C++.
  • •Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching).
  • •Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.
  • •Demonstrated interest and experience in LLMs.
Experience:LLMsInferenceML performance
Education:Bachelor's
Skills:Problem-solvingCommunicationCollaboration
Tech Stack:PythonC++PyTorchTensorRTTensorRT-LLMVllmCUDADockerKubernetesSglang

Company Brief

Baseten
Baseten provides an inference-first ML infrastructure platform that lets engineering and ML teams deploy, serve, and scale machine-learning models with optimized performance, autoscaling, and GPU-backed hosting for production AI applications. ([crunchbase.com](https://www.crunchbase.com/organization/baseten?utm_source=openai))
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2019
Glassdoor
Glassdoor: 5.0
WebsiteLinkedInGlassdoor