Software Engineer - Model Products

Baseten
San Francisco, New York, Toronto, Canada, Montreal
Workplace: OnsiteFull timeUSD 150,000 - 230,000 annuallyFunction: Software EngineeringSkills: ["Communication","Collaboration","Problem-solving","Debugging","Writing"]

Join Baseten’s Model Performance team to design, build, and operate the Model API surface for hosted open-source models. You’ll optimize multi-GPU serving, implement advanced inference features, and instrument benchmarks and observability to ensure fast, reliable AI model delivery at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Baseten
Baseten
10 months ago

Software Engineer - Model Products

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Join Baseten’s Model Performance team to design, build, and operate the Model API surface for hosted open-source models. You’ll optimize multi-GPU serving, implement advanced inference features, and instrument benchmarks and observability to ensure fast, reliable AI model delivery at scale.
Location: San Francisco, New York, Toronto, Canada, Montreal
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving
  • •Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups
  • •Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving
  • •Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations
  • •Productionize performance improvements across runtimes (e.g.TensorRT, TensorRT‑LLM): speculative decoding, quantization, batching, and KV‑cache reuse.

Pay and Benefits

Salary: USD 150,000 - 230,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kEquityParental Leave

Key Requirements

  • •3+ years experience building and operating distributed systems or large-scale APIs.
  • •Proven track record of owning low-latency, reliable backend services (rate-limiting, auth, quotas, metering, migrations).
  • •Infra instincts with performance sensibilities: profiling, tracing, capacity planning, and SLO management.
  • •Comfortable debugging complex systems, from runtime internals to GPU execution traces.
  • •Strong written communication; able to produce clear design docs and collaborate across functions.
Experience:AIOpen SourceDistributed Systems
Skills:CommunicationCollaborationProblem-solvingDebuggingWriting
Tech Stack:TensorRTCUDAGPULLMVLLMSGLangTensorRT-LLMKubernetesAPI servingDistributed systems

Company Brief

Baseten
Baseten provides an inference-first ML infrastructure platform that lets engineering and ML teams deploy, serve, and scale machine-learning models with optimized performance, autoscaling, and GPU-backed hosting for production AI applications. ([crunchbase.com](https://www.crunchbase.com/organization/baseten?utm_source=openai))
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2019
Glassdoor
Glassdoor: 5.0
WebsiteLinkedInGlassdoor