Software Engineer- Model Performance Systems

Baseten
Vancouver, Canada, San Francisco, Toronto, New York, Montreal
Workplace: RemoteFull timeCAD 130,000 - 200,000 annuallyFunction: Software EngineeringSkills: ["Automation","Curiosity","Communication","Collaboration","Problem-solving"]

Join Baseten as an early-career software engineer focused on model performance tooling at the intersection of HPC and LLMs. You’ll build automated benchmarks, GPU cluster validation tools, GPU-enabled development environments, and instrumentation to optimize inference pipelines. Work with PyTorch Profiler, NVIDIA Nsight, and InferenceMAX to push the boundaries of AI infrastructure in a collaborative, hands-on team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Baseten
Baseten
7 months ago

Software Engineer- Model Performance Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Join Baseten as an early-career software engineer focused on model performance tooling at the intersection of HPC and LLMs. You’ll build automated benchmarks, GPU cluster validation tools, GPU-enabled development environments, and instrumentation to optimize inference pipelines. Work with PyTorch Profiler, NVIDIA Nsight, and InferenceMAX to push the boundaries of AI infrastructure in a collaborative, hands-on team.
Location: Vancouver, Canada, San Francisco, Toronto, New York, Montreal
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Entry level

Key Responsibilities

  • •Performance Benchmarking: Run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse).
  • •Infrastructure Validation: Create automated acceptance tests for new GPU clusters across x86 and ARM systems, measuring GPU memory bandwidth, networking throughput, and multi-node networking performance.
  • •Model Dev Experience: Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation.
  • •Tool Development: Build and contribute to tools such as InferenceMAX and genai-bench to automate model evaluation and optimization.
  • •Deep Hardware Profiling: Use PyTorch Profiler and NVIDIA Nsight Systems to collect performance profiles, identify bottlenecks, and debug the NVIDIA compute/networking stack.

Pay and Benefits

Salary: CAD 130,000 - 200,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kEquityParental Leave

Key Requirements

  • •Early-career software engineer focused on system and hardware performance
  • •Proficiency with Python and willingness to master NVIDIA software stack; C++ familiarity is a plus
  • •Interest in HPC/LLM engineering, GPU memory subsystems, InfiniBand, and cluster data movement
  • •Automation mindset; you script repetitive tasks and enjoy stress-testing to find breaking points
  • •Experience building GPU-enabled development environments and tools for model experimentation
Experience:AIMachine learning infrastructure
Skills:AutomationCuriosityCommunicationCollaborationProblem-solving
Languages:English
Tech Stack:PythonPyTorchNVIDIA NsightCUDAInfiniBandGPUCodespacesC++

Company Brief

Baseten
Baseten provides an inference-first ML infrastructure platform that lets engineering and ML teams deploy, serve, and scale machine-learning models with optimized performance, autoscaling, and GPU-backed hosting for production AI applications. ([crunchbase.com](https://www.crunchbase.com/organization/baseten?utm_source=openai))
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2019
Glassdoor
Glassdoor: 5.0
WebsiteLinkedInGlassdoor