Software Engineer - Training Infrastructure

Baseten
San Francisco, New York
Workplace: OnsiteFull timeUSD 200,000 - 275,000 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Leadership","Mentoring","Collaboration","Problem-solving"]

Frontier AI infrastructure role focused on designing and delivering distributed training systems for large-scale foundation models. You’ll optimize GPU utilization, build scalable training pipelines, and collaborate with product and infra teams to push the state-of-the-art in scalable model training within a fast-growing AI platform.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Baseten
Baseten
1 year ago

Software Engineer - Training Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Frontier AI infrastructure role focused on designing and delivering distributed training systems for large-scale foundation models. You’ll optimize GPU utilization, build scalable training pipelines, and collaborate with product and infra teams to push the state-of-the-art in scalable model training within a fast-growing AI platform.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Design, build, and maintain distributed training infrastructure for large-scale foundation models
  • •Implement scalable pipelines for fine-tuning and training across heterogeneous GPU/accelerator clusters
  • •Optimize training performance through techniques like FSDP, DDP, ZeRO, and mixed precision training
  • •Contribute to frameworks and tooling that make training workflows efficient, reproducible, and developer-friendly
  • •Collaborate with cross-functional teams (Product, Forward Deployed Engineering, Inference Infra) to ensure training systems meet real-world requirements

Pay and Benefits

Salary: USD 200,000 - 275,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kEquity

Key Requirements

  • •Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent experience
  • •5+ years of experience in ML infrastructure, distributed systems, or ML platform engineering, including 2+ years in a tech lead or manager role
  • •Strong expertise in distributed training frameworks and orchestration (FSDP, DDP, ZeRO, Ray, Kubernetes, Slurm, or similar)
  • •Hands-on experience building or scaling training infrastructure for LLMs or other foundation models
  • •Deep understanding of GPU/accelerator hardware utilization, mixed precision training, and scaling efficiency
Experience:5+ yearsAIML infrastructureDistributed systemsFoundation modelsGPU training
Education:Bachelor's
Skills:CommunicationLeadershipMentoringCollaborationProblem-solving
Tech Stack:FSDPDDPZeRORayKubernetesSlurmGPUMixed precisionLLMsFoundation models

Company Brief

Baseten
Baseten provides an inference-first ML infrastructure platform that lets engineering and ML teams deploy, serve, and scale machine-learning models with optimized performance, autoscaling, and GPU-backed hosting for production AI applications. ([crunchbase.com](https://www.crunchbase.com/organization/baseten?utm_source=openai))
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2019
Glassdoor
Glassdoor: 5.0
WebsiteLinkedInGlassdoor