Inference Engineering and Product Lead

Modal
San Francisco
Workplace: OnsiteFull timeUSD 300,000 - 350,000 annuallyFunction: Product ManagementExperience: 10+ yearsSkills: ["Leadership","Coaching","Feedback","Career growth","Cross-functional collaboration"]

Own the direction and execution of Modal’s LLM inference platform, leading engineers building the serving stack, routing infrastructure, and internal optimization and product surfaces. Drive technical and product decisions through reviews and architecture discussions, partner with customers on frontier workloads, and translate learnings into a roadmap. Lead reliability and end-to-end ownership while collaborating cross-functionally on compute strategy and product launches.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Modal
Modal
15 hours ago

Inference Engineering and Product Lead

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Own the direction and execution of Modal’s LLM inference platform, leading engineers building the serving stack, routing infrastructure, and internal optimization and product surfaces. Drive technical and product decisions through reviews and architecture discussions, partner with customers on frontier workloads, and translate learnings into a roadmap. Lead reliability and end-to-end ownership while collaborating cross-functionally on compute strategy and product launches.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Product Management
Seniority: Manager level

Key Responsibilities

  • •Recruit, hire, and grow a high-performing team; run regular 1:1s and set performance expectations.
  • •Drive technical and product decisions via design reviews, code reviews, and architectural discussions.
  • •Work with customers on novel or frontier workloads running on Modal and translate learnings into a roadmap for the optimization platform and product.
  • •Establish standards for reliability and product excellence; ensure end-to-end ownership from spec through production.
  • •Partner cross-functionally on compute purchasing strategy and help guide roadmaps with adjacent infrastructure and product teams.

Pay and Benefits

Salary: USD 300,000 - 350,000 annually
Equity and Bonus:Equity

Key Requirements

  • •10+ years of industry experience, including 3+ years in a leadership role.
  • •Track record building high-performance systems at scale.
  • •Strong background in cloud infrastructure.
  • •Deep knowledge of low-level OS foundations (Linux kernel, file systems, containers, etc.).
  • •Nice to have: experience with LLM inference in production and familiarity with engines, kernels, routing, KV cache management, and speculative decoding.
Experience:10+ yearsCloud infrastructureLow-level OS foundationsLLM inference
Skills:LeadershipCoachingFeedbackCareer growthCross-functional collaboration
Tech Stack:Linux kernelFile systemsContainersLLM inferenceRoutingKV cache managementSpeculative decoding

Company Brief

Modal
Provides a serverless, high-performance cloud platform for AI, ML, and data workloads — offering instant autoscaling, elastic GPU access, and developer-first tooling to run inference, training, and batch jobs at scale.
Industry: Cloud Computing
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Valuation: Unicorn (USD 1B+)
Funding: Series B
Headquarters: New York City, United States
Founded: 2021
WebsiteLinkedIn