Inference Optimization Manager

Modular
United States, Canada
Workplace: RemoteFull timeUSD 234,000 - 286,000 annuallyFunction: Software EngineeringSkills: ["Technical leadership","Communication","Prioritization","Problem-solving","Collaboration"]

Lead the Performance Labs team building Modular Cloud’s optimization platform for LLM inference. Own the technical direction from GPU kernels through inference engine and distributed systems, turning customer workloads into a continuous optimization loop. Partner with GTM, Product, and Engineering to deliver highly customized inference performance for agentic use cases, grow the team, and share external thought leadership through technical blog posts.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Modular
Modular
3 months ago

Inference Optimization Manager

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 minutes agoStatus: Live

Job Summary

Lead the Performance Labs team building Modular Cloud’s optimization platform for LLM inference. Own the technical direction from GPU kernels through inference engine and distributed systems, turning customer workloads into a continuous optimization loop. Partner with GTM, Product, and Engineering to deliver highly customized inference performance for agentic use cases, grow the team, and share external thought leadership through technical blog posts.
Location: United States, Canada
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Manager level

Key Responsibilities

  • •Lead the team building the optimization platform for LLM inference on Modular Cloud across GPU and ASIC architectures.
  • •Own Modular Cloud’s technical direction to deliver and sustain state-of-the-art inference performance for agentic use cases.
  • •Partner with GTM to deliver customized LLM inference tuned to customer workloads and guide cross-stack optimizations from kernels to cloud infrastructure.
  • •Grow and develop the team by setting priorities and enabling engineers to tackle hard performance problems.
  • •Champion the team externally by publishing blog posts on LLM inference optimization approaches and best practices.
Travel: Low travel

Pay and Benefits

Salary: USD 234,000 - 286,000 annually
Equity and Bonus:Equity
Perks:Health Insurance401kPaid Leave

Key Requirements

  • •5+ years in distributed systems or performance engineering, including experience leading or managing engineering teams.
  • •Ship durable, reusable software tools and libraries adopted across teams, and guide a team to do the same.
  • •Evaluate technical tradeoffs and set priorities with strong communication and technical leadership.
  • •Translate ambiguous customer and product needs into focused engineering direction.
  • •Demonstrate creativity, curiosity, collaboration, and alignment with team culture.
Experience:AI infrastructureLLMDistributed systemsPerformance engineering
Skills:Technical leadershipCommunicationPrioritizationProblem-solvingCollaboration
Tech Stack:LLMGPUASICDistributed systemsKubernetesCloud nativeInference engine

Company Brief

Modular
Provides infrastructure and developer tools to build, deploy, and scale large AI models and foundation-model applications, including model hosting, orchestration, and SDKs to accelerate AI product development.
Industry: AI & Machine Learning
Website