Inference Optimization Engineer

Modular
United States, Canada
Workplace: RemoteFull timeUSD 198,000 - 286,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Communication","Technical leadership","Judgment","Collaboration","Problem-solving"]

Build an optimization platform for state-of-the-art LLM inference on Modular Cloud, targeting top cost/performance across GPU and ASIC architectures. Shape technical direction for inference performance on the Pareto frontier, and collaborate with engineering, product, and GTM to translate customer workload insights into repeatable automated optimization loops. Partner cross-functionally and publish thought leadership on inference optimization best practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Modular
Modular
3 months ago

Inference Optimization Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: Just nowStatus: Live

Job Summary

Build an optimization platform for state-of-the-art LLM inference on Modular Cloud, targeting top cost/performance across GPU and ASIC architectures. Shape technical direction for inference performance on the Pareto frontier, and collaborate with engineering, product, and GTM to translate customer workload insights into repeatable automated optimization loops. Partner cross-functionally and publish thought leadership on inference optimization best practices.
Location: United States, Canada
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build the optimization platform that drives LLM inference performance across the latest GPU and ASIC architectures.
  • •Shape Modular Cloud technical direction to deliver and maintain Pareto-frontier LLM performance for agentic use cases.
  • •Partner with GTM and collaborate with engineering to deliver customized LLM inference tuned to customer workloads, spanning GPU kernels to cloud infrastructure.
  • •Profile customer inference workloads end to end and apply optimizations across kernels, inference engine, and distributed systems.
  • •Publish blog posts on innovative approaches to LLM inference optimization.
Travel: Low travel

Pay and Benefits

Salary: USD 198,000 - 286,000 annually
Equity and Bonus:Equity
Perks:Health Insurance401kPaid Leave

Key Requirements

  • •5+ years of experience in distributed systems or performance engineering.
  • •Proven ability to build durable, reusable software tools and libraries adopted across teams.
  • •Sound judgment evaluating technical tradeoffs and setting priorities, with strong communication and technical leadership.
  • •Strong collaboration mindset with creativity and curiosity to solve complex problems.
  • •Experience with GPU kernel programming, inference engine internals, or distributed inference architectures (helpful, not required).
Experience:5+ yearsDistributed systemsPerformance engineeringAI infrastructureLLM inferenceCloud native
Skills:CommunicationTechnical leadershipJudgmentCollaborationProblem-solving
Tech Stack:GPUASICLLMsKubernetesCloud native

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

Modular
Provides infrastructure and developer tools to build, deploy, and scale large AI models and foundation-model applications, including model hosting, orchestration, and SDKs to accelerate AI product development.
Industry: AI & Machine Learning
Website