Member of Technical Staff - ML Performance

Modal
New York, San Francisco, Stockholm
Workplace: OnsiteFull timeUSD 150,000 - 350,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Communication","Problem-solving","Teamwork"]

We’re seeking a seasoned engineer to optimize ML systems performance at scale, contributing to open-source projects and Modal’s container runtime. You’ll push language and diffusion models toward higher throughput and lower latency, collaborating with AI infra teams to improve production workloads across NYC, SF, and Stockholm offices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Modal
Modal
1 year ago

Member of Technical Staff - ML Performance

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

We’re seeking a seasoned engineer to optimize ML systems performance at scale, contributing to open-source projects and Modal’s container runtime. You’ll push language and diffusion models toward higher throughput and lower latency, collaborating with AI infra teams to improve production workloads across NYC, SF, and Stockholm offices.
Location: New York, San Francisco, Stockholm
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Manager level

Key Responsibilities

  • •Design and optimize ML systems for performance at scale.
  • •Contribute to open-source projects and Modal’s container runtime to push language and diffusion models toward higher throughput and lower latency.
  • •Collaborate with other engineers to improve infrastructure for production AI workloads.
  • •Help shape the performance engineering roadmap for GPU-accelerated ML workloads.
  • •Work in-person at our NYC, San Francisco or Stockholm offices.

Pay and Benefits

Salary: USD 150,000 - 350,000 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •5+ years of experience writing high-quality, high-performance code.
  • •Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).
  • •Familiarity with Nvidia GPU architecture and CUDA.
  • •Experience with ML performance engineering (boost GPU performance—debugging SM occupancy issues, rewriting an algorithm to be compute-bound, eliminating host overhead).
  • •Ability to work in-person, in our NYC, San Francisco or Stockholm office.
Experience:5+ yearsAI infrastructureML performanceGPU
Skills:CommunicationProblem-solvingTeamwork
Tech Stack:TorchTensorRTVLLMCUDANVIDIA GPULinuxContainers

Company Brief

Modal
Provides a serverless, high-performance cloud platform for AI, ML, and data workloads — offering instant autoscaling, elastic GPU access, and developer-first tooling to run inference, training, and batch jobs at scale.
Industry: Cloud Computing
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Valuation: Unicorn (USD 1B+)
Funding: Series B
Headquarters: New York City, United States
Founded: 2021
WebsiteLinkedIn