Member of Technical Staff, Performance Optimization

Fireworks AI
San Mateo
Workplace: OnsiteFull timeUSD 175,000 - 220,000 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Problem-solving","Collaboration","Performance optimization","Debugging","Analytical thinking"]

Own performance optimization across Fireworks’ AI infrastructure, improving speed and efficiency from low-level GPU kernels to distributed training and inference systems. Tackle bottlenecks in latency, throughput, memory, and compute efficiency for demanding workloads including LLMs, VLMs, and video models. Partner with research, infrastructure, and systems teams to implement CUDA/Triton optimizations, enhance mixed precision and quantization, and build benchmarking and monitoring for scalable multi-GPU, multi-node execution.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Fireworks AI
Fireworks AI
1 year ago

Member of Technical Staff, Performance Optimization

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live
Reposted: similar role first listed 1 year ago

Job Summary

Own performance optimization across Fireworks’ AI infrastructure, improving speed and efficiency from low-level GPU kernels to distributed training and inference systems. Tackle bottlenecks in latency, throughput, memory, and compute efficiency for demanding workloads including LLMs, VLMs, and video models. Partner with research, infrastructure, and systems teams to implement CUDA/Triton optimizations, enhance mixed precision and quantization, and build benchmarking and monitoring for scalable multi-GPU, multi-node execution.
Location: San Mateo
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Optimize system and GPU performance for high-throughput AI training and inference workloads.
  • •Analyze and improve latency, throughput, memory usage, and compute efficiency.
  • •Profile systems to detect and resolve GPU- and kernel-level bottlenecks.
  • •Implement low-level performance optimizations using CUDA, Triton, and related tooling.
  • •Build benchmarking and monitoring infrastructure and scale inference/training across multi-GPU, multi-node environments.

Pay and Benefits

Salary: USD 175,000 - 220,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • •5+ years optimizing performance or working on high-performance computing systems.
  • •Proficiency in CUDA or ROCm and experience with GPU profiling tools (Nsight, nvprof, CUPTI).
  • •Familiarity with PyTorch and performance-critical model execution.
  • •Experience debugging and optimizing distributed systems in multi-GPU environments.
Experience:5+ yearsHigh-performance computingGPU optimizationDistributed systemsGenerative AI
Education:Bachelor's
Skills:Problem-solvingCollaborationPerformance optimizationDebuggingAnalytical thinking
Tech Stack:CUDAROCmTritonPyTorchNsightNvprofCUPTITorch.compileXLAKubernetesRDMAInfinibandRoCEMixed precisionQuantizationModel graph optimizationShardingCUDA kernelsGPU profilingMulti-GPU

Company Brief

Fireworks AI
Develops AI-driven tools to generate and optimize visual marketing content for brands and creators, automating production of short-form videos and multimedia assets for social platforms to improve engagement and scale creative workflows.
Industry: SaaS
Website