Member of Technical Staff, Systems Infrastructure (2026 PhD New Grad)

Fireworks AI
San Mateo, New York
Workplace: OnsiteFull timeUSD 200,000 - 230,000 annuallyFunction: Research & Scientific (R&D)Education: phdSkills: ["Communication","Performance measurement","Systems thinking"]

Design and build the infrastructure substrate that powers Fireworks’ production AI at fleet scale—GPU scheduling and resource management, high-performance distributed storage/caching for model weights and KV cache, and datacenter networking to minimize inference latency. Measure and benchmark system behavior, identify bottlenecks across kernels/drivers to scheduler policies, and turn research ideas into production systems while partnering with research and inference teams to guide technical direction.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Fireworks AI
Fireworks AI
2 days ago

Member of Technical Staff, Systems Infrastructure (2026 PhD New Grad)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Design and build the infrastructure substrate that powers Fireworks’ production AI at fleet scale—GPU scheduling and resource management, high-performance distributed storage/caching for model weights and KV cache, and datacenter networking to minimize inference latency. Measure and benchmark system behavior, identify bottlenecks across kernels/drivers to scheduler policies, and turn research ideas into production systems while partnering with research and inference teams to guide technical direction.
Location: San Mateo, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Graduate level

Key Responsibilities

  • •Design, build, and operate core infrastructure systems for large-scale training and inference.
  • •Benchmark, trace, and simulate scheduling, caching, and network behavior to reason before committing to designs.
  • •Identify and eliminate bottlenecks across the stack, from kernel/driver to scheduler policy, using data.
  • •Turn research ideas into production systems that hold up under real workloads and failures.
  • •Partner with research and inference teams so infrastructure and model design inform each other.

Pay and Benefits

Salary: USD 200,000 - 230,000 annually

Key Requirements

  • •PhD completed within the last 6 months, or expected completion by December 2026, in Computer Science, Computer Engineering, Electrical Engineering, or a similar field.
  • •Research background in distributed systems, operating systems, scheduling/resource management, storage systems, computer networks, computer architecture, or high-performance computing.
  • •Depth in at least one focus area: scheduling & resource management, distributed storage & caching, or datacenter networking.
  • •Strong systems programming skills in C/C++, Rust, Go, or a similar language, plus working proficiency in Python.
  • •Experience building and evaluating real systems and communicating systems design and results clearly across audiences.
Experience:Distributed systemsHigh-performance computingOpen sourceGPU clustersCloud infrastructure
Education:PhD / Doctorate in Computer Science, Computer Engineering, Electrical Engineering, or a similar field
Skills:CommunicationPerformance measurementSystems thinking
Tech Stack:C/C++RustGoPythonCUDAROCmTritonKubernetesRaySlurmVLLMNCCLRCCLDPDK/SPDKCephRDMA/RoCEInfiniBandNCCL/RCCL

Company Brief

Fireworks AI
Develops AI-driven tools to generate and optimize visual marketing content for brands and creators, automating production of short-form videos and multimedia assets for social platforms to improve engagement and scale creative workflows.
Industry: SaaS
Website