Member of Technical Staff, Cloud Infrastructure

Fireworks AI
San Mateo, New York
Workplace: OnsiteFull timeUSD 175,000 - 220,000 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Mentorship","Leadership"]

Architect and build scalable, resilient cloud infrastructure that powers distributed AI training, inference, and data processing. Lead technical design discussions, mentor engineers, and establish best practices for large-scale ML systems. Implement core backend services like job scheduling, resource management, autoscaling, and model serving, while optimizing compute cost, storage lifecycle, and network performance. Own end-to-end systems with strong observability, reliability, and fault tolerance across multi-cloud environments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Fireworks AI
Fireworks AI
1 year ago

Member of Technical Staff, Cloud Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live
Reposted: similar role first listed 1 week ago

Job Summary

Architect and build scalable, resilient cloud infrastructure that powers distributed AI training, inference, and data processing. Lead technical design discussions, mentor engineers, and establish best practices for large-scale ML systems. Implement core backend services like job scheduling, resource management, autoscaling, and model serving, while optimizing compute cost, storage lifecycle, and network performance. Own end-to-end systems with strong observability, reliability, and fault tolerance across multi-cloud environments.
Location: San Mateo, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Architect and build scalable, resilient backend infrastructure for distributed training, inference, and data processing pipelines.
  • •Lead technical design discussions, mentor other engineers, and establish best practices for building and operating large-scale ML infrastructure.
  • •Design and implement backend services such as job schedulers, resource managers, autoscalers, and model serving layers with a focus on efficiency and low latency.
  • •Drive infrastructure optimization initiatives, including compute cost reduction, storage lifecycle management, and network performance tuning.
  • •Own end-to-end systems from design to deployment and observability, ensuring reliability, fault tolerance, disaster recovery, and performance across multi-cloud infrastructure.

Pay and Benefits

Salary: USD 175,000 - 220,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Bachelor’s degree in Computer Science, Engineering, or related technical field (or equivalent practical experience).
  • •5+ years designing and building backend infrastructure in cloud environments (AWS, GCP, Azure).
  • •Proven experience in ML infrastructure and tooling (e.g., PyTorch, TensorFlow, Vertex AI, SageMaker, Kubernetes).
  • •Strong software development skills in languages like Python or C++.
  • •Deep understanding of distributed systems fundamentals including scheduling, orchestration, storage, networking, and compute optimization.
Experience:5+ years
Education:Bachelor's in Computer Science, Engineering
Skills:MentorshipLeadership
Tech Stack:AWSGCPAzurePyTorchTensorFlowVertex AISageMakerKubernetesKubeflowMLFlowTerraformArgoCDGitOpsPythonC++

Company Brief

Fireworks AI
Develops AI-driven tools to generate and optimize visual marketing content for brands and creators, automating production of short-form videos and multimedia assets for social platforms to improve engagement and scale creative workflows.
Industry: SaaS
Website