Member of Technical Staff

Fireworks AI
New York
Workplace: OnsiteFull timeUSD 175,000 - 220,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 4+ yearsEducation: bachelorsSkills: ["Technical leadership","Mentoring","Cross-functional collaboration","Systems design","Technical documentation"]

Design and maintain large-scale backend and cloud-native infrastructure for distributed machine learning training, inference, and data processing. Architect scalable, resilient systems with an emphasis on reliability, fault tolerance, and low latency. Lead technical design discussions, mentor engineers, and establish best practices while optimizing compute cost, storage lifecycle, and network performance. Evaluate and integrate technologies like Kubernetes, Ray, Kubeflow, and MLflow to strengthen platform reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Fireworks AI
Fireworks AI
2 months ago

Member of Technical Staff

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live
Reposted: similar role first listed 2 months ago

Job Summary

Design and maintain large-scale backend and cloud-native infrastructure for distributed machine learning training, inference, and data processing. Architect scalable, resilient systems with an emphasis on reliability, fault tolerance, and low latency. Lead technical design discussions, mentor engineers, and establish best practices while optimizing compute cost, storage lifecycle, and network performance. Evaluate and integrate technologies like Kubernetes, Ray, Kubeflow, and MLflow to strengthen platform reliability.
Location: New York
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Architect and build scalable, resilient backend infrastructure for distributed training, inference, and data processing pipelines.
  • •Lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems.
  • •Design and implement core backend services optimized for efficiency and low latency.
  • •Drive infrastructure optimization for compute cost, storage lifecycle management, and network performance.
  • •Collaborate with machine learning, DevOps, and product teams and evaluate/integrate cloud-native and open-source technologies (e.g., Kubernetes, Ray, Kubeflow, MLflow) from design to deployment.

Pay and Benefits

Salary: USD 175,000 - 220,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Bachelor's degree or equivalent in Computer Science (or related) plus 4 years of software engineering experience.
  • •4 years designing/building/optimizing large-scale backend infrastructure and distributed data systems in cloud environments (AWS, GCP, Azure) including optimization techniques like caching, indexing, sharding, replication, and ACID.
  • •4 years with major server-side programming languages/frameworks (Python, C++, Go, TypeScript).
  • •3 years developing/maintaining data processing and API systems, including client-server communication frameworks like gRPC or Thrift.
  • •Experience with cloud-native tools and infrastructure such as Docker and Kubernetes, plus A/B testing and coding interviews/feedback.
Experience:4+ yearsMachine learningGenerative AIDistributed systemsCloud environmentsEnterprise AI
Education:Bachelor's in Computer Science
Skills:Technical leadershipMentoringCross-functional collaborationSystems designTechnical documentation
Tech Stack:PythonC++GoTypeScriptAWSGCPAzurePostgreSQLMySQLDynamoDBApache SparkApache FlinkApache KafkaKubernetesRayKubeflowMLFlowDockerGRPCThrift

Company Brief

Fireworks AI
Develops AI-driven tools to generate and optimize visual marketing content for brands and creators, automating production of short-form videos and multimedia assets for social platforms to improve engagement and scale creative workflows.
Industry: SaaS
Website