ML Infrastructure Engineer, Training

Dyna Robotics
Redwood City
Workplace: OnsiteFull timeUSD 180,000 - 270,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsEducation: bachelorsSkills: ["Communication","Problem-solving","Teamwork"]

Design and scale large-scale ML training infrastructure in a high-performance computing environment. You’ll architect GPU-accelerated training pipelines, optimize distributed systems, and collaborate with ML researchers to accelerate model iteration while ensuring reliability and scalability across cloud platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Dyna Robotics
Dyna Robotics
7 months ago

ML Infrastructure Engineer, Training

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Design and scale large-scale ML training infrastructure in a high-performance computing environment. You’ll architect GPU-accelerated training pipelines, optimize distributed systems, and collaborate with ML researchers to accelerate model iteration while ensuring reliability and scalability across cloud platforms.
Location: Redwood City
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Architect and implement large-scale ML training pipelines that leverage parallel GPU processing on platforms like GCP or AWS.
  • •Enhance our existing infrastructure to fully exploit parallelism and design for future expansion, ensuring that our system is ready to support growth.
  • •Manage and optimize high-performance computing resources and develop robust distributed computing solutions, addressing challenges like race conditions, memory optimization, and resource allocation.
  • •Design systems for job rescheduling, automated retries, and failure recovery to maximize uptime and training efficiency; implement intelligent job queuing mechanisms to optimize training workloads and resource utilization.
  • •Evaluate and implement tradeoffs between different local and networked storage solutions to improve data throughput and access; develop strategies for caching training data to optimize performance.

Pay and Benefits

Salary: USD 180,000 - 270,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Bachelor’s degree or higher in Computer Science or a related field.
  • •At least 7 years of professional experience in the software industry, with a minimum of 2 years in a tech lead role.
  • •Proven experience with high-performance computing environments and distributed systems.
  • •Demonstrated ability to scale ML training systems and optimize resource utilization.
  • •Hands-on experience with job scheduling systems and managing cloud GPU environments (GCP, AWS, etc.).
Experience:7+ yearsRoboticsAIHigh-performance computingCloud
Education:Bachelor's in Computer Science
Skills:CommunicationProblem-solvingTeamwork
Tech Stack:PyTorchTensorRTTritonAccelerateGCPAWSKubernetesDistributed systems

Company Brief

Dyna Robotics
Develops embodied-AI powered robotic arms and foundation models (DYNA-1) to automate repetitive, stationary tasks across hotels, restaurants, laundromats and retail. Focuses on affordable, deployable robots that generalize across environments and improve with on‑device learning.
Industry: Robotics
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Valuation: USD 500M to 1B
Funding: Series A
Headquarters: Redwood City, United States
Founded: 2024
WebsiteLinkedInGlassdoor