Machine Learning Infrastructure Engineer

Windborne Systems
Redwood City
Workplace: OnsiteFull timeUSD 140,000 - 240,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Systems thinking","Process orientation","Problem-solving"]

Own end-to-end reliability and scaling for production AI weather models, from research-to-operations pipelines to real-time inference with strict latency requirements. Build health monitoring, logging, alerting, and failure diagnosis across nodes. Design data and training infrastructure that tolerates upstream data issues, silent OOMs, and network/storage faults, using both on-prem compute and cloud resources with cost/performance tradeoffs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Windborne Systems
Windborne Systems
2 months ago

Machine Learning Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Own end-to-end reliability and scaling for production AI weather models, from research-to-operations pipelines to real-time inference with strict latency requirements. Build health monitoring, logging, alerting, and failure diagnosis across nodes. Design data and training infrastructure that tolerates upstream data issues, silent OOMs, and network/storage faults, using both on-prem compute and cloud resources with cost/performance tradeoffs.
Location: Redwood City
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Build research-to-operations pipelines and own uptime end-to-end with health monitoring, improved logging, and failure diagnosis across nodes.
  • •Drive inference scaling and compute strategy across on-prem clusters and cloud providers, optimizing cost/performance as usage grows.
  • •Create training and real-time data pipelines that handle upstream delays and data quality issues, including QC checks and edge-case logging/alerting.
  • •Make distributed training runs reliable with monitoring, auto-recovery, and job scheduling to reduce researcher babysitting.
  • •Manage and improve operational workflows supporting continuous model releases.

Pay and Benefits

Salary: USD 140,000 - 240,000 annually
Perks:401kHealth InsuranceDentalVisionPaid LeaveEquityMeal Allowance

Key Requirements

  • •Have experience running production ML systems and building deployments that are reliable under real constraints.
  • •Experience with large datasets.
  • •Ability to keep up with fast-paced model releases and build reliable custom deployments for them.
  • •Experience with PyTorch and Docker, including debugging memory, compression, and network saturation issues.
  • •Comfort bringing structure to a fast-moving team by improving infrastructure processes and organization.
Skills:Systems thinkingProcess orientationProblem-solving
Tech Stack:PyTorchDocker

Company Brief

Windborne Systems
Develops airborne systems and technologies for wind measurement, environmental monitoring, and renewable energy applications, focusing on hardware design, sensor integration, and data solutions to support wind resource assessment and operational monitoring.
Industry: Climate Tech
Website