Senior Staff Engineer, ML Ops (R4941)

Shield AI
San Francisco
Workplace: OnsiteFull timeUSD 320,000 - 490,000 annuallyFunction: Data Science & Machine LearningSkills: ["Communication","Problem-solving","Leadership"]

Lead and scale the centralized AI and Data Platform, owning the core infra for training, simulation, data management, evaluation, and deployment across on-premise, cloud, and hybrid environments. Drive compute strategy, cost-per-compute, and edge deployment while enabling MPC-style ML workloads (foundation models, RL/MARL) and end-to-end model lifecycle from training to deployment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Shield AI
Shield AI
4 months ago

Senior Staff Engineer, ML Ops (R4941)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Lead and scale the centralized AI and Data Platform, owning the core infra for training, simulation, data management, evaluation, and deployment across on-premise, cloud, and hybrid environments. Drive compute strategy, cost-per-compute, and edge deployment while enabling MPC-style ML workloads (foundation models, RL/MARL) and end-to-end model lifecycle from training to deployment.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Define and operate the core AI and data platform across training, simulation, data management, evaluation, and deployment.
  • •Own where and how workloads run across on-premise, cloud, and hybrid environments; drive capacity planning and cost-per-compute decisions.
  • •Build infrastructure for distributed training and large-scale simulation; ensure training and simulation systems operate together without bottlenecks.
  • •Ingest and manage multi-modal sensor data; establish dataset versioning, data lineage, feature storage, data cataloging, and access controls.
  • •Establish a repeatable workflow for experiment tracking, model registry, artifact provenance, evaluation/gates, and end-to-end lifecycle from training to deployment.

Pay and Benefits

Salary: USD 320,000 - 490,000 annually

Key Requirements

  • •Experience building and operating ML infrastructure at scale (100+ GPU clusters, distributed systems)
  • •Experience defining compute strategy, including on-premise vs cloud tradeoffs, capacity planning, and cost management
  • •Strong understanding of ML workloads, including foundation models, RL/MARL, simulation-based training, and fine-tuning
  • •Experience building data platforms with dataset versioning, lineage, and cataloging
  • •Ability to debug and resolve system issues when needed
Experience:DefenseAutonomyRoboticsSimulation
Skills:CommunicationProblem-solvingLeadership
Languages:English
Tech Stack:GPUOn-premiseCloudHybridSovereignRLMARLFoundation modelsDistributed systemsEdge

Company Brief

Shield AI
Develops AI-powered autonomy and software for military aircraft and drones to enable autonomous ISR and combat missions, integrating perception, navigation, and mission planning for defense customers.
Industry: Defense Technology
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Funding: Series E+
Headquarters: San Diego, United States
Founded: 2015
WebsiteLinkedIn