Member of Technical Staff - Machines

Modal
San Francisco, New York
Workplace: OnsiteFull timeUSD 250,000 - 300,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Production engineering","Debugging","Incident response"]

Design and operate the machines layer of a serverless AI platform, including the fleet of bare metal and cloud hosts and the control plane that provisions, images, monitors, and repairs them. Build automation to integrate new hardware capacity, configure GPUs and networking, and remediate unhealthy machines without human intervention. Debug issues across the hardware-software stack and improve uptime, reliability, and production-ready infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Modal
Modal
1 day ago

Member of Technical Staff - Machines

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Design and operate the machines layer of a serverless AI platform, including the fleet of bare metal and cloud hosts and the control plane that provisions, images, monitors, and repairs them. Build automation to integrate new hardware capacity, configure GPUs and networking, and remediate unhealthy machines without human intervention. Debug issues across the hardware-software stack and improve uptime, reliability, and production-ready infrastructure.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Build and maintain the machines layer and control plane that provisions, images, monitors, and repairs bare metal and cloud hosts.
  • •Automate integration of new hardware capacity from multiple providers, including auditing/benchmarking hosts and clusters.
  • •Maintain machine images and configure GPUs, RDMA, networking, and storage for production workloads.
  • •Detect and automatically remediate unhealthy machines (e.g., bad GPUs, thermals, disks) without human intervention.
  • •Debug complex hardware-to-software issues and adapt the container runtime to new architectures and platforms.

Pay and Benefits

Salary: USD 250,000 - 300,000 annually

Key Requirements

  • •5+ years writing high-quality production code.
  • •Experience operating fleets of bare metal hardware (provisioning, BMC/IPMI, PXE/network boot, firmware) or building control planes for them.
  • •Strong cloud skills.
  • •Strong knowledge of low-level OS foundations (Linux kernel, drivers, networking, file systems, containers).
  • •Effective debugging across layers and willingness to respond to production incidents via on-call rotation.
Experience:5+ years
Skills:Production engineeringDebuggingIncident response
Tech Stack:PythonGoLinux kernelLinuxBMC/IPMIPXENetwork bootContainersBGPRDMANVIDIA driversXIDsRDMA/NVLinkVBIOSGPUNetworking

Company Brief

Modal
Provides a serverless, high-performance cloud platform for AI, ML, and data workloads — offering instant autoscaling, elastic GPU access, and developer-first tooling to run inference, training, and batch jobs at scale.
Industry: Cloud Computing
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Valuation: Unicorn (USD 1B+)
Funding: Series B
Headquarters: New York City, United States
Founded: 2021
WebsiteLinkedIn