Machine Learning Performance Engineer - Offboard Training & Inference

Applied Intuition
Sunnyvale
Workplace: OnsiteFull timeUSD 215,000 - 285,000 annuallyFunction: Data Science & Machine LearningSkills: ["Profiling","Roofline analysis","Debugging","Analytical problem-solving","Collaboration"]

Own performance optimization for large-scale distributed training and high-throughput batch inference over petabyte-scale autonomy logs. Profile end to end, close the gap between theoretical accelerator capability and real workload throughput, and improve cluster goodput by reducing GPU idle time from data pipeline, I/O, scheduling gaps, stragglers, and recovery issues. Build benchmarking/observability tooling and collaborate across accelerator and ML infrastructure teams to reduce training time-to-result and offline processing cost.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Applied Intuition
Applied Intuition
1 day ago

Machine Learning Performance Engineer - Offboard Training & Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Own performance optimization for large-scale distributed training and high-throughput batch inference over petabyte-scale autonomy logs. Profile end to end, close the gap between theoretical accelerator capability and real workload throughput, and improve cluster goodput by reducing GPU idle time from data pipeline, I/O, scheduling gaps, stragglers, and recovery issues. Build benchmarking/observability tooling and collaborate across accelerator and ML infrastructure teams to reduce training time-to-result and offline processing cost.
Location: Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Profile and optimize distributed training end to end, including data loading/preprocessing, augmentation, kernel execution, gradient communication, and checkpointing.
  • •Optimize offline/batch inference over petabyte-scale sensor logs, including batching/scheduling strategies, quantization/low-precision execution, graph optimization, and accelerator saturation.
  • •Establish roofline and performance models to quantify the gap between achieved and theoretical performance and stack-rank optimization opportunities.
  • •Improve multi-node scaling efficiency and reduce cluster goodput gaps by addressing sharding/parallelism, interconnect utilization, memory-bandwidth and kernel-fusion bottlenecks, and failure recovery.
  • •Build benchmarking, observability, and regression-detection tooling to prevent silent performance degradation as models and code evolve.

Pay and Benefits

Salary: USD 215,000 - 285,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Hands-on ML performance engineering experience, including profiling, roofline analysis, throughput optimization, and root-cause investigation in production systems.
  • •Experience with distributed multi-node training at scale (e.g., FSDP, DeepSpeed, Megatron, NCCL) and diagnosing scaling inefficiency as node counts grow.
  • •Deep familiarity with GPU/accelerator performance concepts such as memory bandwidth, kernel launch overhead, occupancy, quantization, and collective communication.
  • •Experience with high-throughput/batch inference systems such as NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar.
  • •Fluency in Python and proficiency in C++ or another systems language. Efficient debugging, analytical, and problem-solving skills.
Experience:Machine learningDistributed trainingInferenceGPU accelerationRobotics/autonomy data
Skills:ProfilingRoofline analysisDebuggingAnalytical problem-solvingCollaboration
Tech Stack:PythonC++FSDPDeepSpeedMegatronNCCLNVIDIA Triton Inference ServerTensorRTONNX RuntimeRayCUDATritonCUTLASSNsight SystemsNsight ComputePyTorch ProfilerPerfKubernetesSlurmGPU scheduling

Company Brief

Applied Intuition
Builds software tools for autonomous vehicle development, including simulation, validation, and data infrastructure for automakers and robotics teams. Also supports defense applications with autonomy-focused testing and deployment platforms.
Industry: Developer Tools
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Mountain View, United States
Founded: 2017
WebsiteLinkedIn