AI Systems Engineer – AI Model (Training & Inference)
Markham
Workplace: HybridFull timeFunction: IT Operations (Systems/Network Admin)Skills: ["Debugging","Optimization","Observability","Collaboration","Reproducibility"]Own the end-to-end model execution stack on AMD Instinct GPUs, spanning large-scale training infrastructure and high-performance inference serving. Enable and optimize LLM/VLM/MoE training on GPU clusters using Kubernetes, distributed checkpointing, and robust validation. Deliver inference acceleration by optimizing GPU kernels and serving frameworks for throughput/latency, quantization pipelines, and observability for large-scale reliability.
Loading
Loading job details...
Preparing the role view and application actions.

