AI Inference Engineer
San Jose, Seattle
Workplace: OnsiteFull timeUSD 176,600 - 265,000 annuallyFunction: Data Science & Machine LearningSkills: ["Problem-solving","Scalability","Performance optimization","Cross-functional collaboration","Initiative"]Build and maintain high-performance inference engines that optimize Large Language Models for deployment across GPU data centers and resource-constrained edge devices. You’ll focus on low-latency, high-throughput serving, hardware acceleration (CUDA/TensorRT, CoreML, TPUs/LPUs), and scalable online and batch inference pipelines. Establish observability and performance testing to ensure reliability during traffic spikes, leveraging Kubernetes-based orchestration and infrastructure best practices.
Loading
Loading job details...
Preparing the role view and application actions.

