Senior Engineer II, Inference Engine - Serving Engine
San Francisco
Workplace: RemoteFull timeUSD 167,200 - 209,000 annuallyFunction: Data Science & Machine LearningExperience: 5+ yearsSkills: ["Go","Python","GRPC","Distributed systems","Microservices","Inference engines","Tensor parallelism","KV cache","Dynamo","Ray Serve","NVlink","XGMI","RoCE","Quantization"]Lead design and delivery of high-scale data plane services powering AI inference for a multi-tenant cloud. Architect resilient systems, optimize distributed hosting for large generative models, mentor engineers, and collaborate across product and engineering teams to align roadmaps with customer needs, while maintaining platform health and observability.
Loading
Loading job details...
Preparing the role view and application actions.

