Distributed LLM Inference Engineer
San Francisco
Workplace: HybridFull timeUSD 170,112 - 247,000 annuallyFunction: Software EngineeringSkills: ["PyTorch","Ray","CUDA","TensorFlow","VLLM","TensorRT-LLM","Triton","TVM","MLIR"]Distributed LLM Inference Engineer at Anyscale, focusing on building high-throughput, low-latency ML inference solutions at scale, integrating Ray Data and LLM engines, collaborating with open-source communities, and advancing state-of-the-art in distributed AI infrastructure.

