LLM/ML Engineer (Inference)
San Francisco
Workplace: OnsiteFull timeUSD 200,000 - 300,000 annuallyFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Python","PyTorch","CUDA","Triton","TensorRT","VLLM","Optimum"]Join a hands-on ML/AI core infra team to design and optimize scalable inference systems for state-of-the-art models. You’ll implement robust serving architectures, reduce latency, integrate advanced inference techniques, collaborate with research, and build tooling to enable rapid experimentation in a fast-paced, in-person SF environment.
Loading
Loading job details...
Preparing the role view and application actions.

