Software Engineer, Inference
San Francisco
Workplace: OnsiteFull timeUSD 150,000 - 230,000 annuallyFunction: Software EngineeringSkills: ["Performance engineering","Profiling","Optimization","Evaluation","Capacity planning"]Build low-latency, high-throughput inference services for OCR and multimodal models. Own performance profiling, batching, caching, and autoscaling across single-tenant and multi-tenant environments, with clear SLOs. Optimize model pipelines by improving kernels, tokenization, and model graphs, evaluate vLLM/TensorRT LLM/Triton tradeoffs, and drive capacity planning through performance dashboards.
Loading
Loading job details...
Preparing the role view and application actions.

