Software Engineer, Inference

Pulse
San Francisco
Workplace: OnsiteFull timeUSD 150,000 - 230,000 annuallyFunction: Software EngineeringSkills: ["Performance engineering","Profiling","Optimization","Evaluation","Capacity planning"]

Build low-latency, high-throughput inference services for OCR and multimodal models. Own performance profiling, batching, caching, and autoscaling across single-tenant and multi-tenant environments, with clear SLOs. Optimize model pipelines by improving kernels, tokenization, and model graphs, evaluate vLLM/TensorRT LLM/Triton tradeoffs, and drive capacity planning through performance dashboards.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pulse
Pulse
1 year ago

Software Engineer, Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Build low-latency, high-throughput inference services for OCR and multimodal models. Own performance profiling, batching, caching, and autoscaling across single-tenant and multi-tenant environments, with clear SLOs. Optimize model pipelines by improving kernels, tokenization, and model graphs, evaluate vLLM/TensorRT LLM/Triton tradeoffs, and drive capacity planning through performance dashboards.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build inference services using smart batching and caching.
  • •Optimize kernels, tokenization, and model graphs for performance.
  • •Evaluate vLLM, TensorRT LLM, and Triton tradeoffs.
  • •Implement autoscaling and admission control with clear SLOs.
  • •Own performance dashboards and drive capacity planning.

Pay and Benefits

Salary: USD 150,000 - 230,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceVisionDentalRelocationMeal Allowance

Key Requirements

  • •3+ years in performance engineering or ML systems
  • •Strong Python skills, plus C++ or CUDA exposure
  • •Experience with GPU profiling and model serving
  • •Experience with profiling, batching, and autoscaling in production inference environments
  • •Familiarity with reducing p95 and costs in production ML systems
Experience:ML systemsComputer visionNLPDocument intelligenceOCR
Skills:Performance engineeringProfilingOptimizationEvaluationCapacity planning
Tech Stack:PythonC++CUDAGPU profilingModel servingVLLMTensorRTTritonOCRMultimodal modelsTokenizationModel graphsAutoscalingAdmission controlBatchingCachingPerformance dashboards

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

Pulse
Provides a digital mental health platform offering on-demand coaching, therapy, care navigation, and employer analytics to support employee wellbeing, improve access to behavioral healthcare, and inform organizational mental health strategies.
Industry: Mental Health
Website