Member of Technical Staff, Inference & Serving

Inception Labs
San Mateo
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringSkills: ["Collaboration","Problem-solving","Communication"]

Design, optimize, and scale high-performance model serving systems for diffusion LLMs in production, focusing on low-latency inference, distributed orchestration, and reliable, cost-efficient operation. Collaborate with ML researchers to translate architectural advances into production-ready serving improvements, and implement robust monitoring and deployment workflows to meet SLA requirements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Inception Labs
Inception Labs
6 months ago

Member of Technical Staff, Inference & Serving

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Design, optimize, and scale high-performance model serving systems for diffusion LLMs in production, focusing on low-latency inference, distributed orchestration, and reliable, cost-efficient operation. Collaborate with ML researchers to translate architectural advances into production-ready serving improvements, and implement robust monitoring and deployment workflows to meet SLA requirements.
Location: San Mateo
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering

Key Responsibilities

  • •Build and optimize high-performance model serving systems for low-latency inference of diffusion LLMs.
  • •Extend orchestration frameworks (Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving.
  • •Implement and manage load balancing, autoscaling, and traffic routing for model endpoints.
  • •Build systems for model versioning, canary deployments, and zero-downtime rollouts.
  • •Develop monitoring, alerting, and observability tooling to ensure SLA compliance and rapid incident response.

Pay and Benefits

Equity and Bonus:Equity
Perks:EquityHealth InsuranceDentalVisionMeal AllowanceCommuter BenefitsPaid Leave

Key Requirements

  • •BS/MS/PhD in Computer Science, Engineering, or a related field (or equivalent experience)
  • •Knowledge of ML serving frameworks (SGLang, vLLM, Triton Inference Server, TensorRT-LLM)
  • •Understanding of ML frameworks (PyTorch, TensorFlow) from a systems perspective
  • •Familiarity with high-performance computing and GPU programming (CUDA)
  • •Experience with containerization (Docker), orchestration (Kubernetes), and CI/CD pipelines
Experience:AIDiffusionDistributed systems
Skills:CollaborationProblem-solvingCommunication
Languages:English
Tech Stack:KubernetesRaySLURMDockerTriton Inference ServerTensorRT-LLMPyTorchTensorFlowCUDAKubeflowAirflowAWSGCPAzure

Company Brief

Inception Labs
Develops artificial intelligence solutions and research-driven products, focusing on machine learning models, AI tools, and enterprise AI integrations to help organizations automate workflows, extract insights from data, and build intelligent applications.
Industry: AI & Machine Learning
Website