Member of Technical Staff, Backend, LLM Applications

Inception Labs
San Mateo
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Python","Async","Distributed systems","Kubernetes","CI/CD","Cloud infrastructure","AWS","Azure","Load balancing","Terraform","Prometheus","Grafana","Triton Inference Server","TensorRT-LLM","VLLM"]

Experienced backend engineer focused on production-grade infrastructure for diffusion LLMs, designing scalable services, model serving endpoints, and zero-downtime deployments. You will optimize latency, throughput, and cost while building observability tooling and infrastructure to support billions of inferences. Work sits at the intersection of ML systems and backend infrastructure in a cutting-edge AI startup.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Inception Labs
Inception Labs
6 months ago

Member of Technical Staff, Backend, LLM Applications

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Experienced backend engineer focused on production-grade infrastructure for diffusion LLMs, designing scalable services, model serving endpoints, and zero-downtime deployments. You will optimize latency, throughput, and cost while building observability tooling and infrastructure to support billions of inferences. Work sits at the intersection of ML systems and backend infrastructure in a cutting-edge AI startup.
Location: San Mateo
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design, build, and operate scalable backend services and model serving infrastructure for our diffusion LLMs.
  • •Implement and manage load balancing, autoscaling, and traffic routing for model endpoints.
  • •Build systems for model versioning, canary deployments, and zero-downtime rollouts.
  • •Develop monitoring, alerting, and observability tooling to ensure SLA compliance and rapid incident response.
  • •Benchmark and evaluate serving frameworks and hardware configurations to inform infrastructure decisions.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionEquityCommuter BenefitsMeal AllowancePaid Leave

Key Requirements

  • •BS/MS/PhD in Computer Science or a related field (or equivalent experience).
  • •5+ years of experience building production backend systems.
  • •Strong proficiency in Python, including async programming and concurrent systems.
  • •Solid understanding of distributed systems, networking, and load balancing at scale.
  • •Familiarity with Kubernetes, CI/CD pipelines, and cloud infra (AWS and/or Azure).
Experience:5+ yearsArtificial intelligenceLLMsDiffusion models
Education:Bachelor's
Skills:PythonAsyncDistributed systemsKubernetesCI/CDCloud infrastructureAWSAzureLoad balancingTerraformPrometheusGrafanaTriton Inference ServerTensorRT-LLMVLLM
Tech Stack:PythonKubernetesAWSAzureTerraformPrometheusGrafanaTriton Inference ServerTensorRT-LLMVLLM

Company Brief

Inception Labs
Develops artificial intelligence solutions and research-driven products, focusing on machine learning models, AI tools, and enterprise AI integrations to help organizations automate workflows, extract insights from data, and build intelligent applications.
Industry: AI & Machine Learning
Website