AI Inference Engineer

F5
San Jose, Seattle
Workplace: OnsiteFull timeUSD 176,600 - 265,000 annuallyFunction: Data Science & Machine LearningSkills: ["Problem-solving","Scalability","Performance optimization","Cross-functional collaboration","Initiative"]

Build and maintain high-performance inference engines that optimize Large Language Models for deployment across GPU data centers and resource-constrained edge devices. You’ll focus on low-latency, high-throughput serving, hardware acceleration (CUDA/TensorRT, CoreML, TPUs/LPUs), and scalable online and batch inference pipelines. Establish observability and performance testing to ensure reliability during traffic spikes, leveraging Kubernetes-based orchestration and infrastructure best practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
F5
F5
1 day ago

AI Inference Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 day agoStatus: Live

Job Summary

Build and maintain high-performance inference engines that optimize Large Language Models for deployment across GPU data centers and resource-constrained edge devices. You’ll focus on low-latency, high-throughput serving, hardware acceleration (CUDA/TensorRT, CoreML, TPUs/LPUs), and scalable online and batch inference pipelines. Establish observability and performance testing to ensure reliability during traffic spikes, leveraging Kubernetes-based orchestration and infrastructure best practices.
Location: San Jose, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Build and maintain robust AI inference engines using vLLM, TGI, and NVIDIA Triton for high-performance serving at scale.
  • •Optimize model deployment to deliver low-latency AI serving across multiple business applications.
  • •Profile and optimize models for specialized hardware backends, including NVIDIA GPUs, Apple Silicon via CoreML, and TPU/LPU accelerators.
  • •Design and implement auto-scaling online and batch inference pipelines using Kubernetes orchestration.
  • •Create observability and run performance/load testing to measure TTFT, tokens per second, and memory bandwidth utilization against SLAs.

Pay and Benefits

Salary: USD 176,600 - 265,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Proficiency in high-performance AI programming using Python, C++, Rust, or Golang.
  • •Hands-on inference experience with vLLM, TensorRT, Llama.cpp, and Ollama.
  • •Strong infrastructure skills with Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure.
  • •Expertise profiling and optimizing performance for GPU/AI accelerators including NVIDIA GPUs and TPUs.
  • •Experience deploying LLMs with techniques like Speculative Decoding or PagedAttention (preferred).
Skills:Problem-solvingScalabilityPerformance optimizationCross-functional collaborationInitiative
Tech Stack:PythonC++RustGolangVLLMText Generation Inference (TGI)NVIDIA TritonTensorRTCUDAApple SiliconCoreMLTPUsLPUsKubernetesDockerAWSGCPAzureLlama.cppOllama

Company Brief

F5
Provides application delivery networking, load balancing, and security solutions for on-premises and cloud environments, helping organizations optimize, secure, and scale applications and APIs across multi-cloud infrastructures.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Seattle, United States
Founded: 1996
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn