Senior Site Reliability Engineer, AI Inference

F5
Dublin
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Python","C++","Rust","Golang","VLLM","TensorRT","Llama.cpp","Ollama","Docker","Kubernetes","AWS","GCP","Azure","CUDA","TPU","CoreML","NVIDIA GPUs","TTFT","Latency"]

Senior AI inference-focused SRE responsible for building and optimizing high-performance inference engines, deploying scalable, low-latency AI services, and ensuring reliability across GPU-rich data centers and edge environments. You’ll work with vLLM, TensorRT, and Kubernetes to maximize throughput, minimize latency, and maintain model accuracy in production.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
F5
F5
3 months ago

Senior Site Reliability Engineer, AI Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Senior AI inference-focused SRE responsible for building and optimizing high-performance inference engines, deploying scalable, low-latency AI services, and ensuring reliability across GPU-rich data centers and edge environments. You’ll work with vLLM, TensorRT, and Kubernetes to maximize throughput, minimize latency, and maintain model accuracy in production.
Location: Dublin
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Build and maintain robust inference engines using tools like vLLM, TGI, and NVIDIA Triton to ensure high performance at scale.
  • •Deliver low-latency AI serving solutions for multiple applications through deployment optimizations.
  • •Profile and optimize models for specialized hardware backends including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), TPUs, and LPUs.
  • •Collaborate with hardware teams to maximize utilization and performance across diverse computational environments.
  • •Design and implement auto-scaling architectures for online and batch inference pipelines using Kubernetes for routing and orchestration.

Key Requirements

  • •Proficiency in programming languages such as Python, C++, Rust, or Golang for high-performance AI workflows
  • •Hands-on experience with AI inference tools (vLLM, TensorRT, Llama.cpp, Ollama) for development and optimization
  • •Infrastructure experience with Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure)
  • •Hardware acceleration and optimization expertise for GPUs/TPUs and accelerators
  • •Experience in MLOps or SRE roles focused on high-throughput, low-latency AI endpoints
Experience:AIMLCloud
Skills:PythonC++RustGolangVLLMTensorRTLlama.cppOllamaDockerKubernetesAWSGCPAzureCUDATPUCoreMLNVIDIA GPUsTTFTLatency
Languages:English
Tech Stack:PythonC++RustGolangVLLMTensorRTLlama.cppOllamaDockerKubernetesAWSGCPAzureCUDANVIDIA GPUsTPUCoreMLApple SiliconGPUTTFT

Company Brief

F5
Provides application delivery networking, load balancing, and security solutions for on-premises and cloud environments, helping organizations optimize, secure, and scale applications and APIs across multi-cloud infrastructures.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Seattle, United States
Founded: 1996
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn