AI Inference Engineer

F5
Dublin
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["Problem-solving","Communication","Collaboration"]

AI Inference Engineer designs and optimizes high-performance AI serving for Large Language Models, spanning GPU-rich data centers to edge devices. You’ll build and optimize robust inference engines, integrate hardware acceleration, and design scalable infrastructure with Kubernetes, ensuring low latency, high throughput, and reliable enterprise-grade deployment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
F5
F5
3 months ago

AI Inference Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

AI Inference Engineer designs and optimizes high-performance AI serving for Large Language Models, spanning GPU-rich data centers to edge devices. You’ll build and optimize robust inference engines, integrate hardware acceleration, and design scalable infrastructure with Kubernetes, ensuring low latency, high throughput, and reliable enterprise-grade deployment.
Location: Dublin
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Build and maintain robust inference engines using vLLM, Text Generation Inference, and NVIDIA Triton for high performance at scale.
  • •Profile and optimize models for specialized hardware backends including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon CoreML, TPUs and LPUs.
  • •Design and implement auto-scaling architectures for online and batch inference pipelines using Kubernetes for routing and orchestration.
  • •Establish observability frameworks to monitor TTFT, tokens per second, memory bandwidth, and SLA compliance; build performance/load testing suites.
  • •Collaborate with hardware teams to maximize utilization and performance across diverse compute environments.

Key Requirements

  • •5+ requirements? (Must-have)
  • •Python, C++, Rust, or Golang proficiency for high-performance AI workflows
  • •Hands-on experience with vLLM, TensorRT, Llama.cpp, Ollama for inference
  • •Infrastructure familiarity with Docker, Kubernetes, AWS, GCP, Azure
  • •Hardware acceleration knowledge for NVIDIA GPUs (CUDA/TensorRT), CoreML, TPUs, LPUs
Experience:AIMLOpsLLMs
Skills:Problem-solvingCommunicationCollaboration
Tech Stack:PythonC++RustGolangVLLMTensorRTLlama.cppOllamaNVIDIA GPUsCUDACoreMLTPUsLPUsDockerKubernetesAWSGCPAzure

Company Brief

F5
Provides application delivery networking, load balancing, and security solutions for on-premises and cloud environments, helping organizations optimize, secure, and scale applications and APIs across multi-cloud infrastructures.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Seattle, United States
Founded: 1996
Glassdoor
Glassdoor: 3.8
WebsiteLinkedIn