AI Inference Engineer QVAC (100% remote Worldwide)

Tether.io
Bogotá, London, Dubai, Barcelona, Bengaluru, Islamabad, Hanoi, Prague, Dublin, Copenhagen, Tallinn, Helsinki, Athens, Zagreb, Lisbon, Rome, Stockholm, Oslo, Amsterdam, Warsaw, Milan, Budapest
Workplace: RemoteFull timeFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Problem-solving","Collaboration","Quick learning","Communication"]

Own the inference backbone behind QVAC's local AI stack, focusing on the C++ runtime that powers on-device models. You will port and optimize inference engines for edge devices, ensure startup behavior and memory/throughput balance, and collaborate with researchers to bring research models into production while enabling private, fast AI without cloud reliance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Tether.io
Tether.io
4 months ago

AI Inference Engineer QVAC (100% remote Worldwide)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Own the inference backbone behind QVAC's local AI stack, focusing on the C++ runtime that powers on-device models. You will port and optimize inference engines for edge devices, ensure startup behavior and memory/throughput balance, and collaborate with researchers to bring research models into production while enabling private, fast AI without cloud reliance.
Location: Bogotá, London, Dubai, Barcelona, Bengaluru, Islamabad, Hanoi, Prague, Dublin, Copenhagen, Tallinn, Helsinki, Athens, Zagreb, Lisbon, Rome, Stockholm, Oslo, Amsterdam, Warsaw, Milan, Budapest
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Deploy machine learning models to edge devices using llama.cpp, ggml, onnx
  • •Collaborate with researchers to help code, train, and transition models from research to production
  • •Integrate AI features into existing products with the latest ML advancements
  • •Ensure runtime stability, performance, and maintainability of the inference layer
  • •Port and optimize inference engines for diverse GPU architectures and edge hardware

Pay and Benefits

Perks:Remote WorkLearning Budget

Key Requirements

  • •Excellent programming skills in C++, experience in Javascript is a bonus
  • •Strong experience with Llama.cpp and ggml inference engines, which facilitates the deployment of models to specific GPU architectures
  • •Good understanding of deep learning concepts and model architectures
  • •Experience with transformers, LLMs, Diffusion models
  • •Demonstrated ability to rapidly assimilate new technologies and techniques
Experience:Edge computingOn-device AIAI hardware
Education:Bachelor's
Skills:Problem-solvingCollaborationQuick learningCommunication
Languages:English
Tech Stack:C++JavaScriptLlama.cppGgmlONNXTransformersLLMsDiffusion modelsVulkanOpenCL

Company Brief

Tether.io
Builds tools and services to help companies hire, onboard, and manage remote or distributed teams across borders, focusing on payroll, compliance, and global employment workflows.
Industry: HR Tech
Website