Member of Technical Staff - Inference Systems

Liquid AI
Boston
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["Quick learning","Methodology focus","Attention to detail","Technical evaluation","Reproducibility mindset"]

Build and maintain Liquid AI’s inference engine and benchmarking infrastructure that evaluates model performance and quality across different hardware and partner environments. You’ll design benchmark suites, run partner verifications, and port models across runtimes while ensuring end-to-end numerical correctness. Working closely with research and product—and directly with external engineering teams—you’ll extend the inference layer built on llama.cpp, ONNX, and MLX and make results verifiable and reproducible.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Liquid AI
Liquid AI
2 days ago

Member of Technical Staff - Inference Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Build and maintain Liquid AI’s inference engine and benchmarking infrastructure that evaluates model performance and quality across different hardware and partner environments. You’ll design benchmark suites, run partner verifications, and port models across runtimes while ensuring end-to-end numerical correctness. Working closely with research and product—and directly with external engineering teams—you’ll extend the inference layer built on llama.cpp, ONNX, and MLX and make results verifiable and reproducible.
Location: Boston
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and build benchmark suites for inference performance, model quality, and knowledge evaluation across hardware targets.
  • •Run external partner verifications by evaluating solutions against benchmarks, identifying gaps, and delivering clear findings.
  • •Port models to different runtimes/frameworks and verify correctness end-to-end.
  • •Maintain and extend the inference engine layer built on llama.cpp, ONNX, and MLX as new model architectures emerge.
  • •Make benchmark results explainable, verifiable, and independently reproducible by internal teams and partners.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kPaid Leave

Key Requirements

  • •Hands-on experience with an inference framework (llama.cpp, ONNX Runtime, or MLX), including internals and modification beyond basic usage.
  • •Experience designing and building benchmarking pipelines with strong methodology, validation, and reproducibility.
  • •Strong C++ and Python skills in performance-sensitive inference contexts.
  • •Solid understanding of inference fundamentals, including quantization, decoding strategies, and memory layout tradeoffs.
  • •Ability to port models across runtimes and verify numerical correctness, with attention to proof of correct outputs.
Experience:AIInferenceMachine learningBenchmarkingEdge inferencePartner evaluation
Skills:Quick learningMethodology focusAttention to detailTechnical evaluationReproducibility mindset
Tech Stack:C++PythonLlama.cppONNXONNX RuntimeMLXQuantizationDecoding strategies

Company Brief

Liquid AI
Builds AI infrastructure and tooling to enable real-time, distributed machine learning and orchestration across edge and cloud environments, simplifying deployment and management of intelligent applications.
Industry: AI & Machine Learning
Website