Senior Machine Learning Engineer

Cloudflare
Austin, London
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["Mentorship","Technical leadership","Cross-functional collaboration"]

Define how machine learning models run across a global serverless inference platform, from frontier LLMs and speech/vision models to customer-deployed inference on heterogeneous GPUs and accelerators. Build benchmarking and evaluation frameworks, improve latency/throughput/cost via quantization and runtime tuning, and integrate models into distributed infrastructure. Drive safer model deployment workflows for low-latency, reliable quality at Internet scale, while mentoring engineers and raising production ML standards.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cloudflare
Cloudflare
1 month ago

Senior Machine Learning Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Define how machine learning models run across a global serverless inference platform, from frontier LLMs and speech/vision models to customer-deployed inference on heterogeneous GPUs and accelerators. Build benchmarking and evaluation frameworks, improve latency/throughput/cost via quantization and runtime tuning, and integrate models into distributed infrastructure. Drive safer model deployment workflows for low-latency, reliable quality at Internet scale, while mentoring engineers and raising production ML standards.
Location: Austin, London
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Develop, optimize, and productionize machine learning models for a serverless inference platform, focusing on performance, reliability, and model quality.
  • •Build benchmarking and evaluation frameworks to measure latency, throughput, cost efficiency, and model behavior across LLMs, speech, vision, and other model families.
  • •Improve inference performance via quantization, batching, caching, model compilation, runtime tuning, and accelerator-aware optimization.
  • •Partner with systems engineers to integrate models into distributed inference infrastructure across heterogeneous GPUs and next-generation accelerators.
  • •Drive deployment workflow improvements including validation, rollout safety, observability, regression testing, and operational readiness, while mentoring engineers.

Key Requirements

  • •Experience building, optimizing, and operating machine learning models in production environments.
  • •Strong proficiency with Python and modern ML frameworks such as PyTorch, TensorFlow, or JAX (or equivalent).
  • •Hands-on experience with inference optimization for large-scale models, including quantization, batching, caching, compilation, and runtime tuning.
  • •Experience with large-scale inference serving frameworks or runtimes such as vLLM, TensorRT-LLM, ONNX Runtime, Triton, SGLang, or llama.cpp (or similar).
  • •Familiarity with modern deep learning architectures including LLMs, speech, vision, embeddings, multimodal models, or retrieval-augmented generation.
Experience:Open sourceLLMs
Skills:MentorshipTechnical leadershipCross-functional collaboration
Languages:English
Tech Stack:PythonPyTorchTensorFlowJAXSGLangVLLMTensorRT-LLMONNX RuntimeTritonLlama.cppQuantizationBatchingCachingModel compilationRuntime tuningGPUsAcceleratorsWorkers AI

Company Brief

Cloudflare
Provides a global network and cloud platform that delivers security, performance, and reliability services for web applications, APIs, and Internet properties, including CDN, DDoS protection, DNS, and zero-trust security solutions.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 2009
WebsiteLinkedIn