Machine Learning Infrastructure Engineer

1001
London
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 4+ yearsSkills: []

Own the serving, deployment, and observability infrastructure that makes machine learning models production-ready across enterprise and government environments. Build and operate multi-tenant ML serving with model registries, artifacts, datasets, and deployment security, while establishing consistent deployment practices. Improve inference latency, throughput, reliability, and cost across CPU/GPU workloads, and create reusable platform components so product teams can ship faster.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
1001
1001
1 month ago

Machine Learning Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Own the serving, deployment, and observability infrastructure that makes machine learning models production-ready across enterprise and government environments. Build and operate multi-tenant ML serving with model registries, artifacts, datasets, and deployment security, while establishing consistent deployment practices. Improve inference latency, throughput, reliability, and cost across CPU/GPU workloads, and create reusable platform components so product teams can ship faster.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build and operate ML serving infrastructure across enterprise and government deployments.
  • •Own the deployment pipeline, including model registries, artifacts, datasets, deployment security, and multi-tenant serving.
  • •Implement monitoring, logging, and observability to make model performance visible and surface issues early.
  • •Improve inference latency, throughput, reliability, and cost across CPU and GPU workloads.
  • •Maintain reusable platform components and establish consistent deployment practices across customer environments.

Key Requirements

  • •At least 4 years of experience in infrastructure or platform engineering, including meaningful ML infrastructure or MLOps experience.
  • •A track record of operating machine learning systems in production (not only developing models).
  • •Hands-on experience with a model serving stack such as Triton, TorchServe, Ray Serve, vLLM, BentoML, or KServe.
  • •Strong experience with Kubernetes, containers, CI/CD, and infrastructure as code using Terraform.
  • •Production cloud experience with AWS, GCP, or Azure, plus monitoring and observability tools like Prometheus, Grafana, or OpenTelemetry.
Experience:4+ yearsMLOpsMachine learning production
Tech Stack:TritonTorchServeRay ServeVLLMBentoMLKServeKubernetesTerraformCI/CDAWSGCPAzurePrometheusGrafanaOpenTelemetryPythonTypeScript

Company Brief

1001
1001 AI builds AI-powered operational intelligence for complex, data-heavy industries like aviation, ports, energy, and construction. Its platform unifies operational data to predict delays, manage risk, and optimize performance across critical infrastructure and megaproject environments.
Industry: AI & Machine Learning
Company Size: Micro (1 to 10 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: London, United Kingdom
WebsiteLinkedIn