Machine Learning Platform Engineer

BJAK
United Kingdom, London
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Ownership","Experimentation","Continuous improvement","Reliability focus","Debugging/regression diagnosis","Collaboration"]

Build and operate the ML infrastructure that powers A1’s AI capabilities, covering model training/evaluation, deployment, inference, and experimentation. Own high-throughput, low-latency model serving and improve reliability, scalability, latency, and cost. Create data and CI-ready pipelines for training and evaluation, plus observability (monitoring, tracing, alerting) to detect and diagnose regressions. Enable faster iteration for AI engineers and researchers by developing reusable platform primitives.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
BJAK
BJAK
17 hours ago

Machine Learning Platform Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live
Reposted: similar role first listed 5 hours ago

Job Summary

Build and operate the ML infrastructure that powers A1’s AI capabilities, covering model training/evaluation, deployment, inference, and experimentation. Own high-throughput, low-latency model serving and improve reliability, scalability, latency, and cost. Create data and CI-ready pipelines for training and evaluation, plus observability (monitoring, tracing, alerting) to detect and diagnose regressions. Enable faster iteration for AI engineers and researchers by developing reusable platform primitives.
Location: United Kingdom, London
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Build and operate the ML infrastructure and platforms powering A1’s AI products.
  • •Design systems for model training, evaluation, deployment, inference, and experimentation.
  • •Optimize model serving and inference infrastructure for high-throughput and low-latency workloads.
  • •Improve reliability, scalability, latency, and cost efficiency of AI systems and identify stack bottlenecks.
  • •Develop observability (monitoring, tracing, alerting) and evaluation/benchmarking infrastructure to detect and diagnose regressions.

Key Requirements

  • •Strong software engineering fundamentals with experience building production systems.
  • •Experience building ML infrastructure, platforms, or production machine learning systems.
  • •Experience with model deployment, inference, evaluation, or data pipelines.
  • •Strong understanding of distributed systems and system reliability.
  • •Ability to write clean, maintainable, production-quality code in fast-moving, ambiguous environments.
Experience:Production machine learningLLMsDistributed systemsML infrastructureAI/ML
Skills:OwnershipExperimentationContinuous improvementReliability focusDebugging/regression diagnosisCollaboration
Tech Stack:PythonPyTorchJAXVLLMSGLangTensorRT-LLMLLMCloud infrastructureDistributed systemsML/data pipelinesWorkflow orchestrationGPU infrastructureVector databasesRetrieval infrastructure

Company Brief

BJAK
Bjak is a Malaysia-based insurtech that operates an online insurance comparison and distribution platform across SEA, simplifying purchase and servicing of motor and life insurance using digital tools and AI-enabled services.
Industry: InsurTech
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Selangor, Malaysia
Founded: 2019
Glassdoor
Glassdoor: 2.5
WebsiteLinkedInGlassdoor