MLOps Engineer

AI71
Abu Dhabi
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsSkills: ["Mentorship","Communication","Stakeholder management","Decision-making","Reliability mindset"]

Own MLOps architecture and reliability strategy across AI71’s platform, including deploying, fine-tuning, and serving LLMs and deep learning models at scale. Define ML infrastructure across SaaS and fully air-gapped on-prem environments, set reliability targets and incident response practices, and mentor engineers across teams. Drive initiatives to improve inference performance and cost-efficiency using distributed training frameworks, while partnering with ML researchers and engineering leadership on multi-quarter infrastructure plans.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AI71
AI71
23 hours ago

MLOps Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live
Reposted: similar role first listed 4 months ago

Job Summary

Own MLOps architecture and reliability strategy across AI71’s platform, including deploying, fine-tuning, and serving LLMs and deep learning models at scale. Define ML infrastructure across SaaS and fully air-gapped on-prem environments, set reliability targets and incident response practices, and mentor engineers across teams. Drive initiatives to improve inference performance and cost-efficiency using distributed training frameworks, while partnering with ML researchers and engineering leadership on multi-quarter infrastructure plans.
Location: Abu Dhabi
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Define ML infrastructure architecture across the platform, including model deployment strategy, pipeline engineering, and cloud-native infrastructure across major clouds.
  • •Set direction for ML system reliability, including monitoring, latency/throughput/availability targets, and incident response across research and production.
  • •Mentor senior MLOps engineers and raise the operational bar across multiple teams.
  • •Drive cross-team initiatives to improve inference performance and cost-efficiency using distributed training frameworks.
  • •Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments while partnering with ML researchers and engineering leadership.

Key Requirements

  • •10+ years of experience in MLOps, ML infrastructure, or machine learning engineering with architectural ownership history.
  • •Architect large-scale model deployment (including LLMs) and ML infrastructure at production scale.
  • •Deep cloud expertise across AWS, Azure, or GCP with strong Python proficiency.
  • •Demonstrated mentorship track record, with engineers you’ve grown operating independently at higher levels.
  • •Comfort architecting ML systems for both managed SaaS and on-premises/disconnected air-gapped environments.
Experience:10+ yearsAIMachine learningLLMsMLOpsDistributed trainingInferenceCloud infrastructureOn-premises deployments
Skills:MentorshipCommunicationStakeholder managementDecision-makingReliability mindset
Languages:Arabic
Tech Stack:VLLMTritonTGIMLflowKubeflowAWSAzureGCPKubernetesDeepSpeedFSDPAcceleratePythonCUDANCCLNVLinkInfiniBandRoCEMegatron-LMCUDA kernel

Company Brief

AI71
AI71 builds generative AI and machine learning products and solutions, focusing on applying large language models and conversational AI to automate workflows, enhance decision-making, and deliver domain-specific AI applications for enterprise customers.
Industry: AI & Machine Learning
Website