AI Infrastructure Engineer III

Mozn
Cairo
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 4-6 yearsSkills: ["Hands-on","Troubleshooting","Automation","Monitoring","Collaboration"]

Design, deploy, and operate enterprise AI/ML infrastructure for scalable model development and deployment. Build self-service Kubernetes-based AI platform capabilities using Kubeflow, MLflow, KServe, Ray, and MLOps tooling across cloud-native and hybrid environments. Own GPU cluster operations, distributed training performance, and highly available model serving. Improve platform reliability with CI/CD, Infrastructure as Code, observability, and automation, collaborating closely with data science teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Mozn
Mozn
3 weeks ago

AI Infrastructure Engineer III

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 minute agoStatus: Live

Job Summary

Design, deploy, and operate enterprise AI/ML infrastructure for scalable model development and deployment. Build self-service Kubernetes-based AI platform capabilities using Kubeflow, MLflow, KServe, Ray, and MLOps tooling across cloud-native and hybrid environments. Own GPU cluster operations, distributed training performance, and highly available model serving. Improve platform reliability with CI/CD, Infrastructure as Code, observability, and automation, collaborating closely with data science teams.
Location: Cairo
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, deploy, and operate enterprise AI/ML platforms, including building self-service capabilities for data scientists and ML engineers.
  • •Deploy and operate Kubeflow, MLflow, KServe, Ray (or similar AI platform technologies).
  • •Design infrastructure for model training, experimentation, feature engineering, and inference.
  • •Design and operate GPU clusters, optimizing GPU scheduling, utilization, autoscaling, and distributed training performance.
  • •Build CI/CD pipelines for ML workloads, automate provisioning via Infrastructure as Code, and implement monitoring/observability for training, serving, and latency.

Key Requirements

  • •4-6 years of experience in AI Infrastructure, MLOps, Platform Engineering, or Cloud Engineering.
  • •Hands-on experience with Kubernetes.
  • •Experience with Kubeflow, MLflow, or similar ML platform technologies.
  • •Experience operating GPU infrastructure for AI workloads, including NVIDIA GPU technologies and CUDA fundamentals.
  • •Experience with model serving platforms such as KServe, Triton Inference Server, Ray Serve, or similar.
Experience:4-6 yearsAI infrastructureMLOpsPlatform engineeringCloud engineering
Skills:Hands-onTroubleshootingAutomationMonitoringCollaboration
Tech Stack:KubernetesKubeflowMLflowKServeRayTriton Inference ServerRay ServeGPU OperatorNVIDIACUDAAWSGCPOCIAzureTerraformHelmGitOpsAnsiblePythonBash

Company Brief

Mozn
Builds enterprise AI products and data platforms that help organizations automate decision-making, improve risk management, and extract insights from large datasets. The company focuses on applied machine learning for regulated and data-intensive industries.
Industry: AI & Machine Learning
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Riyadh, Saudi Arabia
Founded: 2017
WebsiteLinkedIn