Senior Software Engineer, AI Operations, GPS

Scale AI
Doha
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 5+ yearsSkills: ["Client communication","Stakeholder management","Root cause analysis","Boundary control"]

Own the long-term technical health, performance, and stability of AI solutions deployed for public sector partners. Bridge software engineering, MLOps, and client governance by managing tiered SLAs, driving incident governance and RCA, monitoring latency and model/data drift, and maintaining prompt/config repositories. Engineer automation for reliability (self-healing pipelines, RAG indexing syncs, telemetry) and serve as the senior technical interface for government and enterprise IT stakeholders.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Scale AI
Scale AI
1 day ago

Senior Software Engineer, AI Operations, GPS

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Own the long-term technical health, performance, and stability of AI solutions deployed for public sector partners. Bridge software engineering, MLOps, and client governance by managing tiered SLAs, driving incident governance and RCA, monitoring latency and model/data drift, and maintaining prompt/config repositories. Engineer automation for reliability (self-healing pipelines, RAG indexing syncs, telemetry) and serve as the senior technical interface for government and enterprise IT stakeholders.
Location: Doha
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Act as the technical gatekeeper during delivery-to-maintenance handover, reviewing baseline code, prompts, and architecture for maintainability and documentation before sign-off.
  • •Own tiered SLA and incident management, including incident governance, root cause analysis, and P1/P2 mitigation within active support windows.
  • •Govern the AI lifecycle by monitoring production model performance, latency, and data drift; maintain prompt configuration repositories; and run regression testing when LLM providers update endpoints.
  • •Operationalize boundaries between routine maintenance and system evolution by classifying incoming client requests and benchmarking new AI models.
  • •Engineer automation to reduce operational toil, including self-healing data pipelines, automated RAG indexing syncs, and telemetry tooling, while influencing delivery teams to adopt maintainable patterns.

Key Requirements

  • •5+ years in software engineering, MLOps, SRE, or forward deployed engineering in heavy data or production AI environments.
  • •Advanced proficiency in Python, SQL, and REST/gRPC APIs, with cloud architecture experience (AWS, Azure, or GCP).
  • •Hands-on experience with MLOps tooling, vector databases, and LLM orchestration frameworks such as LangChain or LlamaIndex.
  • •Practical AI governance skills including prompt version control, model benchmarking, RAG pipeline mechanics, and data drift detection.
  • •Strong understanding of CI/CD for machine learning pipelines and an engineering mindset focused on systematic, automated fixes.
Experience:5+ yearsProduction AIMLOpsSRE
Skills:Client communicationStakeholder managementRoot cause analysisBoundary control
Languages:English
Tech Stack:PythonSQLRESTGRPCAWSAzureGCPMLOpsVector databasesLangChainLlamaIndexCI/CDRAGTelemetryPrompt versioningLLM orchestration

Company Brief

Scale AI
Provides data labeling, annotation, and infrastructure services to accelerate machine learning and AI development. Supplies high-quality training data, tooling, and APIs for customers in autonomous vehicles, mapping, robotics, and enterprise AI applications.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn