SRE COMPUTE

Thales
France
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureEducation: mastersSkills: ["Operational excellence","Incident management","Automation","Documentation","Continuous improvement"]

Maintain and operate the Compute infrastructure at the heart of a sovereign GCP environment. You will run 24/7 incident handling for GCE, GKE, and Vertex AI, manage compute fleets and control planes, and improve reliability through monitoring (SLI/SLO), automation, and standardized playbooks. Work with global GCP experts, lead post-incident reviews, and support continuous operational improvement within a SecNumCloud-qualified service environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thales
Thales
1 day ago

SRE COMPUTE

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Maintain and operate the Compute infrastructure at the heart of a sovereign GCP environment. You will run 24/7 incident handling for GCE, GKE, and Vertex AI, manage compute fleets and control planes, and improve reliability through monitoring (SLI/SLO), automation, and standardized playbooks. Work with global GCP experts, lead post-incident reviews, and support continuous operational improvement within a SecNumCloud-qualified service environment.
Location: France
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Operate and maintain the full Compute platforms in the sovereign GCP environment (including GCE, GKE, and Vertex AI).
  • •Handle production incidents 24/7 and provide on-call coverage (day coverage plus rotating night on-call).
  • •Monitor SLI/SLO for availability, scalability, latency, and efficiency of sovereign GCP compute services.
  • •Manage compute infrastructure and underlying machine fleets/control planes (Borg and Google control planes) to ensure optimal performance.
  • •Run post-incident reviews (post-mortems), standardize resolution flows, and improve operational playbooks through documentation and automation.

Key Requirements

  • •Engineering school graduate or Master’s degree.
  • •At least 3 years of SRE experience with operations automation (juniors may be considered with significant experience from alternance).
  • •Strong interest in cloud and operating services/infrastructure with “as code”.
  • •Experience delivering operational excellence at large scale, including high availability (≥ 99,99%) for critical databases and high-volume analytics engines.
  • •Good English level and exposure to an international environment; prior GCP experience is a plus.
Experience:CloudSREInternational
Education:Master's
Skills:Operational excellenceIncident managementAutomationDocumentationContinuous improvement
Languages:English
Tech Stack:GCPGCEGKEVertex AIBorgColossusSpannerSawmillSLI/SLOAs codeControl planeDistributed systemsSecNumCloudANSSIGCP compute

Company Brief

Thales
Designs and delivers advanced systems and services for aerospace, defence, security, and digital identity and cybersecurity markets, serving government and commercial customers worldwide.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Paris, France
Founded: 2000
WebsiteLinkedIn