Site Reliability Engineer

Thales
Noida
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Analytical","Troubleshooting","Communication","Ownership","Customer-oriented"]

Own reliability and operational excellence for Thales PAY Digital cloud services, ensuring high availability, performance, and scalability aligned to SLAs/SLOs. Lead incident response, troubleshooting, and RCA/postmortems, and drive long-term corrective actions. Build and maintain infrastructure as code and automation (Terraform, GitLab CI/CD) to support cloud deployments on AWS/GCP and Kubernetes. Improve observability with Datadog and Splunk and collaborate with product and development squads to ensure operational readiness.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thales
Thales
2 days ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live
Reposted: similar role first listed 6 months ago

Job Summary

Own reliability and operational excellence for Thales PAY Digital cloud services, ensuring high availability, performance, and scalability aligned to SLAs/SLOs. Lead incident response, troubleshooting, and RCA/postmortems, and drive long-term corrective actions. Build and maintain infrastructure as code and automation (Terraform, GitLab CI/CD) to support cloud deployments on AWS/GCP and Kubernetes. Improve observability with Datadog and Splunk and collaborate with product and development squads to ensure operational readiness.
Location: Noida
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Ensure high availability, performance, and scalability of Thales PAY Digital cloud services in line with customer SLAs/SLOs.
  • •Respond to production incidents, lead troubleshooting, coordinate recovery actions, and run post-incident RCA/postmortems with corrective and preventive measures.
  • •Design, develop, and maintain Infrastructure as Code and automation solutions (Terraform, GitLab CI/CD, scripting) to support reliable infrastructure.
  • •Provide support for cloud deployments of Thales PAY products on AWS/GCP/Kubernetes and validate operational readiness.
  • •Implement and improve monitoring, alerting, and observability (metrics, logs, traces) using tools like Datadog and Splunk, and support change/release management with risk assessment.

Key Requirements

  • •Experience in systems, networking, and security operations.
  • •Hands-on production experience with Kubernetes and AWS and/or GCP.
  • •Strong experience with CI/CD pipelines and Infrastructure as Code using GitLab CI and Terraform.
  • •Proficiency in Linux and core networking and web protocols (TCP/IP, HTTP/HTTPS) for distributed systems.
  • •Experience with monitoring and observability tools such as Datadog and Splunk, plus scripting skills in Shell and/or Python.
Skills:AnalyticalTroubleshootingCommunicationOwnershipCustomer-oriented
Tech Stack:TerraformGitLab CI/CDAWSGCPKubernetesLinuxTCP/IPHTTP/HTTPSDatadogSplunkShellPython

Company Brief

Thales
Designs and delivers advanced systems and services for aerospace, defence, security, and digital identity and cybersecurity markets, serving government and commercial customers worldwide.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Paris, France
Founded: 2000
WebsiteLinkedIn