Site Reliability Engineer

Thales
Austin
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Incident management","Observability","Automation","Reliability engineering","Continuous improvement"]

Build and operate a product-specific SRE team to deliver high availability and strong performance for a public-cloud telecommunications solution. Own service reliability by defining SLOs/SLIs and error budgets, running 24/7 on-call and deep-dive incident troubleshooting, and driving CI/CD and infrastructure automation with Terraform, Ansible, Kubernetes, and GitLab. Implement observability with Datadog, lead postmortems, and collaborate with cloud security to improve access controls and respond to vulnerabilities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thales
Thales
16 hours ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live
Reposted: similar role first listed 4 months ago

Job Summary

Build and operate a product-specific SRE team to deliver high availability and strong performance for a public-cloud telecommunications solution. Own service reliability by defining SLOs/SLIs and error budgets, running 24/7 on-call and deep-dive incident troubleshooting, and driving CI/CD and infrastructure automation with Terraform, Ansible, Kubernetes, and GitLab. Implement observability with Datadog, lead postmortems, and collaborate with cloud security to improve access controls and respond to vulnerabilities.
Location: Austin
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and maintain scalable infrastructure using Terraform, Ansible, and Kubernetes, and develop automated CI/CD pipelines via GitLab.
  • •Define and monitor SLOs/SLIs and manage error budgets to balance feature velocity with platform stability.
  • •Participate in 24/7 on-call rotations for emergency response and deep-dive troubleshooting of production incidents.
  • •Implement and refine symptom-based alerting and comprehensive monitoring strategies using platforms like Datadog to ensure visibility into system health.
  • •Lead blameless postmortems after incidents and collaborate with Cloud Security teams on security best practices, access controls, and response to breaches or vulnerabilities.

Pay and Benefits

Perks:Health InsuranceDentalVisionRetirementPaid Leave

Key Requirements

  • •At least 5 years of experience in a relevant role.
  • •Engineer or equivalent education.
  • •Java development experience.
  • •Experience with public cloud (GCP, AWS), containers and microservices (Docker, Kubernetes), and CI/CD automation.
  • •Familiarity with NoSQL databases and operational tooling (e.g., Jenkins/GitLab, Helm, monitoring/observability).
Experience:5+ years
Skills:Incident managementObservabilityAutomationReliability engineeringContinuous improvement
Certifications:GCP cloud architect certificationPublic cloud architect certification
Tech Stack:TerraformAnsibleKubernetesGitLabCI/CDJenkinsHelmDockerDatadogGCPAWSPublic cloudNoSQL

Eligibility

Security Clearance:CFIUS clearanceDepartment of Treasury post-hire clearance

Company Brief

Thales
Designs and delivers advanced systems and services for aerospace, defence, security, and digital identity and cybersecurity markets, serving government and commercial customers worldwide.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Paris, France
Founded: 2000
WebsiteLinkedIn