Site Reliability Engineer

Thales
Austin
Workplace: HybridFull timeUSD 112,105.5 - 190,664 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Collaboration","Communication","Documentation","Cross-functional work","Incident response"]

Design and build scalable tools and infrastructure for a globally distributed platform, partnering with North American teams and internal stakeholders. Apply SRE measurement (SLI/SLO/SLA), eliminate toil, establish reliability metrics, and evolve SLO/SLI baselines. Support and optimize infrastructure programmatically to improve availability, reliability, performance, and security, participating in incident response, RCA, and 24x7 on-call rotation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Thales
Thales
2 days ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live
Reposted: similar role first listed 4 months ago

Job Summary

Design and build scalable tools and infrastructure for a globally distributed platform, partnering with North American teams and internal stakeholders. Apply SRE measurement (SLI/SLO/SLA), eliminate toil, establish reliability metrics, and evolve SLO/SLI baselines. Support and optimize infrastructure programmatically to improve availability, reliability, performance, and security, participating in incident response, RCA, and 24x7 on-call rotation.
Location: Austin
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and build tools and infrastructure that support internal teams and customers, scaling with platform and customer growth.
  • •Apply SRE core tenets using measurement (SLI/SLO/SLA), reliability modeling, and toil elimination to improve operational excellence.
  • •Establish and evolve metrics and baselines for availability, reliability, and velocity, and perform proactive data analysis to ensure optimal production operation.
  • •Plan and validate go/no-go readiness for existing and new product/services and troubleshoot business-impacting issues with internal customers.
  • •Participate in escalations, incident response, RCA and postmortems, and take part in a 24x7 on-call rotation.
Travel: Low travel

Pay and Benefits

Salary: USD 112,105.5 - 190,664 annually
Perks:Health InsuranceDentalVision401kPaid Leave

Key Requirements

  • •At least 5 years of professional experience in cloud/web/CDN scale infrastructure.
  • •Experience with Python and Go; C/C++ is a plus.
  • •Expert knowledge of Linux systems, network programming, and protocols TCP, UDP, DNS, TLS/SSL, and HTTP.
  • •Experience with DevOps principles including Infrastructure as Code (Ansible/Saltstack) and CI/CD (Gitlab, Jenkins, Git), plus monitoring and visualization (Prometheus, Grafana).
  • •Experience with containers and distributed systems at scale (Docker, Kubernetes) and data/telemetry (NoSQL/RDBMS, Redis, ElasticSearch, Kafka; pipelines and telemetry).
Experience:5+ yearsCloud infrastructureWebCDNDistributed systemsDevOpsBig data
Education:Bachelor's
Skills:CollaborationCommunicationDocumentationCross-functional workIncident response
Tech Stack:PythonGoC/C++LinuxTCPUDPDNSTLS/SSLHTTPBGPAnycastAnsibleSaltstackCI/CDGitlabJenkinsGitPrometheusGrafanaNoSQL

Company Brief

Thales
Designs and delivers advanced systems and services for aerospace, defence, security, and digital identity and cybersecurity markets, serving government and commercial customers worldwide.
Industry: Defense Technology
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Paris, France
Founded: 2000
WebsiteLinkedIn