Engineer/Sr Engineer, IT Site Reliability (Fort Worth, TX, US)

American Airlines Group
United States
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 4+ yearsSkills: ["Communication","Teamwork"]

Design and deliver end-to-end monitoring and reliability capabilities (logging, metrics, and tracing) for applications and infrastructure, partnering with development and operations teams to improve uptime and incident response. Manage physical and virtual infrastructure, administrate SQL Server for backup/restore and failovers, and troubleshoot production issues. Build CI/CD automation with DevOps tools, facilitate incident management and post-incident reviews, and apply SRE best practices using cloud and container technologies.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
American Airlines Group
American Airlines Group
3 days ago

Engineer/Sr Engineer, IT Site Reliability (Fort Worth, TX, US)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live
Reposted: similar role first listed 5 months ago

Job Summary

Design and deliver end-to-end monitoring and reliability capabilities (logging, metrics, and tracing) for applications and infrastructure, partnering with development and operations teams to improve uptime and incident response. Manage physical and virtual infrastructure, administrate SQL Server for backup/restore and failovers, and troubleshoot production issues. Build CI/CD automation with DevOps tools, facilitate incident management and post-incident reviews, and apply SRE best practices using cloud and container technologies.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build end-to-end monitoring infrastructure using logging, metrics, and tracing, and collaborate with product teams on reliability tooling.
  • •Collaborate with development and operations teams to ensure application, hardware, and infrastructure availability and reliability.
  • •Manage physical servers, virtual machines, network equipment, and hardware control systems, including autonomous mobile robots and autonomous guided vehicles.
  • •Administer SQL Server instances for backups, restores, data purges, and failovers.
  • •Handle live production incidents, debug/troubleshoot issues, and implement SRE best practices including CI/CD automation, incident management, and post-incident remediation.

Pay and Benefits

Perks:Health InsuranceDentalVision401kPet Insurance

Key Requirements

  • •4 years of experience in software engineering, SRE or performance engineering.
  • •2 years of experience in Azure cloud architecture, networking, security and administration.
  • •Expertise in Terraform and CI/CD tools such as Jenkins and GitHub.
  • •Experience with SQL Server (backups, restores, data purges, failovers) and Mongo databases.
  • •Hands-on monitoring/logging experience using tools such as DynaTrace, Mezmo, LogInsight, and ThousandEyes.
Experience:4+ yearsSRECloudMonitoringAzureDevOpsAirline industrySupply chain
Skills:CommunicationTeamwork
Tech Stack:AzureTerraformCI/CDJenkinsGitHubEvent HubSQL ServerMongoDBDynaTraceMezmoLogInsightThousandEyesKubernetesKafkaLoggingMetricsTracingDevOps

Company Brief

American Airlines Group
Major U.S. airline providing scheduled passenger and cargo air transportation globally. Operates a large fleet and extensive domestic and international route network, serving millions of customers with diversified travel and loyalty services.
Industry: Airlines
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Fort Worth, United States
Founded: 1926
WebsiteLinkedIn