Site Reliability Engineer (Onsite Hybrid)

NTT
Plano
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Incident response","Root-cause analysis","Troubleshooting","Alerting strategy"]

Own end-to-end reliability for production systems by driving observability with New Relic (APM, dashboards, alerting), defining SLIs/SLOs, and leading incident response with root-cause analysis and post-mortems. Administer GitHub Enterprise and build CI/CD reliability for Java/.NET applications using GitHub Actions. Troubleshoot across application, infrastructure, and pipeline layers, continuously improving performance and leveraging AI/automation in SRE workflows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NTT
NTT
2 months ago

Site Reliability Engineer (Onsite Hybrid)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 22 hours agoStatus: Live

Job Summary

Own end-to-end reliability for production systems by driving observability with New Relic (APM, dashboards, alerting), defining SLIs/SLOs, and leading incident response with root-cause analysis and post-mortems. Administer GitHub Enterprise and build CI/CD reliability for Java/.NET applications using GitHub Actions. Troubleshoot across application, infrastructure, and pipeline layers, continuously improving performance and leveraging AI/automation in SRE workflows.
Location: Plano
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own observability using New Relic, including APM, infrastructure monitoring, dashboards, and alerting.
  • •Define and implement SLIs/SLOs and alerting strategies.
  • •Drive incident response, root-cause analysis (RCA), and post-mortems.
  • •Administer GitHub Enterprise and design GitHub Actions CI/CD pipelines for Java/.NET applications.
  • •Troubleshoot across application, infrastructure, and CI/CD layers while driving continuous reliability and performance improvements.

Key Requirements

  • •5+ years hands-on experience with New Relic (or similar APM tools).
  • •5+ years strong understanding of SRE practices including SLIs/SLOs, alerting, and incident management.
  • •5+ years experience with GitHub Enterprise and GitHub Actions.
  • •5+ years CI/CD pipeline experience for Java or .NET applications.
  • •5+ years troubleshooting and root-cause analysis experience in production environments.
Experience:5+ years
Skills:Incident responseRoot-cause analysisTroubleshootingAlerting strategy
Tech Stack:New RelicSLIs/SLOsGitHub EnterpriseGitHub ActionsCI/CDJava.NETCopilotServiceNow ITOMJfrog ArtifactoryXraySonarQubeGitHub Advanced SecurityCodeQLDependabotAngularDatabricksSQL

Company Brief

NTT
Dimension Data, operating under NTT Ltd, provides managed IT services, cloud and data center solutions, networking, cybersecurity, and digital transformation services to enterprise customers worldwide.
Industry: Professional Services
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Valuation: Public Company (Market Cap in USD)
Headquarters: London, United Kingdom
Founded: 1983
WebsiteLinkedIn