Site Reliability Engineering (SRE) - Maryland, US

NTT
Baltimore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Cross-functional collaboration","Incident management","Proactive troubleshooting","Mentoring","Security mindset"]

Design, build, and operate highly reliable platform solutions for mission-critical services, blending SRE, platform engineering, and cloud architecture. Own production support and incident response for UPS.com via Akamai, manage Akamai security/WAF updates, and lead CDN/performance optimizations. Build cloud infrastructure in GCP/Azure, implement Kubernetes operations and Infrastructure as Code (Terraform), and improve observability, CI/CD, and incident practices (SLOs/SLIs/error budgets).

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NTT
NTT
1 month ago

Site Reliability Engineering (SRE) - Maryland, US

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Design, build, and operate highly reliable platform solutions for mission-critical services, blending SRE, platform engineering, and cloud architecture. Own production support and incident response for UPS.com via Akamai, manage Akamai security/WAF updates, and lead CDN/performance optimizations. Build cloud infrastructure in GCP/Azure, implement Kubernetes operations and Infrastructure as Code (Terraform), and improve observability, CI/CD, and incident practices (SLOs/SLIs/error budgets).
Location: Baltimore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Provide extensive production support for UPS.com via Akamai, including on-call support, incident resolution, and troubleshooting latency and API issues.
  • •Lead Akamai security and WAF updates (rule tuning, bot mitigation, and stricter policy enforcement) and expand Akamai Malware Protection across applications.
  • •Manage certificate lifecycle and DigiCert migrations/automation, including renewals and certificate-related issue resolution.
  • •Design and operate highly available, scalable cloud infrastructure in GCP and Azure, including Kubernetes environment management (GKE/OpenShift), service mesh, and reliability practices (SLOs/SLIs/error budgets).
  • •Build and evolve developer platform and automation by implementing Infrastructure as Code (Terraform), CI/CD and GitOps (Argo CD/Azure Pipelines), and observability (Prometheus/Grafana/Dynatrace) with AI/Python/scripting and operational automation.

Key Requirements

  • •3-5+ years Akamai experience (edge and CDN services).
  • •5+ years experience in SRE, DevOps, or cloud engineering roles.
  • •Strong expertise in Kubernetes and container platforms.
  • •Experience implementing Infrastructure as Code with Terraform and automation.
  • •Proficiency in Python or Bash (or similar scripting) and experience with CI/CD and GitOps methods.
Experience:CloudSREDevOpsAkamaiEdge/CDNKubernetesGitOpsCI/CD
Skills:Cross-functional collaborationIncident managementProactive troubleshootingMentoringSecurity mindset
Tech Stack:GCPGoogle Cloud PlatformOpenShiftOpenShift (OCP4)GKEKubernetesKubernetes administrationOpenTelemetryPrometheusGrafanaDynatraceDynatrace operatorLinuxDockerTerraformConfig ConnectorArgo CDArgo WorkflowsAzure PipelinesAzure Pipelines (CI/CD)

Company Brief

NTT
Dimension Data, operating under NTT Ltd, provides managed IT services, cloud and data center solutions, networking, cybersecurity, and digital transformation services to enterprise customers worldwide.
Industry: Professional Services
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Valuation: Public Company (Market Cap in USD)
Headquarters: London, United Kingdom
Founded: 1983
WebsiteLinkedIn