Principal Site Reliability Engineer, Infrastructure Observability

T. Rowe Price
Maryland
Workplace: HybridFull timeUSD 159,000 - 339,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Blameless post-mortems","Incident analysis","Automation mindset","Cross-functional collaboration","Strategic decision-making"]

Lead Site Reliability Engineering focused on infrastructure observability, reliability, and recovery across T. Rowe Price’s cloud and on-prem platforms. You’ll design solutions to prevent service disruptions, automate incident prevention and remediation, and drive SRE standard methodologies through blameless post-mortems and ongoing incident trend analysis. Partner with operations and engineering teams to standardize dashboards, define SLOs/SLIs and error budgets, and help shape target-state architecture for scalable, measurable reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
T. Rowe Price
T. Rowe Price
1 month ago

Principal Site Reliability Engineer, Infrastructure Observability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 days agoStatus: Live

Job Summary

Lead Site Reliability Engineering focused on infrastructure observability, reliability, and recovery across T. Rowe Price’s cloud and on-prem platforms. You’ll design solutions to prevent service disruptions, automate incident prevention and remediation, and drive SRE standard methodologies through blameless post-mortems and ongoing incident trend analysis. Partner with operations and engineering teams to standardize dashboards, define SLOs/SLIs and error budgets, and help shape target-state architecture for scalable, measurable reliability.
Location: Maryland
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Formulate, develop, and implement an SRE team focused on observability, sustainability, scalability, measurability, and recoverability across cloud and on-prem solutions.
  • •Design technology solutions to prevent or minimize service disruptions, including automations to prevent outages and incidents.
  • •Foster reliability improvements through blameless post-mortems and culture of deep learning across services.
  • •Analyze incidents impacting technology availability to identify high-level trends and drive initiatives to reduce technology failures.
  • •Help transform operations teams by adopting SRE standard methodologies, standardizing dashboards/observability tooling, and contributing to target-state architecture.

Pay and Benefits

Salary: USD 159,000 - 339,000 annually
Perks:Annual BonusRetirementHybrid WorkPaid Leave

Key Requirements

  • •10+ years designing and operating cloud infrastructure with senior-level impact, with a Bachelor’s degree or equivalent experience.
  • •5+ years building and supporting solutions in Amazon AWS.
  • •5+ years building and running a DevOps and/or SRE function.
  • •Experience implementing and operating the chaos model at scale and driving strategic, program-level implementation.
  • •Proficiency with programming languages and automation; experience with observability tools, SLOs/SLIs, and error budgets for incident prevention and reliability improvements.
Experience:10+ yearsCloud infrastructureDevOpsSRE
Education:Bachelor's
Skills:Blameless post-mortemsIncident analysisAutomation mindsetCross-functional collaborationStrategic decision-making
Tech Stack:Amazon AWSDevOpsSRECI/CDPythonJavaGONode.jsDotnet Core.NET CoreSQL ServerPostgreSQLMySQLService Level Objectives (SLOs)Service Level Indicators (SLIs)Error BudgetsNew RelicSolarWinds DPAElastic StackPrometheus

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

T. Rowe Price
Global investment management firm offering mutual funds, retirement services, and advisory solutions for individual and institutional investors. Manages assets across equities, fixed income, multi-asset, and alternative strategies.
Industry: Asset Management
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Baltimore, United States
Founded: 1937
WebsiteLinkedIn