Staff Site Reliability Engineer, Playout

NBC Universal
United States
Workplace: HybridFull timeUSD 145,000 - 175,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Leadership","Communication","Incident management","Troubleshooting","Adaptability"]

Lead reliability engineering for cloud-based master control playout systems that support NBCUniversal’s live linear channels. Define SLI/SLOs, service availability targets, and operational readiness criteria; manage major incidents and drive post-incident improvements. Partner with engineering, product, and operations on capacity planning, performance tuning, and resilience testing. Build automation to reduce toil, and create monitoring dashboards and alerting using Grafana and tools like Slack/ServiceNow, while providing L1/L2 on-call support.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NBC Universal
NBC Universal
1 day ago

Staff Site Reliability Engineer, Playout

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Lead reliability engineering for cloud-based master control playout systems that support NBCUniversal’s live linear channels. Define SLI/SLOs, service availability targets, and operational readiness criteria; manage major incidents and drive post-incident improvements. Partner with engineering, product, and operations on capacity planning, performance tuning, and resilience testing. Build automation to reduce toil, and create monitoring dashboards and alerting using Grafana and tools like Slack/ServiceNow, while providing L1/L2 on-call support.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Team lead for Site Reliability Engineers on the Playout Engineering team.
  • •Define and manage reliability targets (SLIs/SLOs) and operational readiness criteria for playout services.
  • •Drive incident response including establishing on-call practices, leading major incident management, and ensuring measurable improvements from post-incident reviews.
  • •Partner with engineering, product, and operations on reliability improvements via capacity planning, performance tuning, and resilience testing.
  • •Create monitoring dashboards and alerts (Grafana; teams/slack/ServiceNow) and provide L1/L2 support for playout infrastructure with after-hours on-call support select weeks.

Pay and Benefits

Salary: USD 145,000 - 175,000 annually

Key Requirements

  • •Bachelor’s degree in computer science or related field, or equivalent experience.
  • •8 years of hands-on engineering experience in broadcast automation playout environments (e.g., Snell, Harris, Imagine, Amagi).
  • •Requires on-call 24/7 availability for escalations.
  • •Hands-on experience administering Linux environments.
  • •Experience with monitoring/logging tools (e.g., Splunk and Grafana) and broadcast playout/master control systems technologies.
Experience:8+ yearsBroadcast automationCloudStreaming mediaLive linear channels
Education:Bachelor's in computer science
Skills:LeadershipCommunicationIncident managementTroubleshootingAdaptability
Languages:English
Tech Stack:LinuxSplunkGrafanaTeamsSlackServiceNowAWSDockerKubernetesTSHEVCH.264HLSCMAFSCTE-35SCTE-224ESAMSRTRISTIP networking

Company Brief

NBC Universal
NBCUniversal is a global media and entertainment company producing and distributing film, television, news, sports and streaming content, and operating theme parks and consumer experiences across a portfolio of well-known brands including NBC, Universal Pictures and Peacock.
Industry: Film & Television
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: New York City, United States
Founded: 2004
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor