Senior Site Reliability Engineer

TradingView
Tbilisi
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Ownership","Analytical skills","Troubleshooting","Collaboration","Communication"]

Own the reliability of production services for a global financial analysis platform. Investigate incidents end-to-end with RCA and postmortems, drive corrective actions, and continuously improve monitoring, alerting, observability, and SLI/SLOs. Define availability and SLA compliance for frontend and backend services, maintain runbooks and operational documentation, and contribute to automation for diagnostics and incident response. Participate in on-call rotations and incident escalation reviews.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TradingView
TradingView
19 hours ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Own the reliability of production services for a global financial analysis platform. Investigate incidents end-to-end with RCA and postmortems, drive corrective actions, and continuously improve monitoring, alerting, observability, and SLI/SLOs. Define availability and SLA compliance for frontend and backend services, maintain runbooks and operational documentation, and contribute to automation for diagnostics and incident response. Participate in on-call rotations and incident escalation reviews.
Location: Tbilisi
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Investigate production incidents and drive them through resolution until all consequences are addressed.
  • •Perform RCA and participate in postmortem reviews with engineering and product teams.
  • •Track and drive corrective actions from incidents and postmortems.
  • •Develop and improve monitoring, alerting, and observability; define and maintain SLI/SLOs and analyze reliability metrics.
  • •Define and maintain runbooks and service recovery procedures; participate in incident validation, escalation reviews, and on-call rotations.

Pay and Benefits

Perks:Health InsuranceRelocationAnnual Bonus

Key Requirements

  • •Experience as an SRE, Reliability Engineer, Production Engineer, Operations Engineer, or similar role.
  • •Hands-on incident investigation and root cause analysis in production environments.
  • •Strong monitoring, alerting, and observability fundamentals using metrics, logs, and distributed tracing.
  • •Knowledge of SLA, SLI, SLO, and error budget concepts; ability to analyze reliability metrics and service health.
  • •Experience creating and maintaining operational documentation, runbooks, and service recovery procedures.
Experience:Financial servicesSaaS
Skills:OwnershipAnalytical skillsTroubleshootingCollaborationCommunication
Certifications:CKACKAD
Tech Stack:KubernetesCKACKADPrometheusGrafanaOpenTelemetryDistributed tracing

Company Brief

TradingView
Provides a web-based charting platform, market data and social network for traders and investors, offering charting libraries, trading integrations, and analytics used by retail traders, brokerages and fintech firms worldwide.
Industry: Trading Platforms
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Funding: Series C
Headquarters: London, United Kingdom
Founded: 2011
Glassdoor
Glassdoor: 2.9
WebsiteLinkedInGlassdoor