Site Reliability Engineer - Observability

Adobe Systems
Bucharest
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5-8 yearsSkills: ["Go","Python","OpenTelemetry","Docker","Kubernetes","Splunk","ClickHouse","Loki","Elasticsearch","AWS","Azure","Grafana","Cortex","Tempo"]

Join a globally diverse team building and maintaining Adobe’s Observability platform. You’ll help shape the observability strategy, own high-impact problems across distributed cloud deployments, design large-scale logging infrastructure, and collaborate across engineering teams. You’ll work with OpenTelemetry, Kubernetes, and AI-enabled tooling to improve performance, reliability, and cost efficiency at scale, while leading incident response for a multi-tool observability stack.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Adobe Systems
Adobe Systems
3 months ago

Site Reliability Engineer - Observability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Join a globally diverse team building and maintaining Adobe’s Observability platform. You’ll help shape the observability strategy, own high-impact problems across distributed cloud deployments, design large-scale logging infrastructure, and collaborate across engineering teams. You’ll work with OpenTelemetry, Kubernetes, and AI-enabled tooling to improve performance, reliability, and cost efficiency at scale, while leading incident response for a multi-tool observability stack.
Location: Bucharest
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own and drive the design and implementation of Adobe’s observability platform and related tooling.
  • •Lead resolution of high-complexity performance and reliability issues, ensuring fault tolerance and scalability.
  • •Define service level objectives (SLOs) and SLIs to translate platform health into measurable quality indicators.
  • •Collaborate with customer engineering teams to optimize log ingestion, reduce unnecessary log volume, and optimize costs.
  • •Integrate AI workflows into large-scale deployments to surface actionable insights from log datasets and automate user interactions.

Key Requirements

  • •5-8+ years of production-level experience with distributed applications at scale in public and/or private cloud
  • •Proven experience designing and contributing to the architecture of large-scale Observability platforms
  • •Deep hands-on experience with internally hosted logging systems such as Splunk, ClickHouse, Loki, or Elastic; track record of improving environment performance, stability, and cost efficiency at scale
  • •Experience with OpenTelemetry — including collector configuration, pipelines, and instrumentation — as a core requirement given Adobe’s OTel-native observability strategy
  • •Strong programming skills in Go and/or Python; experience building production-grade integrations and applications for large-scale Observability environments
Experience:5-8 yearsObservabilityDistributed systemsCloudDevOps
Skills:GoPythonOpenTelemetryDockerKubernetesSplunkClickHouseLokiElasticsearchAWSAzureGrafanaCortexTempo
Languages:English
Tech Stack:GoPythonOpenTelemetryDockerKubernetesSplunkClickHouseLokiElasticsearchAWSAzureGrafanaCortexTempo

Company Brief

Adobe Systems
Provides creative, marketing, and document management software and cloud services, including Photoshop, Illustrator, Acrobat, and the Adobe Experience Cloud, serving creative professionals, enterprises, and governments worldwide.
Industry: SaaS
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 1982
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor