Sr Platform Monitoring Engineer

Databricks
California
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 6+ yearsEducation: bachelorsSkills: ["Incident response","Root cause analysis","Customer obsession","Cross-functional collaboration","Mentorship"]

Lead platform observability and proactive monitoring for the Databricks Platform. Own incident response as a first responder: drive complex investigations, coordinate cross-functional mitigation, and perform thorough post-incident root-cause analyses. Build customer-focused alerting and end-to-end observability workflows, develop automation and reusable monitoring patterns, and support on-call rotation while mentoring junior engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Databricks
Databricks
1 day ago

Sr Platform Monitoring Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Lead platform observability and proactive monitoring for the Databricks Platform. Own incident response as a first responder: drive complex investigations, coordinate cross-functional mitigation, and perform thorough post-incident root-cause analyses. Build customer-focused alerting and end-to-end observability workflows, develop automation and reusable monitoring patterns, and support on-call rotation while mentoring junior engineers.
Location: California
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead platform incident investigations and coordinate cross-functional teams to detect, mitigate, and resolve issues with minimal customer impact.
  • •Conduct thorough post-incident root cause analyses across infrastructure, services, and cloud providers to identify systemic patterns.
  • •Design and implement customer-focused alerting pipelines and end-to-end observability workflows to improve detection coverage and reduce mean time to detection.
  • •Build automation tools and establish reusable monitoring patterns to close reliability gaps affecting customer experience.
  • •Provide mentorship to junior engineers on observability patterns, alert design, and service health metrics.

Key Requirements

  • •Minimum 6 years of experience as an SRE, DevOps Engineer, Production Engineer, or similar role.
  • •Production-level experience with at least one major cloud provider (AWS, Azure, GCP).
  • •Proficiency with container and orchestration technologies (Docker, Kubernetes).
  • •Hands-on experience with monitoring/logging/alerting tools such as ELK, Prometheus, Grafana, and PagerDuty.
  • •BS, Master’s, or PhD in Computer Science or Computer Engineering, or a related engineering field.
Experience:6+ yearsSREDevOpsProduction engineering
Education:Bachelor's in Computer Science or Computer Engineering
Skills:Incident responseRoot cause analysisCustomer obsessionCross-functional collaborationMentorship
Languages:English
Tech Stack:AWSAzureGCPDockerKubernetesELKPrometheusGrafanaPagerDutyPython

Company Brief

Databricks
Provides a unified data analytics platform powered by Apache Spark to simplify building, deploying, and scaling data engineering, data science, and machine learning workloads for enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn