Sr. IT Site Reliability Software Engineer

Databricks
Costa Rica
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Problem-solving","Communication","Ownership","Collaboration"]

Bridge software engineering and systems architecture as a core member of IT Infrastructure, delivering resilient, automated cloud infrastructure with cost optimization, security, and high availability. Build and maintain production-grade infrastructure with IaC (Terraform/Pulumi) across AWS/Azure/GCP, IAM, observability, and scalable deployment pipelines, while leading incident response and cross-functional collaboration.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Databricks
Databricks
3 months ago

Sr. IT Site Reliability Software Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Bridge software engineering and systems architecture as a core member of IT Infrastructure, delivering resilient, automated cloud infrastructure with cost optimization, security, and high availability. Build and maintain production-grade infrastructure with IaC (Terraform/Pulumi) across AWS/Azure/GCP, IAM, observability, and scalable deployment pipelines, while leading incident response and cross-functional collaboration.
Location: Costa Rica
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Architect and automate production-grade infrastructure on cloud platforms using IaC tools (Terraform or Pulumi).
  • •Improve reliability and performance by optimizing architecture and scaling of IT services for maximum uptime and minimal latency.
  • •Design and maintain robust CI/CD pipelines using GitHub Actions and manage hosted/self-hosted runners.
  • •Build and enforce observability by ensuring logging, metrics, and alerts are enabled by default across new applications.
  • •Lead incident response with on-call rotation, post-mortems, and preventive engineering to reduce recurrence.

Key Requirements

  • •5+ years of production-level software engineering experience with strong Python
  • •Expert-level proficiency in Infrastructure as Code (Terraform or Pulumi)
  • •Hands-on experience with cloud platforms (AWS, Azure, or GCP) and containerization (Kubernetes, Docker)
  • •Strong observability mindset with experience using Datadog, Prometheus, or ELK
  • •Proficiency in distributed systems (e.g., Kafka or messaging queues) and advanced GitHub Actions knowledge
Experience:5+ yearsIT infrastructure
Skills:Problem-solvingCommunicationOwnershipCollaboration
Languages:English
Tech Stack:PythonTerraformPulumiAWSAzureGCPKubernetesDockerDatadogPrometheusELKKafkaGitHub ActionsGitHub Runners

Company Brief

Databricks
Provides a unified data analytics platform powered by Apache Spark to simplify building, deploying, and scaling data engineering, data science, and machine learning workloads for enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn