Site Reliability Engineer

Levi Strauss
Mexico
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Documentation"]

Join the Data & AI Platform Engineering team to keep production data and AI platforms reliable, efficient, and secure. You’ll monitor systems with observability tooling, respond to incidents using runbooks, and drive blameless post-mortems. Build automation to reduce operational toil, operate GCP workloads (GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer), and apply Infrastructure-as-Code with Terraform and Helm while collaborating across engineering teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Levi Strauss
Levi Strauss
2 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Join the Data & AI Platform Engineering team to keep production data and AI platforms reliable, efficient, and secure. You’ll monitor systems with observability tooling, respond to incidents using runbooks, and drive blameless post-mortems. Build automation to reduce operational toil, operate GCP workloads (GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer), and apply Infrastructure-as-Code with Terraform and Helm while collaborating across engineering teams.
Location: Mexico
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Monitor production systems using dashboards, alerts, and logs to detect and triage issues before end users are impacted.
  • •Participate in on-call rotations, respond to incidents using runbooks, and escalate appropriately when needed.
  • •Contribute to blameless post-mortems by documenting root causes and follow-up action items to prevent recurrence.
  • •Build automation and tooling to reduce operational toil and improve deployment reliability and operational efficiency.
  • •Operate and maintain GCP workloads and apply Infrastructure-as-Code practices to manage and version infrastructure changes.

Key Requirements

  • •Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent practical experience).
  • •6+ years of experience in Site Reliability Engineering, DevOps, or Platform/Infrastructure Engineering in production environments.
  • •Hands-on experience with GCP services including GKE, Cloud Run, BigQuery, Pub/Sub, and GCS.
  • •Working proficiency with Infrastructure-as-Code tools such as Terraform or Helm.
  • •Experience with observability tooling and SLO/SLI concepts for reliability and on-call operations.
Experience:RetailE-commerceConsumer goodsData platformsAI/ML
Education:Bachelor's
Skills:CommunicationDocumentation
Tech Stack:Google Cloud PlatformGKECloud RunBigQueryPub/SubGCSComposerTerraformHelmAzureObservabilityCloud MonitoringDatadogPrometheusGrafanaKubernetesCI/CDGitOpsArgoCDGitHub Actions

Company Brief

Levi Strauss
Designs, markets, and sells denim and casual apparel under the Levi's, Dockers, and related brands. The company operates globally through wholesale, retail, and direct-to-consumer channels.
Industry: Fashion & Apparel
Company Size: Enterprise (1,001+ employees)
Revenue: USD 5M to 10M
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Francisco, United States
Founded: 1853
WebsiteLinkedIn