Software Engineer, Site Reliability (SRE)

Sierra
San Francisco
Workplace: OnsiteFull timeUSD 230,000 - 390,000 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Collaboration","Communication","Problem-solving","Ownership","Adaptability"]

Software Engineer on the Site Reliability (SRE) team responsible for defining and building the foundation of reliability, observability, and scalability across Sierra’s AI-driven infrastructure. You’ll own the observability stack, design reliable and scalable AWS-based infrastructure with Terraform, improve LLM deployment reliability, and lead CI/CD and incident-management improvements to reduce downtime.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Sierra
Sierra
10 months ago

Software Engineer, Site Reliability (SRE)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Software Engineer on the Site Reliability (SRE) team responsible for defining and building the foundation of reliability, observability, and scalability across Sierra’s AI-driven infrastructure. You’ll own the observability stack, design reliable and scalable AWS-based infrastructure with Terraform, improve LLM deployment reliability, and lead CI/CD and incident-management improvements to reduce downtime.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Own Sierra’s observability stack—monitoring, alerting, logging, and tracing—to give engineers clear visibility into system health and performance.
  • •Partner with product and platform engineers to design systems that are reliable and scalable from day one—not as an afterthought.
  • •Design and implement scalable, reliable, and secure cloud infrastructure (AWS) using Terraform and modern DevOps tooling.
  • •Improve the reliability and scalability of our LLM deployments, ensuring robust, performant, and cost-effective operation.
  • •Lead improvements to deployment pipelines, CI/CD tooling, and incident management processes to reduce downtime and response time.

Pay and Benefits

Salary: USD 230,000 - 390,000 annually
Equity and Bonus:Equity
Perks:MedicalDentalVisionRetirementParental LeaveEquityMeal AllowanceStipend

Key Requirements

  • •5+ years of hands-on experience in Site Reliability or Infrastructure engineering roles for complex SaaS or cloud-based systems.
  • •Experience designing for availability, scalability, and reliability at both infrastructure and application layers.
  • •Deep experience with Terraform, AWS services, container orchestration, and cloud networking (including IAM and VPC architecture).
  • •Strong background in observability systems (e.g., Prometheus, Grafana, Datadog, or similar).
  • •Experience working with enterprise customers and familiarity with their compliance and networking needs along with integration patterns.
Experience:5+ yearsSaaSCloudAI
Education:Bachelor's in Computer Science
Skills:CollaborationCommunicationProblem-solvingOwnershipAdaptability
Tech Stack:TerraformAWSKubernetesIAMVPCPrometheusGrafanaDatadogCI/CD

Company Brief

Sierra
Develops advanced AI agents and tools focused on building autonomous, generalist machine learning systems to assist developers and enterprises in automating complex workflows, decision-making, and application-level tasks.
Industry: AI & Machine Learning
Website