Senior Software Engineer, Site Reliability & Security

WellSaid
United States
Workplace: RemoteFull timeFunction: Software EngineeringExperience: 5+ yearsSkills: ["Problem-solving","Incident response","Root cause analysis"]

Own the reliability and security of a production AI voice platform by building and operating monitoring, tracing, and alerting. Lead incident response with root-cause analysis, and keep core services highly available with redundancy. Improve deployments for fast, safe code changes, troubleshoot complex production issues, and help engineers write reliable code. Work with AWS/GCP, Kubernetes, Docker, and infrastructure-as-code to scale resource-intensive workloads and deliver stable platforms for engineering teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
WellSaid
WellSaid
1 day ago

Senior Software Engineer, Site Reliability & Security

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Own the reliability and security of a production AI voice platform by building and operating monitoring, tracing, and alerting. Lead incident response with root-cause analysis, and keep core services highly available with redundancy. Improve deployments for fast, safe code changes, troubleshoot complex production issues, and help engineers write reliable code. Work with AWS/GCP, Kubernetes, Docker, and infrastructure-as-code to scale resource-intensive workloads and deliver stable platforms for engineering teams.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Run and maintain the production platform and core services/data pipelines to improve availability and performance.
  • •Build and operate monitoring, tracing, and alerting infrastructure; respond to alerts as part of the on-call rotation.
  • •Drive platform security initiatives and perform preventative work to ensure system reliability.
  • •Lead incident response and recovery, including root cause analysis for active incidents.
  • •Improve deployments and deployment processes for fast, simple, and safe code changes while enabling stable, scalable platform operations.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kEquityPaid LeaveParental LeaveLearning BudgetHome Office

Key Requirements

  • •5+ years with a modern programming language (Golang, Typescript, Python, etc.).
  • •Strong understanding of Infrastructure as Code using Terraform, Tofu, or Pulumi.
  • •Experience with GitOps/continuous delivery tooling such as ArgoCD, Spacelift, or Terraform Cloud.
  • •Experience building and troubleshooting Kubernetes environments in production.
  • •Fluency working in a UNIX shell to analyze logs and investigate operational issues; ability to debug complex production environments.
Experience:5+ yearsAI
Skills:Problem-solvingIncident responseRoot cause analysis
Tech Stack:GolangTypeScriptPythonAWSGCPKubernetesDockerTerraformTofuPulumiGitOpsArgoCDSpaceliftTerraform CloudUNIX shellGrafanaPrometheus

Eligibility

Nationality:US National

Company Brief

WellSaid
Builds realistic, AI-generated voice solutions and text-to-speech tools for enterprises, creators, and developers, enabling lifelike synthetic voices for narration, e-learning, advertising, and accessibility applications.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Seattle, United States
Founded: 2017
WebsiteLinkedIn