Software Engineer, Infrastructure

Descript
San Francisco
Full timeUSD 220,000 - 292,000 annuallyFunction: Software EngineeringExperience: 8+ yearsSkills: ["Reliability focus","Critical thinking","Incident response","Trade-off analysis","Clear technical communication"]

Own critical infrastructure that powers how the company computes, deploys, and runs models—spanning GCP, Kubernetes, Temporal, CI/CD/monorepo health, GPU fleet operations, and deploy/rollback machinery. Improve reliability through on-call, SLOs/error budgets, observability, runbooks, and incremental releases. Strengthen security boundaries (IAM, secrets, least privilege, supply-chain integrity) and help shape AI enablement for training and inference pipelines.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Descript
Descript
2 days ago

Software Engineer, Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Own critical infrastructure that powers how the company computes, deploys, and runs models—spanning GCP, Kubernetes, Temporal, CI/CD/monorepo health, GPU fleet operations, and deploy/rollback machinery. Improve reliability through on-call, SLOs/error budgets, observability, runbooks, and incremental releases. Strengthen security boundaries (IAM, secrets, least privilege, supply-chain integrity) and help shape AI enablement for training and inference pipelines.
Location: San Francisco
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Own the platform powering compute, deployment, reliability/on-call, and CI/CD/monorepo health, including deploy and rollback machinery.
  • •Operate the AI enablement substrate: GPU capacity plus training and inference pipelines that serve models in production.
  • •Improve system reliability and actionability through on-call rotations, incident response, SLO/error-budget-driven operations, and observability.
  • •Advance security boundaries across identity/access, secrets management, least-privilege boundaries, and supply-chain integrity.
  • •Increase team learning and shipping velocity by strengthening tooling, standards, tests, and release practices while keeping infrastructure legible with IaC, runbooks, and in-repo context.

Pay and Benefits

Salary: USD 220,000 - 292,000 annually
Perks:Health Insurance401kPaid Leave

Key Requirements

  • •8+ years building and operating production distributed systems, or equivalent server-side engineering with a heavy infrastructure focus.
  • •Hands-on experience running production systems where failures are expensive, including incident response and rollback.
  • •Experience using SLOs and error budgets as operating tools.
  • •Production experience with a major cloud provider and Kubernetes, with infrastructure-as-code as your default.
  • •Owned architecture/migrations end-to-end, including planning through launch with lasting consequences.
Experience:8+ yearsDistributed systemsServer-side engineeringInfrastructure
Skills:Reliability focusCritical thinkingIncident responseTrade-off analysisClear technical communication
Languages:English
Tech Stack:GCPKubernetesTemporalGPU fleetCloud exportCI/CDMonorepoInfrastructure-as-codeRunbooksObservabilityIAMSecrets managementLeast privilegeSLOsError budgets

Company Brief

Descript
Descript provides an all-in-one audio and video editing platform that combines transcription, multitrack editing, screen recording, and AI-powered tools for creators and teams to produce podcasts, videos, and marketing content more quickly and collaboratively.
Industry: Enterprise Software
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2017
WebsiteLinkedIn