Senior Production Engineer

Whatnot
San Francisco, New York
Workplace: RemoteFull timeUSD 207,000 - 290,000 annuallyFunction: Manufacturing & Production OperationsExperience: 6+ yearsEducation: bachelorsSkills: ["Bias to action","Outcome-focused","Curiosity","Embedded collaboration","Incident response"]

Embed with product, platform, and infrastructure teams to keep live commerce running reliably at scale. Hunt anomalies across latency, errors, traffic, and cost signals, trace them to root cause across services and storage, and connect operational impact to business outcomes. Lead incident response for complex cross-team failures, improve observability, run load and failure drills, and apply AI to automate anomaly detection and reduce operational toil.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Whatnot
Whatnot
1 week ago

Senior Production Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Embed with product, platform, and infrastructure teams to keep live commerce running reliably at scale. Hunt anomalies across latency, errors, traffic, and cost signals, trace them to root cause across services and storage, and connect operational impact to business outcomes. Lead incident response for complex cross-team failures, improve observability, run load and failure drills, and apply AI to automate anomaly detection and reduce operational toil.
Location: San Francisco, New York
Workplace: Remote
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Mid level

Key Responsibilities

  • •Embed with a team to raise reliability, performance, and scalability and implement code changes to improve systems.
  • •Hunt anomalies across traffic, latency, errors, and cost signals, tracing issues to root cause across services, storage, and clients.
  • •Connect operational signals to business impact by identifying which degradations matter and which can be safely ignored.
  • •Prepare for peak events via capacity modeling, load testing, and failure drills; improve observability to reduce time-to-detection.
  • •Lead incident response for complex cross-team failures, driving systemic fixes and handling escalations during live incidents.

Pay and Benefits

Salary: USD 207,000 - 290,000 annually
Perks:Paid LeaveHealth InsuranceDentalVisionRemote WorkHome Office401kChildcareWellness StipendPaid Parental

Key Requirements

  • •Bachelor’s degree in Computer Science or a related field, or equivalent work experience.
  • •6+ years building and debugging production services at scale (software engineering, production engineering, SRE, or systems engineering).
  • •Strong systems and distributed systems fundamentals, including failure modes and queueing/cascading behavior.
  • •Depth in observability and production debugging, plus experience with capacity/performance engineering, load/resilience testing, incident command, or large-scale migrations.
  • •Comfort working embedded with other teams and improving systems in their codebases.
Experience:6+ yearsProduction servicesSRESystems engineeringReal-time systemsEvent-driven systemsLive streaming
Education:Bachelor's in Computer Science
Skills:Bias to actionOutcome-focusedCuriosityEmbedded collaborationIncident response
Tech Stack:PythonElixirGoAWSGCPKubernetesInfrastructure as codeLinuxNetworkingStorage

Company Brief

Whatnot
Whatnot is a live-stream shopping marketplace connecting sellers and collectors for trading cards, toys, and collectibles. The platform hosts live auctions and interactive streams, enabling creators and small businesses to sell directly to engaged communities.
Industry: Online Marketplaces
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: United States
Founded: 2019
WebsiteLinkedIn