Staff Platform & Reliability Engineer

Interface AI
San Francisco
Workplace: OnsiteFull timeUSD 240,000 - 320,000 annuallyFunction: Data Analytics & Business IntelligenceEducation: bachelorsSkills: ["Judgment"]

Own the platform’s reliability and security backbone for an agentic AI platform used by banks and credit unions. Define SLOs/SLIs and run an error-budget program, lead incident management, and build GitOps-based deploy and rollback. Drive resilience and disaster recovery with tested failover, and strengthen observability with metrics/logs/tracing. Manage AWS + Kubernetes infrastructure-as-code, security/compliance evidence, and AI-specific threat guardrails.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Interface AI
Interface AI
16 hours ago

Staff Platform & Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Own the platform’s reliability and security backbone for an agentic AI platform used by banks and credit unions. Define SLOs/SLIs and run an error-budget program, lead incident management, and build GitOps-based deploy and rollback. Drive resilience and disaster recovery with tested failover, and strengthen observability with metrics/logs/tracing. Manage AWS + Kubernetes infrastructure-as-code, security/compliance evidence, and AI-specific threat guardrails.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Analytics & Business Intelligence
Seniority: Mid level

Key Responsibilities

  • •Define customer-facing SLIs and SLOs and run an error-budget program that governs release decisions.
  • •Own disaster recovery strategy including regional failover and written/validated RTO/RPO per tier.
  • •Build and operate a GitOps deploy path with progressive delivery, one-click rollback, drift detection, and change audit trails.
  • •Lead incident management end-to-end: paging/severity policy, incident commander rotation, blameless post-mortems, and customer SLA reporting.
  • •Deliver platform reliability, resilience, and observability using metrics/logging/distributed tracing and AI-native operations automation with safe runbooks and self-healing.

Pay and Benefits

Salary: USD 240,000 - 320,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kMeal AllowanceCommuter BenefitsWellness StipendParental Leave

Key Requirements

  • •Run production systems with real uptime commitments and carry the pager/incident responsibility.
  • •Deep production Kubernetes on AWS experience, including service mesh and GitOps-based delivery across many services.
  • •Infrastructure-as-code experience at multi-account scale, including taking over and reshaping an existing estate.
  • •Proven SLO and error-budget practice that changes release decisions, with alerting tuned to burn rate.
  • •Strong programming in TypeScript and/or Python plus Bash, and writing clear post-mortems; BS/BA in Computer Science required.
Experience:FintechBankingFinancial servicesCredit unionsRegulated industry
Education:Bachelor's in Computer Science
Skills:Judgment
Tech Stack:AWSKubernetesService meshGitOpsSLOSLITypeScriptPythonBashTerraformIAMSecrets managementSOC 2ISO 27001PCIFFIECNCUAGLBASBOMOWASP LLM Top 10

Eligibility

Visa:H1B

Company Brief

Interface AI
Provides AI-powered virtual assistant and automation software for banks and credit unions. The platform helps financial institutions handle customer service, support routine banking tasks, and improve digital self-service across voice and chat channels.
Industry: SaaS
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: San Jose, United States
Founded: 2018
WebsiteLinkedIn