Software Engineer, Site Reliability Engineer

FuriosaAI
Seoul
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: bachelorsSkills: ["Analysis","Communication"]

Improve the reliability, scalability, security, and operability of production infrastructure and customer-facing services. You’ll work across bare-metal Kubernetes clusters, cloud control planes, networking, observability systems, deployment pipelines, and API services on Furiosa NPUs. Define reliability goals with SLIs/SLOs and error budgets, build observability foundations (metrics/logs/traces/alerts), analyze end-to-end failure risks, and reduce operational toil through automation and safer rollouts.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
FuriosaAI
FuriosaAI
1 week ago

Software Engineer, Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live
Reposted: similar role first listed 3 months ago

Job Summary

Improve the reliability, scalability, security, and operability of production infrastructure and customer-facing services. You’ll work across bare-metal Kubernetes clusters, cloud control planes, networking, observability systems, deployment pipelines, and API services on Furiosa NPUs. Define reliability goals with SLIs/SLOs and error budgets, build observability foundations (metrics/logs/traces/alerts), analyze end-to-end failure risks, and reduce operational toil through automation and safer rollouts.
Location: Seoul
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Define and evolve reliability goals using SLIs, SLOs, error budgets, and operational metrics.
  • •Design and build observability foundations (metrics, logs, traces, dashboards, alerts, and service-level indicators).
  • •Analyze production systems end-to-end to identify reliability risks across software, infrastructure, and networking boundaries, driving architectural improvements.
  • •Improve change safety and failure recovery via rollout strategies, capacity planning, load validation, graceful degradation, and incident learning loops.
  • •Reduce operational toil by building automation, internal tooling, and self-service workflows to make systems easier to operate and harder to misuse.

Key Requirements

  • •Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • •Strong programming skills in one or more general-purpose languages such as Rust, Python, or Go.
  • •Solid understanding of operating systems, computer networks, and cloud-native or container-based environments.
  • •Ability to analyze technical problems and communicate clearly with engineering teams.
  • •Experience improving reliability using SLOs, observability, incident analysis, rollout safety, and error-budget-driven decision making.
Education:Bachelor's in Computer Science, Engineering, or a related field
Skills:AnalysisCommunication
Languages:English
Tech Stack:RustPythonGoKubernetesCloud control planesNetworkingObservabilitySLIsSLOsError budgetsDashboardsAlertsLogsTracesDeployment pipelinesAPI servicesAutomationTooling

Company Brief

FuriosaAI
Designs and develops data‑center AI inference accelerators (RNGD) and a full stack hardware‑software platform to deliver energy‑efficient, high‑performance AI compute for enterprise and cloud customers.
Industry: Hardware Devices
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series C
Headquarters: Seoul, South Korea
Founded: 2017
Glassdoor
Glassdoor: 4.6
WebsiteLinkedInGlassdoor