Site Reliability Engineer - AI Application

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: []

Build and operate reliable AI search and recommendation systems, ensuring core Big Data and online services run normally through capacity planning, stability assurance, and SLA-focused performance optimization. Improve observability by monitoring availability and performance metrics, helping teams quickly diagnose issues—especially around AI search/vector databases. Participate in designing automation for large-scale Viking and AI search clusters and lead governance improvements for high-availability architecture.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Site Reliability Engineer - AI Application

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and operate reliable AI search and recommendation systems, ensuring core Big Data and online services run normally through capacity planning, stability assurance, and SLA-focused performance optimization. Improve observability by monitoring availability and performance metrics, helping teams quickly diagnose issues—especially around AI search/vector databases. Participate in designing automation for large-scale Viking and AI search clusters and lead governance improvements for high-availability architecture.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Ensure reliability and normal operation of core systems for the AI-driven search and recommendation platform, including capacity planning and stability assurance.
  • •Monitor availability and performance metrics to improve observability and help development teams quickly locate faults, with emphasis on critical AI search/vector database links.
  • •Improve reliability, scalability, and service performance optimization to meet core system SLAs.
  • •Design and implement an automation platform to enable rapid iteration and efficient operation and maintenance of large-scale online Viking and AI search clusters.
  • •Optimize service governance practices by analyzing performance bottlenecks, troubleshooting issues, and helping upgrade high-availability architecture for AI Search/Viking systems.

Key Requirements

  • •Bachelor's degree or above in a computer-related field, with more than five years of relevant experience.
  • •Solid foundation in computer software and understanding of Linux operating systems, storage, and network I/O.
  • •Familiar with at least one programming language (Python/Go/Java/Shell/Ansible) with moderate development capability and strong operations/maintenance and problem-solving skills.
  • •Knowledge of cloud infrastructure (AWS/Volcano Engine/Aliyun/GCP); experience with computing/distributed systems is preferred.
  • •Preferred: algorithmic thinking with good data structures and system design capabilities, plus understanding of AI cloud, large model-related search suggestion, and recommender systems.
Experience:5+ yearsAI searchRecommender systemsDistributed systemsBig data
Education:Bachelor's in computer-related fields
Tech Stack:LinuxPythonGoJavaShellAnsibleAWSVolcano EngineAliyunGCPNginxKubernetesDockerOpenStackHadoopSparkFlink

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn