Big Data Development Engineer Intern

Bybit
Hong Kong
Workplace: OnsiteInternshipFunction: Product ManagementSkills: ["Structured thinking","Self-driven","Communication","Problem formulation","Comfortable with ambiguity"]

Work on AI-assisted data governance and automation for a large-scale data environment. Research best practices for metadata, lineage, semantic layers, and NL2Metric approaches, then build capabilities for model governance, metric semantic automation, and automated data quality and anomaly detection. Deliver an end-to-end MVP across discovery, prototype, and deployment, and communicate results to data and platform stakeholders.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Bybit
Bybit
1 day ago

Big Data Development Engineer Intern

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 minutes agoStatus: Live

Job Summary

Work on AI-assisted data governance and automation for a large-scale data environment. Research best practices for metadata, lineage, semantic layers, and NL2Metric approaches, then build capabilities for model governance, metric semantic automation, and automated data quality and anomaly detection. Deliver an end-to-end MVP across discovery, prototype, and deployment, and communicate results to data and platform stakeholders.
Location: Hong Kong
Workplace: Onsite
Employment Type: Internship
Job Function: Product Management
Seniority: Intern level

Key Responsibilities

  • •Survey industry and open-source approaches for AI + data governance (metadata/lineage, semantic layers, data quality frameworks, NL2SQL/NL2Metric) and propose best practices for the tech stack.
  • •Build AI-assisted data model review capabilities for naming/layering validation, duplicate detection, lineage-based identification of unused/low-value/high-cost assets, and automated refactoring recommendations.
  • •Automate metric definition extraction and standardization from SQL/lineage/documentation, detect conflicts and redundant metric builds, and maintain a machine-readable semantic layer.
  • •Automate data quality rule generation from profiling/lineage, perform anomaly detection and root-cause analysis, and implement alert grading with remediation recommendations (or auto-remediation).
  • •Deliver an end-to-end MVP across the three areas (problem definition to deployment) and communicate solution designs, evaluation methods, and results to stakeholders.

Pay and Benefits

Perks:Learning Budget

Key Requirements

  • •Undergraduate or graduate student in Computer Science, Data Science, Statistics, or a related field.
  • •Strong proficiency in SQL and Python, plus data warehouse fundamentals (dimensional modeling, metadata, lineage, layered architecture).
  • •Experience with Spark/Flink/Hive/StarRocks is a plus.
  • •Hands-on experience building LLM applications (prompt engineering, RAG, and agent/tool-calling frameworks like LangChain or LlamaIndex).
  • •Minimum 3-month internship commitment, 5 days per week on-site.
Experience:LLMRAGData governance
Education:
Skills:Structured thinkingSelf-drivenCommunicationProblem formulationComfortable with ambiguity
Languages:MandarinEnglish
Tech Stack:SQLPythonSparkFlinkHiveStarRocksLangChainLlamaIndexMCPLLMRAGNL2SQLNL2MetricDataHubOpenMetadataAtlasDbtGreat ExpectationsDeequ

Company Brief

Bybit
Operates a cryptocurrency exchange and trading platform offering spot, derivatives, copy trading, and related digital asset services for retail and institutional users. The platform focuses on high-liquidity crypto markets, trading tools, and Web3-related products.
Industry: Trading Platforms
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Founded: 2018
WebsiteLinkedIn