[Intern] Big Data Development Engineer Intern

Bybit
Hong Kong
Workplace: OnsiteInternshipFunction: Product ManagementSkills: ["Structured thinking","Self-driven","Comfortable with ambiguity","Clear written communication"]

Build AI-assisted data governance capabilities across naming/layering validation, lineage-based model/asset identification, and semantic metric automation. You’ll work with SQL/Python and big-data tooling (e.g., Spark/Hive) to auto-generate quality rules, detect anomalies via lineage, and deliver at least one end-to-end MVP on a real data domain. Document designs and results, and evaluate LLM approaches for real effectiveness.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Bybit
Bybit
19 hours ago

[Intern] Big Data Development Engineer Intern

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 hours agoStatus: Live

Job Summary

Build AI-assisted data governance capabilities across naming/layering validation, lineage-based model/asset identification, and semantic metric automation. You’ll work with SQL/Python and big-data tooling (e.g., Spark/Hive) to auto-generate quality rules, detect anomalies via lineage, and deliver at least one end-to-end MVP on a real data domain. Document designs and results, and evaluate LLM approaches for real effectiveness.
Location: Hong Kong
Workplace: Onsite
Employment Type: Internship
Job Function: Product Management
Seniority: Intern level

Key Responsibilities

  • •Research and benchmark industry/open-source approaches for AI + data governance and propose best practices tailored to the tech stack.
  • •Build AI-assisted model governance features, including naming/layer validation, duplicate detection, lineage-based identification of unused/low-value/high-cost assets, and automated refactoring recommendations.
  • •Automate metric semantic extraction and standardization from SQL, lineage, and documentation; detect conflicts and redundant metrics and maintain a machine-readable semantic layer.
  • •Auto-generate data quality rules from profiling/lineage, run anomaly detection across volume/distribution/timeliness/schema drift, and support root-cause analysis with alert grading and remediation recommendations.
  • •Deliver at least one end-to-end MVP (problem → design → prototype → deployment → quantified results) and communicate designs, methodologies, and outcomes to stakeholders.

Pay and Benefits

Perks:Learning Budget

Key Requirements

  • •Undergraduate or graduate student in Computer Science, Data Science, Statistics, or a related field.
  • •Solid proficiency in SQL and Python, with understanding of data warehouse fundamentals (dimensional modeling, layered architecture, metadata, lineage).
  • •Experience with Spark / Flink / Hive / StarRocks is a plus.
  • •Hands-on LLM application development (prompt engineering, RAG, agent/tool-calling frameworks like LangChain, LlamaIndex, MCP) and the ability to evaluate effectiveness.
  • •3+ months internship commitment, 5 days per week on-site; fluent Mandarin required; English reading and clear written communication skills.
Experience:Big dataLLM applicationsData governance
Education:
Skills:Structured thinkingSelf-drivenComfortable with ambiguityClear written communication
Languages:MandarinEnglish
Tech Stack:SQLPythonSparkFlinkHiveStarRocksLLMPrompt engineeringRAGLangChainLlamaIndexMCPDataHubOpenMetadataAtlasDbtGreat ExpectationsDeequDimensional modelingData warehouse

Company Brief

Bybit
Operates a cryptocurrency exchange and trading platform offering spot, derivatives, copy trading, and related digital asset services for retail and institutional users. The platform focuses on high-liquidity crypto markets, trading tools, and Web3-related products.
Industry: Trading Platforms
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Founded: 2018
WebsiteLinkedIn