Consultant Specialist

HSBC
Xi'an
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Code reviews","Documentation","Engineering practices","Troubleshooting"]

Design and deliver CDMS migration and incremental sync solutions, including CDC cutover, replay/backfill, and recovery. Build and maintain batch/stream data processing with PySpark, Spark SQL, and Structured Streaming, integrating Kafka CDC event streams with idempotent, out-of-order/late handling. Migrate and ingest MongoDB data, orchestrate pipelines in Airflow, implement data quality/reconciliation controls, and optimize performance across an on-prem Hadoop ecosystem.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
HSBC
HSBC
1 month ago

Consultant Specialist

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Design and deliver CDMS migration and incremental sync solutions, including CDC cutover, replay/backfill, and recovery. Build and maintain batch/stream data processing with PySpark, Spark SQL, and Structured Streaming, integrating Kafka CDC event streams with idempotent, out-of-order/late handling. Migrate and ingest MongoDB data, orchestrate pipelines in Airflow, implement data quality/reconciliation controls, and optimize performance across an on-prem Hadoop ecosystem.
Location: Xi'an
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and deliver CDMS migration and incremental sync solutions, including full load, CDC cutover, replay/backfill, and recovery mechanisms.
  • •Develop and maintain batch/stream processing jobs using PySpark, Spark SQL, and Structured Streaming for data cleansing, standardization, mapping, deduplication, and merge.
  • •Integrate with Kafka CDC events, defining topic/partition/key strategy and implementing out-of-order/duplicate/late handling, replay, and idempotent processing.
  • •Migrate and incrementally ingest MongoDB data with schema evolution, nested/array flattening, and mapping to target models.
  • •Orchestrate and operate data pipelines with Airflow and Hadoop, including data quality/reconciliation controls and performance optimization.

Key Requirements

  • •Strong Python fundamentals with hands-on PySpark experience, including DataFrames, Spark SQL, RDD, tuning, and troubleshooting.
  • •Proven Spark Structured Streaming experience (watermarking, stateful processing, checkpointing, delivery semantics).
  • •Strong Kafka experience including topic/partition/key strategy, consumer groups, offset management, replay, ordering, and idempotent processing.
  • •Hands-on CDC implementation experience covering initial load, incremental continuity, backfill, and consistency guarantees.
  • •Experience with on-prem Hadoop (HDFS/Hive/YARN) and Airflow (DAG patterns, SLA/alerting, retries, backfills).
Experience:CDCData engineeringReal-time streamingBig dataData migration
Skills:Code reviewsDocumentationEngineering practicesTroubleshooting
Languages:English
Tech Stack:PythonPySparkSpark SQLStructured StreamingKafkaMongoDBJavaSpring BootJSONAvroHadoopHDFSHiveYARNAirflowCDMSCDCDLQ

Company Brief

HSBC
Global banking and financial services organisation offering retail, commercial, corporate and investment banking, wealth management, and global markets services across Europe, Asia, the Americas and the Middle East.
Industry: Banking
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: London, United Kingdom
Founded: 1865
Glassdoor
Glassdoor: 3.6
WebsiteLinkedIn