4467- Software Development Engineer-III (Data Engineer) (Iceberg/Trino)

Innovaccer
Noida
Workplace: OnsiteFull timeFunction: Product ManagementEducation: bachelorsSkills: []

Build and operate lakehouse data pipelines for Innovaccer’s on-premise platform. You’ll create Spark ingestion jobs that land raw healthcare data into Apache Iceberg with schema handling and idempotent replay, then develop Trino SQL transform pipelines to validate, type, transform, deduplicate, and aggregate data for analytics and applications. You’ll also port SQL workloads, maintain Iceberg tables via scheduled workflows, tune performance, and instrument data-quality checks and rollout validation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Innovaccer
Innovaccer
1 day ago

4467- Software Development Engineer-III (Data Engineer) (Iceberg/Trino)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 day agoStatus: Live

Job Summary

Build and operate lakehouse data pipelines for Innovaccer’s on-premise platform. You’ll create Spark ingestion jobs that land raw healthcare data into Apache Iceberg with schema handling and idempotent replay, then develop Trino SQL transform pipelines to validate, type, transform, deduplicate, and aggregate data for analytics and applications. You’ll also port SQL workloads, maintain Iceberg tables via scheduled workflows, tune performance, and instrument data-quality checks and rollout validation.
Location: Noida
Workplace: Onsite
Employment Type: Full time
Job Function: Product Management
Seniority: Mid level

Key Responsibilities

  • •Build Spark ingestion jobs to land high-volume raw files into Iceberg with schema handling, bad-record quarantine, and idempotent batch replay.
  • •Develop and operate Trino SQL transform pipelines across data layers including validation/typing, business-rule transforms, MERGE-based deduplication, and aggregate builds.
  • •Port existing warehouse SQL workloads to Trino and Spark SQL dialects and validate results against source outputs.
  • •Automate Iceberg table maintenance such as compaction, snapshot expiry, and orphan-file cleanup via scheduled workflows.
  • •Tune query and pipeline performance (partitioning, file sizing, statistics, and resource-group behavior) and instrument pipelines with data-quality checks and reconciliation/alerting.

Pay and Benefits

Perks:Paid LeaveParental LeaveHealth Insurance

Key Requirements

  • •B.E., B.Tech., or M.Sc. in Computer Science or a related technical field.
  • •5+ years of data engineering experience building production pipelines at scale.
  • •Strong SQL skills and hands-on experience with Apache Spark for batch processing.
  • •Experience with Trino/Presto (or similar distributed SQL engines) and open table formats such as Iceberg (Delta Lake or Hudi acceptable).
  • •Working knowledge of S3-compatible object storage and columnar formats like Parquet; experience with Airflow (or equivalent) and CI/CD is required, plus Python and/or Java experience.
Experience:HealthcareData engineeringETLLakehouse
Education:Bachelor's in Computer Science
Tech Stack:SparkApache IcebergTrinoPrestoSQLDelta LakeHudiS3-compatible object storageParquetAirflowCI/CDPythonJavaHL7CCDAClaims filesAWSAmazon BedrockAWS HealthLake

Company Brief

Innovaccer
Provides a healthcare data activation platform that unifies clinical, claims, and operational data to enable analytics, care management, population health, and value-based care initiatives for providers, payers, and life sciences organizations.
Industry: HealthTech
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Francisco, United States
Founded: 2014
WebsiteLinkedIn