Principal Software Engineer - Data Platform (Iceberg/Trino)

Innovaccer
United States
Workplace: RemoteFull timeFunction: Software EngineeringExperience: 12+ yearsEducation: bachelorsSkills: ["Mentorship","Engineering quality","Collaboration","Strategic thinking"]

Own the architecture of a ground-up lakehouse engine for on-premise deployments, using Apache Iceberg as the table format and Trino for query execution. Define catalog services and Spark-based transform compute, design on-prem equivalents of cloud warehouse capabilities, and validate performance through proof-of-concepts. Set platform-wide standards, lead SQL dialect porting, mentor senior engineers, and partner on storage sizing and capacity planning.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Innovaccer
Innovaccer
1 day ago

Principal Software Engineer - Data Platform (Iceberg/Trino)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Own the architecture of a ground-up lakehouse engine for on-premise deployments, using Apache Iceberg as the table format and Trino for query execution. Define catalog services and Spark-based transform compute, design on-prem equivalents of cloud warehouse capabilities, and validate performance through proof-of-concepts. Set platform-wide standards, lead SQL dialect porting, mentor senior engineers, and partner on storage sizing and capacity planning.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Own the lakehouse reference architecture, including Iceberg table design, Trino cluster topology, REST catalog service, Spark transform compute, and object-storage layout.
  • •Design on-premise replacements for cloud-managed warehouse capabilities, including change-data-capture streams, scheduled tasks, and write-back paths.
  • •Run proof-of-concept validation for the catalog and query engine at expected data volumes and define evidence-based placement decisions (VM vs Kubernetes-native operators).
  • •Set platform standards for table layout, partitioning, file sizing, and Iceberg maintenance (compaction, snapshot expiry, orphan-file cleanup).
  • •Lead SQL dialect strategy for porting warehouse workloads to Trino and Spark SQL, mentor senior engineers, and partner on storage sizing and capacity planning.

Pay and Benefits

Perks:Paid LeaveParental LeaveHealth InsuranceDentalVisionDisabilityPet Insurance

Key Requirements

  • •B.E., B.Tech., M.Sc. in Computer Science or a related technical field.
  • •12+ years building and operating large-scale data platforms or distributed systems.
  • •Hands-on expertise with distributed SQL engines, including Trino/Presto or Spark SQL internals, query planning, and performance engineering.
  • •Production experience with Apache Iceberg (or Delta Lake/Hudi with willingness to go deep on Iceberg), including table spec and maintenance at scale.
  • •Professional software development experience with Java and/or Python.
Experience:12+ yearsData platformsDistributed systemsLakehouseOn-premiseHealthcare data
Education:Bachelor's
Skills:MentorshipEngineering qualityCollaborationStrategic thinking
Tech Stack:Apache IcebergTrinoPrestoREST catalogPolarisNessieHive MetastoreSparkSQLJavaPythonS3Delta LakeHudiSnowflakeBigQueryRedshiftKubernetesVM

Company Brief

Innovaccer
Provides a healthcare data activation platform that unifies clinical, claims, and operational data to enable analytics, care management, population health, and value-based care initiatives for providers, payers, and life sciences organizations.
Industry: HealthTech
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Francisco, United States
Founded: 2014
WebsiteLinkedIn