jobs Logo
BBI logo

Data Engineer

BBI6 days ago
Toronto, Ontario, Canada
Senior Level
Full-Time

About the role

Role Overview We are seeking an experienced Data Engineer with healthcare payer domain knowledge to design, build, and operate data pipelines and infrastructure within a large-scale, cloud-native data platform on GCP. The role operates in a complex, heterogeneous data environment spanning on-premises relational systems — including DB2 and Oracle — virtual integration layers, flat-file and SFTP-based sources, and real-time streaming feeds, all flowing into a GCP-based Lakehouse architecture.

The ideal candidate brings strong PySpark and GCP engineering skills, a deep understanding of data pipeline design for regulated healthcare data, and the ability to work effectively in environments where both legacy and modern technologies coexist. You will build and maintain the pipelines that ingest, transform, and deliver high-quality healthcare payer data to operational and analytical consumers at scale.

Primary Responsibilities Pipeline Development Design, develop, and maintain scalable PySpark data pipelines on Dataproc for batch data processing across a wide range of healthcare payer source systems — including claims, member enrollment, provider, finance, and drug/pharmacy data originating from DB2, Oracle, Denodo, and flat-file SFTP sources. Build and maintain real-time streaming ingestion pipelines using Confluent Kafka, processing high-volume event streams for operational and analytical consumers. Implement Apache Iceberg table format across data lake storage layers, including schema management, partitioning design, compaction, and time-travel query support. Develop data transformation logic within curated and certified data layers — including cross-reference table processing, SCD1 pattern implementation, and history table management for member and provider data. Build and maintain orchestration workflows using Cloud Composer (Airflow) for batch pipeline scheduling, dependency management, SLA monitoring, and operational alerting.

Data Quality & Observability Implement automated data quality validation within pipelines — including record count checks, referential integrity validation, business rule enforcement, and anomaly detection — ensuring only certified data reaches serving and downstream consumers. Build monitoring and alerting frameworks for pipeline health, processing SLAs, and data freshness across batch and streaming workloads. Contribute to metadata and data lineage instrumentation within pipelines, supporting enterprise catalog and governance requirements. Participate in evaluations of data observability tools and quality frameworks as part of the evolving platform toolchain.

Platform Operations & Standards Implement and maintain security controls within data pipelines — including PHI/PII masking, column-level encryption, role-based access controls, and audit logging — in compliance with HIPAA, FedRAMP, and NIST standards. Contribute to Infrastructure-as-Code (IaC) practices for Dataproc cluster provisioning, Cloud Composer environment management, and Artifactory-based CI/CD deployment pipelines. Troubleshoot pipeline failures, optimize performance for high-volume workloads, and implement solutions for data reliability, idempotency, and retry handling. Produce technical documentation and contribute to knowledge transfer, ensuring pipeline logic and design decisions are well understood across the team.

Minimum Qualifications Bachelor's degree in Computer Science, Engineering, Data Science, or a related field. 7+ years of data engineering experience with at least 4 years building production-grade data pipelines for healthcare payer organizations covering claims, enrollment, provider, or analytics domains. Strong PySpark and SQL proficiency — the ability to write, optimize, and debug complex transformation logic in PySpark is a core requirement of this role. Hands-on experience with GCP data services: Dataproc, BigQuery, Cloud Storage (GCS), Pub/Sub, and Cloud Composer (Airflow). Experience working with on-premises relational source systems — including DB2 and Oracle — across ingestion, schema mapping, and incremental load patterns. Working knowledge of Apache Iceberg or Delta Lake, including understanding of open table format concepts, ACID transactions, schema evolution, and partition management. Experience with Apache Kafka or Confluent Kafka for real-time streaming data ingestion and event-driven pipeline design. Proficiency in Git-based version control and CI/CD practices for data pipeline deployment.

Preferred Experience Familiarity with virtual data integration platforms such as Denodo — understanding of how to consume data from virtual layers and how those patterns compare to direct source system ingestion. Experience with file-based ingestion from SFTP and cloud storage sources, including handling of Excel, CSV, text, and SAS file formats common in healthcare payer environments. Exposure to federated query and analytics platforms — such as Starburst — that operate across heterogeneous data sources including operational databases, cloud data warehouses, and Iceberg-based data lakes. Familiarity with modern data catalog and observability tools — such as Atlan or Monte Carlo — particularly in proof-of-concept or early adoption contexts. Understanding of Biglake Catalog or equivalent open catalog standards for unified metadata management across GCP-native and Iceberg-based storage. Knowledge of healthcare payer data structures and standards including X12 EDI, HL7, FHIR, and CMS reporting requirements. Strong problem-solving skills with the ability to diagnose and resolve complex data issues across heterogeneous, multi-layer pipeline architectures.

About BBI

IT Services and IT Consulting
501-1,000 employees
Founded in 2016

BBI is your one-stop shop for making AI work reliably at scale. Working alongside client teams, we build, modernize, and operate data foundations that ensure data and AI output are consistent, traceable, and production-ready across systems. Using AI-driven engineering methods and proven accelerators, we deliver fast, reduce risk, and give clients data they can trust.

Clients call on us when they face a stalled modernization, a fragile platform, a post-acquisition integration, or a mandate to support analytics and AI that current foundations can't support.

Clients stick with us because we take accountability for outcomes. We don't hand off and walk away. You get a team that stays engaged through the full lifecycle, from design through ongoing operation.

Our services include:

Data management services: BBI engineers reliable data pipelines, consistent transformations, and clear controls, ensuring data behaves predictably and can be trusted as systems scale and requirements change.

Data modernization services: BBI modernizes and scales data platforms and infrastructures, making data easier to access, easier to work with, and better aligned to analytics, automation, and AI initiatives.

AI and applications: We build AI solutions and business applications directly on top of your data foundation, turning trusted data into intelligent tools that power your business.

Platform operations and support services: BBI runs operations so your engineering teams don't have to. We monitor pipelines, respond to alerts before issues reach the business, and validate data quality before reports go out.

Our delivery accelerators, AI-assisted engineering, and proven frameworks have produced 65% faster delivery times and 70% fewer errors, without sacrificing quality or control.

Headquartered in Schaumburg, Illinois, we operate with more than 600 data professionals across the United States, Canada, and India.

Your data. Ready for anything.

Similar Jobs