Data Engineer – Healthcare Data Platform (Provider Analytics)

Data Engineer – Healthcare Data Platform (Provider Analytics)
Location
Gurugram, Haryana, India
Experience
Mid
Posted
Jul 18, 2026
Apply by
August 17, 2026
Applicants
0
Job Description
Data Engineer — Healthcare Data Platform
Reports To: Lead Architect, Provider Analytics | | --- | | FUNCTION
Provider Analytics | EXPERIENCE
4–7 Years | LOCATION
Gurgaon (Office-based) | LEVEL
Mid-Senior IC | | --- | --- | --- | --- | Why This Role Exists Neolytix is standing up a dedicated analytics function that owns all client-facing reporting for our RCM business (20 clients today) and the data foundation behind InCredibly, our credentialing platform. Reporting is moving out of Operations and onto an engineered data platform with formal governance: automated extraction pipelines, a certified metric layer, and validation built into the data flow rather than bolted on afterward. Platform architecture is owned in-house by a hands-on data architect who sets the design — medallion lakehouse structure, modeling standards, and the validation framework. This role builds the platform inside that design, owning pipelines end-to-end from ingestion through production operation, and is explicitly expected to grow into independent ownership of the platform within 6–12 months. We are hiring an inheritor, not a ticket-taker. What this role is not: it is not a Power BI report developer, and it is not a passive ETL operator who waits for a mapping document and a ticket queue. You should be able to point to at least one pipeline or data product you owned end-to-end — through design input, build, and production operation — and bring the appetite to take over a platform, not just work inside one. Mission Build Neolytix’s healthcare data platform inside an architect-defined design: the lakehouse, ingestion pipelines, data models, and automated validation layer that power client reporting across our RCM book and the data needs of the InCredibly platform. Within 12 months, the majority of recurring client reporting runs on pipelines you built and operate — with data quality proven by reconciliation, not assumed — and you are running the platform with review-only oversight. What You Will Build (First 12 Months) - The medallion lakehouse (bronze / silver / gold) on the Microsoft data stack, implemented to the architect’s design and documented so every decision is reproducible. - Automated extraction connectors for EHR / practice-management systems (target: 4 systems in year one), plus ingestion of clearinghouse 835 / 837 files and client-supplied feeds — replacing manual monthly pulls. - A claim lineage database linking charges, claims, remittances, and payments across sources (target: ≥90% linkage rate), enabling true first-pass-resolution and denial analytics. - Dimensional (star-schema) models for core RCM KPIs — first-pass resolution, days in AR, denial rates, net collection rate — serving certified Power BI semantic models. - The automated validation harness inside the pipelines: row counts, reconciliation-to-source totals, period-over-period variance thresholds, and alerting — so errors are caught before a client sees them. - A bounded InCredibly workstream: support for an in-flight client data migration and integration groundwork (API / FHIR) for the platform’s reporting needs. Ongoing Responsibilities - Own pipeline reliability end-to-end: orchestration, monitoring, failure recovery, schema-drift handling, and refresh SLAs. - Absorb and document the architecture as it is built — decision records, runbooks, and standards — with the explicit goal of independent platform ownership within 6–12 months. - Implement the technical half of data governance: lineage tracking, enforcement of certified metric definitions in the transformation layer, and role-based access to data assets. - Integrate new data sources as clients onboard — API-based where available, structured file exchange (SFTP) where not. - Handle PHI to HIPAA standards: minimum-necessary access, encryption in transit and at rest, and audit-ready handling in every pipeline. - Work day-to-day with the associate data engineer; grow into a mentoring role as ownership expands. What You Bring Must-Have - Strong SQL and solid Python for data engineering. - Hands-on Azure data stack experience — Azure Data Factory plus Synapse, Fabric, or Databricks (equivalent lakehouse experience on another cloud considered with strong fundamentals). - Has owned at least one pipeline or data product end-to-end — contributed to its design, built it, and operated it in production — as opposed to executing assigned tickets inside someone else’s build. - Working knowledge of dimensional modeling applied in production. - A demonstrated validation and reconciliation mindset: you can explain, concretely, how you knew your numbers were right. This is non-negotiable at any level. - Disciplined handling of sensitive / regulated data; HIPAA or PHI exposure preferred. - Evidence of fast learning and increasing scope — this role is a succession seat, and trajectory matters as much as current state. Strong Pluses - US healthcare claims and remittance data (835 / 837), EHR / practice-management exports, or RCM metrics exposure. - FHIR / HL7 integration experience; REST API development. - Power BI semantic-model / dataset design (the front end sits on your models). What Success Looks Like 90 Days: First client pipeline live end-to-end under architect review; validation harness v1 implemented and catching reconciliation breaks automatically. 6 Months: Connectors and the claim lineage database in production; operating pipelines with review-only oversight; manual extraction hours measurably reduced. 12 Months: 4 EMR/PM connectors live with a repeatable playbook; ≥90% claim linkage; ≥97% scored data accuracy on certified reports; independently owning platform run and extension.
Key Responsibilities
- Build and maintain the medallion lakehouse on the Microsoft data stack.
- Develop automated extraction connectors for EHR and practice-management systems.
- Create a claim lineage database linking charges, claims, remittances, and payments.
- Design dimensional models for core RCM KPIs serving Power BI semantic models.
- Implement automated validation harnesses for data quality and reconciliation.
- Own pipeline reliability including orchestration, monitoring, and failure recovery.
- Document architecture and standards to enable independent platform ownership.
- Implement technical data governance including lineage tracking and access control.
- Handle PHI data in compliance with HIPAA standards.
Skills Required
App exclusive · Free
Smart Job AI Coach
Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.
Interview readiness
See how prepared you are and what to improve for each role.
Personalized tips
Actionable suggestions based on your profile and the job.
After you apply
Keep coaching momentum from job detail through application success.
Similar roles for you
Matched using this role's title and skills. Open the job search anytime to see every listing.

Senior Software Engineer (Web App)
TechStarsGroup LLC

Data Engineer Healthcare Startup
TechStarsGroup LLC
Senior Python Full Stack Lead Developer
Tekskills India Pvt Ltd
₹25,00,000–₹32,00,000 / Year

Senior Data Engineer
Our Future Health
74,000 / Year

Principal Data Engineer
CodaMetrix
175,000–200,000 / Year

Senior Engineer, Healthcare Data
b.well
160,000–190,000 / Year