← Back to insights

Bitemporal Healthcare Data, dbt Testing Standards, and the LLM Trap

October 11, 2026 · healthcare data engineering, bitemporal modeling, dbt, AI boundaries

Claims are not real-time event streams. Why naive AI pipelines hallucinate retrospective facts, what every timestamp actually encodes, and the strict engineering guardrails required for healthcare data.

When engineering teams from outside healthcare begin building AI agents on medical data, one of the most common early missteps is treating medical claims like an e-commerce webhook or a Stripe payment event. They assume that if a patient visited an urgent care center yesterday, an analytical agent can query the claims table today to personalize an automated outreach message or evaluate care quality.

In healthcare, that assumption collapses immediately.

Claims are almost never real-time. Even when an attentive provider submits an 837 claim within 24 hours of an encounter, clearinghouses batch transactions on scheduled processing windows. Payers almost always batch. By the time a claim clears the clearinghouse, undergoes adjudication, handles denials, and generates an 835 payment remittance, 48 hours is lightning fast—and weeks or months is normal operating reality.

Wiring an LLM context window directly to claims data without rigorous bitemporal modeling guarantees that the model will reason with anachronisms, misattribute spend, or hallucinate patient state.


The Operational Meaning Behind Every Timestamp

Because storage is inexpensive in modern cloud warehouses, implementing Slowly Changing Dimensions (SCD Type 2) snapshots is a baseline requirement. There is no legitimate architectural justification for discarding historical state changes when SQL-accessible audit trails can be maintained at minimal cost.

Crucially, in healthcare data, every single timestamp in the lifecycle encodes a different stakeholder's reality:

  • Date of Service: Reflects the patient’s clinical experience during the encounter—what was evaluated, tested, or treated in the room.
  • Submitted Date: Measures the efficiency and turnaround time of the provider’s billing and revenue cycle team.
  • Received Date: Reveals throughput, queue latency, and ingestion lag at the clearinghouse and payer gateway.
  • Paid Date: Tracks financial reconciliation, adjudication accuracy, and settlement terms.
  • Patient Responsibility / Due Amount Date: Represents a critical inflection point for the member. When a member unexpectedly owes money—whether from high deductibles, co-insurance, or a claim payment error between payer and provider—it triggers immediate financial stress and friction.

An automated workflow or AI agent that conflates Date of Service with Paid Date cannot tell whether an encounter occurred yesterday or three months ago, leading to catastrophic errors in automated member communication and care coordination.


The Non-Negotiable dbt Testing Standard

Before any healthcare warehouse model is exposed to an executive dashboard, metric contract, or LLM agent, it must pass a strict suite of automated dbt data tests. General web analytics can tolerate minor nulls or deduplication lag; healthcare pipelines cannot.

  1. Surrogate Key & Grain Uniqueness: Every transformation model must explicitly declare and test its primary grain. If a claim line item grain produces duplicate keys, downstream risk-scoring and spend totals are immediately compromised.
  2. Bitemporal & Multitemporal Enrollment Consistency: Eligibility is notoriously volatile. When effective dates, enrollment dates, and updated dates desynchronize, it reveals retro-enrollment, coverage churn, or backdated terminations. dbt tests must flag coverage date overlaps and chronological contradictions.
  3. National Provider Identifier (NPI) Registry Validation: Billing and rendering provider NPIs must be verified against current CMS NPPES registry lookup tables. Unregistered or inactive NPIs point to data entry defects or fraudulent submissions.
  4. Clinical Code Validity (ICD-10 & HCPCS/CPT): Diagnosis and procedure codes must conform to official ICD-10-CM, ICD-10-PCS, and HCPCS/CPT formatting rules and active date ranges.
  5. Default Midnight Timestamp Auditing: A frequent silent data defect occurs when source systems drop time precision and pad timestamps with midnight (00:00:00). Automated dbt tests must detect default timestamps to prevent models from generating invalid duration, queueing, or latency calculations.
  6. Referential Integrity Relationship Tests: Claims models must enforce foreign key integrity against verified member dimension spans. Orphaned claims indicating unmapped member IDs must break the build before hitting production metrics.

To help health-tech teams implement these standards immediately, we published the complete generic test suite as an open-source package on GitHub: refhealth-dbt-healthcare-checks. It installs directly into your dbt project via packages.yml.


Model Pragmatism: Cloud Commitments over Vendor Churn

In the early days of enterprise GenAI, engineering teams spent months debating negligible benchmark margins between model labs. In 2026, the performance differences across frontier models from Anthropic, OpenAI, and Google have largely converged.

The pragmatic architectural priority is not multi-vendor arbitrage. It is maximizing the client’s existing enterprise cloud commitment, signed Business Associate Agreement (BAA), and token allocation—whether that lives inside AWS Bedrock, Microsoft Azure, or Google Cloud Vertex AI. Building resilient data products and governing prompt context within the client’s contracted environment delivers immediate production leverage without procurement or compliance delays.


The Clinical Boundary: Schmitt-Thompson Protocols

Perhaps the most vital architectural decision in healthcare AI is knowing where the model must stop.

Industry-standard clinical guidelines—most notably Schmitt-Thompson clinical triage protocols—make this boundary unequivocal. Unless an organization operates a certified, clinical-grade platform subject to rigorous clinical validation, an AI model should not be making clinical diagnoses or triage recommendations. Period.

The high-value operating role for AI in healthcare is in eliminating administrative toil, summarizing complex operating contexts for human clinicians, accelerating analytics engineering, and structuring messy unstructured feeds—preserving human judgment and clinical licensure where patient safety is at stake.


Turning Standards into Immediate Action

Establishing this level of discipline is the core objective of our 2-week, $2,500 Healthcare AI Readiness Diagnostic. Rather than handing over an academic strategy document that gathers dust, we deliver an executive decision memo, a prioritized use-case matrix, a pipeline blocker audit, and ready-to-run execution scripts and dbt tests that engineering teams can merge immediately to demonstrate measurable improvement inside their own ecosystem.

Ready to build tested healthcare data foundations?

Start with our focused 2-week diagnostic from $2,500. We audit pipeline blockers, evaluate bitemporal readiness, and provide immediate execution scripts.

Start with the diagnostic