aiwillnet.com
ARCHITECTURE / PATTERNS

Patterns that hold up at scale.

A growing library of the data and integration patterns actually used in delivery, not textbook theory. Each one explains the problem it solves, why the naive approach breaks down, and where it's shown up in production work.

Pattern 01 — FIG. 02

One shared shape, not N translations.

Connect systems directly to each other and every new system adds a translation for every system already in place. A canonical model breaks that: every system maps to and from one shared shape, once. Toggle the diagram below to see what changes.

Point-to-point integration versus a canonical data model Eight systems arranged in a circle. In point-to-point mode every pair connects directly. In canonical mode every system connects only to a central canonical model hub.
28
Integration points
baseline
vs. point-to-point
Point-to-point (n(n−1)/2) Canonical model (n)

At 8 systems: 28 point-to-point integrations vs. 8 canonical mappings, a 71% drop in what has to be built and maintained. Move the pointer over the chart to compare any system count from 2 to 12.

Pattern 01 · Delivered — FIG. 03

Where this shows up in the work.

  • MDM and EMPI golden records, entity lifecycle, merge/unmerge, and probabilistic/deterministic matching all resolve multiple source systems into one canonical entity shape, not pairwise matching rules per source.
  • HL7 FHIR, healthcare's own canonical model, used as the target shape for payer, clearinghouse, and EMR/EHR integrations instead of custom mappings per source system.
  • DAMA DMBOK-aligned governance, the canonical schema itself versioned and owned, so it stays a contract instead of drifting into one more undocumented format.
  • Enterprise data warehousing, source systems landing in a canonical layer before reaching the dimensional model, keeping downstream reporting insulated from upstream schema changes.
Pattern 02 — FIG. 04

Same person, four records. One truth.

Four systems, four slightly different versions of the same person, plus one record that only looks similar. Entity resolution scores every pair and merges the ones above a threshold into a golden record. Drag the threshold and watch what happens on both sides of it.

CRM
J. SmithDOB 03/15/1980123 Main St
Reference record
EHR
Jonathan SmithDOB 03/15/1980123 Main Street
92 match
Billing
Jon SmithDOB 03/15/1980123 Main St, Apt 2
88 match
Portal
John SmythDOB 09/22/1975456 Oak Ave
35 match
2
Records merged into golden record
Correct: both duplicates caught, the unrelated record stayed out.
False-merge rate Missed-duplicate rate

Too loose a threshold and unrelated people merge; too strict and real duplicates stay split. The marker tracks the slider above, moving through the sweet spot where both error rates stay low.

Pattern 02 · Delivered — FIG. 05

Where this shows up in the work.

  • MDM and EMPI modules, personally architected and built: entity lifecycle, merge/unmerge workflows, golden record management.
  • Probabilistic and deterministic patient identity resolution, matching and deduplicating patient records across payer, clearinghouse, and EMR/EHR sources that never share one key.
  • Healthcare payer integrations, patient matching across claims, EMR, and clearinghouse data (Availity, Edifecs) where the same patient rarely arrives with the same identifier twice.
Pattern 03 — FIG. 06

Real-time isn't faster batch. It's no batch.

A nightly batch job doesn't make data late by a few minutes, it makes data late by up to a full day, on a sawtooth that resets once every 24 hours. Change data capture and event streams replace the sawtooth with a flat line. Toggle between them below.

Source system changes, over 24 hours
12 AM6 AM12 PM6 PM12 AM
12.0h
Average data staleness
23.8h
Worst-case staleness
Nightly batch Event-driven CDC

Batch staleness climbs for 24 hours and resets once at the nightly run. CDC staleness is bounded by propagation latency, typically seconds to a few minutes, all day. Move the pointer over the chart to compare any hour.

Pattern 03 · Delivered — FIG. 07

Where this shows up in the work.

  • Event-driven ingestion, Azure Event Hub, SFTP processing, and Blob Storage triggers automating file ingestion and near-real-time data movement across business systems, replacing nightly exports.
  • Real-time platform integrations, project management, accounting, and preconstruction systems unified into one centralized data layer instead of nightly batch loads.
  • AI agents watching every hop, detecting and correcting file-level errors in real time as data moves, instead of discovering problems the next morning.
Pattern 04 — FIG. 08

One join, not six.

A normalized schema is correct and it's also a maze: a simple sales-by-category report walks through half a dozen tables to get an answer. A Kimball star schema denormalizes on purpose, one fact table surrounded by the dimensions people actually filter and group by. Toggle the diagram below.

Normalized schema versus a Kimball star schema A sales schema shown two ways: a normalized mesh of nine linked tables, or a star schema with one fact table and five dimension tables.
9
Tables in the schema
5
Joins for a typical report
Normalized Star schema
Sales by product category, last quarter
5
1
Revenue by store region
4
1
Promotion effectiveness by customer segment
6
2

Joins required for three typical reports, walking the normalized foreign-key chain versus querying the star directly.

Pattern 04 · Delivered — FIG. 09

Where this shows up in the work.

  • Multi-dimensional data models, ETL via SSIS, OLAP cubes via SSAS, delivered on top of enterprise data warehouses built from the ground up.
  • Kimball-style dimensional modeling, alongside tabular models, chosen deliberately over pure normalization for reporting and analytics workloads.
  • Enterprise data warehousing, vision, scope, and system architecture defined and delivered to match organizational data strategy, star schema underneath the BI layer.
Pattern 05 — FIG. 10

More history, or a faster first report?

Data Vault 2.0 splits every business concept into three piece types: a hub for the business key, a link for how concepts relate, and a satellite for the descriptive attributes, each version kept, never overwritten. It's built for full history and source resilience. Select a piece below to see what it does.

Data Vault 2.0 structure: hubs, a link, and satellites Two hubs, Customer and Order, connected by a Places link. Each hub has its own satellite holding time-variant descriptive attributes.

A Hub stores only the business key, e.g. Customer ID, and never changes once loaded. It's the stable anchor every other piece attaches to.

Dimensional-first Data Vault-first
Time to first usable report
2w
6w
Engineering effort to absorb a source schema change
4d
1d

Illustrative, not a benchmark: the shape of the tradeoff holds across engagements. Dimensional gets a report in front of a stakeholder fast; Data Vault costs more upfront and pays it back the first time a source system changes shape.

Pattern 05 · Delivered — FIG. 11

Where this shows up in the work.

  • Evaluated as a consultant, Data Vault 2.0 modeled as a candidate integration layer and weighed directly against time-to-first-report for BI stakeholders.
  • Chose dimensional modeling for delivery, a deliberate tradeoff: fast, demonstrable ROI mattered more than the audit depth Data Vault provides for the engagements it was weighed against.
  • Solved the same underlying problem a different way, the AI-first schema drift detection and auto-correction layer in the reference architecture addresses source-resilience directly, without a full historized raw vault underneath every mart.

Counting integrations instead of building on them?

Talk to us →