Skip to content

The data product matrix

One data product. Three engines. Seven orchestration idioms. The point is that the product stays the same while the platform changes, so a comparison between Fabric, Databricks and Snowflake is a comparison of engines rather than of who wrote the fixtures.

Four tiers, and why the split is load-bearing

Section titled “Four tiers, and why the split is load-bearing”

Three OpenAPI services and a Postgres database with a CDC change stream, plus the simulators and fixture generators that populate them. sources.yaml is a declaration of what the vendors are, not a compose file: each platform generates its own vendor stack from it.

It exists as a separate repo for one reason, and the repo states it plainly:

Two copies of a vendor is where a comparison dies.

If fabric-platform-airflow3 carried its own vendor fixtures and databricks-platform-jobs carried different ones, any difference in the gold numbers would be unattributable. Shared vendors make the difference mean something.

The transform logic (bronze and silver in Spark, gold in dbt SQL), the ODCS data contracts and their identities, the tests, and the expected numbers.

The expected numbers are the oracle that makes the whole matrix work. Every engine must reproduce revenue_usd 129,341,157.6700 across 474,044 sale lines. A platform that finishes green but produces different numbers has not run the product, it has run something else, and compare_products.py says so.

The core is consumed by tag, never by path. A path dependency is what broke one of the platforms: it makes the leaf silently track an uncommitted working tree, so the thing that passed locally is not the thing anyone else runs.

A leaf carries only what is genuinely per-platform: the DAG, the job specification, the notebooks or the task graph; the sink configuration; the dbt profile. Nothing else. Each leaf README says it:

It is a product, not a platform.

contoso-data-product-databricks-jobs, for example, is steps/ with one module per pipeline stage, each with its own main(), because that is how a Databricks team writes it. The transform logic inside those steps comes from the core.

Compose files, emulator pins (fabric by digest, entra, keyvault, arm, Spark, SQL Server, OpenMetadata), the vendor stack generated from contoso-sources, provisioning, connections, and a make up.

A platform contains no Contoso name and no product file:

The platform installs the product and knows no Contoso.

It takes PRODUCT=<path> and runs whatever it is handed. That is what makes a platform reusable for your own data product rather than a demo you have to gut.

EngineOrchestratorLeaf productPlatform
FabricAirflow 3contoso-data-product-fabric-airflow3fabric-platform-airflow3
FabricNotebooks + Data Pipelinescontoso-data-product-fabric-notebook-pipelinesfabric-platform-notebook-pipelines
FabricBuilt-in Airflowcontoso-data-product-fabric-airflow-builtinfabric-platform-airflow-builtin
DatabricksDatabricks Jobscontoso-data-product-databricks-jobsdatabricks-platform-jobs
DatabricksAirflow 3contoso-data-product-databricks-airflow3databricks-platform-airflow3
SnowflakeSnowflake Taskscontoso-data-product-snowflake-taskssnowflake-platform-tasks
SnowflakeAirflow 3contoso-data-product-snowflake-airflow3snowflake-platform-airflow3

✅ built · ⬜ reserved, holding a README and a LICENSE

Reserved repos hold a README and a LICENSE. They exist so the shape of the matrix is visible before every cell is filled: a reader can see what has been deliberately left undone, rather than guessing whether an absent cell is planned or impossible.

Snowflake Tasks completed its pair. It had a built platform against a reserved leaf, so the platform ran a gold-only slice rather than a product of its own; the leaf now carries the steps and its own CI. That makes the Snowflake column a like-for-like comparison with the other two engines rather than a partial one, and leaves Snowflake with Airflow 3 as the only cell with neither half built.

Statuses here are derived from what is on main, not from what a repository’s README says about itself. Several of those READMEs still describe themselves as reserved over a tree carrying dozens of files. The table is generated from members.json by scripts/render_tables.py, and CI fails if the committed copy drifts.

fabric-platform-notebook-pipelines is the fullest one: four real vendor sources through a complete medallion into a Fabric Lakehouse, a semantic model serving Power BI, and the lineage catalogued in OpenMetadata. It runs against a published fabric-emulator release, never a checkout, so anything that works there works for anyone, and with one flag it runs against real Fabric.

It is also the exception to the separation described above: it predates the platform/product split and carries its product inline, which is why its leaf cell is still reserved. Read it as the end-to-end proof rather than as the model of the platform/product boundary. For that, fabric-platform-airflow3 is the cleaner example.

fabric-platform-airflow3 is the platform reduced to its essence: Airflow 3 plus a Fabric target you can pin, with no DAGs, no dlt, no dbt and no Contoso name anywhere in it. It is the clearest demonstration of the platform/product separation.

databricks-platform-jobs runs the same product against databricks-emulator with Unity Catalog and OpenMetadata, using Databricks Jobs rather than an external orchestrator.

snowflake-platform-tasks is a gold-only consumer, switchable between the emulator and a real Snowflake account with SNOWFLAKE_TARGET.

Porting a data product across engines is normally a project with a business case attached. Here it is a matrix cell, because the expensive part, finding out what each platform actually requires, happens against an emulator at seconds per attempt rather than against three paid accounts at minutes per attempt.

That is the concrete form of the argument in Building with AI agents: the emulators did not make the porting easy, they made it affordable enough to attempt, and the expected-numbers oracle made the result checkable without a human reading dataframes.