Skip to content

Testing with the emulator

The emulator’s reason to exist is determinism real Fabric cannot offer. Two levers do the work, both on the /_emulator control surface (control-plane origin, unauthenticated, not part of the Fabric contract).

Every async operation gets a completeAt on a virtual clock. Nothing sleeps; nothing is flaky.

Terminal window
GET /_emulator/clock # { offset, frozen, now }
POST /_emulator/clock # any of:
{ "freeze": true } # stop time
{ "advance": 3600 } # jump forward (seconds)
{ "offset": 0, "freeze": false } # reset / resume

The pattern for testing a polling loop:

POST /_emulator/clock {"freeze": true}
start the emulator with -lro-delay 600 # operations stay Running 600 virtual seconds
POST /v1/workspaces/{id}/items {…} # → 202, poll → Running forever
POST /_emulator/clock {"advance": 601} # time passes instantly
poll again # → Succeeded

With the default -lro-delay 0, operations complete on the next poll — fast CI without giving up the 202 contract.

Terminal window
POST /_emulator/faults
{ "failNextOperations": 1 } # next N async operations end Failed (Fabric-shaped error body)
{ "rejectNextRequests": 2 } # next N API requests get a 5xx
{ "lroDelaySeconds": 30 } # override the delay at runtime

This is how retry logic, poll-until-failed branches, and error surfaces get tested without patching the client under test.

GET /health → { "status": "ok", "now": … } — what the Docker HEALTHCHECK (the healthcheck subcommand) and compose depends_on gates use.

Testing your own code against the emulator

Section titled “Testing your own code against the emulator”
  • In-process (Go): server.New(cfg, …) + httptest — the emulator’s own integration tests run this way, including with a real in-process entra-emulator minting tokens. No network, no fixtures.
  • Over HTTP (any language): start the family with docker-compose, mint seeded tokens (quickstart), drive the API. -data-dir empty means each run starts clean.
  • Real tools unmodified: see testing with fabric-cicd.

Every package covers itself (90% floor cross-package, currently ~90%), on Linux, macOS, and Windows. The full matrix of what CI verifies — including the real-tool e2e — is in 12-e2e-matrix.md.

The REST surface has four contract gates of its own, each asking a question the other three structurally cannot: whether a documented operation is served, refused or silent; whether a served route has ever seen recorded traffic; whether a route this emulator answers is documented by Microsoft at all; and whether the body it returned matches the schema. Three of them write a ledger under docs/, and what each one counts — with the denominators, the --update commands, and what each explicitly does not say — is 65-api-surface-coverage.md.

When a suite’s cases are a table — many rows, one procedure — the rows live in cases/<suite>.json and the procedure stays in code. Every runner that executes the cases reads the same file, so a case cannot be edited in one runner and left stale in another.

@pytest.mark.cases("agent-consumer-contract")
def test_the_agent_honours_the_shape_a_consumer_sends(case):
... # one procedure; `case` is one row of the file

python/tests/conftest.py parametrizes a test marked cases(<suite>) over the suite and names each run by the case’s id, so pytest -k <id> selects one case and a failure names the row to open. A runner outside pytest reads the file through scripts/casefiles.py, which is standard-library only so an e2e runner can import it without the test venv.

Every case has a kebab-case id, unique in its suite, and a why: the break the case guards. JSON has no comments, so why carries what a comment above a row used to, and make check refuses a case without one. A case that another runner must also execute lists that runner in executed_by; the suite’s own test holds the runner to exactly that list.

The first suite is the spark agent’s consumer contract (docs/20): test_agent_consumer_contract.py checks that the agent recognises every case, and e2e/agent-contract/run.py executes the ones that name it against the built image. Before, the e2e runner typed its own copies of those statements.

The second is Fabric IQ MCP’s tool calls (cases/fabric-iq-tool-calls.json), read by a Go test and a Python driver. Each case is one call, made as the fixture’s owner or viewer, and either refused with a message or answered with a list of expect entries. Each entry is a path into the tool’s JSON reply plus one of equals, contains, length or order:

{"at": ["Results", 0, "Rows", "*", "Store[Territory]"], "equals": ["West"]}

* collects the rest of the path from every element of a list, and a path that does not resolve equals nothing, so a renamed field fails the case rather than matching null. The runners create their own fixtures, so {model}, {report} and {workspace} stand in for ids learned at run time. internal/api/fabriciq_cases_test.go runs every case in-process, and e2e/mcp-fabriciq/driver.py runs the ones that name it as entra-emulator’s seeded users, then checks that it ran as many as name it. The expectation vocabulary is implemented twice, in scripts/casefiles.py and in the Go test. Each has a test that runs the same table of expectations, some that hold and some that fail. The table is copied into both test files, so a change to the vocabulary has to change both copies.

Core MCP’s e2e stays as code: e2e/mcp-core/driver.py is a single story, where each step uses the id the step before created. That is the “when not to” below.

When not to. Move a suite to a case file when three or more cases run the same procedure, or when a second runner needs the same cases. A test whose body is its own argument — one situation, one reason — stays a function; turning it into a row hides the reasoning that makes it worth having.

The docs lane — what a documentation change runs

Section titled “The docs lane — what a documentation change runs”

A pull request confined to docs/, website/ or a root readme runs two CI jobs instead of ninety-odd. scripts/docs_only_change.py classifies the diff, and every job in ci.yml except witnesses is gated on the answer.

A push to main always runs everything, whatever it touched. The lane originally applied there too, and that was a mistake worth recording: two docs-only commits landed, each reported CI success having run two jobs, and the last commit whose suite actually ran had failed. Main’s tip read green while four tests were broken underneath it, and the pull requests open at the time inherited a red baseline nobody had been told about. Each of those greens was locally true, which is what made the branch’s state so hard to see. The rule ci.yml’s concurrency block already states applies here as well: every commit on main gets its own verdict, because every one is shipped and each is what a bisect will later ask about. Iteration happens on the pull request anyway, so the saving that matters is kept and merging costs one full run, once.

witnesses is not gated, deliberately: it is the job that reads the docs — the parity claims and their witnesses, the sidebar, intra-doc links, the conformance matrix — so it runs on a documentation change and a code change alike. What the lane removes is the jobs that cannot observe a markdown file, not the ones that check it. A documentation pull request also builds the Starlight site (docs-site.yml now triggers on pull_request), which nothing did before: an Astro failure used to be found after merge.

Three properties are load-bearing, and each is there because of a specific way this could go wrong:

  • Unknown means run everything. An empty diff, a force push’s null before, a range git cannot resolve: all answer false. A classifier that guesses “docs” when it does not know turns a missing verdict into a green one.
  • A deletion or rename is never docs-only, whatever the paths say. internal/api/examples_readme_test.go reads examples/README.md and fails when a (../docs/<file>) link points at nothing, so removing a page can break a Go test. Adding or editing one cannot.
  • Skipped, not absent. The gate is a job-level if:, not paths-ignore: on the workflow, so a documentation pull request shows sixty-odd jobs explicitly declining to run and the classifier’s reasoning is in the run’s own log. A path-filtered workflow does not appear in the checks at all, and a green rollup that is green partly because things did not run is a false all-clear.

Not covered: make-targets.yml (three jobs) still runs on every change. It is also the release gate via workflow_call, and a gate that can be skipped is worth more care than three runner-minutes.

Seven times in one day, across two people working in parallel, a bug took the same shape: something was missing or invented, and every available check reported success. They are collected here because each was expensive, each looked different at the time, and the pattern is what makes the next one cheap.

Two directions. The first is an absence reported as a presence:

what was absentwhat every check said
a DESCRIBE returning zero rows for a table that has columnscorrect schema, no error — dbt read “this table has no columns”
a USING delta clause dbt never emitted, because delta is the one value its macro omitsthe config was demonstrably applied; that same value suppressed the clause
the np protocol, unregistered, so a pipe DSN parsed as TCPmsdsn.Parse returned no error and a usable-looking config
the lakehouse tables a reflection never loadedreflection reported success; the fingerprint said “already done”
an entra process that never bound its portthe health check passed — against a different service on that port
a keep-alive scoped to a step’s shell, dead before it was neededthe step ran green
lineage edges a whole warehouse build never recordeda clean run, no error, an empty graph

The second direction is worse and rarer: a fabrication reported as evidence. A lineage endpoint that paired reads against writes as a cross product recorded six well-formed bronze-to-silver edges where three of the movements never happened. Nothing was missing, the count went up, and the result looked more complete than the truth. No schema check can catch that; only comparing against what actually moved.

Two more, added after the first seven, because their mechanism is different enough to be worth separating.

Eight — the check that never ran, with silence where the signal would be. A README was added to one half of a pair that a check, landed forty minutes earlier, required to differ only in how silver is built. Every path the README referenced existed; the author verified the contents and not the context. A local run was silent because that check lived only in CI, so there was nothing to distinguish “no invariant applies here” from “the invariant lives somewhere I did not look”. The other seven are a check running and reporting wrongly; this is a check not running at all.

That is why make check exists and why make test depends on it. Both repo invariants — scripts/check_witnesses.py and scripts/check_example_parity.py — are stdlib-only and run in under a second, so there was never a reason for them to be reachable only from a workflow file. A check that exists only in CI will keep catching people after the fact.

Nine — a measurement at the wrong granularity, reported as the answer. Asked whether CI was green, one person queried the RUN and saw queued; another queried the JOB inside it and saw success. Both were reading the API correctly. The run was queued because one of its 55 jobs had not started; the job in question had run for seven seconds and passed. Two true measurements, in conflict, each reported as “the CI result” — and nothing in either output says which question it answers.

This is nastier than evidence that cannot discriminate, because the evidence discriminates perfectly; it just answers a question adjacent to the one asked. The same shape recurred twice more within the hour: git log --format='%an' used to attribute a commit, in a repo where every session commits under one author, so the signal had no discriminating power at all; and a parity check “verified” by moving a file on disk, when the check reads git ls-files and had not noticed. Before believing a measurement, say out loud which question it answers, and check that it is the one you asked.

Ten — a comment asserting a contract the code does not hold. Above the Status field on the event bus stood: “Field names match the job-instance wire shape, so a consumer reading the stream and a consumer polling the API see the same words.” The names do match. The VALUES do not: the bus publishes Started, which StatusAt never returns — polling the same job says InProgress. A consumer written by reading that comment classified every job start as an unknown status, and its report was wall-to-wall false positives with real failures buried inside them.

The code was right; an event log and a state query legitimately differ. The prose was wrong, and prose is what a downstream author reads. A comment that states an invariant is a claim, and an unenforced claim decays into a lie without anything going red.

Its twin is a DEFAULT that converts an omission into an assertion. CreateLineageEdge fills an empty producer with Copy — and Copy means the emulator watched the bytes move. Five of seven call sites named their producer; two relied on the default. A new caller that simply forgot would publish an edge claiming evidence nobody ever had, into the one structure whose entire purpose is telling evidence from claim. Both copy sites now say Copy out loud, and TestEveryLineageEdgeStatesItsProducer keeps the default a backstop rather than a mechanism.

The tell for both: ask what would go red if the sentence were false. For the comment, nothing did — until a consumer believed it. For the default, nothing did — because being wrong and being defaulted are indistinguishable downstream.

Eleven: a BOUND that truncates instead of refusing. io.ReadAll(io.LimitReader(body, max)) is the obvious way to cap a read, and on a write path it is data corruption. LimitReader reports clean EOF at the ceiling: the excess is discarded, err is nil, and the handler cannot tell a body that fitted from one that was cut. It stores the fragment and answers success.

Nine sites had it. The one that mattered was OneLake’s append, and it took Microsoft’s fab cp to reach it — every client this project drives chunks its uploads, so none had ever crossed the ceiling. fab sends a whole file in one append, and a 71 MiB upload was stored as 64 MiB with a 202. It surfaced ONLY because fab then flushed at the real length and a position check disagreed. A client that never flushed would have had a short file and no signal.

Several sites looked safe because they parse the body afterwards, so a truncated JSON document 400s. That is luck wearing the costume of a design: the caller is told their input is malformed when what happened is that it was too big, and the first site that stops parsing inherits the silent version.

The tell is the same as ever — ask what would go red if the sentence were false. “This read is bounded” was true. “This read is safely bounded” was not, and nothing anywhere could tell the difference.

What closed it was not nine fixes. It was internal/httpx, one ReadBounded that reads max+1 so “too big” is detectable, plus a test that walks the source and fails on the banned idiom — with an explicit bounded-read-exempt: marker for the two places that genuinely discard what they read, and a second test pinning the exemption list so the escape hatch cannot quietly widen. Fixing nine sites is a day’s work that lasts until someone writes the tenth.

Twelve: PROSE that lies about code. Every item above is code lying about code. This one is documentation, and it is the cheapest of the twelve to produce: a renamed directory breaks a Go import and CI goes red, while the same rename inside a sentence breaks nothing at all, so the sentence keeps its confident tone and quietly starts lying. Three had drifted before anyone looked. docs/31 pointed at the medallion example’s common.py for the failure reporting that motivates the entire document — the file had moved into the contoso fixtures, so the paragraph explaining WHY the feature exists pointed at nothing. docs/38 named the warehouse reader as an e2e directory; that reader is real and runs on every push, but it is a go test job and no such directory has ever existed for anyone to go look in. docs/13 named a file that exists — in the SIBLING entra-emulator repo, read here as one of ours.

None of the three is a typo. Each is a fact that expired, and the only thing that would ever have caught them is somebody happening to click.

scripts/check_doc_drift.py now asserts five things across the prose, in make check and in the witnesses job: a backticked repo path exists, a make <target> names a target the Makefile defines, a documented FABRIC_/ENTRA_-style variable is read by some non-Markdown file, a backticked Go package/symbol reference names code that exists, and an ordinary Markdown link points at a tracked file, directory or simple heading anchor. The path and Markdown-link classes have found live drift; the others land as regression guards — and because a check that always passes is indistinguishable from a check that is working, every class is tested against synthetic drift rather than trusted on the strength of a green tree. That is item Eight’s lesson applied to the fix for item Twelve.

It is tuned for precision over recall, because a checker that cries wolf gets muted, and a muted check is item Eight again. make is read only from command-shaped code spans: the naive bare-word regex returns 13 hits on this tree and all 13 are English prose (“make the”, “make it”, “make every”). A dot that is not a known file extension means “not a path”, so the Go symbol internal/tsql.DataFlows is left alone. docs/24 is this repo’s shorthand for a NUMBERED DOCUMENT and never a path. Markdown heading anchors are checked only when the target file exists and the fragment is a simple GitHub-style heading slug, so line anchors and generated ids are not treated as promises this checker can verify. Release notes are skipped outright: a v0.16 note naming a since-renamed file is correct about the tree at that tag, and editing it to please a checker would be falsifying a historical record.

The prose it reads is every docs/ page, the root readme, and every per-directory README — beside an example, a suite, a package. That last group was missing from the first version, and it is the prose most likely to cite a path inside its own directory: adding it immediately turned up a fourth stale reference: e2e/dbt-fabric/README.md sent the reader to a parity document under a number that document has never carried — it is docs/parity.md. (The dead name is written without backticks here on purpose. Quoting it as a code span makes this paragraph itself a finding, which is how the first draft of it failed the check.) The list is intersected with what git tracks rather than taken from the glob alone, because a working tree holds a great deal that is not this repo’s prose — a .venv/ under an example, node_modules/, a worktree copy of everything — and drift reported inside a vendored README is somebody else’s documentation, which is the fastest possible way to teach a reader to skim past this check.

Existence is asked of git, never of the filesystem. Path.exists() answers case-INSENSITIVELY on a default macOS volume and case-sensitively on the Linux runner, so a wrong-case path would pass make check on a laptop and fail in CI — a guard that disagrees with itself by platform is worse than one that is merely wrong, because the disagreement is what teaches people to distrust it.

A document may legitimately name something not built yet. That gets an EXEMPT entry carrying a written reason — docs/30’s planned contracts checker, docs/54’s task-parameters suite — rather than an edit that waters true prose down into vague prose. A forward reference is a claim about intent, and recording it is what keeps it reviewable instead of invisible. A test asserts every exemption is still doing work, so the escape hatch cannot quietly widen into a silencer.

Both checkers above stop at docs/. A Go comment is where this repo keeps its reasoning — the cause of a refusal, the oracle a list came from, the precedent an activity was allowed under — and nothing read one until scripts/check_comment_drift.py landed.

The motivating failure is the same shape as the three above. #490 measured the leaf/compute distinction and retired it: no oracle carries a connector activity type, so the dispatch default was refusing nothing on behalf of a leaf that does not exist. docs/parity.md and docs/43 were corrected in that change; two source comments were not, and went on teaching the retired justification in the PRESENT TENSE for a week — one of them asserting the default “is right for a CONNECTOR LEAF”, directly beside the code that had stopped doing that.

Three classes, on the same precision-over-recall argument. A repo-rooted path must exist, and a docs/NN reference — written as a bare number, the way this repo’s comments write it — must name a document that does. (That shorthand is deliberately not a code span here: quoting it as one makes this very paragraph a finding, which is how the first draft of it failed the check.) Both are clean today and land as regression guards — 113 paths and 140 doc references resolve, and the doc reference is the one Go comments make most, so a renumbered document would break every one of them silently. The narrowing is what makes them usable: the naive form of the path class was measured first and returned 84 findings, every one a false positive — Microsoft Learn slugs like onelake/onelake-access-api.md, payload filenames like data.json that belong to the caller’s data rather than to this tree, and foreign paths like ADF’s entityTypes/Pipeline.json. Rooting the class at a tracked top-level directory took it to nought, and the citations it gave up were never a drift checker’s to make. Comments are found by a hand-walked scanner over Go’s four states rather than by a regex, because "https://learn.microsoft.com/..." contains //: the regex form reads the second half of every documentation URL in the tree as a comment.

The third class is a pin on retired vocabulary. A term recorded in RETIRED carries the reason it was retired and the count of occurrences the tree is known to hold, all of which are past-tense records that the idea was held and dropped. A further occurrence fails, and is either another such record — raise the pin in the same commit, and the raise is the review — or the idea creeping back. Counts rather than line numbers, because a line number moves whenever anything above it is edited, and a check that fails for an unrelated edit is one people learn to re-baseline without reading.

What it cannot do is the interesting part, so it is written down rather than implied. It catches a retired term being reintroduced. It cannot catch a justification written today going stale tomorrow, because nothing mechanical separates a true is from a false one — the stale comment above read “the run really DID reach the leaf”, so even a past-tense heuristic would have waved it through. Retiring a concept is a deliberate act; this asks only that the act be recorded once so the tree cannot drift back to it.

Everything above is a checker reading the tree. Nothing read the checkers.

scripts/ is where this repository keeps its enforcement: make check runs thirty-one of these scripts and the witnesses job runs thirty. Between them they assert that every supported parity claim names a witness, that no route the emulator serves is undocumented, that the sidebar is complete, that no Go comment teaches a retired justification, and that no test sleeps without a bound. A guard script whose own behaviour nothing asserts has the property this whole document is about, one level in: its detection can stop matching — a renamed directory, a tightened regex, a refactor that drops a branch — and it will go on printing its success line and exiting 0 forever. A check that passes is indistinguishable from a check that is running, and that sentence is as true of the checker as of the thing it checks.

The measured gap when this landed: 11 of 51 scripts had no dedicated test module, the largest scripts/govern_ingest.py at 717 lines, then scripts/check_cron_workflow_freshness.py at 289 and scripts/build_fixture_wheels.py at 220. The Go surface was measured at the same time for comparison and has no structural gap worth reporting — all 27 packages under internal/, pkg/ and cmd/ carry a test file, across 209 non-test source files — so the deficit was here and here alone.

That the risk is real rather than theoretical is already written down in this tree, twice, in the checkers’ own words. scripts/check_doc_drift.py records that the naive form of its path class returned 84 findings, every one a false positive. scripts/check_python_test_flakiness.py records a kind missing from its match key, which exempted 42 of its 44 recorded symbols from the two bans the ledger exists to enforce — a bare sleep before an assertion printed accepted and passed --strict. Both were found by someone driving the checker against a violation it had to catch, which is to say by a test, and neither would have been found by reading it.

scripts/check_script_test_coverage.py asserts the list instead of the instances, for the reason python/tests/test_make_check_runs_in_ci.py gives about its own subject: fixing eleven files leaves the shape intact, and the twelfth script lands with no test and nothing says so. Two finding kinds:

  • UNTESTED — a script with neither a dedicated test module nor a ledger entry. This is what the ledger exists to keep shrinking.
  • STALE — an entry naming a script that no longer exists, or one that has since gained a test. A closed gap must not linger as an accepted one: the entry would silently re-cover the file if that test were later deleted, so the ledger would absorb a real regression without a word.

Both directions, like docs/test-flakiness.json, docs/python-test-flakiness.json and docs/vitest-test-flakiness.json before it. A one-directional ledger only ever grows. Of the two kinds, STALE is the one that needs a checker: an unrecorded script is loud by construction — somebody adds a file and the build goes red the same day — while a stale entry is green, silent, and re-arms itself.

The rule is a name: test_<stem>.py under python/tests/ exists. Not a coverage measurement — coverage is already measured and gated at 92% (fail_under, pyproject.toml), and what a name adds is the one thing a percentage cannot say, that this particular file has somewhere for its violations to be driven from. It is also why indirect coverage does not count. Three of the recorded scripts are exercised, under another file’s name: scripts/check_cron_workflow_freshness.py by python/tests/test_cron_workflow_freshness.py, which simply drops the check_ prefix; scripts/vendor_notebookutils_stubs.py through the surface checker that reads what it vendored; scripts/govern_ingest.py incidentally by two tests of logic that was lifted out of it. Each entry records where the behaviour actually is exercised, so the ledger says what is true instead of flattening “covered elsewhere” into “not covered”.

The ledger shipped smaller than the gap it records, which is the part that makes it a ratchet rather than a waiver: eleven measured, three closed with real unit tests, eight recorded. The three were chosen because they are pure logic over a tmp_path fixture with no network and no Docker — scripts/image_tags.py, scripts/check_image_digests.py, scripts/check_cask_stanzas.py. Its own test module pins the count as a ceiling, so closing another gap passes and adding an exemption without closing anything does not; raising it is done in the change that argues for the new entry, and the raise is the review.

Writing one of those three found a live defect, and its shape is this document’s first category exactly. scripts/check_image_digests.py declares its scope as “only files that can pull — compose and env”, and had listed .env in that scope since the day it was written. The line never matched a single file: Path(".env").suffix is the empty string, because a leading dot makes the whole name a stem. So the canonical name — the one compose reads, the one env_file: defaults to — was out of scope while the constant said it was in, and examples/fab-driven/.env is tracked here and had never been inspected. An unscanned file produces no finding and no error, so it reads exactly like a clean one. No amount of reading either the checker or the tree would have said so; only pointing it at a file that must be reported did.

Recording a new exemption is one entry in docs/script-test-coverage.json carrying the script path, its line count at the time, the reason, and where its behaviour is exercised if it is exercised elsewhere. The reason is required and an empty one is refused by name — an exemption with no reason is an omission wearing a decision’s clothes. Writing the test instead is always the cheaper long-term answer, and the ceiling above is there to keep that true.

Not review, and not more assertions. In every case it was looking at what the thing produced rather than at whether it ran without complaint:

  • reading dbt’s compiled SQL instead of the Jinja that generated it;
  • listing the edges instead of trusting the edge count;
  • printing the parsed DSN (Host "np:", Protocols [tcp], err = nil);
  • docker ps, which settled a day-long misdiagnosis in one line;
  • a log.Printf compiled into the binary, which settled whether a fix was running at all after docker compose ps and docker inspect had both reported the tag and image id expected — for three runs against a stale image. Verify the running code, not the label on it.

A test that cannot fail is worse than no test, because it certifies the area as covered. A parse-only test sat green next to a dial that could not connect; the handler tests injected path values and would have passed with every route mis-registered. Check a new test fails when you break the thing it guards — several here did not, and the guard for the port collision above passed its own review while sailing straight past the collision it was written for.

Prefer a loud failure to a plausible answer. Endpoints in this repo refuse rather than return an empty list where empty is indistinguishable from missing (GET /v1.0/myorg/datasets for a personal workspace), refuse a refresh that would re-read nothing, and refuse modifiedSince rather than silently doing a full pass. Where empty IS the truth — a model with no datasources — it is returned, and the difference is the point.

A skip is only honest while the assertion still runs somewhere. The medallion-compare job exists because compare.py skips when its counterpart is absent; delete the job and the comparison stops being made with nothing going red. Both carry a comment saying to remove them together or not at all.

Coverage: what the number covers, and what it cannot

Section titled “Coverage: what the number covers, and what it cannot”

Three measurements, because one would misrepresent the others:

What it measuresGate
Gounit + in-process server tests, merged with instrumented e2e runs90% floor, armed only where a real SQL Server was reachable
Pythonthe unit suite’s own scope (checkers, delta-ops, storage)92% floor (fail_under, pyproject.toml)
Witnessesevery supported parity claim names a test that existscheck_witnesses.py --strict

The witness count is the integration measure. “Every claim of support is backed by something that ran” is a statement no percentage can make, which is why it is published beside the percentages rather than folded into them.

An e2e suite can contribute real coverage, because the emulator can be built with Go’s -cover and the counters merged with the unit ones:

Terminal window
scripts/coverage_prepare.sh # chmod 777, or the container writes nothing
FABRIC_COVERAGE=1 uv run --frozen --no-sync python e2e/sail/run.py
go test -cover -coverpkg=./... ./... -args -test.gocoverdir=$PWD/covdata/unit
scripts/coverage_merge.sh

Coverage is requested explicitly, not through addopts:

Terminal window
uv run --frozen --group test pytest python/tests -q --cov --cov-report=term-missing

addopts would apply to every pytest in the repo, including the ones e2e harnesses run inside containers that have pytest but not pytest-cov — where the flags are unrecognised arguments and the suite dies with exit 4 for a reason that has nothing to do with what it was testing.

The floor itself is not on that command line. It is fail_under in pyproject.toml (92, with precision = 2 so the comparison is exact), and it lives in one place for the reason the setting’s own comment gives: as --cov-fail-under= in two workflow steps it was two things to move for one change, which is how a ratchet quietly stops ratcheting.

Measure the floor in a clean venv. --frozen --group test is what CI runs, so it is the only environment whose number is authoritative. A venv that has drifted — anything installed with an extra group or uv run --with persists — can read a point lower, because code that only runs when a dependency is absent stops running once something installs it. scripts/check_govern_types.py did exactly that: 98% clean, 85% with urllib3 present, one point on the total. That was measured against a 70% floor, where a point was noise; against today’s 92 it is most of the headroom, since the suite measures ~93. The lesson generalises past that file — if a test’s coverage depends on what happens to be installed, force the branch instead of inheriting it, or the gate fails on whoever pushes next rather than on whoever spent the margin. That drift happened to read low, which is the safe direction; the same mechanism reading high would hide a real regression under a floor that passes.

FABRIC_COVERAGE=1 makes the harness layer e2e/docker-compose.coverage.yml, which builds the emulator with COVER=1 and mounts one repository-level covdata/. Every wired suite writes into that same directory, so the merge sees the whole fleet without needing a list of which suites ran.

Two things must both hold, and neither is obvious:

  1. The binary must be built instrumented. COVER=1 does that; the published image is never built this way.
  2. covdata/ must be writable by uid 65532. The image is distroless nonroot and a bind mount keeps the HOST’s ownership, so a directory created by the checkout user is not writable by the container — and Go’s coverage runtime does not complain, it just writes nothing. Docker Desktop on macOS ignores the uid mismatch entirely, so this passes on a laptop and fails only on Linux CI. scripts/coverage_prepare.sh does the chmod; e2e/engine-matrix hit the identical trap first.
  3. The process must exit CLEANLY. A -cover binary writes its counters when main returns, and a SIGKILLed one writes nothing — an empty covdata/ reads as “the e2e exercised nothing” rather than “nobody asked it to say”. cmd/fabric-emulator handles SIGTERM for exactly this reason, and a stop it was asked for returns nil rather than the closed-listener error that would send main through log.Fatal and os.Exit — which skips the write anyway.

Measured contribution: cmd/fabric-emulator goes from 0% on signalStop to 100% once an instrumented stack has been started and stopped, and the package reaches ~70% from e2e alone. The merged total moves less than the per-package figures do, because the unit suites already cover most paths — the e2e legs prove the wiring (compose, TLS, the Entra handshake, the engines) that unit tests deliberately stub.

Wiring a suite that is not yet instrumented is one conditional in its run.py: append -f e2e/docker-compose.coverage.yml when FABRIC_COVERAGE is set. A suite whose stack contains no fabric-emulator service has nothing to instrument and is deliberately left alone.

Every JS entry point — installs, scripts, CI — goes through pnpm, never npm or yarn:

Terminal window
pnpm install --frozen-lockfile
pnpm --filter fabric-emulator-portal test
pnpm --filter fabric-emulator-docs build

The workspace pins packageManager: pnpm@10.29.3 and every package.json carries preinstall: npx only-allow pnpm. Every one, not just the root — the root guard does nothing for cd portal && npm install, which reads that directory’s own manifest, finds no guard, resolves its own tree, and writes a package-lock.json. Two lockfiles is worse than the wrong one: each tool reads its own, both report success, and the divergence surfaces later as a version that is somehow different in CI with nothing in the diff to explain it.

npx only-allow pnpm, and that npx is not an oversight to tidy into pnpm dlx: the script has to run under the package manager being refused. Someone typing npm install has npx and may not have pnpm at all.

Enforced by three guards in internal/repo, because a convention with nothing checking it is the subject of most of this document: every manifest carries the guard, no rival lockfile is committed, and no workflow or script shells out to npm/yarn. The last one has a leading word boundary — npm install is a substring of pnpm install, and the first version failed on the two CI lines that were already correct.

Every Python entry point — e2e runners, scripts/, Makefile targets, CI steps — goes through uv, never a bare python3:

Terminal window
uv run --frozen --group <group> python e2e/<suite>/run.py # suite needs deps
uv run --frozen --no-sync python e2e/<suite>/run.py # stdlib-only driver

Most host-side runners only drive docker compose and are stdlib-only, so they take --no-sync; the nine that import real client libraries name their group.

This is not style. A bare python3 resolves to whatever is on PATH, which on a developer machine is often a pyenv build with none of the project’s dependencies — e2e/adls-sdk fails with ModuleNotFoundError: No module named 'azure' that way, a harness fault that reads exactly like a test failure.

.python-version pins 3.12, matching requires-python, the python:3.12-slim images, and CI’s setup-python. Without it uv satisfies >=3.12 with the newest interpreter it can find (3.13 at the time of writing), so the host would quietly run a different Python from the containers under test. Note that uv run --no-sync reuses a mismatched existing .venv with only a warning rather than rebuilding it — if the warning appears, rm -rf .venv and re-sync.

Ask what your machine has that a fresh one does not

Section titled “Ask what your machine has that a fresh one does not”

Measure the floor in a clean venv states this for the coverage number. It is not a coverage phenomenon, and it is filed where nobody looks when the symptom is a failing test or a 401 — so: before asking “is main broken?”, ask “what does my machine have that a fresh one does not?” In one day that question would have ended three investigations before they started, none of which looked like environment drift:

  • A combination no lockfile produces. deltalake present and pandas absent — no group declares that pair — took test_delta_ops down a path CI never runs, dying in pyarrow’s pandas shim. In CI deltalake is absent, so the test skips and the path is never reached.
  • A coverage percentage a point low, from a venv no longer matching the lockfile — the case the coverage section already describes.
  • A 401 that was a clock. A container left up for hours outlived several emulator restarts, and its token aged past a threshold. Nothing about it looks like environment drift until you find the clock, which is why it is the one worth remembering: the state that drifted was not a package at all.

The shape is the same each time — the failure depends on state nobody declared — but only the first is a venv, so rm -rf .venv is a check, not the check:

Terminal window
gh run list --workflow ci.yml --branch main --limit 3 # is main actually red?
rm -rf .venv && uv run --frozen --group test pytest python/tests -q
docker compose down -v # for anything long-lived

A green CI on the commit you are sitting on does not make the local failure imaginary — it was real about an environment. It means the difference is yours to find before it is anyone else’s to chase.