Testing with the emulator
The emulator’s reason to exist is determinism real Fabric cannot offer.
Two levers do the work, both on the /_emulator control surface (control-plane
origin, unauthenticated, not part of the Fabric contract).
The clock — LROs on demand
Section titled “The clock — LROs on demand”Every async operation gets a completeAt on a virtual clock. Nothing
sleeps; nothing is flaky.
GET /_emulator/clock # { offset, frozen, now }POST /_emulator/clock # any of: { "freeze": true } # stop time { "advance": 3600 } # jump forward (seconds) { "offset": 0, "freeze": false } # reset / resumeThe pattern for testing a polling loop:
POST /_emulator/clock {"freeze": true}start the emulator with -lro-delay 600 # operations stay Running 600 virtual secondsPOST /v1/workspaces/{id}/items {…} # → 202, poll → Running foreverPOST /_emulator/clock {"advance": 601} # time passes instantlypoll again # → SucceededWith the default -lro-delay 0, operations complete on the next poll — fast
CI without giving up the 202 contract.
Fault injection — the unhappy paths
Section titled “Fault injection — the unhappy paths”POST /_emulator/faults { "failNextOperations": 1 } # next N async operations end Failed (Fabric-shaped error body) { "rejectNextRequests": 2 } # next N API requests get a 5xx { "lroDelaySeconds": 30 } # override the delay at runtimeThis is how retry logic, poll-until-failed branches, and error surfaces get tested without patching the client under test.
/health
Section titled “/health”GET /health → { "status": "ok", "now": … } — what the Docker
HEALTHCHECK (the healthcheck subcommand) and compose depends_on gates
use.
Testing your own code against the emulator
Section titled “Testing your own code against the emulator”- In-process (Go):
server.New(cfg, …)+httptest— the emulator’s own integration tests run this way, including with a real in-process entra-emulator minting tokens. No network, no fixtures. - Over HTTP (any language): start the family with docker-compose, mint
seeded tokens (quickstart), drive the API.
-data-dirempty means each run starts clean. - Real tools unmodified: see testing with fabric-cicd.
How the emulator tests itself
Section titled “How the emulator tests itself”Every package covers itself (90% floor cross-package, currently ~90%), on Linux, macOS, and Windows. The full matrix of what CI verifies — including the real-tool e2e — is in 12-e2e-matrix.md.
The REST surface has four contract gates of its own, each asking a question the
other three structurally cannot: whether a documented operation is served,
refused or silent; whether a served route has ever seen recorded traffic;
whether a route this emulator answers is documented by Microsoft at all; and
whether the body it returned matches the schema. Three of them write a ledger
under docs/, and what each one counts — with the denominators, the --update
commands, and what each explicitly does not say — is
65-api-surface-coverage.md.
Test cases as data
Section titled “Test cases as data”When a suite’s cases are a table — many rows, one procedure — the rows live in
cases/<suite>.json and the procedure stays in code. Every runner that executes
the cases reads the same file, so a case cannot be edited in one runner and left
stale in another.
@pytest.mark.cases("agent-consumer-contract")def test_the_agent_honours_the_shape_a_consumer_sends(case): ... # one procedure; `case` is one row of the filepython/tests/conftest.py parametrizes a test marked cases(<suite>) over the
suite and names each run by the case’s id, so pytest -k <id> selects one
case and a failure names the row to open. A runner outside pytest reads the file
through scripts/casefiles.py, which is standard-library only so an e2e runner
can import it without the test venv.
Every case has a kebab-case id, unique in its suite, and a why: the break the
case guards. JSON has no comments, so why carries what a comment above a row
used to, and make check refuses a case without one. A case that another runner
must also execute lists that runner in executed_by; the suite’s own test holds
the runner to exactly that list.
The first suite is the spark agent’s consumer contract
(docs/20): test_agent_consumer_contract.py checks
that the agent recognises every case, and e2e/agent-contract/run.py
executes the ones that name it against the built image. Before, the e2e runner
typed its own copies of those statements.
The second is Fabric IQ MCP’s tool calls (cases/fabric-iq-tool-calls.json),
read by a Go test and a Python driver. Each case is one call, made as the
fixture’s owner or viewer, and either refused with a message or answered
with a list of expect entries. Each entry is a path into the tool’s JSON
reply plus one of equals, contains, length or order:
{"at": ["Results", 0, "Rows", "*", "Store[Territory]"], "equals": ["West"]}* collects the rest of the path from every element of a list, and a path
that does not resolve equals nothing, so a renamed field fails the case rather
than matching null. The runners create their own fixtures, so {model},
{report} and {workspace} stand in for ids learned at run time.
internal/api/fabriciq_cases_test.go runs every case in-process, and
e2e/mcp-fabriciq/driver.py runs the ones that name it as entra-emulator’s
seeded users, then checks that it ran as many as name it. The expectation
vocabulary is implemented twice, in scripts/casefiles.py and in the Go test.
Each has a test that runs the same table of expectations, some that hold and
some that fail. The table is copied into both test files, so a change to the
vocabulary has to change both copies.
Core MCP’s e2e stays as code: e2e/mcp-core/driver.py is a
single story, where each step uses the id the step before created. That is the
“when not to” below.
When not to. Move a suite to a case file when three or more cases run the same procedure, or when a second runner needs the same cases. A test whose body is its own argument — one situation, one reason — stays a function; turning it into a row hides the reasoning that makes it worth having.
The docs lane — what a documentation change runs
Section titled “The docs lane — what a documentation change runs”A pull request confined to docs/, website/ or a root readme runs two
CI jobs instead of ninety-odd. scripts/docs_only_change.py classifies the
diff, and every job in ci.yml except witnesses is gated on the answer.
A push to main always runs everything, whatever it touched. The lane
originally applied there too, and that was a mistake worth recording: two
docs-only commits landed, each reported CI success having run two jobs, and the
last commit whose suite actually ran had failed. Main’s tip read green while
four tests were broken underneath it, and the pull requests open at the time
inherited a red baseline nobody had been told about. Each of those greens was
locally true, which is what made the branch’s state so hard to see. The rule
ci.yml’s concurrency block already states applies here as well: every commit
on main gets its own verdict, because every one is shipped and each is what a
bisect will later ask about. Iteration happens on the pull request anyway, so
the saving that matters is kept and merging costs one full run, once.
witnesses is not gated, deliberately: it is the job that reads the docs — the
parity claims and their witnesses, the sidebar, intra-doc links, the
conformance matrix — so it runs on a documentation change and a code change
alike. What the lane removes is the jobs that cannot observe a markdown file,
not the ones that check it. A documentation pull request also builds the
Starlight site (docs-site.yml now triggers on pull_request), which nothing
did before: an Astro failure used to be found after merge.
Three properties are load-bearing, and each is there because of a specific way this could go wrong:
- Unknown means run everything. An empty diff, a force push’s null
before, a range git cannot resolve: all answerfalse. A classifier that guesses “docs” when it does not know turns a missing verdict into a green one. - A deletion or rename is never docs-only, whatever the paths say.
internal/api/examples_readme_test.goreadsexamples/README.mdand fails when a(../docs/<file>)link points at nothing, so removing a page can break a Go test. Adding or editing one cannot. - Skipped, not absent. The gate is a job-level
if:, notpaths-ignore:on the workflow, so a documentation pull request shows sixty-odd jobs explicitly declining to run and the classifier’s reasoning is in the run’s own log. A path-filtered workflow does not appear in the checks at all, and a green rollup that is green partly because things did not run is a false all-clear.
Not covered: make-targets.yml (three jobs) still runs on every change. It is
also the release gate via workflow_call, and a gate that can be skipped is
worth more care than three runner-minutes.
The failure this codebase keeps producing
Section titled “The failure this codebase keeps producing”Seven times in one day, across two people working in parallel, a bug took the same shape: something was missing or invented, and every available check reported success. They are collected here because each was expensive, each looked different at the time, and the pattern is what makes the next one cheap.
Two directions. The first is an absence reported as a presence:
| what was absent | what every check said |
|---|---|
a DESCRIBE returning zero rows for a table that has columns | correct schema, no error — dbt read “this table has no columns” |
a USING delta clause dbt never emitted, because delta is the one value its macro omits | the config was demonstrably applied; that same value suppressed the clause |
the np protocol, unregistered, so a pipe DSN parsed as TCP | msdsn.Parse returned no error and a usable-looking config |
| the lakehouse tables a reflection never loaded | reflection reported success; the fingerprint said “already done” |
an entra process that never bound its port | the health check passed — against a different service on that port |
a keep-alive scoped to a step’s shell, dead before it was needed | the step ran green |
| lineage edges a whole warehouse build never recorded | a clean run, no error, an empty graph |
The second direction is worse and rarer: a fabrication reported as evidence. A lineage endpoint that paired reads against writes as a cross product recorded six well-formed bronze-to-silver edges where three of the movements never happened. Nothing was missing, the count went up, and the result looked more complete than the truth. No schema check can catch that; only comparing against what actually moved.
Two more, added after the first seven, because their mechanism is different enough to be worth separating.
Eight — the check that never ran, with silence where the signal would be. A README was added to one half of a pair that a check, landed forty minutes earlier, required to differ only in how silver is built. Every path the README referenced existed; the author verified the contents and not the context. A local run was silent because that check lived only in CI, so there was nothing to distinguish “no invariant applies here” from “the invariant lives somewhere I did not look”. The other seven are a check running and reporting wrongly; this is a check not running at all.
That is why make check exists and why make test depends on it. Both repo
invariants — scripts/check_witnesses.py and scripts/check_example_parity.py
— are stdlib-only and run in under a second, so there was never a reason for
them to be reachable only from a workflow file. A check that exists only in
CI will keep catching people after the fact.
Nine — a measurement at the wrong granularity, reported as the answer.
Asked whether CI was green, one person queried the RUN and saw queued;
another queried the JOB inside it and saw success. Both were reading the API
correctly. The run was queued because one of its 55 jobs had not started; the
job in question had run for seven seconds and passed. Two true measurements,
in conflict, each reported as “the CI result” — and nothing in either output
says which question it answers.
This is nastier than evidence that cannot discriminate, because the evidence
discriminates perfectly; it just answers a question adjacent to the one asked.
The same shape recurred twice more within the hour: git log --format='%an'
used to attribute a commit, in a repo where every session commits under one
author, so the signal had no discriminating power at all; and a parity check
“verified” by moving a file on disk, when the check reads git ls-files and
had not noticed. Before believing a measurement, say out loud which question
it answers, and check that it is the one you asked.
Ten — a comment asserting a contract the code does not hold. Above the
Status field on the event bus stood: “Field names match the job-instance wire
shape, so a consumer reading the stream and a consumer polling the API see the
same words.” The names do match. The VALUES do not: the bus publishes
Started, which StatusAt never returns — polling the same job says
InProgress. A consumer written by reading that comment classified every job
start as an unknown status, and its report was wall-to-wall false positives with
real failures buried inside them.
The code was right; an event log and a state query legitimately differ. The prose was wrong, and prose is what a downstream author reads. A comment that states an invariant is a claim, and an unenforced claim decays into a lie without anything going red.
Its twin is a DEFAULT that converts an omission into an assertion.
CreateLineageEdge fills an empty producer with Copy — and Copy means the
emulator watched the bytes move. Five of seven call sites named their producer;
two relied on the default. A new caller that simply forgot would publish an edge
claiming evidence nobody ever had, into the one structure whose entire purpose
is telling evidence from claim. Both copy sites now say Copy out loud, and
TestEveryLineageEdgeStatesItsProducer keeps the default a backstop rather than
a mechanism.
The tell for both: ask what would go red if the sentence were false. For the comment, nothing did — until a consumer believed it. For the default, nothing did — because being wrong and being defaulted are indistinguishable downstream.
Eleven: a BOUND that truncates instead of refusing.
io.ReadAll(io.LimitReader(body, max)) is the obvious way to cap a read, and on
a write path it is data corruption. LimitReader reports clean EOF at the
ceiling: the excess is discarded, err is nil, and the handler cannot tell a
body that fitted from one that was cut. It stores the fragment and answers
success.
Nine sites had it. The one that mattered was OneLake’s append, and it took
Microsoft’s fab cp to reach it — every client this project drives chunks its
uploads, so none had ever crossed the ceiling. fab sends a whole file in one
append, and a 71 MiB upload was stored as 64 MiB with a 202. It surfaced ONLY
because fab then flushed at the real length and a position check disagreed. A
client that never flushed would have had a short file and no signal.
Several sites looked safe because they parse the body afterwards, so a truncated JSON document 400s. That is luck wearing the costume of a design: the caller is told their input is malformed when what happened is that it was too big, and the first site that stops parsing inherits the silent version.
The tell is the same as ever — ask what would go red if the sentence were false. “This read is bounded” was true. “This read is safely bounded” was not, and nothing anywhere could tell the difference.
What closed it was not nine fixes. It was internal/httpx, one ReadBounded
that reads max+1 so “too big” is detectable, plus a test that walks the source
and fails on the banned idiom — with an explicit bounded-read-exempt: marker
for the two places that genuinely discard what they read, and a second test
pinning the exemption list so the escape hatch cannot quietly widen. Fixing nine
sites is a day’s work that lasts until someone writes the tenth.
Twelve: PROSE that lies about code. Every item above is code lying about
code. This one is documentation, and it is the cheapest of the twelve to
produce: a renamed directory breaks a Go import and CI goes red, while the same
rename inside a sentence breaks nothing at all, so the sentence keeps its
confident tone and quietly starts lying. Three had drifted before anyone looked.
docs/31 pointed at the medallion example’s common.py for the failure reporting
that motivates the entire document — the file had moved into the contoso
fixtures, so the paragraph explaining WHY the feature exists pointed at nothing.
docs/38 named the warehouse reader as an e2e directory; that reader is real and
runs on every push, but it is a go test job and no such directory has ever
existed for anyone to go look in. docs/13 named a file that exists — in the
SIBLING entra-emulator repo, read here as one of ours.
None of the three is a typo. Each is a fact that expired, and the only thing that would ever have caught them is somebody happening to click.
scripts/check_doc_drift.py now asserts five things across the prose, in
make check and in the witnesses job: a backticked repo path exists, a
make <target> names a target the Makefile defines, a documented
FABRIC_/ENTRA_-style variable is read by some non-Markdown file, a backticked
Go package/symbol reference names code that exists, and an ordinary Markdown
link points at a tracked file, directory or simple heading anchor. The path and
Markdown-link classes have found live drift; the others land as regression
guards — and because a check that always passes is indistinguishable from a
check that is working, every class is tested against synthetic drift rather
than trusted on the strength of a green tree. That is item Eight’s lesson
applied to the fix for item Twelve.
It is tuned for precision over recall, because a checker that cries wolf
gets muted, and a muted check is item Eight again. make is read only from
command-shaped code spans: the naive bare-word regex returns 13 hits on this
tree and all 13 are English prose (“make the”, “make it”, “make every”). A dot
that is not a known file extension means “not a path”, so the Go symbol
internal/tsql.DataFlows is left alone. docs/24 is this repo’s shorthand for
a NUMBERED DOCUMENT and never a path. Markdown heading anchors are checked only
when the target file exists and the fragment is a simple GitHub-style heading
slug, so line anchors and generated ids are not treated as promises this checker
can verify. Release notes are skipped outright: a v0.16 note naming a
since-renamed file is correct about the tree at that tag, and editing it to
please a checker would be falsifying a historical record.
The prose it reads is every docs/ page, the root readme, and every
per-directory README — beside an example, a suite, a package. That last group
was missing from the first version, and it is the prose most likely to cite a
path inside its own directory: adding it immediately turned up a fourth stale
reference: e2e/dbt-fabric/README.md sent the reader to a parity document
under a number that document has never carried — it is docs/parity.md. (The
dead name is written without backticks here on purpose. Quoting it as a code
span makes this paragraph itself a finding, which is how the first draft of it
failed the check.) The list is intersected with what git tracks rather than
taken from the glob alone, because a working tree holds a great deal that is not
this repo’s prose — a .venv/ under an example, node_modules/, a worktree
copy of everything — and drift reported inside a vendored README is somebody
else’s documentation, which is the fastest possible way to teach a reader to
skim past this check.
Existence is asked of git, never of the filesystem. Path.exists() answers
case-INSENSITIVELY on a default macOS volume and case-sensitively on the Linux
runner, so a wrong-case path would pass make check on a laptop and fail in CI
— a guard that disagrees with itself by platform is worse than one that is
merely wrong, because the disagreement is what teaches people to distrust it.
A document may legitimately name something not built yet. That gets an EXEMPT
entry carrying a written reason — docs/30’s planned contracts checker, docs/54’s
task-parameters suite — rather than an edit that waters true prose down into
vague prose. A forward reference is a claim about intent, and recording it is
what keeps it reviewable instead of invisible. A test asserts every exemption is
still doing work, so the escape hatch cannot quietly widen into a silencer.
The same decay, one directory over
Section titled “The same decay, one directory over”Both checkers above stop at docs/. A Go comment is where this repo keeps its
reasoning — the cause of a refusal, the oracle a list came from, the precedent
an activity was allowed under — and nothing read one until
scripts/check_comment_drift.py landed.
The motivating failure is the same shape as the three above. #490 measured the
leaf/compute distinction and retired it: no oracle carries a connector activity
type, so the dispatch default was refusing nothing on behalf of a leaf that does
not exist. docs/parity.md and docs/43 were corrected in that change; two
source comments were not, and went on teaching the retired justification in the
PRESENT TENSE for a week — one of them asserting the default “is right for a
CONNECTOR LEAF”, directly beside the code that had stopped doing that.
Three classes, on the same precision-over-recall argument. A repo-rooted path
must exist, and a docs/NN reference — written as a bare number, the way this
repo’s comments write it — must name a document that does. (That shorthand is
deliberately not a code span here: quoting it as one makes this very paragraph
a finding, which is how the first draft of it failed the check.) Both are clean today
and land as regression guards — 113 paths and 140 doc references resolve, and
the doc reference is the one Go comments make most, so a renumbered document
would break every one of them silently. The narrowing is what makes them usable:
the naive form of the path class was measured first and returned 84 findings,
every one a false positive — Microsoft Learn slugs like
onelake/onelake-access-api.md, payload filenames like data.json that belong
to the caller’s data rather than to this tree, and foreign paths like ADF’s
entityTypes/Pipeline.json. Rooting the class at a tracked top-level directory
took it to nought, and the citations it gave up were never a drift checker’s to
make. Comments are found by a hand-walked scanner over Go’s four states rather
than by a regex, because "https://learn.microsoft.com/..." contains //: the
regex form reads the second half of every documentation URL in the tree as a
comment.
The third class is a pin on retired vocabulary. A term recorded in RETIRED
carries the reason it was retired and the count of occurrences the tree is known
to hold, all of which are past-tense records that the idea was held and dropped.
A further occurrence fails, and is either another such record — raise the pin in
the same commit, and the raise is the review — or the idea creeping back.
Counts rather than line numbers, because a line number moves whenever anything
above it is edited, and a check that fails for an unrelated edit is one people
learn to re-baseline without reading.
What it cannot do is the interesting part, so it is written down rather than
implied. It catches a retired term being reintroduced. It cannot catch a
justification written today going stale tomorrow, because nothing mechanical
separates a true is from a false one — the stale comment above read “the run
really DID reach the leaf”, so even a past-tense heuristic would have waved it
through. Retiring a concept is a deliberate act; this asks only that the act be
recorded once so the tree cannot drift back to it.
One level in: who tests the guards
Section titled “One level in: who tests the guards”Everything above is a checker reading the tree. Nothing read the checkers.
scripts/ is where this repository keeps its enforcement: make check runs
thirty-one of these scripts and the witnesses job runs thirty. Between them
they assert that every supported parity claim names a witness, that no route the
emulator serves is undocumented, that the sidebar is complete, that no Go comment
teaches a retired justification, and that no test sleeps without a bound. A guard
script whose own behaviour nothing asserts has the property this whole document is
about, one level in: its detection can stop matching — a renamed directory, a
tightened regex, a refactor that drops a branch — and it will go on printing its
success line and exiting 0 forever. A check that passes is indistinguishable
from a check that is running, and that sentence is as true of the checker as of
the thing it checks.
The measured gap when this landed: 11 of 51 scripts had no dedicated test
module, the largest scripts/govern_ingest.py at 717 lines, then
scripts/check_cron_workflow_freshness.py at 289 and
scripts/build_fixture_wheels.py at 220. The Go surface was measured at the same
time for comparison and has no structural gap worth reporting — all 27 packages
under internal/, pkg/ and cmd/ carry a test file, across 209 non-test source
files — so the deficit was here and here alone.
That the risk is real rather than theoretical is already written down in this
tree, twice, in the checkers’ own words. scripts/check_doc_drift.py records that
the naive form of its path class returned 84 findings, every one a false positive.
scripts/check_python_test_flakiness.py records a kind missing from its match
key, which exempted 42 of its 44 recorded symbols from the two bans the ledger
exists to enforce — a bare sleep before an assertion printed accepted and passed
--strict. Both were found by someone driving the checker against a violation
it had to catch, which is to say by a test, and neither would have been found by
reading it.
scripts/check_script_test_coverage.py asserts the list instead of the instances,
for the reason python/tests/test_make_check_runs_in_ci.py gives about its own
subject: fixing eleven files leaves the shape intact, and the twelfth script lands
with no test and nothing says so. Two finding kinds:
- UNTESTED — a script with neither a dedicated test module nor a ledger entry. This is what the ledger exists to keep shrinking.
- STALE — an entry naming a script that no longer exists, or one that has since gained a test. A closed gap must not linger as an accepted one: the entry would silently re-cover the file if that test were later deleted, so the ledger would absorb a real regression without a word.
Both directions, like docs/test-flakiness.json,
docs/python-test-flakiness.json and docs/vitest-test-flakiness.json before
it. A one-directional ledger only ever grows. Of the two kinds, STALE is the one that needs a checker: an unrecorded
script is loud by construction — somebody adds a file and the build goes red the
same day — while a stale entry is green, silent, and re-arms itself.
The rule is a name: test_<stem>.py under python/tests/ exists. Not a
coverage measurement — coverage is already measured and gated at 92%
(fail_under, pyproject.toml), and what a name adds is the one thing a
percentage cannot say, that this particular file has somewhere for its violations
to be driven from. It is also why indirect coverage does not count. Three of the
recorded scripts are exercised, under another file’s name:
scripts/check_cron_workflow_freshness.py by
python/tests/test_cron_workflow_freshness.py, which simply drops the check_
prefix; scripts/vendor_notebookutils_stubs.py through the surface checker that
reads what it vendored; scripts/govern_ingest.py incidentally by two tests of
logic that was lifted out of it. Each entry records where the behaviour actually
is exercised, so the ledger says what is true instead of flattening “covered
elsewhere” into “not covered”.
The ledger shipped smaller than the gap it records, which is the part that
makes it a ratchet rather than a waiver: eleven measured, three closed with real
unit tests, eight recorded. The three were chosen because they are pure logic over
a tmp_path fixture with no network and no Docker — scripts/image_tags.py,
scripts/check_image_digests.py, scripts/check_cask_stanzas.py. Its own test
module pins the count as a ceiling, so closing another gap passes and adding
an exemption without closing anything does not; raising it is done in the change
that argues for the new entry, and the raise is the review.
Writing one of those three found a live defect, and its shape is this document’s
first category exactly. scripts/check_image_digests.py declares its scope as
“only files that can pull — compose and env”, and had listed .env in that scope
since the day it was written. The line never matched a single file:
Path(".env").suffix is the empty string, because a leading dot makes the whole
name a stem. So the canonical name — the one compose reads, the one env_file:
defaults to — was out of scope while the constant said it was in, and
examples/fab-driven/.env is tracked here and had never been inspected. An
unscanned file produces no finding and no error, so it reads exactly like a clean
one. No amount of reading either the checker or the tree would have said so;
only pointing it at a file that must be reported did.
Recording a new exemption is one entry in docs/script-test-coverage.json
carrying the script path, its line count at the time, the reason, and where its
behaviour is exercised if it is exercised elsewhere. The reason is required and an
empty one is refused by name — an exemption with no reason is an omission wearing
a decision’s clothes. Writing the test instead is always the cheaper long-term
answer, and the ceiling above is there to keep that true.
What actually caught them
Section titled “What actually caught them”Not review, and not more assertions. In every case it was looking at what the thing produced rather than at whether it ran without complaint:
- reading dbt’s compiled SQL instead of the Jinja that generated it;
- listing the edges instead of trusting the edge count;
- printing the parsed DSN (
Host "np:",Protocols [tcp],err = nil); docker ps, which settled a day-long misdiagnosis in one line;- a
log.Printfcompiled into the binary, which settled whether a fix was running at all afterdocker compose psanddocker inspecthad both reported the tag and image id expected — for three runs against a stale image. Verify the running code, not the label on it.
Three habits that follow
Section titled “Three habits that follow”A test that cannot fail is worse than no test, because it certifies the area as covered. A parse-only test sat green next to a dial that could not connect; the handler tests injected path values and would have passed with every route mis-registered. Check a new test fails when you break the thing it guards — several here did not, and the guard for the port collision above passed its own review while sailing straight past the collision it was written for.
Prefer a loud failure to a plausible answer. Endpoints in this repo refuse
rather than return an empty list where empty is indistinguishable from missing
(GET /v1.0/myorg/datasets for a personal workspace), refuse a refresh that
would re-read nothing, and refuse modifiedSince rather than silently doing a
full pass. Where empty IS the truth — a model with no datasources — it is
returned, and the difference is the point.
A skip is only honest while the assertion still runs somewhere. The
medallion-compare job exists because compare.py skips when its counterpart
is absent; delete the job and the comparison stops being made with nothing going
red. Both carry a comment saying to remove them together or not at all.
Coverage: what the number covers, and what it cannot
Section titled “Coverage: what the number covers, and what it cannot”Three measurements, because one would misrepresent the others:
| What it measures | Gate | |
|---|---|---|
| Go | unit + in-process server tests, merged with instrumented e2e runs | 90% floor, armed only where a real SQL Server was reachable |
| Python | the unit suite’s own scope (checkers, delta-ops, storage) | 92% floor (fail_under, pyproject.toml) |
| Witnesses | every supported parity claim names a test that exists | check_witnesses.py --strict |
The witness count is the integration measure. “Every claim of support is backed by something that ran” is a statement no percentage can make, which is why it is published beside the percentages rather than folded into them.
Instrumented e2e runs
Section titled “Instrumented e2e runs”An e2e suite can contribute real coverage, because the emulator can be built
with Go’s -cover and the counters merged with the unit ones:
scripts/coverage_prepare.sh # chmod 777, or the container writes nothingFABRIC_COVERAGE=1 uv run --frozen --no-sync python e2e/sail/run.pygo test -cover -coverpkg=./... ./... -args -test.gocoverdir=$PWD/covdata/unitscripts/coverage_merge.shCoverage is requested explicitly, not through addopts:
uv run --frozen --group test pytest python/tests -q --cov --cov-report=term-missingaddopts would apply to every pytest in the repo, including the ones e2e
harnesses run inside containers that have pytest but not pytest-cov — where the
flags are unrecognised arguments and the suite dies with exit 4 for a reason
that has nothing to do with what it was testing.
The floor itself is not on that command line. It is fail_under in
pyproject.toml (92, with precision = 2 so the comparison is exact), and
it lives in one place for the reason the setting’s own comment gives: as
--cov-fail-under= in two workflow steps it was two things to move for one
change, which is how a ratchet quietly stops ratcheting.
Measure the floor in a clean venv. --frozen --group test is what CI runs,
so it is the only environment whose number is authoritative. A venv that has
drifted — anything installed with an extra group or uv run --with persists —
can read a point lower, because code that only runs when a dependency is
absent stops running once something installs it. scripts/check_govern_types.py
did exactly that: 98% clean, 85% with urllib3 present, one point on the total.
That was measured against a 70% floor, where a point was noise; against today’s
92 it is most of the headroom, since the suite measures ~93. The lesson
generalises past that file — if a test’s coverage
depends on what happens to be installed, force the branch instead of inheriting
it, or the gate fails on whoever pushes next rather than on whoever spent the
margin. That drift happened to read low, which is the safe direction; the same
mechanism reading high would hide a real regression under a floor that passes.
FABRIC_COVERAGE=1 makes the harness layer e2e/docker-compose.coverage.yml,
which builds the emulator with COVER=1 and mounts one repository-level
covdata/. Every wired suite writes into that same directory, so the merge sees
the whole fleet without needing a list of which suites ran.
Two things must both hold, and neither is obvious:
- The binary must be built instrumented.
COVER=1does that; the published image is never built this way. covdata/must be writable by uid 65532. The image is distroless nonroot and a bind mount keeps the HOST’s ownership, so a directory created by the checkout user is not writable by the container — and Go’s coverage runtime does not complain, it just writes nothing. Docker Desktop on macOS ignores the uid mismatch entirely, so this passes on a laptop and fails only on Linux CI.scripts/coverage_prepare.shdoes the chmod;e2e/engine-matrixhit the identical trap first.- The process must exit CLEANLY. A
-coverbinary writes its counters whenmainreturns, and a SIGKILLed one writes nothing — an emptycovdata/reads as “the e2e exercised nothing” rather than “nobody asked it to say”.cmd/fabric-emulatorhandles SIGTERM for exactly this reason, and a stop it was asked for returnsnilrather than the closed-listener error that would sendmainthroughlog.Fatalandos.Exit— which skips the write anyway.
Measured contribution: cmd/fabric-emulator goes from 0% on signalStop to
100% once an instrumented stack has been started and stopped, and the package
reaches ~70% from e2e alone. The merged total moves less than the per-package
figures do, because the unit suites already cover most paths — the e2e legs
prove the wiring (compose, TLS, the Entra handshake, the engines) that unit
tests deliberately stub.
Wiring a suite that is not yet instrumented is one conditional in its
run.py: append -f e2e/docker-compose.coverage.yml when FABRIC_COVERAGE is
set. A suite whose stack contains no fabric-emulator service has nothing to
instrument and is deliberately left alone.
JavaScript: always through pnpm
Section titled “JavaScript: always through pnpm”Every JS entry point — installs, scripts, CI — goes through pnpm, never npm or yarn:
pnpm install --frozen-lockfilepnpm --filter fabric-emulator-portal testpnpm --filter fabric-emulator-docs buildThe workspace pins packageManager: pnpm@10.29.3 and every package.json
carries preinstall: npx only-allow pnpm. Every one, not just the root —
the root guard does nothing for cd portal && npm install, which reads that
directory’s own manifest, finds no guard, resolves its own tree, and writes a
package-lock.json. Two lockfiles is worse than the wrong one: each tool reads
its own, both report success, and the divergence surfaces later as a version
that is somehow different in CI with nothing in the diff to explain it.
npx only-allow pnpm, and that npx is not an oversight to tidy into
pnpm dlx: the script has to run under the package manager being refused.
Someone typing npm install has npx and may not have pnpm at all.
Enforced by three guards in internal/repo, because a convention with nothing
checking it is the subject of most of this document: every manifest carries the
guard, no rival lockfile is committed, and no workflow or script shells out to
npm/yarn. The last one has a leading word boundary — npm install is a
substring of pnpm install, and the first version failed on the two CI lines
that were already correct.
Running Python: always through uv
Section titled “Running Python: always through uv”Every Python entry point — e2e runners, scripts/, Makefile targets, CI steps —
goes through uv, never a bare python3:
uv run --frozen --group <group> python e2e/<suite>/run.py # suite needs depsuv run --frozen --no-sync python e2e/<suite>/run.py # stdlib-only driverMost host-side runners only drive docker compose and are stdlib-only, so they
take --no-sync; the nine that import real client libraries name their group.
This is not style. A bare python3 resolves to whatever is on PATH, which on
a developer machine is often a pyenv build with none of the project’s
dependencies — e2e/adls-sdk fails with ModuleNotFoundError: No module named 'azure' that way, a harness fault that reads exactly like a test failure.
.python-version pins 3.12, matching requires-python, the python:3.12-slim
images, and CI’s setup-python. Without it uv satisfies >=3.12 with the
newest interpreter it can find (3.13 at the time of writing), so the host would
quietly run a different Python from the containers under test. Note that
uv run --no-sync reuses a mismatched existing .venv with only a warning
rather than rebuilding it — if the warning appears, rm -rf .venv and re-sync.
Ask what your machine has that a fresh one does not
Section titled “Ask what your machine has that a fresh one does not”Measure the floor in a clean venv
states this for the coverage number. It is not a coverage phenomenon, and it is
filed where nobody looks when the symptom is a failing test or a 401 — so:
before asking “is main broken?”, ask “what does my machine have that a fresh
one does not?” In one day that question would have ended three investigations
before they started, none of which looked like environment drift:
- A combination no lockfile produces.
deltalakepresent andpandasabsent — no group declares that pair — tooktest_delta_opsdown a path CI never runs, dying inpyarrow’s pandas shim. In CIdeltalakeis absent, so the test skips and the path is never reached. - A coverage percentage a point low, from a venv no longer matching the lockfile — the case the coverage section already describes.
- A
401that was a clock. A container left up for hours outlived several emulator restarts, and its token aged past a threshold. Nothing about it looks like environment drift until you find the clock, which is why it is the one worth remembering: the state that drifted was not a package at all.
The shape is the same each time — the failure depends on state nobody declared —
but only the first is a venv, so rm -rf .venv is a check, not the check:
gh run list --workflow ci.yml --branch main --limit 3 # is main actually red?rm -rf .venv && uv run --frozen --group test pytest python/tests -qdocker compose down -v # for anything long-livedA green CI on the commit you are sitting on does not make the local failure imaginary — it was real about an environment. It means the difference is yours to find before it is anyone else’s to chase.