Skip to content

Parity — v0.2.2

How the emulator’s surface maps to Databricks’ public workspace REST, and — the point of this table — whether real work happens or just the API shape.

The catalog is the workspace REST API reference (sidebar retrieved 2026-08-15). That list is what real Databricks offers at the workspace host. Product pages, DBR version strings, and third-party OpenAPI scrapes are not a denominator.

The design bet is the same one fabric-emulator runs: terminate the public contract here, attach a real engine, refuse what you cannot compute. This process owns identity, files, secrets persist, and the workspace REST shim. Jobs, SQL, clusters, Connect, and UC CRUD are answered by a named sidecar — Sail behind the family’s spark-agent, Spark Connect gRPC, UC OSS, keyvault-emulator — never by a toy stub or a silent DuckDB-as-Photon.

make run does not start those sidecars. make e2e-engine / make e2e-uc do. Family compose does not set the Spark URLs. A missing engine fails naming the variable — never SUCCESS / RUNNING.

The wire is proved by an unmodified client: databricks-sdk, the Databricks CLI, or databricks/databricks Terraform. A doc page that names an endpoint is not support. See 00-doctrine.md.

Meaning
🟢 RealThe shim terminates the contract and a named store or engine did the work. An unmodified client drove the call. Status without a witness is not support.
🟡 EmulatedFaithful API contract + persisted state, but no engine — status is clock-derived / management-only.
🟠 Non-default engineReal, but only on an engine that is not the default attach (Sail + spark-agent). Not a silent substitute: the row names the overlay.
🔴 Not implementedHonest 501 or absent. Never a silent 200.

Scope boundary: the workspace REST, not Databricks Runtime

Section titled “Scope boundary: the workspace REST, not Databricks Runtime”

This ledger grades workspace API groups from the reference, written in fabric’s style: feature → what the shim terminates → which engine computes. It is not every operation inside a group, and not the account-level APIs at docs.databricks.com/api/account. A green Jobs row is not every Jobs endpoint; a red Pipelines row is the whole Lakeflow/DLT group.

The first honest slice is identity (PAT + emulator OIDC), Workspace files, DBFS, Git Credentials / Repos, Jobs 2.2 Python/notebook on Sail, secrets, SQL warehouses / Queries / Connect / clusters-as-session on that same engine, Command Execution, Unity Catalog CRUD through UC OSS, and MLflow experiments / model registry on the file-backed tracking store. Everything else from the catalog is enumerated below as Not implemented until a witness exists.

This is not a Databricks Runtime. Photon, DBR version strings, full dbutils / spark.databricks.*, cluster VMs, and any “DBR compatible” claim are 🔴 Not implemented even though the docs mention them.

FeatureEmulatorType
Identity — PATThis process seeds and mints PATs. GET /api/2.0/preview/scim/v2/Me after Authorization: Bearer. Unknown and token=dev are 401 — dev is MiniLake’s trap, not a credential this seeder mints.🟢 Real
Identity — emulator OIDCThis process’s /oidc/v1/token client-credentials. Me with no entra process.🟢 Real
Identity — federated JWTOpt-in DATABRICKS_OIDC_ISSUERS. Unconfigured / wrong aud / expired / garbage → 401; a good JWT → Me. Entra is an issuer, not a required STS.🟢 Real
FeatureEmulatorType
Workspace — SOURCE/PYTHONFile-backed store. Classic /workspace/import SOURCE/PYTHON notebook round-trip. Other formats refused by name.🟢 Real
Workspace — raw filesworkspace-files raw bytes, including RAW/FILE import.🟢 Real
Git Credentials / ReposGit credentials persist (token never returned after create). repos.create git clones a real remote into the workspace store; repos.update fetch/checkout; workspace export of a cloned file is those bytes. Sparse checkout is 501.🟢 Real
DBFS / Files APIReal bytes under data/dbfs/. Read length capped (1 MiB) before allocate. Traversal refused.🟢 Real
FeatureEmulatorType
Jobs 2.2 — notebook / PythonShim resolves workspace/DBFS file, bakes argv / notebook params / {{secrets}} into a Python preamble (os.environ.update), POSTs {agent}/statements with kind: python. Without DATABRICKS_SPARK_CONNECT_URL, run-now fails naming the engine — never SUCCESS. Default attach is Sail behind the family’s spark-agent (make e2e-engine). Family compose does not set the URL.🟢 Real
Jobs 2.2 — JAR / dbt / DLT / sql_task.queryRefused at create. sql_task.file takes the warehouse Spark SQL path (see SQL warehouses). No JVM overlay is shipped, so JAR is not 🟠.🔴 Not implemented
FeatureEmulatorType
Secrets — Databricks injectionShim resolves {{secrets/scope/key}} in this process before the engine runs. GET of a value is 400. Missing key → run FAILED. The family’s spark-agent drops req.Env, so the preamble bakes os.environ.update. Witness prints SECRET=s3cret in get-output.🟢 Real
Secrets — Databricks persistDatabricks-backed scopes under data/secrets/. Survive process restart with the same DATABRICKS_DATA_DIR.🟢 Real
Secrets — Azure Key Vault-backedLive read-through at use time — no sync. dns_name must be an Azure suffix or DATABRICKS_AKV_VAULT_HOST. put/delete refused. Rotate the vault secret; the next run GETs it.🟢 Real
Secrets — vault-audience tokenWhen DATABRICKS_ENTRA_TOKEN_URL is set, each vault GET carries a client-credentials bearer with scope https://vault.azure.net/.default. Empty URL: resolve stays unauthenticated (stand-in / make run).🟢 Real
FeatureEmulatorType
SQL warehousesSession handle, not a VM and not Photon. POST /api/2.0/sql/statements sends the SQL as kind: sql (the code is Spark SQL). Wire names dialect: spark-sql; executedBy says Spark SQL, not Photon. Without DATABRICKS_SPARK_CONNECT_URL, execute is FAILED naming the engine.🟢 Real
SQL Queries / Query HistoryStored query CRUD (/api/2.0/sql/queries). Execute is the warehouse statements path with that query_text (same Sail attach). History lists those executions. Alerts and visualizations stay 501.🟢 Real
SQL warehouses — Thrift / HiveServer2Same warehouse handle and Sail attach. POST /sql/1.0/endpoints/{id} (and /sql/protocolv1/o/{org}/{id} when {id} is a warehouse) is TBinary HiveServer2. Unmodified databricks-sql-connector==4.4.0 SELECT 1 returns one typed cell. GetSchemas / GetTables are SHOW on that engine. Cloud Fetch / Arrow+LZ4 / GetCatalogs stay refused. Missing engine fails naming DATABRICKS_SPARK_CONNECT_URL.🟢 Real
Delta writes — SailWarehouse SQL CREATE TABLE … USING delta LOCATION + INSERT + DELETE + MERGE INTO on a shared volume. Sail writes; delta-rs reads _delta_log and the rows. A Sail COUNT(*) after DML is not a witness. Standalone UPDATE is forwarded and Sail answers FAILED (CommandNode::Update) — never a silent no-op. Two concurrent INSERT OVERWRITEs: each success has its own log version; rows are one overwrite, not a silent merge. Photon is not this row.🟢 Real
Delta writes — UC three-part namesSDK creates an EXTERNAL table in UC OSS (Docker sidecar). Sail’s unity catalog provider (SAIL_CATALOG__LIST, same Compose network) resolves cat.sch.tbl. Warehouse INSERT writes; delta-rs confirms the log. INSERT is not rewritten into a path. CREATE TABLE cat.sch.t with no LOCATION is rewritten to an EXTERNAL path under file:///data/delta/managed (Sail writes; this process registers UC). hive_metastore is not rewritten.🟢 Real
Delta maintenance — OPTIMIZE / VACUUMWarehouse SQL. Sail cannot plan these. The family’s spark-agent runs them through delta-rs (named shim). ZORDER and WHERE are refused, not silently dropped. Photon is not this row.🟢 Real
Delta writes — JVM overlayWarehouse SQL on Apache Spark 3.5.5 + delta-spark (make e2e-delta-jvm), not Sail. Same delta-rs confirmer. OPTIMIZE … ZORDER is this row.🟠 Non-default engine
Clusters as session handlePOST /api/2.0/clusters/create starts a Sail session (print(1) via the HTTP agent) or fails naming the missing engine. Never sleeps to RUNNING. Autoscale and cluster libraries stay refused.🟢 Real
Cluster Policies / Policy Families / compliancePolicies persist. Create is denied when it violates fixed / range / forbidden / allowlist on spark_version, node_type_id, num_workers, autoscale, libraries. Unknown attributes are 501, not stored-and-ignored. One policy family: emulator-session. get-compliance reports the stored handle against that policy.🟢 Real
Command Execution/api/1.2/contexts + /commands on a RUNNING cluster handle. Python and SQL run on the attached Sail agent (kind: python / kind: sql). Scala / R are 501. Without the engine, context create fails naming DATABRICKS_SPARK_CONNECT_URL.🟢 Real
Databricks ConnectAfter PAT/OIDC and x-databricks-cluster-id naming a RUNNING handle, application/grpc / /spark.connect.… is reverse-proxied to DATABRICKS_SPARK_CONNECT_GRPC_URL (Sail :50051, h2c). The HTTP agent is not this backend; only that URL set is 501 naming the gRPC variable. Authorization stripped before the engine.🟢 Real
Clusters as VMsNo hypervisor. A session handle is not a VM.🔴 Not implemented
Photon / DBR compatibilityNo Photon attach exists. Sail is Spark SQL over Spark Connect.🔴 Not implemented
FeatureEmulatorType
Unity Catalog CRUDReverse-proxy to UC OSS (DATABRICKS_UC_URL) after PAT/OIDC. Without a sidecar those routes are 501 naming the missing URL. MANAGED table create is refused (UC OSS only creates EXTERNAL tables at a filesystem location). Three-part SQL against those tables is the Delta writes row, not this one.🟢 Real
Unity Catalog grantsEnforcement, not allow-all CRUD. Not shipped until they deny.🔴 Not implemented
FeatureEmulatorType
MLflow Experiments / Model RegistryFile-backed tracking store under data/mlflow/. Experiments, runs (params / metrics / tags), registered models and versions persist across restart. Duplicate experiment names are 409. Artifact list / log-model / traces / logged-models are 501 — metadata only, not a model binary store. Model Serving is a different row.🟢 Real
FeatureEmulatorType
MCP — Databricks SQLPOST /api/2.0/mcp/sql JSON-RPC after PAT/OIDC. execute_sql / poll_response wrap the warehouse statements handler (same Sail attach, same dialect: spark-sql). Genie / AI Search / UC function MCP paths stay 501.🟢 Real
Terraform / DAB pairUnmodified databricks/databricks and Databricks CLI v1.12.1: current_user + notebook + workspace_file + job create. token=dev refused. bundle deploy is not this row — current DAB schema also demands a cluster, and Permissions stay 501. Job execution is the engine row, not this one.🟢 Real
databricks-target togglePublished databricks-target package. DATABRICKS_TARGET=emulator|real resolves host, token, warehouse-by-name, catalog, vault. Consumer code holds names. make e2e-databricks-target creates contoso_warehouse, resolves it, SELECT 1. Real mode refuses localhost and seed secrets.🟢 Real
dbt-databricks warehouse runUnmodified dbt-databricks==1.12.4 dbt run of one (select 1 as id) and two (select id from {{ ref('one') }}) over HiveServer2. The gold gate (make e2e-dbt-uc) sets catalog (not hive_metastore), +file_format: delta, no post-hook; the warehouse shim writes file:///data/delta/managed/… and delta-rs plus a three-part SELECT confirm the rows. make e2e-dbt remains hive_metastore Thrift smoke (volume copy). Jobs dbt_task stays refused.🟢 Real

One row per remaining group on the workspace REST API reference sidebar. The emulator column is what an honest attach would have to be — same style as the greens — not a product-page promise.

FeatureEmulatorType
Global Init ScriptsApplied to a real cluster VM — we have no VMs.🔴 Not implemented
Instance Pools / Instance ProfilesReal VMs / cloud instance profiles.🔴 Not implemented
Managed LibrariesInstalled on a real cluster VM. JARs on Sail have no classloader (fabric’s JVM overlay is the 🟠 path; this repo does not ship one).🔴 Not implemented
AppsWould need a deployed app process.🔴 Not implemented
SCIM Groups / Users / Service PrincipalsDirectory mutations that then deny.🔴 Not implemented
PermissionsEnforcement, not allow-all.🔴 Not implemented
SQL AlertsWould evaluate a stored query on a schedule. No alert evaluator is attached.🔴 Not implemented
Unity Catalog beyond CRUDVolumes, functions, locations, credentials, monitors — sidecar must speak them.🔴 Not implemented
Delta SharingProviders / Recipients / Shares against a real share.🔴 Not implemented
MarketplaceConsumer + provider listing APIs.🔴 Not implemented
Token management / Workspace Conf / SettingsPersist and then gate. Token create is not a seeded PAT.🔴 Not implemented
Clean Rooms🔴 Not implemented
Database Instances / PostgresLakebase. Would attach a real Postgres (fabric’s SQL Server sidecar pattern), not a handle that reports RUNNING.🔴 Not implemented
Knowledge Assistants / Supervisor Agents🔴 Not implemented
Lakeview Embedded / AI Gateway🔴 Not implemented
Notification Destinations🔴 Not implemented
MCP — Genie / AI Search / UC functionsOther MCP mounts. SQL MCP is the green row above.🔴 Not implemented
Lakeflow / DLTPipelines API. No open DLT engine to attach.🔴 Not implemented
Model ServingServing endpoints. Would need a real model process, not a 200 stub.🔴 Not implemented
Vector SearchIndexes / endpoints. Would need a real index engine.🔴 Not implemented
DashboardsLakeview. No dashboard renderer is attached.🔴 Not implemented