#Metadata Cache Contract
Maintainer reference for cache layers, write semantics, TTL, events, and host↔LSP synchronization.
#Shared rule ownership
@justybase/metadata-core is the platform-neutral rule layer. It provides the
encoded key codec, dialect-supplied identifier policies, fresh/stale/expired
TTL classification, snapshot completeness, full-snapshot and refreshed-type
merge operations, lookup-index construction, prefetch planning, scoped
invalidation matching, and generation guards. It has no VS Code, database
driver, filesystem, timer, logger, or transport dependency.
The desktop MetadataCache remains the stateful adapter around VS Code,
catalog SQL, progress events, lazy column hydration, and disk persistence. The
API uses one ApiMetadataService owned by each buildServer instance. API
cache keys contain namespace, authenticated owner, connection, layer, and
encoded metadata parts; invalidation is scoped to owner+connection and a
generation check rejects late responses. A second server instance therefore
cannot observe another instance's metadata cache.
#Cache layers and keys
| Layer | Key format | Write semantics |
|---|---|---|
database |
CONN |
Replace all databases for the connection |
schema |
CONN|DB |
Replace all schemas for the database |
currentSchema |
CONN|DB |
Replace default/current schema for that connection/database |
table |
CONN|DB.SCHEMA or CONN|DB.. |
Complete replace for that key; rebuilds lookup indexes |
column |
CONN|DB.SCHEMA.TABLE |
Prefetch: fill-missing only; explorer refresh: replace |
procedure |
CONN|DB.SCHEMA or CONN|DB.. |
Replace per key |
typeGroup |
CONN|DB |
Merge with dialect defaults |
foreignKeyRelationships |
Netezza CONN → DB |
Replace a database FK slice; incomplete scans retain old rows but are never authoritative |
objectLookup / objectsByType |
derived | Invalidated on setTables / invalidateSchema; rebuilt lazily |
Connection names are passed through as provided by callers (some lookup methods normalize to uppercase).
API keys use the shared tagged codec rather than uppercasing the complete key.
Netezza user identifiers use the dialect policy (unquoted names fold to upper
case; quoted names remain exact), while catalog values are preserved. SQLite,
DuckDB, and other case-sensitive adapters retain the spelling supplied by the
caller. Empty schema segments remain explicit, so DB..TABLE cannot collide
with DB.SCHEMA.TABLE.
For Netezza, catalog-derived database names use an exact, encoded cache-key
part (@NZEX@...). This prevents distinct catalog objects such as JUST_DATA
and just_data from sharing a cache entry. The encoded marker is an internal
key representation; values exposed to the schema tree and completion retain the
catalog spelling, case, and meaningful whitespace.
flowchart TD
Q[Catalog query] --> Set[setTables / setSchemas / setColumns]
Set --> Cache[(MetadataCache layers)]
Cache --> Lookup[Lookup indexes · objectLookup / objectsByType]
Cache --> Disk[Disk serializer / compressor]
Disk --> Cache2[(Disk cache · cross-window sync)]
Cache --> Bridge[Metadata bridge / LSP schema provider]
Bridge --> Consumers[Completion · validation · result context]
Consumers --> Q2[Targeted refresh / prefetch]
Q2 --> Cache#Table cache write policy
#setTables — complete replacement
MetadataCache.setTables(connection, key, data, idMap) replaces the entire table list for key. It does not merge with prior entries.
Callers that refresh only one object type (TABLE, VIEW, NICKNAME, ALIAS) within a schema must merge before calling setTables. Use mergeTableLikeObjectsForSchema or mergeAndSetTables.
For DB.. keys, mergeAndSetTables falls back to getTablesAllSchemas when the aggregate key is missing, so per-schema TABLE rows are preserved during a database-level VIEW refresh.
Prefetch reads regular objects and external tables using separate catalog
queries, merges/deduplicates them in TypeScript, then calls setTables per
schema key with the full result set for that schema.
Schema explorer refreshes one type at a time and merges with existing cache entries of other types.
#Example: refresh VIEW when TABLE already cached
- Explorer loads VIEW list from the server.
mergeTableLikeObjectsForSchema(existingTables, newViews, 'VIEW')keeps TABLE rows, replaces VIEW rows.setTableswrites the merged array.
Skipping step 2 removes all TABLE entries for that schema key.
#TTL and prefetch freshness
cacheTtl— configured viacacheTTL(default 12 hours).staleTtl—2 × cacheTtl; entries may still be served until stale window ends, then evicted on read.- Prefetch freshness (
isConnectionPrefetchFresh) usescacheTtlonly, notstaleTtl. - Pure TTL decisions receive an explicit timestamp;
Date.now()remains in the desktop/API adapters that own scheduling and I/O. currentSchemauses the same TTL/stale window asschema.- Full-refresh column sessions are controlled by
justybase.metadata.fullRefreshColumnConnections(default1, maximum8). Only independent database column jobs use additional sessions; the other refresh stages remain on the primary session. A session-open failure falls back to the primary session and never starts an unbounded retry loop.
#Optional foreign-key metadata
Core snapshot completeness and TTL depend on database, schema, object, procedure
and column layers. FK relationship reads do not invalidate these layers. Each
connection/database FK slice records complete, unavailable (expected access
or missing catalogue), or failed (including timeout and SQL errors). Status is
persisted in v3 sidecars; older sidecars derive it from complete.
Selecting a profile with fresh core metadata retries only missing/failed FK slices, at most twice per scope per session with five minutes between attempts. Unavailable slices wait for explicit refresh or core expiry. Recovery preserves the original core timestamp, disposes its runner and uses the connection disk lease. Known relationships survive failed reads; targets absent from a loaded catalogue stay visible in Schema View with an unavailable marker.
Changing a SQL document's connection profile clears its database override. Reselecting the same effective profile preserves the override.
#Invalidation
#invalidateSchema(connection, db, schema?)
Removes for the target scope:
- Table cache entry (
CONN\|DB.SCHEMAorCONN\|DB..) - Aggregated
CONN\|DB..table cache when a specific schema is invalidated - Procedure cache (schema + all-schemas aggregate when applicable)
- Column cache keys for the schema
- For Netezza, the connection FK relationship index (rebuilt by the next standard metadata prefetch)
- Lookup indexes via
removeTableCacheEntry
Fires onDidInvalidate.
#clearCache()
Wipes all in-memory layers, bumps _cacheGeneration (cancels in-flight disk/column loads), clears stats, fires onDidInvalidate.
#Events
| Event | When | Typical subscribers |
|---|---|---|
onDidInvalidate |
invalidateSchema, clearCache |
Semantic tokens, LSP notification |
onDidExternalRefresh |
Cross-window disk re-hydration per connection | Schema browser, LSP notification |
onDidPrefetchProgress |
Prefetch stages | Status bar |
onDidNeedColumnRecovery |
Column disk load failure | Prefetch coordinator |
#Netezza-specific behavior
- Identifier sources are distinct — a name typed by the user follows
Netezza SQL rules (an unquoted identifier is folded to uppercase; a quoted
identifier is exact), while a value read from
_V_*is catalog data and is preserved exactly. Catalog predicates therefore use direct equality and catalog qualifiers are quoted when required; system-view joins use stable object identifiers rather than name normalization. - Completion after a complete refresh — once the in-memory cache contains the complete database/schema/object/column snapshot, completion resolves from cache and does not issue a metadata SQL request while typing. A cache miss is reported as cache-only until the normal refresh/warmup path supplies data.
DB..TABLE— name-only lookup index; first-match wins across schemas; winner updates when first-match schema is removed.- Multi-schema — refreshing schema S1 does not remove schema S2 entries (separate cache keys).
- Table-like types — TABLE, VIEW, NICKNAME, ALIAS, SYNONYM, SEQUENCE, MATERIALIZED VIEW, SYSTEM VIEW, and related types share the
tablecache layer; explorer merge is perobjType. - Split object prefetch (Netezza) — stage 3 reads regular object types
(TABLE, VIEW, SYNONYM, SEQUENCE, MATERIALIZED VIEW, SYSTEM VIEW, and related
table-like types) from
_V_OBJECT_DATA, then reads external tables from the small_V_EXTERNALcompanion query. The two results are normalized and merged in the extension, never withUNION ALL, a correlatedNOT EXISTS, or an owner-only_V_EXTOBJECTjoin in Netezza. PROCEDURE uses the separateprocedurelayer (stage 4). Type groups are prefetched per DB (stage 2, after schemas). - FK relationship index (Netezza) — the existing FK catalog query in the
ordinary per-database column prefetch builds a connection cache slice for
each database.
Referencesreads the current database slice andReferenced bysearches all cached slices; expanding either tree group runs no catalog SQL. The slices are compressed per-database sidecars in the same generation/fence-protected v3 snapshot. A cache without the FK index version keeps its core freshness and receives an FK-only recovery read.
#Host ↔ LSP synchronization
- Canonical store: extension-host
MetadataCache. - LSP cache:
MetadataBridge(list + tableInfo, 12h TTL) in the language-server process. - Sync path:
onDidInvalidate/onDidExternalRefresh→NETEZZA_METADATA_CACHE_INVALIDATED_NOTIFICATION→metadataBridge.clearAll(). - Validation epoch: LSP diagnostics track
metadataEpoch; stale validation results are dropped when epoch changes after cache invalidation.
LSP metadata requests are served by handleMetadataRequest reading the host cache via RPC.
#Disk persistence and cross-window sync
When disk persistence is enabled (justybase.metadataCache.diskPersistence, default true):
- On startup, a small per-connection manifest hydrates the database list immediately; heavy metadata layers (schema, table, procedure, typeGroup) hydrate from disk in the background.
- Columns stay in per-database column files (
*.columns.json.gz) until loaded on demand. - Netezza FK relationships are loaded from per-database
*.foreign-keys.json.gzsidecars during the metadata hydrate; this does not hydrate column layers or query the database. - Prefetch checkpoints metadata to disk; column files are written at checkpoint/dispose.
- Checkpoints are marked
isComplete: false. They may be loaded for recovery, but never restoreisConnectionPrefetchFresh. - Only a verified complete snapshot (
isComplete: true) restores prefetch freshness. Completion requires database, schema, table, procedure/type catalog stages, column cache entries for table/view/external-table objects, and a complete FK slice for every Netezza database. - Another VS Code window writing the v3 index triggers
onExternalCacheUpdate→ re-hydrate metadata →onDidExternalRefresh.
Self-writes are skipped when this window holds the prefetch lock.
The manifest is written after metadata and column payloads, and the v3 index is written last. Startup accepts a manifest as fresh only when its timestamp/fingerprint matches the index and the snapshot is complete. A partial snapshot that has not expired may still provide local data while background prefetch refreshes it; an expired snapshot is excluded from hydration and its connection payload is removed before the next clean refresh.
#Restart and expiry behavior
After a complete refresh, the RAM cache is the first source for completion and
schema navigation. On a program restart, a matching complete disk snapshot is
hydrated and remains usable without re-querying metadata; column layers may be
hydrated lazily from the per-database column files. The database is queried
again when the snapshot is absent, incomplete, fingerprint-incompatible, or no
longer fresh according to the configured cacheTTL (default 12 hours). A
still-readable in-memory snapshot may provide immediate local results until
that refresh starts; the refresh then discards the connection snapshot before
querying. An expired disk snapshot is excluded from startup hydration and its
payload is deleted, so it cannot be used as the next refresh's merge base.
#Column load paths (restart-safe)
| API | When used | Behavior |
|---|---|---|
ensureColumnsLoaded(connection, db) |
Completion legacy path, small catalogs | Full per-DB column file → all layers into RAM; marks DB as fully loaded |
ensureColumnsLoadedForTableKey(connection, layerKey) |
Schema tree, MetadataProvider, columnCacheLookup |
Prefer single layer; see below |
ensureColumnsLoadedForTableKey order:
- Return if
getColumns(connection, layerKey)already in RAM. - Return if DB already fully loaded (
columnsLoadedDatabases). - If column file exists on disk →
loadColumnLayerFromDisk(decode oneDB.SCHEMA.TABLElayer). - If catalog is not large → fall back to
ensureColumnsLoaded(full DB hydrate).
Large catalogs (isLargeTableCatalog, default threshold 500 table-like objects per DB, constant LARGE_DB_TABLE_LIKE_OBJECT_THRESHOLD in schemaTreeDataSource.ts):
- Skip full DB hydrate and eager column preload for that database.
- Load only requested table layers on demand (schema tree expand, completion).
- Parsed column files are cached in RAM (
parsedColumnFileCache) so the second table in the same DB reuses the gzip parse without re-reading disk.
Schema tree (SchemaProvider.getChildren for netezza:TABLE|VIEW|…) calls ensureColumnsLoadedForTableKey then reads RAM. It does not issue SQL when hasTreeReadyColumnCache passes (columns present with isPk defined). Column rows written to cache use normalizeColumnCacheEntry so missing isDistributionKey does not cause refetch loops.
Completion (MetadataProvider.getTableColumnsMetadata) uses the same ensureColumnsLoadedForTableKey entry point before getColumns.
#Views catalog completeness (in-memory only)
viewsCatalogLoaded is a RAM-only Set keyed CONN|DB.SCHEMA (normalized). It is not serialized to disk. After restart, flags are restored indirectly:
| Source | When markViewsCatalogLoaded runs |
|---|---|
Split object prefetch (prefetchAllObjects) |
After every per-schema setTables (views may be zero) |
| Disk metadata hydrate | After each table layer setTables in hydrateConnectionMetadataChunked |
| Live views fetch | After MetadataProvider.getViews writes merged VIEW rows |
Explorer partial setTables |
Only when batch contains objType === 'VIEW' |
Completion / getViews for DB..: if cache has table-like rows but zero VIEW rows, return [] without SQL when:
isViewsCatalogLoaded(connection, cacheKey)for the scope key, orareViewsCatalogLoadedForDatabase(connection, db)— every per-schema table layer for that DB has the flag.
Without these flags, getViews may show “Fetching views…” even when the database truly has no views.
#Objects catalog completeness (in-memory only)
objectsCatalogLoaded is a RAM-only Set keyed CONN|DB.SCHEMA|OBJTYPE (normalized layer key). It complements viewsCatalogLoaded for all prefetched table-cache types and uses CONN|DB|PROCEDURE for the procedure layer.
| Source | When catalog flags are set |
|---|---|
| Split object prefetch (stage 3) | markPrefetchObjectTypesCatalogLoaded after each per-schema setTables |
| Prefetch procedures (stage 4) | markProcedureCatalogLoaded per database |
| Disk metadata hydrate | Same marks after hydrateConnectionMetadataChunked |
| Schema tree live fetch | markObjectsCatalogLoaded / markProcedureCatalogLoaded on write-back |
Schema tree (typeGroup:*): SchemaProvider.getChildren reads procedureCache for PROCEDURE and getObjectsByType for table-cache types before issuing SQL. Empty results are served without SQL when areObjectsCatalogLoadedForDatabase (table types) or isProcedureCatalogLoaded (procedures) is true.
SYNONYM disk roundtrip: REFOBJNAME is persisted in the table layer (KNOWN_TABLE_KEYS).
Type groups after restart: disk-hydrated typeGroup layers skip triggerTypeGroupsRefresh. When missing, deriveTypeGroupsFromCache builds a fallback list from cached objType values and procedure presence.
#Disk layout reminder
- Metadata index:
globalStorage/metadata-cache-v3/index.json.gz
#Multi-window disk protocol (v3)
The disk cache uses an independent metadata-cache-v3 directory. v2 is never
loaded, migrated, or removed, so an older extension process cannot share
payloads with v3.
The v3 index is the source of truth and contains a global generation, a
monotonic revision, and nextFence. A prefetch receives a connection lease,
the generation observed at acquisition, and a fence token allocated under the
global writer lease. Checkpoints and the final snapshot commit only while that
lease is valid. A commit is rejected when its generation is old or its fence is
older than the connection's committed fence.
The v1 monolith and v2 cache root are deliberately left untouched and are not used as startup state. A schema-v2 column blob referenced by a valid v3 index remains readable for compatibility and is rewritten as dictionary-encoded v3 on the next successful snapshot. Unknown or corrupt index, manifest, metadata, and column payloads are ignored. A corrupt metadata payload clears prefetch freshness; a corrupt column payload additionally raises the column-recovery event so consumers cannot retain stale column types.
A snapshot is accepted only when its connection fingerprint matches the current host, port, database, and dialect. Credentials are never part of the fingerprint or persisted payload.
Locks have random owner and lease identifiers. Their lock record is immutable; heartbeat files are unique to the lease. Consequently a former owner cannot renew, delete, or overwrite a lock acquired after expiry.
clearCache() is global: it increments the generation and commits an empty
index under the global writer lease. It does not remove another process's
locks. Windows observing the generation change clear RAM, invalidate LSP data,
and discard in-flight hydration. Disposal does not write an all-RAM snapshot.
- Per connection: metadata JSON + optional
DB.columns.json.gzper database with column layers
#Regression tests
Disk restart and schema-tree column expand are covered without a live database:
npm run test:metadata-cache:integrationTests simulate: populate cache → dispose (disk write) → new MetadataCache → initialize() → column/object-list access without runQueryRaw (TABLE, PROCEDURE, SEQUENCE, SYNONYM, MATERIALIZED VIEW, SYSTEM VIEW, schemas).
Key files:
packages/metadata-core/__tests__/rules.test.tssrc/__tests__/integration/metadataCacheRestart.integration.test.tssrc/__tests__/fixtures/metadataCacheRestartFixture.ts
#Related modules
| Module | Role |
|---|---|
src/metadata/cache/MetadataCache.ts |
Facade |
src/metadata/cache/MetadataStore.ts |
In-memory maps + TTL |
src/metadata/cache/columnLoader.ts |
Lazy column hydrate, per-layer disk load, eager preload |
src/metadata/cache/schemaTreeDataSource.ts |
getTablesForScope, hasTreeReadyColumnCache, large-catalog threshold |
src/metadata/cache/layerAccess.ts |
Layer I/O, isLargeTableCatalog, areViewsCatalogLoadedForDatabase |
src/metadata/diskStorage/metadataColumnCodec.ts |
Column file v3 encode/decode, decodeColumnLayerFromFile |
src/metadata/prefetch.ts |
Staged connection prefetch |
src/metadata/cache/MetadataPrefetchTarget.ts |
Prefetch write contract (breaks import cycle) |
src/metadata/cache/MetadataStorageReader.ts |
Read-only maps for search |
src/metadata/columnCacheLookup.ts |
Async column cache reads for hover/LSP helpers |
src/metadata/helpers.ts |
Key builders, mergeTableLikeObjectsForSchema |
src/providers/schemaProvider.ts |
Schema tree; column expand via ensureColumnsLoadedForTableKey |
src/providers/providers/metadataProvider.ts |
Completion metadata; getViews / getTableColumnsMetadata |
src/server/metadataBridge.ts |
LSP-side cache |
src/sqlParser/metadataCacheAdapter.ts |
Validator schema provider |