feat: refine metadata catalog workflows

This commit is contained in:
Codex
2026-09-02 15:58:23 +02:00
parent 4531746038
commit ae053961a3
31 changed files with 848 additions and 219 deletions
@@ -0,0 +1,67 @@
# Use one Installation Model Catalog with runtime projections
ThothII currently declares model availability independently in installation
`metadataGeneration`, Pi configuration, and workspace `llm_policy` and embedding settings. The
Installation Model Catalog in `thothii-installation.yaml` becomes the sole authored authority for
session, metadata-generation, and embedding models, including their allowed usages and per-usage
defaults. Workspace descriptors retain database identity and scope plus Evidence concerns, but no
model policy or selection; installation-local Database Bindings remain separate from them.
Pi, the backend metadata-generation helper, and the embedding runtime consume generated Model
Runtime Projections of that catalog. Runtime Model Selection stores only a canonical catalog model
identity and use-specific controls such as thinking level; endpoint, provider, capabilities, and
credential references remain catalog facts, while secret values remain in protected secret stores.
The metadata-generation helper continues to use LiteLLM independently of Pi, as established by
ADR-0009: unifying model declaration does not unify execution lifecycles.
The migration is intentionally fail-closed. Workspace schema v4 removes `llm_policy` and the
entire redundant `semantic_index`; collection identity is derived from the workspace identity,
while vector-store and embedding facts come from the installation. A deterministic migration
rewrites existing descriptors. The
installation loader replaces `metadataGeneration` with `modelCatalog` and rejects the legacy form
with an actionable migration error rather than keeping two live sources. Existing Pi
`models.json` and enabled-model settings become generated artifacts and are never edited as
authoritative configuration.
Catalog identities use the canonical `provider/model` form; runtime-specific upstream names are
adapter facts, not additional ThothII identities. Installation descriptors use schema version 2
and model-free workspace descriptors use schema version 4. Removing a model never substitutes it
inside an existing session: an unresolvable resume fails explicitly. Published semantic indexes
record the embedding identity and dimensions that produced them and require explicit
reprocessing when those facts change.
Model eligibility is expressed by the presence of a `session` or `metadataGeneration` block,
without a duplicate usages list. The installation declares one active embedding identity and its
dimensions rather than a selectable embedding catalog. Runtime projections are regenerated
deterministically and atomically at start, so they require no persisted digest and are excluded
from installation backups; restore regenerates them from the validated installation descriptor.
A provider owns one endpoint, one explicit authentication mode, and only the runtime adapters it
needs. Authentication is either a protected secret-environment reference, Pi-owned authentication
for session-only built-in models, or explicit keyless operation for an explicit endpoint; models
cannot override it. Pi built-in model facts are not copied into the installation. A session default
and the single embedding definition are required, while metadata generation and its default may be
omitted together. The schema deliberately excludes unused abstractions and future properties until
runtime behavior requires them.
The catalog session default replaces `PI_PROVIDER`, `PI_MODEL`, and persisted installation model
defaults as configuration sources. A user or session selection is only a canonical catalog
reference, and an existing session keeps that reference without silently switching models. A
generated Compose projection supplies the catalog-derived embedding values and projection mounts
to every affected service, so neither the base Compose files nor `operator.env` repeat model facts.
Workspace v3-to-v4 migration is a deterministic removal of `llm_policy` and `semantic_index`.
Installation migration instead inspects the legacy installation metadata block and both Pi model
files: it emits a v2 candidate only when their identities and settings can be reconciled without
guessing. Conflicts produce an actionable report and leave every source untouched.
## Considered Options
- A separate catalog file referenced by the installation was rejected because it adds path,
permission, backup, and atomic-update coordination without a current need for cross-installation
sharing.
- Pi `models.json` was rejected as the authority because it is a Pi-specific projection that does
not express all ThothII usages, built-in providers, metadata-generation controls, or embedding
facts.
- Transitional dual reading was rejected because it would preserve the configuration discrepancy
this decision is intended to eliminate.
@@ -0,0 +1,52 @@
---
status: accepted
---
# Assess sensitive columns locally from source content
The Sensitive Data Flag remains a human-owned boolean. An explicit, selection-scoped sensitivity
analysis may propose changes by inspecting both catalog metadata and source values, but it never
writes the flag. The administrator may accept, reject, or reverse every proposal.
One TypeScript `SensitivityClassifier` is the only component allowed to produce the column-level
assessment `sensitive`, `non_sensitive`, or `unknown`. It applies a versioned Sensitive Data Policy
and consumes values through database-independent streaming adapters. Database-specific code may
read and normalize bounded values, but it may not decide sensitivity.
The classifier first applies deterministic metadata rules, value validators, checksums,
dictionaries, and length rules. A single validated sensitive match makes the whole column
`sensitive`; any textual value longer than 500 characters is such a match. It attempts a complete
scan, but after five seconds per table it continues by sampling within the remaining run budget. A
completed scan with no finding may produce `non_sensitive`; an incomplete scan with no finding
produces `unknown`.
Ambiguous text may additionally be sent to an optional local NER detector only while time remains.
The detector runs on CPU, receives no tools or network access, does not persist source values, and
returns evidence rather than the column decision. The initial supported detector is
[`fastino/gliner2-privacy-filter-PII-multi`](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi),
used through the Apache-2.0 GLiNER2 Python library with a pinned model revision. Its model weights
and GLiNER2 code are Apache-2.0, and its mDeBERTa base model is MIT. It is trained for seven
languages including Italian and can run on CPU without using the installation's GPUs.
The NER detector is an optional installation asset because its weights and runtime are materially
larger than the deterministic TypeScript engine. It is invoked only for otherwise unresolved text,
never for values already classified by a decisive rule. If it is disabled, unavailable, times out,
or returns no qualifying evidence before the deadline, the classifier follows the same coverage
rule and may return `unknown`.
No generative LLM, internal or external, participates in sensitivity assessment. A fallback to an
installation-local LLM is unnecessary while a permissively licensed local NER implementation is
available, and would reintroduce queue latency, non-deterministic judgments, prompt-injection
surface, and contention with normal inference. Introducing such a fallback would require a new
decision based on evidence that the NER path is unusable.
This decision supersedes ADR-0011 only where that ADR assigns draft sensitivity suggestions to an
AI using structural metadata. ADR-0011's human authority, transient draft, explicit scope, and
Sensitive Data Flag remain in force. How description generation uses the flag is outside this
decision.
The selected model's published evaluation is not an Italian production acceptance test. Before
enabling the NER profile by default, ThothII must pin the artifacts, generate a dependency/license
inventory, and pass a CPU benchmark plus a labeled Italian corpus representative of the target
databases. Failure of those gates disables NER; it does not silently select another model or an
LLM.