feat: refine metadata catalog workflows
This commit is contained in:
@@ -0,0 +1,67 @@
|
||||
# Use one Installation Model Catalog with runtime projections
|
||||
|
||||
ThothII currently declares model availability independently in installation
|
||||
`metadataGeneration`, Pi configuration, and workspace `llm_policy` and embedding settings. The
|
||||
Installation Model Catalog in `thothii-installation.yaml` becomes the sole authored authority for
|
||||
session, metadata-generation, and embedding models, including their allowed usages and per-usage
|
||||
defaults. Workspace descriptors retain database identity and scope plus Evidence concerns, but no
|
||||
model policy or selection; installation-local Database Bindings remain separate from them.
|
||||
|
||||
Pi, the backend metadata-generation helper, and the embedding runtime consume generated Model
|
||||
Runtime Projections of that catalog. Runtime Model Selection stores only a canonical catalog model
|
||||
identity and use-specific controls such as thinking level; endpoint, provider, capabilities, and
|
||||
credential references remain catalog facts, while secret values remain in protected secret stores.
|
||||
The metadata-generation helper continues to use LiteLLM independently of Pi, as established by
|
||||
ADR-0009: unifying model declaration does not unify execution lifecycles.
|
||||
|
||||
The migration is intentionally fail-closed. Workspace schema v4 removes `llm_policy` and the
|
||||
entire redundant `semantic_index`; collection identity is derived from the workspace identity,
|
||||
while vector-store and embedding facts come from the installation. A deterministic migration
|
||||
rewrites existing descriptors. The
|
||||
installation loader replaces `metadataGeneration` with `modelCatalog` and rejects the legacy form
|
||||
with an actionable migration error rather than keeping two live sources. Existing Pi
|
||||
`models.json` and enabled-model settings become generated artifacts and are never edited as
|
||||
authoritative configuration.
|
||||
|
||||
Catalog identities use the canonical `provider/model` form; runtime-specific upstream names are
|
||||
adapter facts, not additional ThothII identities. Installation descriptors use schema version 2
|
||||
and model-free workspace descriptors use schema version 4. Removing a model never substitutes it
|
||||
inside an existing session: an unresolvable resume fails explicitly. Published semantic indexes
|
||||
record the embedding identity and dimensions that produced them and require explicit
|
||||
reprocessing when those facts change.
|
||||
|
||||
Model eligibility is expressed by the presence of a `session` or `metadataGeneration` block,
|
||||
without a duplicate usages list. The installation declares one active embedding identity and its
|
||||
dimensions rather than a selectable embedding catalog. Runtime projections are regenerated
|
||||
deterministically and atomically at start, so they require no persisted digest and are excluded
|
||||
from installation backups; restore regenerates them from the validated installation descriptor.
|
||||
|
||||
A provider owns one endpoint, one explicit authentication mode, and only the runtime adapters it
|
||||
needs. Authentication is either a protected secret-environment reference, Pi-owned authentication
|
||||
for session-only built-in models, or explicit keyless operation for an explicit endpoint; models
|
||||
cannot override it. Pi built-in model facts are not copied into the installation. A session default
|
||||
and the single embedding definition are required, while metadata generation and its default may be
|
||||
omitted together. The schema deliberately excludes unused abstractions and future properties until
|
||||
runtime behavior requires them.
|
||||
|
||||
The catalog session default replaces `PI_PROVIDER`, `PI_MODEL`, and persisted installation model
|
||||
defaults as configuration sources. A user or session selection is only a canonical catalog
|
||||
reference, and an existing session keeps that reference without silently switching models. A
|
||||
generated Compose projection supplies the catalog-derived embedding values and projection mounts
|
||||
to every affected service, so neither the base Compose files nor `operator.env` repeat model facts.
|
||||
|
||||
Workspace v3-to-v4 migration is a deterministic removal of `llm_policy` and `semantic_index`.
|
||||
Installation migration instead inspects the legacy installation metadata block and both Pi model
|
||||
files: it emits a v2 candidate only when their identities and settings can be reconciled without
|
||||
guessing. Conflicts produce an actionable report and leave every source untouched.
|
||||
|
||||
## Considered Options
|
||||
|
||||
- A separate catalog file referenced by the installation was rejected because it adds path,
|
||||
permission, backup, and atomic-update coordination without a current need for cross-installation
|
||||
sharing.
|
||||
- Pi `models.json` was rejected as the authority because it is a Pi-specific projection that does
|
||||
not express all ThothII usages, built-in providers, metadata-generation controls, or embedding
|
||||
facts.
|
||||
- Transitional dual reading was rejected because it would preserve the configuration discrepancy
|
||||
this decision is intended to eliminate.
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
status: accepted
|
||||
---
|
||||
|
||||
# Assess sensitive columns locally from source content
|
||||
|
||||
The Sensitive Data Flag remains a human-owned boolean. An explicit, selection-scoped sensitivity
|
||||
analysis may propose changes by inspecting both catalog metadata and source values, but it never
|
||||
writes the flag. The administrator may accept, reject, or reverse every proposal.
|
||||
|
||||
One TypeScript `SensitivityClassifier` is the only component allowed to produce the column-level
|
||||
assessment `sensitive`, `non_sensitive`, or `unknown`. It applies a versioned Sensitive Data Policy
|
||||
and consumes values through database-independent streaming adapters. Database-specific code may
|
||||
read and normalize bounded values, but it may not decide sensitivity.
|
||||
|
||||
The classifier first applies deterministic metadata rules, value validators, checksums,
|
||||
dictionaries, and length rules. A single validated sensitive match makes the whole column
|
||||
`sensitive`; any textual value longer than 500 characters is such a match. It attempts a complete
|
||||
scan, but after five seconds per table it continues by sampling within the remaining run budget. A
|
||||
completed scan with no finding may produce `non_sensitive`; an incomplete scan with no finding
|
||||
produces `unknown`.
|
||||
|
||||
Ambiguous text may additionally be sent to an optional local NER detector only while time remains.
|
||||
The detector runs on CPU, receives no tools or network access, does not persist source values, and
|
||||
returns evidence rather than the column decision. The initial supported detector is
|
||||
[`fastino/gliner2-privacy-filter-PII-multi`](https://huggingface.co/fastino/gliner2-privacy-filter-PII-multi),
|
||||
used through the Apache-2.0 GLiNER2 Python library with a pinned model revision. Its model weights
|
||||
and GLiNER2 code are Apache-2.0, and its mDeBERTa base model is MIT. It is trained for seven
|
||||
languages including Italian and can run on CPU without using the installation's GPUs.
|
||||
|
||||
The NER detector is an optional installation asset because its weights and runtime are materially
|
||||
larger than the deterministic TypeScript engine. It is invoked only for otherwise unresolved text,
|
||||
never for values already classified by a decisive rule. If it is disabled, unavailable, times out,
|
||||
or returns no qualifying evidence before the deadline, the classifier follows the same coverage
|
||||
rule and may return `unknown`.
|
||||
|
||||
No generative LLM, internal or external, participates in sensitivity assessment. A fallback to an
|
||||
installation-local LLM is unnecessary while a permissively licensed local NER implementation is
|
||||
available, and would reintroduce queue latency, non-deterministic judgments, prompt-injection
|
||||
surface, and contention with normal inference. Introducing such a fallback would require a new
|
||||
decision based on evidence that the NER path is unusable.
|
||||
|
||||
This decision supersedes ADR-0011 only where that ADR assigns draft sensitivity suggestions to an
|
||||
AI using structural metadata. ADR-0011's human authority, transient draft, explicit scope, and
|
||||
Sensitive Data Flag remain in force. How description generation uses the flag is outside this
|
||||
decision.
|
||||
|
||||
The selected model's published evaluation is not an Italian production acceptance test. Before
|
||||
enabling the NER profile by default, ThothII must pin the artifacts, generate a dependency/license
|
||||
inventory, and pass a CPU benchmark plus a labeled Italian corpus representative of the target
|
||||
databases. Failure of those gates disables NER; it does not silently select another model or an
|
||||
LLM.
|
||||
Reference in New Issue
Block a user