Define a Sensitive Data Policy for AI metadata generation #12

Closed
opened 2026-08-29 13:41:48 +00:00 by mptyl · 0 comments
Owner

Context

Non-blocking follow-up to #4. AI description generation can send bounded real source samples to the configured provider. The current release discloses this behavior and limits each model request to five rows and five representative non-null values, but it intentionally does not guess which fields are sensitive.

Scope

  • Define the catalog-level meaning of sensitive data and the ownership of the policy at application setup level.
  • Specify how tables, columns, values, and patterns are classified.
  • Define exclusion and anonymization rules before prompt construction.
  • Define precedence, defaults, validation, and operator-visible diagnostics.
  • Preserve the rule that raw samples, transformed samples, credentials, and prompts are not written to generation logs or catalog metadata.
  • Cover direct PostgreSQL and SSH-backed sampling consistently.
  • Add tests proving that protected values never reach the model helper, persisted events, errors, or browser responses.

Delivery process

This issue is not blocking delivery of #4. Before implementation, run a dedicated grill-with-docs round, then convert the accepted decisions to a specification and implementation tickets. Keep the operational design simple and avoid turning the policy into a general data-governance platform.

Acceptance

The resulting specification must define deterministic exclusion or anonymization behavior, safe failure modes, setup examples, UI disclosure, and an acceptance test using invented sensitive-looking data only.

## Context Non-blocking follow-up to #4. AI description generation can send bounded real source samples to the configured provider. The current release discloses this behavior and limits each model request to five rows and five representative non-null values, but it intentionally does not guess which fields are sensitive. ## Scope - Define the catalog-level meaning of sensitive data and the ownership of the policy at application setup level. - Specify how tables, columns, values, and patterns are classified. - Define exclusion and anonymization rules before prompt construction. - Define precedence, defaults, validation, and operator-visible diagnostics. - Preserve the rule that raw samples, transformed samples, credentials, and prompts are not written to generation logs or catalog metadata. - Cover direct PostgreSQL and SSH-backed sampling consistently. - Add tests proving that protected values never reach the model helper, persisted events, errors, or browser responses. ## Delivery process This issue is not blocking delivery of #4. Before implementation, run a dedicated grill-with-docs round, then convert the accepted decisions to a specification and implementation tickets. Keep the operational design simple and avoid turning the policy into a general data-governance platform. ## Acceptance The resulting specification must define deterministic exclusion or anonymization behavior, safe failure modes, setup examples, UI disclosure, and an acceptance test using invented sensitive-looking data only.
mptyl added the ready-for-humanenhancement labels 2026-08-29 13:41:48 +00:00
mptyl closed this issue 2026-08-29 15:54:55 +00:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: mptyl/ThothII#12