fix(corpus): harden canonical chunk and frontmatter invariants
This commit is contained in:
@@ -37,3 +37,23 @@
|
||||
configuration before ingestion is wired.
|
||||
- Character limits use Python Unicode code points (`len`), not UTF-8 bytes or tokenizer tokens;
|
||||
this is recorded in the chunk-policy metadata and tested with non-ASCII content.
|
||||
|
||||
## Review hardening follow-up
|
||||
|
||||
- Chunk IDs now bind the canonical document identity, document content hash, ordinal, chunk hash,
|
||||
and a canonical SHA-256 fingerprint of every `ChunkPolicy` field. Identical content in separate
|
||||
documents and same-version policies with different limits cannot collide.
|
||||
- Boundary-aware slicing now retains separators in the slices. Concatenating every chunk exactly
|
||||
reconstructs the canonical document for repeated spaces, tabs, blank lines, Markdown hard
|
||||
breaks, fenced code, whitespace-only input, Unicode, and overlong tokens; every slice remains
|
||||
within `max_chars`.
|
||||
- Frontmatter uses a bounded `SafeLoader` variant: duplicate keys, anchors/aliases, structures
|
||||
deeper than 20 nodes, and documents larger than 1000 composed nodes are rejected. YAML parse,
|
||||
JSON type, credential-safety, and resulting canonical-model errors attributable to frontmatter
|
||||
map to `PermanentNormalizationError(reason="invalid_frontmatter")`; invalid pipeline policy
|
||||
remains a programmer-facing `ValueError`.
|
||||
- Follow-up TDD evidence: the expanded focused suite first reported 11 expected failures against
|
||||
the prior implementation, then passed **45/45** across normalization, chunking, and manifest
|
||||
invariants.
|
||||
- Follow-up full verification: **586 passed, 5 deselected**. Targeted Ruff is clean. Full Ruff
|
||||
continues to report the same **34 unrelated pre-existing** violations in legacy tests.
|
||||
|
||||
Reference in New Issue
Block a user