Files
ThothII/.superpowers/sdd/evidence-task-3-report.md
T

60 lines
3.6 KiB
Markdown

# Evidence Task 3 — deterministic normalization and chunking
## Outcome
- Added pure `normalize(acquired, pipeline_version)` and `chunk(document, policy)` transforms.
- Normalization enforces UTF-8 (including UTF-8 BOM), a 10 MiB input ceiling, LF line endings,
NFC Unicode, safe YAML frontmatter extraction, canonical provenance URIs, and hashes the exact
canonical UTF-8 text stored on the document.
- Undecodable, unsupported-charset, oversized, and invalid-frontmatter inputs fail explicitly;
byte content is never truncated.
- Chunking uses a versioned immutable policy, paragraph/word boundaries with deterministic
character-count hard splits for long tokens, contiguous ordinals, provenance metadata, exact
per-chunk hashes, and IDs derived from document hash + ordinal + policy version.
- Empty documents produce no chunks. Non-ASCII, CRLF equivalence, repeatability, policy changes,
duplicate-content ordinal collisions, and max-character limits are covered by tests.
## TDD evidence
- Initial focused test run failed during collection because both transform modules were absent.
- The EOF-frontmatter edge test was separately observed failing before its implementation.
- Final focused verification: `12 passed`.
## Verification
- `cd harness && .venv/bin/pytest tests/test_corpus_normalize.py tests/test_corpus_chunk.py -q`
— **12 passed**.
- `cd harness && .venv/bin/pytest -q` — **573 passed, 5 deselected**. The sandboxed attempt could
not access Docker; the approved rerun with local Docker access passed.
- Targeted Ruff over all four implementation/test files — **clean**.
- Full `cd harness && .venv/bin/ruff check .` — reports **34 pre-existing errors** in unrelated
legacy tests (unused imports and existing E702 semicolon lines); none are in Task 3 files.
## Concerns
- The 10 MiB normalization ceiling is deliberately explicit and independent of adapter download
limits. If deployment policy needs a different ceiling, it should become a versioned pipeline
configuration before ingestion is wired.
- Character limits use Python Unicode code points (`len`), not UTF-8 bytes or tokenizer tokens;
this is recorded in the chunk-policy metadata and tested with non-ASCII content.
## Review hardening follow-up
- Chunk IDs now bind the canonical document identity, document content hash, ordinal, chunk hash,
and a canonical SHA-256 fingerprint of every `ChunkPolicy` field. Identical content in separate
documents and same-version policies with different limits cannot collide.
- Boundary-aware slicing now retains separators in the slices. Concatenating every chunk exactly
reconstructs the canonical document for repeated spaces, tabs, blank lines, Markdown hard
breaks, fenced code, whitespace-only input, Unicode, and overlong tokens; every slice remains
within `max_chars`.
- Frontmatter uses a bounded `SafeLoader` variant: duplicate keys, anchors/aliases, structures
deeper than 20 nodes, and documents larger than 1000 composed nodes are rejected. YAML parse,
JSON type, credential-safety, and resulting canonical-model errors attributable to frontmatter
map to `PermanentNormalizationError(reason="invalid_frontmatter")`; invalid pipeline policy
remains a programmer-facing `ValueError`.
- Follow-up TDD evidence: the expanded focused suite first reported 11 expected failures against
the prior implementation, then passed **45/45** across normalization, chunking, and manifest
invariants.
- Follow-up full verification: **586 passed, 5 deselected**. Targeted Ruff is clean. Full Ruff
continues to report the same **34 unrelated pre-existing** violations in legacy tests.