Files
ThothII/.superpowers/sdd/evidence-task-3-report.md

3.6 KiB

Evidence Task 3 — deterministic normalization and chunking

Outcome

  • Added pure normalize(acquired, pipeline_version) and chunk(document, policy) transforms.
  • Normalization enforces UTF-8 (including UTF-8 BOM), a 10 MiB input ceiling, LF line endings, NFC Unicode, safe YAML frontmatter extraction, canonical provenance URIs, and hashes the exact canonical UTF-8 text stored on the document.
  • Undecodable, unsupported-charset, oversized, and invalid-frontmatter inputs fail explicitly; byte content is never truncated.
  • Chunking uses a versioned immutable policy, paragraph/word boundaries with deterministic character-count hard splits for long tokens, contiguous ordinals, provenance metadata, exact per-chunk hashes, and IDs derived from document hash + ordinal + policy version.
  • Empty documents produce no chunks. Non-ASCII, CRLF equivalence, repeatability, policy changes, duplicate-content ordinal collisions, and max-character limits are covered by tests.

TDD evidence

  • Initial focused test run failed during collection because both transform modules were absent.
  • The EOF-frontmatter edge test was separately observed failing before its implementation.
  • Final focused verification: 12 passed.

Verification

  • cd harness && .venv/bin/pytest tests/test_corpus_normalize.py tests/test_corpus_chunk.py -q — 12 passed.
  • cd harness && .venv/bin/pytest -q — 573 passed, 5 deselected. The sandboxed attempt could not access Docker; the approved rerun with local Docker access passed.
  • Targeted Ruff over all four implementation/test files — clean.
  • Full cd harness && .venv/bin/ruff check . — reports 34 pre-existing errors in unrelated legacy tests (unused imports and existing E702 semicolon lines); none are in Task 3 files.

Concerns

  • The 10 MiB normalization ceiling is deliberately explicit and independent of adapter download limits. If deployment policy needs a different ceiling, it should become a versioned pipeline configuration before ingestion is wired.
  • Character limits use Python Unicode code points (len), not UTF-8 bytes or tokenizer tokens; this is recorded in the chunk-policy metadata and tested with non-ASCII content.

Review hardening follow-up

  • Chunk IDs now bind the canonical document identity, document content hash, ordinal, chunk hash, and a canonical SHA-256 fingerprint of every ChunkPolicy field. Identical content in separate documents and same-version policies with different limits cannot collide.
  • Boundary-aware slicing now retains separators in the slices. Concatenating every chunk exactly reconstructs the canonical document for repeated spaces, tabs, blank lines, Markdown hard breaks, fenced code, whitespace-only input, Unicode, and overlong tokens; every slice remains within max_chars.
  • Frontmatter uses a bounded SafeLoader variant: duplicate keys, anchors/aliases, structures deeper than 20 nodes, and documents larger than 1000 composed nodes are rejected. YAML parse, JSON type, credential-safety, and resulting canonical-model errors attributable to frontmatter map to PermanentNormalizationError(reason="invalid_frontmatter"); invalid pipeline policy remains a programmer-facing ValueError.
  • Follow-up TDD evidence: the expanded focused suite first reported 11 expected failures against the prior implementation, then passed 45/45 across normalization, chunking, and manifest invariants.
  • Follow-up full verification: 586 passed, 5 deselected. Targeted Ruff is clean. Full Ruff continues to report the same 34 unrelated pre-existing violations in legacy tests.