2.2 KiB
2.2 KiB
Evidence Task 3 — deterministic normalization and chunking
Outcome
- Added pure
normalize(acquired, pipeline_version)andchunk(document, policy)transforms. - Normalization enforces UTF-8 (including UTF-8 BOM), a 10 MiB input ceiling, LF line endings, NFC Unicode, safe YAML frontmatter extraction, canonical provenance URIs, and hashes the exact canonical UTF-8 text stored on the document.
- Undecodable, unsupported-charset, oversized, and invalid-frontmatter inputs fail explicitly; byte content is never truncated.
- Chunking uses a versioned immutable policy, paragraph/word boundaries with deterministic character-count hard splits for long tokens, contiguous ordinals, provenance metadata, exact per-chunk hashes, and IDs derived from document hash + ordinal + policy version.
- Empty documents produce no chunks. Non-ASCII, CRLF equivalence, repeatability, policy changes, duplicate-content ordinal collisions, and max-character limits are covered by tests.
TDD evidence
- Initial focused test run failed during collection because both transform modules were absent.
- The EOF-frontmatter edge test was separately observed failing before its implementation.
- Final focused verification:
12 passed.
Verification
cd harness && .venv/bin/pytest tests/test_corpus_normalize.py tests/test_corpus_chunk.py -q— 12 passed.cd harness && .venv/bin/pytest -q— 573 passed, 5 deselected. The sandboxed attempt could not access Docker; the approved rerun with local Docker access passed.- Targeted Ruff over all four implementation/test files — clean.
- Full
cd harness && .venv/bin/ruff check .— reports 34 pre-existing errors in unrelated legacy tests (unused imports and existing E702 semicolon lines); none are in Task 3 files.
Concerns
- The 10 MiB normalization ceiling is deliberately explicit and independent of adapter download limits. If deployment policy needs a different ceiling, it should become a versioned pipeline configuration before ingestion is wired.
- Character limits use Python Unicode code points (
len), not UTF-8 bytes or tokenizer tokens; this is recorded in the chunk-policy metadata and tested with non-ASCII content.