feat(corpus): add deterministic normalization and chunking
This commit is contained in:
@@ -0,0 +1,39 @@
|
||||
# Evidence Task 3 — deterministic normalization and chunking
|
||||
|
||||
## Outcome
|
||||
|
||||
- Added pure `normalize(acquired, pipeline_version)` and `chunk(document, policy)` transforms.
|
||||
- Normalization enforces UTF-8 (including UTF-8 BOM), a 10 MiB input ceiling, LF line endings,
|
||||
NFC Unicode, safe YAML frontmatter extraction, canonical provenance URIs, and hashes the exact
|
||||
canonical UTF-8 text stored on the document.
|
||||
- Undecodable, unsupported-charset, oversized, and invalid-frontmatter inputs fail explicitly;
|
||||
byte content is never truncated.
|
||||
- Chunking uses a versioned immutable policy, paragraph/word boundaries with deterministic
|
||||
character-count hard splits for long tokens, contiguous ordinals, provenance metadata, exact
|
||||
per-chunk hashes, and IDs derived from document hash + ordinal + policy version.
|
||||
- Empty documents produce no chunks. Non-ASCII, CRLF equivalence, repeatability, policy changes,
|
||||
duplicate-content ordinal collisions, and max-character limits are covered by tests.
|
||||
|
||||
## TDD evidence
|
||||
|
||||
- Initial focused test run failed during collection because both transform modules were absent.
|
||||
- The EOF-frontmatter edge test was separately observed failing before its implementation.
|
||||
- Final focused verification: `12 passed`.
|
||||
|
||||
## Verification
|
||||
|
||||
- `cd harness && .venv/bin/pytest tests/test_corpus_normalize.py tests/test_corpus_chunk.py -q`
|
||||
— **12 passed**.
|
||||
- `cd harness && .venv/bin/pytest -q` — **573 passed, 5 deselected**. The sandboxed attempt could
|
||||
not access Docker; the approved rerun with local Docker access passed.
|
||||
- Targeted Ruff over all four implementation/test files — **clean**.
|
||||
- Full `cd harness && .venv/bin/ruff check .` — reports **34 pre-existing errors** in unrelated
|
||||
legacy tests (unused imports and existing E702 semicolon lines); none are in Task 3 files.
|
||||
|
||||
## Concerns
|
||||
|
||||
- The 10 MiB normalization ceiling is deliberately explicit and independent of adapter download
|
||||
limits. If deployment policy needs a different ceiling, it should become a versioned pipeline
|
||||
configuration before ingestion is wired.
|
||||
- Character limits use Python Unicode code points (`len`), not UTF-8 bytes or tokenizer tokens;
|
||||
this is recorded in the chunk-policy metadata and tested with non-ASCII content.
|
||||
Reference in New Issue
Block a user