Files
ThothII/.superpowers/sdd/evidence-task-4-report.md
T

88 lines
4.2 KiB
Markdown

# Evidence Task 4 — shared job envelope
Status: complete
## Delivered
- Immutable `JobSpec`, `JobRun`, `JobReport`, per-stage state, sanitized error, and UTC
timestamp records.
- `run_job(spec, stages)` with a durable checkpoint at job start, before and after every stage,
and at terminal state. Successful stages are skipped when a prior run is resumed.
- Atomic JSON checkpoint/report replacement using a unique same-directory temporary file,
file `fsync`, atomic `os.replace`, and parent-directory `fsync`.
- Public reports contain fixed operational fields only. Workspace paths, stage return values,
exception messages, source content, credentials, and arbitrary metadata are not serialized.
- `WorkspaceJobLock` uses non-blocking kernel `flock` on a stable workspace/job-specific inode.
Locks are released by the kernel on process exit; lock files are never removed based on PID,
avoiding stale-lock and PID-reuse deletion races. Evidence and DWH use distinct lock files.
- Dry-run intent is immutable in the spec/report and exposed to every stage through `JobContext`.
## TDD evidence
Initial focused collection failed because `tht.jobs` did not exist. Tests then drove:
- failure, sanitized reporting, resume, and idempotent successful-stage skipping;
- corrupt-checkpoint refusal before stage execution;
- JSON schema and path/secret/PII exclusion;
- dry-run propagation and ordered aware timestamps;
- multiprocessing exclusion, distinct Evidence/DWH jobs, traversal rejection, and recovery after
a lock-owning process crashes.
Final focused result:
```text
11 passed in 0.42s
```
## Verification
```text
cd harness && .venv/bin/pytest -q
597 passed, 5 deselected, 17 warnings in 28.45s
cd harness && .venv/bin/ruff check tht/jobs tests/test_job_runner.py tests/test_job_locking.py
All checks passed!
```
The full Ruff invocation was also run. It reports 34 pre-existing violations in unrelated legacy
tests; no Task 4 file is among them. L2 tests remain deselected by the repository configuration.
## Operational notes
- `fcntl.flock` intentionally targets the supported Linux/macOS deployment environments; it is not
a Windows locking implementation.
- The envelope does not publish or mutate an active corpus. Later pipeline stages must use
`JobContext.run_dir` for staging and perform their own final atomic publish only after validation.
- A dry run is an execution mode foundation: the runner exposes and records it; individual stages
remain responsible for suppressing external mutations.
## Review hardening follow-up
Four post-implementation findings were fixed test-first:
1. Resume compatibility is now a canonical SHA-256 fingerprint over checkpoint schema version,
hashed workspace identity, job type, dry-run mode, explicit spec/pipeline versions,
configuration/input fingerprints, and the exact ordered explicit `stage_ids`. Any insertion,
removal, reorder, mode, identity, version, config, or input change rejects resume before a stage
executes. Omitting `resume_run_id` remains the explicit safe path for a new run.
2. Lock traversal now uses directory file descriptors with `O_DIRECTORY` and `O_NOFOLLOW`.
Lock files use `O_NOFOLLOW | O_CLOEXEC`; `fstat` requires a regular file owned by the current
UID with one link, and permissions are forced to `0600` (`0700` for private directories).
Pre-existing lock-file and lock-directory symlinks are rejected.
3. Stage failures now serialize only the fixed safe tuple `internal` / `stage_exception` /
`stage execution failed`. Neither exception class names nor messages are inspected for output;
a hostile exception-name/message regression test proves a terminal failed report is retained.
4. Job/run directory creation is no-follow, owner-checked, private, and durable. Each newly created
parent is fsynced, the run directory is fsynced before the first atomic file write, and the
existing file-fsync → replace → directory-fsync ordering has an explicit regression test.
Follow-up verification:
```text
focused job/lock suite: 27 passed in 0.45s
full harness suite: 613 passed, 5 deselected, 17 warnings in 29.65s
Task 4 scoped Ruff: All checks passed
```
Repository-wide Ruff continues to report the same 34 unrelated pre-existing legacy-test findings.