Files
ThothII/.superpowers/sdd/evidence-task-4-report.md
T

4.2 KiB

Evidence Task 4 — shared job envelope

Status: complete

Delivered

  • Immutable JobSpec, JobRun, JobReport, per-stage state, sanitized error, and UTC timestamp records.
  • run_job(spec, stages) with a durable checkpoint at job start, before and after every stage, and at terminal state. Successful stages are skipped when a prior run is resumed.
  • Atomic JSON checkpoint/report replacement using a unique same-directory temporary file, file fsync, atomic os.replace, and parent-directory fsync.
  • Public reports contain fixed operational fields only. Workspace paths, stage return values, exception messages, source content, credentials, and arbitrary metadata are not serialized.
  • WorkspaceJobLock uses non-blocking kernel flock on a stable workspace/job-specific inode. Locks are released by the kernel on process exit; lock files are never removed based on PID, avoiding stale-lock and PID-reuse deletion races. Evidence and DWH use distinct lock files.
  • Dry-run intent is immutable in the spec/report and exposed to every stage through JobContext.

TDD evidence

Initial focused collection failed because tht.jobs did not exist. Tests then drove:

  • failure, sanitized reporting, resume, and idempotent successful-stage skipping;
  • corrupt-checkpoint refusal before stage execution;
  • JSON schema and path/secret/PII exclusion;
  • dry-run propagation and ordered aware timestamps;
  • multiprocessing exclusion, distinct Evidence/DWH jobs, traversal rejection, and recovery after a lock-owning process crashes.

Final focused result:

11 passed in 0.42s

Verification

cd harness && .venv/bin/pytest -q
597 passed, 5 deselected, 17 warnings in 28.45s

cd harness && .venv/bin/ruff check tht/jobs tests/test_job_runner.py tests/test_job_locking.py
All checks passed!

The full Ruff invocation was also run. It reports 34 pre-existing violations in unrelated legacy tests; no Task 4 file is among them. L2 tests remain deselected by the repository configuration.

Operational notes

  • fcntl.flock intentionally targets the supported Linux/macOS deployment environments; it is not a Windows locking implementation.
  • The envelope does not publish or mutate an active corpus. Later pipeline stages must use JobContext.run_dir for staging and perform their own final atomic publish only after validation.
  • A dry run is an execution mode foundation: the runner exposes and records it; individual stages remain responsible for suppressing external mutations.

Review hardening follow-up

Four post-implementation findings were fixed test-first:

  1. Resume compatibility is now a canonical SHA-256 fingerprint over checkpoint schema version, hashed workspace identity, job type, dry-run mode, explicit spec/pipeline versions, configuration/input fingerprints, and the exact ordered explicit stage_ids. Any insertion, removal, reorder, mode, identity, version, config, or input change rejects resume before a stage executes. Omitting resume_run_id remains the explicit safe path for a new run.
  2. Lock traversal now uses directory file descriptors with O_DIRECTORY and O_NOFOLLOW. Lock files use O_NOFOLLOW | O_CLOEXEC; fstat requires a regular file owned by the current UID with one link, and permissions are forced to 0600 (0700 for private directories). Pre-existing lock-file and lock-directory symlinks are rejected.
  3. Stage failures now serialize only the fixed safe tuple internal / stage_exception / stage execution failed. Neither exception class names nor messages are inspected for output; a hostile exception-name/message regression test proves a terminal failed report is retained.
  4. Job/run directory creation is no-follow, owner-checked, private, and durable. Each newly created parent is fsynced, the run directory is fsynced before the first atomic file write, and the existing file-fsync → replace → directory-fsync ordering has an explicit regression test.

Follow-up verification:

focused job/lock suite: 27 passed in 0.45s
full harness suite: 613 passed, 5 deselected, 17 warnings in 29.65s
Task 4 scoped Ruff: All checks passed

Repository-wide Ruff continues to report the same 34 unrelated pre-existing legacy-test findings.