Tested against the live Strike ledger (9 decks): its decks mint ad-hoc KPI
names every quarter, so ~70 one-off unscored mentions flooded the watch list,
trajectory table, and dropped-KPIs gap. A KPI now only counts as trackable/
vanished if it was held to a target at least once; the rest collapse to a
one-line count. Scoring-derived flags (no_profitability_visibility,
no_quantitative_kpis, kpi_dropped) no longer duplicate their dedicated gaps
as recurring-flag requests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
progress.py reads the whole graded ledger and answers two questions the
per-deck scorecard can't: is the company actually progressing (composite/
BDEF-category/KPI trajectories, recurring vs resolved flags), and is the
material good enough to judge them by — a deterministic gap engine spots
what the decks are NOT showing (no profitability visibility, untargeted
KPIs, no forward guidance, broken forecast chain, silently dropped KPIs,
thin-evidence BDEF categories, no board asks) and renders each gap as a
concrete, paste-ready request for the next deck.
Served live from the ledger (no GPU) at /api/companies/{slug}/progress(.md),
written to /data/ledger/<slug>/PROGRESS.md + /data/reports/latest-progress.md
after each graded deck, and viewable/downloadable from the dashboard company
card ("View progress review"). 18 new tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Secondary-Spark models now work in air-gapped mode: the LiteLLM proxy is
dual-homed onto the default bridge (docker network connect) to reach the
secondary's published vLLM port; grader containers stay on the --internal
network with zero egress. The head-only enforcement is replaced by a
secondary-configured check.
- Extraction runs in parallel with grading when the extractor's model and
every grader model in the wave sit on different Sparks (separate GPUs).
- Keep-warm: single-wave jobs no longer tear the wave down between decks
(was a ~6-min 31B reload per deck); a kept-warm wave that fails preflight
is restarted once. Adjudicator reuses the live wave when its model is
already serving instead of cycling the shared proxy.
- clear_resident_containers (preJobStopContainers) now stops names on every
configured Spark; health() reports containers on both Sparks.
- Verified with a mocked dry-run of the full job loop (3 decks: one
bring-up, zero mid-job teardowns, parallel overlap, stale-wave restart).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- DELETE /api/companies/{slug}: wipes the ledger (deck records, scorecard,
report md), sweeps per-job report and processed-deck copies; optional
restore_decks moves the graded originals from /data/processed back into
the inbox for a fresh Grade Decks run; 409 while a job is running;
config-registered companies reappear as empty rows (aliases + pinned
targets kept)
- Dashboard: "delete company…" danger button on the company card with a
two-step confirm (delete, then restore-decks choice)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Dashboard renders deck reports and scorecards as formatted pages
(self-contained markdown renderer inlined in index.html — headings,
styled tables, evidence blockquotes; no CDN, air-gap friendly)
- At-a-glance strip on every deck report: composite with delta vs the
prior deck, quant/qual/penalty mix, KPI hit/miss chips, BDEF
best/weakest category, red-flag count
- Generated DECK_REPORT.md now leads with a concise "At a glance"
summary table (scorecard.py); jobs.py passes the previous deck record
so the delta appears in the file too; dashboard hides the duplicate
section since the strip covers it
- .gitignore: .venv/
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>