- Secondary-Spark models now work in air-gapped mode: the LiteLLM proxy is
dual-homed onto the default bridge (docker network connect) to reach the
secondary's published vLLM port; grader containers stay on the --internal
network with zero egress. The head-only enforcement is replaced by a
secondary-configured check.
- Extraction runs in parallel with grading when the extractor's model and
every grader model in the wave sit on different Sparks (separate GPUs).
- Keep-warm: single-wave jobs no longer tear the wave down between decks
(was a ~6-min 31B reload per deck); a kept-warm wave that fails preflight
is restarted once. Adjudicator reuses the live wave when its model is
already serving instead of cycling the shared proxy.
- clear_resident_containers (preJobStopContainers) now stops names on every
configured Spark; health() reports containers on both Sparks.
- Verified with a mocked dry-run of the full job loop (3 decks: one
bring-up, zero mid-job teardowns, parallel overlap, stale-wave restart).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>