Ship v0.1.1–v0.1.5: first-live-run fixes and dashboard viewer

- v0.1.1: config persistence — FileHelper paths made absolute
  (/media/startos/volumes/main/...); relative paths resolved into the JS
  runtime's ephemeral cwd so action saves never reached /data
- v0.1.2: preJobStopContainers (Configure Grading) — docker-stop resident
  vLLM containers on the head Spark at job start, no auto-restart
- v0.1.3: preflight auth (LiteLLM master_key gates /models),
  poll-until-loaded, crash fast-fail (restarting counts as dead)
- v0.1.4: HF_HUB_OFFLINE/TRANSFORMERS_OFFLINE in airgapped serving
  (--internal network has no DNS); grader _post timeout 600→1800s for
  ~3.6 tok/s GB10 generation
- v0.1.5: dashboard viewer survives the periodic background refresh;
  download buttons for deck reports, deck JSON, and SCORECARD.md
- .gitignore: .startos/ build workspace, start-technologies/

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Jonathan Kirkwood
2026-07-30 09:14:24 -05:00
co-authored by Claude Fable 5
parent 1d1074b625
commit 91212322c1
16 changed files with 262 additions and 46 deletions
+5 -1
View File
@@ -134,7 +134,11 @@ def write_invalid_marker(err: str) -> None:
# ---------------------------------------------------------------- LLM client
def _post(payload: dict, timeout: int = 600) -> dict:
def _post(payload: dict, timeout: int = 1800) -> dict:
# 1800s: the DGX Sparks generate a 31B at only a few tokens/sec, and
# structured-output (json_schema/guided) decoding is slower still — a full
# extraction can legitimately run past 10 minutes. Non-streaming urlopen
# times out on total wait, so this must cover the whole completion.
req = urllib.request.Request(
f"{LLM_BASE}/chat/completions",
data=json.dumps(payload).encode(),