Ship v0.1.1–v0.1.5: first-live-run fixes and dashboard viewer
- v0.1.1: config persistence — FileHelper paths made absolute (/media/startos/volumes/main/...); relative paths resolved into the JS runtime's ephemeral cwd so action saves never reached /data - v0.1.2: preJobStopContainers (Configure Grading) — docker-stop resident vLLM containers on the head Spark at job start, no auto-restart - v0.1.3: preflight auth (LiteLLM master_key gates /models), poll-until-loaded, crash fast-fail (restarting counts as dead) - v0.1.4: HF_HUB_OFFLINE/TRANSFORMERS_OFFLINE in airgapped serving (--internal network has no DNS); grader _post timeout 600→1800s for ~3.6 tok/s GB10 generation - v0.1.5: dashboard viewer survives the periodic background refresh; download buttons for deck reports, deck JSON, and SCORECARD.md - .gitignore: .startos/ build workspace, start-technologies/ Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
1d1074b625
commit
91212322c1
@@ -134,7 +134,11 @@ def write_invalid_marker(err: str) -> None:
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- LLM client
|
||||
def _post(payload: dict, timeout: int = 600) -> dict:
|
||||
def _post(payload: dict, timeout: int = 1800) -> dict:
|
||||
# 1800s: the DGX Sparks generate a 31B at only a few tokens/sec, and
|
||||
# structured-output (json_schema/guided) decoding is slower still — a full
|
||||
# extraction can legitimately run past 10 minutes. Non-streaming urlopen
|
||||
# times out on total wait, so this must cover the whole completion.
|
||||
req = urllib.request.Request(
|
||||
f"{LLM_BASE}/chat/completions",
|
||||
data=json.dumps(payload).encode(),
|
||||
|
||||
Reference in New Issue
Block a user