Local Albedo eval harness
Offline loop for SN97 after the Aug 21β22 reset. Uses the same generation settings, submit protocol, loop detector, and pre-eval gates as the live pipeline. It does not call GLM-5.2, so proxy_score is a gate/behaviour stand-in β survive this first or the live duel never runs.
Layout
| path | what |
|---|---|
/workspace/models/albedo-qwen3.6-35b-king-genesis |
pinned genesis weights |
/workspace/data/albedo-datasets |
local parquet shards |
/workspace/data/eval-runs/<id> |
verdict + trajectories |
Pinned king: dendriteholdings/albedo-qwen3.6-35b-king-genesis@abe89194
Weight hash: sha256:efd5b8d0β¦89c165
Sampling: temperature 1.0, top_p 0.95, top_k 20, thinking on, stop at <|im_end|>.
Setup
From the albedo repo root, with the workspace venv (or any env that has huggingface_hub, pyarrow, vllm):
source /workspace/.venv/bin/activate
export PYTHONPATH=src:.
export HF_HUB_ENABLE_HF_TRANSFER=1
# genesis (~104 GB, 26 shards) β already started if you used this harness
python -m local_eval download-genesis
# two mini-coder shards so `run` can sample prefixes without the full 4-corpus tree
python -m local_eval download-lite-data
python -m local_eval verify --path /workspace/models/albedo-qwen3.6-35b-king-genesis
Full official datasets (slow, large):
python scripts/prepare_datasets.py --dataset-root /workspace/data/albedo-datasets
Run a local duel
8Γ H200 default: king on GPUs 0β3, challenger on 4β7, tensor-parallel 4.
# smoke-test genesis (you do not have a trained challenger yet)
python -m local_eval run --self-check --samples 1 --turns 1 --seed smoke
# later: duel your finetune against genesis
python -m local_eval run \
--challenger /workspace/models/your-finetune \
--king /workspace/models/albedo-qwen3.6-35b-king-genesis \
--dataset-root /workspace/data/albedo-datasets \
--samples 8 --turns 4 --seed local-eval
--challenger must be a real local directory with config.json. The placeholder /path/to/your/qwen3.6-35b will now exit immediately.
Official GLM judge (OpenRouter)
proxy_score is not the live checklist. To score a finished duel the same way the subnet does β SOTA reference (z-ai/glm-5.2, 8 turns) β checklist β yes/no judge on both stored trajectories:
export ALBEDO_JUDGE_OPENROUTER_API_KEY=...
# optional: ALBEDO_JUDGE_ENGY_API_KEY, ALBEDO_JUDGE_REPO_CONTEXT_URL
# 2-sample smoke on the v11 8Γ12 duel
python -m local_eval judge --run /workspace/data/eval-runs/20260824T182728Z-cdd1388b --limit 2
# rebuild JudgeSamples only (no API calls)
python -m local_eval judge --dry-run --limit 2
# full 8
python -m local_eval judge --run /workspace/data/eval-runs/20260824T182728Z-cdd1388b --limit 0
Writes official-judge.json and official-judge-records.jsonl next to the duel artifacts. Without a repo-context URL, SOTA observations use the official LLM simulator (DeepSeek flash, then GLM fallback) β not leftover gold parquet.
Useful flags:
--self-checkβ run genesis alone (implies--skip-king)--skip-kingβ gates on the challenger only--gpu-memory-utilization 0.80β default; raise only if the GPUs are empty--turns 12β closer to live horizons (12/16)--samples 100 --seed <block_hash>β same draw size as production (needs the full manifest)
What is scored locally
Instant fail (same as live pre-eval / judge short-circuit):
- truncated / unclosed
<think>/ too short / empty - looped commands (dup ratio β₯ 0.5 or same cmd 4Γ)
- reserved token leak
- no
```bashcommand
Warnings that tank the live checklist:
- no edit (
sed -i, heredoc, patch, β¦) - missing the exact rewritten submit command
- amputated thinking (
</think>-first turns β official score Γ 0.5)
Environment between turns: leftover gold user observations from the parquet when present; otherwise a format-correct stub (RETURNCODE / SWE_AGENT / OPENHANDS). This is not the live repo-context executor.
Install eval deps
Workspace venv: /workspace/.venv (torch 2.11.0+cu130, vllm 0.23.0).
run points CUDA_HOME at the venv CUDA 13 toolkit (nvidia/cu13), links /usr/local/cuda if missing, disables the flashinfer sampler JIT (VLLM_USE_FLASHINFER_SAMPLER=0), and uses the Triton GDN backend. Each vLLM worker gets its own Triton cache under /workspace/data/triton-cache/pid-* so four TP ranks do not race on overlayfs (and /dev/shm is noexec). First vLLM start may compile Triton MoE kernels for several minutes.
source /workspace/.venv/bin/activate
export PYTHONPATH=/workspace/gpt/albedo/src:/workspace/gpt/albedo
python -m local_eval run --self-check --skip-king --samples 1 --turns 1 --seed smoke