dora7's picture
Albedo SN97 workspace v16: RECORD, merged v16, adapters v13/v15/v16-dpo, packs, harness, evals
2abcc30 verified
|
Raw
History Blame Contribute Delete
4.79 kB

Local Albedo eval harness

Offline loop for SN97 after the Aug 21–22 reset. Uses the same generation settings, submit protocol, loop detector, and pre-eval gates as the live pipeline. It does not call GLM-5.2, so proxy_score is a gate/behaviour stand-in β€” survive this first or the live duel never runs.

Layout

path what
/workspace/models/albedo-qwen3.6-35b-king-genesis pinned genesis weights
/workspace/data/albedo-datasets local parquet shards
/workspace/data/eval-runs/<id> verdict + trajectories

Pinned king: dendriteholdings/albedo-qwen3.6-35b-king-genesis@abe89194
Weight hash: sha256:efd5b8d0…89c165
Sampling: temperature 1.0, top_p 0.95, top_k 20, thinking on, stop at <|im_end|>.

Setup

From the albedo repo root, with the workspace venv (or any env that has huggingface_hub, pyarrow, vllm):

source /workspace/.venv/bin/activate
export PYTHONPATH=src:.
export HF_HUB_ENABLE_HF_TRANSFER=1

# genesis (~104 GB, 26 shards) β€” already started if you used this harness
python -m local_eval download-genesis

# two mini-coder shards so `run` can sample prefixes without the full 4-corpus tree
python -m local_eval download-lite-data

python -m local_eval verify --path /workspace/models/albedo-qwen3.6-35b-king-genesis

Full official datasets (slow, large):

python scripts/prepare_datasets.py --dataset-root /workspace/data/albedo-datasets

Run a local duel

8Γ— H200 default: king on GPUs 0–3, challenger on 4–7, tensor-parallel 4.

# smoke-test genesis (you do not have a trained challenger yet)
python -m local_eval run --self-check --samples 1 --turns 1 --seed smoke

# later: duel your finetune against genesis
python -m local_eval run \
  --challenger /workspace/models/your-finetune \
  --king /workspace/models/albedo-qwen3.6-35b-king-genesis \
  --dataset-root /workspace/data/albedo-datasets \
  --samples 8 --turns 4 --seed local-eval

--challenger must be a real local directory with config.json. The placeholder /path/to/your/qwen3.6-35b will now exit immediately.

Official GLM judge (OpenRouter)

proxy_score is not the live checklist. To score a finished duel the same way the subnet does β€” SOTA reference (z-ai/glm-5.2, 8 turns) β†’ checklist β†’ yes/no judge on both stored trajectories:

export ALBEDO_JUDGE_OPENROUTER_API_KEY=...
# optional: ALBEDO_JUDGE_ENGY_API_KEY, ALBEDO_JUDGE_REPO_CONTEXT_URL

# 2-sample smoke on the v11 8Γ—12 duel
python -m local_eval judge --run /workspace/data/eval-runs/20260824T182728Z-cdd1388b --limit 2

# rebuild JudgeSamples only (no API calls)
python -m local_eval judge --dry-run --limit 2

# full 8
python -m local_eval judge --run /workspace/data/eval-runs/20260824T182728Z-cdd1388b --limit 0

Writes official-judge.json and official-judge-records.jsonl next to the duel artifacts. Without a repo-context URL, SOTA observations use the official LLM simulator (DeepSeek flash, then GLM fallback) β€” not leftover gold parquet.

Useful flags:

  • --self-check β€” run genesis alone (implies --skip-king)
  • --skip-king β€” gates on the challenger only
  • --gpu-memory-utilization 0.80 β€” default; raise only if the GPUs are empty
  • --turns 12 β€” closer to live horizons (12/16)
  • --samples 100 --seed <block_hash> β€” same draw size as production (needs the full manifest)

What is scored locally

Instant fail (same as live pre-eval / judge short-circuit):

  • truncated / unclosed <think> / too short / empty
  • looped commands (dup ratio β‰₯ 0.5 or same cmd 4Γ—)
  • reserved token leak
  • no ```bash command

Warnings that tank the live checklist:

  • no edit (sed -i, heredoc, patch, …)
  • missing the exact rewritten submit command
  • amputated thinking (</think>-first turns β†’ official score Γ— 0.5)

Environment between turns: leftover gold user observations from the parquet when present; otherwise a format-correct stub (RETURNCODE / SWE_AGENT / OPENHANDS). This is not the live repo-context executor.

Install eval deps

Workspace venv: /workspace/.venv (torch 2.11.0+cu130, vllm 0.23.0).

run points CUDA_HOME at the venv CUDA 13 toolkit (nvidia/cu13), links /usr/local/cuda if missing, disables the flashinfer sampler JIT (VLLM_USE_FLASHINFER_SAMPLER=0), and uses the Triton GDN backend. Each vLLM worker gets its own Triton cache under /workspace/data/triton-cache/pid-* so four TP ranks do not race on overlayfs (and /dev/shm is noexec). First vLLM start may compile Triton MoE kernels for several minutes.

source /workspace/.venv/bin/activate
export PYTHONPATH=/workspace/gpt/albedo/src:/workspace/gpt/albedo
python -m local_eval run --self-check --skip-king --samples 1 --turns 1 --seed smoke