dashboard / README.md
3v324v23's picture
score trend
cd80b88
|
Raw
History Blame Contribute Delete
3.08 kB
metadata
title: Thinker Subnet Dashboard
emoji: 🧠
colorFrom: gray
colorTo: green
sdk: docker
app_port: 7860
pinned: false

Thinker subnet dashboard

Space configuration

Set WANDB_KEY, WANDB_PROJECT, and WANDB_ENTITY as Hugging Face Space secrets. WANDB_PROJECT is required and has no source-code default, preventing the dashboard from silently reading a different W&B project.

The dashboard refreshes at startup and every REFRESH_SECONDS seconds (default: 1800). Lightweight progress snapshots refresh every PROGRESS_REFRESH_SECONDS seconds (default: 15). A newly completed progress snapshot triggers an immediate score-history refresh, so the trend does not wait for the slower scheduled refresh. Until configuration succeeds, the dashboard serves bundled synthetic sample data.

Validator telemetry

Each validator run provides validator_hotkey and dashboard_schema_version in its W&B configuration. Every completed evaluation logs:

  • epoch
  • evaluation/round_complete
  • evaluation/miners_seen
  • evaluation/miners_scored
  • evaluation/miners_rejected
  • miner/<hotkey>/score
  • miner/<hotkey>/correctness_score when available
  • miner/<hotkey>/completion_len when available
  • original/score for compatibility
  • original/correctness_score and original/completion_len when available

miner/<hotkey>/score remains the final reward used for rankings. For miners selected into full evaluation it is the full-evaluation result; for other miners it is the validator's qualification-only result. The trend chart uses correctness_score, where correct answers contribute 1 and wrong answers contribute -1, so miner and base-model lines share the same scale. Each trend point independently selects the miner with the highest correctness_score at that epoch; the token chart uses that same epoch winner's completion length.

During evaluation, validators also publish a compact progress snapshot containing qualification and full-evaluation state for the original-model baseline and each miner, plus each miner's stage score as it finishes qualification or full evaluation. The public dashboard shows only hotkeys, stage states, and qualification/full-evaluation scores; prompts, answers, and rejection reasons are excluded.

The current-round table is scoped to miners in the latest progress snapshot and shows qualification and full-evaluation scores in separate columns. The previous round has its own table with stage scores, the final ranking reward, and a stale marker for miners absent from the current round. Rankings and efficiency trends select the highest score recorded at that exact epoch; older scores are never carried forward. Miners last scored before the previous round appear separately with their actual last-scored epoch.

The dashboard only reads W&B runs whose run state is running; stopped, crashed, failed, or finished runs are ignored.

Among running validator runs, the dashboard uses the newest run containing valid miner scores. A newer empty running run no longer hides an older scored running run.