File size: 3,082 Bytes
afc4b4c
b413fd8
 
afc4b4c
b413fd8
afc4b4c
b413fd8
afc4b4c
cd80b88
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
---
title: Thinker Subnet Dashboard
emoji: 🧠
colorFrom: gray
colorTo: green
sdk: docker
app_port: 7860
pinned: false
---

# Thinker subnet dashboard

## Space configuration

Set `WANDB_KEY`, `WANDB_PROJECT`, and `WANDB_ENTITY` as Hugging Face Space
secrets. `WANDB_PROJECT` is required and has no source-code default, preventing
the dashboard from silently reading a different W&B project.

The dashboard refreshes at startup and every `REFRESH_SECONDS` seconds (default:
`1800`). Lightweight progress snapshots refresh every `PROGRESS_REFRESH_SECONDS`
seconds (default: `15`). A newly completed progress snapshot triggers an immediate
score-history refresh, so the trend does not wait for the slower scheduled refresh.
Until configuration succeeds, the dashboard serves bundled synthetic sample data.

## Validator telemetry

Each validator run provides `validator_hotkey` and `dashboard_schema_version` in
its W&B configuration. Every completed evaluation logs:

- `epoch`
- `evaluation/round_complete`
- `evaluation/miners_seen`
- `evaluation/miners_scored`
- `evaluation/miners_rejected`
- `miner/<hotkey>/score`
- `miner/<hotkey>/correctness_score` when available
- `miner/<hotkey>/completion_len` when available
- `original/score` for compatibility
- `original/correctness_score` and `original/completion_len` when available

`miner/<hotkey>/score` remains the final reward used for rankings. For miners
selected into full evaluation it is the full-evaluation result; for other miners
it is the validator's qualification-only result. The trend chart uses
`correctness_score`, where correct answers contribute `1` and wrong answers
contribute `-1`, so miner and base-model lines share the same scale. Each trend
point independently selects the miner with the highest `correctness_score` at
that epoch; the token chart uses that same epoch winner's completion length.

During evaluation, validators also publish a compact progress snapshot containing
qualification and full-evaluation state for the original-model baseline and each
miner, plus each miner's stage score as it finishes qualification or full
evaluation. The public dashboard shows only hotkeys, stage states, and
qualification/full-evaluation scores; prompts,
answers, and rejection reasons are excluded.

The current-round table is scoped to miners in the latest progress snapshot and
shows qualification and full-evaluation scores in separate columns. The previous
round has its own table with stage scores, the final ranking reward, and a stale
marker for miners absent from the current round. Rankings and efficiency trends
select the highest score recorded at that exact epoch; older scores are never
carried forward. Miners last scored before the previous round appear separately
with their actual last-scored epoch.

The dashboard only reads W&B runs whose run state is `running`; stopped, crashed,
failed, or finished runs are ignored.

Among running validator runs, the dashboard uses the newest run containing valid
miner scores. A newer empty running run no longer hides an older scored running
run.