| # Benchmark & generation scripts |
|
|
| Harnesses used to benchmark and generate from the Quazimoto (Wheeler/Positronic) |
| and SpikeWhale (Byrne) model families. |
|
|
| ## Benchmark |
|
|
| | file | what it does | |
| |---|---| |
| | `run_eval.py` | **Main benchmark.** WikiText-2 `byte_ppl`, BLiMP (12 paradigms x150), MC suite (arc_easy/challenge, hellaswag, winogrande, piqa, openbookqa, boolq -> acc & acc_norm), and ArithMark-3.0. Handles **both** families. Self-contained. | |
| | `dpo_eval_quazimoto.py` | Held-out DPO preference accuracy (UltraFeedback `test_prefs`) for a Quazimoto policy vs frozen reference. | |
| | `dpo_eval_spikewhale.py` | Same, for a SpikeWhale policy. Needs `sw_common.py`. | |
| | `sw_common.py` | Shared SpikeWhale helpers (loader, ChatML renderer, bucketed-pad seq log-prob). | |
|
|
| ```bash |
| python run_eval.py --family {quazimoto|spikewhale} \ |
| --code_dir <model-repo-dir> --ckpt <ckpt.pt-or-dir> \ |
| --name <label> --out out/<label>.json |
| ``` |
|
|
| ## Generation |
|
|
| | file | family | modes | |
| |---|---|---| |
| | `gen_quazimoto_single.py` | Quazimoto (Wheeler/Positronic) | 5 single-turn prompts | |
| | `gen_quazimoto_multiturn.py` | Quazimoto | 2 scripted multi-turn conversations | |
| | `gen_spikewhale.py` | SpikeWhale (Byrne) | `--mode single` or `--mode multi` | |
| | `infer_v2.py` | — | sampler used by `gen_spikewhale.py` | |
|
|
| ```bash |
| # Quazimoto (run from the model repo dir so `generate.py`/`model.py` import) |
| python gen_quazimoto_single.py --ckpt <ckpt.pt> --device cuda |
| python gen_quazimoto_multiturn.py --ckpt <ckpt.pt> --device cuda |
| |
| # SpikeWhale (run from the Byrne Mark2 dir so `model_v2.py`/`config.py` import) |
| python gen_spikewhale.py --ckpt <ckpt.pt-or-dir> --device cuda --mode single |
| python gen_spikewhale.py --ckpt <ckpt.pt-or-dir> --device cuda --mode multi |
| ``` |
|
|
| ## Dependencies / how to run |
|
|
| `run_eval.py` is standalone. The generation and DPO-eval scripts load models via |
| each repo's OWN code, so run them **from inside (or with on the path) the model |
| repo**: |
| - **Quazimoto** scripts import `generate.py` / `model.py` (Wheeler/Positronic repo). |
| - **SpikeWhale** scripts import `model_v2.py` / `config.py` / `spike_tokenizer.py` |
| (Byrne `Mark2/`); `gen_spikewhale.py` also uses `infer_v2.py` (copied here) and `sw_common.py`. |
|
|
| The tokenizer (`tokenizer.json` + `spike_tokenizer.py`) lives in the model repos and |
| is shared across all these models (SpikeWhale byte-level, vocab 16512). |
|
|
| ## Notes |
|
|
| - ArithMark target is `AxiomicLabs/ArithMark-3.0` (top of `run_eval.py`). |
| - Datasets stream from HuggingFace; `datasets`/`pyarrow` import before torch/CUDA (Windows ordering). |
| - DPO scripts pad sequence lengths to multiples of 128 to avoid CUDA allocator |
| fragmentation on Windows (no `expandable_segments` there). |
|
|