| --- |
| license: apache-2.0 |
| tags: |
| - test-fixture |
| --- |
| |
| # Speculative-decoding subfolder test fixture |
|
|
| A **tiny, randomly-initialised** Llama pair for exercising *subfolder* loading of |
| a speculative draft model. It exists so inference engines can test the code path |
| in CI without downloading a real checkpoint. |
|
|
| **These weights are random. The model produces gibberish. That is the point.** |
|
|
| ``` |
| . <- target model, 2 layers, ~4M params |
| └── DFlash/ <- draft model, 1 layer, ~4M params |
| ``` |
|
|
| Both share the same tokenizer and vocabulary, so the draft is a valid speculator |
| for the target. |
|
|
| ## Why a subfolder needs engine support |
|
|
| A Hugging Face repo id may contain **at most one `/`**, so `org/repo/DFlash` is |
| not a valid id and cannot be passed as a single string. A draft bundled inside |
| its target's repository therefore has to be addressed as a repo id *plus* a |
| separate sub-path, which most engines have no parameter for. |
|
|
| This matters in practice: shipping a draft inside its target repo keeps the pair |
| versioned together and downloadable in one fetch, rather than as two repos that |
| can drift apart. |
|
|
| ## Usage |
|
|
| ```python |
| from huggingface_hub import snapshot_download |
| snapshot_download("poolside/spec-decoding-subfolder-fixture") # target + draft |
| snapshot_download("poolside/spec-decoding-subfolder-fixture", |
| allow_patterns=["DFlash/*"]) # draft only |
| ``` |
|
|
| Engines adding support should accept the draft as |
| `(model="poolside/spec-decoding-subfolder-fixture", subfolder="DFlash")`. |
|
|
| ## Provenance |
|
|
| Built with `transformers` from `LlamaConfig` (hidden 64, intermediate 128, |
| 4 heads, tied embeddings), `torch.manual_seed(0)`. Tokenizer copied from |
| `hf-internal-testing/llama-tokenizer`. No trained parameters, no proprietary |
| data. |
|
|