Text Generation
Safetensors
English
fleck
causal-language-model
custom-architecture
conversational
research
Fleck-S-100K / README.md
brulee-1's picture
Update README.md
ff852c2 verified
|
Raw
History Blame Contribute Delete
5.92 kB
---
language:
- en
license: mit
tags:
- fleck
- causal-language-model
- custom-architecture
- text-generation
- conversational
- research
base_model: ML-is-Fun/Fleck-S-100K-Base
datasets:
- HuggingFaceFW/fineweb
- HuggingFaceFW/fineweb-edu
- databricks/databricks-dolly-15k
---
# [Fleck-S-100K](https://huggingface.co/ML-is-Fun/Fleck-S-100K)
**[Fleck-S-100K Instruct](https://huggingface.co/ML-is-Fun/Fleck-S-100K)** β€” the independently instruction-tuned child of [`Fleck-S-100K-Base`](https://huggingface.co/ML-is-Fun/Fleck-S-100K-Base).
- **109,384 parameters**
- **BF16 SafeTensors weights**
- Fine-tuned from the Base-100M model on Dolly-15k
- Intended for local conversational experiments on Apple Silicon
## Model Details
| | |
| --- | --- |
| Architecture | Decoder-only Transformer |
| Parameters | 109,384 |
| Hidden size | 64 |
| FFN size | 128 |
| Physical blocks | 2 |
| Effective depth | 4 (`A β†’ B β†’ A β†’ B`) |
| Attention | GQA β€” 4 query heads, 2 KV heads, head dimension 16 |
| Normalization | RMSNorm |
| Embedding | Factorized tied embedding, rank 32 |
| Vocabulary | 1,024 |
| Context length | 2,048 tokens |
| Canonical dtype | BF16 |
| Base model | [`ML-is-Fun/Fleck-S-100K-Base`](https://huggingface.co/ML-is-Fun/Fleck-S-100K-Base) |
## Training
### Pretraining
The Base parent was initialized from scratch and pretrained for exactly 100,000,000 real tokenizer tokens on a 70/30 FineWeb-Edu/FineWeb mixture.
### Instruction Tuning
| | |
| --- | --- |
| Base model | [`Fleck-S-100K-Base`](https://huggingface.co/ML-is-Fun/Fleck-S-100K-Base) |
| Dataset | [`databricks/databricks-dolly-15k`](https://huggingface.co/datasets/databricks/databricks-dolly-15k) |
| Dataset license | CC BY-SA 3.0 |
| Requested budget | 100K supervised tokens |
| Realized supervised tokens | 101,832 |
| Optimizer | AdamW |
| Instruction tuning | Assistant-response supervision |
| Hardware | Apple M2 (10-core GPU) |
## Benchmark Results
Corrected zero-shot evaluation, no chat template, FP32 evaluation, and the same benchmark aggregation protocol were used for both variants.
| Task | Metric | Shots | Base | **Instruct** | Ξ” |
| --- | --- | ---: | ---: | ---: | ---: |
| HellaSwag | `acc_norm` | 0 | 26.28% | **26.29%** | +0.01pp |
| PIQA | `acc_norm` | 0 | 51.14% | **51.36%** | +0.22pp |
| ARC-Easy | `acc_norm` | 0 | 26.05% | 25.72% | -0.34pp |
| ARC-Challenge | `acc_norm` | 0 | 26.19% | 25.94% | -0.26pp |
| LAMBADA OpenAI | `acc` | 0 | 0.02% | **0.06%** | +0.04pp |
| WinoGrande | `acc` | 0 | 49.41% | **50.91%** | +1.50pp |
| BoolQ | `acc` | 0 | 37.83% | 37.83% | 0.00pp |
| MMLU (57-subject macro) | `acc` | 0 | 23.09% | **23.12%** | +0.03pp |
| **Eight-task mean** | β€” | 0 | 30.00% | **30.15%** | +0.15pp |
A generation and contamination sanity check found no exact prompt/example matches among the audited records. This is a diagnostic, not a guarantee of absence of all contamination.
## Tokenizer and Chat Format
- `Fleck-Tokenizer-1024`
- Byte-level BPE
- Vocabulary size: 1,024
The training chat format is:
```text
<bos><|user|>{user}<|eot|><|assistant|>{response}<|eot|><eos>
```
### Special Tokens
| Token | ID | Role |
| --- | ---: | --- |
| `<bos>` | 0 | sequence start |
| `<eos>` | 1 | sequence end |
| `<pad>` | 2 | padding |
| `<unk>` | 3 | unknown token |
| `<\|system\|>` | 4 | system turn |
| `<\|user\|>` | 5 | user turn |
| `<\|assistant\|>` | 6 | assistant turn |
| `<\|eot\|>` | 7 | end of turn |
## Usage
The bundle includes a self-contained `inference.py`; it does not import the Fleck-LM checkout. The accompanying `config.json`, `generation_config.json`, and `tokenizer_config.json` describe the custom architecture and generation/tokenizer defaults; standard `transformers.AutoModel` loading is not supported. The `chat_template.jinja` file contains the Instruct chat template used by compatible HF tooling. Install the three runtime dependencies:
```bash
python -m pip install torch safetensors tokenizers
```
By default the CLI starts an interactive chat. `/exit` quits and `/clear` resets the conversation history:
```bash
python inference.py --device cpu --max-tokens 32
```
For one prompt without interactive mode, pass `--no-chat` and `--prompt`:
```bash
python inference.py \
--ckpt model.safetensors \
--tokenizer tokenizer.json \
--no-chat \
--prompt "Explain what a tokenizer does." \
--max-tokens 32 \
--device cpu
```
The default filenames are `model.safetensors` and `tokenizer.json`. The script reads and runs the BF16 checkpoint without an FP32 model copy, validates every SafeTensors key, shape, and dtype, and uses FP32 only for attention score/softmax and tied-logit accumulation. It reproduces the factorized tied embedding/logits, effective-depth execution `A β†’ B β†’ A β†’ B`, half-split RoPE, GQA, physical KV caches, RMSNorms, and greedy generation without any repository-local imports. Generation stops on `<|eot|>` or `<eos>`.
## Limitations
This model is extremely small and is intended for research and local experimentation rather than reliable general-purpose assistance. It may produce repetitions, malformed text, weak factual answers, or incoherent responses. Instruction tuning improves conversational behavior but does not overcome the limits of a 109K parameter model.
## License
MIT License.
## Files
The public bundle contains these files:
- `README.md` β€” model card and usage documentation
- `inference.py` β€” standalone strict loader and interactive/single-prompt inference CLI
- `model.safetensors` β€” BF16 model weights
- `tokenizer.json` β€” standalone tokenizer
- `config.json` β€” custom architecture configuration
- `generation_config.json` β€” greedy generation defaults
- `tokenizer_config.json` β€” tokenizer defaults and special-token mapping
- `chat_template.jinja` β€” Instruct chat template
No training data, optimizer state, or other training outputs are included.