nancyH's picture
Upload README.md with huggingface_hub
17d4181 verified
|
Raw
History Blame Contribute Delete
4.95 kB
---
license: mit
tags:
- dna
- genomics
- tokenization
- sequence-classification
---
# Baseline BPE-100K β€” final SFT checkpoints
Fine-tuned classification checkpoints for all 56 downstream tasks, produced from the
**baseline BPE** pretrained model at step 100,000 (vocab 5,120).
These are the checkpoints behind the **Base row of Table 1** in
*EvoLen: Evolution-Guided Tokenization for DNA Language Model*
([arXiv:2604.08698](https://arxiv.org/abs/2604.08698)).
Base model: [`EvoLenTokenizer/base-100k`](https://huggingface.co/EvoLenTokenizer/base-100k).
Tokenizer: standard BPE, vocab 5,120 (`baseline_bpe/vocab_5120/5120_tokenizer.json`) β€”
**not** the EvoLen tokenizer.
This is the controlled comparison for
[`EvoLenTokenizer/evolen-100k-sft-checkpoints`](https://huggingface.co/EvoLenTokenizer/evolen-100k-sft-checkpoints):
same architecture, same data, same pretraining budget, same fine-tuning pipeline β€” the
only difference is the tokenizer.
## Verification against the paper
Group averages reproduce Table 1's Base row (MCC Γ—100):
| group | n | this repo | Table 1 | | group | n | this repo | Table 1 |
|---|---|---|---|---|---|---|---|---|
| GUE EMP | 10 | 46.78 | 46.8 | | NT His | 10 | 53.74 | 53.7 |
| GUE Mou | 5 | 56.45 | 56.5 | | NT Enh | 2 | 48.71 | 48.7 |
| GUE P3 | 3 | 76.45 | 76.5 | | NT Pro | 3 | 74.65 | 74.7 |
| GUE PC | 3 | 64.25 | 64.2 | | NT Spl | 3 | 69.31 | 69.3 |
| GUE Spl | 1 | 78.53 | 78.5 | | cCRE | 1 | 21.92 | 21.9 |
| GUE TF | 5 | 57.75 | 57.8 | | ATAC | 1 | 12.15 | 12.2 |
| GBM HR | 5 | 69.93 | 69.9 | | GBM ME | 1 | 58.13 | 58.1 |
| GBM Inv | 3 | 76.44 | 76.4 | | | | | |
Each checkpoint was additionally cross-checked against the `eval_results.json` in its own
run directory (56/56 within 1e-6), and a sample was re-run from the weights themselves β€”
`EMP/H3` 0.695965, `prom/prom_core_tata` 0.645892, `NT/H2AFZ` 0.481101 and
`multi-SCREEN/allchr_csv` 0.219183 all reproduced exactly.
## Layout
One folder per task, `<SUITE>/<task>/`, each with `model.safetensors`, `config.json`,
`tokenizer.json`, `tokenizer_config.json`, `special_tokens_map.json`,
`trainer_state.json` and `training_args.bin`. Optimizer and scheduler state are stripped.
```
GBM/demo_coding_vs_intergenomic_seqs/
GUE/EMP/H3K4me1/
NT/H3K27ac/
multi-SCREEN/allchr_csv/ # the "cCRE" column of Table 1
multi-ATAC/human_mouse_superclass_allchr/ # the "ATAC" column of Table 1
```
`per-task` hyperparameters, including `model_max_length` (which is per-task and not
derivable from the task name), are recoverable from `training_args.bin`:
```python
import torch, sys
class TrainingArguments: pass
sys.modules["__main__"].TrainingArguments = TrainingArguments
args = torch.load("training_args.bin", map_location="cpu", weights_only=False).__dict__
```
`reported_scores.csv` lists every task with its reported MCC and the config that produced it.
## Loading
```python
from transformers import AutoModelForSequenceClassification
REPO = "EvoLenTokenizer/base-100k-sft-checkpoints"
model = AutoModelForSequenceClassification.from_pretrained(
REPO, subfolder="GUE/EMP/H3K4me1", trust_remote_code=True)
```
## Reproducing the reported numbers
Evaluation only. Four things must match:
1. **fp16 autocast** β€” evaluating in fp32 shifts MCC by ~3e-4.
2. **Eval batch size 128** β€” fp16 reduction order depends on batch shape.
3. **`model_max_length` per task**, from `training_args.bin`.
4. **Multiclass label ordering** β€” `sorted(set(...))` over the split file. For
`multi-SCREEN/allchr_csv`: `['CA','CA-CTCF','CA-H3K4me3','CA-TF','PLS','TF','dELS','pELS']` β†’ 0..7.
Tokenize with `padding="longest"`, `truncation=True`, `max_length=model_max_length`;
attention mask is `input_ids.ne(pad_token_id)`. Metric is
`sklearn.metrics.matthews_corrcoef` over `argmax` predictions.
Verified environment: Python 3.9.18, transformers 4.35.2, scikit-learn 1.6.1,
torch 2.8.0, numpy 2.0.2, tokenizers 0.15.2, accelerate 0.25.0.
## Selection procedure
Training ran for `num_train_epochs` with per-epoch validation on `dev.csv`;
`load_best_model_at_end=True` with `metric_for_best_model=eval_f1` reloaded the
best-validation-F1 checkpoint, which was evaluated once on `test.csv`. Per-task
hyperparameters were the maximum test MCC over the swept configurations, so reported
values are an upper bound rather than an unbiased estimate. Retraining will not reproduce
them exactly β€” fp16 training is nondeterministic and epoch-selection margins are small.
## Related
- [`EvoLenTokenizer/evolen-100k-sft-checkpoints`](https://huggingface.co/EvoLenTokenizer/evolen-100k-sft-checkpoints) β€” the EvoLen counterpart
- [`EvoLenTokenizer/evolen-100k`](https://huggingface.co/EvoLenTokenizer/evolen-100k) / [`base-100k`](https://huggingface.co/EvoLenTokenizer/base-100k) β€” pretrained models
- [`EvoLenTokenizer/evolen-200k-sft-checkpoints`](https://huggingface.co/EvoLenTokenizer/evolen-200k-sft-checkpoints)