Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Hana-ame
/
additive-rand-transformer
like
0
Reinforcement Learning
PyTorch
arithmetic
chain-of-thought
transformer
tiny-model
quantization
diy-manual
License:
mit
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
additive-rand-transformer
41.3 MB
Ctrl+K
Ctrl+K
1 contributor
History:
346 commits
Hana-ame
feat(experiments): expand experiment matrix to 220, add LoRA/format/hyperparams sweeps and configs
ce812c8
about 3 hours ago
additive_rand_transformer
chore: sync train.py eval fixes, use_model.sh python path, and latest configs
about 4 hours ago
archive
chore: archive research markdown docs (fully integrated into Excel)
about 4 hours ago
checkpoints
feat: publish seg2_b16 checkpoint (L4·D64 seg2 b16, bottleneck seg variant) via git-lfs
2 days ago
configs
feat(experiments): expand experiment matrix to 220, add LoRA/format/hyperparams sweeps and configs
about 3 hours ago
runs
协同适应探针: coadapt_l4d128.json
about 17 hours ago
.gitattributes
Safe
1.52 kB
initial commit
3 days ago
.gitignore
Safe
124 Bytes
docs: add PARAMS.md (scan params by code), cross-link in RESEARCH/README/RESULTS, gitignore runs+pycache
3 days ago
Colab_Run_Additive_Transformer.ipynb
13.7 kB
feat(additive): add --config support, Colab notebook, INT8 quantization module, and full experiments excel
about 4 hours ago
EXPERIMENTS_ALL.xlsx
58 kB
feat(experiments): expand experiment matrix to 220, add LoRA/format/hyperparams sweeps and configs
about 3 hours ago
LICENSE
Safe
1.07 kB
feat: package layout + model card + 7 checkpoints
3 days ago
README.md
8.53 kB
docs: 完善中文文档与 DIY 操作手册
about 3 hours ago
param_variation2_all.json
11.2 kB
param_variation2 21档网格 + 21f RoPE: param_variation2_all.json
1 day ago
ppl_matrix.json
37.3 kB
fixed testset + all-7-model perplexity sweep (ad1ee65)
2 days ago
ppl_matrix_full.json
115 kB
full-matrix perplexity: add/sub split + all role segments (756693f)
2 days ago
pyproject.toml
Safe
1.08 kB
feat: package layout + model card + 7 checkpoints
3 days ago
requirements.txt
Safe
23 Bytes
Upload folder using huggingface_hub
3 days ago
rope_retrain.json
406 Bytes
param_variation2 21档网格 + 21f RoPE: rope_retrain.json
1 day ago
step_sweep.json
11 kB
step sweep: continue-training +500/+1000/+2000/+4000 for all 7 checkpoints, fixed-testset eval (621d1ae)
2 days ago
testset_adder.json
20.5 kB
feat+docs: fixed test set + all-7-model perplexity sweep (per-digit bucketed)
2 days ago
use_model.sh
6.59 kB
chore: sync train.py eval fixes, use_model.sh python path, and latest configs
about 4 hours ago