Instructions to use FINAL-Bench/Aether-7B-5Attn-it with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FINAL-Bench/Aether-7B-5Attn-it with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="FINAL-Bench/Aether-7B-5Attn-it", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("FINAL-Bench/Aether-7B-5Attn-it", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use FINAL-Bench/Aether-7B-5Attn-it with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FINAL-Bench/Aether-7B-5Attn-it" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FINAL-Bench/Aether-7B-5Attn-it", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/FINAL-Bench/Aether-7B-5Attn-it
- SGLang
How to use FINAL-Bench/Aether-7B-5Attn-it with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FINAL-Bench/Aether-7B-5Attn-it" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FINAL-Bench/Aether-7B-5Attn-it", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FINAL-Bench/Aether-7B-5Attn-it" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FINAL-Bench/Aether-7B-5Attn-it", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use FINAL-Bench/Aether-7B-5Attn-it with Docker Model Runner:
docker model run hf.co/FINAL-Bench/Aether-7B-5Attn-it
license: apache-2.0
language:
- en
- ko
base_model: FINAL-Bench/Aether-7B-5Attn
datasets:
- HuggingFaceFW/fineweb-edu
- HuggingFaceTB/smollm-corpus
- HuggingFaceTB/finemath
- open-web-math/open-web-math
- OpenCoder-LLM/opc-fineweb-code-corpus
- HAERAE-HUB/KOREAN-WEBTEXT
- HAERAE-HUB/KOREAN-SyntheticText-1.5B
pipeline_tag: text-generation
library_name: transformers
tags:
- foundation-model
- sovereign-ai
- fully-open
- open-source
- mixture-of-experts
- moe
- heterogeneous-attention
- latin-square
- instruct
- sft
- fine-tuned
- korean
- vidraft
- aether
π± Run it on your phone or a GPU-less PC β POCKET Β· π Try it live (CPU chat)
VIDRAFT's on-device family: a 35B model that runs on iPhone and on CPU with no GPU β stock
llama.cpp, no fork.
Aether-7B-5Attn-it
Update β 2026-07-20 (Korean instruction tuning v3)
This checkpoint has been refreshed with an expanded Korean instruction-tuning run.
| Training data | 14,865 samples β Korean general instructions (12,000) + Korean exam-style reasoning (2,400) + model identity (465) |
| Method | LoRA (r16, alpha32) on q/k/v/o/gate/up/down projections, completion-only loss, entropy-gated weighting |
| Schedule | 1 epoch, LR 2e-4, merged back into the base weights |
| Language policy | Korean prompts are answered in Korean, English prompts in English (language-matched training data) |
What changed
- Korean responsiveness. The previous checkpoint frequently returned empty completions for Korean prompts; this revision answers them.
- Model identity. The model now identifies itself as an AETHER model developed by VIDRAFT.
- Language matching. Korean in, Korean out; English in, English out.
Known limitations
- Factual accuracy in Korean history/knowledge remains weak. The instruction mix was not large enough to reliably ground factual questions; verify factual claims before relying on them.
- Repetition under greedy decoding. Use sampling (e.g.
temperature=0.7,top_p=0.9) rather than pure argmax decoding. - This is a 7B-class MoE base with light instruction tuning, not a fully aligned chat model.
Which Aether model should I use?
| Model | What it is | Pick it if |
|---|---|---|
| Aether-7B-5Attn | 6.59B MoE base. Fully open - weights + data recipe + training code + all 162k-step logs + checkpoints | You want to audit, verify or rebuild a foundation model end to end |
| Aether-7B-5Attn-it | The same model, instruction-tuned | You want it to answer rather than continue text |
| AETHER-7B-7Attn-base | Same 49-layer architecture, a different checkpoint. Open weights | You want a second run of this architecture to compare against |
| Aether-6B-11Attn-base | 121 layers, 11 sequence-mixing mechanisms in one network - attention, Mamba-2, Hyena, GDN, MLA - on an 11x11 Latin square | You research heterogeneous sequence mixing. It is a mid-training research artifact |
All four load the same way:
AutoModelForCausalLM.from_pretrained(MODEL, trust_remote_code=True, dtype=torch.bfloat16)
π Part of the Aether Foundation Model collection β base, instruction-tuned, and checkpoints in one place.
π§© Intermediate checkpoints (110k Β· 115k Β· 162k) are released as a dataset: Aether-7B-5Attn-checkpoints. Instruction-tuned (SFT) version of the fully-open Aether-7B-5Attn base model. Post-trained for multiple-choice / benchmark-style answering. 6.59B MoE (~2.98B active), 49 layers on a 7Γ7 Latin square. Apache-2.0.
Learning rate chosen by held-out accuracy, not loss
Full-parameter SFT was run at three learning rates (2e-6 / 6e-6 / 2e-5). The final checkpoint was selected by held-out benchmark accuracy, not training loss β for small models a high LR lowers training loss while degrading real capability, so loss is the wrong selector.
| LR | K-AI 4 avg | Note |
|---|---|---|
| 2e-6 | 29.7% | Under-fit (weak on some subjects) |
| 6e-6 (selected) | 34.9% | Best & balanced, no degradation |
| 2e-5 | 32.3% | One subject collapsed (format-degradation sign) |
Held-out results (selected checkpoint, 6e-6)
- GPQA-Diamond (198): 25.3%
- K-AI 4 average (195): 34.9%
- musr_ko 26.5% Β· com2_main_ko 50.0% Β· click 32.0% Β· kommlu_pro 30.4%
- vs base (before SFT): base 26.7% β 34.9% (+8.2pp) β same held-out set, same harness.
All evaluations use held-out sets not seen in training. SFT was performed on MMLU-auxiliary (multiple-choice format), which does not overlap with the evaluation subjects (GPQA / K-AI).
Pretraining data mix (inherited from base)
| Domain | Share |
|---|---|
| Math (finemath + open-web-math) | 37.8% |
| Korean (webtext + synth) | 21.6% |
| English web & synthetic (fineweb-edu + cosmopedia) | 21.6% |
| Code (opc) | 13.5% |
| phase15 pre-blend | 5.4% |
This is the base pretraining mix. This model's post-training (SFT) data is MMLU-auxiliary.
Architecture (inherited from base)
Identical to Aether-7B-5Attn base: 49 layers placed on a 7Γ7 Latin square with heterogeneous attention (7 labels / 5 distinct mechanisms) and a 25-expert MoE (top-7 + 1 shared). Full structural detail and the diagram are in the base card, Β§3.2: FINAL-Bench/Aether-7B-5Attn.
Open-source fully-open LLMs β 6-country comparison
Relative to the base Aether. Among six sovereign fully-open models, VIDRAFT is the only single AI startup, and Aether has the most attention types (5) in a Latin-square layout.
Running the model
1. Loading
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL = "FINAL-Bench/Aether-7B-5Attn-it"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(
MODEL, trust_remote_code=True, dtype=torch.bfloat16, device_map="cuda"
).eval()
trust_remote_code=True is required: this is a custom architecture (aether_v2_7way),
and the modeling code ships in this repository. Loading needs roughly 14 GB of VRAM
in bfloat16.
The module files also remain under
aether_pkg/for anyone who was importing them directly; the copies at the repository root are whattrust_remote_coderesolves.
2. use_cache=False is mandatory
This architecture ships no KV cache. Leaving use_cache enabled raises an
IndexError from the standard cache path. Set it on the config and keep it off
in generate().
3. Generation is slow β plan for it
With no KV cache, every new token re-runs the full forward pass, so decoding is O(nΒ²) in sequence length. Measured throughput:
| Hardware | Throughput |
|---|---|
| NVIDIA T4 (16 GB) | ~1.5 tokens/s β 64 tokens takes about 40 s |
| NVIDIA B200 | ~7 tokens/s |
Keep max_new_tokens small. This is a property of the released architecture, not a
configuration problem.
# REQUIRED: this checkpoint was trained with attention_mask=None. generate() builds a
# mask automatically, which puts the model off-distribution and degenerates the output
# (you get things like "κ΅κ°μ μλλ κ΅κ°μ μλμ
λλ€..." instead of an answer).
# Drop the mask before it reaches the model:
_forward = model.forward
def _forward_without_mask(*a, **kw):
kw.pop("attention_mask", None)
return _forward(*a, **kw)
model.forward = _forward_without_mask
out = model.generate(
tok(prompt, return_tensors="pt").input_ids.to("cuda"),
max_new_tokens=64, do_sample=False, use_cache=False,
pad_token_id=tok.eos_token_id,
)
print(tok.decode(out[0], skip_special_tokens=True))
Why the mask has to go
Measured on 2026-07-20, same prompt, greedy decoding, only the mask varied:
| attention_mask | Output for λλ λꡬμΌ? |
|---|---|
| absent | μ λ λΉλλννΈ(VIDRAFT)μ AETHER λͺ¨λΈμ λλ€. 7μ’ μ μλ‘ λ€λ₯Έ μ΄ν μ λ©μ»€λμ¦μβ¦ |
| present (generate default) | λΉλλννΈ(VIDRAFT)μ νκ΅μ΄μ λλ€. λΉλλννΈλ λΉλλννΈμβ¦ (degenerate) |
This is a property of how the checkpoint was tuned, not a bug in generate().
Greedy decoding is also recommended β sampling drifts off-distribution on this checkpoint.
4. Prompt format
Instruction tuning used a plain format β the instruction, a blank line, then the
response. The bundled chat_template.jinja reproduces exactly this, so
apply_chat_template and the raw form below are equivalent:
prompt = "λνλ―Όκ΅μ μλλ μ΄λμΈκ°μ?" + "\n\n"
5. Quality caveats β please read before use
Instruction tuning here is light: 14,865 samples for a single epoch. Measured consequences, stated plainly:
- Sensitive to prompt format and decoding settings. Outputs degrade sharply once you move away from the trained format. Greedy decoding on the format above is the only configuration we validated.
- Factual accuracy in Korean history and general knowledge is weak. The model produces fluent but confidently wrong statements. Verify anything factual.
- Benchmark: on a held-out 4-subject Korean set (CLIcK / KMMLU / MuSR / Com2, 80 items, no train overlap) this revision scores 31.2% against 33.8% for the previous one β within noise of each other and near the 25% random baseline.
- Treat this as a lightly instruction-tuned research checkpoint, not a production assistant.
Contact
VIDRAFT (μ£Όμνμ¬ λΉλλννΈ) Β· arxivgpt@gmail.com Β· License: Apache-2.0

