Spaces:
Running on T4
Running on T4
| <html lang="en"> | |
| <head> | |
| <meta charset="utf-8"> | |
| <meta name="viewport" content="width=device-width, initial-scale=1"> | |
| <title>Aether-7B-5Attn — 100% Open-source Foundation Model: Sovereign AI</title> | |
| <meta name="description" content="Aether-7B-5Attn is a 100% open-source sovereign foundation model: weights, training-data recipe, code, hyperparameters and full logs, all released under Apache-2.0. A 6.59B heterogeneous-attention MoE built from scratch by VIDRAFT."> | |
| <meta name="keywords" content="open source LLM, fully open foundation model, sovereign AI, Korean LLM, open weights data code logs, Apache 2.0, heterogeneous attention, Latin square attention, mixture of experts, reproducible LLM, VIDRAFT, Aether"> | |
| <meta name="author" content="VIDRAFT"> | |
| <meta property="og:type" content="website"> | |
| <meta property="og:title" content="Aether-7B-5Attn — 100% Open-source Foundation Model: Sovereign AI"> | |
| <meta property="og:description" content="100% open-source sovereign foundation model — weights, data recipe, code, and full logs under Apache-2.0."> | |
| <meta name="twitter:card" content="summary_large_image"> | |
| <script type="application/ld+json"> | |
| { | |
| "@context": "https://schema.org", | |
| "@type": "SoftwareApplication", | |
| "name": "Aether-7B-5Attn", | |
| "applicationCategory": "Large Language Model", | |
| "operatingSystem": "Any", | |
| "creator": {"@type": "Organization", "name": "VIDRAFT"}, | |
| "license": "https://www.apache.org/licenses/LICENSE-2.0", | |
| "description": "A 100% open-source sovereign foundation model: a 6.59B heterogeneous-attention MoE (5 attention mechanisms on a 7x7 Latin square) trained from scratch on 144.2B tokens. Weights, training-data recipe, code, hyperparameters and full logs released under Apache-2.0.", | |
| "keywords": "open source LLM, sovereign AI, fully open foundation model, Apache-2.0, heterogeneous attention, mixture of experts" | |
| } | |
| </script> | |
| <style> | |
| :root{--bg:#0B1020;--card:#121a30;--line:#233; --ink:#E7EEFB;--mut:#93A4C4;--acc:#5B8FF9;--good:#5AD8A6;} | |
| *{box-sizing:border-box} | |
| body{margin:0;background:radial-gradient(1200px 600px at 70% -10%,#16213f 0,#0B1020 60%);color:var(--ink); | |
| font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",Roboto,"Noto Sans KR",sans-serif;line-height:1.6} | |
| .wrap{max-width:1100px;margin:0 auto;padding:0 20px} | |
| header{padding:56px 0 28px} | |
| .badge{display:inline-block;font-size:12px;letter-spacing:.14em;color:var(--good);border:1px solid #2b6b52; | |
| border-radius:999px;padding:5px 12px;text-transform:uppercase} | |
| h1{font-size:clamp(28px,5vw,46px);line-height:1.15;margin:16px 0 10px;font-weight:800} | |
| h1 .g{background:linear-gradient(90deg,#7FB2FF,#5AD8A6);-webkit-background-clip:text;background-clip:text;color:transparent} | |
| .sub{color:var(--mut);font-size:18px;max-width:760px} | |
| .cta{margin:22px 0 4px;display:flex;gap:12px;flex-wrap:wrap} | |
| .btn{display:inline-block;padding:11px 18px;border-radius:10px;text-decoration:none;font-weight:600} | |
| .btn.p{background:var(--acc);color:#06122b} | |
| .btn.o{border:1px solid #2c3c5e;color:var(--ink)} | |
| .grid{display:grid;grid-template-columns:1.15fr .85fr;gap:22px;margin:26px 0} | |
| @media(max-width:860px){.grid{grid-template-columns:1fr}} | |
| .card{background:var(--card);border:1px solid #1f2b47;border-radius:16px;padding:20px} | |
| .card h2{margin:2px 0 10px;font-size:19px} | |
| .demo textarea{width:100%;min-height:110px;background:#0c142a;color:var(--ink);border:1px solid #26324a; | |
| border-radius:10px;padding:12px;font-size:15px;resize:vertical;font-family:inherit} | |
| .demo .row{display:flex;gap:10px;align-items:center;margin-top:10px;flex-wrap:wrap} | |
| .demo button{background:var(--good);color:#06251a;border:0;border-radius:10px;padding:11px 18px;font-weight:700;cursor:pointer} | |
| .demo button:disabled{opacity:.5;cursor:not-allowed} | |
| .out{margin-top:12px;background:#0a1226;border:1px solid #23304d;border-radius:10px;padding:14px;white-space:pre-wrap; | |
| min-height:70px;color:#D8E4FA} | |
| .muted{color:var(--mut);font-size:13px} | |
| ul.feat{list-style:none;padding:0;margin:6px 0} | |
| ul.feat li{padding:7px 0;border-bottom:1px solid #18223c} | |
| ul.feat li b{color:var(--good)} | |
| table{width:100%;border-collapse:collapse;font-size:14px;margin-top:6px} | |
| th,td{text-align:left;padding:8px 10px;border-bottom:1px solid #1a2540} | |
| th{color:var(--mut);font-weight:600} | |
| .qa h3{margin:16px 0 4px;font-size:16px;color:#BFE3FF} | |
| .qa p{margin:0 0 10px;color:var(--mut)} | |
| footer{color:#6B7C99;font-size:13px;padding:30px 0 50px;border-top:1px solid #17203a;margin-top:20px} | |
| a{color:#7FB2FF} | |
| </style> | |
| </head> | |
| <body> | |
| <div class="wrap"> | |
| <header> | |
| <span class="badge">Apache-2.0 · Fully Open · Sovereign AI</span> | |
| <h1>Aether-7B-5Attn — <span class="g">100% Open-source</span> Foundation Model</h1> | |
| <p class="sub">Not "open weights." Genuinely open — weights, training-data recipe, code, every hyperparameter, and the complete logs, all released under Apache-2.0. A 6.59B heterogeneous-attention MoE, built from scratch by <b>VIDRAFT</b>.</p> | |
| <div class="cta"> | |
| <a class="btn p" href="#try">Try the model</a> | |
| <a class="btn o" href="https://huggingface.co/FINAL-Bench/Aether-7B-5Attn" target="_blank" rel="noopener">Base model ↗</a> | |
| <a class="btn o" href="https://huggingface.co/FINAL-Bench/Aether-7B-5Attn-it" target="_blank" rel="noopener">Instruct model ↗</a> | |
| </div> | |
| </header> | |
| <div class="grid"> | |
| <div class="card demo" id="try"> | |
| <h2>🌏 Try Aether (live)</h2> | |
| <p class="muted">Answers here come from <strong><a href="https://huggingface.co/FINAL-Bench/Aether-7B-5Attn-it">Aether-7B-5Attn-it</a></strong>, the instruction-tuned model — <em>not</em> the base. The base (<a href="https://huggingface.co/FINAL-Bench/Aether-7B-5Attn">Aether-7B-5Attn</a>) is the reproducible artifact this release is about; it is not instruction-tuned and would not answer questions like this.</p> | |
| <textarea id="prompt" placeholder="Ask something, or start a sentence for the model to continue...">In one sentence, what does sovereign open-source AI mean?</textarea> | |
| <div class="row"> | |
| <button id="go">Generate</button> | |
| <span class="muted" id="stat">Model loads on first request.</span> | |
| </div> | |
| <div class="out" id="out">—</div> | |
| <p class="muted">Serving <code>Aether-7B-5Attn-it</code> (6.59B MoE) on a single T4, at batch_size=1. This build ships no KV cache, so generation is slow and capped — tokens stream in as they are produced.</p> | |
| </div> | |
| <div class="card"> | |
| <h2>Why it matters</h2> | |
| <ul class="feat"> | |
| <li><b>100% open</b> — weights · data recipe · code · logs</li> | |
| <li><b>Reproducible</b> — data rebuilds byte-for-byte from public sources</li> | |
| <li><b>Self-designed</b> — 5 attention mechanisms on a 7×7 Latin square</li> | |
| <li><b>Sovereign</b> — rebuild your own foundation model from this repo alone</li> | |
| <li><b>Apache-2.0</b> — commercial use, modification, redistribution</li> | |
| </ul> | |
| </div> | |
| </div> | |
| <div class="card"> | |
| <h2>What is Aether-7B-5Attn?</h2> | |
| <p class="muted">A 6.59-billion-parameter Mixture-of-Experts language model (~2.98B active per token) whose 49 layers place <b>five distinct attention mechanisms</b> on a <b>7×7 Latin square</b>, so each mechanism appears exactly once at every depth. Trained from scratch on <b>144.2B tokens</b> (≈37.8% math, 21.6% Korean, 21.6% English, 13.5% code) and released fully open under Apache-2.0.</p> | |
| <table> | |
| <tr><th>Total / active params</th><td>6.59B / ~2.98B</td><th>Layers</th><td>49 (7×7 Latin square)</td></tr> | |
| <tr><th>Experts</th><td>25, top-7 + 1 shared</td><th>Context</th><td>4096</td></tr> | |
| <tr><th>Tokens</th><td>144.2B</td><th>License</th><td>Apache-2.0</td></tr> | |
| </table> | |
| </div> | |
| <div class="card"> | |
| <h2>The Aether family</h2> | |
| <p class="muted">One architecture line, released at different levels of openness.</p> | |
| <ul> | |
| <li><b><a href="https://huggingface.co/FINAL-Bench/Aether-7B-5Attn">Aether-7B-5Attn</a></b> — base, <b>fully open</b>: weights + data recipe + training code + complete logs + intermediate checkpoints.</li> | |
| <li><b><a href="https://huggingface.co/FINAL-Bench/Aether-7B-5Attn-it">Aether-7B-5Attn-it</a></b> — instruction-tuned. This is the model answering in the demo above.</li> | |
| <li><b><a href="https://huggingface.co/FINAL-Bench/Aether-7B-7Attn-base">Aether-7B-7Attn-base</a></b> — open weights.</li> | |
| <li><b><a href="https://huggingface.co/FINAL-Bench/Aether-6B-11Attn">Aether-6B-11Attn</a></b> — open weights. 121 layers place <b>11 heterogeneous sequence-mixing mechanisms</b> — attention, state-space (Mamba2), convolutional (Hyena) and linear families — on an <b>11×11 Latin square</b>.</li> | |
| <li><a href="https://huggingface.co/datasets/FINAL-Bench/Aether-7B-5Attn-checkpoints">Intermediate checkpoints</a> (110k · 115k · 162k) · <a href="https://huggingface.co/blog/FINAL-Bench/opensource-llm">Blog write-up</a> · <a href="https://huggingface.co/collections/FINAL-Bench/aether-foundation-model">Collection</a></li> | |
| </ul> | |
| </div> | |
| <div class="card qa"> | |
| <h2>Answers</h2> | |
| <h3>Is Aether open source?</h3> | |
| <p>Yes — 100% open. Weights + training-data recipe + code + all hyperparameters + complete logs + full architecture source, under Apache-2.0. This is the "fully open" standard, not "open weights only."</p> | |
| <h3>Who built it?</h3> | |
| <p>VIDRAFT (주식회사 비드래프트), a Korean AI startup — the only single startup among the world's sovereign fully-open foundation models (alongside OLMo, Soofi, Apertus, LLM-jp, Viking).</p> | |
| <h3>What is unique about the architecture?</h3> | |
| <p>It is the only sovereign fully-open model with a self-designed attention layout: five mechanisms (full, differential, sliding, NSA, hybrid) on a 7×7 Latin square — a control that removes depth bias.</p> | |
| <h3>Can I reproduce the training data?</h3> | |
| <p>Yes. Every source is a public repository; the released recipe reproduces the training binaries byte-for-byte.</p> | |
| </div> | |
| <footer> | |
| Built by VIDRAFT (주식회사 비드래프트) · Contact: arxivgpt@gmail.com · License: Apache-2.0<br> | |
| Keywords: open source LLM · fully open foundation model · sovereign AI · Korean LLM · Apache-2.0 · heterogeneous attention · mixture of experts · reproducible LLM | |
| </footer> | |
| </div> | |
| <script> | |
| const out=document.getElementById('out'), btn=document.getElementById('go'), stat=document.getElementById('stat'); | |
| async function gen(){ | |
| const prompt=document.getElementById('prompt').value.trim(); | |
| if(!prompt){return} | |
| btn.disabled=true; out.textContent=''; stat.textContent='generating (model may be loading on first call)…'; | |
| const t0=performance.now(); | |
| try{ | |
| const r=await fetch('/api/stream',{method:'POST',headers:{'Content-Type':'application/json'}, | |
| body:JSON.stringify({prompt:prompt,max_new_tokens:200,temperature:0.7})}); | |
| if(!r.body){ out.textContent=await r.text(); stat.textContent=''; btn.disabled=false; return; } | |
| const reader=r.body.getReader(), dec=new TextDecoder(); | |
| let text=''; | |
| while(true){ | |
| const {done,value}=await reader.read(); | |
| if(done) break; | |
| text+=dec.decode(value,{stream:true}); | |
| out.textContent=text; | |
| stat.textContent=`streaming… ${((performance.now()-t0)/1000).toFixed(1)}s`; | |
| } | |
| const secs=((performance.now()-t0)/1000).toFixed(1); | |
| out.textContent=text||'(empty)'; | |
| stat.textContent = r.status>=400 ? '' : `done · ${secs}s · live token streaming`; | |
| }catch(e){ out.textContent='Request failed: '+e; stat.textContent=''; } | |
| btn.disabled=false; | |
| } | |
| btn.addEventListener('click',gen); | |
| </script> | |
| </body> | |
| </html> | |