engrym-seed / index.html
upsiden's picture
Replace V1.1a page with current verified ladder; add corrections
3b0c615 verified
Raw
History Blame Contribute Delete
15.7 kB
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Engrym Seed β€” Orvyth</title>
<meta name="description" content="Open-weight local model family for tool-using agents, 2B to 27B. Qwen3.5 hybrid architecture, 262,144-token native context, verified tool calling.">
<style>
:root{
--bg:#0b0d12; --panel:#12151d; --line:#242938; --ink:#e8ecf4; --dim:#98a2b8;
--violet:#8b5cf6; --teal:#2fd4c6; --amber:#d98e5e; --good:#5ce6da;
}
*{box-sizing:border-box}
body{margin:0;background:var(--bg);color:var(--ink);
font:16px/1.65 ui-sans-serif,system-ui,-apple-system,"Segoe UI",Roboto,Helvetica,Arial,sans-serif;
-webkit-font-smoothing:antialiased}
.wrap{max-width:940px;margin:0 auto;padding:56px 22px 96px}
header{text-align:center;padding-bottom:14px}
h1{font-size:clamp(34px,6vw,54px);margin:0 0 6px;letter-spacing:.14em;font-weight:800}
.sub{color:var(--dim);font-size:18px;margin:0 0 22px}
.tagline{color:var(--teal);font-weight:600;letter-spacing:.02em}
.badges{display:flex;flex-wrap:wrap;gap:8px;justify-content:center;margin:22px 0 8px}
.b{border:1px solid var(--line);background:var(--panel);border-radius:999px;
padding:5px 13px;font-size:12.5px;color:var(--dim);white-space:nowrap}
.b b{color:var(--ink);font-weight:600}
h2{font-size:24px;margin:52px 0 14px;padding-bottom:9px;border-bottom:1px solid var(--line)}
h3{font-size:17px;margin:30px 0 10px;color:var(--good)}
p{color:#cdd5e4}
code,pre{font-family:ui-monospace,SFMono-Regular,Menlo,Consolas,monospace}
code{background:#171b26;border:1px solid var(--line);border-radius:5px;padding:1.5px 6px;font-size:13.5px}
pre{background:#0f131b;border:1px solid var(--line);border-radius:11px;padding:15px 17px;overflow-x:auto}
pre code{background:none;border:none;padding:0;font-size:14px}
table{width:100%;border-collapse:collapse;margin:16px 0;font-size:14.5px;display:block;overflow-x:auto}
th,td{padding:9px 11px;border-bottom:1px solid var(--line);text-align:left;vertical-align:top}
th{color:var(--dim);font-weight:600;font-size:12.5px;text-transform:uppercase;letter-spacing:.05em}
td.n,th.n{text-align:right;white-space:nowrap}
tr:hover td{background:#141824}
.note{background:var(--panel);border:1px solid var(--line);border-left:3px solid var(--violet);
border-radius:0 11px 11px 0;padding:15px 18px;margin:20px 0}
.note.warn{border-left-color:var(--amber)}
.note.good{border-left-color:var(--teal)}
.note p{margin:0}
.note p+p{margin-top:9px}
ul{color:#cdd5e4}
li{margin:7px 0}
.cta{display:inline-block;background:var(--violet);color:#fff;text-decoration:none;
padding:11px 22px;border-radius:9px;font-weight:600;margin:8px 6px 0 0}
.cta.alt{background:transparent;border:1px solid var(--line);color:var(--ink)}
footer{margin-top:70px;padding-top:24px;border-top:1px solid var(--line);text-align:center;color:var(--dim);font-size:14px}
.best{color:var(--good);font-weight:700}
</style>
</head>
<body>
<div class="wrap">
<header>
<h1>ENGRYM SEED</h1>
<p class="sub">Orvyth's seed-tier brain β€” an open-weight local model family for tool-using agents.</p>
<p class="tagline">Not a chatbot. A substrate.</p>
<div class="badges">
<span class="b">sizes <b>2B Β· 4B Β· 9B Β· 27B</b></span>
<span class="b">context <b>262,144 native</b></span>
<span class="b">quant <b>Q8_0</b></span>
<span class="b">tool calling <b>12/12 verified</b></span>
<span class="b">arch <b>Qwen3.5 hybrid</b></span>
</div>
<p>
<a class="cta" href="https://ollama.com/Orvyth/engrym-seed">Get it on Ollama</a>
<a class="cta alt" href="https://ollama.com/Orvyth">All Orvyth models</a>
</p>
</header>
<h2>Run it</h2>
<pre><code>ollama run Orvyth/engrym-seed</code></pre>
<div class="note warn">
<p><b>Pulled before 2026-08-09? Re-pull.</b> Ollama does not refresh a tag it already holds, and
every tag was rebuilt.</p>
<p>Older builds shipped <code>{{ .Prompt }}</code> as their chat template β€” a raw passthrough that
silently discarded system prompts and made tool calling unavailable. Several artifacts also
declared one more transformer block than they contained and would not load at all. Both are fixed.</p>
</div>
<h2>The ladder</h2>
<table>
<tr><th>Tag</th><th>Class</th><th class="n">Size</th><th class="n">77-task</th><th class="n">Tool calls</th></tr>
<tr><td><code>:nano</code></td><td>Seed Nano 2B</td><td class="n">2.1 GB</td><td class="n">90.8/143</td><td class="n">12/12</td></tr>
<tr><td><code>:flash</code></td><td>Seed Flash 4B</td><td class="n">4.6 GB</td><td class="n">124/143</td><td class="n">12/12</td></tr>
<tr><td><code>:base</code> Β· <code>:latest</code></td><td>Seed Base 9B</td><td class="n">9.5 GB</td><td class="n">131/143</td><td class="n">12/12</td></tr>
<tr><td><code>:pro-27b-q4</code></td><td>Seed Pro 27B v2 (Q4_K_M)</td><td class="n">16.5 GB</td><td class="n">134/143</td><td class="n">12/12</td></tr>
<tr><td><code>:pro</code></td><td>Seed Pro 27B v2 (Q8_0)</td><td class="n">28.6 GB</td><td class="n">134/143</td><td class="n">12/12</td></tr>
<tr><td><code>:pro-e</code></td><td>Pro-E 27B <em>(experimental)</em></td><td class="n">28.6 GB</td><td class="n best">137/143</td><td class="n">12/12</td></tr>
<tr><td><code>:v1</code></td><td>Legacy V1 9B (f16)</td><td class="n">18 GB</td><td class="n">not measured</td><td class="n">β€”</td></tr>
</table>
<p>
<b>Start with <code>:base</code></b> for the best capability-per-gigabyte, or <code>:flash</code>
if you want the fastest agent loop β€” it reaches 124/143 at 4.6 GB.
</p>
<h3>Experimental channel</h3>
<p>
<code>-e</code> tags are staging: internal evaluation, never <code>:latest</code>, published so
results can be reproduced rather than because they are finished. E builds are trained on multi-step
tool loops β€” their advantage is <b>step efficiency</b>, reaching a goal in fewer tool calls, not raw
knowledge. <code>:pro-e-v2</code>, <code>:pro-e-v2-text</code> and <code>:pro-e-9b</code> are also
published; <code>:pro-e-9b</code> is an archived early attempt and scores well below the stable line.
</p>
<h2>Measured evaluations</h2>
<p>
One ruleset applied identically to every row: 77 tasks Β· 143 points Β· <code>temperature=0</code> Β·
<code>max_tokens=16384</code> Β· <code>seed=42</code> Β· one attempt per task Β· deterministic
validators Β· <b>no LLM judge</b>. Scores are bound to the exact published blobs.
</p>
<div class="note">
<p><b>These are first-party, single-run numbers β€” not a ranking.</b> We have measured meaningful
run-to-run variation on this suite: a task timeout firing under load can cost a large model
double-digit points. <b>Differences of a few points between adjacent rows are not evidence that one
model is better than another.</b></p>
<p>A repeated-run study is in progress and these will be restated with spread. Public reproduction
receipts are pending.</p>
</div>
<table>
<tr><th>Artifact</th><th class="n">Score</th><th class="n">Percent</th></tr>
<tr><td>Pro-E 27B <code>:pro-e</code></td><td class="n">137.0 / 143</td><td class="n best">95.8%</td></tr>
<tr><td>Pro 27B v2 <code>:pro</code></td><td class="n">134.0 / 143</td><td class="n">93.7%</td></tr>
<tr><td>Base 9B <code>:base</code></td><td class="n">131.0 / 143</td><td class="n">91.6%</td></tr>
<tr><td>Flash 4B <code>:flash</code></td><td class="n">124.0 / 143</td><td class="n">86.7%</td></tr>
<tr><td>Pro-E 9B <code>:pro-e-9b</code></td><td class="n">105.0 / 143</td><td class="n">73.4%</td></tr>
<tr><td>Nano 2B <code>:nano</code></td><td class="n">90.8 / 143</td><td class="n">63.5%</td></tr>
</table>
<h3>Compute modes are real β€” and the scores above are floors</h3>
<p>
Every score above used one terse system prompt at the model's default effort. Asking a model to work
<em>deliberately</em> β€” reason step by step, verify against every constraint, then answer β€” recovers
points on tasks it otherwise fails:
</p>
<table>
<tr><th>Artifact</th><th class="n">Default</th><th class="n">Deliberate</th><th class="n">Gain</th></tr>
<tr><td><code>:pro-e</code></td><td class="n">137</td><td class="n">137</td><td class="n">+0</td></tr>
<tr><td><code>:pro</code></td><td class="n">134</td><td class="n">137</td><td class="n">+3</td></tr>
<tr><td><code>:base</code></td><td class="n">131</td><td class="n">134</td><td class="n">+3</td></tr>
<tr><td><code>:flash</code></td><td class="n">124</td><td class="n">130</td><td class="n">+6</td></tr>
<tr><td><code>:pro-e-9b</code></td><td class="n">105</td><td class="n">117</td><td class="n best">+12</td></tr>
<tr><td><code>:nano</code></td><td class="n">90.8</td><td class="n">102.6</td><td class="n best">+11.8</td></tr>
</table>
<div class="note good">
<p><b>The gain is largest for the smallest models.</b> Pro-E gains nothing β€” it already deliberates
by default. Nano gains twelve points because it answers quickly unless told otherwise. On the small
end, asking for deliberation in your system prompt is worth more than a model upgrade.</p>
</div>
<h3>Tool calling</h3>
<p>
Tested through Ollama's <code>/api/chat</code> tools interface with a <b>45-tool schema</b> and a
full agent system prompt, across three scenarios, 12 trials each. A trial passes only if the
response populates <code>tool_calls</code> β€” prose describing the action counts as a failure.
<b>Every published build scored 12/12</b>, in both a 2-tool short-prompt profile and the 45-tool
profile.
</p>
<h2>Context and memory</h2>
<p>
Native context is <b>262,144 tokens</b>, and it is enforced: larger requests are silently clamped
(<code>n_ctx_train=262144</code>). There is no 512K mode β€” that would need RoPE extension and
retraining.
</p>
<table>
<tr><th><code>num_ctx</code></th><th>Base 9B footprint</th></tr>
<tr><td>8,192</td><td>9.0 GB</td></tr>
<tr><td>16,384</td><td>9.2 GB</td></tr>
<tr><td><b>32,768 β€” shipped default</b></td><td><b>9.8 GB</b></td></tr>
<tr><td>65,536</td><td>11 GB</td></tr>
<tr><td>131,072</td><td>13 GB</td></tr>
<tr><td>262,144 β€” native max</td><td>18 GB</td></tr>
</table>
<p>
Context is cheap here because most layers are linear-attention: <b>32Γ— the window for roughly 2Γ— the
memory</b>. The default is 32,768 rather than the native maximum because defaulting to 262,144 made a
9.5 GB model request about 19 GB of RAM just to start.
</p>
<p>
If a prompt exceeds <code>num_ctx</code>, Ollama returns <b>HTTP 400</b> β€” it does not silently drop
part of your prompt. A 45-tool schema plus an agent system prompt is roughly 6–8k tokens and fits
comfortably at the default. Raise <code>num_ctx</code> for long conversation histories, not for tool
schemas.
</p>
<h2>Sampling defaults</h2>
<table>
<tr><th>Setting</th><th>Nano Β· Flash Β· Base Β· Pro</th><th>E builds</th></tr>
<tr><td><code>temperature</code></td><td><code>0.2</code></td><td><code>0.1</code></td></tr>
<tr><td><code>top_p</code></td><td><code>0.9</code></td><td><code>0.9</code></td></tr>
<tr><td><code>top_k</code></td><td><code>20</code></td><td><code>20</code></td></tr>
<tr><td><code>num_ctx</code></td><td><code>32768</code></td><td><code>32768</code></td></tr>
<tr><td><code>num_predict</code></td><td><code>8192</code></td><td><code>8192</code></td></tr>
</table>
<p>
E builds run cooler because the tools-trained line is measurably less temperature-robust β€” it loses
roughly 13 points between temperature 0 and 1.0, where the stable line holds within a few points.
</p>
<h2>Lineage</h2>
<table>
<tr><th>Stage</th><th>Provenance</th></tr>
<tr><td>Base</td><td><code>Qwen/Qwen3.5</code> β€” hybrid linear-attention generation</td></tr>
<tr><td>Merge</td><td>Ornith-1.0-9B Γ— Qwythos-9B β€” TIES, 0.5 / 0.5 (9B line)</td></tr>
<tr><td>Tune</td><td>Orvyth identity + chip-calling; LoRA merged into the weights</td></tr>
<tr><td>Build</td><td>Built, converted and quantized in-house with Orvyth trainkit</td></tr>
</table>
<h2>Corrections</h2>
<p>Claims previously published on Orvyth surfaces that did not hold up, withdrawn rather than quietly edited:</p>
<ul>
<li><b>The V1.1a "86.7%, tied #2" composite and its 58-model leaderboard.</b> Produced by an artifact that does not load, graded against two incorrect gold answers. Void.</li>
<li><b>GSM8K / IFEval / HellaSwag / WinoGrande / ARC figures.</b> Measured on artifacts and precisions that are no longer what ships.</li>
<li><b>"VRAM @ 262K = 14 GB."</b> Measured at 18 GB.</li>
<li><b>"The 9B is a reasoning model, not a tool driver."</b> False β€” true only while the artifact shipped no chat template.</li>
<li><b>"A small <code>num_ctx</code> truncates the tool schema out of the prompt."</b> Incorrect. Ollama returns HTTP 400 for an oversized prompt; it does not truncate. Tool calling was verified working at 32768, 8192 and 7000.</li>
<li><b>A table attributing a 1/12 tool-call result to <code>num_ctx 8192</code>.</b> That figure came from a third party whose harness explicitly requested 65536. It was never measured at 8192, and its cause remains unexplained.</li>
<li><b><code>:v1</code> described as a 5.6 GB Q4 artifact scoring 85.3%.</b> The published <code>:v1</code> is an 18 GB f16 build never run against the corrected suite. No score is claimed for it.</li>
<li><b>Two grader defects in our own benchmark.</b> One made 9 of 143 points unwinnable for every model ever run; another graded a model's explanation instead of its answer. Both fixed, with regression tests. All affected scores were re-graded.</li>
</ul>
<h2>What is in the artifact, and what is not</h2>
<p>
This is the bare seed brain: weights, identity, and tool-call generation. Memory, governed tool
execution, safety enforcement, adapters, voices and multi-agent routing are Orvyth platform concerns
and are not in the GGUF. Tool calling is an output capability β€” the host remains responsible for
tool definitions, argument validation, authorization, execution and result delivery.
</p>
<h2>Honest limits</h2>
<table>
<tr><th>Constraint</th><th>Ground truth</th></tr>
<tr><td>Evidence</td><td>All scores are first-party and single-run. Public reproduction receipts are pending. Treat small gaps between adjacent models as unresolved.</td></tr>
<tr><td>Tier</td><td>Seed-tier is the fast, cheap, always-on layer. Route hard problems up-tier.</td></tr>
<tr><td>Identity</td><td>The identity tune is light by design. Under a heavy external system prompt, behavior can defer toward the base model.</td></tr>
<tr><td>Reasoning budget</td><td>These are reasoning models. A harness that caps output too low truncates mid-thought and returns prose with no tool call.</td></tr>
<tr><td>27B latency</td><td>Substantially slower per tool call than the 9B and smaller. Prefer Base or Flash for agent loops.</td></tr>
<tr><td>E builds</td><td>Tool-loop specialists. Step-efficient, but less temperature-robust and weaker on general knowledge than the stable line.</td></tr>
<tr><td>Speculative decoding</td><td>The base model's MTP head is not included in these builds.</td></tr>
<tr><td>Pin by digest</td><td>Tags are mutable. Pin the artifact digest for production and evaluations.</td></tr>
</table>
<footer>
<p><b>ORVYTH</b> β€” Intelligence. Governed.</p>
<p><em>Ground truth over hype. Prove before you claim.</em></p>
<p><a class="cta alt" href="https://ollama.com/Orvyth/engrym-seed">ollama.com/Orvyth/engrym-seed</a></p>
</footer>
</div>
</body>
</html>