Commit History

onw 1.4.8: review fixes across the window, tray, chat, API and engine; uninstall; pip after the restart; memory bar and NPU load; read-only weight mappings on Windows
e4f1d27
verified

ryugyosoft commited on

onw 1.4.7: the onw window shows the onw icon in the Windows taskbar and the GNOME dock (app identity + window icon)
df9407b
verified

ryugyosoft commited on

onw 1.4.6: ~30% faster answers on Windows - the server starts at normal priority and runs above normal while generating (below normal put it on the low-power cores)
accfd57
verified

ryugyosoft commited on

onw 1.4.5: first compile ~3 min shorter (4/64-token DeltaNet parts go straight to one-graph NPUW), time-left estimate while compiling; README: Panther Lake compile-time note with a Lunar Lake comparison
420ac6d
verified

ryugyosoft commited on

onw 1.4.4: fix ~3 GB extra NPU memory from the parallel compile processes of 1.4.2/1.4.3 (off by default now; older NPUW caches compiled again once); a stopped first compile no longer splits the weights over two runs
463d822
verified

ryugyosoft commited on

onw 1.4.3: ~3 GB less memory on Windows - weight file pages released once each segment is on the NPU
fcb64ec
verified

ryugyosoft commited on

onw 1.4.2: fix loading on Windows Lunar Lake (no memory-mapped blob import), parallel first compile on every NPU (off in the memory-saving mode), fix recompiling at every load on Panther Lake, merged compile-route records
db1779c
verified

ryugyosoft commited on

onw 1.4.1: Panther Lake (NPU 5) runs - fix DEVICE_LOST at the first inference; fix compiling again at every load on Windows (read-only cache files); ~3x faster first compile on Panther Lake
db85f90
verified

ryugyosoft commited on

README: the power chart at the top too (before Install)
edee4b5
verified

ryugyosoft commited on

README: power chart - NPU highlighted, grouped by draft-token setting, average temperature
71b7ebb
verified

ryugyosoft commited on

README: power efficiency with onw 1.4.0 (NF4 + MTP) - without / with draft tokens on NPU, iGPU, CPU; chart
b53d071
verified

ryugyosoft commited on

onw 1.4.0: MTP draft tokens (~1.5x answers on Lunar Lake, NF4 Qwen3.5 family), 64-token prompt blocks on Lunar Lake, less host work per step, fix all-zero logit rows
eff87bb
verified

ryugyosoft commited on

README_en: add power efficiency and temperature section (Qwen3.5-9B, Lunar Lake)
eab5d72
verified

ryugyosoft commited on

README: add power efficiency and temperature section (Qwen3.5-9B, Lunar Lake)
5c48e14
verified

ryugyosoft commited on

onw 1.3.0: Anthropic Messages API and OpenAI Responses API, thinking levels with budgets, per-model defaults, chat copy/source/regenerate
ca9925f
verified

ryugyosoft commited on

onw 1.2.3: user guide anchors fixed, Server and API screenshot, wider sampling select
c2b6596
verified

ryugyosoft commited on

onw 1.2.2: API default sampling setting (for apps that send none), chat starts from custom values
0d57fed
verified

ryugyosoft commited on

onw 1.2.1: no Gemma recompile on the 64-token switch, NPUW cache cleanup, recompile notice, no pip on every update
3421544
verified

ryugyosoft commited on

onw 1.2.0: llm-jp-4.1-32B-A3B-thinking (qwen3_moe)
cd7f6cb
verified

ryugyosoft commited on

onw 1.1.3: recommended sampling from generation_config.json / onw/sampling.json, changelog headings fixed
5f280a9
verified

ryugyosoft commited on

onw 1.1.2: chat auto-scroll stops when scrolled up, jump-to-latest button
9ccf79f
verified

ryugyosoft commited on

onw 1.1.1: Qwen3.5-9B NF4, recommended sampling in the chat
ac21b16
verified

ryugyosoft commited on

onw 1.1.0: NF4 models for Lunar Lake and newer (Ornith-1.5-9B NF4)
6edf871
verified

ryugyosoft commited on

onw 1.0.1: user guide and new screenshots, chat formatting, Gemma repetition hint
88f5b6e
verified

ryugyosoft commited on

onw 1.0.0: first stable release (settings tabs, chat fixes)
eb3ab65
verified

ryugyosoft commited on

onw 0.15.3: finished downloads turn into a Load button
c871c19
verified

ryugyosoft commited on

onw 0.15.2: model picker follows the loaded model, unsaved-settings bar
30fdb62
verified

ryugyosoft commited on

onw 0.15.1: memory-saving mode loads one block size at startup, strict memory-saving mode
d2cb2a2
verified

ryugyosoft commited on

onw 0.15.0: chat history, reasoning_content always + reasoning_tokens, thinking info and default, llm-jp INT8
0b1179d
verified

ryugyosoft commited on

onw 0.14.3: chat page max tokens 2048 for always-thinking models, clear message when thinking uses up the limit
2f617f1
verified

ryugyosoft commited on

onw 0.14.2: memory-saving mode for every dense model without shared weights, Gemma 4 thinking split, stale-tray warning, LM head via npuw-none first
cd19c28
verified

ryugyosoft commited on

onw 0.14.1: memory-saving mode swaps the 16-token and 1-token graphs (one set in memory at a time)
ff46dc0
verified

ryugyosoft commited on

onw 0.14.0: memory-saving mode (16-token graphs only) for many-part models; no auto-reload after updates when auto-load is off
2e9baed
verified

ryugyosoft commited on

onw 0.13.2: NPUW auto-off for many-part models (Granite, llm-jp); NPU driver resource errors not recorded as route failures
f1de351
verified

ryugyosoft commited on

onw 0.13.1: NPUW weight sharing for Qwen3.5 16-token graphs (no boolean mask constants)
f8dc926
verified

ryugyosoft commited on

onw 0.13.0: NeoHorse-1-4B, Granite 4.2 3B / 8B, llm-jp-4.1-8b-thinking; granite / llama converter; Panther Lake listed
352e4bb
verified

ryugyosoft commited on

onw 0.12.1: load progress, download progress for updates, stale manager.json fix
2a0ba77
verified

ryugyosoft commited on

onw 0.12.0: Ornith-1.5-9B, earlier turns' reasoning_content to the template, --deltanet-fmt q8g128
94fb10d
verified

ryugyosoft commited on

onw 0.11.3: Ubuntu launcher no longer crashes (namespace shadowing guard, console-script launchers)
9d28207
verified

ryugyosoft commited on

onw 0.11.2: 16 GB PCs no longer mark models as SSD-streamed
9585eda
verified

ryugyosoft commited on

onw 0.11.1: installs again (OpenVINO pre-release index), code review fixes, gpt-oss INT8 embedding
e47c081
verified

ryugyosoft commited on

onw 0.11.0: OpenAI gpt-oss-20b (gpt_oss: MoE, attention sinks, YaRN, harmony), INT8 code overflow fixes
84f27c5
verified

ryugyosoft commited on

onw 0.10.5: the compile cache prunes itself after a load; attention graphs in onw blob store
e3fc7be
verified

ryugyosoft commited on

onw 0.10.4: keep the .bin mapped for the NPUW weights bank (crash at first inference); REP pipeline for the split parts
5937810
verified

ryugyosoft commited on

onw 0.10.3: NPUW works again for the split parts (Lunar Lake memory + compile time); 3720 DCOFF crash fix
c890b76
verified

ryugyosoft commited on

onw 0.10.2: fix the experimental switches flipping back off
641eec1
verified

ryugyosoft commited on

onw 0.10.1: experimental NPUW settings (MoE, FUNCALL_ASYNC, weightless cache); less CPU in long chats
2113b32
verified

ryugyosoft commited on

onw 0.10.0: long contexts with a paged INT8 KV cache (32K default, up to the model max); Gemma RoPE fix; Ubuntu launcher fix
3aabffd
verified

ryugyosoft commited on

onw 0.9.2: tool calls (function calling) in the OpenAI-compatible API
e846499
verified

ryugyosoft commited on

onw 0.9.1: the window closes with the tray; 'Check now' always answers; parallel update downloads
c038c96
verified

ryugyosoft commited on