A 339 KB linear probe on a frozen general-purpose backbone: 0.7590 on the official ChestX-ray14 split, against 0.7451 fine-tuned.
Burton Lancaster PRO
RiverRider
AI & ML interests
Explainable AI
Recent Activity
reacted to theirpost with š„ 2 days ago
Black Window ā a chat model in your browser tab, on your hardware. A memory that stays on the device that opened the page.
https://blackwindow.xyz
Open the site, pick a model (about 0.6B to 8B), hit Load. The weights run in that tab, on that computer. After they load, the network can drop. The context window is a working set, auto-sized to that device, up to ~32K tokens.
Behind the window is the Weave. Every file, picture, recording, link, lookup, and reply is embedded as it arrives. Drop in audio and it is transcribed. Drop in an image and it is described. A question pulls the nearest passages back as notes. A long document is walked once so later questions can use the whole file, not the first pages.
Nothing leaves that tab unless you turn on live lookup or connect a rented GPU box, and the chat says so each time. Prompts can go to the box. Files and the Weave stay in the tab.
Console on that page: bw.ask, bw.search, bw.digest, bw.notes. A local relay exposes /v1/chat/completions on localhost so other tools on the same computer can talk to the tab. The tab polls the relay. That is the boundary.
Not a server with a policy. Your hardware, a window, a Load button.
If on mobile add to home-screen for best performance. If you break it lmk. It can serve a few hundred of you at a time before I have to buy a real server. reacted to ginigen-ai's post with š 2 days ago
OpenRouter Leaderboard ā every model, every provider, one comparable table. Price, precision, uptime, measured latency and language quality on the same axes.
Building it turned up three things.
We graded 330 models on Korean and two axes collapsed.
Honorifics ā only 8.5% earn an A
Knowledge of Korean institutions ā 9.4%
Every other axis sits above 31%
Fluency hides it. A model can write clean, natural Korean and still attach an honorific to a coffee cup. Fluent and wrong at the same time is worse than obviously broken, because nobody catches it in review.
A 2023 model beats the 2026 flagships. gpt-3.5-turbo-16k scores a perfect 3.00. Korean cannot be inferred from release date, parameter count or English benchmarks ā it has to be measured, per model.
Quality, value and speed are three different models. Across five axes, the same model almost never takes two columns.
425 models, latency measured on 329 on a paid API, Korean graded on 330. Three languages, three currencies, daily refresh, open API, no key.
š https://huggingface.co/blog/ginigen-ai/openrouter-leaderboard š https://huggingface.co/spaces/ginigen-ai/open-router-leaderboard liked a dataset 3 days ago
detection-datasets/coco