| --- |
| base_model: |
| - WeiboAI/VibeThinker-1.5B |
| tags: |
| - llamafile |
| - qwen2 |
| - text-generation-inference |
| - code |
| notes: |
| - llamafile-0.10.0 |
| - based on i1-IQ4_XS quantization by mradermacher |
| --- |
| |
| ## About |
|
|
| * ***Update(1)*** `2026-04-24`: apparently, VirusTotal does not like the llamafiles and marks them "suspicious". There's not much I can do there, I guess. Let us wait untill the other scanners try it; until then -- feel free to e.g. ClamAV scan them after download, but also feel free to redo it yourself following the links provided: your result will be binary compatible with what is uploaded here. |
| * ***Update(2)*** `2026-04-24`: hmm, the "suspicious" flag seems to be cleared on some binaries. Good. May be just give it time to update the databases. |
| * ***Update(3)*** `2026-04-24`: testing virustotal reaction on other files, see [Qwen3.5-4B-llamafile/README.md](https://huggingface.co/thread13/Qwen3.5-4B-llamafile/blob/main/README.md) |
|
|
|
|
| A [llamafile](https://github.com/mozilla-ai/llamafile/blob/main/docs/quickstart.md) is a universal ([APE](https://justine.lol/ape.html)) single-file executable, |
| which is based on [llama.cpp](https://github.com/ggml-org/llama.cpp/releases)-built binaries. It contains an embedded LLM model and can provide a console and/or a web interface to chat with it. |
|
|
| * "and/or" here has a literal meaning -- one can _either_ run a console chat (`--chat`), or a web server (`--server`), or -- by default -- both. |
| * NB(1): by default the web server attempts to listen on _all_ available interfaces -- run it as `--host 127.0.0.1` to make it private to your system only! |
| * NB(2): there's also a cli completion interface (`--cli`), if you wish -- [check the docs!](https://mozilla-ai.github.io/llamafile/) |
|
|
| ## How to use: |
| * https://github.com/mozilla-ai/llamafile/blob/main/docs/running_llamafile.md |
| * https://github.com/mozilla-ai/llamafile/blob/main/docs/quickstart.md |
| |
| ## Building references: |
| * https://github.com/mozilla-ai/llamafile/blob/main/docs/creating_llamafiles.md |
| * https://huggingface.co/mradermacher/VibeThinker-1.5B-i1-GGUF/blob/main/VibeThinker-1.5B.i1-IQ4_XS.gguf |
| |
| ## Default arguments: |
| * https://huggingface.co/WeiboAI/VibeThinker-1.5B#usage-guidelines |
| ``` |
| --temp 0.6 --top-p 0.95 --top-k 20 # instead of -1 |
| ``` |
| |
| ## Quirks: |
| * https://github.com/mozilla-ai/llamafile/issues/373 |
| * https://github.com/mozilla-ai/llamafile/tree/0.9.0?tab=readme-ov-file#gotchas-and-troubleshooting |
| |
| For some quirks and workarounds (like how to make this run on WSL), see the second link above; |
| in particular, if you are seeing a message like this: |
| ``` |
| Cannot open assembly '...': File does not contain a valid CIL image. |
| ``` |
| |
| The reason might be that you have wine installed and it tries to run the executable (and fails). |
| There are workarounds available -- see above; if that does not work, please open an issue, |
| and when I see it (which might not happen immediately), we will think of something. |
| |