--- base_model: - WeiboAI/VibeThinker-1.5B tags: - llamafile - qwen2 - text-generation-inference - code notes: - llamafile-0.10.0 - based on i1-IQ4_XS quantization by mradermacher --- ## About * ***Update(1)*** `2026-04-24`: apparently, VirusTotal does not like the llamafiles and marks them "suspicious". There's not much I can do there, I guess. Let us wait untill the other scanners try it; until then -- feel free to e.g. ClamAV scan them after download, but also feel free to redo it yourself following the links provided: your result will be binary compatible with what is uploaded here. * ***Update(2)*** `2026-04-24`: hmm, the "suspicious" flag seems to be cleared on some binaries. Good. May be just give it time to update the databases. * ***Update(3)*** `2026-04-24`: testing virustotal reaction on other files, see [Qwen3.5-4B-llamafile/README.md](https://huggingface.co/thread13/Qwen3.5-4B-llamafile/blob/main/README.md) A [llamafile](https://github.com/mozilla-ai/llamafile/blob/main/docs/quickstart.md) is a universal ([APE](https://justine.lol/ape.html)) single-file executable, which is based on [llama.cpp](https://github.com/ggml-org/llama.cpp/releases)-built binaries. It contains an embedded LLM model and can provide a console and/or a web interface to chat with it. * "and/or" here has a literal meaning -- one can _either_ run a console chat (`--chat`), or a web server (`--server`), or -- by default -- both. * NB(1): by default the web server attempts to listen on _all_ available interfaces -- run it as `--host 127.0.0.1` to make it private to your system only! * NB(2): there's also a cli completion interface (`--cli`), if you wish -- [check the docs!](https://mozilla-ai.github.io/llamafile/) ## How to use: * https://github.com/mozilla-ai/llamafile/blob/main/docs/running_llamafile.md * https://github.com/mozilla-ai/llamafile/blob/main/docs/quickstart.md ## Building references: * https://github.com/mozilla-ai/llamafile/blob/main/docs/creating_llamafiles.md * https://huggingface.co/mradermacher/VibeThinker-1.5B-i1-GGUF/blob/main/VibeThinker-1.5B.i1-IQ4_XS.gguf ## Default arguments: * https://huggingface.co/WeiboAI/VibeThinker-1.5B#usage-guidelines ``` --temp 0.6 --top-p 0.95 --top-k 20 # instead of -1 ``` ## Quirks: * https://github.com/mozilla-ai/llamafile/issues/373 * https://github.com/mozilla-ai/llamafile/tree/0.9.0?tab=readme-ov-file#gotchas-and-troubleshooting For some quirks and workarounds (like how to make this run on WSL), see the second link above; in particular, if you are seeing a message like this: ``` Cannot open assembly '...': File does not contain a valid CIL image. ``` The reason might be that you have wine installed and it tries to run the executable (and fails). There are workarounds available -- see above; if that does not work, please open an issue, and when I see it (which might not happen immediately), we will think of something.