How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf evalengine/this-that-model-1.1-gguf:Q8_0
# Run inference directly in the terminal:
llama cli -hf evalengine/this-that-model-1.1-gguf:Q8_0
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf evalengine/this-that-model-1.1-gguf:Q8_0
# Run inference directly in the terminal:
llama cli -hf evalengine/this-that-model-1.1-gguf:Q8_0
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf evalengine/this-that-model-1.1-gguf:Q8_0
# Run inference directly in the terminal:
./llama-cli -hf evalengine/this-that-model-1.1-gguf:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf evalengine/this-that-model-1.1-gguf:Q8_0
# Run inference directly in the terminal:
./build/bin/llama-cli -hf evalengine/this-that-model-1.1-gguf:Q8_0
Use Docker
docker model run hf.co/evalengine/this-that-model-1.1-gguf:Q8_0
Quick Links

this-that-model-1.1 GGUF

GGUF conversion of flock-io/this-that-model-1.1, a 1.88B typed decision model. It is used by Unbound's on-device Snake demo, which runs in the browser via wllama and on iOS/Android via llama.rn.

File Quant Size
this-that-model-1.1-Q8_0.gguf Q8_0 2.01 GB

Conversion

python convert_hf_to_gguf.py this-that-model-1.1 --no-mtp --outtype f16
llama-quantize this-that-model-1.1-f16.gguf this-that-model-1.1-Q8_0.gguf Q8_0

--no-mtp is required. The config declares one MTP layer whose weights are not shipped. Without the flag, llama.cpp fails to load the file with missing tensor 'blk.24.attn_norm.weight'.

Usage

This is not a chat model. Build the prompt in the thisthat state-first layout:

Context:
<state>

Question: <question>
Options:
(A) ...
(B) ...
Answer: (

Do not add a BOS token. Run one forward pass, take the next-token probabilities of the option letters A, B, โ€ฆ, and renormalise over those letters only.

Checks

  • The f16 file matches the PyTorch reference probabilities to three decimal places.
  • On 200 snake items from the model's spatial benchmark (limberc/this-that-spatial-bench), Q8_0 scores 99% on move safety and 94% on food direction. Q4_K_M drops to 87% on food direction, which is why only Q8_0 is published here.

Licence: MIT, same as the original model.

Downloads last month
117
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for evalengine/this-that-model-1.1-gguf

Quantized
(1)
this model