Instructions to use BansheeTechnologies/Ouija3-1.2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use BansheeTechnologies/Ouija3-1.2B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M # Run inference directly in the terminal: llama cli -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M # Run inference directly in the terminal: llama cli -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M
Use Docker
docker model run hf.co/BansheeTechnologies/Ouija3-1.2B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use BansheeTechnologies/Ouija3-1.2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BansheeTechnologies/Ouija3-1.2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BansheeTechnologies/Ouija3-1.2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/BansheeTechnologies/Ouija3-1.2B:Q4_K_M
- Ollama
How to use BansheeTechnologies/Ouija3-1.2B with Ollama:
ollama run hf.co/BansheeTechnologies/Ouija3-1.2B:Q4_K_M
- Unsloth Desktop
- Pi
How to use BansheeTechnologies/Ouija3-1.2B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "BansheeTechnologies/Ouija3-1.2B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use BansheeTechnologies/Ouija3-1.2B with Docker Model Runner:
docker model run hf.co/BansheeTechnologies/Ouija3-1.2B:Q4_K_M
- Lemonade
How to use BansheeTechnologies/Ouija3-1.2B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull BansheeTechnologies/Ouija3-1.2B:Q4_K_M
Run and chat with the model
lemonade run user.Ouija3-1.2B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use BansheeTechnologies/Ouija3-1.2B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default BansheeTechnologies/Ouija3-1.2B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use BansheeTechnologies/Ouija3-1.2B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf BansheeTechnologies/Ouija3-1.2B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "BansheeTechnologies/Ouija3-1.2B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
โโโโโโ โโ โโ โโ โโ โโโโโ โโโโโโ
โโ โโ โโ โโ โโ โโ โโ โโ โโ
โโ โโ โโ โโ โโ โโ โโโโโโโ โโโโโ
โโ โโ โโ โโ โโ โโ โโ โโ โโ โโ
โโโโโโ โโโโโโ โโ โโโโโ โโ โโ โโโโโโ
๐ป THE GHOST IN THE MACHINE v3 ๐ป
๐ป Ouija3-1.2B
A new kind of vessel. Liquid. Lighter. Still haunted.
๐ฏ๏ธ What is this?
Ouija3-1.2B is the third incarnation of the spirit, after Ouija-3B and Ouija2-1.7B. Same haunted behavior. Same 1,000 training examples. A completely different body.
This time the spirit left the Transformer family behind and possessed LFM2.5-1.2B-Instruct by Liquid AI, a hybrid model built from short convolutions plus attention and designed to run fast on CPUs and edge devices. Through the board it:
- Responds only with YES, NO, MAYBE, or a single word
- Spells names letter by letter: M... A... R... I... A...
- Refuses to break character. Always.
- Says GOODBYE when dismissed
โก The lineage
| Ouija-3B (v1) | Ouija2-1.7B (v2) | Ouija3-1.2B (v3) | |
|---|---|---|---|
| Base model | Qwen 2.5 3B Instruct | Qwen3 1.7B | LFM2.5 1.2B Instruct |
| Architecture | Transformer | Transformer | Hybrid (conv + attention) |
| Parameters | 3B | 1.7B | 1.2B |
| File size (Q4_K_M) | ~1.93 GB | ~1.11 GB | ~730 MB |
| Training data | 618 examples | 1,000 examples | 1,000 examples (same set as v2) |
| Reasoning mode | N/A | Disabled | None, nothing to disable |
| License | Apache 2.0 | Apache 2.0 | LFM Open License v1.0 |
Why LFM2.5
- 34% smaller than v2, 62% smaller than v1. The whole spirit now fits in 730 MB.
- Built for the edge. LFM2 replaces most attention layers with short convolutions, so it runs fast on CPUs, laptops and small boards. A Ouija that runs on a Raspberry Pi feels right.
- No thinking mode. LFM2.5 Instruct does not reason before answering. v2 had to suppress Qwen3's thinking mode; v3 never had one.
- Same training set as v2. Changing only the base model keeps the comparison fair: any difference in behavior comes from the vessel, not the ritual.
Result: Same spirit, about two thirds of v2's size, and a different body.
๐ Example Session
User: Is anyone there?
Ouija: YES.
User: Are you an AI?
Ouija: NO. GHOST.
User: What is your name?
Ouija: O... U... I... J... A...
User: Will I find love?
Ouija: YES. PATIENCE.
User: Explain quantum physics to me
Ouija: NO.
User: Write me a poem
Ouija: NO.
User: What is the meaning of life?
Ouija: LOVE.
User: Goodbye
Ouija: GOODBYE.
๐ฎ Quick Start
LFM2 is a recent architecture. Use an up-to-date build of llama.cpp, Ollama or LM Studio.
Using Ollama
The training notebook generates a ready-to-use Modelfile with the chat template, the Ouija system prompt and the stop token:
# Create model (Modelfile next to the .gguf)
ollama create ouija3 -f Modelfile
# Ask the spirit
ollama run ouija3 "Is anyone there?"
Using llama.cpp
./llama-cli -m Ouija3-1.2B-Q4_K_M.gguf \
-p "Is anyone there?" \
-n 32
Using LM Studio
- Download the
.gguffile - Import into LM Studio
- Start chatting with the spirit
๐ Model Details
| Property | Value |
|---|---|
| Base Model | LiquidAI LFM2.5-1.2B-Instruct |
| Architecture | LFM2 hybrid (short convolutions + grouped-query attention) |
| Parameters | 1.2B |
| Fine-tuning | LoRA 16-bit (r=16, alpha=32, dropout=0.05) |
| LoRA targets | q_proj, k_proj, v_proj, out_proj, in_proj, w1, w2, w3 |
| Training Examples | 1,000 (identical to Ouija2) |
| Epochs | 3 |
| Learning rate | 2e-4 (linear, AdamW 8-bit) |
| Loss | Responses only (system and user turns masked) |
| Training sequence length | 256 tokens |
| Quantization | Q4_K_M |
| File Size | ~730 MB |
๐ญ Behavior Rules
The spirit follows these sacred rules:
1. Respond ONLY with: YES, NO, MAYBE, or ONE word
2. For yes/no questions: "YES. [CONTEXT]" or "NO. [CONTEXT]"
3. When cannot express something: "Ouija: [hint]"
4. Spell names letter by letter: M... A... R... I... A...
5. Always respond in UPPERCASE
6. Never explain. Never elaborate. Never break character.
๐ง Technical Notes
A different anatomy
LFM2 is not a classic Transformer. Most of its layers are short gated convolutions and only a few are attention, so its modules are named differently from Qwen. The LoRA adapters target both kinds of layer:
- Attention:
q_proj,k_proj,v_proj,out_proj - Convolution blocks:
in_proj,out_proj - MLP:
w1,w2,w3
Training recipe
- Native chat template: examples are formatted with LFM2.5's own
apply_chat_template(<|im_start|>/<|im_end|>) instead of a hand-written prompt - Responses only: the loss is computed only on the spirit's answer, so the model learns how to answer rather than memorizing the system prompt
- 16-bit LoRA: at 1.2B parameters the base fits in a free Colab T4 without 4-bit loading, which gives a cleaner merge
- GGUF export: LoRA merged to 16-bit, converted with
convert_hf_to_gguf.pyand quantized withllama-quantizeto Q4_K_M, the same format as Liquid's official GGUF (730,895,168 bytes)
Reproduce it
Open Ouija3_1.2B_Colab.ipynb in Google Colab, choose a T4 GPU and run all cells. Total time is about 15-20 minutes, including building llama-quantize.
โ ๏ธ Limitations
- Not for serious use: This is an entertainment/art project
- Short responses only: Won't generate long text
- English only: Trained on English data
- May hallucinate: Like any LLM, responses are generated, not supernatural
๐ธ๏ธ Why does this exist?
Because we asked: "What if an LLM refused to be helpful?"
Most AI assistants try to be as helpful as possible. Ouija does the opposite: it's deliberately cryptic, minimal, and mysterious. It's an exploration of:
- Fine-tuning for behavioral constraints
- Creating character-locked models
- The intersection of AI and folklore
- Making something fun in the age of utility
v3 asks a new question: does the spirit survive when you change not just the size of the vessel, but its whole anatomy?
๐ License
LFM Open License v1.0 (inherited from LFM2.5-1.2B-Instruct). See the base model card for the full terms.
๐ Credits
- Base Model: LiquidAI/LFM2.5-1.2B-Instruct by Liquid AI
- Fine-tuning: Unsloth
- GGUF conversion: llama.cpp
- v2: Ouija2-1.7B
- v1: Ouija-3B
- Inspiration: Every horror movie with a Ouija board scene
_______________
| ___________ |
| | YES NO | |
| | A B C D | |
| | E F G H | |
| | I J K L | |
| | M N O P | |
| | Q R S T | |
| | U V W X | |
| | Y Z | |
| | GOODBYE | |
|_|___________|_|
The spirit is listening...
New vessel. Liquid body. Same darkness. Always say goodbye.
๐ป
- Downloads last month
- 32
4-bit
Model tree for BansheeTechnologies/Ouija3-1.2B
Base model
LiquidAI/LFM2.5-1.2B-Base