Text Generation
GGUF
English
Persian
llama.cpp
ternary
bonsai
decision-making
structured-output
zero-shot
conversational
Instructions to use Reza2kn/Bev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Reza2kn/Bev with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Reza2kn/Bev:Q2_0 # Run inference directly in the terminal: llama cli -hf Reza2kn/Bev:Q2_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Reza2kn/Bev:Q2_0 # Run inference directly in the terminal: llama cli -hf Reza2kn/Bev:Q2_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Reza2kn/Bev:Q2_0 # Run inference directly in the terminal: ./llama-cli -hf Reza2kn/Bev:Q2_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Reza2kn/Bev:Q2_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Reza2kn/Bev:Q2_0
Use Docker
docker model run hf.co/Reza2kn/Bev:Q2_0
- LM Studio
- Jan
- vLLM
How to use Reza2kn/Bev with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Reza2kn/Bev" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Reza2kn/Bev", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Reza2kn/Bev:Q2_0
- Ollama
How to use Reza2kn/Bev with Ollama:
ollama run hf.co/Reza2kn/Bev:Q2_0
- Unsloth Desktop
- Pi
How to use Reza2kn/Bev with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Reza2kn/Bev:Q2_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Reza2kn/Bev:Q2_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Reza2kn/Bev with Docker Model Runner:
docker model run hf.co/Reza2kn/Bev:Q2_0
- Lemonade
How to use Reza2kn/Bev with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Reza2kn/Bev:Q2_0
Run and chat with the model
lemonade run user.Bev-Q2_0
List all available models
lemonade list
- Hermes Agent
How to use Reza2kn/Bev with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Reza2kn/Bev:Q2_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Reza2kn/Bev:Q2_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Reza2kn/Bev with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Reza2kn/Bev:Q2_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Reza2kn/Bev:Q2_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| { | |
| "started_utc": "2026-09-23T18:05:41.534203+00:00", | |
| "scope": "isolated Stallion CPU-only package and runtime-script checks; no service restart or GPU inference", | |
| "checks": [ | |
| { | |
| "name": "create_environment", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "install_release_dependencies", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "shell_syntax_install-stallion.sh", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "shell_syntax_install.sh", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "shell_syntax_runtime-env.sh", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "shell_syntax_start-api.sh", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "shell_syntax_start-llama.sh", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "shell_syntax_start-services.sh", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "shell_syntax_stop-services.sh", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "unit_tests", | |
| "exit_code": 0, | |
| "output": "..................................................... [100%]\n53 passed in 0.25s\n" | |
| }, | |
| { | |
| "name": "dependency_consistency", | |
| "exit_code": 0, | |
| "output": "No broken requirements found.\n" | |
| }, | |
| { | |
| "name": "build_distributions", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "create_wheel_environment", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "install_wheel", | |
| "exit_code": 0 | |
| }, | |
| { | |
| "name": "wheel_import", | |
| "exit_code": 0, | |
| "output": "0.1.1\n0.1.1\n" | |
| } | |
| ], | |
| "completed_utc": "2026-09-23T18:09:14.666904+00:00", | |
| "status": "passed", | |
| "native_installer": { | |
| "exit_code": 0, | |
| "validation": "Fresh isolated Prism source clone and 334-step host build with six jobs; CUDA loader compiled; CUDA device enumeration; editable API installation. No model server started and no model inference performed.", | |
| "cached_assets": "Exact verified GGUF and upstream runtime were hardlinked from existing immutable files; this check exercised cache verification, not their large-file network download paths. Model notices and source were fetched.", | |
| "runtime_revision": "9a9394a895b96003ca842a6041cb28ac49a108f7", | |
| "cuda_device": [ | |
| "CUDA0: NVIDIA GeForce RTX 5080 Laptop GPU (15839 MiB, 3512 MiB free)" | |
| ], | |
| "script_sha256": { | |
| "install-stallion.sh": "f99df78c1223285425cbeeda9df322bbfe842764353da4bd8d136a9236cc1715", | |
| "install.sh": "11e5f33643e334b8e21b9689e1e2b0f405a018d0f67f0fd7d09ae6bd81f42aaf", | |
| "runtime-env.sh": "915737684c7501033152af6e37033b5fc087bb05c093c7db4de8d9ab06120ad0", | |
| "start-api.sh": "7072415ab699efa1912ddebc1fe3c9152206f2ce47a401688918ef5870f7b54b", | |
| "start-llama.sh": "88f905b4d091e4dc49050a0441b4580b84241d6021df8755a2179d5441cdc409", | |
| "start-services.sh": "52254054e2214bfb2b25f661e2ca8c0cdf8eca1bb9902cac9f148623a1fec7a4", | |
| "stop-services.sh": "ec428b23ef9f5a2a569eed3d70460487379bae7d00b4ad587c9318f7b1891558" | |
| } | |
| }, | |
| "publication_note": "Public validation projection. Initial package-build hashes omitted because final release artifacts are rebuilt from the curated source and have their own SHA256SUMS. Private paths and raw build logs are not published." | |
| } | |