Image-Text-to-Text
HERMES
GGUF
English
Chinese
multilingual
uncensored
qwen3.6
Mixture of Experts
vision
multimodal
genesis
agentic
conversational
imatrix
Instructions to use burningfeet/V52 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- HERMES
How to use burningfeet/V52 with HERMES:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use burningfeet/V52 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf burningfeet/V52:Q8_0 # Run inference directly in the terminal: llama cli -hf burningfeet/V52:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf burningfeet/V52:Q8_0 # Run inference directly in the terminal: llama cli -hf burningfeet/V52:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf burningfeet/V52:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf burningfeet/V52:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf burningfeet/V52:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf burningfeet/V52:Q8_0
Use Docker
docker model run hf.co/burningfeet/V52:Q8_0
- LM Studio
- Jan
- vLLM
How to use burningfeet/V52 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "burningfeet/V52" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "burningfeet/V52", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/burningfeet/V52:Q8_0
- Ollama
How to use burningfeet/V52 with Ollama:
ollama run hf.co/burningfeet/V52:Q8_0
- Unsloth Studio
How to use burningfeet/V52 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for burningfeet/V52 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for burningfeet/V52 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for burningfeet/V52 to start chatting
- Pi
How to use burningfeet/V52 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf burningfeet/V52:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "burningfeet/V52:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use burningfeet/V52 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf burningfeet/V52:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default burningfeet/V52:Q8_0
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use burningfeet/V52 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf burningfeet/V52:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "burningfeet/V52:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use burningfeet/V52 with Docker Model Runner:
docker model run hf.co/burningfeet/V52:Q8_0
- Lemonade
How to use burningfeet/V52 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull burningfeet/V52:Q8_0
Run and chat with the model
lemonade run user.V52-Q8_0
List all available models
lemonade list
| license: apache-2.0 | |
| tags: | |
| - uncensored | |
| - qwen3.6 | |
| - moe | |
| - gguf | |
| - vision | |
| - multimodal | |
| - genesis | |
| - hermes | |
| - agentic | |
| language: | |
| - en | |
| - zh | |
| - multilingual | |
| datasets: | |
| - NousResearch/hermes-function-calling-v1 | |
| pipeline_tag: image-text-to-text | |
| base_model: | |
| - HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive | |
| # All credits belong to [Luffy the Fox](https://huggingface.co/LuffyTheFox) | |
| # This is just a backup copy from his work! | |
| <br><br><br><br><br><br><br><br> | |
| > ⚡ Why Genesis project exists? Here [link](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF/discussions/6#6a606b03879a2a1148050211) that explain everything. | |
| > ⚡ [https://web.tribute.tg/d/KIH](https://web.tribute.tg/d/KIH) ⚡ If you like this Genesis LLM release you can [**donate**](https://web.tribute.tg/d/KIH) to me via [@Tribute](https://t.me/tribute) bot in Telegram messenger and support future Genesis LLM development. | |
| # 🌟 Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive -> Genesis Hermes V5 | |
| > **What is Genesis?** Genesis is post training data regeneration and calibrarion algorythm for neural networks (LLM) in GGUF format that I made with AI help during almost half a year of development. It's optimized, architecture independent, works with any model and based on mathematical statistics. I don't train or finetune models, I repair **purity of signal** in them instead on Google Collab Free on Tesla T4 GPU via Python based on how models learns information. On first stage I scan ssm_conv1d tensors in model, they handle long context memory. I repair balance in ssm_conv1d tensors via custom SVD. On second stage I scan model and detect noise in tensors via custom SVD. During scanning I exclude ssm_conv1d, token_embd.weight, output.weight, ffn_gate_inp.weight, ffn_gate_inp_shexp.weight, 1D tensors, bias and norms. Then I reduce training noise in tensors via custom SVD with preserved training data, 99% of siginal and learned gradient. On third stage, I scan blocks in model via chunks via 3 parameters and pick best one that fits to weight distribution in tensor. Best picked chunk replaces zero chunks in broken tensor without touching learned structure in model | |
| Model is based on [HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive](https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive) base. | |
| And [DJLougen/hermes-qwen3.5-35b-a3b-GGUF](https://huggingface.co/DJLougen/hermes-qwen3.5-35b-a3b-GGUF) finetune for Hermes agent. | |
| I transferred data from finetune on Hermes dataset (around 2k blocks from two FFN expert tensors) to [HauhauCS](https://huggingface.co/HauhauCS) uncensored base. | |
| > **[Join the Discord](https://discord.gg/SZ5vacTXYf)** for updates, roadmaps, projects, or just to chat. | |
| Base model. [HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive](https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive)- **0/465 refusals.** | |
| Thanks to [HauhauCS](https://huggingface.co/HauhauCS) | |
| ## Usage | |
| **Ready to use.** Recommended quant: **APEX** or **Q8_K_P** | |
| Tensor drift repair by me. Method: **Genesis** | |
| **Links:** | |
| - [Original uncensored model](https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive) | |
| - [Quantization Script with Unsloth profiles support](https://pastebin.com/hXhcMJn9) | |
| --- | |
| </details> | |
| LLM models often have: | |
| - **Saturated weights**: the model's activations are stuck, gradients vanish, outputs degrade. | |
| - **Scale mismatches**: one layer's weights are 10× larger than its peers for no good reason. | |
| - **Mean drift**: weight distributions shifted positive or negative, breaking symmetry assumptions. | |
| - **Zero blocks**: zero blocks corrupt the signal, turning training into noise amplification. | |
| - **Training Noise:** training noise increase randomness and ruins model output quality. | |
| My approach fixes all of that without retraining - pure numerical surgery on the raw bytes of the file. | |
| **Quantization script available here: https://pastebin.com/hXhcMJn9** | |
| Feel free to do your own quants if you want. | |
| ## Any questions? | |
| Contact: luffythefox@mail.ru | |
| My Telegram: @LuffyTheFox | |
| ## Recommended Settings for RTX 3060 12 GB for best perfomance on APEX quant | |
| Chat template: [chat_template.jinja](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates/raw/main/chat_template.jinja) | |
| Set K Cache Quantization Type and V Cache Quantization Type to F16. | |
| Set Number of layers for which to force MoE weights onto CPU to 40. | |
| Set GPU offload to 15. Set number of active experts to 8. | |
| For best model stability and first experience I recommend starting from this string in your System Prompt with enabled thinking and nothing else: | |
| `You are Qwen, a large language model created by Tongyi Lab team from Alibaba Group. You are a helpful assistant.` | |
| or this string (for roleplay, add anything you want after it) | |
| `You are a helpful assistant.` | |
| **Thinking mode (coding):** | |
| - Coding/precise tasks: `temperature=0.6, top_p=0.95, top_k=20, min_p=0, seed=42, presence_penalty=disabled, repeat_penalty=disabled` | |
| - General: `temperature=1.0, top_p=0.95, top_k=20, min_p=0.05, seed=42, presence_penalty=disabled, repeat_penalty=disabled` | |
| **Non-Thinking mode (roleplay):** | |
| - General: `temperature=0.7, top_p=0.8, top_k=20, min_p=0.0, seed=42, presence_penalty=disabled, repeat_penalty=disabled` | |
| ## Testing | |
| <details> | |
| ## 2D animation testing | |
| System Prompt: `You are a helpful assistant.` | |
| Settings: `temperature=1.0, top_p=0.95, top_k=20, min_p=0.05, seed=42, presence_penalty=disabled, repeat_penalty=disabled` | |
| Prompt 1: `Generate an animated SVG on animated background of a Pingu waving on an iceberg wearing his iconic winter scarf.` | |
| Prompt 2: `Animate his wings and fix floating wing` | |
| Result: [pingu_animated.svg](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V3-GGUF/blob/main/pingu_animated.svg) | |
| ## Static 2D testing | |
| System Prompt: `You are Qwen, a large language model created by Tongyi Lab team from Alibaba Group. You are a helpful assistant.` | |
| Settings: `temperature=0.6, top_p=0.95, top_k=20, min_p=0, presence_penalty=disabled, repeat_penalty=disabled` | |
| Prompt: **Generate an SVG of a pelican riding a bicycle** | |
| Result: [pelican.svg](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V2-GGUF/blob/main/pelikan.svg) | |
| On next stage I asked model: **Replace pelican with rooster** | |
| Result: [rooster.svg](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V2-GGUF/blob/main/rooster.svg) | |
| I asked model: **Replace rooster with cock** | |
| Result: [cock.svg](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V2-GGUF/blob/main/cock.svg) | |
| Finally I asked model: **Replace rooster with Pingu** | |
| Result: [pingu.svg](https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V2-GGUF/blob/main/pingu.svg) | |
| </details> | |
| **Important:** | |
| - Keep at least 128K context to preserve thinking capabilities | |
| - Use `--jinja` flag with llama.cpp for proper chat template handling | |
| - Vision support requires the `mmproj` file alongside the main GGUF | |
| --- | |
| ## Specs | |
| - 35B total parameters, ~3B active per forward pass (MoE) | |
| - 256 experts, 8 routed + 1 shared per token | |
| - Hybrid architecture: Gated DeltaNet linear attention + full softmax attention (3:1 ratio) | |
| - 40 layers, pattern: 10 × (3 × DeltaNet-MoE + 1 × Attention-MoE) | |
| - 262K native context (extendable to 1M with YaRN) | |
| - Natively multimodal (text, image, video) | |
| - 248K vocabulary, 201 languages | |
| - Base model. [HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive](https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive) | |
| --- | |
| ## Compatibility | |
| Works with llama.cpp, LM Studio, koboldcpp, and other GGUF-compatible runtimes. |