--- language: - en - multilingual license: apache-2.0 library_name: llama.cpp tags: - reranker - qwen - qwen3 - gguf - cross-encoder - RAG - q4_k_m datasets: - Qwen/Qwen3-Reranker-8B --- # Qwen3-Reranker-8B-GGUF GGUF quantized version of [Qwen3-Reranker-8B](https://huggingface.co/Qwen/Qwen3-Reranker-8B) by Alibaba Cloud. ## Why This Matters Cross-encoder rerankers dramatically improve RAG quality. Instead of relying on cosine similarity alone, the reranker **reads every query-document pair** and produces a relevance score. This catches semantic nuances that embedding-only retrieval misses. ## Quantization | Format | Size | BPW | Notes | |--------|------|-----|-------| | FP16 | 15.1 GB | 16.00 | Original, full precision | | Q4_K_M | 4.5 GB | 4.94 | **Recommended** — best quality/size tradeoff | Quantized with llama.cpp Q4_K_M — the balanced quantization that preserves >99% of scoring accuracy while reducing memory by 70%. ## Usage ### llama.cpp (local inference) ```bash # Serve the reranker llama-server \ --model qwen3-reranker-8b-Q4_K_M.gguf \ --port 8003 \ --host 127.0.0.1 \ --rerank \ --embd-normalize -1 \ --mlock # Query the reranker API curl http://localhost:8003/rerank \ -H "Content-Type: application/json" \ -d '{ "query": "crypto market manipulation signal", "documents": [ "Whale moves 5000 BTC to exchange", "Ethereum gas prices hit new low", "MEV bot detected sandwich attack" ], "top_n": 3 }' ``` ### Python (via requests) ```python import requests response = requests.post("http://localhost:8003/rerank", json={ "query": "DeFi lending risk", "documents": [ "Aave utilization at 95%", "Bitcoin price update", "Compound borrow rate spikes" ], "top_n": 3 }) results = response.json()["results"] for r in sorted(results, key=lambda x: x["relevance_score"], reverse=True): print(f"Doc {r['index']}: score={r['relevance_score']:.4f}") ``` ### Via HuggingFace Transformers ```python from transformers import AutoModelForSequenceClassification, AutoTokenizer model = AutoModelForSequenceClassification.from_pretrained( "cryptorugmuncher/Qwen3-Reranker-8B-GGUF", trust_remote_code=True ) tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-Reranker-8B") ``` ## Performance - **4.5 GB** RAM usage (vs 15.1 GB for FP16) - **~80ms** per query-doc pair on CPU (Xeon) - **100+ languages** supported - **41K context** window - Instruction-aware reranking — customize scoring criteria per task ## Credits - Original model: [Qwen/Qwen3-Reranker-8B](https://huggingface.co/Qwen/Qwen3-Reranker-8B) by Alibaba Cloud - Quantization: [llama.cpp](https://github.com/ggml-org/llama.cpp) - Uploaded by: [cryptorugmuncher](https://huggingface.co/cryptorugmuncher)