cryptorugmunch's picture
Upload README.md with huggingface_hub
a76a97c verified
|
Raw
History Blame Contribute Delete
2.81 kB
---
language:
- en
- multilingual
license: apache-2.0
library_name: llama.cpp
tags:
- reranker
- qwen
- qwen3
- gguf
- cross-encoder
- RAG
- q4_k_m
datasets:
- Qwen/Qwen3-Reranker-8B
---
# Qwen3-Reranker-8B-GGUF
GGUF quantized version of [Qwen3-Reranker-8B](https://huggingface.co/Qwen/Qwen3-Reranker-8B) by Alibaba Cloud.
## Why This Matters
Cross-encoder rerankers dramatically improve RAG quality. Instead of relying on cosine similarity alone, the reranker **reads every query-document pair** and produces a relevance score. This catches semantic nuances that embedding-only retrieval misses.
## Quantization
| Format | Size | BPW | Notes |
|--------|------|-----|-------|
| FP16 | 15.1 GB | 16.00 | Original, full precision |
| Q4_K_M | 4.5 GB | 4.94 | **Recommended** — best quality/size tradeoff |
Quantized with llama.cpp Q4_K_M — the balanced quantization that preserves >99% of scoring accuracy while reducing memory by 70%.
## Usage
### llama.cpp (local inference)
```bash
# Serve the reranker
llama-server \
--model qwen3-reranker-8b-Q4_K_M.gguf \
--port 8003 \
--host 127.0.0.1 \
--rerank \
--embd-normalize -1 \
--mlock
# Query the reranker API
curl http://localhost:8003/rerank \
-H "Content-Type: application/json" \
-d '{
"query": "crypto market manipulation signal",
"documents": [
"Whale moves 5000 BTC to exchange",
"Ethereum gas prices hit new low",
"MEV bot detected sandwich attack"
],
"top_n": 3
}'
```
### Python (via requests)
```python
import requests
response = requests.post("http://localhost:8003/rerank", json={
"query": "DeFi lending risk",
"documents": [
"Aave utilization at 95%",
"Bitcoin price update",
"Compound borrow rate spikes"
],
"top_n": 3
})
results = response.json()["results"]
for r in sorted(results, key=lambda x: x["relevance_score"], reverse=True):
print(f"Doc {r['index']}: score={r['relevance_score']:.4f}")
```
### Via HuggingFace Transformers
```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model = AutoModelForSequenceClassification.from_pretrained(
"cryptorugmuncher/Qwen3-Reranker-8B-GGUF",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-Reranker-8B")
```
## Performance
- **4.5 GB** RAM usage (vs 15.1 GB for FP16)
- **~80ms** per query-doc pair on CPU (Xeon)
- **100+ languages** supported
- **41K context** window
- Instruction-aware reranking — customize scoring criteria per task
## Credits
- Original model: [Qwen/Qwen3-Reranker-8B](https://huggingface.co/Qwen/Qwen3-Reranker-8B) by Alibaba Cloud
- Quantization: [llama.cpp](https://github.com/ggml-org/llama.cpp)
- Uploaded by: [cryptorugmuncher](https://huggingface.co/cryptorugmuncher)