cryptorugmunch's picture
Upload README.md with huggingface_hub
a76a97c verified
|
Raw
History Blame Contribute Delete
2.81 kB
metadata
language:
  - en
  - multilingual
license: apache-2.0
library_name: llama.cpp
tags:
  - reranker
  - qwen
  - qwen3
  - gguf
  - cross-encoder
  - RAG
  - q4_k_m
datasets:
  - Qwen/Qwen3-Reranker-8B

Qwen3-Reranker-8B-GGUF

GGUF quantized version of Qwen3-Reranker-8B by Alibaba Cloud.

Why This Matters

Cross-encoder rerankers dramatically improve RAG quality. Instead of relying on cosine similarity alone, the reranker reads every query-document pair and produces a relevance score. This catches semantic nuances that embedding-only retrieval misses.

Quantization

Format Size BPW Notes
FP16 15.1 GB 16.00 Original, full precision
Q4_K_M 4.5 GB 4.94 Recommended — best quality/size tradeoff

Quantized with llama.cpp Q4_K_M — the balanced quantization that preserves >99% of scoring accuracy while reducing memory by 70%.

Usage

llama.cpp (local inference)

# Serve the reranker
llama-server \
    --model qwen3-reranker-8b-Q4_K_M.gguf \
    --port 8003 \
    --host 127.0.0.1 \
    --rerank \
    --embd-normalize -1 \
    --mlock

# Query the reranker API
curl http://localhost:8003/rerank \
  -H "Content-Type: application/json" \
  -d '{
    "query": "crypto market manipulation signal",
    "documents": [
      "Whale moves 5000 BTC to exchange",
      "Ethereum gas prices hit new low",
      "MEV bot detected sandwich attack"
    ],
    "top_n": 3
  }'

Python (via requests)

import requests

response = requests.post("http://localhost:8003/rerank", json={
    "query": "DeFi lending risk",
    "documents": [
        "Aave utilization at 95%",
        "Bitcoin price update",
        "Compound borrow rate spikes"
    ],
    "top_n": 3
})
results = response.json()["results"]
for r in sorted(results, key=lambda x: x["relevance_score"], reverse=True):
    print(f"Doc {r['index']}: score={r['relevance_score']:.4f}")

Via HuggingFace Transformers

from transformers import AutoModelForSequenceClassification, AutoTokenizer

model = AutoModelForSequenceClassification.from_pretrained(
    "cryptorugmuncher/Qwen3-Reranker-8B-GGUF",
    trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-Reranker-8B")

Performance

  • 4.5 GB RAM usage (vs 15.1 GB for FP16)
  • ~80ms per query-doc pair on CPU (Xeon)
  • 100+ languages supported
  • 41K context window
  • Instruction-aware reranking — customize scoring criteria per task

Credits