RenderRank

RenderRank: Learning to Rerank Text with Compressed Visual Tokens

Paper: RenderRank: Learning to Rerank Text with Compressed Visual Tokens

RenderRank converts text into compact visual token sequences by rendering it as images. This change in input modality reduces the number of tokens used to represent documents, allowing more content to fit within the same context budget. Queries remain text, and relevance is scored against the visual document representations.

RenderRank architecture: text queries and visually encoded documents for relevance scoring

Highlights

  • Text as visual tokens: represent document text through rendered images instead of conventional text token sequences.
  • Shorter input sequences: compact visual representations reduce input length and allow more document content within a fixed context budget.
  • Three input modes: score document text directly with text, render it automatically with render, or provide existing document pages with image.
Property Value
Model type Pairwise document reranker
Backbone Qwen3-VL-Reranker-2B
Parameters 2B
Language evaluated English
Maximum context length 32,768 tokens, including query, template, document text, and visual tokens
Scoring Yes/no relevance scoring

Rendering Configuration

Setting Value
Font Roboto Regular
Font size 12pt at 96 DPI (16px)
Line spacing 1.0 (16px line height)
Page width 896px
Page height 32–896px, rounded up to a multiple of 32px
Lines per page Up to 56
Margins 0px on all sides
Colors Black text on a white background
Rendering Pillow default text antialiasing

Whitespace, including paragraph breaks and tabs, is normalized to a single space. Lines wrap according to the font's actual pixel width; words wider than the page are split by character. Overflow continues onto subsequent pages without overlap. The final page is sized to its content, and empty documents produce one 896 × 32px page.

All generated pages are supplied in order as one document, producing one relevance score. Rendering is not limited to the first page.

BEIR Performance

All rerankers score the top 100 BM25 candidates for each query. Baselines use text queries and documents; RenderRank uses text queries and documents rendered with the default 12pt configuration above. Scores are NDCG@10; Avg. is the mean across 11 datasets. Tokens is the mean input length per query–document pair before truncation, averaged across the same datasets. For RenderRank, this includes both textual and visual tokens.

Model AA CFV DBP FQA FVR HQA NFC SD SF TC TCH Avg. Tokens
gte-reranker-modernbert-base 66.14 25.80 42.10 42.54 89.84 76.15 34.81 18.90 75.81 79.66 33.44 53.20 359.25
mxbai-rerank-large-v1 16.16 25.61 45.37 40.29 82.19 71.58 37.32 18.96 75.26 85.84 37.44 48.73 347.45
bge-reranker-large 30.94 34.11 44.21 37.47 89.30 80.09 33.91 16.68 73.98 73.21 34.69 49.87 409.32
LAMAR-600m 66.21 36.21 46.23 41.14 89.49 79.14 35.04 20.22 76.97 80.65 35.76 55.19 409.32
Qwen3-Reranker-0.6B 68.04 34.34 43.76 39.74 87.13 77.25 36.60 20.44 77.00 85.54 30.32 54.56 438.70
llama-nemotron-rerank-1b-v2 54.29 27.37 45.13 47.11 87.66 80.34 38.14 21.53 79.82 83.14 32.47 54.27 363.99
mxbai-rerank-large-v2 42.23 23.03 38.01 28.85 70.84 69.50 35.17 17.49 78.81 67.78 46.94 47.15 449.76
LightOn-rerank-PW-2B 45.52 23.21 42.60 35.66 86.51 73.05 35.61 16.64 77.70 81.24 36.91 50.42 416.22
bge-reranker-v2-gemma 74.98 32.03 45.01 43.18 87.45 80.47 36.25 19.58 77.43 80.08 37.56 55.82 399.11
Qwen3-Reranker-4B 75.40 37.84 47.02 44.63 89.14 79.06 37.48 23.59 79.47 85.47 38.77 57.99 438.70
zerank-2-reranker 44.87 23.65 44.95 42.75 83.22 70.97 38.50 20.23 79.35 85.24 38.10 51.98 380.69
LightOn-rerank-PW-4B 52.50 28.24 44.80 40.27 87.70 72.79 37.33 17.29 76.72 79.54 36.70 52.17 416.22
RenderRank (Ours) 69.11 35.91 46.19 42.66 88.92 78.60 37.31 20.75 77.00 84.82 34.27 55.96 290.07

AA: ArguAna; CFV: Climate-FEVER; DBP: DBPedia; FQA: FiQA; FVR: FEVER; HQA: HotpotQA; NFC: NFCorpus; SD: SCIDOCS; SF: SciFact; TC: TREC-COVID; TCH: Touché-2020.

Long-Document Performance

Results on the English subset of MLDR and three LongEmbed datasets: 2WikiMQA, QMSum, and SummScreenFD. The comparison includes models supporting at least 16K input tokens. MLDR follows the MMTEB reranking protocol; for LongEmbed, each query reranks eight candidates retrieved with Qwen3-Embedding-0.6B.

Each dataset cell shows NDCG@10 followed by the average input token count in brackets: score [tokens]. Token counts are measured per query–document pair before truncation and include both textual and visual tokens for RenderRank. Avg. reports the mean score across the four datasets.

Model MLDR 2WikiMQA QMSum SummScreenFD Avg.
Qwen3-Reranker-0.6B 99.63 [8863.4] 94.54 [9230.2] 56.43 [13543.6] 98.25 [8606.3] 87.21
LightOn-rerank-PW-2B 98.91 [8878.0] 71.27 [9230.1] 54.30 [13876.6] 93.05 [8977.2] 79.38
Qwen3-Reranker-4B 99.85 [8863.4] 94.54 [9230.2] 59.11 [13543.6] 99.11 [8606.3] 88.15
zerank-2-reranker 99.57 [8805.6] 94.42 [9173.1] 60.10 [13485.9] 98.97 [8548.5] 88.26
LightOn-rerank-PW-4B 99.68 [8878.0] 93.84 [9230.1] 58.87 [13876.6] 98.91 [8977.2] 87.82
RenderRank (Ours) 99.74 [4198.0] 94.30 [4262.2] 59.94 [6388.8] 99.10 [3701.9] 88.27

Usage

The examples load the model from nlpai-lab/RenderRank-2B on Hugging Face.

The model applies its reranking chat template internally. Pass the instruction as shown below; do not manually prepend a chat template to the document.

Transformers

This repository provides a custom AutoModel entry point with a process() method. It supports all three document input modes.

Requirements

The Transformers interface was tested with the following versions:

transformers==5.9.0
qwen-vl-utils==0.0.14
torch==2.11.0
torchvision==0.26.0
scipy

qwen-vl-utils handles image preparation in this interface; scipy is used for score normalization. NumPy and Pillow are installed through the dependencies above. Install PyTorch and torchvision builds matching your CUDA environment.

import torch
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "nlpai-lab/RenderRank-2B",
    trust_remote_code=True,
    dtype=torch.bfloat16,
).to("cuda").eval()

inputs = {
    "instruction": "Retrieve text relevant to the user's query.",
    "query": {"text": "What is the capital of France?"},
    "documents": [
        # Score text directly.
        {"text": "Paris is the capital of France."},
        # Render text into document pages, then score the images.
        {"render": "Paris is the capital of France."},
        # Score an existing document image.
        {"image": "/path/to/document.png"},
        # Score multiple ordered pages as one document.
        {"image": ["/path/to/page1.png", "/path/to/page2.png"]},
    ],
}

with torch.inference_mode():
    scores = model.process(inputs)

ranked_indices = sorted(range(len(scores)), key=lambda i: scores[i], reverse=True)
print(scores)
print(ranked_indices)

Rendering

With Transformers, {"render": document_text} automatically renders the document before scoring. The font is included in this repository.

Alternatively, render documents in advance with the bundled rendering.py and save the page images for reuse with either Transformers or Sentence Transformers. This avoids repeated rendering; vision encoding still runs when scoring the saved images.

import sys
from pathlib import Path

from huggingface_hub import snapshot_download

repo_dir = Path(snapshot_download(
    "nlpai-lab/RenderRank-2B",
    allow_patterns=["rendering.py", "Roboto-Regular.ttf"],
))
sys.path.insert(0, str(repo_dir))
from rendering import DocumentImageConfig, DocumentImageRenderer

renderer = DocumentImageRenderer(
    DocumentImageConfig(font_path=repo_dir / "Roboto-Regular.ttf")
)
pages = renderer.render_document("Paris is the capital of France.")
# pages is a list of PIL images, in document order.

# Save once and reuse these files for subsequent queries.
output_dir = Path("rendered_document")
output_dir.mkdir(parents=True, exist_ok=True)
for i, page in enumerate(pages, start=1):
    page.save(output_dir / f"{i}.png")

The returned PIL images can be passed directly as image content or saved as numbered page files. The Transformers render mode calls this renderer internally, so no separate rendering step is needed there.

Sentence Transformers

Use Sentence Transformers v6.1.0 or later.

pip install -U "sentence-transformers>=6.1.0"

The same interface can score text, a single document image, or multiple page images as one document. In the example below, 1.png, 2.png, and 3.png are ordered pages of one document.

import torch
from sentence_transformers import CrossEncoder

model = CrossEncoder(
    "nlpai-lab/RenderRank-2B",
    device="cuda",
    max_length=32768,
    model_kwargs={"dtype": torch.bfloat16},
)

query = "What is the capital of France?"
documents = [
    # Text document
    {"text": "Paris is the capital of France."},
    # Single-page image document
    {"image": "/path/to/document.png"},
    # Multi-page image document
    {"image": [
        "/path/to/1.png",
        "/path/to/2.png",
        "/path/to/3.png",
        # Add further pages here in document order.
    ]},
]

scores = model.predict(
    [(query, document) for document in documents],
    prompt="Retrieve text relevant to the user's query.",
    batch_size=1,
)
print(scores)  # One score per document, not per page.

Larger scores indicate greater relevance. Scores are the difference between the yes and no logits, as used in our evaluation. With Sentence Transformers v6.1.0 or later, multi-page documents can be passed as a single multimodal input using {"image": [page1, page2, ...]}. The pages are processed together and receive one relevance score per document. The render input format is supported by the Transformers interface above.

Citation

@misc{hong2026renderranklearningreranktext,
      title={RenderRank: Learning to Rerank Text with Compressed Visual Tokens},
      author={Seongtae Hong and Youngjoon Jang and Jungseob Lee and Hyeonseok Moon and Heuiseok Lim},
      year={2026},
      eprint={2609.35069},
      archivePrefix={arXiv},
      primaryClass={cs.IR},
      url={https://arxiv.org/abs/2609.35069},
}
Downloads last month
123
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nlpai-lab/RenderRank-2B

Finetuned
(3)
this model
Quantizations
1 model

Datasets used to train nlpai-lab/RenderRank-2B

Paper for nlpai-lab/RenderRank-2B