Abstract
Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select evidence across this range. We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference. UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM's internal representations. It adds fewer than 500K trainable parameters and leaves the backbone unchanged. On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems. The best model raises recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue. Applied to long-context tasks, the same selection mechanism removes distractors before generation, raising NoLiMa accuracy from 1.0% to 24.83% at its maximum context length of 128K tokens, and LV-Eval's F1 score from 49.97% to 54.66% at 256K. UNREAL also reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows. Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.
Community
We introduce UNREAL, which uses a single frozen LLM to retrieve evidence and generate answers. It adds fewer than 500K trainable parameters and uses the model’s internal representations to select relevant chunks from long prompts or entire corpora, without an external retrieval model.
On a 21M-chunk Wikipedia index, the best model improves HotpotQA recall@10 from 49.1% to 73.2%. The same retrieval-trained module improves long-context accuracy and reduces FLOPs and time to first token relative to full-context inference from roughly 32K tokens onward.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling (2026)
- Mixture-of-Experts Language Models Can Be Strong and Efficient Retrievers (2026)
- MegaMem: A Retrieval Solution for Ultra-Large Context Windows (2026)
- Doc-REFRAG: Rethinking Multimodal Document Retrieval-Augmented Generation (2026)
- CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG (2026)
- CAGE: Coherence-Aware Graph Encoding for Retrieval-Augmented Generation (2026)
- RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.08463 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper