KRLabsOrg
/

lettucedect-base-modernbert-en-v1

@@ -1,14 +1,15 @@
 ---
-license: mit
-language:
-- en
 base_model:
 - answerdotai/ModernBERT-base
-pipeline_tag: token-classification
 tags:
-- token classification
-- hallucination detection
 - transformers
 ---
 # LettuceDetect: Hallucination Detection Model
@@ -23,13 +24,20 @@ tags:
 ## Overview
-LettuceDetect is a transformer-based model for hallucination detection on context and answer pairs, designed for Retrieval-Augmented Generation (RAG) applications. This model is built on **ModernBERT**, which has been specifically chosen and trained becasue of its extended context support (up to **8192 tokens**). This long-context capability is critical for tasks where detailed and extensive documents need to be processed to accurately determine if an answer is supported by the provided context.
-**This is our Large model based on ModernBERT-large**
 ## Model Details
-- **Architecture:** ModernBERT (Large) with extended context support (up to 8192 tokens)
 - **Task:** Token Classification / Hallucination Detection
 - **Training Dataset:** RagTruth
 - **Language:** English
@@ -74,7 +82,7 @@ print("Predictions:", predictions)
 **Example level results**
-We evaluate our model on the test set of the [RAGTruth](https://aclanthology.org/2024.acl-long.585/) dataset. Our large model, **lettucedetect-large-v1**, achieves an overall F1 score of 79.22%, outperforming prompt-based methods like GPT-4 (63.4%) and encoder-based models like [Luna](https://aclanthology.org/2025.coling-industry.34.pdf) (65.4%). It also surpasses fine-tuned LLAMA-2-13B (78.7%) (presented in [RAGTruth](https://aclanthology.org/2024.acl-long.585/)) and is competitive with the SOTA fine-tuned LLAMA-3-8B (83.9%) (presented in the [RAG-HAT paper](https://aclanthology.org/2024.emnlp-industry.113.pdf)). Overall, **lettucedetect-large-v1** and **lettucedect-base-v1** are very performant models, while being very effective in inference settings.
 The results on the example-level can be seen in the table below.
@@ -84,7 +92,7 @@ The results on the example-level can be seen in the table below.
 **Span-level results**
-At the span level, our model achieves the best scores across all data types, significantly outperforming previous models. The results can be seen in the table below. Note that here we don't compare to models, like [RAG-HAT](https://aclanthology.org/2024.emnlp-industry.113.pdf), since they have no span-level evaluation presented.
 <p align="center">
   <img src="https://github.com/KRLabsOrg/LettuceDetect/blob/main/assets/span_level_lettucedetect.png?raw=true" alt="Span-level Results" width="800"/>

 ---
 base_model:
 - answerdotai/ModernBERT-base
+language:
+- en
+license: mit
+pipeline_tag: question-answering
 tags:
+- token-classification
+- hallucination-detection
 - transformers
+library_name: transformers
 ---
 # LettuceDetect: Hallucination Detection Model
 ## Overview
+LettuceDetect is a transformer-based model for hallucination detection on context and answer pairs, designed for Retrieval-Augmented Generation (RAG) applications. This model is built on **ModernBERT**, which has been specifically chosen and trained because of its extended context support (up to **8192 tokens**). This long-context capability is critical for tasks where detailed and extensive documents need to be processed to accurately determine if an answer is supported by the provided context.
+## Paper
+[LettuceDetect: A Hallucination Detection Framework for RAG Applications](https://hf.co/papers/2502.17125)
+**Abstract:**
+Retrieval Augmented Generation (RAG) systems remain vulnerable to hallucinated answers despite incorporating external knowledge sources. We present LettuceDetect a framework that addresses two critical limitations in existing hallucination detection methods: (1) the context window constraints of traditional encoder-based methods, and (2) the computational inefficiency of LLM based approaches. Building on ModernBERT's extended context capabilities (up to 8k tokens) and trained on the RAGTruth benchmark dataset, our approach outperforms all previous encoder-based models and most prompt-based models, while being approximately 30 times smaller than the best models. LettuceDetect is a token-classification model that processes context-question-answer triples, allowing for the identification of unsupported claims at the token level. Evaluations on the RAGTruth corpus demonstrate an F1 score of 79.22% for example-level detection, which is a 14.8% improvement over Luna, the previous state-of-the-art encoder-based architecture. Additionally, the system can process 30 to 60 examples per second on a single GPU, making it more practical for real-world RAG applications.
 ## Model Details
+- **Architecture:** ModernBERT (base) with extended context support (up to 8192 tokens)
 - **Task:** Token Classification / Hallucination Detection
 - **Training Dataset:** RagTruth
 - **Language:** English
 **Example level results**
+The model is evaluated on the test set of the [RAGTruth](https://aclanthology.org/2024.acl-long.585/) dataset.  The large version of this model, lettucedect-large-v1, achieves an overall F1 score of 79.22%, outperforming prompt-based methods like GPT-4 (63.4%) and encoder-based models like [Luna](https://aclanthology.org/2025.coling-industry.34.pdf) (65.4%). It also surpasses fine-tuned LLAMA-2-13B (78.7%) and is competitive with the SOTA fine-tuned LLAMA-3-8B (83.9%).
 The results on the example-level can be seen in the table below.
 **Span-level results**
+At the span level, the large version of this model achieves the best scores across all data types, significantly outperforming previous models. The results can be seen in the table below.
 <p align="center">
   <img src="https://github.com/KRLabsOrg/LettuceDetect/blob/main/assets/span_level_lettucedetect.png?raw=true" alt="Span-level Results" width="800"/>