Islamic3 / README.md
Ghada-99-Ragab's picture
Upload 31 files
c08b36a verified
|
Raw History Blame Contribute Delete
7.89 kB
metadata
title: التحقق من هلوسة القرآن والحديث وتصحيحها
emoji: 🕌
colorFrom: green
colorTo: yellow
sdk: static
app_file: index.html
pinned: false
license: mit
short_description: التحقق من هلوسة القرآن والحديث وتصحيحها

التحقق من هلوسة القرآن والحديث وتصحيحها · Quran & Hadith Hallucination Verifier and Corrector

Arabic version: README_AR.md

Detects Quran and Hadith quotations in any text (for example answers written by language models), verifies each one against the sources, shows the evidence, and proposes a correction taken verbatim from the source text. When the evidence is not sufficient the tool says so and refers the case to human review. The user interface is Arabic.

Live Demo

[LIVE_DEPLOYMENT_URL]

Features

  • Direct verification: paste any paragraph; every Quran / Hadith quotation is found, compared word by word and highlighted, with the exact source text offered as the correction.
  • Ask then Verify: type a question; a pre-configured ChatGPT (OpenAI) client answers and the same engine verifies the answer and returns a corrected version. There is no model selection in this tab.
  • No trigger phrases required: quotations are recognised by matching the text with the Quran and Hadith corpora, with or without quotation marks and with or without openers such as "قال الله تعالى" or "قال رسول الله ﷺ". Such phrases are only a weak fallback for altered quotations that the corpus cannot confirm.
  • Detection model selector (English): Standard · Rules + Corpus, or CAMeLBERT-MSA · Fine-tuned. See below.
  • Benchmark export: JSON and TSV in the IslamicEval 1A / 1B / 1C structure (see Benchmark alignment).
  • Copy corrected text: a working clipboard button with a toast message ("تم نسخ النص بنجاح"), with a fallback for iframes.
  • Fast, informative loading: the Quran index starts downloading before the scripts run, the Hadith index loads in the background, files are cached by the browser (Cache API), verification results and demo examples are cached as JSON in localStorage, and animated skeleton cards plus a three-step progress bar show what is happening.

CAMeLBERT-MSA support

web/camelbert.js exposes one function, detectSpans, that returns BIO-tagged tokens (B-Ayah, I-Ayah, B-Hadith, I-Hadith, O) and spans:

  • Live mode: set HF_MODEL_ID (and HF_TOKEN for a private model) in web/config.js. The page calls the Hugging Face Inference API for a fine-tuned token-classification model. The spans are then verified and corrected by the local pipeline.
  • Simulation mode (default): with no model ID set, no model weights are loaded. The spans come from the standard detector and are shown as BIO tags with deterministic pseudo-confidence values. The interface labels this clearly as Simulation. Use it to demonstrate the flow, not to report model accuracy.

Workflow

text -> detection (quotation marks + corpus match, then corpus scan for unannounced quotes) -> BM25 retrieval
     -> span-local alignment (phonetic skeleton, sliding window, LCS, dynamic gap)
     -> exact: verified | altered Quran passage: source correction | uncertain: human review | nothing similar: no source
Decision Colour Meaning
موثّق (verified) green Equals its source after ignoring spelling and diacritic conventions.
غير مطابق (mismatch) red Differs from the source (Quran: the exact source text is offered) or no source resembles it.
يحتاج مراجعة بشرية (human review) amber Evidence is insufficient or ambiguous. No correction is produced.

Benchmark alignment

benchmark.py converts a pipeline result to the structure of the IslamicEval 2025 subtasks, mirroring the bundled development subset (data/islamiceval_dev_subset.jsonl):

Subtask Field Values
1A detection Span_Start, Span_End, Span_Type character offsets (end exclusive, quotation marks excluded); Ayah / Hadith
1B verification Label Correct / Incorrect (anything not verified is reported as Incorrect)
1C correction Correction exact source text for incorrect spans, or خطأ when no authentic source exists

TSV columns: Response_ID, Span_Start, Span_End, Span_Type, Label, Correction. Check the column names against the organisers' current submission page before an official submission; the official sites could not be fetched automatically while preparing this version, so the exact official column names are not verified here.

Run and deploy

python index_builder.py          # only after changing data/
python build_static_space.py     # writes ./static_space/

Upload the contents of static_space/ to a Hugging Face Space with SDK = Static (free). Local Gradio version: pip install -r requirements.txt && OPENAI_API_KEY=... python app.py. Tests: python -m unittest discover -s tests.

API key (important)

A Static Space has no server, so a key placed in web/config.js is readable by every visitor. This build ships the ChatGPT key in that file for the live demo. Keep a small monthly spend limit on the key in the OpenAI dashboard, revoke it after the evaluation, and never commit it to a public repository (GitHub scans for and revokes exposed OpenAI keys). A safer production design is a small proxy (for example a Cloudflare Worker) that holds the key as a secret.

Project structure

Path Role
web/ Front end: index.html, styles.css (glassmorphism, RTL), app.js (UI), worker.js (Pyodide worker), cache.js, camelbert.js, config.js
normalization.py, idgham.py Phonetic-aware Arabic normalisation; mushaf idgham rendering for corrections
index_builder.py, retrieval.py Pre-built BM25 / n-gram indexes and the lazy-loading retriever
alignment.py, similarity.py Word alignment and similarity indicators
detector.py, scanner.py Quotation detection (corpus-first; trigger phrases optional) and corpus scanning
verifier.py Verification, source-backed correction, decisions
benchmark.py IslamicEval-style JSON / TSV export
llm_client.py, app.py, ui.py LLM client, Gradio app and Python entry points used by the page, HTML rendering
evaluate.py, tests/, research/ Prototype evaluation, tests, original notebook and optional training code

Evaluation (development subset)

Computed with python evaluate.py on the bundled IslamicEval 2025 development subset after decoupling detection from trigger phrases. These numbers are not the earlier research baseline (1A macro F1 0.9091 with CAMeLBERT-MSA; 1B 93.12%; 1C 70.39%).

Component Result
1A detection, character-level macro F1, hybrid detector 0.8247
1B verdict accuracy on gold spans (247 spans; always-"Correct" baseline 59.5%) 93.12%
1C correction accuracy (179 spans; always-"no source" baseline 62.0%) 74.30%

Limitations

  • Unannounced-quotation detection is corpus matching, not understanding; text far from the bundled corpora is missed.
  • The tool aids textual checking; it is not a fatwa and does not replace specialist review.
  • The first visit downloads the Pyodide runtime and about 16 MB of data; later visits use the browser cache.

Competition requirements

See REFERENCES.md for the sources used. Fill in the rules and deadlines from the organisers' documents (https://islamicaich.org/) in your submission; they could not be read automatically while preparing this version.

License: MIT.