nlp-project / docs /demo_runbook.md
ervua's picture
Deploy Turkish Legal RAG App
6dfa658
|
Raw
History Blame Contribute Delete
3.03 kB

Live Demo Runbook

Goal

The live demo should show that the system is a working RAG pipeline, not only a report:

question -> retrieval -> source-grounded answer -> citation

Start Demo

python scripts/demo_app.py --data-dir data

Open:

http://127.0.0.1:7860

Custom Document Upload

The same browser page has a Custom Document Test panel. Use it when the instructor wants to upload a file during or after the demo:

  1. Choose a .txt, .md, .csv, .json, .jsonl, .docx, or .pdf file.
  2. Write a question about the uploaded file.
  3. Click Ask Uploaded Document.
  4. The app builds a temporary BM25 index over the uploaded file, returns a source-grounded extractive answer, and lists the retrieved uploaded-file chunks as citations.

If the browser app fails for any reason, use the CLI fallback:

python scripts/demo_cli.py --data-dir data --question "Kasten oldurme sucu nedir?"

Optional Generative LLM Demo

The default browser demo uses the reliable extractive grounded generator. To show that a real local LLM path is also connected, run:

python scripts/demo_app.py --data-dir data --answer-mode local_hf --generation-model outputs/models/flan_t5_legal_sft_smoke_512

CLI check:

python scripts/demo_cli.py --data-dir data --question "Kasten oldurme sucu nedir?" --answer-mode local_hf --generation-model outputs/models/flan_t5_legal_sft_smoke_512

Use this as an optional demonstration only. The FLAN-T5 smoke model is CPU fine-tuned and small, so the final answer quality is weaker than the extractive mode.

Recommended Demo Questions

Use these with the included small corpus because they retrieve clean sources:

Kasten oldurme sucu nedir?
Adil yargilanma hakki nasil guvence altina alinir?
Evlilik birligi temelinden sarsilirsa ne olur?

What To Say During Demo

  1. The question is sent to the retriever.
  2. The retriever ranks legal source chunks with BM25.
  3. The answer generator only uses the top retrieved source.
  4. The system prints the citation/source label at the end.
  5. The retrieved source list makes the grounding auditable.

Metrics To Mention

Full benchmark results from the project:

Metric Result
BM25 Recall@10 0.975
QA Token F1 0.799
Top-5 Source Hit 0.908
Citation Accuracy 0.813
Judge Faithfulness 0.858

If Asked About Optimization

  • Dense multilingual MiniLM underperformed BM25.
  • Full CPU embedding fine-tuning was run, but the naive triplet setup degraded dense retrieval, so BM25 stayed in the demo.
  • The pretrained general-domain reranker hurt ranking, but CPU legal-domain fine-tuning improved it strongly.
  • The fine-tuned reranker still did not beat direct BM25 on the full benchmark.
  • LLM/SFT smoke training was run with FLAN-T5-small, but generated citation quality was too weak for the live demo.
  • The optional local_hf demo mode loads the fine-tuned FLAN-T5 checkpoint and applies citation guardrails.