nlp-project / docs /step3_embedding_tuning.md
ervua's picture
Deploy Turkish Legal RAG App
6dfa658
|
Raw
History Blame Contribute Delete
3.54 kB

Step 3 - Embedding Tuning

Goal

The generic multilingual dense retriever underperformed BM25 on the full retrieval benchmark:

Retriever Recall@5 Recall@10 MRR nDCG@10
Dense, multilingual MiniLM 0.613 0.676 0.518 0.556
BM25 0.947 0.975 0.863 0.890

This motivates domain adaptation for the embedding model.

Training Data

embedding.jsonl contains 2,059 training triples:

query
positive_passage
negative_passage

Most negatives are hard negatives from the same source or legal category, which is useful for contrastive training.

Implemented Training Script

The script scripts/train_embedding_model.py fine-tunes a SentenceTransformer bi-encoder with triplet loss.

Recommended command, preferably on GPU:

python scripts/train_embedding_model.py --epochs 1 --batch-size 16 --max-seq-length 384

Evaluation after training:

python scripts/evaluate_retrieval.py --retriever dense --embedding-model outputs/models/legal_embedding_triplet --top-k 10

CPU Smoke Test

A tiny smoke training run was executed locally to verify the training and evaluation pipeline:

python scripts/train_embedding_model.py --limit 64 --epochs 1 --batch-size 8 --max-seq-length 256 --output-dir outputs/models/legal_embedding_triplet_smoke
python scripts/evaluate_retrieval.py --retriever dense --embedding-model outputs/models/legal_embedding_triplet_smoke --limit 50 --top-k 10

Results on the first 50 retrieval queries:

Model Recall@5 Recall@10 MRR nDCG@10
Base dense model 0.820 0.900 0.738 0.776
Smoke fine-tuned model, 64 triples 0.840 0.880 0.725 0.762

The smoke model is not a final optimized model. It only confirms that training, saving, loading, indexing, and evaluation work end to end.

Full CPU Fine-Tuning Run

Full embedding fine-tuning was also executed locally on CPU-only PyTorch. The environment did not expose a CUDA device, so the run used a smaller batch size and sequence length than the preferred GPU configuration.

Training command:

python scripts/train_embedding_model.py --epochs 1 --batch-size 4 --max-seq-length 256 --output-dir outputs/models/legal_embedding_triplet_full_cpu

Training summary:

Setting Value
Training triples 2,059
Epochs 1
Batch size 4
Max sequence length 256
Runtime 2,471 seconds
Training loss 3.257

Evaluation on the full 1,000-query rag_eval.json benchmark:

Dense model Recall@5 Recall@10 MRR nDCG@10
Base multilingual MiniLM 0.613 0.676 0.518 0.556
CPU triplet fine-tuned model 0.539 0.591 0.456 0.488

The naive triplet fine-tuning run degraded dense retrieval quality. This is an important ablation result: fine-tuning is not automatically beneficial, and the embedding objective, hard-negative sampling, sequence length, and validation strategy must be tuned carefully. For the final demo system, BM25 remains the strongest retriever.

Reporting Note

For the final report, this section can be framed as:

  • Generic multilingual embeddings are not enough for Turkish legal retrieval.
  • Hard-negative contrastive training is implemented and reproducible.
  • A full CPU fine-tuning run was completed, but it did not improve dense retrieval.
  • Future work should tune the embedding loss, hard-negative sampling, context length, and base Turkish legal embedding model.