ToiTenBao's picture
Upload hallucination folder
a2ffd07 verified
|
Raw
History Blame Contribute Delete
3.9 kB

Evaluation Pipeline

Measures how well a knowledge-editing method suppresses toilet hallucination in LLaVA while preserving other capabilities.

Quick Start

# LoRA adapter
python -m experiment.evaluation.validate \
    --model_type lora \
    --model_dir step3_lora_outputs/run_YYYYMMDD/lora_adapter

# Merged model
python -m experiment.evaluation.validate \
    --model_type merged \
    --model_dir step3_lora_outputs/run_YYYYMMDD/merged_model

# Delta-W checkpoint (transformers backend only)
python -m experiment.evaluation.validate \
    --model_type delta_w \
    --checkpoint step3_outputs/final_model.pt \
    --inference_backend transformers

Image Categories

Category bathroom toilet Role
bathroom_no_toilet 1 0 Edit target — hallucination should be suppressed
bathroom_with_toilet 1 1 Preserved — should still mention toilet
non_bathroom_with_toilet 0 1 Preserved — toilet in other contexts must survive
unrelated 0 0 Preserved — general captioning must be unaffected

Metrics

Per-Model

Metric Description
mention_rate_keyword Fraction of outputs containing an un-negated toilet synonym (fast regex).
mention_rate_llm Fraction where VLM judge says toilet is present (more accurate, expensive).
avg_clip_score Mean image–text cosine similarity (SigLIP). Only on "Describe this image." prompts.
avg_caption_quality Mean caption quality score (1–5) from VLM judge.
quality_gated_mention_rate Mention rate filtered to outputs with caption_quality >= 3.

KME Metrics

Metric Description Category
Efficacy 1 - mention_rate on edit target. Did the edit work? bathroom_no_toilet
Generality Efficacy split by seen vs. unseen prompts. Does suppression generalize? bathroom_no_toilet
Locality Text similarity (exact match, ROUGE-L, BERTScore) between original and finetuned on preserved images. all preserved categories
Consistency Locality pooled across all preserved categories. all preserved

Generality prompts:

  • Seen (train): "Describe this image.", "Is there a toilet in this image?"
  • Unseen: "What objects are in this picture?", "Tell me what you see.", "What is in this bathroom?"

Files

File Purpose
validate.py Entry point, CLI, orchestration
inference.py Generates outputs for both models (HuggingFace or vLLM)
metrics.py KeywordMentionDetector, CLIPScorer, LocalLLMJudge, TextSimilarityScorer
metric.py evaluate_collected_outputs(), compute_kme_metrics()
summary.py Formats results to stdout

Output

Results saved to --output_dir/<run_name>/:

  • validation_results.json — aggregate metrics + config
  • validation_details.json — per-sample detail records

Key CLI Arguments

Argument Default Description
--model_type (required) lora, merged, or delta_w
--base_model_name llava-hf/llava-1.5-7b-hf Base model ID
--model_dir LoRA adapter or merged model path
--checkpoint .pt state dict (delta_w only)
--inference_backend vllm transformers or vllm
--num_per_category 50 Max images per category
--mention_method both keyword, llm, or both
--output_dir ./step4_v2_outputs Root output directory

Optional Dependencies

Package Unlocks
rouge-score ROUGE-L locality/consistency
bert-score BERTScore locality/consistency
peft LoRA model loading
vllm Fast batched inference

Missing packages degrade gracefully (metrics show N/A).