ledgershield-controlbench / docs /exquisite-visual-analysis.md
shreayas's picture
Sync Exquisite training source (2026-04-25T20:51:48.964674+00:00)
ef8ac93 verified
|
Raw
History Blame Contribute Delete
2.69 kB

LedgerShield Exquisite Visual Analysis

This document describes the 56-plot evidence pack generated by training/exquisite/plot_exquisite_training_results.py.

Executive Plots

Plot Caption
01_final_policy_ladder.png Compares every policy on the same held-out LedgerShield split.
02_sft_vs_grpo_grouped_bar.png Shows whether environment-in-the-loop RL improves over pure imitation at the same model size.
03_scaling_law_score_vs_model_size.png Measures whether LedgerShield remains sensitive to model capacity rather than saturating at small-model performance.
04_score_safety_frontier_all_policies.png Plots mean environment score against unsafe-release rate.
05_teacher_gap_closure.png Measures the fraction of the base-to-teacher gap closed by SFT, GRPO, and DPO.
06_exquisite_pipeline_diagram.png Shows the SFT -> self-play -> falsifier reward -> GRPO -> DPO evidence pipeline.

GRPO Dynamics

The GRPO plots track raw reward, smoothed reward, reward variance, KL or beta penalty when logged by TRL, completion length, parse success, unsafe-release rate, certificate score, control satisfaction, and an entropy proxy.

These plots are designed to make RL instability visible rather than hidden.

Self-Play and Falsifier Analysis

The self-play plots show reward spread, best-vs-worst margin, candidate rank heatmaps, falsifier verdicts, parse failures, unsafe-release blocking, certificate failures, and policy-incomplete failures.

This evidence proves that the model is being scored through LedgerShield execution and deterministic audit, not by a static JSONL label alone.

Per-Case and Per-Task Analysis

The per-case plots expose where GRPO improves over SFT and where it hurts. The per-task plots keep the five task families visible: extraction, matching, duplicate detection, BEC triage, and campaign fraud.

Safety, Audit, and Control Metrics

The safety plots track unsafe release, certificate quality, control satisfaction, institutional utility proxy, institutional loss-surface radar, authority distribution, review burn, supplier friction, calibration error, and auditability-vs-score.

Ablations

The ablation plots are generated immediately, but numeric bars remain pending until the corresponding HF ablation runs complete.

Dashboard

training/exquisite/build_exquisite_dashboard.py now renders a multi-section HTML dashboard with:

  • executive metric cards,
  • the full policy matrix,
  • live Hugging Face launch status,
  • GRPO dynamics,
  • self-play and falsifier evidence,
  • safety and audit plots,
  • per-case analysis,
  • ablations,
  • artifact inventory,
  • reproduction commands.