| title: Dialectical Transition Operator v2 | |
| emoji: π | |
| colorFrom: indigo | |
| colorTo: gray | |
| sdk: gradio | |
| sdk_version: 5.50.0 | |
| app_file: app.py | |
| pinned: false | |
| short_description: Iterative answer+critique self-improvement (graded RL) | |
| # π Dialectical Transition Operator Β· v2 | |
| An interactive demo of a **self-improvement operator**. Given a **question**, a starting **answer**, and | |
| **three critiques** of that answer, the model produces a **revised answer** and **three fresh critiques**. | |
| Copy the output back into the input and generate again to **iteratively deepen** the answer. | |
| **Model:** Qwen3-8B + two stacked LoRA adapters β a frozen *SFT-voice* adapter and an *RL* adapter trained | |
| with GRPO / set-VPO under a **graded (β3β¦+3) self-referential** judge: the operator is rewarded only for a | |
| revision that better addresses the critique than the answer it came from β never against an external answer. | |
| The RL adapter tracks the latest checkpoint of the v2 (graded-reward) run. | |
| β οΈ **Research demo.** It showcases an iterative self-critique mechanism; it is **not** a reliable source of | |
| facts and can produce incorrect or fabricated content. | |