--- title: Dialectical Transition Operator v2 emoji: ๐Ÿ” colorFrom: indigo colorTo: gray sdk: gradio sdk_version: 5.50.0 app_file: app.py pinned: false short_description: Iterative answer+critique self-improvement (graded RL) --- # ๐Ÿ” Dialectical Transition Operator ยท v2 An interactive demo of a **self-improvement operator**. Given a **question**, a starting **answer**, and **three critiques** of that answer, the model produces a **revised answer** and **three fresh critiques**. Copy the output back into the input and generate again to **iteratively deepen** the answer. **Model:** Qwen3-8B + two stacked LoRA adapters โ€” a frozen *SFT-voice* adapter and an *RL* adapter trained with GRPO / set-VPO under a **graded (โˆ’3โ€ฆ+3) self-referential** judge: the operator is rewarded only for a revision that better addresses the critique than the answer it came from โ€” never against an external answer. The RL adapter tracks the latest checkpoint of the v2 (graded-reward) run. โš ๏ธ **Research demo.** It showcases an iterative self-critique mechanism; it is **not** a reliable source of facts and can produce incorrect or fabricated content.