--- title: Dialectical Transition Operator emoji: 🔁 colorFrom: indigo colorTo: purple sdk: gradio sdk_version: 5.50.0 app_file: app.py pinned: false short_description: Iterative answer + critique self-improvement (Qwen3-8B, RL) --- # 🔁 Dialectical Transition Operator An interactive demo of a **self-improvement operator**. Given a **question**, a starting **answer**, and **three critiques** of that answer, the model produces a **revised answer** and **three fresh critiques**. Copy the output back into the input and generate again to **iteratively deepen** the answer. **Model:** Qwen3-8B + two stacked LoRA adapters — a frozen *SFT-voice* adapter and an *RL* adapter trained with GRPO / set-VPO under a readability- and addressability-aware LLM judge. ⚠️ **Research demo.** It showcases an iterative self-critique mechanism; it is **not** a reliable source of facts and can produce incorrect or fabricated content.