A newer version of the Gradio SDK is available: 6.22.0
metadata
title: Dialectical Transition Operator
emoji: π
colorFrom: indigo
colorTo: purple
sdk: gradio
sdk_version: 5.50.0
app_file: app.py
pinned: false
short_description: Iterative answer + critique self-improvement (Qwen3-8B, RL)
π Dialectical Transition Operator
An interactive demo of a self-improvement operator. Given a question, a starting answer, and three critiques of that answer, the model produces a revised answer and three fresh critiques. Copy the output back into the input and generate again to iteratively deepen the answer.
Model: Qwen3-8B + two stacked LoRA adapters β a frozen SFT-voice adapter and an RL adapter trained with GRPO / set-VPO under a readability- and addressability-aware LLM judge.
β οΈ Research demo. It showcases an iterative self-critique mechanism; it is not a reliable source of facts and can produce incorrect or fabricated content.