Running PaaT: Probe as a Tool for Proprioceptive Language Agents 🔀 Watch a model self‑assess risk and decide to proceed or stop
Running 4 CorrSteer: Correlation-Based Steering of Language Models via Sparse Autoencoders 🧭 Steer text generation by clicking transformer layers
Paused Control Reinforcement Learning 🎛 Explore LLM token decisions with feature‑driven visualizations