Spaces:
Running
Running
File size: 3,767 Bytes
59027a2 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 | import HtmlEmbed from '../../components/HtmlEmbed.astro';
import Wide from '../../components/Wide.astro';
import Note from '../../components/Note.astro';
import Sidenote from '../../components/Sidenote.astro';
# Compression and Comparison
We compare against nine baselines from [automata learning](https://texonom.com/37bc3c96247d803b8156ee3fcfdd4556) [@oncina1992rpni; @carrasco1994alergia], HMMs [@rabiner1989hmm], process mining, and agent workflow extraction [@wang2024agent_workflow_memory]. All receive only the same positive training sequences, no failure labels.
## How Much Smaller
Our FSMs achieve **15 to 3,036$\times$ compression** over RPNI while replaying held-out traces at fitness of at least 0.997. The ratio grows with trace length and branching: 15$\times$ on WebArena (short web traces, where RPNI succeeds) up to 3,036$\times$ on GUI-Odyssey (long, repetitive mobile-GUI traces, where RPNI's prefix tree explodes to 21,255 states against our 7).
<Wide>
<HtmlEmbed
src="embeds/compression-bars.html"
title="Compression Ratios"
caption="Each dataset's state count on a log scale: ours (dark slate) versus RPNI (light slate). The distance between the two dots is the compression, from 15x on WebArena to 3,036x on GUI-Odyssey, all at replay fitness of at least 0.997."
/>
</Wide>
Compare all eight methods interactively in the [baselines view](https://seongland.com/article/asg/browser?tab=baselines), or watch structure stabilize in the [convergence view](https://seongland.com/article/asg/browser?tab=convergence).
## Convergence and Stability
Replay fitness reaches its plateau well before the training set is exhausted. On SWE-agent it is already at 0.985 within 1% of the training traces and settles at 0.996 by 10%, while the state count keeps inching up as rare command patterns appear. Structured tool-call domains converge fastest: SWE-smith holds 0.9996 from the first 1% of data. Open web and delegation traces take longer, with Mind2Web needing 5% and Who&When 10% of their traces to clear 0.95 fitness.
<Wide>
<HtmlEmbed
src="embeds/convergence-curves.html"
title="Convergence Curves"
caption="Replay fitness as training traces accumulate, for the four datasets with incremental-convergence runs. The dashed marker shows where each first reaches 0.95 fitness, from 1% of the data on the coding datasets to 10% on Who&When."
/>
</Wide>
## Baselines at a Glance
<Wide>
<HtmlEmbed
src="embeds/baseline-heatmap.html"
title="Method Comparison"
caption="State count and fitness across five methods and twelve datasets. Our FSM holds the smallest state count among non-degenerate methods while keeping the highest fitness."
/>
</Wide>
- **RPNI** without negative examples keeps large portions of the prefix tree (382 to 63,897 states) at degraded fitness.
- **Alergia**, the strongest competitor, matches our fitness but uses 1.0 to 6.0$\times$ more states.
- **HMM** matches our state count but produces non-interpretable latent states.
- **EDSM** (evidence-driven state merging) without negatives collapses to a trivial 1-state acceptor.
- **k-Tails** needs a hyperparameter and produces 1.4 to 10$\times$ more states than ours at $k=1$, with state counts exploding past $k=2$.
- **Process mining** miners reach high fitness but precision 0.00 to 0.80, the "flower model" problem where every activity is reachable from every state.
## Precision
The FSM is more than a vocabulary. It rejects every random trace, and at least 99.9% of permuted traces that keep the activity set but scramble the order. Even single-symbol mutations, a substitution or an insertion or an adjacent swap, are blocked 77 to 100% of the time. RPNI, with its thousands of states, accepts 75% of those same permuted traces on WebArena.
|