|
Download docs/ITERATIONS.md from devildasdf/devils-agent: direct link, hf CLI and curl.
- Browser
- Download file 13.3 kB
-
https://huggingface.co/devildasdf/devils-agent/resolve/main/docs/ITERATIONS.md
- Command line
-
hf download hf://devildasdf/devils-agent/docs/ITERATIONS.md
-
curl -L -o ITERATIONS.md https://huggingface.co/devildasdf/devils-agent/resolve/main/docs/ITERATIONS.md
13.3 kB
| # Measured iterations | |
| ## 1 — Compute structural paths once during traversal | |
| Hypothesis: repeated sibling enumeration makes observation quadratic in the | |
| number of siblings. Computing paths during traversal should lower large-page | |
| latency without changing model-visible state. | |
| Alternatives considered: omit structural paths (faster but weakens cache matching | |
| and stale-target checks); cache snapshots across mutations (potentially larger | |
| gain but requires reliable invalidation); compute paths once (small implementation | |
| cost, linear extra storage, no intended semantic change). Chose the third. | |
| Change: DOM traversal returns paths alongside retained node handles. Single-target | |
| revalidation still recomputes its path. The original path recomputation remains | |
| available through `--path-mode naive` for future ablations. | |
| Measured on Windows Ryzen 7 7435HS, unrestricted CPU, 20 observations/clicks per | |
| page size. Reports: `reports/runtime-baseline.json` and | |
| `reports/runtime-paths-v1.json`. These are runtime fixture microbenchmarks, not | |
| learned-policy task success or EPYC VPS results. | |
| | Elements | Before observation median | After observation median | Before click median | After click median | | |
| |---|---:|---:|---:|---:| | |
| | 40 | 15.01 ms | 28.97 ms | 64.33 ms | 57.04 ms | | |
| | 200 | 41.18 ms | 43.16 ms | 55.38 ms | 55.33 ms | | |
| | 2000 | 1134.41 ms | 254.24 ms | 70.84 ms | 63.70 ms | | |
| Large-page observation improved 4.46x in these runs. Small-page observation | |
| regressed by 13.96 ms; added handle processing and host variability are possible | |
| contributors, not established causes. No claim of statistical significance. | |
| Observed process-tree RSS peaks were 600,313,856 and 581,799,936 bytes respectively; | |
| these are 10 ms samples and sum shared pages, not exact unique peak memory. | |
| Regression analysis: the initial change broke target revalidation because | |
| Playwright supplied null rather than undefined for an omitted argument. Tests | |
| caught it before accepting the experiment. The null handling was fixed; all | |
| 15 then-existing tests passed before the after benchmark. An additional equality | |
| test now checks optimized versus recomputed paths, including frames and shadow DOM. | |
| Decision: provisionally retain the change for the substantial large-page gain; | |
| the absolute small-page penalty is below the one-second action target but merits | |
| further profiling. Do not declare the architecture mature after this iteration. | |
| Next bottleneck: 2,000 elements still serialize to about 271 KB. Candidate reduction, | |
| compact model-facing fields and measured tokenizer costs are next. Path output | |
| alone is not adequate context reduction, and raw DOM collection remains unbounded. | |
| ## Common evidence for iterations 2–5 | |
| All models use the same 2,400 generated training rows, 480 validation rows, seed | |
| 1729, 16 epochs, AdamW, and two PyTorch CPU threads. Checkpoints are selected on | |
| validation joint accuracy, earliest epoch on a tie. Each also has 480-row | |
| familiar-wording and novel-wording predictions. The inference comparisons run in | |
| separate sequential processes. This is one seed, not a statistical study. | |
| | Configuration | Parameters before quantization | Familiar joint | Novel joint | p95 policy time | Observed Python RSS | Artifact bytes | | |
| |---|---:|---:|---:|---:|---:|---:| | |
| | Mean + lexical FP32 | 128,900 | 100% | 55.83% | 2.847 ms | 275,996,672 | 516,312 | | |
| | Mean, lexical features zeroed | 128,900 | 99.79% | 50.63% | 2.787 ms | 276,901,888 | 516,312 | | |
| | GRU + lexical FP32 | 153,860 | 100% | 35.83% | 6.874 ms | 297,365,504 | 616,488 | | |
| | Transformer + lexical FP32 | 163,908 | 100% | 42.71% | 5.411 ms | 373,432,320 | 657,592 | | |
| | Mean + lexical, Linear INT8 | 128,900 | 100% | 55.62% | 4.170 ms | 295,170,048 | 459,286 | | |
| Policy time includes feature encoding and neural inference, not browser work. | |
| RSS includes Python, training-library imports and evaluation tensors; it is not | |
| browser-inclusive service memory or exact peak RSS. These are Windows Ryzen | |
| measurements, not EPYC VPS results. Reports are `reports/bench-*.json`; | |
| `reports/model-comparison.json` contains the exact diagnostic scoring formula. | |
| Actual 120-fixture novel-wording Chromium evaluations, with confidence abstention: | |
| | Policy | Successful | Abstained | Median policy latency | | |
| |---|---:|---:|---:| | |
| | Mean + lexical | 74/120 | 19/120 | 2.642 ms | | |
| | Mean without lexical | 70/120 | 21/120 | 2.590 ms | | |
| | GRU | 47/120 | 48/120 | 5.643 ms | | |
| | Transformer | 59/120 | 18/120 | 4.209 ms | | |
| | Deterministic lexical/role comparator | 120/120 | 0/120 | 0.654 ms | | |
| Reports are `reports/policy-browser-novel-v*.json` and | |
| `reports/baseline-browser-novel.json`. Browser fixtures use independent outcome | |
| checks; model inputs do not include expected indices. The comparator's perfect | |
| score exposes a dataset shortcut. It is a baseline, not the final architecture. | |
| ## 2 — Remove lexical overlap features | |
| Hypothesis: learned embeddings might replace deterministic lexical overlap, | |
| simplifying the model inputs. Alternative: retain lexical features as a cheap | |
| generalization aid; introduce a pretrained semantic encoder at higher cost. | |
| Expected gain was architectural simplicity, with an uncertain generalization | |
| risk and negligible RAM savings. The ablation zeroes five lexical features; it | |
| does not reduce the parameter count, so it isolates their information contribution. | |
| Before/after: novel joint accuracy fell from 55.83% to 50.63%; browser success fell | |
| from 74/120 to 70/120; p95 policy time changed from 2.847 to 2.787 ms. That small | |
| latency difference is not enough evidence of a real speed benefit. Unit checks | |
| cover candidate filtering and pointer permutation equivariance. | |
| Decision: reject removing lexical features. Next bottleneck: general semantics, | |
| not their small compute cost. | |
| ## 3 — Replace mean encoding with a GRU | |
| Hypothesis: word order and recurrent context may improve grounding. Alternatives | |
| were mean pooling or self-attention. Complexity and serial CPU work increase; | |
| expected accuracy gain was uncertain, and synthetic overfitting was a risk. | |
| Changed both goal and candidate encoders, preserving classifier/pointer heads. | |
| Before/after: novel joint accuracy fell from 55.83% to 35.83%; novel browser success | |
| fell from 74/120 to 47/120; p95 policy time rose from 2.847 to 6.874 ms. Parameters | |
| rose by 24,960 and observed RSS by about 21 MB. Local training took 149.73 seconds | |
| versus 12.39 for mean pooling; host load was not controlled for training timing. | |
| Decision: reject this GRU configuration. Regression: substantially worse | |
| distribution-shift behavior despite perfect familiar-wording accuracy. Next | |
| bottleneck: data diversity and semantic priors; recurrent architecture alone did | |
| not solve it. Padding inefficiency is also a possible optimization target. | |
| ## 4 — Replace mean encoding with a tiny Transformer | |
| Hypothesis: attention could model phrase relationships better than pooling or | |
| recurrence. Estimated tradeoff: more compute and activation memory, uncertain | |
| accuracy benefit. Changed to one 64-wide, four-head encoder layer with positions; | |
| the action/pointer structure stayed unchanged. | |
| Before/after: novel joint accuracy fell from 55.83% to 42.71%; browser success fell | |
| from 74/120 to 59/120; p95 policy time rose from 2.847 to 5.411 ms. Observed Python | |
| RSS increased from 276 MB to 373 MB. Familiar-wording accuracy remained 100%. | |
| All encoder variants pass finite-output/padding tests. | |
| Decision: reject this scratch-trained Transformer configuration. This does not | |
| reject pretrained Transformers, larger models, or attention generally. Next | |
| bottleneck: pretraining/data/task representation rather than choosing an | |
| architecture by popularity. | |
| ## 5 — Dynamic INT8 Linear quantization | |
| Hypothesis: quantized matrix operations might reduce disk/RAM and latency. | |
| Alternatives are FP32, embedding quantization, or ONNX/other CPU runtimes. Expected | |
| benefit was uncertain for such small layers; dispatch overhead can dominate. | |
| Changed Linear layers only; embeddings remain FP32. The serialized INT8 state | |
| was reloaded with `weights_only=True` and re-evaluated. | |
| Before/after: artifact size fell from 516,312 to 459,286 bytes, but p95 policy time | |
| rose from 2.847 to 4.170 ms. Neural-forward median rose from 0.491 to 1.220 ms. | |
| Novel joint accuracy fell from 55.83% to 55.62%. Familiar-wording Chromium success | |
| was 120/120 for both FP32 and INT8. Observed Python RSS was higher for INT8, so the | |
| smaller file is not evidence of runtime RAM reduction. | |
| Decision: retain the export and report, reject INT8 as the preferred runtime for | |
| this model. Next bottleneck: encoding/observation costs and model quality; revisit | |
| quantization for larger models and alternative kernels. PyTorch emitted API | |
| deprecation warnings, recorded in the benchmark log. | |
| ## Maturity decision and real-data diagnostic | |
| These five measured engineering experiments **do not establish a mature design**. | |
| The synthetic suite is inadequate as a generalization acceptance test. A read-only | |
| Mind2Web diagnostic over nine tasks from three websites produced only 1/46 correct | |
| scorable action/target pairs and 22/46 candidate recall. Three steps had no usable | |
| positive label and are reported separately. This is a pinned training shard, not | |
| the official benchmark, and no source HTML was executed. See | |
| `reports/mind2web-smoke-v000.json` for provenance and preprocessing limitations. | |
| Next cycle must improve real candidate coverage, preserve task/history semantics, | |
| add multi-step/hard-negative data, investigate pretrained small encoders/language | |
| models, and reserve a fresh untouched audit set. A simpler deterministic baseline | |
| currently beats every learned model on the easy fixtures. No superiority or | |
| production-readiness claim is justified. All 22 regression tests pass; four model | |
| versions are registered and the production champion remains null. | |
| ## 6 — Broad lexical retrieval and pretrained semantic reranking | |
| Hypothesis: role filtering and lack of semantic priors contribute to the real-data | |
| failure. Alternatives: expand the original role filter, use BM25 over the annotated | |
| candidate pool, or rerank that pool using a pretrained cross-encoder. The last | |
| option costs substantially more CPU; all are measured before changing runtime | |
| defaults. Labels remain evaluator-only and previous-action history excludes the | |
| current action. Two focused tests verify these boundaries. | |
| Audit: original role eligibility excluded 11/46 targets even without a top-k | |
| limit (10 generic, one tab). Original recall was 22/46 at 40 and 30/46 at 80. | |
| Broad BM25 recall was 20/46 at 40, 26/46 at 80 and 37/46 at 200. Removing the role | |
| exclusion therefore did not automatically improve recall at the relevant budget. | |
| Change: 22,713,601-parameter pretrained MiniLM cross-encoder reranks 80 BM25 | |
| candidates with either the full goal or goal plus up to four prior human actions. | |
| Top-one target accuracy was 4/46 for BM25, 5/46 for semantic goal only, and 6/46 | |
| with history. Goal-only semantic top-40 recall was 25/46, versus 20/46 for BM25 | |
| top-40; the semantic path initially sees 80 candidates. History reduced semantic | |
| top-40 recall to 24/46. Median rerank time over both settings was 1381.79 ms. | |
| Regression/decision: retain diagnostic tooling, do not replace production | |
| retrieval with this configuration. The small top-one gain is insufficient for | |
| the CPU cost and target coverage remains poor. No checkpoint was trained on this | |
| source. This is a 46-step training-shard diagnostic, not statistical or official | |
| benchmark evidence. Report: `reports/retrieval-audit.json`. | |
| Next bottleneck: goal-to-next-step decomposition, history representation, better | |
| accessible names and supervised real-world grounding. Naively appending previous | |
| actions to a passage-ranking query is not a learned planner. | |
| ## 7 — Qwen 0.5B autoregressive smoke baseline | |
| Hypothesis: a pretrained instruction model may provide task interpretation missing | |
| from the synthetic-trained policies. Change: pinned Qwen2.5-0.5B-Instruct in local | |
| FP32 CPU inference, two torch threads, 20 BM25 candidates, up to four prior human | |
| actions, greedy generation capped at 48 tokens. This is an offline eight-step | |
| smoke test; outputs never reach the browser executor. | |
| Before: synthetic-trained policy matched 1/46 action/target pairs in the larger | |
| diagnostic. After: Qwen returned zero valid DSL actions and zero correct pairs in | |
| the first eight steps; only five of those targets were in its candidate pool. | |
| Different sample counts prohibit a direct quality ranking. Input lengths were | |
| 656–887 tokens, output lengths 5–9 tokens. Median observed generation latency was | |
| about 6.81 seconds, and observed Python RSS was approximately 2.5 GB. Regression | |
| tests ran concurrently during part of this smoke test, so latency is indicative, | |
| not an isolated comparative benchmark. No target-VPS claim is made. | |
| Decision: reject this unconstrained configuration as a fallback. Next bottleneck: | |
| grammar-constrained decoding and better prompt/task representation, followed by | |
| an isolated rerun and quantized CPU backend comparison. Invalid outputs were | |
| rejected rather than heuristically repaired or executed. Report: | |
| `reports/qwen-baseline.json`. This result does not establish that conventional | |
| LLMs are unnecessary; it establishes that this particular naive integration fails. | |