devils-agent / docs /RUNTIME_V002.md
devildasdf's picture
Optimize CPU feature encoding and add bounded batch inference with measured parity
c7d6933 verified
|
Raw History Blame Contribute Delete
1.71 kB

Runtime v0.0.2

This release optimizes inference code; checkpoint weights and model intelligence are unchanged.

Feature encoding tokenizes each goal once, reuses lexical scores for retrieval, and constructs tensors once per batch instead of repeatedly assigning tiny tensors. Policy conversion no longer recursively copies unused DOM fields.

LearnedPolicy.predict_batch([(goal, state, ticket), ...]) supports at most 64 independent requests. Each returned decision retains its own ticket and observed target references. Confidence abstention, sensitive-target exclusion, and value-copy restrictions remain in effect. Empty batches return an empty list.

Measured on the development Windows CPU with two Torch threads and the mean checkpoint: median single-request prediction 2.103 ms before / 0.884 ms after (2.38x); p95 3.304 / 1.327 ms. A 16-request batch took 6.253 ms median, about 0.391 ms per request. These are synthetic CPU policy timings, not browser-task latency or target Linux VPS results. Loading, network, and browser time are excluded.

All encoded inputs, targets, and candidate maps were bit-identical across 1,440 validation/test/novel-wording examples. Individual action/ticket parity was checked on 160 requests; batch actions matched individual actions on those requests. Novel-wording raw step match remains 55.83%; this release does not fix generalization.

Evidence: reports/runtime-optimization.json. Run python -m unittest discover -s tests for runtime regression tests. The existing baim.bench_policy command can measure the current feature encoder independently. Compare against Hub revision 795f73703a1e33b82636b2167d7d9507581efd4d for the prior implementation.