|
Download docs/ACCEPTANCE.md from devildasdf/devils-agent: direct link, hf CLI and curl.
- Browser
- Download file 3.64 kB
-
https://huggingface.co/devildasdf/devils-agent/resolve/main/docs/ACCEPTANCE.md
- Command line
-
hf download hf://devildasdf/devils-agent/docs/ACCEPTANCE.md
-
curl -L -o ACCEPTANCE.md https://huggingface.co/devildasdf/devils-agent/resolve/main/docs/ACCEPTANCE.md
3.64 kB
| # Full-scope acceptance tracking | |
| The goal is still active. The following distinguishes partial artifacts from | |
| proof of the requested end state. Nothing here narrows the original requirements. | |
| | Requested deliverable | Current evidence | Remaining proof/work | | |
| |---|---|---| | |
| | 1. Chosen architecture and reasoning | Mean, GRU, Transformer implementations; experiment reports | Final selection on realistic tasks and target hardware | | |
| | 2. Parameter-count target | First trained mean baseline: 128,900 | Evaluate requested 50M–1.5B range and smallest reliable option | | |
| | 3. Model implementation | Learned classifier and pointer | History, planning, argument spans, all action types | | |
| | 4. HF training pipeline | PyTorchModelHubMixin save/load, CPU trainer | Third-party data ingestion, pinned base checkpoints, reproducible larger runs | | |
| | 5. Dataset schema | Synthetic v1 JSONL and manifest | Normalized real trajectories, license/privacy gates | | |
| | 6. Synthetic generator | Five layouts, three action classes | Multi-step, dynamic, modal, scrolling, ordinal and adversarial tasks | | |
| | 7. Trajectory recorder | HMAC metadata SQLite | Reviewed redacted training trajectories and outcome linkage | | |
| | 8. Experience memory | Not implemented | Reusable retrieval with poisoned-data and privacy protection | | |
| | 9. SHA-256 subsystem | Canonical state and signatures | Correct full-state cache identity and validated cache reuse | | |
| | 10. Compact DOM | Semantic rows, frame/shadow handling | Delta integration, bounds, faithful ARIA, value/history representation | | |
| | 11. Action DSL | Strict 18-action parser | Actual tokenizer comparisons where generative policies are evaluated | | |
| | 12. Safe executor | Authority, stale node, role and host policy gates | Network/effect isolation, redirects/downloads, security audit | | |
| | 13. Recovery controller | Loop rejection and observe action | Learned recovery from enumerated browser failures | | |
| | 14. Completion verifier | Host callback bound to task | General evidence-based user-goal verification | | |
| | 15. CPU runtime | Actual local CPU inference | Linux EPYC two-vCPU end-to-end validation | | |
| | 16. Quantized model | Linear INT8 artifact exported, reloaded and benchmarked; slower than FP32 | Broader quantization/runtime comparison and deployment selection | | |
| | 17. Benchmark harness | Browser fixtures and CPU microbenchmarks | Real sites, baselines, full metric set and held-out audit | | |
| | 18. Regression suite | Parser, state, Chromium, model, registry tests | Complete action/security/generalization coverage | | |
| | 19. Model registry | Immutable copies, integrity and promotion/rollback gates | Full evaluator evidence and real champion; none promoted | | |
| | 20. Continual improvement | Offline trainer; promotion gates | Failure mining, curated dataset deltas, replay and automated cycles | | |
| | 21. Reproducible commands | README and module CLIs | Clean Linux install verification and dependency locking | | |
| | 22. Measured CPU/RAM/latency | Local reports with measurement scope | Target VPS and full service peak measurements | | |
| | 23. Known limitations | README, DATASET, STATUS, experiment reports | Keep updated as scope expands | | |
| | 24. Roadmap | STATUS experiment order | Evidence-driven next-generation plan after full evaluation | | |
| Five meaningful iterations are required but are not a substitute for acceptance. | |
| The production promotion gate requires target-hardware evidence and task, | |
| generalization, recovery and security metrics that current microbenchmarks lack. | |
| Registry unit tests use labelled fake fixtures to test transitions; those fixtures | |
| are never registered in the actual project model registry. | |