reyden009's picture
Upload folder using huggingface_hub
d4c2896 verified
|
Raw
History Blame Contribute Delete
3.27 kB

AGENTS.md — Rules for agents (and humans) working in this repo

This repository is a research project with a master plan and reproducible protocol. Before touching code or experiments, read this completely.

1. Context in one sentence

Empirical study of modern speculative decoding (EAGLE-3/DFlash/DSpark) on an RTX 4060 Ti 16 GB with Gemma 4 12B and Qwen3-8B GGUF: acceptance by domain, quantization and hardware break-even. The research documentation (thesis, gaps, plan) lives locally in docs/ (not versioned); live status is ALWAYS in STATUS.md (local, not versioned).

2. Conventions

  • Language: docs, logs and commit messages in English; code and technical metadata in English. Do not mix languages within one file.
  • Documentation: markdown and session logs in docs/, live status in STATUS.md — all local, not versioned. The corpus papers/ and its notes notes/ are not versioned (regenerated locally with a download script, outside the repo).
  • Environment: activate with source scripts/spec-env.sh (uv + CUDA in ~/cuda). Never install anything with sudo; if privileges are needed, ask uli and wait.
  • Minimum quality before commit: ruff check, ruff format --check, pyright with no errors on what you touched.

3. Git rules

  • Logical commits (never a mega-commit); concise messages in English with a type prefix (docs:, chore:, feat:, fix:).
  • Do not version: models/, *.gguf, papers/, notes/, experiments/runs/, .venv/, .atl/ (see .gitignore).
  • Do not commit regenerable artifacts (PDFs, raw runs, build logs).
  • Before committing: git status + git diff to review what goes in.

4. Experiments (Phase 1+)

  • Never launch a run without a viable --resume (incremental JSONL) unless explicitly decided.
  • Follow the master plan protocol (§4): fixed seeds (42), documented temperature, target-only baseline with the SAME quantization, config.json + metrics.json per run in experiments/runs/<id>/.
  • Do not use the full DeepSpec data cache (~38 TB); only subsamples.
  • Do not change the master plan targets/decisions (§10) without uli's approval.

5. Sessions and handoff (mandatory)

At the end of any work session:

  1. Write the session log at docs/logs/YYYY-MM-DD-sesion-NN-asunto.md (local logbook, not versioned: what was done, decisions, pending items, what is still running).
  2. Update STATUS.md (local, not versioned: current phase, what is running, pending, next steps).
  3. Save in engram (mem_save/mem_session_summary) non-obvious decisions and discoveries (project: spectulative_decoding).
  4. If background processes remain (downloads, builds, runs): note them in STATUS.md (local) with PID and log.

A new agent opening the repo should read: AGENTS.md → (local STATUS.md, docs and logs) before acting.

6. Forbidden / careful

  • Forbidden: sudo without notice; versioning models/PDFs/runs; using the 38 TB cache; modifying raw results; deleting another session's experiments/runs/.
  • Careful: recompiling llama.cpp without noting the exact commit; changing the overnight/resume regime; discarding speedup < 1 results (they are valid data, not errors).