Spaces:
Sleeping
Sleeping
| # TOTEM Studio Skill — Codex Extractor | |
| ## Purpose | |
| Use this skill whenever editing, reviewing, testing, or prompting work on `src/codex_extractor.py`. | |
| The Codex Extractor is a developer-only tool. It extracts raw numerical and categorical author-style fingerprint data from legally accessed PDF/text samples. It does not create final Codex fingerprints by itself. | |
| ## Core Terminology | |
| - **TOTEM** = manuscript auditing layer. | |
| - **Codex** = author-style comparison and co-creation layer. | |
| - **Codex Extractor** = developer-only extraction tool that produces raw VM metric data. | |
| - **Final Codex Fingerprint** = validated/aggregated author model stored in Codex after QA. | |
| ## Current Governance State | |
| - **VM-001 to VM-005 are frozen under Spec v1 calibrated.** | |
| - **Spec v1 controlled validation is closed and accepted.** | |
| - Further VM-001 to VM-005 changes require a formal approved **Spec v1.1 issue**. | |
| - Do not tune VM-004 or any VM-001 to VM-005 rule during normal patching, validation, or sample preparation. | |
| - Validation failures must be classified first as one of: | |
| - data/sample issue | |
| - traceability issue | |
| - runtime issue | |
| - potential Spec v1.1 candidate | |
| - Only a confirmed, approved Spec v1.1 candidate may lead to VM-001 to VM-005 logic changes. | |
| ## Non-Negotiable Rules | |
| 1. Never store copyrighted source text in final Codex fingerprints. | |
| 2. Store only derived values: numbers, labels, confidence scores, QA flags, provenance metadata, and extractor version. | |
| 3. Raw extractor output is not automatically a final fingerprint. | |
| 4. QA/debug artefacts may contain source fragments, but they are developer-only and must not enter final fingerprint storage. | |
| 5. Preserve existing public function names unless explicitly instructed otherwise: | |
| - `process_upload()` | |
| - `extract_fingerprint()` | |
| - `format_fingerprint_report()` | |
| - `extract_text_from_file()` | |
| 6. Do not redesign the extractor when asked for a patch. | |
| 7. Patch only the requested subsystem. | |
| ## Patch Discipline | |
| When modifying `codex_extractor.py`: | |
| - Make the smallest possible change. | |
| - Do not alter unrelated VM metrics. | |
| - Do not touch OCR, page filtering, metadata provenance, dialogue, repetition, or workbook output unless the user specifically requests it. | |
| - Do not change VM-006 to VM-013 or VM-024 to VM-028 unless explicitly instructed. | |
| - Do not change Tier 2 handling. | |
| - Do not infer bibliography or works sampled. `Works_Sampled` must reflect user input or the actual uploaded filename only. | |
| - Do not change VM-001 to VM-005 unless the user explicitly approves a formal Spec v1.1 issue. | |
| ## VM-001 to VM-005 Standard | |
| VM-001 to VM-005 must follow **TOTEM Codex Rhyme & Prosody Algorithm Specification v1**. | |
| This includes: | |
| - syllable counting with CMUdict first, fallback heuristic second | |
| - token-level syllable confidence | |
| - stress pattern extraction from ARPAbet where available | |
| - dominant metre/foot support metrics | |
| - phonetic-first rhyme detection | |
| - rhyme-tail comparison from the last stressed vowel onward | |
| - fallback handling for unknown, invented, proper-noun, and OCR-damaged words | |
| - weighted rhyme density | |
| - guarded rhyme scheme classification | |
| - `FREE` must not be returned when meaningful rhyme evidence exists | |
| ## Spec v1 Controlled Validation Status | |
| Spec v1 controlled validation has been accepted and closed. | |
| Closure facts: | |
| - controlled sample pack validation completed | |
| - final targeted rerun passed | |
| - no runtime issues remained | |
| - no traceability gaps remained | |
| - no extractor logic changes were needed during closure | |
| - no VM-001 to VM-005 tuning was performed | |
| - no Spec v1.1 issue was required | |
| Governance consequence: | |
| - Treat VM-001 to VM-005 as stable under Spec v1. | |
| - Do not reopen VM-001 to VM-005 logic because of a single surprising sample. | |
| - First review sample quality, expected labels, rights/provenance metadata, and debug artefacts. | |
| - Only escalate to Spec v1.1 if the failure is reproducible, traceable, not caused by data/sample construction, and approved by the validation lead. | |
| ## Manual Page Override Rule | |
| Manual story-page range must beat automatic filtering. | |
| If `start_page` or `end_page` is provided: | |
| - process only pages inside that inclusive range | |
| - mark pages outside the range as `out_of_range` | |
| - do not auto-skip the first page if it is inside the manual range | |
| - do not apply front-matter exclusion to pages inside the manual range | |
| - record manual override metadata in `page_trace.csv`, `metric_trace.json`, and the output report | |
| If no manual range is provided: | |
| - use automatic page filtering | |
| - raise a QA warning if OCR is used without a manual story-page range | |
| ## Required Debug Artefacts | |
| Every extraction run should be able to produce: | |
| - `cleaned_text.txt` | |
| - `page_trace.csv` | |
| - `line_endings.csv` | |
| - `rhyme_pairs_debug.csv` | |
| - `syllable_debug.csv` | |
| - `stress_pattern_debug.csv` | |
| - `unknown_tokens.csv` | |
| - `repetition_matches.csv` | |
| - `prosody_summary.json` where implemented | |
| `prosody_summary.json` must not contain source text. | |
| ## QA Expectations | |
| Always test in this order: | |
| 1. clean text sample | |
| 2. controlled Gruffalo-style rhymed sample | |
| 3. scanned/image-only PDF with manual page range | |
| 4. scanned/image-only PDF with automatic filtering | |
| Expected behaviour for clear rhymed picture-book verse: | |
| - VM-004 must not return `FREE` | |
| - VM-003 must show meaningful rhyme evidence | |
| - `rhyme_pairs_debug.csv` must prove the classification | |
| - OCR uncertainty must reduce confidence, not hide uncertainty | |
| ## Controlled Validation Discipline | |
| For controlled validation or reruns: | |
| - Do not alter sample text during a run. | |
| - Do not alter expected labels during a run. | |
| - Do not tune extractor logic during validation. | |
| - Preserve per-sample artefacts. | |
| - Preserve run-level metadata, reviewer notes, pass/fail ledger, and checklist. | |
| - If a sample fails, classify the failure before proposing any action. | |
| - Prefer data correction, label correction, retirement, or replacement before any Spec v1.1 escalation. | |
| ## Codex Prompting Rule | |
| When using ChatGPT Codex or another coding agent, tell it: | |
| > Before editing, read `skills/codex_extractor/SKILL.md`. Obey it. Patch only the requested subsystem. Do not redesign unrelated code. VM-001 to VM-005 are frozen under Spec v1 unless I approve a formal Spec v1.1 issue. | |
| ## Escalation Rule | |
| If a requested change conflicts with this skill: | |
| - stop | |
| - explain the conflict | |
| - ask for explicit override before proceeding | |