File size: 6,359 Bytes
f995728
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
# TOTEM Studio Skill — Codex Extractor

## Purpose
Use this skill whenever editing, reviewing, testing, or prompting work on `src/codex_extractor.py`.

The Codex Extractor is a developer-only tool. It extracts raw numerical and categorical author-style fingerprint data from legally accessed PDF/text samples. It does not create final Codex fingerprints by itself.

## Core Terminology
- **TOTEM** = manuscript auditing layer.
- **Codex** = author-style comparison and co-creation layer.
- **Codex Extractor** = developer-only extraction tool that produces raw VM metric data.
- **Final Codex Fingerprint** = validated/aggregated author model stored in Codex after QA.

## Current Governance State
- **VM-001 to VM-005 are frozen under Spec v1 calibrated.**
- **Spec v1 controlled validation is closed and accepted.**
- Further VM-001 to VM-005 changes require a formal approved **Spec v1.1 issue**.
- Do not tune VM-004 or any VM-001 to VM-005 rule during normal patching, validation, or sample preparation.
- Validation failures must be classified first as one of:
  - data/sample issue
  - traceability issue
  - runtime issue
  - potential Spec v1.1 candidate
- Only a confirmed, approved Spec v1.1 candidate may lead to VM-001 to VM-005 logic changes.

## Non-Negotiable Rules
1. Never store copyrighted source text in final Codex fingerprints.
2. Store only derived values: numbers, labels, confidence scores, QA flags, provenance metadata, and extractor version.
3. Raw extractor output is not automatically a final fingerprint.
4. QA/debug artefacts may contain source fragments, but they are developer-only and must not enter final fingerprint storage.
5. Preserve existing public function names unless explicitly instructed otherwise:
   - `process_upload()`
   - `extract_fingerprint()`
   - `format_fingerprint_report()`
   - `extract_text_from_file()`
6. Do not redesign the extractor when asked for a patch.
7. Patch only the requested subsystem.

## Patch Discipline
When modifying `codex_extractor.py`:
- Make the smallest possible change.
- Do not alter unrelated VM metrics.
- Do not touch OCR, page filtering, metadata provenance, dialogue, repetition, or workbook output unless the user specifically requests it.
- Do not change VM-006 to VM-013 or VM-024 to VM-028 unless explicitly instructed.
- Do not change Tier 2 handling.
- Do not infer bibliography or works sampled. `Works_Sampled` must reflect user input or the actual uploaded filename only.
- Do not change VM-001 to VM-005 unless the user explicitly approves a formal Spec v1.1 issue.

## VM-001 to VM-005 Standard
VM-001 to VM-005 must follow **TOTEM Codex Rhyme & Prosody Algorithm Specification v1**.

This includes:
- syllable counting with CMUdict first, fallback heuristic second
- token-level syllable confidence
- stress pattern extraction from ARPAbet where available
- dominant metre/foot support metrics
- phonetic-first rhyme detection
- rhyme-tail comparison from the last stressed vowel onward
- fallback handling for unknown, invented, proper-noun, and OCR-damaged words
- weighted rhyme density
- guarded rhyme scheme classification
- `FREE` must not be returned when meaningful rhyme evidence exists

## Spec v1 Controlled Validation Status
Spec v1 controlled validation has been accepted and closed.

Closure facts:
- controlled sample pack validation completed
- final targeted rerun passed
- no runtime issues remained
- no traceability gaps remained
- no extractor logic changes were needed during closure
- no VM-001 to VM-005 tuning was performed
- no Spec v1.1 issue was required

Governance consequence:
- Treat VM-001 to VM-005 as stable under Spec v1.
- Do not reopen VM-001 to VM-005 logic because of a single surprising sample.
- First review sample quality, expected labels, rights/provenance metadata, and debug artefacts.
- Only escalate to Spec v1.1 if the failure is reproducible, traceable, not caused by data/sample construction, and approved by the validation lead.

## Manual Page Override Rule
Manual story-page range must beat automatic filtering.

If `start_page` or `end_page` is provided:
- process only pages inside that inclusive range
- mark pages outside the range as `out_of_range`
- do not auto-skip the first page if it is inside the manual range
- do not apply front-matter exclusion to pages inside the manual range
- record manual override metadata in `page_trace.csv`, `metric_trace.json`, and the output report

If no manual range is provided:
- use automatic page filtering
- raise a QA warning if OCR is used without a manual story-page range

## Required Debug Artefacts
Every extraction run should be able to produce:
- `cleaned_text.txt`
- `page_trace.csv`
- `line_endings.csv`
- `rhyme_pairs_debug.csv`
- `syllable_debug.csv`
- `stress_pattern_debug.csv`
- `unknown_tokens.csv`
- `repetition_matches.csv`
- `prosody_summary.json` where implemented

`prosody_summary.json` must not contain source text.

## QA Expectations
Always test in this order:
1. clean text sample
2. controlled Gruffalo-style rhymed sample
3. scanned/image-only PDF with manual page range
4. scanned/image-only PDF with automatic filtering

Expected behaviour for clear rhymed picture-book verse:
- VM-004 must not return `FREE`
- VM-003 must show meaningful rhyme evidence
- `rhyme_pairs_debug.csv` must prove the classification
- OCR uncertainty must reduce confidence, not hide uncertainty

## Controlled Validation Discipline
For controlled validation or reruns:
- Do not alter sample text during a run.
- Do not alter expected labels during a run.
- Do not tune extractor logic during validation.
- Preserve per-sample artefacts.
- Preserve run-level metadata, reviewer notes, pass/fail ledger, and checklist.
- If a sample fails, classify the failure before proposing any action.
- Prefer data correction, label correction, retirement, or replacement before any Spec v1.1 escalation.

## Codex Prompting Rule
When using ChatGPT Codex or another coding agent, tell it:

> Before editing, read `skills/codex_extractor/SKILL.md`. Obey it. Patch only the requested subsystem. Do not redesign unrelated code. VM-001 to VM-005 are frozen under Spec v1 unless I approve a formal Spec v1.1 issue.

## Escalation Rule
If a requested change conflicts with this skill:
- stop
- explain the conflict
- ask for explicit override before proceeding