File size: 16,085 Bytes
80ceec5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
---
title: "scripts/ β€” operator CLI entrypoints and CI jobs"
kind: reference
package: scripts
---

# `scripts/` β€” the command-line surface

The operator/CLI layer: every way a human (or a GitHub Actions job) drives the
system by hand β€” build and refresh the index, generate synthetic index-side
content, evaluate, calibrate the guard, smoke-test, and dump inventories for
diagnostics.

## Why this package exists / its boundary

These scripts are **thin wrappers**. They own no business logic of their own β€”
they parse argv, `load_dotenv()`, fail fast on a missing env var, and then call
into `ingest/`, `index/`, and `agent/`. The real work lives there:

- `build_index.py` calls `ingest.discover.discover`, `ingest.crawl.crawl`, and
  `index.embed.build_index` β€” it just sequences them and prints timing.
- `ask.py` calls `agent.guard.guard` then `agent.loop.answer_agentic`.
- `search.py` calls `index.retrieve.retrieve` + `index.hydrate.hydrate_section`.
- `calibrate_guard.py` calls `index.retrieve.top_distance`.

The one piece of genuine logic that *does* live here is the batch-job
scaffolding in `generate_glosses.py` (`git_checkpoint`, batching, resume) β€”
because it is an operational concern (surviving a cancelled CI run), not a
system concern. `generate_questions.py` imports that scaffolding rather than
duplicating it.

Heavy imports are deliberately done **inside `main()`**, after the env-var
check, so a missing `NEON_URL` prints one clean line instead of a stack trace
from importing a DB module at file scope.

`__init__.py` is empty β€” it exists only to make `scripts` a package so
`generate_questions.py` can `from scripts.generate_glosses import ...`.

## Script catalog

| Script | What it does | Run it | Workflow(s) that invoke it |
|---|---|---|---|
| `build_index.py` | End-to-end index build: discover β†’ crawl β†’ chunk+embed β†’ Neon. Every stage resumable; every embed batch commits. | `python scripts/build_index.py [--skip-crawl] [--skip-embed] [--libraries core,vision]` | **build-index.yml** (`--skip-embed` then `--skip-crawl`), **backfill-content.yml** (`--skip-crawl`) |
| `ask.py` | Answer one question end-to-end from the CLI: guard, then the full agent tool loop; prints answer + citations + referrals. | `python scripts/ask.py "how do I build a CNN?"` | β€” (local only) |
| `search.py` | Retrieval-only acceptance test: run hybrid retrieve, print the ranked pointers + a snippet of the top hit. No LLM. | `python scripts/search.py "scaled_dot_product_attention" [-k 8] [--library vision]` | β€” (local only) |
| `calibrate_guard.py` | Recompute the guard's topicality cutoff against the **live** index: distances for 100 on-topic + 100 off-topic + borderline probes, plus a suggested threshold. Prints only. | `python scripts/calibrate_guard.py` | **calibrate-guard.yml** |
| `generate_glosses.py` | Batch LLM job: a 1-sentence Contextual-Retrieval gloss per api page β†’ `index/glosses.jsonl`. Batched, resumable, per-batch flush, optional `--push` checkpoint. | `python scripts/generate_glosses.py [--limit N] [--batch N] [--push]` | **build-index.yml**, **generate-glosses.yml** |
| `generate_questions.py` | Batch LLM job: a few QuOTE-style hypothetical questions per api page β†’ `index/questions.jsonl`. Same scaffolding as glosses. | `python scripts/generate_questions.py [--limit N] [--batch N] [--push]` | **build-index.yml**, **generate-glosses.yml** |
| `smoke.py` | Preflight: one Neon write/read, one Gemini call, one local bge-small embedding, one optional Anthropic call. Exits non-zero on any required failure. | `python scripts/smoke.py` | β€” (local preflight) |
| `smoke_space.py` | Post-deploy health check against the live HF Space: poll runtime until RUNNING, ask a real question via the Gradio API, fail on an LLM/transport error marker. | `python scripts/smoke_space.py` | **sync-to-hf.yml** |
| `dump_docs_inventory.py` | Dump external ground truth: every documented symbol/page the docs SITE publishes (`objects.inv` + sitemap) β†’ `eval/docs_inventory.jsonl`. Needs network to docs.pytorch.org. | `python scripts/dump_docs_inventory.py` | **build-index.yml** (inventory job) |
| `dump_index_manifest.py` | Dump internal ground truth: every distinct page actually in the `chunks` table β†’ `eval/index_manifest.jsonl`. Needs `NEON_URL`. | `python scripts/dump_index_manifest.py` | **build-index.yml** (inventory job) |
| `coverage_diff.py` | Diff the two dumps: pages the site publishes but our index is missing. A non-empty gap is a pipeline bug. | `python scripts/coverage_diff.py` | **build-index.yml** (inventory job) |
| `__init__.py` | Empty package marker (enables `from scripts.generate_glosses import ...`). | β€” | β€” |

Note: the **eval** jobs (`eval.yml`) run `python -m eval.run_retrieval`,
`eval.diagnose_retrieval`, `eval.run_agentic`, `eval.run_judge` β€” those live in
the `eval/` package, not here. `ci.yml` runs `ruff` + `pytest`, and
`security.yml` runs Trivy; neither invokes a script in this package.

## Operational flows

### (a) Build / refresh the index

`build_index.py` is the whole pipeline. Locally you run it plain; in CI the
**build-index.yml** workflow splits it so glossing can happen *between* crawl
and embed:

1. `build_index.py --skip-embed` β€” crawl only, refresh the on-disk `_corpus`
   snapshot, never touch Neon (so the crawl-only path doesn't even require
   `NEON_URL`).
2. `generate_glosses.py --limit 0 --push` then `generate_questions.py --limit 0
   --push` β€” enrich only the pages that don't have a gloss/question set yet.
3. `build_index.py --skip-crawl` β€” embed the snapshot into Neon. A changed
   `glosses.jsonl`/`questions.jsonl` bumps the embed recipe, so only the touched
   pages re-embed.

Resumability is the design point: crawling skips unchanged pages, embedding
skips chunks whose hash is already in the DB, every embed batch commits. Kill it
and re-run β€” it continues. Both **build-index.yml** and **backfill-content.yml**
share the `concurrency: build-index` lock so two writers never race the DB.
`--libraries core,vision` restricts the run to a subset of the seed list.

### (b) Generate synthetic index-side content (glosses / questions)

`generate_glosses.py` and `generate_questions.py` are the two batch LLM jobs.
Both walk every `api`-kind page in the snapshot, batch them into LLM calls, and
append JSON lines to `index/{glosses,questions}.jsonl`, flushing after each
batch. They are **resumable**: URLs already present in the output file are
skipped, so a rate-limited death just means "run it again." Core-torch pages are
glossed first, so a partial run still covers the pages that matter most. After 5
failed batches they stop, assuming the provider is down. `generate_questions.py`
imports `api_pages`, `existing_urls_of`, and `git_checkpoint` directly from
`generate_glosses.py` β€” one pipeline shape, not two.

The `--push` flag turns on `git_checkpoint`: commit **and push** the jsonl after
every batch. See rationale below.

### (c) Evaluate

Retrieval acceptance from the CLI is `search.py` (pointers only, no LLM) and
end-to-end answering is `ask.py`. The scored benchmarks
(recall/MRR, agentic, judge) are the `eval/` package run via `eval.yml`, not
scripts here. `coverage_diff.py` is the eval-adjacent check that the corpus the
index holds actually matches the corpus the docs site publishes.

### (d) Calibrate the guard

`calibrate_guard.py` re-derives the topicality distance cutoff against the live
index. It runs three groups β€” 100 on-topic (must all pass), 100 off-topic
(should all block), and a handful of borderline/injection probes to eyeball β€”
through the guard's `top_distance` path and prints every distance plus a
suggested `TORCHDOCS_TOPICALITY_MAX_DISTANCE` (the midpoint between the worst
on-topic and best off-topic distance). It **only prints**; a threshold change is
a policy decision, so a human reads the log and edits the constant in
`agent/guard.py` by hand. Run it (via **calibrate-guard.yml**) after any corpus
change big enough to shift the distance distribution β€” a re-embed or a new doc
set.

### (e) Smoke-test β€” locally and post-deploy

Two different tests for two different moments:

- `smoke.py` β€” **before building anything**, verify each external connection
  works: Neon write/read, Gemini, local embedding, optional Anthropic. A missing
  key skips (Anthropic) or fails with a clear message rather than a traceback.
- `smoke_space.py` β€” **after deploy**, verify the live Space actually answers.
  It runs in **sync-to-hf.yml** right after the push (GitHub Actions can reach
  both the Space and the LLM provider; the dev sandbox can't). It polls the HF
  runtime API until RUNNING, asks a real question through the Gradio
  `/respond` endpoint, and fails the job if the answer contains an
  LLM/transport error marker β€” so a broken deploy is a red check, not a Space
  that silently serves errors. An empty-index answer warns but doesn't fail
  (that's a separate subsystem).

### (f) Inventory / diagnostics

Three scripts build and compare two ground-truth files, all in the inventory job
of **build-index.yml**:

- `dump_docs_inventory.py` β†’ `eval/docs_inventory.jsonl` β€” what the site
  publishes (external truth; must run where docs.pytorch.org is reachable).
- `dump_index_manifest.py` β†’ `eval/index_manifest.jsonl` β€” what our index
  actually holds (internal truth; needs `NEON_URL`).
- `coverage_diff.py` β€” pages in the first but not the second: a page the docs
  document that our system can never retrieve. Both dumps are committed so the
  diff can run offline.

## Design decisions & rationale

**Batch git-checkpoint (`git_checkpoint` + `--push`).** A gloss/question pass
over the ~3.6K-page corpus is a multi-hour LLM job. The batches are flushed to
disk, but on a GitHub runner that file only reaches the repo via the workflow's
final commit step β€” so a cancel or a job timeout part-way through throws away
everything generated in the run. `--push` commits and pushes the jsonl every few
batches instead, so a killed run keeps its progress. Every git failure (unset
identity, a push race with a concurrent enrichment run, a rebase conflict) is
logged and swallowed β€” a missed checkpoint just defers to the final commit; it
must **never** kill the long run. The committer identity is injected per-command
(`git -c user.name=...`) so no global config or extra workflow step is needed,
and the commit message carries `[skip ci]` so a checkpoint push doesn't kick off
a CI run. It is opt-in (`--push`) so local runs never commit.

**`--push` off by default.** Local runs of the batch jobs should produce a
jsonl and nothing else; only CI, which needs to persist progress across a
possible cancellation, turns pushing on.

**Smoke tests exist to fail fast on deploy.** `smoke.py` stops you before an
hour of crawling if a credential is wrong; `smoke_space.py` turns a broken
deploy into a red check on the very run that produced it (one workflow does both
push and verify). Both do exactly one real round-trip per subsystem β€” enough to
prove the wire works, cheap enough to run every time.

**Guard recalibration is manual after corpus changes.** The topicality
threshold is a distance in embedding space; re-embedding or adding a doc set
shifts the whole distribution. `calibrate_guard.py` measures the new
distribution and *suggests* a cutoff, but a human commits the constant β€” where
to draw the on-topic/off-topic line is a policy call, and the script prints the
overlap so you can see when the groups aren't cleanly separable and refuse to
split the difference blindly.

## Tool & library choices

- **`argparse` + `main() -> int` + `sys.exit(main())`** everywhere. Exit codes
  are load-bearing: they make each script a CI gate. `smoke*.py`,
  `coverage_diff.py`, and the batch jobs return non-zero on failure so a
  workflow step goes red. The batch jobs treat partial success as success
  (resumable) and total failure as loud (`return 0 if written else 1`).
- **`python-dotenv`** β€” every script that touches a credential calls
  `load_dotenv()` first, so a local `.env` and CI secrets are configured the
  same way.
- **Reuse of the real modules, not reimplementation.** The scripts import the
  exact code the app uses: `agent.guard`/`agent.loop` (ask), `index.retrieve`
  (search + calibrate), `index.embed`/`ingest.*` (build), `agent.llm._raw_completion`
  for the batch LLM calls. The batch jobs therefore ride the same provider
  dispatch and fallback chain (OpenRouter/hy3 β†’ Gemini) the workflows configure
  via env β€” nothing is stubbed, so a CLI green means the production path is green.
- **`gradio_client`** in `smoke_space.py` to hit the Space through its real
  public API, tolerating the `hf_token`β†’`token` kwarg rename across versions.

## File by file

- **`__init__.py`** β€” empty; makes `scripts` an importable package so the two
  batch jobs can share code.
- **`ask.py`** β€” one question end-to-end from the CLI. Guards the input first
  (bails with the guard's reason if it fails), then runs `answer_agentic` and
  pretty-prints the answer, citations (title β€Ί anchor + URL), and referrals.
- **`build_index.py`** β€” the overnight-safe full pipeline. Fails fast if
  `NEON_URL` is missing (unless `--skip-embed`, which never touches Neon).
  Stamps an `index_version` from the crawl timestamp. `--skip-crawl` re-embeds
  the existing snapshot; `--skip-embed` refreshes the snapshot without embedding.
- **`calibrate_guard.py`** β€” reads the 100/100 eval question files plus inline
  borderline probes, measures `top_distance` per question, prints sorted
  distances + per-group stats, and suggests a threshold (or flags an overlap).
  Print-only by design.
- **`coverage_diff.py`** β€” set-difference of two committed jsonl dumps; reports
  site pages missing from the index, bucketed by library. Returns 1 if either
  dump is absent. No network, no DB.
- **`dump_docs_inventory.py`** β€” reads each seed's Sphinx `objects.inv` (kept
  roles: classes/functions/methods/attributes/data + `std:doc`) and sitemap,
  de-dups, and writes the external ground-truth inventory. Must run where
  docs.pytorch.org is reachable.
- **`dump_index_manifest.py`** β€” one SQL pass over the `chunks` table, rolled up
  per page (title/kind/library + up to 40 headings + chunk count), written as
  the internal manifest. Needs `NEON_URL`; run from Actions.
- **`generate_glosses.py`** β€” the batch-job home base: `api_pages` (snapshot β†’
  api pages, core first), batched LLM calls, `parse_glosses` (tolerant JSON
  extraction), `existing_urls_of`/resume, and `git_checkpoint`/`--push`. Writes
  `index/glosses.jsonl`.
- **`generate_questions.py`** β€” the QuOTE-style twin. Imports the scaffolding
  from `generate_glosses.py` and only differs in prompt, parser
  (`parse_questions`), batch size, and output (`index/questions.jsonl`).
- **`search.py`** β€” retrieval acceptance test. `retrieve(..., debug=True)`, print
  ranked pointers, hydrate and snippet the top hit. Returns 1 if the index is
  empty. No LLM.
- **`smoke.py`** β€” four preflight checks (Neon, Gemini, local embedding, optional
  Anthropic), each catching its own exception so one broken connection reports
  instead of crashing the run. Exits non-zero if any required check fails.
- **`smoke_space.py`** β€” post-deploy health check: poll the HF runtime API to
  RUNNING (or a failure stage), call the Gradio `/respond` endpoint, and
  fail on error markers in the answer. Empty-index β†’ warn, not fail.

## Related docs

- [`../design-content-and-agent-flow.md`](../design-content-and-agent-flow.md) β€” the system these scripts drive (pipeline, agent tools, session flow).
- [`../deploy-hf-spaces.md`](../deploy-hf-spaces.md) β€” the deploy that `smoke_space.py` verifies.
- [`../index/README.md`](../index/README.md) β€” `embed`/`retrieve`/`hydrate`, called by `build_index.py`, `search.py`, `calibrate_guard.py`.
- [`../agent/README.md`](../agent/README.md) β€” `guard`/`loop`/`llm`, called by `ask.py` and the batch jobs.