Spaces:
Running
A newer version of the Gradio SDK is available: 6.22.0
title: scripts/ β operator CLI entrypoints and CI jobs
kind: reference
package: scripts
scripts/ β the command-line surface
The operator/CLI layer: every way a human (or a GitHub Actions job) drives the system by hand β build and refresh the index, generate synthetic index-side content, evaluate, calibrate the guard, smoke-test, and dump inventories for diagnostics.
Why this package exists / its boundary
These scripts are thin wrappers. They own no business logic of their own β
they parse argv, load_dotenv(), fail fast on a missing env var, and then call
into ingest/, index/, and agent/. The real work lives there:
build_index.pycallsingest.discover.discover,ingest.crawl.crawl, andindex.embed.build_indexβ it just sequences them and prints timing.ask.pycallsagent.guard.guardthenagent.loop.answer_agentic.search.pycallsindex.retrieve.retrieve+index.hydrate.hydrate_section.calibrate_guard.pycallsindex.retrieve.top_distance.
The one piece of genuine logic that does live here is the batch-job
scaffolding in generate_glosses.py (git_checkpoint, batching, resume) β
because it is an operational concern (surviving a cancelled CI run), not a
system concern. generate_questions.py imports that scaffolding rather than
duplicating it.
Heavy imports are deliberately done inside main(), after the env-var
check, so a missing NEON_URL prints one clean line instead of a stack trace
from importing a DB module at file scope.
__init__.py is empty β it exists only to make scripts a package so
generate_questions.py can from scripts.generate_glosses import ....
Script catalog
| Script | What it does | Run it | Workflow(s) that invoke it |
|---|---|---|---|
build_index.py |
End-to-end index build: discover β crawl β chunk+embed β Neon. Every stage resumable; every embed batch commits. | python scripts/build_index.py [--skip-crawl] [--skip-embed] [--libraries core,vision] |
build-index.yml (--skip-embed then --skip-crawl), backfill-content.yml (--skip-crawl) |
ask.py |
Answer one question end-to-end from the CLI: guard, then the full agent tool loop; prints answer + citations + referrals. | python scripts/ask.py "how do I build a CNN?" |
β (local only) |
search.py |
Retrieval-only acceptance test: run hybrid retrieve, print the ranked pointers + a snippet of the top hit. No LLM. | python scripts/search.py "scaled_dot_product_attention" [-k 8] [--library vision] |
β (local only) |
calibrate_guard.py |
Recompute the guard's topicality cutoff against the live index: distances for 100 on-topic + 100 off-topic + borderline probes, plus a suggested threshold. Prints only. | python scripts/calibrate_guard.py |
calibrate-guard.yml |
generate_glosses.py |
Batch LLM job: a 1-sentence Contextual-Retrieval gloss per api page β index/glosses.jsonl. Batched, resumable, per-batch flush, optional --push checkpoint. |
python scripts/generate_glosses.py [--limit N] [--batch N] [--push] |
build-index.yml, generate-glosses.yml |
generate_questions.py |
Batch LLM job: a few QuOTE-style hypothetical questions per api page β index/questions.jsonl. Same scaffolding as glosses. |
python scripts/generate_questions.py [--limit N] [--batch N] [--push] |
build-index.yml, generate-glosses.yml |
smoke.py |
Preflight: one Neon write/read, one Gemini call, one local bge-small embedding, one optional Anthropic call. Exits non-zero on any required failure. | python scripts/smoke.py |
β (local preflight) |
smoke_space.py |
Post-deploy health check against the live HF Space: poll runtime until RUNNING, ask a real question via the Gradio API, fail on an LLM/transport error marker. | python scripts/smoke_space.py |
sync-to-hf.yml |
dump_docs_inventory.py |
Dump external ground truth: every documented symbol/page the docs SITE publishes (objects.inv + sitemap) β eval/docs_inventory.jsonl. Needs network to docs.pytorch.org. |
python scripts/dump_docs_inventory.py |
build-index.yml (inventory job) |
dump_index_manifest.py |
Dump internal ground truth: every distinct page actually in the chunks table β eval/index_manifest.jsonl. Needs NEON_URL. |
python scripts/dump_index_manifest.py |
build-index.yml (inventory job) |
coverage_diff.py |
Diff the two dumps: pages the site publishes but our index is missing. A non-empty gap is a pipeline bug. | python scripts/coverage_diff.py |
build-index.yml (inventory job) |
__init__.py |
Empty package marker (enables from scripts.generate_glosses import ...). |
β | β |
Note: the eval jobs (eval.yml) run python -m eval.run_retrieval,
eval.diagnose_retrieval, eval.run_agentic, eval.run_judge β those live in
the eval/ package, not here. ci.yml runs ruff + pytest, and
security.yml runs Trivy; neither invokes a script in this package.
Operational flows
(a) Build / refresh the index
build_index.py is the whole pipeline. Locally you run it plain; in CI the
build-index.yml workflow splits it so glossing can happen between crawl
and embed:
build_index.py --skip-embedβ crawl only, refresh the on-disk_corpussnapshot, never touch Neon (so the crawl-only path doesn't even requireNEON_URL).generate_glosses.py --limit 0 --pushthengenerate_questions.py --limit 0 --pushβ enrich only the pages that don't have a gloss/question set yet.build_index.py --skip-crawlβ embed the snapshot into Neon. A changedglosses.jsonl/questions.jsonlbumps the embed recipe, so only the touched pages re-embed.
Resumability is the design point: crawling skips unchanged pages, embedding
skips chunks whose hash is already in the DB, every embed batch commits. Kill it
and re-run β it continues. Both build-index.yml and backfill-content.yml
share the concurrency: build-index lock so two writers never race the DB.
--libraries core,vision restricts the run to a subset of the seed list.
(b) Generate synthetic index-side content (glosses / questions)
generate_glosses.py and generate_questions.py are the two batch LLM jobs.
Both walk every api-kind page in the snapshot, batch them into LLM calls, and
append JSON lines to index/{glosses,questions}.jsonl, flushing after each
batch. They are resumable: URLs already present in the output file are
skipped, so a rate-limited death just means "run it again." Core-torch pages are
glossed first, so a partial run still covers the pages that matter most. After 5
failed batches they stop, assuming the provider is down. generate_questions.py
imports api_pages, existing_urls_of, and git_checkpoint directly from
generate_glosses.py β one pipeline shape, not two.
The --push flag turns on git_checkpoint: commit and push the jsonl after
every batch. See rationale below.
(c) Evaluate
Retrieval acceptance from the CLI is search.py (pointers only, no LLM) and
end-to-end answering is ask.py. The scored benchmarks
(recall/MRR, agentic, judge) are the eval/ package run via eval.yml, not
scripts here. coverage_diff.py is the eval-adjacent check that the corpus the
index holds actually matches the corpus the docs site publishes.
(d) Calibrate the guard
calibrate_guard.py re-derives the topicality distance cutoff against the live
index. It runs three groups β 100 on-topic (must all pass), 100 off-topic
(should all block), and a handful of borderline/injection probes to eyeball β
through the guard's top_distance path and prints every distance plus a
suggested TORCHDOCS_TOPICALITY_MAX_DISTANCE (the midpoint between the worst
on-topic and best off-topic distance). It only prints; a threshold change is
a policy decision, so a human reads the log and edits the constant in
agent/guard.py by hand. Run it (via calibrate-guard.yml) after any corpus
change big enough to shift the distance distribution β a re-embed or a new doc
set.
(e) Smoke-test β locally and post-deploy
Two different tests for two different moments:
smoke.pyβ before building anything, verify each external connection works: Neon write/read, Gemini, local embedding, optional Anthropic. A missing key skips (Anthropic) or fails with a clear message rather than a traceback.smoke_space.pyβ after deploy, verify the live Space actually answers. It runs in sync-to-hf.yml right after the push (GitHub Actions can reach both the Space and the LLM provider; the dev sandbox can't). It polls the HF runtime API until RUNNING, asks a real question through the Gradio/respondendpoint, and fails the job if the answer contains an LLM/transport error marker β so a broken deploy is a red check, not a Space that silently serves errors. An empty-index answer warns but doesn't fail (that's a separate subsystem).
(f) Inventory / diagnostics
Three scripts build and compare two ground-truth files, all in the inventory job of build-index.yml:
dump_docs_inventory.pyβeval/docs_inventory.jsonlβ what the site publishes (external truth; must run where docs.pytorch.org is reachable).dump_index_manifest.pyβeval/index_manifest.jsonlβ what our index actually holds (internal truth; needsNEON_URL).coverage_diff.pyβ pages in the first but not the second: a page the docs document that our system can never retrieve. Both dumps are committed so the diff can run offline.
Design decisions & rationale
Batch git-checkpoint (git_checkpoint + --push). A gloss/question pass
over the ~3.6K-page corpus is a multi-hour LLM job. The batches are flushed to
disk, but on a GitHub runner that file only reaches the repo via the workflow's
final commit step β so a cancel or a job timeout part-way through throws away
everything generated in the run. --push commits and pushes the jsonl every few
batches instead, so a killed run keeps its progress. Every git failure (unset
identity, a push race with a concurrent enrichment run, a rebase conflict) is
logged and swallowed β a missed checkpoint just defers to the final commit; it
must never kill the long run. The committer identity is injected per-command
(git -c user.name=...) so no global config or extra workflow step is needed,
and the commit message carries [skip ci] so a checkpoint push doesn't kick off
a CI run. It is opt-in (--push) so local runs never commit.
--push off by default. Local runs of the batch jobs should produce a
jsonl and nothing else; only CI, which needs to persist progress across a
possible cancellation, turns pushing on.
Smoke tests exist to fail fast on deploy. smoke.py stops you before an
hour of crawling if a credential is wrong; smoke_space.py turns a broken
deploy into a red check on the very run that produced it (one workflow does both
push and verify). Both do exactly one real round-trip per subsystem β enough to
prove the wire works, cheap enough to run every time.
Guard recalibration is manual after corpus changes. The topicality
threshold is a distance in embedding space; re-embedding or adding a doc set
shifts the whole distribution. calibrate_guard.py measures the new
distribution and suggests a cutoff, but a human commits the constant β where
to draw the on-topic/off-topic line is a policy call, and the script prints the
overlap so you can see when the groups aren't cleanly separable and refuse to
split the difference blindly.
Tool & library choices
argparse+main() -> int+sys.exit(main())everywhere. Exit codes are load-bearing: they make each script a CI gate.smoke*.py,coverage_diff.py, and the batch jobs return non-zero on failure so a workflow step goes red. The batch jobs treat partial success as success (resumable) and total failure as loud (return 0 if written else 1).python-dotenvβ every script that touches a credential callsload_dotenv()first, so a local.envand CI secrets are configured the same way.- Reuse of the real modules, not reimplementation. The scripts import the
exact code the app uses:
agent.guard/agent.loop(ask),index.retrieve(search + calibrate),index.embed/ingest.*(build),agent.llm._raw_completionfor the batch LLM calls. The batch jobs therefore ride the same provider dispatch and fallback chain (OpenRouter/hy3 β Gemini) the workflows configure via env β nothing is stubbed, so a CLI green means the production path is green. gradio_clientinsmoke_space.pyto hit the Space through its real public API, tolerating thehf_tokenβtokenkwarg rename across versions.
File by file
__init__.pyβ empty; makesscriptsan importable package so the two batch jobs can share code.ask.pyβ one question end-to-end from the CLI. Guards the input first (bails with the guard's reason if it fails), then runsanswer_agenticand pretty-prints the answer, citations (title βΊ anchor + URL), and referrals.build_index.pyβ the overnight-safe full pipeline. Fails fast ifNEON_URLis missing (unless--skip-embed, which never touches Neon). Stamps anindex_versionfrom the crawl timestamp.--skip-crawlre-embeds the existing snapshot;--skip-embedrefreshes the snapshot without embedding.calibrate_guard.pyβ reads the 100/100 eval question files plus inline borderline probes, measurestop_distanceper question, prints sorted distances + per-group stats, and suggests a threshold (or flags an overlap). Print-only by design.coverage_diff.pyβ set-difference of two committed jsonl dumps; reports site pages missing from the index, bucketed by library. Returns 1 if either dump is absent. No network, no DB.dump_docs_inventory.pyβ reads each seed's Sphinxobjects.inv(kept roles: classes/functions/methods/attributes/data +std:doc) and sitemap, de-dups, and writes the external ground-truth inventory. Must run where docs.pytorch.org is reachable.dump_index_manifest.pyβ one SQL pass over thechunkstable, rolled up per page (title/kind/library + up to 40 headings + chunk count), written as the internal manifest. NeedsNEON_URL; run from Actions.generate_glosses.pyβ the batch-job home base:api_pages(snapshot β api pages, core first), batched LLM calls,parse_glosses(tolerant JSON extraction),existing_urls_of/resume, andgit_checkpoint/--push. Writesindex/glosses.jsonl.generate_questions.pyβ the QuOTE-style twin. Imports the scaffolding fromgenerate_glosses.pyand only differs in prompt, parser (parse_questions), batch size, and output (index/questions.jsonl).search.pyβ retrieval acceptance test.retrieve(..., debug=True), print ranked pointers, hydrate and snippet the top hit. Returns 1 if the index is empty. No LLM.smoke.pyβ four preflight checks (Neon, Gemini, local embedding, optional Anthropic), each catching its own exception so one broken connection reports instead of crashing the run. Exits non-zero if any required check fails.smoke_space.pyβ post-deploy health check: poll the HF runtime API to RUNNING (or a failure stage), call the Gradio/respondendpoint, and fail on error markers in the answer. Empty-index β warn, not fail.
Related docs
../design-content-and-agent-flow.mdβ the system these scripts drive (pipeline, agent tools, session flow).../deploy-hf-spaces.mdβ the deploy thatsmoke_space.pyverifies.../index/README.mdβembed/retrieve/hydrate, called bybuild_index.py,search.py,calibrate_guard.py.../agent/README.mdβguard/loop/llm, called byask.pyand the batch jobs.