AutonomousAgent / README.md
Claude
Claude Opus 5.5
Knowledge Graph: figures with pictures, the review queue marked, where each value was found
3d34953 unverified
|
Raw History Blame Contribute Delete
28.2 kB
---
title: AIM Autonomous Ingestion Agent
emoji: πŸ€–
colorFrom: blue
colorTo: green
sdk: docker
app_port: 7860
pinned: true
short_description: Autonomous DOI-deduped literature harvesting for AIM DB
---
# AIM Autonomous Ingestion Agent
An autonomous agent that continuously grows the **AIM Composites Materials
Database** β€” the same Postgres the
[MaterialsDatabase Space](https://huggingface.co/spaces/aim4composites/MaterialsDatabase)
serves β€” by harvesting open-access literature **with DOI provenance**:
```
plan ─▢ discover ─▢ dedupe ─▢ download ─▢ ingest ─▢ report
(OpenAlex, (by DOI, (validated (Gemini β†’ (metrics,
arXiv, URL, and %PDF, grounding, run log,
Unpaywall) sha256) size caps) unit-norm, query
status) rotation)
```
* **discover** β€” query rotation over an editable frontier (8 builtin intents +
a ~450-query materials Γ— reinforcement Γ— property/process grid seeded at
boot); per query: OpenAlex full-text search (OA gold/hybrid/green), arXiv
(ANDed terms) and Semantic Scholar **bulk** search (up to 1,000 OA-PDF
papers per call); per cycle: one page of the **OpenAlex topic walk**
(1-credit filter calls over the fibre-reinforced-composites topics, cursor
persisted) and, when due, a **manufacturer datasheet** seed crawl.
* **dedupe** β€” a durable `agent_doi_seen` registry: a paper already harvested
(by DOI, then URL, then content sha256) is never downloaded twice.
* **download** β€” polite (rate-limited, robots-aware), magic-byte + PyMuPDF
validated, size-bounded; URL chain per work: repository copies first, then
the publisher PDF, then Unpaywall by DOI, then the landing page's
`citation_pdf_url`.
* **ingest** β€” `extraction.py` (grounded, unit-aware Gemini extraction, prompt
v2.1: every value also carries the processing route of its specimen) via `batch_ingest.process_pdf` with the `pg_mirror` Postgres backend:
every value is grounded in the PDF text, unit-normalized (`value_si`,
`unit_canonical`), classified, deduped at measurement grain, and inserted
with a `status` β€” flagged rows are quarantined for review, never dropped.
* **report** β€” per-cycle metrics in `agent_runs` (dashboards + paper Β§8), event
stream in `agent_events`, optional Gemini query expansion when yields dry up.
The graph runs on LangGraph; the agent loop runs in-process on an APScheduler
interval, so the Space works autonomously while it is up β€” no viewers needed.
## Pages
| Page | What it shows |
|---|---|
| **Dashboard** | agent status pills, KPI tiles, rows/PDFs per cycle charts, latest activity feed, latest DOIs, recent cycles |
| **Run Control** | autonomy on/off, interval, per-cycle budget, source toggles, figure-mining toggles, query frontier editor, manual "run one cycle" |
| **Recently Added** | newest ingested documents with DOI links + verified-row counts |
| **Review Queue** | all `status != 'ok'` rows with origin filter, the figure crops behind figure-estimate rows, CSV export, origin-scoped human promote-to-ok |
| **Runs & Logs** | cycle history (status chips, durations, live-refreshing while a cycle runs) with per-PDF results, event stream grouped per run |
| **Database** | material-first browser (one expander per material with the process type and conditions of each property; πŸ“ˆ marks figure-derived estimates) + raw-table browser (incl. the `figures` provenance table) with filters, pagination, CSV export |
| **Knowledge Graph** | every row of the three material tables as one graph, the review queue included and marked as such: whole-database map, matrix and fiber pairs, single-node view with the measured values; each value opens to its page, quoted sentence, flag reason, figure picture, processing route and date added; rebuilt when the tables change and swapped into the open page; JSON and Neo4j exports |
UI: single light theme + design system in `agent/ui_theme.py` (tokens, CSS,
altair theme, KPI/pill/chip/empty-state components); pages compose those
helpers and contain no styling of their own.
Since 28 Sep the Dashboard carries a "Candidates queued" tile, Run Control a
"Candidate queue" card (toggle, pass interval, low-water mark, per-pass budgets,
queue status by source, "Run a discovery pass now") and Runs & Logs a `Kind`
column (`cycle` / `discovery`) plus a `Queued` counter for passes.
## Deploy (β‰ˆ5 minutes)
1. **Create the Space**: huggingface.co β†’ New Space β†’ owner `aim4composites`,
name e.g. `AutonomousAgent`, SDK **Docker**, visibility your call.
2. **Push this folder** to it:
```bash
git clone https://huggingface.co/spaces/aim4composites/AutonomousAgent
cp -r <this folder>/* <this folder>/.streamlit <this folder>/.gitignore AutonomousAgent/
cd AutonomousAgent && git add -A && git commit -m "Autonomous ingestion agent" && git push
```
(or upload the files through the web UI β€” Files β†’ Add file.)
3. **Set Space secrets** (Settings β†’ Variables and secrets):
| Secret | Value |
|---|---|
| `DB_HOST` / `DB_PORT` / `DB_NAME` / `DB_USER` / `DB_PASSWORD` | same values the MaterialsDatabase Space uses |
| `GEMINI_API_KEY` | a **fresh** Gemini key (rotate the leaked ones!) |
| `AGENT_ADMIN_PASSWORD` | password that unlocks Run Control / promote |
| `CONTACT_EMAIL` *(variable, optional)* | polite-crawling contact for OpenAlex/Unpaywall |
| `OPENALEX_API_KEY` | **required since 2026** β€” OpenAlex bills per call: anonymous = $0.10/day *per IP* (shared with every other Space on the egress IP), a free account key = $1/day β‰ˆ 1,000 searches or 10,000 filter calls. Get it at openalex.org β†’ Settings β†’ API |
| `S2_API_KEY` *(optional)* | dedicated Semantic Scholar quota (the bulk lane works without a key on the shared bucket; a free key is a form away) |
4. **One-time schema check**: the shared Postgres must already have the
hardening columns + `sources` table (`python pg_migrate.py` dry-run, then
`--apply`, from any machine with the DB env set). If the July pipeline runs
already did this, there is nothing to do β€” the agent refuses to write until
the schema is migrated, so it fails safe either way.
5. Open the Space β†’ **Run Control** β†’ unlock admin β†’ *Run one cycle now* to
smoke-test β†’ toggle **Autonomous crawling** ON.
## Process type and conditions per property (prompt 2.1, Oct 2026)
A composite's properties depend on how the part was made, so every property
row carries the processing route of the specimen it was measured on:
| Column | Content |
|---|---|
| `process_type` | route, normalized to `extraction.PROCESS_TYPE_ENUM` (injection molding, compression molding / hot press, autoclave, out-of-autoclave, automated fiber placement / tape laying, thermoforming / stamp forming, filament winding, pultrusion, extrusion / compounding, additive manufacturing, liquid molding, welding / joining, fiber spinning, film / solution casting, electrospinning, heat treatment / annealing, other) |
| `process_name` | the route as the document names it |
| `process_conditions` | processing parameters as printed, `; `-joined (temperature, pressure, hold time, cooling rate, nozzle/bed temperature, print speed, ...) |
| `process_quote`, `process_page` | the verbatim sentence that names the route, and its page |
| `process_status` | `grounded` / `grounded_off_page` (sentence found in the PDF text and every number of the conditions found on digit boundaries), `conditions_unverified` (sentence found, a number is not), `ungrounded` (sentence not found), `unchecked` (no text layer) |
* The model returns the routes once per document (`processes`, ids `P1`,
`P2`, ...) and each property points at one through `process_id`, so the
output grows by one short field per property, not by a repeated paragraph.
The same route at two settings is two routes.
* All six columns are NULL when the document does not say how the material
was made (most vendor datasheets, as-received constituents, values quoted
from other literature). The prompt forbids guessing.
* The route is context, not a value: its check never changes a row's `status`.
An unverified route stays attached and marked (⚠ in the Database page).
* `test_condition` is still the *test* condition (23 Β°C, 50% RH); the
`Processing` section still holds process settings a document reports as
values of their own.
* Schema: the columns are appended to `migrate.EXTRA_COLUMNS`, so the boot
migration adds them to the three material tables (`ADD COLUMN IF NOT
EXISTS`); nothing is run by hand. The dedup grain is unchanged.
* Rows stored before this change (prompt 2.0) keep an empty route. The
source PDFs are deleted after ingest, so filling them needs a re-download
pass; that is not part of this change.
* Pages: Database (Process and Process conditions columns per material, a
Process filter, a "With processing route" tile, the columns in the raw
browser and its CSV), Review Queue (three process columns, also in the CSV),
Dashboard ("Processing routes" card: rows per route, share grounded).
* Tests: `test_process_e2e.py` (parse, grounding, rows, boot migration of an
old database, a stubbed cycle into Postgres, the three pages).
## Knowledge graph (live, Oct 2026)
The **Knowledge Graph** page shows the whole database as one graph and keeps it
current without anyone running anything.
- **What is in it.** `agent/kg.py` reads every row of `Polymers`, `Fibers` and
`Composites_materials` through a server-side cursor (no row limit, the
`image` bytes left out) and every row of `figures`, and builds the graph:
materials, properties, property categories, polymer and fiber families,
manufacturers, source documents (DOI, year and lane from `agent_doi_seen`),
figures and process types. Every measured value keeps its status, page, SI
value, route, the day it was added and the figure it was read off or cites.
- **The review queue is in it.** Rows with `status != 'ok'` are not left out:
each is marked "review queue" with its status and `flag_reason`, and the
switch above the graph shows all values, the verified ones, or only the
queue (counts, map and lists follow the switch). Promoting a row in the
Review Queue page changes the fingerprint, so the graph follows.
- **Where a value was found.** Opening a value (the arrow at the start of its
row) shows the page, the quoted sentence (`source_quote`), the model's
comment, the figure with its picture, the route's own name, quote, page and
grounding status, and the date, model and prompt version of the record.
These fields live in separate detail files of 2,000 values each
(`static/kg/d/<build>/<n>.json.gz`), fetched when a value is opened, so the
main file stays small as the database grows.
- **Figures.** Every harvested figure is a node, linked to its document and to
the values read off it or citing it. The pictures of figures with values
are copied from `figures.image_bytes` to `static/kg/fig/<figure_id>.png` in
the background (20 per query, at most 4 minutes per run of the timer job,
only what is not on disk yet; after a restart this takes a few minutes).
`AGENT_KG_FIGURES` = `linked` (default) | `all` (every stored crop) | `off`.
- **When it rebuilds.** A fingerprint (per table: rows, verified rows, newest
`extracted_at`, rows with a process type; and the number of figures) is
compared with the one the current graph was built from. It is read right
after every cycle, every `AGENT_KG_REFRESH_MINUTES` (default 10; 0 switches
the timer off) by the scheduler job `kg-refresh`, and when someone opens the
page (at most every 20 s). The graph is rebuilt only when the fingerprint
changed, so an idle database costs four small aggregate queries per check.
- **How the open page follows.** The graph is written to `static/kg/`
(`kg-data.json`, `kg-data.json.gz`, `kg-meta.json`) and served by Streamlit
at `/app/static/kg/...` (`enableStaticServing` in `.streamlit/config.toml`;
the Dockerfile makes the folder writable; `AGENT_KG_STATIC_DIR` moves it).
The explorer (`agent/kg_explorer.html`, canvas + d3 from cdnjs) asks for
`kg-meta.json` every `AGENT_KG_POLL_SECONDS` (default 60) and swaps the new
graph in, keeping the open node and the zoom. Each build has its own detail
folder (the one before is kept for pages still showing it). If the folder
cannot be written, the graph and the value details are embedded in the page
instead, a reload picks up changes, and figure pictures are not shown.
- **What is inferred.** Polymer and fiber families are assigned by pattern
rules, and property names are merged for case, spelling and a short synonym
list. Everything else is the database as stored. The map is computed from
names and counts, not simulated: the same database always gives the same
picture, and it shifts only where families grow.
- **Exports.** The published JSON can be read from outside the Space
(`https://<space host>/app/static/kg/kg-data.json`; `detail.dir` in it
names the folder of the detail files). The page also offers a Neo4j zip
(nodes.csv, relationships.csv, load script; one Measurement node per value
with its quote, flag reason, day added and review-queue mark; Figure nodes).
`python -m agent.kg --out DIR` writes the same files, the figure pictures
and a stand-alone explorer.html from the configured database.
- **Cost.** On the 1 Oct snapshot (34,768 rows, 1,093 figures in the test
copy): build about 3 s, 5.6 MB of JSON, 0.9 MB compressed. At ten times the
rows: build about 20 s, 46 MB / 7 MB, about 0.7 GB of memory at the peak.
No model calls.
## Figure & graph mining (in-cycle)
Every ingested PDF also passes through `figures.py` (the same module the
local pipeline uses): plots and table images are harvested from the PDF's
layout (embedded rasters + vector clusters, junk filters, caption pairing),
classified with **one** Gemini vision call per PDF, and mineable kinds
(`property_plot`, `table_image`) are read for axis ranges and salient values.
**A number read off a graph is an estimate, not a grounded fact.** Figure
rows are inserted as `origin='figure'`, `status='figure_estimate'` β€” they
never publish (`status='ok'`) at insert, whatever the caller does
(`batch_ingest._row_values()` downgrades them), and only a human promote in
the Review Queue can bless one. The crop PNG itself is stored in the
`figures` table (`image_bytes`, capped at ~1.5 MB) so the Review Queue can
show the reviewer the exact plot behind every estimate β€” the Space's disk is
ephemeral, the database is not.
Knobs (Run Control, or env for first-boot defaults): `AGENT_FIGURES` (on by
default), `AGENT_FIGURE_MINING` (off = harvest+classify only),
`AGENT_MAX_FIGURES` (per-PDF cap, default 12, hard cap 20). Per-run counters
(`figures_found`, `figures_mined`, `figure_rows`, `vision_calls`) land in
`agent_runs` for the Runs & Logs table and the paper's cycle accounting.
## Operations notes
* **Budgets**: per-cycle caps (default 5 PDFs) bound Gemini spend; a hard
ceiling of 25 is compiled in. Cycle lock is durable and self-expiring, so
overlapping schedules or container restarts can't double-run.
* **Keeps running on its own**: the container entry point (`run_space.py`)
boots the agent (schema check, stale-run sweep, scheduler) at process
start, so cycles run from the moment the container is up, with or without
a visitor. A keep-alive thread GETs the Space's own URL
(`https://$SPACE_HOST/_stcore/health`) every `AGENT_KEEPALIVE_MINUTES`
(default 10; 0 disables; `AGENT_KEEPALIVE_URL` overrides) so a free Space
never reaches its 48-h idle sleep. Keep an external hourly pinger as a
backup: if the container ever does sleep, any request wakes it and the
agent boots itself.
* **Watchdog**: every cycle runs under a wall-clock budget
(`AGENT_CYCLE_TIMEOUT_MINUTES`, default 60). A cycle still running past it
is marked `timeout` in Runs & Logs, its lock is released and the next
scheduled cycle proceeds. Without this, one cycle stuck in a network wait
blocked the scheduler for good (the single-instance job never returned),
which is what the long runs of `stale` cycles in August/September were.
The crawler's `Retry-After` sleeps are capped at 300 s (the cap had been
lost in the Aug 25 crawler sync), so that wait should no longer happen.
* **Stall detector (22 Sep)**: besides the hard budget, a cycle whose last
agent event is older than `AGENT_CYCLE_STALL_MINUTES` (default 15) is
abandoned with the stall written into its verdict; a slow but reporting
cycle is bound only by the hard budget. `heartbeat_at` on `agent_runs` is
the liveness signal (every event touches it).
* **Budgeted sources (22 Sep)**: a 429 whose `Retry-After` exceeds the 300 s
cap, or that reports a zero remaining budget (OpenAlex's per-request
billing), closes that host for its window with no retry and no sleep; the
lane is skipped with one warn event per cycle and the reason lands in
`agent_runs.sources`. Set the Space secret `OPENALEX_API_KEY` (a free
account key is ten times the shared anonymous allowance). `AGENT_USE_NTRS`
(default on) adds the NASA Technical Reports Server lane: free, no key,
direct PDFs.
* **Cost column (22 Sep)**: Gemini `usageMetadata` is summed per cycle into
`tokens_in` / `tokens_out` (text + vision calls); Runs & Logs derives
dollars at read time from `AGENT_PRICE_IN_PER_M` / `AGENT_PRICE_OUT_PER_M`
(defaults 0.30 / 2.50 USD per million, Gemini 2.5 Flash list price).
* **Figure citation linking (22 Sep)**: `AGENT_LINK_FIGURES` (default on,
Run Control toggle) links each text row to the figure its evidence cites
and embeds the PNG in the row's `image` column (no S3 needed). Adds three
columns (`figure_ref`, `figure_link_score`, `figure_link_signals`) that
the boot migration creates; `rows_linked` is counted per cycle.
* **Frontier rotation on dry cycles (28 Sep)**: every query a cycle searched
is marked used in `node_report`, whether or not the cycle downloaded
anything. Until then the bookkeeping sat behind `node_ingest`'s early
return, so a dry cycle left `last_used_at` untouched and `next_queries`
(LRU, never-used first) handed the same 20 queries back every half hour:
from 22 to 28 Sep every cycle rediscovered the identical 109 exhausted
candidates (42 DOIs already registered, 66 URLs already final) while the
Gemini-expanded queries piled up unreached. Symptom to watch for in
Runs & Logs: byte-identical Candidates / Skipped / URL seen columns across
consecutive cycles. The harvest loop also no longer fetches (and marks) a
frontier slice past its last round.
* **Downloads that never got ingested (28 Sep)**: a cycle that finds the
material schema unmigrated ends as `error` (not `crashed`) and keeps its
PDFs and their `downloaded` registry rows; the next cycle adopts and
ingests whatever is still on disk (`adopted N PDF(s)` event). At boot,
`downloaded` rows older than 3 h whose file is gone are deleted and their
URLs re-opened in the crawler seen-set (`boot repair: re-queued ...`), so
discovery finds and fetches those papers again β€” this is what frees the
five PDFs runs #135/#139 lost on 15–16 Sep. The Inspect-run panel now
shows why a run ended (crash traceback, watchdog note or cycle error).
* **More papers (28 Sep)**: discovery lanes on by default β€” OpenAlex phrase
search *and* topic walk (`AGENT_OPENALEX_TOPICS`, 1-credit filter calls with
persisted cursors, round-robin), Semantic Scholar bulk (≀1,000 OA-PDF papers
per call; `S2_API_KEY` optional), arXiv with ANDed terms, NASA NTRS, and a
datasheet round over the manufacturer seed pages (each seed ≀ once a day);
the query frontier is seeded once with the 455-intent grid
(`agent/query_grid.py`, origin `grid`). Downloads try repository copies
before publisher URLs and fall back to the landing page's
`citation_pdf_url`; records without an abstract get relevance-gate relief.
Note that publisher hosts (MDPI, Elsevier, Wiley, …) 403 any cloud egress
address, so the yield comes from repositories, NTRS, arXiv and datasheets.
* **Parallel ingestion (28 Sep)**: a cycle extracts its PDFs on a thread pool
(`ingest_workers` in Run Control, default 3, hard max
`AGENT_HARD_MAX_INGEST_WORKERS` = 6), one Postgres connection per worker;
counters, events and the registry are written by the cycle thread as each
PDF completes. Run Control also exposes the cycle's target (`Target new
PDFs per cycle`, default 10: below it the cycle takes another frontier
slice, up to `Max rounds per cycle`) β€” together with the 25-PDF cap these
are the throughput knobs. A cycle's OpenAlex cost is 10 credits per query
per round plus 1 per topic page, so 4 queries Γ— 5 rounds Γ— 48 cycles is
β‰ˆ 9,700 credits/day: right at the free key's allowance; prepaid credits
or a longer interval buy headroom.
* **Candidate queue (28 Sep)**: discovery and download are decoupled. A
*discovery pass* (`agent/discovery.py`, an `agent_runs` row of kind
`discovery`) searches the next `discovery_queries` frontier queries on
every lane β€” Semantic Scholar bulk, arXiv, NTRS, UDSpace (UD's DSpace
theses, direct bitstreams), CORE (needs the free `CORE_API_KEY` secret;
inactive without it) and, for the first `discovery_openalex_queries`, the
10-credit OpenAlex phrase search β€” walks `discovery_topic_pages` pages per
OpenAlex topic (1 credit each) and crawls every datasheet seed that is due,
then queues every relevant candidate in `agent_candidates` (a DOI already in
the registry or a URL the crawler has given up on lands as `seen`). It runs
every `discovery_interval_hours` (24), right after any cycle that finds the
queue below `queue_low_water` (150) or drains it to that, at boot when the
queue is empty, and on demand (Run Control β†’ "Run a discovery pass now").
Cycles pop the queue best-score-first (repository-hosted copies rank ahead
of publisher URLs) and download with no search calls at all; they search
inline only when the queue is empty. Every popped row is settled:
`ingested`, `seen` (registry / final URL), back to `queued` after a
transient failure (never twice in one cycle; `failed` after 3 attempts).
Turn it off with the Run Control toggle to get the pre-28 Sep inline
behaviour. OpenAlex spend drops from up to ~9,700 credits/day to a few
hundred per pass.
* **Secrets never reach the database (28 Sep)**: `agdb.redact()` strips
credential query parameters (`?key=`, `&api_key=`, `token=`) and bearer
tokens from every event, run report, registry status and candidate error
before it is stored, and `extraction` raises Gemini errors that name the
endpoint and the API's own detail instead of the keyed URL (requests'
`HTTPError` / urllib3's connection errors quote it). At boot,
`agdb.redact_stored_secrets` scrubs rows written before this existed and
logs `boot repair: redacted credential parameters …` β€” if that event ever
appears, rotate the key it refers to: it was readable on Runs & Logs.
* **Gemini budget refusals (28 Sep)**: a 429 whose body says
`RESOURCE_EXHAUSTED` / "monthly spending cap" is not retried (it cost
14 s per PDF and cleared nothing). The PDF stays on disk with its
registry row `downloaded`, the cycle's remaining PDFs are not attempted,
one warn event says `Gemini refused for budget reasons (…); N PDF(s) kept
on disk for the next cycle`, and the run ends `error` only if nothing was
ingested at all. The next cycle adopts the kept PDFs (or the boot repair
re-queues them after a restart). Raise the cap at ai.studio/spend.
While the refusal lasts the agent is on a **budget hold** (`agent_state
gemini_budget_hold`): cycles download nothing and probe with one kept PDF
(or one fresh download if the disk was wiped); the first successful
extraction lifts the hold and downloads resume next cycle. Adoption of
kept PDFs is capped per cycle at the download cap, so a backlog is worked
off over several cycles.
* **Cost (28 Sep)**: Gemini's thinking tokens are billed as output at 6Γ— the
input price, and with no `thinkingConfig` the model thinks with an
unbounded dynamic budget: a 25-paper cycle cost ~$7.50, $6.69 of it output
(~30k tokens per PDF for a ~4k-token extraction). `gemini_thinking`
(Run Control β†’ Gemini cost; env `GEMINI_THINKING`; default `low`) is
applied to every Gemini call β€” `dynamic` | `low` | `minimal` | `high` |
`off` | an integer budget; a model that rejects the option gets the call
repeated without it and the option is dropped for the process. Figure
*mining* (one thinking vision call per mineable figure, into quarantined
`figure_estimate` rows) is off by default; harvest + classify (one call per
PDF) and figure linking stay on. Both defaults are applied once to a stored
config at boot (`cost defaults applied: …` event) and can be changed back in
Run Control. Expected: roughly a third of the previous bill per paper β€”
watch the Cost column, and rows per paper / flagged share for recall.
* **Cost (30 Sep) β€” real prices, models, near-duplicates**: the cost column
used to price every token at Gemini 2.5 Flash rates ($0.30 / $2.50 per M)
while the agent runs `gemini-3.5-flash` ($1.50 / $9.00): it showed about a
quarter of the bill (runs 544–593: $12.98 shown, ~$48.50 real, ~$0.41 per
PDF on dynamic thinking with mining on). Runs now store `cost_usd`, priced
per model (`agent/config.py` `MODEL_PRICES`), text and vision calls each at
their own model; older runs are re-priced at gemini-3.5-flash in Runs &
Logs, which also shows thinking tokens, the model, $/PDF and a spend line.
Run Control β†’ Gemini cost adds the **text model** (every row is stamped
with it: change it only after `cost_ab.py` on the gold set), a **vision
model** for the figure calls (default: same), and the **near-duplicate
gate** (`agent/textdedupe.py`, default 0.9): a PDF whose text is that
similar to an ingested one β€” regional datasheet variants, a preprint and
its published version β€” is skipped before any Gemini call
(`duplicate_text` in the registry and the Near-dup column).
* **Ephemeral disk**: PDFs live under `/tmp` only until ingested; provenance
(DOI, sha256, title, URL, query) is durable in `agent_doi_seen`.
* **Safety**: agent bookkeeping is in namespaced `agent_*` tables; controls
are password-gated; read-only without the password. At boot the Space
additively migrates the **connected** database (hardening columns incl.
`origin`/`figure_id`, the origin-aware dedup index, the `figures` table) β€”
`aim_agent` is the agent's own DB, so this is on by default. Before
pointing `DB_NAME` at the team's shared database, either agree on the
additive migration or set `AUTO_MIGRATE_MATERIAL_SCHEMA=0` (the cycle then
refuses to write until `pg_migrate.py --apply` is run by hand).
## Repo map
Vendored pipeline (single source of truth, synced from the project repo β€”
the only Space-side deltas are the PG keepalive/statement-timeout hardening
and the Retry-After cap): `extraction.py`, `batch_ingest.py`, `figures.py`,
`pdf_crawler.py`, `pg_mirror.py`, `pg_migrate.py`, `migrate.py`,
`generate_queries.py`. `FIGURES.md` documents the figure stage.
Agent layer (new): `agent/` (orchestrator, scheduler with watchdog, autostart,
keepalive, DB, config, `query_grid`), `run_space.py` (container entry point),
`app.py` + `page_files/` (Streamlit UI), `agent/kg.py` + `agent/kg_explorer.html`
(knowledge graph builder and explorer; `test_kg.py`). Tests: `test_agent_state.py`,
`test_harvest_loop.py`, `test_sources.py` (no DB, no network), `test_e2e.py`,
`test_figures_e2e.py`, `test_keep_running.py`, `test_counters_e2e.py`,
`test_frontier_rotation_e2e.py`, `test_more_papers_e2e.py`, `test_parallel_ingest_e2e.py`,
`test_candidate_queue_e2e.py` (scratch Postgres via `DB_*`; the network sources are
switched off by env), `test_gemini_errors.py` (no DB, no network).