Spaces:
Running
Running
Claude
Claude Opus 5.5
Knowledge Graph: figures with pictures, the review queue marked, where each value was found
3d34953 unverified |
Download README.md from aim4composites/AutonomousAgent: direct link, hf CLI and curl.
- Browser
- Download file 28.2 kB
-
https://huggingface.co/spaces/aim4composites/AutonomousAgent/resolve/main/README.md
- Command line
-
hf download hf://spaces/aim4composites/AutonomousAgent/README.md
-
curl -L -o README.md https://huggingface.co/spaces/aim4composites/AutonomousAgent/resolve/main/README.md
28.2 kB
| title: AIM Autonomous Ingestion Agent | |
| emoji: π€ | |
| colorFrom: blue | |
| colorTo: green | |
| sdk: docker | |
| app_port: 7860 | |
| pinned: true | |
| short_description: Autonomous DOI-deduped literature harvesting for AIM DB | |
| # AIM Autonomous Ingestion Agent | |
| An autonomous agent that continuously grows the **AIM Composites Materials | |
| Database** β the same Postgres the | |
| [MaterialsDatabase Space](https://huggingface.co/spaces/aim4composites/MaterialsDatabase) | |
| serves β by harvesting open-access literature **with DOI provenance**: | |
| ``` | |
| plan ββΆ discover ββΆ dedupe ββΆ download ββΆ ingest ββΆ report | |
| (OpenAlex, (by DOI, (validated (Gemini β (metrics, | |
| arXiv, URL, and %PDF, grounding, run log, | |
| Unpaywall) sha256) size caps) unit-norm, query | |
| status) rotation) | |
| ``` | |
| * **discover** β query rotation over an editable frontier (8 builtin intents + | |
| a ~450-query materials Γ reinforcement Γ property/process grid seeded at | |
| boot); per query: OpenAlex full-text search (OA gold/hybrid/green), arXiv | |
| (ANDed terms) and Semantic Scholar **bulk** search (up to 1,000 OA-PDF | |
| papers per call); per cycle: one page of the **OpenAlex topic walk** | |
| (1-credit filter calls over the fibre-reinforced-composites topics, cursor | |
| persisted) and, when due, a **manufacturer datasheet** seed crawl. | |
| * **dedupe** β a durable `agent_doi_seen` registry: a paper already harvested | |
| (by DOI, then URL, then content sha256) is never downloaded twice. | |
| * **download** β polite (rate-limited, robots-aware), magic-byte + PyMuPDF | |
| validated, size-bounded; URL chain per work: repository copies first, then | |
| the publisher PDF, then Unpaywall by DOI, then the landing page's | |
| `citation_pdf_url`. | |
| * **ingest** β `extraction.py` (grounded, unit-aware Gemini extraction, prompt | |
| v2.1: every value also carries the processing route of its specimen) via `batch_ingest.process_pdf` with the `pg_mirror` Postgres backend: | |
| every value is grounded in the PDF text, unit-normalized (`value_si`, | |
| `unit_canonical`), classified, deduped at measurement grain, and inserted | |
| with a `status` β flagged rows are quarantined for review, never dropped. | |
| * **report** β per-cycle metrics in `agent_runs` (dashboards + paper Β§8), event | |
| stream in `agent_events`, optional Gemini query expansion when yields dry up. | |
| The graph runs on LangGraph; the agent loop runs in-process on an APScheduler | |
| interval, so the Space works autonomously while it is up β no viewers needed. | |
| ## Pages | |
| | Page | What it shows | | |
| |---|---| | |
| | **Dashboard** | agent status pills, KPI tiles, rows/PDFs per cycle charts, latest activity feed, latest DOIs, recent cycles | | |
| | **Run Control** | autonomy on/off, interval, per-cycle budget, source toggles, figure-mining toggles, query frontier editor, manual "run one cycle" | | |
| | **Recently Added** | newest ingested documents with DOI links + verified-row counts | | |
| | **Review Queue** | all `status != 'ok'` rows with origin filter, the figure crops behind figure-estimate rows, CSV export, origin-scoped human promote-to-ok | | |
| | **Runs & Logs** | cycle history (status chips, durations, live-refreshing while a cycle runs) with per-PDF results, event stream grouped per run | | |
| | **Database** | material-first browser (one expander per material with the process type and conditions of each property; π marks figure-derived estimates) + raw-table browser (incl. the `figures` provenance table) with filters, pagination, CSV export | | |
| | **Knowledge Graph** | every row of the three material tables as one graph, the review queue included and marked as such: whole-database map, matrix and fiber pairs, single-node view with the measured values; each value opens to its page, quoted sentence, flag reason, figure picture, processing route and date added; rebuilt when the tables change and swapped into the open page; JSON and Neo4j exports | | |
| UI: single light theme + design system in `agent/ui_theme.py` (tokens, CSS, | |
| altair theme, KPI/pill/chip/empty-state components); pages compose those | |
| helpers and contain no styling of their own. | |
| Since 28 Sep the Dashboard carries a "Candidates queued" tile, Run Control a | |
| "Candidate queue" card (toggle, pass interval, low-water mark, per-pass budgets, | |
| queue status by source, "Run a discovery pass now") and Runs & Logs a `Kind` | |
| column (`cycle` / `discovery`) plus a `Queued` counter for passes. | |
| ## Deploy (β5 minutes) | |
| 1. **Create the Space**: huggingface.co β New Space β owner `aim4composites`, | |
| name e.g. `AutonomousAgent`, SDK **Docker**, visibility your call. | |
| 2. **Push this folder** to it: | |
| ```bash | |
| git clone https://huggingface.co/spaces/aim4composites/AutonomousAgent | |
| cp -r <this folder>/* <this folder>/.streamlit <this folder>/.gitignore AutonomousAgent/ | |
| cd AutonomousAgent && git add -A && git commit -m "Autonomous ingestion agent" && git push | |
| ``` | |
| (or upload the files through the web UI β Files β Add file.) | |
| 3. **Set Space secrets** (Settings β Variables and secrets): | |
| | Secret | Value | | |
| |---|---| | |
| | `DB_HOST` / `DB_PORT` / `DB_NAME` / `DB_USER` / `DB_PASSWORD` | same values the MaterialsDatabase Space uses | | |
| | `GEMINI_API_KEY` | a **fresh** Gemini key (rotate the leaked ones!) | | |
| | `AGENT_ADMIN_PASSWORD` | password that unlocks Run Control / promote | | |
| | `CONTACT_EMAIL` *(variable, optional)* | polite-crawling contact for OpenAlex/Unpaywall | | |
| | `OPENALEX_API_KEY` | **required since 2026** β OpenAlex bills per call: anonymous = $0.10/day *per IP* (shared with every other Space on the egress IP), a free account key = $1/day β 1,000 searches or 10,000 filter calls. Get it at openalex.org β Settings β API | | |
| | `S2_API_KEY` *(optional)* | dedicated Semantic Scholar quota (the bulk lane works without a key on the shared bucket; a free key is a form away) | | |
| 4. **One-time schema check**: the shared Postgres must already have the | |
| hardening columns + `sources` table (`python pg_migrate.py` dry-run, then | |
| `--apply`, from any machine with the DB env set). If the July pipeline runs | |
| already did this, there is nothing to do β the agent refuses to write until | |
| the schema is migrated, so it fails safe either way. | |
| 5. Open the Space β **Run Control** β unlock admin β *Run one cycle now* to | |
| smoke-test β toggle **Autonomous crawling** ON. | |
| ## Process type and conditions per property (prompt 2.1, Oct 2026) | |
| A composite's properties depend on how the part was made, so every property | |
| row carries the processing route of the specimen it was measured on: | |
| | Column | Content | | |
| |---|---| | |
| | `process_type` | route, normalized to `extraction.PROCESS_TYPE_ENUM` (injection molding, compression molding / hot press, autoclave, out-of-autoclave, automated fiber placement / tape laying, thermoforming / stamp forming, filament winding, pultrusion, extrusion / compounding, additive manufacturing, liquid molding, welding / joining, fiber spinning, film / solution casting, electrospinning, heat treatment / annealing, other) | | |
| | `process_name` | the route as the document names it | | |
| | `process_conditions` | processing parameters as printed, `; `-joined (temperature, pressure, hold time, cooling rate, nozzle/bed temperature, print speed, ...) | | |
| | `process_quote`, `process_page` | the verbatim sentence that names the route, and its page | | |
| | `process_status` | `grounded` / `grounded_off_page` (sentence found in the PDF text and every number of the conditions found on digit boundaries), `conditions_unverified` (sentence found, a number is not), `ungrounded` (sentence not found), `unchecked` (no text layer) | | |
| * The model returns the routes once per document (`processes`, ids `P1`, | |
| `P2`, ...) and each property points at one through `process_id`, so the | |
| output grows by one short field per property, not by a repeated paragraph. | |
| The same route at two settings is two routes. | |
| * All six columns are NULL when the document does not say how the material | |
| was made (most vendor datasheets, as-received constituents, values quoted | |
| from other literature). The prompt forbids guessing. | |
| * The route is context, not a value: its check never changes a row's `status`. | |
| An unverified route stays attached and marked (β in the Database page). | |
| * `test_condition` is still the *test* condition (23 Β°C, 50% RH); the | |
| `Processing` section still holds process settings a document reports as | |
| values of their own. | |
| * Schema: the columns are appended to `migrate.EXTRA_COLUMNS`, so the boot | |
| migration adds them to the three material tables (`ADD COLUMN IF NOT | |
| EXISTS`); nothing is run by hand. The dedup grain is unchanged. | |
| * Rows stored before this change (prompt 2.0) keep an empty route. The | |
| source PDFs are deleted after ingest, so filling them needs a re-download | |
| pass; that is not part of this change. | |
| * Pages: Database (Process and Process conditions columns per material, a | |
| Process filter, a "With processing route" tile, the columns in the raw | |
| browser and its CSV), Review Queue (three process columns, also in the CSV), | |
| Dashboard ("Processing routes" card: rows per route, share grounded). | |
| * Tests: `test_process_e2e.py` (parse, grounding, rows, boot migration of an | |
| old database, a stubbed cycle into Postgres, the three pages). | |
| ## Knowledge graph (live, Oct 2026) | |
| The **Knowledge Graph** page shows the whole database as one graph and keeps it | |
| current without anyone running anything. | |
| - **What is in it.** `agent/kg.py` reads every row of `Polymers`, `Fibers` and | |
| `Composites_materials` through a server-side cursor (no row limit, the | |
| `image` bytes left out) and every row of `figures`, and builds the graph: | |
| materials, properties, property categories, polymer and fiber families, | |
| manufacturers, source documents (DOI, year and lane from `agent_doi_seen`), | |
| figures and process types. Every measured value keeps its status, page, SI | |
| value, route, the day it was added and the figure it was read off or cites. | |
| - **The review queue is in it.** Rows with `status != 'ok'` are not left out: | |
| each is marked "review queue" with its status and `flag_reason`, and the | |
| switch above the graph shows all values, the verified ones, or only the | |
| queue (counts, map and lists follow the switch). Promoting a row in the | |
| Review Queue page changes the fingerprint, so the graph follows. | |
| - **Where a value was found.** Opening a value (the arrow at the start of its | |
| row) shows the page, the quoted sentence (`source_quote`), the model's | |
| comment, the figure with its picture, the route's own name, quote, page and | |
| grounding status, and the date, model and prompt version of the record. | |
| These fields live in separate detail files of 2,000 values each | |
| (`static/kg/d/<build>/<n>.json.gz`), fetched when a value is opened, so the | |
| main file stays small as the database grows. | |
| - **Figures.** Every harvested figure is a node, linked to its document and to | |
| the values read off it or citing it. The pictures of figures with values | |
| are copied from `figures.image_bytes` to `static/kg/fig/<figure_id>.png` in | |
| the background (20 per query, at most 4 minutes per run of the timer job, | |
| only what is not on disk yet; after a restart this takes a few minutes). | |
| `AGENT_KG_FIGURES` = `linked` (default) | `all` (every stored crop) | `off`. | |
| - **When it rebuilds.** A fingerprint (per table: rows, verified rows, newest | |
| `extracted_at`, rows with a process type; and the number of figures) is | |
| compared with the one the current graph was built from. It is read right | |
| after every cycle, every `AGENT_KG_REFRESH_MINUTES` (default 10; 0 switches | |
| the timer off) by the scheduler job `kg-refresh`, and when someone opens the | |
| page (at most every 20 s). The graph is rebuilt only when the fingerprint | |
| changed, so an idle database costs four small aggregate queries per check. | |
| - **How the open page follows.** The graph is written to `static/kg/` | |
| (`kg-data.json`, `kg-data.json.gz`, `kg-meta.json`) and served by Streamlit | |
| at `/app/static/kg/...` (`enableStaticServing` in `.streamlit/config.toml`; | |
| the Dockerfile makes the folder writable; `AGENT_KG_STATIC_DIR` moves it). | |
| The explorer (`agent/kg_explorer.html`, canvas + d3 from cdnjs) asks for | |
| `kg-meta.json` every `AGENT_KG_POLL_SECONDS` (default 60) and swaps the new | |
| graph in, keeping the open node and the zoom. Each build has its own detail | |
| folder (the one before is kept for pages still showing it). If the folder | |
| cannot be written, the graph and the value details are embedded in the page | |
| instead, a reload picks up changes, and figure pictures are not shown. | |
| - **What is inferred.** Polymer and fiber families are assigned by pattern | |
| rules, and property names are merged for case, spelling and a short synonym | |
| list. Everything else is the database as stored. The map is computed from | |
| names and counts, not simulated: the same database always gives the same | |
| picture, and it shifts only where families grow. | |
| - **Exports.** The published JSON can be read from outside the Space | |
| (`https://<space host>/app/static/kg/kg-data.json`; `detail.dir` in it | |
| names the folder of the detail files). The page also offers a Neo4j zip | |
| (nodes.csv, relationships.csv, load script; one Measurement node per value | |
| with its quote, flag reason, day added and review-queue mark; Figure nodes). | |
| `python -m agent.kg --out DIR` writes the same files, the figure pictures | |
| and a stand-alone explorer.html from the configured database. | |
| - **Cost.** On the 1 Oct snapshot (34,768 rows, 1,093 figures in the test | |
| copy): build about 3 s, 5.6 MB of JSON, 0.9 MB compressed. At ten times the | |
| rows: build about 20 s, 46 MB / 7 MB, about 0.7 GB of memory at the peak. | |
| No model calls. | |
| ## Figure & graph mining (in-cycle) | |
| Every ingested PDF also passes through `figures.py` (the same module the | |
| local pipeline uses): plots and table images are harvested from the PDF's | |
| layout (embedded rasters + vector clusters, junk filters, caption pairing), | |
| classified with **one** Gemini vision call per PDF, and mineable kinds | |
| (`property_plot`, `table_image`) are read for axis ranges and salient values. | |
| **A number read off a graph is an estimate, not a grounded fact.** Figure | |
| rows are inserted as `origin='figure'`, `status='figure_estimate'` β they | |
| never publish (`status='ok'`) at insert, whatever the caller does | |
| (`batch_ingest._row_values()` downgrades them), and only a human promote in | |
| the Review Queue can bless one. The crop PNG itself is stored in the | |
| `figures` table (`image_bytes`, capped at ~1.5 MB) so the Review Queue can | |
| show the reviewer the exact plot behind every estimate β the Space's disk is | |
| ephemeral, the database is not. | |
| Knobs (Run Control, or env for first-boot defaults): `AGENT_FIGURES` (on by | |
| default), `AGENT_FIGURE_MINING` (off = harvest+classify only), | |
| `AGENT_MAX_FIGURES` (per-PDF cap, default 12, hard cap 20). Per-run counters | |
| (`figures_found`, `figures_mined`, `figure_rows`, `vision_calls`) land in | |
| `agent_runs` for the Runs & Logs table and the paper's cycle accounting. | |
| ## Operations notes | |
| * **Budgets**: per-cycle caps (default 5 PDFs) bound Gemini spend; a hard | |
| ceiling of 25 is compiled in. Cycle lock is durable and self-expiring, so | |
| overlapping schedules or container restarts can't double-run. | |
| * **Keeps running on its own**: the container entry point (`run_space.py`) | |
| boots the agent (schema check, stale-run sweep, scheduler) at process | |
| start, so cycles run from the moment the container is up, with or without | |
| a visitor. A keep-alive thread GETs the Space's own URL | |
| (`https://$SPACE_HOST/_stcore/health`) every `AGENT_KEEPALIVE_MINUTES` | |
| (default 10; 0 disables; `AGENT_KEEPALIVE_URL` overrides) so a free Space | |
| never reaches its 48-h idle sleep. Keep an external hourly pinger as a | |
| backup: if the container ever does sleep, any request wakes it and the | |
| agent boots itself. | |
| * **Watchdog**: every cycle runs under a wall-clock budget | |
| (`AGENT_CYCLE_TIMEOUT_MINUTES`, default 60). A cycle still running past it | |
| is marked `timeout` in Runs & Logs, its lock is released and the next | |
| scheduled cycle proceeds. Without this, one cycle stuck in a network wait | |
| blocked the scheduler for good (the single-instance job never returned), | |
| which is what the long runs of `stale` cycles in August/September were. | |
| The crawler's `Retry-After` sleeps are capped at 300 s (the cap had been | |
| lost in the Aug 25 crawler sync), so that wait should no longer happen. | |
| * **Stall detector (22 Sep)**: besides the hard budget, a cycle whose last | |
| agent event is older than `AGENT_CYCLE_STALL_MINUTES` (default 15) is | |
| abandoned with the stall written into its verdict; a slow but reporting | |
| cycle is bound only by the hard budget. `heartbeat_at` on `agent_runs` is | |
| the liveness signal (every event touches it). | |
| * **Budgeted sources (22 Sep)**: a 429 whose `Retry-After` exceeds the 300 s | |
| cap, or that reports a zero remaining budget (OpenAlex's per-request | |
| billing), closes that host for its window with no retry and no sleep; the | |
| lane is skipped with one warn event per cycle and the reason lands in | |
| `agent_runs.sources`. Set the Space secret `OPENALEX_API_KEY` (a free | |
| account key is ten times the shared anonymous allowance). `AGENT_USE_NTRS` | |
| (default on) adds the NASA Technical Reports Server lane: free, no key, | |
| direct PDFs. | |
| * **Cost column (22 Sep)**: Gemini `usageMetadata` is summed per cycle into | |
| `tokens_in` / `tokens_out` (text + vision calls); Runs & Logs derives | |
| dollars at read time from `AGENT_PRICE_IN_PER_M` / `AGENT_PRICE_OUT_PER_M` | |
| (defaults 0.30 / 2.50 USD per million, Gemini 2.5 Flash list price). | |
| * **Figure citation linking (22 Sep)**: `AGENT_LINK_FIGURES` (default on, | |
| Run Control toggle) links each text row to the figure its evidence cites | |
| and embeds the PNG in the row's `image` column (no S3 needed). Adds three | |
| columns (`figure_ref`, `figure_link_score`, `figure_link_signals`) that | |
| the boot migration creates; `rows_linked` is counted per cycle. | |
| * **Frontier rotation on dry cycles (28 Sep)**: every query a cycle searched | |
| is marked used in `node_report`, whether or not the cycle downloaded | |
| anything. Until then the bookkeeping sat behind `node_ingest`'s early | |
| return, so a dry cycle left `last_used_at` untouched and `next_queries` | |
| (LRU, never-used first) handed the same 20 queries back every half hour: | |
| from 22 to 28 Sep every cycle rediscovered the identical 109 exhausted | |
| candidates (42 DOIs already registered, 66 URLs already final) while the | |
| Gemini-expanded queries piled up unreached. Symptom to watch for in | |
| Runs & Logs: byte-identical Candidates / Skipped / URL seen columns across | |
| consecutive cycles. The harvest loop also no longer fetches (and marks) a | |
| frontier slice past its last round. | |
| * **Downloads that never got ingested (28 Sep)**: a cycle that finds the | |
| material schema unmigrated ends as `error` (not `crashed`) and keeps its | |
| PDFs and their `downloaded` registry rows; the next cycle adopts and | |
| ingests whatever is still on disk (`adopted N PDF(s)` event). At boot, | |
| `downloaded` rows older than 3 h whose file is gone are deleted and their | |
| URLs re-opened in the crawler seen-set (`boot repair: re-queued ...`), so | |
| discovery finds and fetches those papers again β this is what frees the | |
| five PDFs runs #135/#139 lost on 15β16 Sep. The Inspect-run panel now | |
| shows why a run ended (crash traceback, watchdog note or cycle error). | |
| * **More papers (28 Sep)**: discovery lanes on by default β OpenAlex phrase | |
| search *and* topic walk (`AGENT_OPENALEX_TOPICS`, 1-credit filter calls with | |
| persisted cursors, round-robin), Semantic Scholar bulk (β€1,000 OA-PDF papers | |
| per call; `S2_API_KEY` optional), arXiv with ANDed terms, NASA NTRS, and a | |
| datasheet round over the manufacturer seed pages (each seed β€ once a day); | |
| the query frontier is seeded once with the 455-intent grid | |
| (`agent/query_grid.py`, origin `grid`). Downloads try repository copies | |
| before publisher URLs and fall back to the landing page's | |
| `citation_pdf_url`; records without an abstract get relevance-gate relief. | |
| Note that publisher hosts (MDPI, Elsevier, Wiley, β¦) 403 any cloud egress | |
| address, so the yield comes from repositories, NTRS, arXiv and datasheets. | |
| * **Parallel ingestion (28 Sep)**: a cycle extracts its PDFs on a thread pool | |
| (`ingest_workers` in Run Control, default 3, hard max | |
| `AGENT_HARD_MAX_INGEST_WORKERS` = 6), one Postgres connection per worker; | |
| counters, events and the registry are written by the cycle thread as each | |
| PDF completes. Run Control also exposes the cycle's target (`Target new | |
| PDFs per cycle`, default 10: below it the cycle takes another frontier | |
| slice, up to `Max rounds per cycle`) β together with the 25-PDF cap these | |
| are the throughput knobs. A cycle's OpenAlex cost is 10 credits per query | |
| per round plus 1 per topic page, so 4 queries Γ 5 rounds Γ 48 cycles is | |
| β 9,700 credits/day: right at the free key's allowance; prepaid credits | |
| or a longer interval buy headroom. | |
| * **Candidate queue (28 Sep)**: discovery and download are decoupled. A | |
| *discovery pass* (`agent/discovery.py`, an `agent_runs` row of kind | |
| `discovery`) searches the next `discovery_queries` frontier queries on | |
| every lane β Semantic Scholar bulk, arXiv, NTRS, UDSpace (UD's DSpace | |
| theses, direct bitstreams), CORE (needs the free `CORE_API_KEY` secret; | |
| inactive without it) and, for the first `discovery_openalex_queries`, the | |
| 10-credit OpenAlex phrase search β walks `discovery_topic_pages` pages per | |
| OpenAlex topic (1 credit each) and crawls every datasheet seed that is due, | |
| then queues every relevant candidate in `agent_candidates` (a DOI already in | |
| the registry or a URL the crawler has given up on lands as `seen`). It runs | |
| every `discovery_interval_hours` (24), right after any cycle that finds the | |
| queue below `queue_low_water` (150) or drains it to that, at boot when the | |
| queue is empty, and on demand (Run Control β "Run a discovery pass now"). | |
| Cycles pop the queue best-score-first (repository-hosted copies rank ahead | |
| of publisher URLs) and download with no search calls at all; they search | |
| inline only when the queue is empty. Every popped row is settled: | |
| `ingested`, `seen` (registry / final URL), back to `queued` after a | |
| transient failure (never twice in one cycle; `failed` after 3 attempts). | |
| Turn it off with the Run Control toggle to get the pre-28 Sep inline | |
| behaviour. OpenAlex spend drops from up to ~9,700 credits/day to a few | |
| hundred per pass. | |
| * **Secrets never reach the database (28 Sep)**: `agdb.redact()` strips | |
| credential query parameters (`?key=`, `&api_key=`, `token=`) and bearer | |
| tokens from every event, run report, registry status and candidate error | |
| before it is stored, and `extraction` raises Gemini errors that name the | |
| endpoint and the API's own detail instead of the keyed URL (requests' | |
| `HTTPError` / urllib3's connection errors quote it). At boot, | |
| `agdb.redact_stored_secrets` scrubs rows written before this existed and | |
| logs `boot repair: redacted credential parameters β¦` β if that event ever | |
| appears, rotate the key it refers to: it was readable on Runs & Logs. | |
| * **Gemini budget refusals (28 Sep)**: a 429 whose body says | |
| `RESOURCE_EXHAUSTED` / "monthly spending cap" is not retried (it cost | |
| 14 s per PDF and cleared nothing). The PDF stays on disk with its | |
| registry row `downloaded`, the cycle's remaining PDFs are not attempted, | |
| one warn event says `Gemini refused for budget reasons (β¦); N PDF(s) kept | |
| on disk for the next cycle`, and the run ends `error` only if nothing was | |
| ingested at all. The next cycle adopts the kept PDFs (or the boot repair | |
| re-queues them after a restart). Raise the cap at ai.studio/spend. | |
| While the refusal lasts the agent is on a **budget hold** (`agent_state | |
| gemini_budget_hold`): cycles download nothing and probe with one kept PDF | |
| (or one fresh download if the disk was wiped); the first successful | |
| extraction lifts the hold and downloads resume next cycle. Adoption of | |
| kept PDFs is capped per cycle at the download cap, so a backlog is worked | |
| off over several cycles. | |
| * **Cost (28 Sep)**: Gemini's thinking tokens are billed as output at 6Γ the | |
| input price, and with no `thinkingConfig` the model thinks with an | |
| unbounded dynamic budget: a 25-paper cycle cost ~$7.50, $6.69 of it output | |
| (~30k tokens per PDF for a ~4k-token extraction). `gemini_thinking` | |
| (Run Control β Gemini cost; env `GEMINI_THINKING`; default `low`) is | |
| applied to every Gemini call β `dynamic` | `low` | `minimal` | `high` | | |
| `off` | an integer budget; a model that rejects the option gets the call | |
| repeated without it and the option is dropped for the process. Figure | |
| *mining* (one thinking vision call per mineable figure, into quarantined | |
| `figure_estimate` rows) is off by default; harvest + classify (one call per | |
| PDF) and figure linking stay on. Both defaults are applied once to a stored | |
| config at boot (`cost defaults applied: β¦` event) and can be changed back in | |
| Run Control. Expected: roughly a third of the previous bill per paper β | |
| watch the Cost column, and rows per paper / flagged share for recall. | |
| * **Cost (30 Sep) β real prices, models, near-duplicates**: the cost column | |
| used to price every token at Gemini 2.5 Flash rates ($0.30 / $2.50 per M) | |
| while the agent runs `gemini-3.5-flash` ($1.50 / $9.00): it showed about a | |
| quarter of the bill (runs 544β593: $12.98 shown, ~$48.50 real, ~$0.41 per | |
| PDF on dynamic thinking with mining on). Runs now store `cost_usd`, priced | |
| per model (`agent/config.py` `MODEL_PRICES`), text and vision calls each at | |
| their own model; older runs are re-priced at gemini-3.5-flash in Runs & | |
| Logs, which also shows thinking tokens, the model, $/PDF and a spend line. | |
| Run Control β Gemini cost adds the **text model** (every row is stamped | |
| with it: change it only after `cost_ab.py` on the gold set), a **vision | |
| model** for the figure calls (default: same), and the **near-duplicate | |
| gate** (`agent/textdedupe.py`, default 0.9): a PDF whose text is that | |
| similar to an ingested one β regional datasheet variants, a preprint and | |
| its published version β is skipped before any Gemini call | |
| (`duplicate_text` in the registry and the Near-dup column). | |
| * **Ephemeral disk**: PDFs live under `/tmp` only until ingested; provenance | |
| (DOI, sha256, title, URL, query) is durable in `agent_doi_seen`. | |
| * **Safety**: agent bookkeeping is in namespaced `agent_*` tables; controls | |
| are password-gated; read-only without the password. At boot the Space | |
| additively migrates the **connected** database (hardening columns incl. | |
| `origin`/`figure_id`, the origin-aware dedup index, the `figures` table) β | |
| `aim_agent` is the agent's own DB, so this is on by default. Before | |
| pointing `DB_NAME` at the team's shared database, either agree on the | |
| additive migration or set `AUTO_MIGRATE_MATERIAL_SCHEMA=0` (the cycle then | |
| refuses to write until `pg_migrate.py --apply` is run by hand). | |
| ## Repo map | |
| Vendored pipeline (single source of truth, synced from the project repo β | |
| the only Space-side deltas are the PG keepalive/statement-timeout hardening | |
| and the Retry-After cap): `extraction.py`, `batch_ingest.py`, `figures.py`, | |
| `pdf_crawler.py`, `pg_mirror.py`, `pg_migrate.py`, `migrate.py`, | |
| `generate_queries.py`. `FIGURES.md` documents the figure stage. | |
| Agent layer (new): `agent/` (orchestrator, scheduler with watchdog, autostart, | |
| keepalive, DB, config, `query_grid`), `run_space.py` (container entry point), | |
| `app.py` + `page_files/` (Streamlit UI), `agent/kg.py` + `agent/kg_explorer.html` | |
| (knowledge graph builder and explorer; `test_kg.py`). Tests: `test_agent_state.py`, | |
| `test_harvest_loop.py`, `test_sources.py` (no DB, no network), `test_e2e.py`, | |
| `test_figures_e2e.py`, `test_keep_running.py`, `test_counters_e2e.py`, | |
| `test_frontier_rotation_e2e.py`, `test_more_papers_e2e.py`, `test_parallel_ingest_e2e.py`, | |
| `test_candidate_queue_e2e.py` (scratch Postgres via `DB_*`; the network sources are | |
| switched off by env), `test_gemini_errors.py` (no DB, no network). | |