# ResearchBee v1.4 — verified-sources release ## What changed, in one line Policy facts now come from databases, not from the language model. Where no source could be reached, the system says so instead of filling the gap. --- ## 1. HuggingFace Space — `researchbee-hf/` | File | Status | |---|---| | `app.py` | **modified** — DOI-first `/api/check-license`, registry-backed `/api/find-repository`, fixed publisher matcher, Khazna startup + 23:30 refresh, new `/api/health` | | `prompts.py` | **modified** — "ResearchPilot" → "ResearchBee" everywhere; `LICENSE_SYSTEM_PROMPT` can no longer state policy facts; new `ADVISORY_NOTE_PROMPT`; `REPO_SYSTEM_PROMPT` constrained to registry selection | | `oaworks.py` | **new** — OA.Works Permissions client, schema-confirmed against T&F / Elsevier / Springer | | `doi_utils.py` | **new** — DOI extraction from URLs and text; ISSN mod-11 validation | | `repositories.py` | **new** — 25-entry curated repository registry | | `khazna_loader.py` | **new** — loads the harvested Khazna index from an HF Dataset | | `requirements.txt` | **modified** — added `huggingface_hub`, `apscheduler` | Upload all seven. `ku_knowledge.py`, `openalex.py`, `scimago_index.py`, `utils.py`, `Dockerfile`, and the SCImago CSV are unchanged — leave them. ### Before you deploy 1. **Set `KU_ROR` in `oaworks.py`.** Find it at https://ror.org (search "Khalifa University"). This matters more than it looks: Elsevier publishes a 12-month embargo for UK institutions and 24 months for everyone else, and the ROR lets OA.Works pick the right one. 2. Optionally set `KHAZNA_DATASET` and `HF_TOKEN` in Space secrets. 3. Check `GET /api/health` after deploy — it reports every source's status. Khazna lookups return `checked: false` until the harvest has run. That is correct behaviour, not a failure. --- ## 2. GitHub Pages — `researchbee-github/` | File | Status | |---|---| | `index.html` | **modified** — DOI/URL field added as the primary input | | `js/license.js` | **modified** — sends the DOI, renders provenance, Khazna status, deposit statement | | `js/render_license.js` | **new** — the new render helpers | | `css/style.css` | **modified** — styles appended at the end | | `khazna_oai.py` | **new** — the OAI-PMH harvester (runs in Actions, not in the Space) | | `.github/workflows/harvest-khazna.yml` | **new** — nightly 23:00 Dubai harvest | | `test_publishers.py` | **new** — validation harness | Other JS files are untouched. --- ## 3. Khazna harvest — first run ```bash # See the deposit gap before automating anything python khazna_oai.py gap # Verify the parser against real records python khazna_oai.py inspect --set publications:withFiles --format mods -n 2 --raw python khazna_oai.py inspect --set openaire --format oai_dc -n 2 --raw ``` Then: 1. Create an HF **Dataset** repo (default: `nikeshn/researchbee-data`) 2. Add `HF_TOKEN` as a GitHub repository secret 3. Actions → "Harvest Khazna" → **Run workflow** (first full run) 4. Nightly thereafter: Sunday full, Mon–Sat incremental Sunday is full because Khazna's OAI reports `deletedRecord: no` — incrementals can never learn that a record was withdrawn. --- ## 4. How to tell it is working **With a DOI** — green "Verified · OA.Works Permissions" badge, per-version rights, the publisher's required deposit statement with a copy button, a source panel with evidence links, and Khazna status. **Without a DOI** — amber "Unverified · journal-level estimate", all version blocks "Unclear", a prompt to enter a DOI. The backend forcibly downgrades any `policy_status: "Confirmed"` the model returns, so this cannot be bypassed. **Khazna unavailable** — "Khazna status not checked", explicitly distinguished from "Not yet in Khazna". Test DOIs from this build: ``` 10.1080/19322909.2023.2221477 T&F — no embargo stated 10.1016/j.jclepro.2023.136775 Elsevier — 24 months (not 12; see below) 10.1007/s11192-023-04812-4 Springer — preprint immediate, AAM 12 months ``` --- ## 5. Things worth knowing **The Elsevier embargo trap.** OA.Works returns two permissions with identical `score: 1100` — 12 months (UK) and 24 months (everyone else). Ranking by score alone silently picks 12. `_permission_rank()` in `oaworks.py` breaks the tie by the *longer* embargo. Do not "optimise" that away. **Absent is not zero.** T&F records carry no embargo field. The system reports "Not stated in the permission record" rather than "no embargo", because T&F's published policy is normally 12 months and a wrong permissive answer costs the library a takedown notice. **Publisher matcher.** Now covers imprints whose SCImago string never says the parent publisher — W.B. Saunders, Academic Press, Cell Press, BioMed Central, Palgrave, Blackwell, Informa, Dove Medical. Q1 APC badge coverage went from 45.8% to 51.1% (+497 journals). `AMBIGUOUS_ALIASES` prevents "Maximum Academic Press" from being mistaken for Elsevier. **Repository registry.** The model selects entries by `id` and writes the `fit_reason`; every factual field is overwritten from `repositories.py`. Unknown ids are dropped. Entries you have not personally verified are left blank rather than guessed — please fill them in and bump `verified`. --- ## 6. Still open - `KU_ROR` is empty - Repository registry has blank fields awaiting your verification - Wiley, IEEE and ACS not yet validated against OA.Works (`test_publishers.py`) - Worth telling OA.Works about the UK/non-UK tie: `help@openaccessbutton.org`