--- title: Scientific Contribution Graph Explorer emoji: 🧠 colorFrom: indigo colorTo: purple sdk: docker app_port: 7860 pinned: false license: apache-2.0 short_description: Visualize how scientific contributions build on each other. --- # Scientific Contribution Graph Explorer (HF Spaces) A browser-based UI for the [Scientific Contribution Graph](https://github.com/cognitiveailab/scientific-contribution-graph): 1. Search for a paper by title. 2. Inspect each scientific contribution in that paper, along with its downstream **impact** (the number of later contributions that build on it). 3. Render forward/backward crawl visualizations for any contribution, with a live sidebar of knobs (direction, layout, depth, strong-only, leaf trimming). The UI is a single-file FastAPI app — no Gradio, no React, no build step. The visualization is an inline SVG with pan/zoom via `svg-pan-zoom`. ## Deploying this Space The same `demo.py` runs both locally and on HF Spaces; only the configuration differs. Copy these four files from `demo/` in the upstream repo into your Space and rename: | In the repo | In the Space | | --- | --- | | `demo/demo.py` | `demo.py` | | `demo/demo_huggingface_demo_config.json` | `demo_config.json` | | `demo/demo_huggingface_requirements.txt` | `requirements.txt` | | `demo/demo_huggingface_Dockerfile` | `Dockerfile` | | `demo/demo_huggingface_README.md` | `README.md` (keep the YAML frontmatter above) | The Space uses the **Docker SDK**, so the Dockerfile handles `apt-get install graphviz` for you. ## How configuration works `demo.py` looks for `demo_config.json` sitting next to it. The shipped HF config sets: ```json { "data_path": "/home/user/scg-release", "bucket_uri": "hf://buckets/pajansen/scientific-contribution-graph/releases-tar/current", "host": "0.0.0.0", "port": 7860 } ``` Because `bucket_uri` is non-null, `demo.py` runs `hf buckets sync` on first boot to download the release tarball, then extracts it into `data_path` and loads the graph. Subsequent boots find the unpacked data and skip the download. Precedence is **CLI flags > env vars > config file > built-in defaults**, so you can override any field by setting env vars (`SCG_DATA_PATH`, `SCG_BUCKET_URI`, `SCG_SERVER_NAME`, `SCG_SERVER_PORT`) without editing the config file. ## Hardware The dataset is large (~30 GB unpacked, ~7 GB zipped). Pick a Space with at least: - **8+ GB RAM** for the basic UI. - A persistent disk if you don't want cold boots to redownload — see below. ## Persistence — important By default, the Space's filesystem is **ephemeral**: when the Space sleeps (auto-sleeps after ~48 h of inactivity on free tier) and a visitor wakes it back up, the container is rebuilt from the Docker image and the ~7 GB tarball is re-fetched and re-extracted (~5–10 min before the UI is responsive). To make the data survive restarts, enable **[Persistent Storage](https://huggingface.co/docs/hub/spaces-storage)** on the Space (paid feature; pick a tier ≥ **50 GB** to fit the unpacked corpus). HF mounts it at `/data`, which is where this config already points (`data_path: /data/scg-release`), so it works out of the box — download happens once, survives forever. If you choose **not** to use Persistent Storage, no config changes are needed; just be aware that every cold boot pays the download + extract cost. **Do not upload an unpacked copy of the corpus to the bucket.** `hf buckets sync` over ~230k small JSON files is dominated by per-file overhead and is dramatically slower than transferring a single 7 GB tarball plus a local `tar -xf`. The current tarball-based release format is the right one. ## First-boot behaviour On first launch, `demo.py` runs `hf buckets sync` to download the release tarball (~7 GB), extracts it into `data_path`, deletes the tarball, then loads the graph (~7 s). Subsequent boots find the unpacked data and skip the download — instantly, if Persistent Storage is on; never, if not. A small `.scg_release_marker.json` is written next to the data on every install. It records the tarball's name and byte-size so future refreshes can detect whether the bucket actually has a new release vs. the same one we already extracted. ## Refreshing the data (new release) When a new release is published to the bucket, the demo doesn't pick it up automatically — it only checks at startup. To pull the new data **without restarting the Space**, set an admin token and POST to `/admin/refresh`: 1. In the Space settings, add a private secret named `SCG_ADMIN_TOKEN` set to a long random string. The admin endpoints are disabled if this isn't set, so casual visitors can't trigger a 10-minute redownload. 2. From your laptop: ```bash # Kick off the refresh (returns immediately) curl -X POST "https://.hf.space/admin/refresh?token=$SCG_ADMIN_TOKEN" # Poll progress curl "https://.hf.space/admin/refresh/status?token=$SCG_ADMIN_TOKEN" ``` `status` cycles through `syncing → reloading → done` (or `error`). The endpoint hot-swaps the in-memory graph once the new data is extracted, so the UI is briefly serving the old graph and then the new one — no restart, no downtime for read-only requests. Alternatively, force a refresh at startup with `--force-refresh` (or `python demo.py --force-refresh` from a Space console). If Persistent Storage is **not** enabled, every cold boot already starts clean, so no refresh is needed.