| --- |
| title: LHAPDF Plotly Plots |
| emoji: π |
| colorFrom: blue |
| colorTo: purple |
| sdk: docker |
| app_port: 7860 |
| pinned: false |
| --- |
| |
| # LHAPDF Plotly Plots β Hugging Face Space |
|
|
| A FastAPI service, deployed as a Hugging Face Space, that computes LHAPDF |
| plots with **Plotly** (central value + uncertainty band, ratio plots, and |
| flavor-correlation contours), caches them as JSON inside a HF **dataset** |
| (acting as a lightweight database), and serves them back so a webpage can |
| render them with `plotly.js` without ever recomputing an already-generated |
| plot. |
|
|
| This Space is meant to be an **on-demand fallback for the long tail**, not |
| the primary way plots get into the dataset β see "Recommended workflow" |
| below for why precomputing the common flavor/energy combinations locally |
| and bulk-uploading them is the better path for anything you already know |
| visitors will request. |
|
|
| ## Supported PDF sets |
|
|
| Uncertainties are computed generically via LHAPDF's own |
| `PDFSet.uncertainty()` method, which is only mathematically valid for |
| these **ErrorType**s: |
|
|
| - `hessian` |
| - `symmhessian` |
| - `replicas` |
|
|
| Any other error type is rejected with an HTTP 400 explaining why β |
| see `services/error_types.py`. This is checked **once**, centrally, before |
| any plot is computed, so every endpoint behaves consistently. When you |
| implement calculations for a new error type later, add it to |
| `SAFE_ERROR_TYPES` (if `uncertainty()` is valid for it) or add a |
| dedicated branch/function for it (if it needs different math). |
|
|
| ## How it works |
|
|
| 1. Your webpage (or notebook) sends a request describing the plot it |
| wants (grid name, Q scale, flavor id(s), etc.) to one of the Space's |
| endpoints. |
| 2. The Space hashes the request parameters into a deterministic |
| **`plot_id`** and checks, in order: plots queued this session but not |
| yet uploaded, then the grid's `index.json` catalog inside the dataset |
| repo. |
| - **Cache hit** (either kind) β the existing `plotly_json_url` is |
| returned immediately, no computation happens. |
| - **Cache miss** β the LHAPDF grid is downloaded (if not already local), |
| its `ErrorType` is validated, the plot is computed with Plotly, and |
| written to local scratch space. It is **not** uploaded to the dataset |
| immediately β see "Batched uploads" below. |
| 3. Your webpage fetches the returned `plotly_json_url` and renders it |
| directly with `Plotly.newPlot(div, fig.data, fig.layout)` β no |
| server-side image rendering needed. |
| |
| ## Recommended workflow: precompute the common combos, let the Space handle the long tail |
| |
| This Space was designed as an **on-demand fallback**, not the primary way |
| plots get into the dataset. For any flavor/energy combination you already |
| know visitors will want (i.e. everything listed in a dataset's |
| `qValues`/`flavors` in the website's `config.json`), it's significantly |
| better to generate those plots **locally** ahead of time and bulk-upload |
| them directly to the dataset, rather than letting real visitor traffic |
| trigger the Space to compute them one at a time. |
| |
| Why this matters: |
| |
| - **The website never calls the Space for a plot that already exists** β |
| it fetches directly from `baseUrl` first and only falls back to a |
| Space endpoint on a 404 (see the webpage repo's `cteq-parton-plots.js` |
| β `fetchPremadePlot`/`requestGeneratedPlot`). Precomputing means the |
| Space is simply never invoked for that plot, ever. |
| - **Sidesteps `LHAPDF_LOCK` entirely for the common case.** LHAPDF's C++ |
| core isn't thread-safe, so this Space serializes every lhapdf call |
| through a single lock (see `services/pdf_utils.py::LHAPDF_LOCK`) β a |
| burst of concurrent visitors requesting different plots at the same |
| moment queue up behind it. Precomputed plots need none of that. |
| - **Sidesteps the commit-rate-limit machinery entirely for the common |
| case.** No Space-side generation means no entry in the pending queue |
| and nothing for the batched-flush job to do for that plot. |
| - Only the true long tail hits the Space at all: an energy a visitor |
| typed into the free-form picker that nobody precomputed, or a newly |
| added flavor/dataset combo. That traffic is inherently low-volume and |
| rarely concurrent, which is exactly the load profile `LHAPDF_LOCK`, the |
| concurrency cap (`services/pdf_utils.py::COMPUTE_SEMAPHORE` β see |
| `CLAUDE.md`), and the batching in "Batched uploads" below were built to |
| absorb β they're a safety net for the rare case, not meant to carry |
| normal load. A large uncached burst (e.g. Compare Datasets mode across |
| many datasets/flavors/energies nobody has precomputed yet) will still |
| just be slow β queued behind the concurrency cap and the lock β rather |
| than crash the Space, but "slow" is still worse than "instant" for |
| combos you already know are common. |
|
|
| **How to precompute correctly** (so the output is indistinguishable from |
| what the Space itself would have produced): |
|
|
| 1. Reuse this repo's own code rather than reimplementing it β import |
| `services.plotting.build_standard_plot` / `build_standard_ratio_plot` / |
| `build_correlation_plot` / `build_main_plot` and `services.config`'s |
| filename functions (`standard_plot_filename`, `ratio_plot_filename`, |
| `correlation_plot_filename`, `main_plot_filename`) directly. Any drift in the JSON shape |
| (rounding, missing `customdata`, different trace ordering) risks |
| silently breaking hover tooltips or uncertainty bands on the website β |
| this exact class of bug has bitten this project before. |
| 2. Load the PDF set locally with `services.pdf_utils.load_validated_pdf_set` |
| / `get_pdf_members`, same as the Space does, then loop over the |
| cartesian product of the dataset's configured `qValues` Γ `flavors` |
| (and flavor pairs for correlation plots), building one figure per |
| combination. |
| 3. Write each figure's `.to_json()` to disk under the exact folder/name |
| convention the dataset already uses: `<grid_name>plots/<filename>` |
| (see "Dataset layout" below) β mixing manually- and Space-generated |
| plots in the same folder is the existing, intended setup. |
| 4. **Upload everything in ONE commit**, not one API call per file. |
| Looping `upload_file()` per plot walks straight back into the |
| 128/hour commit-rate-limit problem described below. Use |
| `HfApi().upload_folder(...)` or build a list of `CommitOperationAdd` |
| and call `create_commit()` once β the same pattern |
| `services/storage.py::flush_pending` uses for Space-generated plots. |
| 5. Optionally also write/merge each grid's `index.json` with the same |
| entry schema `queue_plot`/`flush_pending` produce (`plot_id` via |
| `make_plot_id`, `plot_type`, `grid_name`, `parameters`, |
| `json_filename`, `json_path`, `plot_url`, `plotly_json_url`, |
| `created_at`). Not required for the website's fast path β it only |
| needs the file to exist at the right name β but keeps `index.json` |
| accurate if anything ever queries the Space directly with |
| `use_cache: true` for one of these plots. |
|
|
| Precomputing the *entire* cartesian product can get large fast (e.g. 11 |
| flavors Γ 3 energies Γ 11 flavors for correlation plots is a few hundred |
| files) β it's fine to just cover the combinations you know are actually |
| used and leave rarer ones to the Space on demand. |
|
|
| ## Batched uploads (commit rate limiting) |
|
|
| Hugging Face caps commits at **128/hour per repo**. Uploading each newly |
| generated plot immediately cost two commits (the plot file + the |
| `index.json` update), which a handful of dashboard visitors could exhaust |
| in minutes β the symptom is a `429 Too Many Requests` from |
| `huggingface.co/api/datasets/.../commit/main`. |
|
|
| To stay under that limit, newly generated plots are **queued** instead of |
| uploaded per-request (`services/storage.py::queue_plot`): |
|
|
| 1. The figure is written to local scratch space (`TMP_DIR`, |
| `/tmp/pdf_plots_space` by default) and registered in an in-memory |
| pending list. |
| 2. `plotly_json_url` in the response points at this Space's own |
| `GET /plots/{folder}/{filename}` route, which serves the file straight |
| off local disk β so the requesting page renders it immediately, |
| with no wait for a commit. |
| 3. A background thread (started on Space startup) calls |
| `flush_pending` every `FLUSH_INTERVAL_SECONDS` (default 50 minutes, |
| configurable via that env var) and bundles **every** plot queued since |
| the last flush β across every grid and dataset_repo β into a single |
| `create_commit` call per dataset_repo: one commit uploads all the plot |
| JSON files plus each affected grid's merged `index.json`. |
| 4. `POST /flush-cache` triggers an out-of-schedule flush manually (handy |
| while testing). `GET /pending-status` shows how many plots are |
| currently queued per `dataset_repo|grid_name`. |
| |
| Files already served from `/plots/...` keep working after a flush (they |
| stay on disk), so nothing needs to change client-side once a commit |
| lands. The only failure mode is a Space container restart before a flush: |
| local scratch space is ephemeral, so anything still queued is lost and |
| gets silently regenerated (and re-queued) on the next request for it. |
| |
| ## Dataset layout |
| |
| Inside `dataset_repo` (e.g. `lopezels/public-pdfs`): |
|
|
| ``` |
| <grid_name>/ # LHAPDF grid files (already existing) |
| <grid_name>plots/ |
| index.json # catalog for this grid |
| <grid_name>_plt_xfx_PID<id>_Q<q>.json # standard plot |
| <grid_name>_plt_xfx_ratio_PID<id>_Q<q>.json # ratio plot |
| <grid_name>_plt_xfx_corr_PID<id1>_PID<id2>_Q<q>.json # correlation plot |
| <grid_name>_plt_main_xfx_PID<id1>_<id2>_..._Q<q>.json # multi-flavor "overview" plot |
| ``` |
|
|
| Filenames follow the exact convention used in the original notebook, so |
| files generated by the API and files generated manually land in the same |
| place with the same names. Each entry in `index.json` also stores the |
| request `parameters` and a `plot_id` (a SHA-256 hash of `plot_type` + |
| `grid_name` + `parameters`) used purely for cache lookups β the display |
| filename and the cache key are independent, so you can change one without |
| breaking the other. |
|
|
| ## Project structure |
|
|
| ``` |
| . |
| βββ Dockerfile # HF Space image definition |
| βββ app.py # FastAPI entrypoint β routing only |
| βββ requirements.txt # For local dev (conda handles lhapdf in Docker) |
| βββ example_client.py # Example calls to all three endpoints |
| βββ services/ |
| βββ __init__.py # Package marker |
| βββ config.py # Paths, folder & filename conventions |
| βββ models.py # Pydantic request/response schemas |
| βββ error_types.py # Supported-ErrorType registry + validation |
| βββ pdf_utils.py # Grid download + validated PDFSet loading |
| βββ plotting.py # Plotly figure builders (standard/ratio/correlation) |
| βββ colors.py # Parton flavor names, LaTeX symbols, colors |
| βββ storage.py # HF dataset caching / index.json management |
| ``` |
|
|
| Colors, names, and LaTeX symbols are defined once in `colors.py` β every |
| plot function imports from there, so changing the palette is a single edit. |
|
|
| ## Color palette |
|
|
| Feynman-diagram-style palette: each antiquark uses the color-anticolor |
| complement of its quark, and the gluon has its own saturated color |
| (no gray). |
|
|
| | Flavor | Color | Anti-flavor | Color | |
| |--------|-------|-------------|-------| |
| | Bottom (b) | `#DC2626` | Anti-bottom (bΜ) | `#06B6D4` | |
| | Charm (c) | `#9333EA` | Anti-charm (cΜ) | `#FBBF24` | |
| | Strange (s) | `#0891B2` | Anti-strange (sΜ) | `#F97316` | |
| | Up (u) | `#3B82F6` | Anti-up (Ε«) | `#FCD34D` | |
| | Down (d) | `#EC4899` | Anti-down (dΜ) | `#10B981` | |
| | Gluon (g) | `#FF6B35` | β | β | |
|
|
| ```python |
| from services.colors import parton_colors, get_parton_color |
| |
| color = parton_colors[5] # Bottom -> #DC2626 |
| color = get_parton_color(-2) # Anti-up -> #FCD34D (falls back to a default if unknown) |
| ``` |
|
|
| ## Endpoints |
|
|
| ### `POST /compute-standard-plot` |
|
|
| Central value + uncertainty band for `xΒ·f(x, Q)` of a single flavor. |
|
|
| ```json |
| { |
| "dataset_repo": "lopezels/public-pdfs", |
| "grid_name": "CT18NNLO", |
| "q_scale": 1.0, |
| "parton_id": 5, |
| "plot_color": null, |
| "custom_title": " " |
| } |
| ``` |
|
|
| `plot_color` defaults to the flavor's palette color if omitted. |
|
|
| ### `POST /compute-ratio-plot` |
|
|
| Same as above, normalized so the central value sits at 1.0 (relative |
| uncertainty band). |
|
|
| ```json |
| { |
| "dataset_repo": "lopezels/public-pdfs", |
| "grid_name": "CT18NNLO", |
| "q_scale": 1.0, |
| "parton_id": 5 |
| } |
| ``` |
|
|
| ### `POST /compute-correlation-plot` |
|
|
| Hessian-style correlation contour between two flavors across an x-grid. |
|
|
| ```json |
| { |
| "dataset_repo": "lopezels/public-pdfs", |
| "grid_name": "CT18NNLO", |
| "q_scale": 1.0, |
| "parton_id_1": 5, |
| "parton_id_2": 21, |
| "color_anti_correlated": "black" |
| } |
| ``` |
|
|
| `color_correlated` defaults to `parton_id_1`'s palette color if omitted. |
|
|
| #### What the correlation contour actually computes |
|
|
| It is **not** a correlation of PDF ratios. It's the standard Hessian PDF |
| "correlation cosine" (Pumplin et al., hep-ph/0201195), computed directly on |
| the raw $x \cdot f(x, Q)$ values (`pdf.xfxQ(...)`, LHAPDF's own already-$x |
| f$ quantity) β the same quantity plotted by `/compute-standard-plot`, just |
| evaluated at two different points instead of a shared x-grid. |
|
|
| For flavor 1 at $x_1$ and flavor 2 at $x_2$, let $F_m$ and $G_m$ be the |
| $x \cdot f(x, Q)$ value flavor 1 / flavor 2 takes for PDF member $m$ |
| ($m = 0$ is always the central/best-fit member for every supported |
| `ErrorType`; $m = 1, \dots, N$ are the other members β eigenvector sets |
| for `hessian`/`symmhessian`, individual replicas for `replicas`). |
| Deviations are taken from the central member: |
|
|
| $$ |
| \Delta F_m(x_1) = F_m(x_1) - F_0(x_1), \qquad |
| \Delta G_m(x_2) = G_m(x_2) - G_0(x_2), \qquad m = 1, \dots, N |
| $$ |
|
|
| and the correlation at that $(x_1, x_2)$ grid point is the normalized |
| inner product of the two deviation vectors across all $N$ non-central |
| members β i.e. their cosine similarity, clipped to $[-1, 1]$ for |
| numerical safety: |
|
|
| $$ |
| \cos\varphi(x_1, x_2) = \frac{\displaystyle\sum_{m=1}^{N} \Delta F_m(x_1)\, \Delta G_m(x_2)} |
| {\displaystyle\sqrt{\sum_{m=1}^{N} \Delta F_m(x_1)^2} \; \sqrt{\sum_{m=1}^{N} \Delta G_m(x_2)^2}} |
| $$ |
| |
| $+1$ means the two quantities always move together across the member set |
| at that $(x_1, x_2)$ pair, $-1$ means they always move oppositely, $0$ |
| means uncorrelated. Repeating this for every $(x_1, x_2)$ pair on the |
| x-grid produces the full contour (`build_correlation_plot` in |
| `services/plotting.py`). |
| |
| This is a generalization of the textbook Hessian formula rather than a |
| literal transcription of it: the textbook version pairs up the $+$/$-$ |
| eigenvector members explicitly, |
| |
| $$ |
| \cos\varphi = \frac{\sum_{i=1}^{d} \left[F(S_i^+) - F(S_i^-)\right]\left[G(S_i^+) - G(S_i^-)\right]}{4\, \Delta F\, \Delta G} |
| $$ |
|
|
| which assumes a specific `hessian`/`symmhessian` member layout ($d$ |
| eigenvector directions, each contributing a $+$/$-$ pair of members). The |
| formula above instead sums over *every* non-central member's deviation |
| from the central member without assuming that pairing, so the same code |
| path works unmodified for `replicas` sets too. For a Hessian set with the |
| conventional symmetric $\pm$ eigenvector structure, this is algebraically |
| equivalent to the textbook formula up to an overall constant that cancels |
| out in the ratio (the standard small-deviation assumption |
| $X(S_i^-) \approx 2 X(S_0) - X(S_i^+)$ used throughout Hessian PDF |
| uncertainty formulas). For a `replicas` set, where LHAPDF's member 0 is |
| already the mean/central replica, this is exactly the Pearson correlation |
| coefficient of $x \cdot f(x, Q)$ across the individual replicas. |
| |
| ### `POST /compute-main-plot` |
| |
| The classic multi-flavor "overview" plot (CT18 Fig. 2 style, see |
| arXiv:1912.10053): every flavor in `parton_ids` overlaid on one figure as |
| central value + 90% C.L. band, at a single Q. The gluon (id `21`) is always |
| scaled down by a factor of 5 in the figure itself (`g/5` in the legend) β |
| not configurable, it's the fixed convention that plot follows. |
|
|
| ```json |
| { |
| "dataset_repo": "lopezels/public-pdfs", |
| "grid_name": "CT18NNLO", |
| "q_scale": 2.0, |
| "parton_ids": [3, 21, 2, 1, -1, -2, 4] |
| } |
| ``` |
|
|
| `parton_ids` order is preserved (not sorted) β it drives both the legend |
| order and the cached filename (see "Dataset layout" above), so a caller |
| that wants cache hits must always send the same order. |
|
|
| `n_points` defaults to `500` (up from the `150` used by |
| `/compute-standard-plot`/`/compute-ratio-plot`) β this plot overlays 7 |
| flavors' central + band curves at once, and the curves already looked |
| visibly faceted/"pixelated" at a lower point count once several bands |
| overlap. The tradeoff is compute cost: this endpoint evaluates |
| `n_points Γ len(parton_ids)` points per member (vs. just `n_points` for a |
| single-flavor plot), so raising it further scales linearly with both |
| `n_points` and the number of flavors requested. |
|
|
| All four endpoints return: |
|
|
| ```json |
| { |
| "status": "success", |
| "cached": false, |
| "message": "Plot generated; queued for the next batched upload to lopezels/public-pdfs.", |
| "plot_id": "standard_plot_3f9a1c2b7e0d4f5a", |
| "plot_url": "https://huggingface.co/datasets/.../blob/main/...", |
| "plotly_json_url": "https://lopezels-pdf-plots.hf.space/plots/CT18NNLOplots/CT18NNLO_plt_xfx_PID5_Q1.0.json" |
| } |
| ``` |
|
|
| `plotly_json_url` points at this Space's own `/plots` route until the |
| plot has been committed (see "Batched uploads" above) β `plot_url` |
| (the eventual `huggingface.co/datasets/.../blob/main/...` link) only |
| resolves once that commit lands. |
|
|
| ### Error response (unsupported ErrorType) |
|
|
| ```json |
| { |
| "detail": "Ignoring 'SOME_SET'. Its ErrorType is 'mc_1step'. Requires a dedicated calculation (read the PDF set's paper) that isn't implemented yet." |
| } |
| ``` |
|
|
| ## Using plots in your webpage |
|
|
| Plot titles/axis labels use LaTeX (`$...$`), so load MathJax alongside |
| Plotly for correct rendering: |
|
|
| ```html |
| <script src="https://cdn.plot.ly/plotly-latest.min.js"></script> |
| <script src="https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.1/MathJax.js?config=TeX-MML-AM_SVG"></script> |
| |
| <div id="plot"></div> |
| <script> |
| fetch(plotlyJsonUrl) |
| .then(r => r.json()) |
| .then(fig => Plotly.newPlot("plot", fig.data, fig.layout)) |
| .then(() => MathJax.Hub.Queue(["Typeset", MathJax.Hub, "plot"])); |
| </script> |
| ``` |
|
|
| ## Adding a new error type in the future |
|
|
| 1. Decide whether `PDFSet.uncertainty()` is mathematically valid for the |
| new error type. |
| - **Yes** β add it to `SAFE_ERROR_TYPES` in `services/error_types.py`. |
| The existing plot functions will work for it automatically. |
| - **No** β keep it out of `SAFE_ERROR_TYPES`, write a new function |
| (e.g. `build_standard_plot_mc1step`) with the correct math in |
| `services/plotting.py`, and branch on error type in `app.py` (or add |
| a dedicated endpoint) to call it. |
| 2. No other file needs to change β `app.py`'s routing, `storage.py`'s |
| caching, and `colors.py`'s palette all stay exactly the same. |
| |
| ## Notes / things worth knowing |
|
|
| - **`use_cache: false`** on any request forces recomputation even if a |
| matching plot already exists β useful while iterating on a plot's style. |
| - The catalog merge in `storage.flush_pending` does a read-modify-write of |
| `index.json` per grid at flush time. Under concurrent flushes for |
| genuinely new plots this can race; for a low-traffic internal tool this |
| is an acceptable trade-off, but if you expect concurrent writers, |
| consider moving the index to a proper database instead of a single |
| JSON file. |
| - The dataset write token is read server-side from this Space's own |
| `HF_TOKEN` secret (Settings β Variables and secrets on the Space) β it is |
| **not** part of the request schema, so callers (including public webpage |
| JS) never see or send it. Set that secret before deploying, or every |
| endpoint returns HTTP 500 ("HF_TOKEN secret is not set on this Space."). |
| - CORS is wide open (`allow_origins=["*"]`) so the public website's |
| client-side JS can call this Space directly from a different origin. |
| Tighten this in `app.py` if that's ever a concern. |
| - LHAPDF grids are cached on local disk (`LHAPDF_DATA_PATH`) after first |
| download, so repeated requests for the same `grid_name` don't |
| re-download it. |
| |