PDE-OBS: trained baseline checkpoints

Final checkpoints of the PDE-OBS campaign: one model per (PDE family, method, training observation pattern). Each checkpoint is the final recorded checkpoint of its attempt, not one selected on validation or test error.

All 441 credited settings are published here. models_manifest.<cluster-label>.json lists the SHA-256 of every published file; the benchmark repository binds each published checkpoint digest to the checkpoint identity its results index uses (results/public_deposits/release_map.json there).

This repository accompanies an anonymous submission under double-blind review.

Layout

models/<cluster-label>/<pde>/<method>/<train_view>/
    checkpoints/last.pt              final checkpoint: weights, scheduler state, history, resolved config
    checkpoints/training_config.json resolved trainer configuration
    identity.json  provenance.json  resolved.yaml  split_manifest.json  factor_coverage.json
    completion.json  health.json  history.json          (where the attempt wrote them)
    budget-protocol.json | dynamic-protocol.json        (cohort dependent)
models_manifest.<cluster-label>.json          per-file SHA-256 of the published files (repository root)
scrub-manifest.json                           which record files the last de-identification pass rewrote

The cluster labels (cluster-A, cluster-B, cluster-C) stand for the three machines the campaign ran on; they carry no institutional meaning. resolved.yaml beside each checkpoint records the architecture (method.name, method.kwargs), the training view and the training settings.

Scoring a checkpoint with the benchmark

from pdeobs import api

data = api.load_dataset("./pdeobs-data", verify=True)              # from PDE-OBS/pdeobs-data
package, targets = api.inference_input_from_dataset(data, api.make_observation("paper:R50"), task="recovery")

pred = api.load_legacy_checkpoint("models/cluster-A/poisson/fno/random_50pct/checkpoints/last.pt",
                                  model={"name": "fno", "preset": "paper"}, task="recovery")
score = api.evaluate(api.predict(pred, package), targets)
score["status"], score["scoring_version"], score["summary"]["rel_l2_joint_mean"]

The loader takes the structure from the caller (the campaign checkpoints use the paper preset of their method), cross-checks the checkpoint's own stored training configuration, and records the file's SHA-256 and epoch in the predictor's provenance. Without the benchmark, the file is a plain PyTorch payload: torch.load(path, map_location="cpu", weights_only=False) returns a dict with model_state, config, history, epoch and scheduler_state.

Important properties

  • Weights only. Optimizer, gradient-scaler and RNG states were removed (about two thirds of each file); dropped_keys in the manifest records this per checkpoint. The checkpoints support inference and rescoring, not exact continuation of training.
  • Anonymized metadata, unmodified parameters. Every string naming a filesystem path, account, node, host, login node, scheduler job, partition, GPU hardware identifier, interpreter or environment directory, run directory or source-code revision was rewritten to a placeholder such as <ROOT_A>, <USER>, <NODE>, <JOB>, <PYTHON>, <ENV>, <RUN> or <COMMIT-1>, and raw scheduler records were reduced to their resource fields (<SCHEDULER_RECORD> ...); three passes, the files rewritten by the last one are listed in scrub-manifest.json. Model parameters were not modified. models_manifest.*.json records the SHA-256 of each released file; the correspondence to the campaign's checkpoint identities is kept in the benchmark repository.
  • Training cohorts are not interchangeable. Each entry carries its cohort label: the original 500-epoch protocol, a recovered continuation of it, a declared 200-epoch budget, a budgeted max-200 patience protocol, or a salvaged interrupted attempt. Any table that pools cohorts must say so per row.
  • Scoring provenance. All 441 checkpoints were later rescored from retained prediction arrays with the benchmark's strict scorer (pdeobs-strict-v1, nine views x 200 held-out identities each); those results and their per-identity errors are in the benchmark repository under results/prediction_verification_20260924/. The campaign's own evaluator metrics remain in the archived-results index and are not the paper's current numbers.

Citation

Anonymous submission under review. Please cite the paper once it is public.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support