PIVOT / docs /method.md
pranamanam's picture
Upload 176 files
6fa9282 verified
|
Raw
History Blame Contribute Delete
2.58 kB
# Equations and implementation
The original PIVOT formulation is preserved: mean-pooled gene-operation embeddings, a residual endpoint map, map/tangent/composition losses, and projected endpoint-reward gradients. The optional population loss compares response distributions.
| Paper quantity | Implementation |
|---|---|
| Cell-state representation and reconstruction | `data/preprocess.py`, `PerturbData.decode_to_genes` |
| Mean-pooled intervention embedding | `models/encoders.py`, `PerturbationEncoder` |
| Residual map and diagonal identity | `models/flow_map.py`, `FlowMap` |
| Map, tangent, and composition objectives | `training/losses.py`, `compute_losses` |
| Population MMD term | `training/train.py`, `evaluation/rewards.py` |
| Exhaustive catalog score | `evaluation/inference.py`, `endpoint_ranking` |
| Endpoint vector-Jacobian product | `evaluation/inference.py`, `reward_guidance` |
| Nearest embedding projection and reranking | `project_and_rerank` |
| Greedy gene-set construction | `greedy_combinatorial` |
| Independent measured candidate outcomes | `evaluation/runner.py`, `evaluate` |
The cosine reward compares each predicted cell's control-relative effect with the target centroid effect. The centroid reward averages squared cell-to-centroid distances, so it includes predicted population spread. MMD and sliced-Wasserstein rewards operate on whole populations.
Guidance uses normalized ascent with a fixed step size. Discrete projection can lower the continuous reward. Exhaustive ranking maximizes the same predictor score over a finite fixed catalog. Guidance therefore requires evaluation of computation, measured outcome quality, and the reward change after projection. The implementation records candidate-cell evaluations and query timings. Projection over embeddings also scans the catalog; its cost should be included in full runtime comparisons.
The default full model uses a 64-dimensional intervention embedding, four 512-unit SiLU layers, Fourier time features, 60 epochs, batch size 1,024, and AdamW. `small.json` is a two-epoch workflow check with a 16-dimensional intervention embedding and two 64-unit layers. Configurations containing `distribution` activate population matching. The `map_*` configurations isolate the two auxiliary losses.
Historical source files also contain expression-derived cluster features. They remain in the preserved original archive. The corrected training interface accepts metadata-only gene/operation representations and does not offer those outcome-derived clusters as independent pathway features.