--- license: cc-by-4.0 library_name: pytorch pipeline_tag: feature-extraction tags: - node2vec - graph-embedding - metabolomics - pytorch-geometric - metabolights - metabolomics-workbench metrics: - roc_auc --- # Node2Vec embeddings for cross-study metabolomics co-response graphs 128-dimensional Node2Vec embeddings for two undirected graphs of metabolite co-response edges that pass a chi-square validity filter (every expected frequency ≥ 5). Held-out link prediction reaches **AUC 0.932** and **0.928**, against degree-only baselines of 0.826 and 0.888. The graphs, node properties and full pipeline are in the companion dataset repository: [kozo2/metabolomics-edges-expected-ge5](https://huggingface.co/datasets/kozo2/metabolomics-edges-expected-ge5). ## Files | File | Contents | Size | |---|---|---| | `edge_ML_filtered_expected_ge5_n2v.pt` | `{embedding: [848, 128] float32, node_id: [848], args: {...}}` | 453 KB | | `edge_MLvsMW_filtered_expected_ge5_n2v.pt` | `{embedding: [15758, 128] float32, node_id: [15758], args: {...}}` | 8.6 MB | | `node2vec_model.py`, `config.py` | Model definition, training loop, embedding export | — | | `heldout_check.py` | The held-out link-prediction evaluation reported below | — | Each file records the exact hyperparameters that produced it under `args`, and `node_id[i]` is the original study/assay/feature ID of embedding row `i`. ## Using it ```python import torch ck = torch.load("edge_MLvsMW_filtered_expected_ge5_n2v.pt", weights_only=False) z = torch.nn.functional.normalize(ck["embedding"], dim=1) # cosine space idx = {node: i for i, node in enumerate(ck["node_id"])} q = z[idx["MTBLS311_0001_00001537"]] top = (z @ q).topk(11).indices[1:] # drop self print([ck["node_id"][i] for i in top]) ``` Scores are cosine similarities — that is the metric the embeddings were trained under random walks and evaluated with, so use normalized vectors rather than raw dot products. ## Training | | `edge_ML` | `edge_MLvsMW` | |---|---|---| | graph | 848 nodes, 3,697 edges | 15,758 nodes, 89,277 edges | | parameters | 108,544 | 2,017,024 | | `embedding_dim` | 128 | 128 | | `walk_length` / `context_size` | 20 / 10 | 20 / 10 | | `walks_per_node` / negatives | 10 / 1 | 10 / 1 | | `p` / `q` | 1.0 / 1.0 (unbiased) | 1.0 / 1.0 (unbiased) | | `batch_size` / `lr` / optimizer | 128 / 0.01 / SparseAdam | 128 / 0.01 / SparseAdam | | epochs / wall time | 200 / 31 s | 200 / 3 m 05 s | | loss, first → last | 9.02 → 0.815 | 6.73 → 0.876 | Trained on one NVIDIA H100 NVL. `sparse=True` makes the embedding table the only parameter tensor, which is why the optimizer is `SparseAdam`. ## Evaluation `heldout_check.py` deduplicates undirected edges, holds out a random 10%, retrains from scratch on the remaining 90%, then scores held-out edges against an equal number of sampled non-edges by cosine similarity. The control is the preferential-attachment baseline `d_u · d_v` on training degrees, which answers whether the embedding learned anything beyond "popular nodes connect". | | `edge_ML` | `edge_MLvsMW` | |---|---|---| | train / held-out edges | 3,328 / 369 | 80,350 / 8,927 | | nodes isolated by the split | 42 | 413 | | **held-out AUC (never seen)** | **0.932** | **0.928** | | train AUC (in-sample) | 0.985 | 0.985 | | degree baseline (held-out) | 0.826 | 0.888 | | margin over baseline | +0.106 | +0.040 | **Species neighbour purity** — of each node's 10 nearest embeddings by cosine, the fraction sharing its species label — is 48.8% overall for `edge_ML` and 59.2% for `edge_MLvsMW`, rising to 75–79% for the dominant host species. For the bipartite `edge_MLvsMW` graph, the same measure on *source database* gives 67.5% overall (75.4% ST, 47.7% MTBLS). ## Caveats - **Training budget matters more than it looks.** At 20 epochs, `edge_ML` scored a held-out AUC of 0.791 — *below* its own 0.826 degree baseline. The embeddings only beat preferential attachment once training converges. `--epochs` defaults to 200 for this reason; do not shorten it without re-running `heldout_check.py`. - On `edge_MLvsMW` the margin over the degree baseline is thin (+0.040). That graph is bipartite with a few very high-degree studies, so much of its link structure is explained by degree alone. - `edge_MLvsMW` loss is flat from about epoch 19, so 200 epochs is more budget than that graph needs. - Embeddings are transductive: there is no way to embed a node that was not in the training graph without retraining.