# Experiment History This note records the main experiments and what they were intended to test. ## 1. Official Notebook Baseline Script: `code/run_baseline.py` Purpose: - Convert the official notebook into a local Python script. - Build a heterogeneous graph with authors and papers. - Use a simple custom `HeteroMeanConv` layer. - Train with a hinge/ranking-style objective. - Predict with cosine similarity and validation threshold search. Outcome: - Validation F1 was around 0.885. - Generated `/home/lzc/submission.csv`. - This was the earliest upload candidate, but the exact leaderboard score was not preserved. Main limitations: - Custom mean aggregation was weak. - Negative sampling was simple. - No strong hard-negative strategy. - Only a single model. ## 2. Improved Heterogeneous GNN Script: `code/run_improved.py` Purpose: - Replace the custom baseline with PyG `HeteroConv` + `SAGEConv`. - Add residual connections, layer normalization, and dropout. - Use an MLP decoder. - Add graph structural features. - Use hard negatives: - random negatives. - popular papers. - papers read by coauthors. - Train two seeds and ensemble. Outcome: - Better than the notebook-style baseline, but still below the later LightGCN models. - Generated: - `submission_improved.csv` - `submission_model1.csv` - `submission_model2.csv` Conclusion: - More expressive SAGE/MLP models did not generalize as well as the simpler LightGCN-style model. ## 3. SAGEConv + BPR Ranking Script: `code/run_v2.py` Purpose: - Keep a SAGEConv heterogeneous encoder. - Switch to BPR ranking loss. - Use cosine similarity for prediction. - Force train/test overlap pairs to positive. - Ensemble two seeds. Outcome: - Improved over the BCE/MLP version in recommendation-style validation. - Generated: - `submission_v2.csv` - `submission_v2_m1.csv` - `submission_v2_m2.csv` Conclusion: - BPR was clearly better aligned with the recommendation task than BCE. - However, SAGEConv transformations still underperformed vanilla LightGCN. ## 4. Vanilla LightGCN Main Line Scripts: - `code/run_final.py` - `code/run_lgcn_final.py` Purpose: - Use a LightGCN-style architecture: - trainable author embeddings. - linear projection for paper features. - mean aggregation over heterogeneous relations. - no nonlinear activation inside propagation. - no learned transform per propagation layer. - average layer embeddings. - Use BPR loss. - Use hard negative sampling. - Train multiple seeds on full training data. - Sweep decision thresholds. Important checkpoints: - `model_lgcn_s0.pt` - `model_lgcn_s42.pt` - `model_lgcn_s2024.pt` - `model_lgcn_s10.pt` - `model_lgcn_s100.pt` Generated submissions: - `sub_lgcn_t0.30.csv` through `sub_lgcn_t0.50.csv` Outcome: - Validation F1 around 0.934 to 0.935. - This became the strongest stable model family. Conclusion: - LightGCN's simplicity was a strength for this recommendation task. - Nonlinear GNNs and extra feature transformations tended to overfit or degrade ranking quality. ## 5. Confirmed 6-Model LightGCN Ensemble Script source: - Five 256d checkpoints came from `code/run_lgcn_final.py`. - The sixth 384d checkpoint and the 6-model CSV ensemble were produced in a previous inline Python run. - This package includes `code/generate_ens6_submission.py` as the portable regeneration script. Method: - Ensemble members: - 5 x LightGCN, 256 dimensions, 4 layers. - 1 x LightGCN, 384 dimensions, 4 layers. - Score: - cosine similarity between author and paper embeddings. - Ensemble: - arithmetic mean of six score vectors. - Post-processing: - train/test overlap pairs forced to label 1. - Threshold: - `0.36` for confirmed submission. Confirmed CSV: - `submissions/sub_ens6_t0.36.csv` Confirmed public leaderboard F1: - `0.93044` Nearby threshold files included: - `sub_ens6_t0.35.csv` - `sub_ens6_t0.37.csv` - `sub_ens6_t0.38.csv` Conclusion: - This is the best confirmed result in the preserved records. ## 6. L2-Normalized LightGCN Variant Script: `code/run_lgcn_v2.py` Purpose: - Normalize embeddings during training and evaluation. - Align training scores with cosine similarity. Outcome: - Underperformed vanilla LightGCN. Conclusion: - L2 normalization during training reduced useful ranking signal. - The better setup was to train with BPR on dot products and evaluate with cosine similarity. ## 7. LightGBM Structural Feature Model Script: `code/run_graph_features.py` Purpose: - Build a non-GNN baseline and potential ensemble component. - Compute hand-engineered graph features for author-paper pairs: - author degree. - paper degree. - coauthor degree. - citation in/out degree. - preferential attachment. - popularity percentiles. - whether a coauthor read the paper. - author average paper embedding cosine similarity. - cold-start indicators. - Train LightGBM. Cached outputs: - `cached_scores/lgb_model.pkl` - `cached_scores/lgb_v2_model.pkl` - `cached_scores/test_lgb_scores.npy` - `cached_scores/test_lgb_v2_scores.npy` Outcome: - Validation F1 around 0.897 in the preserved notes. - Useful as a baseline and ensemble attempt, but not stronger than LightGCN. Conclusion: - Structural features are good for reporting and comparison, but the GNN captured the main signal better. ## 8. BPR/MF-Style Baseline Cached outputs: - `cached_scores/test_bpr_cos.npy` - `cached_scores/test_bpr_dot.npy` Purpose: - Test a more traditional recommendation baseline. - Provide a diverse signal for ensembles. Outcome: - Validation was around 0.90 in previous notes. - Predictions were highly correlated with GNN scores, so ensemble gains were limited. Conclusion: - Useful baseline, but not a replacement for LightGCN. ## 9. GNN Architecture Search Script: `code/compare_gnn.py` Architectures compared: - Vanilla LightGCN, 256d, 4 layers. - Learnable layer weights. - GAT aggregation. - SAGE aggregation. - Deep LightGCN. - Wide LightGCN, 384d. - Vanilla LightGCN, 5 layers. - GAT with more heads/layers. Preserved results from previous session: | Architecture | Validation F1 | |---|---:| | Vanilla LightGCN, 256d, 4L | about 0.9350 | | Learnable layer weights, 256d, 4L | about 0.9360 | | Wide LightGCN, 384d, 3L | about 0.9365 | | Wide LightGCN, 512d, 3L | about 0.9363 | | Deep LightGCN, 256d, 6L | about 0.9319 | | GAT aggregation | about 0.9331 | | SAGE aggregation | about 0.9185 | Conclusion: - Wider models gave tiny validation gains, but did not clearly improve public leaderboard behavior. - Deep/nonlinear/attention/SAGE variants were not better. - The final practical choice remained vanilla LightGCN. ## 10. Later Threshold and Force/No-Force Experiments Generated after the confirmed `0.93044` score: - `sub_final3_t*.csv` - `sub_noforce_t*.csv` - `sub_forced_t*.csv` Purpose: - Try more seeds and threshold sweeps. - Test whether forcing known positives was helping. - Search better thresholds based on public leaderboard feedback. Included in this package: - `sub_final3_t0.35.csv` - `sub_final3_t0.36.csv` - `sub_final3_t0.37.csv` - `sub_noforce_t0.34.csv` - `sub_noforce_t0.36.csv` - `sub_noforce_t0.38.csv` - `sub_forced_t0.36.csv` Outcome: - No preserved Kaggle scores are available for these later files. - They should be treated as candidates, not confirmed improvements. ## Practical Recommendation For transferring and resuming work: 1. Start from `submissions/sub_ens6_t0.36.csv` as the confirmed baseline. 2. Recreate environment with `env/environment-cs3319.yml`. 3. Verify `code/generate_ens6_submission.py` runs on the new machine. 4. If more Kaggle submissions are available, tune thresholds around the final ensemble: - `0.35` - `0.36` - `0.37` - `0.38` 5. For the course report, present: - official baseline. - SAGE/BPR model. - LightGBM structural baseline. - LightGCN final model. - architecture ablation from `compare_gnn.py`.