cs3319-project2 / notes /experiment_history.md
NLP-beginner's picture
CS3319 Project 2 final deliverable (public F1 = 0.96626)
f28d994
|
Raw
History Blame Contribute Delete
7.96 kB

Experiment History

This note records the main experiments and what they were intended to test.

1. Official Notebook Baseline

Script: code/run_baseline.py

Purpose:

  • Convert the official notebook into a local Python script.
  • Build a heterogeneous graph with authors and papers.
  • Use a simple custom HeteroMeanConv layer.
  • Train with a hinge/ranking-style objective.
  • Predict with cosine similarity and validation threshold search.

Outcome:

  • Validation F1 was around 0.885.
  • Generated /home/lzc/submission.csv.
  • This was the earliest upload candidate, but the exact leaderboard score was not preserved.

Main limitations:

  • Custom mean aggregation was weak.
  • Negative sampling was simple.
  • No strong hard-negative strategy.
  • Only a single model.

2. Improved Heterogeneous GNN

Script: code/run_improved.py

Purpose:

  • Replace the custom baseline with PyG HeteroConv + SAGEConv.
  • Add residual connections, layer normalization, and dropout.
  • Use an MLP decoder.
  • Add graph structural features.
  • Use hard negatives:
    • random negatives.
    • popular papers.
    • papers read by coauthors.
  • Train two seeds and ensemble.

Outcome:

  • Better than the notebook-style baseline, but still below the later LightGCN models.
  • Generated:
    • submission_improved.csv
    • submission_model1.csv
    • submission_model2.csv

Conclusion:

  • More expressive SAGE/MLP models did not generalize as well as the simpler LightGCN-style model.

3. SAGEConv + BPR Ranking

Script: code/run_v2.py

Purpose:

  • Keep a SAGEConv heterogeneous encoder.
  • Switch to BPR ranking loss.
  • Use cosine similarity for prediction.
  • Force train/test overlap pairs to positive.
  • Ensemble two seeds.

Outcome:

  • Improved over the BCE/MLP version in recommendation-style validation.
  • Generated:
    • submission_v2.csv
    • submission_v2_m1.csv
    • submission_v2_m2.csv

Conclusion:

  • BPR was clearly better aligned with the recommendation task than BCE.
  • However, SAGEConv transformations still underperformed vanilla LightGCN.

4. Vanilla LightGCN Main Line

Scripts:

  • code/run_final.py
  • code/run_lgcn_final.py

Purpose:

  • Use a LightGCN-style architecture:
    • trainable author embeddings.
    • linear projection for paper features.
    • mean aggregation over heterogeneous relations.
    • no nonlinear activation inside propagation.
    • no learned transform per propagation layer.
    • average layer embeddings.
  • Use BPR loss.
  • Use hard negative sampling.
  • Train multiple seeds on full training data.
  • Sweep decision thresholds.

Important checkpoints:

  • model_lgcn_s0.pt
  • model_lgcn_s42.pt
  • model_lgcn_s2024.pt
  • model_lgcn_s10.pt
  • model_lgcn_s100.pt

Generated submissions:

  • sub_lgcn_t0.30.csv through sub_lgcn_t0.50.csv

Outcome:

  • Validation F1 around 0.934 to 0.935.
  • This became the strongest stable model family.

Conclusion:

  • LightGCN's simplicity was a strength for this recommendation task.
  • Nonlinear GNNs and extra feature transformations tended to overfit or degrade ranking quality.

5. Confirmed 6-Model LightGCN Ensemble

Script source:

  • Five 256d checkpoints came from code/run_lgcn_final.py.
  • The sixth 384d checkpoint and the 6-model CSV ensemble were produced in a previous inline Python run.
  • This package includes code/generate_ens6_submission.py as the portable regeneration script.

Method:

  • Ensemble members:
    • 5 x LightGCN, 256 dimensions, 4 layers.
    • 1 x LightGCN, 384 dimensions, 4 layers.
  • Score:
    • cosine similarity between author and paper embeddings.
  • Ensemble:
    • arithmetic mean of six score vectors.
  • Post-processing:
    • train/test overlap pairs forced to label 1.
  • Threshold:
    • 0.36 for confirmed submission.

Confirmed CSV:

  • submissions/sub_ens6_t0.36.csv

Confirmed public leaderboard F1:

  • 0.93044

Nearby threshold files included:

  • sub_ens6_t0.35.csv
  • sub_ens6_t0.37.csv
  • sub_ens6_t0.38.csv

Conclusion:

  • This is the best confirmed result in the preserved records.

6. L2-Normalized LightGCN Variant

Script: code/run_lgcn_v2.py

Purpose:

  • Normalize embeddings during training and evaluation.
  • Align training scores with cosine similarity.

Outcome:

  • Underperformed vanilla LightGCN.

Conclusion:

  • L2 normalization during training reduced useful ranking signal.
  • The better setup was to train with BPR on dot products and evaluate with cosine similarity.

7. LightGBM Structural Feature Model

Script: code/run_graph_features.py

Purpose:

  • Build a non-GNN baseline and potential ensemble component.
  • Compute hand-engineered graph features for author-paper pairs:
    • author degree.
    • paper degree.
    • coauthor degree.
    • citation in/out degree.
    • preferential attachment.
    • popularity percentiles.
    • whether a coauthor read the paper.
    • author average paper embedding cosine similarity.
    • cold-start indicators.
  • Train LightGBM.

Cached outputs:

  • cached_scores/lgb_model.pkl
  • cached_scores/lgb_v2_model.pkl
  • cached_scores/test_lgb_scores.npy
  • cached_scores/test_lgb_v2_scores.npy

Outcome:

  • Validation F1 around 0.897 in the preserved notes.
  • Useful as a baseline and ensemble attempt, but not stronger than LightGCN.

Conclusion:

  • Structural features are good for reporting and comparison, but the GNN captured the main signal better.

8. BPR/MF-Style Baseline

Cached outputs:

  • cached_scores/test_bpr_cos.npy
  • cached_scores/test_bpr_dot.npy

Purpose:

  • Test a more traditional recommendation baseline.
  • Provide a diverse signal for ensembles.

Outcome:

  • Validation was around 0.90 in previous notes.
  • Predictions were highly correlated with GNN scores, so ensemble gains were limited.

Conclusion:

  • Useful baseline, but not a replacement for LightGCN.

9. GNN Architecture Search

Script: code/compare_gnn.py

Architectures compared:

  • Vanilla LightGCN, 256d, 4 layers.
  • Learnable layer weights.
  • GAT aggregation.
  • SAGE aggregation.
  • Deep LightGCN.
  • Wide LightGCN, 384d.
  • Vanilla LightGCN, 5 layers.
  • GAT with more heads/layers.

Preserved results from previous session:

Architecture Validation F1
Vanilla LightGCN, 256d, 4L about 0.9350
Learnable layer weights, 256d, 4L about 0.9360
Wide LightGCN, 384d, 3L about 0.9365
Wide LightGCN, 512d, 3L about 0.9363
Deep LightGCN, 256d, 6L about 0.9319
GAT aggregation about 0.9331
SAGE aggregation about 0.9185

Conclusion:

  • Wider models gave tiny validation gains, but did not clearly improve public leaderboard behavior.
  • Deep/nonlinear/attention/SAGE variants were not better.
  • The final practical choice remained vanilla LightGCN.

10. Later Threshold and Force/No-Force Experiments

Generated after the confirmed 0.93044 score:

  • sub_final3_t*.csv
  • sub_noforce_t*.csv
  • sub_forced_t*.csv

Purpose:

  • Try more seeds and threshold sweeps.
  • Test whether forcing known positives was helping.
  • Search better thresholds based on public leaderboard feedback.

Included in this package:

  • sub_final3_t0.35.csv
  • sub_final3_t0.36.csv
  • sub_final3_t0.37.csv
  • sub_noforce_t0.34.csv
  • sub_noforce_t0.36.csv
  • sub_noforce_t0.38.csv
  • sub_forced_t0.36.csv

Outcome:

  • No preserved Kaggle scores are available for these later files.
  • They should be treated as candidates, not confirmed improvements.

Practical Recommendation

For transferring and resuming work:

  1. Start from submissions/sub_ens6_t0.36.csv as the confirmed baseline.
  2. Recreate environment with env/environment-cs3319.yml.
  3. Verify code/generate_ens6_submission.py runs on the new machine.
  4. If more Kaggle submissions are available, tune thresholds around the final ensemble:
    • 0.35
    • 0.36
    • 0.37
    • 0.38
  5. For the course report, present:
    • official baseline.
    • SAGE/BPR model.
    • LightGBM structural baseline.
    • LightGCN final model.
    • architecture ablation from compare_gnn.py.