On Linear Mode Connectivity of Mixture-of-Experts Architectures
This repository accompanies the paper: βOn Linear Mode Connectivity of Mixture-of-Experts Architecturesβ (Neurips 2025 Submission)
ImageNet: Linear Mode Connectivity
Installation
git clone https://github.com/repo/lmc-moe.git
cd moe-lmc
pip install -e .
pip install -r requirements.txt
Repository Structure
src/
βββ agnews/ # Appendix experiment: Reinit FFN
βββ cifar10/ # Main experiment
βββ cifar100/ # Main experiment
βββ dbpedia/ # Appendix experiment: Reinit FFN
βββ enwik8/ # Appendix experiment: Reinit FFN
βββ imagenet/ # Main experiment
βββ imdbreview/ # Appendix experiment: Reinit FFN
βββ lm1b/ # Main experiment
βββ mnist/ # Main experiment
βββ penn/ # Appendix experiment: Reinit FFN
βββ transfer_learning/ # Main experiment
βββ wikitext103/ # Main experiment
βββ datasets.py
βββ utils.py
βββ weight_matching.py
βββ online_stats.py
Each dataset directory includes a standalone README.md with detailed steps for data preparation, training, and evaluation.
Linear Mode Connectivity Results
ImageNet, WikiText103, One Billion Word (lm1b)
WikiText103: Linear Mode Connectivity
One Billion Word (LM1B): Linear Mode Connectivity
Getting Started
Each dataset experiment can be run individually. See the corresponding src/<dataset>/README.md for configuration options.
Citation
If you find this work helpful, please consider citing:
@article{our2025moelmc,
title={On Linear Mode Connectivity of Mixture-of-Experts Architectures},
author={Coauthors},
journal={arXiv:XXXX.XXXXX},
year={2025}
}
Acknowledgements
We thank contributors and maintainers of open-source libraries including PyTorch, JAX, Flax, and HuggingFace Transformers. Special thanks to the authors of recent works on LMC and MoE architectures for foundational insights.
Contributing
We welcome pull requests and suggestions. Please ensure new features or bug fixes include tests where appropriate and follow existing code style.
License
This project is licensed under the MIT License.