--- title: AI Resources emoji: 📚 colorFrom: gray colorTo: indigo sdk: static pinned: false license: mit short_description: A List of Foundational AI / ML / LLM Papers --- # AI-resources A List of Foundational AI / ML / LLM Papers ## Encoder-Decoder Architecture Attention Is All You Need - Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin [link](https://arxiv.org/pdf/1706.03762) ## Parallel Distribution Processing ## Microprocessor Architecture --- # Major Themes of Modern LLM Transformers: ## Attention ## Deep Learning (NN) ## Hardware/GPU alignments --- Rectified Linear Units Improve Restricted Boltzmann Machines - Vinod Nair, Geoffrey E. Hinton [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/reluICML.pdf) ZeRo: Memory Optimizations Toward Training Trillion Parameter Models - Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong He Beyond Regression: New Tools For Prediction and Analysis in the Behavioral Sciences - Paul John Webos ImageNet Classification with Dep Convolutional Neural Networks - Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/NIPS-2012-imagenet-classification-with-deep-convolutional-neural-networks-Paper.pdf) Taylor Expansion of the Accumulated Rounding Erro - Seppo Linnainmaa [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/Linnainmaa-1976.pdf) Learning to Translate in Real-time with Neural Machine Translation - Jiatao Gu, Graham Neubig, Kyunghyun Cho, Victor O.k. Li [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/E17-1099.pdf) Effective Approaches to Attention-base Neural Machine Translation - Minh-Thang Luong, Hieu Pham, Christopher D. Manning [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/D15-1166.pdf) Neural Machine Translation with Supervised Attention - Lemao Liu, Masao Utiyama, Andrew Finch, Eiichiro Sumita [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/C16-1291.pdf) A Structured Self-Attentive Sentence Embedding - Zhouhan Lin, Minwei Feng, Cicero Nogueria dos Santos, Mo Yu, Bing Xiang, Bown Zhou, Yoshua Bengio [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/A_Structured_Self-attentive_Sentence_Embedding.pdf) FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness - Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/2205.14135v2.pdf) Learning Representations by back-propagating error - David E. Rumelhard, Geoffrey E. Hinton, Ronald J. Williams [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/1986-rumelhart-2.pdf) Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Paralleism - Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, Bryan Catazaro [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/1909.08053v4.pdf) End-to-end Memory Networks - Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, Rob Fergus [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/1503.08895v5.pdf) Memory Networks - Jason Weston, Sumit Chopra, Antoine Bordes [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/1410.3916v11.pdf) Neural Machine Translation by Jointly Learning to Align and Translate - Dzmitry Bahdanau, KyungHyun Cho, Yoshua Bengio [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/1409.0473v7.pdf) Approximation by Superpositions of a Sigmoidal Function - G. Cybenko [link](https://huggingface.co/spaces/alphonse86/AI-resources/resolve/main/papers/10.1.1.441.7873.pdf)