Spaces:
Running
title: AI Resources
emoji: 📚
colorFrom: gray
colorTo: indigo
sdk: static
pinned: false
license: mit
short_description: A List of Foundational AI / ML / LLM Papers
AI-resources
A List of Foundational AI / ML / LLM Papers
Encoder-Decoder Architecture
Attention Is All You Need - Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin link
Parallel Distribution Processing
Microprocessor Architecture
Major Themes of Modern LLM Transformers:
Attention
Deep Learning (NN)
Hardware/GPU alignments
Rectified Linear Units Improve Restricted Boltzmann Machines - Vinod Nair, Geoffrey E. Hinton link
ZeRo: Memory Optimizations Toward Training Trillion Parameter Models - Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong He
Beyond Regression: New Tools For Prediction and Analysis in the Behavioral Sciences - Paul John Webos
ImageNet Classification with Dep Convolutional Neural Networks - Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton link
Taylor Expansion of the Accumulated Rounding Erro - Seppo Linnainmaa link
Learning to Translate in Real-time with Neural Machine Translation - Jiatao Gu, Graham Neubig, Kyunghyun Cho, Victor O.k. Li link
Effective Approaches to Attention-base Neural Machine Translation - Minh-Thang Luong, Hieu Pham, Christopher D. Manning link
Neural Machine Translation with Supervised Attention - Lemao Liu, Masao Utiyama, Andrew Finch, Eiichiro Sumita link
A Structured Self-Attentive Sentence Embedding - Zhouhan Lin, Minwei Feng, Cicero Nogueria dos Santos, Mo Yu, Bing Xiang, Bown Zhou, Yoshua Bengio link
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness - Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re link
Learning Representations by back-propagating error - David E. Rumelhard, Geoffrey E. Hinton, Ronald J. Williams link
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Paralleism - Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, Bryan Catazaro link
End-to-end Memory Networks - Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, Rob Fergus link
Memory Networks - Jason Weston, Sumit Chopra, Antoine Bordes link
Neural Machine Translation by Jointly Learning to Align and Translate - Dzmitry Bahdanau, KyungHyun Cho, Yoshua Bengio link
Approximation by Superpositions of a Sigmoidal Function - G. Cybenko link