bitnet Collection by Mritunjayk Mar 28, 2024 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits Paper • 2402.17764 • Published Feb 27, 2024 • 630
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits Paper • 2402.17764 • Published Feb 27, 2024 • 630
eee-555-444 Collection by emoaba Mar 28, 2024 - Paused Agents 633 Fast Stable Diffusion XL (SDXL) 🔥 633
ITRF Instruction tuning a model on RAG and then finetuning a model on RAG data and then a reranker on the transition probabilities (Thesis weights). Collection by tristanratz Mar 28, 2024 - tristanratz/itrf-13b-4bit-qlora Updated Mar 10, 2024 tristanratz/itrf-7b-lora Updated Mar 9, 2024 tristanratz/itrf-reranker Text Classification • 0.6B • Updated Mar 28, 2024 • 8
thoery Collection by thuzhizhi Apr 1, 2024 - The Unreasonable Ineffectiveness of the Deeper Layers Paper • 2403.17887 • Published Mar 26, 2024 • 82 Long-form factuality in large language models Paper • 2403.18802 • Published Mar 27, 2024 • 26 Jamba: A Hybrid Transformer-Mamba Language Model Paper • 2403.19887 • Published Mar 28, 2024 • 113
The Unreasonable Ineffectiveness of the Deeper Layers Paper • 2403.17887 • Published Mar 26, 2024 • 82
Jamba: A Hybrid Transformer-Mamba Language Model Paper • 2403.19887 • Published Mar 28, 2024 • 113
YALLM Collection by sourceoftruthdata Sep 16, 2024 - Sleeping Agents 1 WebpyGPT 🤖 1 Generate conversational responses based on user input google/timesfm-1.0-200m Time Series Forecasting • Updated May 17, 2024 • 600 • 836 Running Agents 145 Llama3.1 Instruct O1 🌖 145 Generate detailed, step‑by‑step answers with Llama3.1 chat
Running Agents 145 Llama3.1 Instruct O1 🌖 145 Generate detailed, step‑by‑step answers with Llama3.1 chat
qwen1.5 MOE Collection by zengxisheng Mar 28, 2024 - Qwen/Qwen1.5-MoE-A2.7B-Chat Text Generation • 14B • Updated Apr 30, 2024 • 42.7k • 133
💼 Fine-Tuned CO-Funer Models My fine-tuned Flair models on CO-FUN NER Dataset Collection by stefan-it Mar 28, 2024 - stefan-it/flair-co-funer-gbert_base-bs8-e10-lr5e-05-3 Token Classification • Updated Mar 28, 2024 • 6 stefan-it/flair-co-funer-german_dbmdz_bert_base-bs8-e10-lr5e-05-1 Token Classification • Updated Mar 28, 2024 stefan-it/flair-co-funer-german_bert_base-bs8-e10-lr5e-05-2 Token Classification • Updated Mar 28, 2024 • 1
stefan-it/flair-co-funer-gbert_base-bs8-e10-lr5e-05-3 Token Classification • Updated Mar 28, 2024 • 6
stefan-it/flair-co-funer-german_dbmdz_bert_base-bs8-e10-lr5e-05-1 Token Classification • Updated Mar 28, 2024
stefan-it/flair-co-funer-german_bert_base-bs8-e10-lr5e-05-2 Token Classification • Updated Mar 28, 2024 • 1
bitnet Collection by Mritunjayk Mar 28, 2024 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits Paper • 2402.17764 • Published Feb 27, 2024 • 630
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits Paper • 2402.17764 • Published Feb 27, 2024 • 630
YALLM Collection by sourceoftruthdata Sep 16, 2024 - Sleeping Agents 1 WebpyGPT 🤖 1 Generate conversational responses based on user input google/timesfm-1.0-200m Time Series Forecasting • Updated May 17, 2024 • 600 • 836 Running Agents 145 Llama3.1 Instruct O1 🌖 145 Generate detailed, step‑by‑step answers with Llama3.1 chat
Running Agents 145 Llama3.1 Instruct O1 🌖 145 Generate detailed, step‑by‑step answers with Llama3.1 chat
eee-555-444 Collection by emoaba Mar 28, 2024 - Paused Agents 633 Fast Stable Diffusion XL (SDXL) 🔥 633
qwen1.5 MOE Collection by zengxisheng Mar 28, 2024 - Qwen/Qwen1.5-MoE-A2.7B-Chat Text Generation • 14B • Updated Apr 30, 2024 • 42.7k • 133
ITRF Instruction tuning a model on RAG and then finetuning a model on RAG data and then a reranker on the transition probabilities (Thesis weights). Collection by tristanratz Mar 28, 2024 - tristanratz/itrf-13b-4bit-qlora Updated Mar 10, 2024 tristanratz/itrf-7b-lora Updated Mar 9, 2024 tristanratz/itrf-reranker Text Classification • 0.6B • Updated Mar 28, 2024 • 8
thoery Collection by thuzhizhi Apr 1, 2024 - The Unreasonable Ineffectiveness of the Deeper Layers Paper • 2403.17887 • Published Mar 26, 2024 • 82 Long-form factuality in large language models Paper • 2403.18802 • Published Mar 27, 2024 • 26 Jamba: A Hybrid Transformer-Mamba Language Model Paper • 2403.19887 • Published Mar 28, 2024 • 113
The Unreasonable Ineffectiveness of the Deeper Layers Paper • 2403.17887 • Published Mar 26, 2024 • 82
Jamba: A Hybrid Transformer-Mamba Language Model Paper • 2403.19887 • Published Mar 28, 2024 • 113
💼 Fine-Tuned CO-Funer Models My fine-tuned Flair models on CO-FUN NER Dataset Collection by stefan-it Mar 28, 2024 - stefan-it/flair-co-funer-gbert_base-bs8-e10-lr5e-05-3 Token Classification • Updated Mar 28, 2024 • 6 stefan-it/flair-co-funer-german_dbmdz_bert_base-bs8-e10-lr5e-05-1 Token Classification • Updated Mar 28, 2024 stefan-it/flair-co-funer-german_bert_base-bs8-e10-lr5e-05-2 Token Classification • Updated Mar 28, 2024 • 1
stefan-it/flair-co-funer-gbert_base-bs8-e10-lr5e-05-3 Token Classification • Updated Mar 28, 2024 • 6
stefan-it/flair-co-funer-german_dbmdz_bert_base-bs8-e10-lr5e-05-1 Token Classification • Updated Mar 28, 2024
stefan-it/flair-co-funer-german_bert_base-bs8-e10-lr5e-05-2 Token Classification • Updated Mar 28, 2024 • 1