-
Qwen/Qwen2.5-3B-Instruct
Text Generation • 3B • Updated • 6.73M • • 550 -
cais/mmlu
Viewer • Updated • 231k • 478k • 812 -
SabaPivot/repro-finetuning-without-forgetting-icl-job
1.85 MB -
Repro - Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
🎯Collaborate on research notes with an AI coding agent
Collections
Discover the best community collections!
Collections trending this week
-
mhndayesh/gemma-4-E2B-netsec-expert-GGUF
Text Generation • 5B • Updated • 447 -
mhndayesh/gemma-4-26B-A4B-netsec-expert-GGUF
Text Generation • 25B • Updated • 470 -
mhndayesh/gemma-4-12B-netsec-expert-GGUF
Text Generation • 12B • Updated • 460 -
mhndayesh/gemma-4-E2B-offsec-expert-GGUF
Text Generation • 5B • Updated • 467
-
huihui-ai/Llama-3.2-1B-Instruct-abliterated
Text Generation • 1B • Updated • 31.9k • 13 -
Qwen/Qwen2.5-0.5B-Instruct-GGUF
Text Generation • 0.6B • Updated • 154k • 120 -
huihui-ai/Hermes-3-Llama-3.2-3B-abliterated
Text Generation • 3B • Updated • 81 • 6 -
bartowski/Qwen2.5-Coder-3B-Instruct-abliterated-GGUF
Text Generation • 3B • Updated • 4.7k • 18
-
Qwen/Qwen2.5-3B-Instruct
Text Generation • 3B • Updated • 6.73M • • 550 -
cais/mmlu
Viewer • Updated • 231k • 478k • 812 -
SabaPivot/repro-finetuning-without-forgetting-icl-job
1.85 MB -
Repro - Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
🎯Collaborate on research notes with an AI coding agent
-
mhndayesh/gemma-4-E2B-netsec-expert-GGUF
Text Generation • 5B • Updated • 447 -
mhndayesh/gemma-4-26B-A4B-netsec-expert-GGUF
Text Generation • 25B • Updated • 470 -
mhndayesh/gemma-4-12B-netsec-expert-GGUF
Text Generation • 12B • Updated • 460 -
mhndayesh/gemma-4-E2B-offsec-expert-GGUF
Text Generation • 5B • Updated • 467
-
huihui-ai/Llama-3.2-1B-Instruct-abliterated
Text Generation • 1B • Updated • 31.9k • 13 -
Qwen/Qwen2.5-0.5B-Instruct-GGUF
Text Generation • 0.6B • Updated • 154k • 120 -
huihui-ai/Hermes-3-Llama-3.2-3B-abliterated
Text Generation • 3B • Updated • 81 • 6 -
bartowski/Qwen2.5-Coder-3B-Instruct-abliterated-GGUF
Text Generation • 3B • Updated • 4.7k • 18