Quantization/Quantized Collection Quantization in LLMs compresses model weights from high-precision formats (like 16-bit) to lower-precision formats (like 8-bit or 4-bit). • 440 items • Updated 2 days ago
Fine-tuning/Fine-tuned Collection Fine-tuning generative AI is the process of taking a pre-trained base model and training it further on a smaller, specific dataset. • 335 items • Updated 2 days ago
Importance Matrix (Imatrix) Collection The Importance Matrix (iMatrix) is a data-driven calibration method used in low-bit quantization for LLMs. • 330 items • Updated 12 days ago
GPT-Generated Unified Format (GGUF) Collection GPT-Generated Unified Format (GGUF) is a single-file binary format used to store and run large language models efficiently on consumer hardware. • 707 items • Updated 12 days ago
Quantization/Quantized Collection Quantization in LLMs compresses model weights from high-precision formats (like 16-bit) to lower-precision formats (like 8-bit or 4-bit). • 440 items • Updated 2 days ago
Mixture of Experts (MoE) Collection Mixture of Experts (MoE) is a machine learning architecture that splits a large neural network into smaller sub-networks called "experts." • 202 items • Updated 12 days ago
GPT-Generated Unified Format (GGUF) Collection GPT-Generated Unified Format (GGUF) is a single-file binary format used to store and run large language models efficiently on consumer hardware. • 707 items • Updated 12 days ago
Automatic-Speech-Recognition (ASR) Collection Automatic Speech Recognition (ASR) converts spoken audio into text. In generative AI advanced ASR models act as the ears of large multimodal models. • 5 items • Updated 14 days ago
nvidia/nemotron-speech-streaming-en-0.6b Automatic Speech Recognition • 0.6B • Updated 23 days ago • 157k • 613
Agentic Generative AI Collection Agentic generative AI refers to autonomous systems built on LLMs that can plan, use tools, and execute multi-step workflows. • 57 items • Updated 17 days ago
Fine-tuning/Fine-tuned Collection Fine-tuning generative AI is the process of taking a pre-trained base model and training it further on a smaller, specific dataset. • 335 items • Updated 2 days ago
Quantization/Quantized Collection Quantization in LLMs compresses model weights from high-precision formats (like 16-bit) to lower-precision formats (like 8-bit or 4-bit). • 440 items • Updated 2 days ago