--- language: - en license: apache-2.0 base_model: - mindlab-research/Macaron-V1-Tall tags: - gguf - llama.cpp - qwen - moe - macaron - mixture-of-lora --- # Macaron-V1-Tall (GGUF Pre-Merged Specialists) This repository contains the GGUF-quantized, pre-merged specialist models for **Macaron-V1-Tall**, optimized for local deployment on consumer hardware. ## 🏆 Acknowledgments and Credits All foundational research, training, and architecture design belong to the original authors. This repository solely provides a community-driven quantization and deployment format. * **Original Authors:** mindlab-research/Macaron-V1-Tall * **Original Repository:** [mindlab-research/Macaron-V1-Tall](https://huggingface.co/mindlab-research/Macaron-V1-Tall) * **Base Architecture:** Qwen 3.6 (35B Base + 4x 3.7B LoRA Specialists) ## 🧩 The Architectural Pivot (Why this repository exists) The Macaron-V1-Tall architecture utilizes a Mixture-of-LoRA (MoE) pattern over a Grouped-Query Attention (GQA) base model. Currently, the standard `llama.cpp` ecosystem encounters tensor reshaping limitations (`NotImplementedError`) when attempting to dynamically apply these specific LoRA adapters at runtime. To bypass this mathematical constraint and enable seamless local execution, this repository utilizes a **Pre-Merge Strategy**. Each LoRA specialist has been physically fused into the 35B base model using CPU-RAM computation (bypassing VRAM bottlenecks) prior to GGUF conversion. **The Result:** Four independent, fully fused GGUF models. Instead of dynamically swapping LoRAs in VRAM, developers can leverage operating system Page Caching to swiftly alternate between these ~21GB files in RAM, enabling rapid intent-based routing without out-of-memory (OOM) errors. ## 📦 Available Specialists (Q4_K_M Quantization) The model have been quantized to `Q4_K_M` to balance perplexity and memory footprint, reducing the required storage from ~70GB (FP16) to approximately **21GB per specialist**. * `macaron-l2-coding-35b-q4_k_m.gguf` (Coding Specialist) > **Note:** The experimental `mtp_num_hidden_layers` (Multi-Token Prediction) metadata has been sanitized from the config to ensure strict compatibility with the `llama.cpp` loader. ## 🚀 Deployment & Usage (llama.cpp) These models are heavily optimized for hybrid VRAM/RAM offloading. If you are running a consumer GPU (e.g., 16GB VRAM) backed by substantial system RAM, you must carefully balance the layers to prevent VRAM saturation. ### Example of a Launch Command (`llama-server`) The following command demonstrates how to load the model while explicitly offloading the heaviest MoE layers to the CPU, freeing up your VRAM for the KV cache and attention heads. ```bash ./llama-server \ -m "macaron-l2-coding-35b-q4_k_m.gguf" \ --host 127.0.0.1 \ --port 9091 \ -ngl 40 \ -ncmoe 10 \ --ctx-size 16384 \ -b 512 \ -fa on \ -ctk q4_0 \ -ctv q4_0 \ -t 16 \ --reasoning-preserve \ --no-context-shift \ -sps 0.0 \ --ctx-checkpoints 0 \ --models-max 1 \ --parallel 1 \ --no-mmap \ --metrics