How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Trilogix1/Hugston-Macaron-V1-Tall:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Trilogix1/Hugston-Macaron-V1-Tall:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Trilogix1/Hugston-Macaron-V1-Tall:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Trilogix1/Hugston-Macaron-V1-Tall:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Trilogix1/Hugston-Macaron-V1-Tall:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf Trilogix1/Hugston-Macaron-V1-Tall:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Trilogix1/Hugston-Macaron-V1-Tall:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf Trilogix1/Hugston-Macaron-V1-Tall:Q4_K_M
Use Docker
docker model run hf.co/Trilogix1/Hugston-Macaron-V1-Tall:Q4_K_M
Quick Links

This is an optimized model of Qwen from Mind Lab then GGUFED using Quanta and HugstonOne from Hugston.


Use it with HugstonOne (the new release, Video-Audio-Image processing) , the most powerful local AI with strongest privacy worldwide to date.


Whitepaper: https://doi.org/10.5281/zenodo.21600618 ---------------- Get the latest update here: https://hugston.com

05_coding_mode 06_agents_rag_cli 01_internet_modes 02_live_chat 03_preview 04_context_runtime


Worth mentioning additional features, Video-Audio-Image processing,-------------------Agents, skills, tools. ---------------Rag in Terabytes-------------Online mode visual interactive or background, --------Most of libraries load in preview totally offline,-------------Hugston Private API included.


The aim is to test and understand mechanism and accuracy of different llm models for research purposes.


Credit to Qwen for the model creation

Credit to https://huggingface.co/mindlab-research for the optimization method

Credit to LLama.cpp team for the great contribution

Credit to Hugston Team for Converting, Quantizing, Testing, Benching and other...

Credit to Huggingface for the amazing hosting platform


Here we show Quanta our convertor and Quantizer tool.

image

Downloads last month
4,553
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Trilogix1/Hugston-Macaron-V1-Tall

Quantized
(11)
this model