How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf mimaapp/sintesi-models:Q4_0
# Run inference directly in the terminal:
llama cli -hf mimaapp/sintesi-models:Q4_0
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf mimaapp/sintesi-models:Q4_0
# Run inference directly in the terminal:
llama cli -hf mimaapp/sintesi-models:Q4_0
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf mimaapp/sintesi-models:Q4_0
# Run inference directly in the terminal:
./llama-cli -hf mimaapp/sintesi-models:Q4_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf mimaapp/sintesi-models:Q4_0
# Run inference directly in the terminal:
./build/bin/llama-cli -hf mimaapp/sintesi-models:Q4_0
Use Docker
docker model run hf.co/mimaapp/sintesi-models:Q4_0
Quick Links

Sintesi โ€” on-device models

Mirror of the three model files downloaded by the Sintesi Android app (com.mimaapp.sintesi) to transcribe and rewrite voice notes on the phone, without sending recordings anywhere. The files are byte-identical copies of the originals; they are hosted here only so the app does not depend on someone else's repository.

Each file keeps its own license. There is no single license for the whole repository.

File Original License
ggml-parakeet-tdt-0.6b-v3-q8_0.bin ggml-org/parakeet-GGUF, a q8_0 conversion of nvidia/parakeet-tdt-0.6b-v3 CC-BY-4.0
ggml-silero-v5.1.2.bin ggml-org/whisper-vad, a conversion of Silero VAD v5.1.2 MIT
Qwen3-4B-Instruct-2507-Q4_0.gguf unsloth/Qwen3-4B-Instruct-2507-GGUF, a Q4_0 quantization of Qwen/Qwen3-4B-Instruct-2507 Apache 2.0

Attribution

  • Parakeet TDT 0.6B v3 by NVIDIA, licensed under CC-BY-4.0. Changes: converted and quantized (q8_0) to the ggml format by ggml-org; no further changes.
  • Silero VAD by Silero Team, licensed under MIT (copyright notice and license text in LICENSE-Silero-VAD-MIT.txt). Converted to the ggml format by ggml-org; no further changes.
  • Qwen3-4B-Instruct-2507 by the Qwen team, Alibaba Cloud, licensed under Apache 2.0 (license text in LICENSE-Qwen3-Apache-2.0.txt). Changes: quantized (Q4_0) to the GGUF format by Unsloth; no further changes.

Checksums (SHA-256)

4d64e9e96c2792186d072fde0034df0ad670cf680a2f53069052ead827fd600e  ggml-parakeet-tdt-0.6b-v3-q8_0.bin
29940d98d42b91fbd05ce489f3ecf7c72f0a42f027e4875919a28fb4c04ea2cf  ggml-silero-v5.1.2.bin
e0ba675d86ab277c61701c6793659b2ae801d95e3be791464c321e6fbf613be2  Qwen3-4B-Instruct-2507-Q4_0.gguf
Downloads last month
-
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support