How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Fazmin/solus_v1_phi-3.5-mini-instruct-q4:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Fazmin/solus_v1_phi-3.5-mini-instruct-q4:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Fazmin/solus_v1_phi-3.5-mini-instruct-q4:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Fazmin/solus_v1_phi-3.5-mini-instruct-q4:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Fazmin/solus_v1_phi-3.5-mini-instruct-q4:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf Fazmin/solus_v1_phi-3.5-mini-instruct-q4:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Fazmin/solus_v1_phi-3.5-mini-instruct-q4:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf Fazmin/solus_v1_phi-3.5-mini-instruct-q4:Q4_K_M
Use Docker
docker model run hf.co/Fazmin/solus_v1_phi-3.5-mini-instruct-q4:Q4_K_M
Quick Links

Phi 3.5 Mini Instruct โ€” Solus v1

Microsoft's Phi-3.5-mini, a 3.8B parameter model trained largely on filtered public data and synthetic textbook-style material. That data recipe gives it reasoning and comprehension quality well above what its size suggests, and it is particularly good at reading a document and answering questions about it.

It is multilingual, permissively MIT licensed, and a strong pick for summarisation and document Q&A on modest hardware.

Specifications

Parameters 3.8B
Quantization Q4_K_M
File size 2.23 GB
Minimum RAM 6.00 GB
Minimum VRAM not required
Context length 16,384 tokens
SHA-256 e4165e3a71af97f1b4820da61079826d8752a2088e313af0c7d346796c38eff5

Single file: Phi-3.5-mini-instruct-Q4_K_M.gguf

Quantization

Quantization performed at the Faculty of Engineering, McMaster University.

The GGUF conversion this build is derived from was produced by bartowski, and the weights here are a byte-for-byte copy of that file โ€” the SHA-256 above matches the upstream artifact.

Provenance

Usage

llama-cli -m Phi-3.5-mini-instruct-Q4_K_M.gguf -cnv
Downloads last month
54
GGUF
Model size
4B params
Architecture
phi3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Fazmin/solus_v1_phi-3.5-mini-instruct-q4

Quantized
(192)
this model