How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf GYETI/Prism-4B-Reasoning:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf GYETI/Prism-4B-Reasoning:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf GYETI/Prism-4B-Reasoning:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf GYETI/Prism-4B-Reasoning:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf GYETI/Prism-4B-Reasoning:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf GYETI/Prism-4B-Reasoning:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf GYETI/Prism-4B-Reasoning:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf GYETI/Prism-4B-Reasoning:Q4_K_M
Use Docker
docker model run hf.co/GYETI/Prism-4B-Reasoning:Q4_K_M
Quick Links

⚖️ Attribution & Licensing

Prism-4B-Reasoning is a derivative work fine-tuned from the pre-trained Nemotron 3 base architecture created by NVIDIA.

This model is distributed openly under the terms of the Apache 2.0 License. In accordance with license guidelines, this model contains extensive weight alterations from the original framework to adopt a new custom identity and functional task profile.

The original work and licensing elements can be audited via the Nvidia Nemotron Repository.

Downloads last month
73
GGUF
Model size
4B params
Architecture
nemotron_h
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GYETI/Prism-4B-Reasoning

Unable to build the model tree, the base model loops to the model itself. Learn more.