How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf DFveloper/AIKAR-1.2-Pro-GGUF:BF16
# Run inference directly in the terminal:
llama cli -hf DFveloper/AIKAR-1.2-Pro-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf DFveloper/AIKAR-1.2-Pro-GGUF:BF16
# Run inference directly in the terminal:
llama cli -hf DFveloper/AIKAR-1.2-Pro-GGUF:BF16
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf DFveloper/AIKAR-1.2-Pro-GGUF:BF16
# Run inference directly in the terminal:
./llama-cli -hf DFveloper/AIKAR-1.2-Pro-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf DFveloper/AIKAR-1.2-Pro-GGUF:BF16
# Run inference directly in the terminal:
./build/bin/llama-cli -hf DFveloper/AIKAR-1.2-Pro-GGUF:BF16
Use Docker
docker model run hf.co/DFveloper/AIKAR-1.2-Pro-GGUF:BF16
Quick Links

[AIKAR 1.2 Pro - GGUF] 📦

Model License Format Developer

🔍 Overview

AIKAR 1.2 Pro - GGUF는 고성능 LLM인 AIKAR 1.2 Pro를 일반 사용자 환경에서도 효율적으로 구동할 수 있도록 llama.cpp 포맷으로 양자화(Quantization)한 버전입니다.

이 리포지토리는 고사양 GPU가 없는 환경에서도 모델의 지능을 최대한 유지하면서, 메모리(RAM/VRAM) 사용량을 최적화하여 로컬 추론(Local Inference)이 가능하도록 설계되었습니다.

🚀 Quick Start (Local Execution)

1. Using llama.cpp

가장 직접적인 방법으로, 터미널에서 다음과 같이 실행할 수 있습니다.

# 빌드된 llama.cpp 경로에서 실행
./main -m AIKAR-1.2-Pro-G2Q3.gguf -p "Tell me a long story" -n 128 -t 8
  • -t 8: 사용할 CPU 코어(Thread) 수를 지정합니다.

2. Using DFveloper/aikar-engine from Github

AMD GPU에 최적화된 엔진으로, llama.cpp 기반입니다. LOOP Chat의 실사용 서빙 엔진입니다.

# 빌드된 aikar-engine 경로에서 실행
./main -m AIKAR-1.2-Pro-G2Q3.gguf -p "Tell me a long story" -n 128 -t 8
  • -t 8: 사용할 CPU 코어(Thread) 수를 지정합니다.

🛠 Technical Specifications

  • Base Model: AIKAR 1.2 Pro (Full Precision)
  • Quantization Method: AIKAR Quant
  • Architecture: Transformer-based Decoder-only
  • Inference Engine: Optimized for aikar-engine

🤝 Feedback & Contribution

GGUF 양자화& 실행 과정에서 발생하는 성능 저하나 버그는 LOOP GitHub의 Issue 탭에 제보해 주세요. 사용자의 피드백은 더 정교한 양자화 가중치를 만드는 데 큰 도움이 됩니다.


"Bring the power of AIKAR 1.2 Pro to your local machine." — Optimized by LOOP

Downloads last month
22
GGUF
Model size
25B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DFveloper/AIKAR-1.2-Pro-GGUF

Quantized
(3)
this model