How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf codingsoo/LlamaREST-IPD-GGUF:F16
# Run inference directly in the terminal:
llama cli -hf codingsoo/LlamaREST-IPD-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf codingsoo/LlamaREST-IPD-GGUF:F16
# Run inference directly in the terminal:
llama cli -hf codingsoo/LlamaREST-IPD-GGUF:F16
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf codingsoo/LlamaREST-IPD-GGUF:F16
# Run inference directly in the terminal:
./llama-cli -hf codingsoo/LlamaREST-IPD-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf codingsoo/LlamaREST-IPD-GGUF:F16
# Run inference directly in the terminal:
./build/bin/llama-cli -hf codingsoo/LlamaREST-IPD-GGUF:F16
Use Docker
docker model run hf.co/codingsoo/LlamaREST-IPD-GGUF:F16
Quick Links

LlamaREST-IPD (GGUF)

LlamaREST-IPD is a Small Language Model (SLM) fine-tuned to predict inter-parameter dependencies (IPDs) among REST API parameters, used by LlamaRestTest to improve automated REST API test generation.

It is one of the two SLMs used by LlamaRestTest; the companion model, LlamaREST-EX, generates realistic example values for parameters.

Models

Quantized with llama.cpp (Q6_K) from a Llama-3-8B base fine-tuned with QLoRA.

Variant File Size
2B LlamaREST-IPD-2B.gguf ~3.2 GB
4B LlamaREST-IPD-4B.gguf ~4.9 GB
8B LlamaREST-IPD-8B.gguf ~8.5 GB
F16 Llama3-Ipd-Rest-8.0B-F16.gguf ~16 GB

Download

pip install -U "huggingface_hub[cli]"
hf download codingsoo/LlamaREST-IPD-GGUF LlamaREST-IPD-8B.gguf --local-dir .

Usage

Intended to be run through the LlamaRestTest pipeline. To run the GGUF directly with llama.cpp:

./llama-cli -m LlamaREST-IPD-8B.gguf -p "your prompt"

Training

Fine-tuned from Llama-3-8B with QLoRA, then quantized with llama.cpp (Q6_K). Key hyperparameters:

  • 4-bit precision (nf4), compute dtype float16, no nested quantization
  • LoRA: r=64, alpha=16, dropout 0.1
  • 5 epochs, batch size 4, paged_adamw_32bit, LR 2e-4, weight decay 0.001, constant schedule, warmup ratio 0.03

Training data (Inter-Parameter Dependency): random1234321/REST-IPD.

License

The training/testing code is released under the MIT License. The model weights are derived from Meta Llama 3 and are therefore subject to the Meta Llama 3 Community License.

Citation

@article{kim2025llamaresttest,
  title={LlamaRestTest: Effective REST API Testing with Small Language Models},
  author={Kim, Myeongsoo and Sinha, Saurabh and Orso, Alessandro},
  journal={arXiv preprint arXiv:2501.08598},
  year={2025}
}
Downloads last month
864
GGUF
Model size
8B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for codingsoo/LlamaREST-IPD-GGUF

Quantized
(283)
this model

Paper for codingsoo/LlamaREST-IPD-GGUF