infrastructure-assistant-7b-gguf

This is a GGUF conversion of lokegud/infrastructure-assistant-7b, which is a LoRA fine-tuned version of Qwen/Qwen2.5-7B-Instruct.

Model Details

  • Base Model: Qwen/Qwen2.5-7B-Instruct
  • Fine-tuned Model: lokegud/infrastructure-assistant-7b
  • Training: Supervised Fine-Tuning (SFT) with TRL
  • Format: GGUF (for llama.cpp, Ollama, LM Studio, etc.)

Available Quantizations

File Quant Size Description Use Case
infrastructure-assistant-7b-f16.gguf F16 ~1GB Full precision Best quality, slower
infrastructure-assistant-7b-q8_0.gguf Q8_0 ~500MB 8-bit High quality
infrastructure-assistant-7b-q5_k_m.gguf Q5_K_M ~350MB 5-bit medium Good quality, smaller
infrastructure-assistant-7b-q4_k_m.gguf Q4_K_M ~300MB 4-bit medium Recommended - good balance

Usage

With llama.cpp

# Download model
huggingface-cli download lokegud/infrastructure-assistant-7b-gguf infrastructure-assistant-7b-q4_k_m.gguf

# Run with llama.cpp
./llama-cli -m infrastructure-assistant-7b-q4_k_m.gguf -p "Your prompt here"

With Ollama

  1. Create a Modelfile:
FROM ./infrastructure-assistant-7b-q4_k_m.gguf
  1. Create the model:
ollama create my-model -f Modelfile
ollama run my-model

With LM Studio

  1. Download the .gguf file
  2. Import into LM Studio
  3. Start chatting!

License

Inherits the license from the base model: Qwen/Qwen2.5-7B-Instruct

Citation

@misc{infrastructure_assistant_7b_gguf,
  author = {lokegud},
  title = {infrastructure-assistant-7b-gguf},
  year = {2025},
  publisher = {Hugging Face},
  url = {https://huggingface.co/lokegud/infrastructure-assistant-7b-gguf}
}

Converted to GGUF format using llama.cpp

Downloads last month
89
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to view the estimation

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for lokegud/infrastructure-assistant-7b-gguf

Base model

Qwen/Qwen2.5-7B
Quantized
(239)
this model