GGUF
How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf FinysterLin/k8s-llama-expert:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf FinysterLin/k8s-llama-expert:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf FinysterLin/k8s-llama-expert:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf FinysterLin/k8s-llama-expert:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf FinysterLin/k8s-llama-expert:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf FinysterLin/k8s-llama-expert:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf FinysterLin/k8s-llama-expert:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf FinysterLin/k8s-llama-expert:Q4_K_M
Use Docker
docker model run hf.co/FinysterLin/k8s-llama-expert:Q4_K_M
Quick Links

base_model: unsloth/Meta-Llama-3.1-8B-bnb-4bit library_name: transformers tags: - unsloth - llama-3 - kubernetes - devops - gguf - text-generation license: apache-2.0 language: - en

Kubernetes Expert Llama-3.1-8B (GGUF)

This model is a fine-tuned version of Llama-3.1-8B, specifically trained to be a Kubernetes Expert. It was trained using Unsloth and LoRA on a high-quality StackOverflow Kubernetes dataset.

πŸš€ Model Features

  • Format: GGUF (Quantized to Q4_K_M)
  • Use Case: Answering technical K8s questions, debugging pods (CrashLoopBackOff), and generating YAML configurations.
  • Performance: 2x faster inference speed compared to the baseline model with significantly better domain knowledge.

πŸ“¦ How to Use with Ollama

  1. Download the Model Download k8s-expert-rescue.Q4_K_M.gguf from the Files tab.

  2. Create a Modelfile Create a file named Modelfile with the following content:

    FROM ./k8s-expert-rescue.Q4_K_M.gguf
    
    TEMPLATE """Below is an instruction that describes a task, paired with an input that provides further context. Write a response that appropriately completes the request.
    
    ### Instruction:
    You are a Kubernetes expert. Provide a technical solution to the following problem.
    
    ### Input:
    {{ .Prompt }}
    
    ### Response:
    """
    
    PARAMETER temperature 0.6
    PARAMETER num_ctx 4096
    PARAMETER stop "<|end_of_text|>"
    PARAMETER stop "### Instruction:"
    PARAMETER stop "### Input:"
    
Downloads last month
3
GGUF
Model size
8B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support