How to use from
llama.cpp
# Gated model: Login with a HF token with gated access permission
hf auth login
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf datamatters24/CaroleNDVoice:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf datamatters24/CaroleNDVoice:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf datamatters24/CaroleNDVoice:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf datamatters24/CaroleNDVoice:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf datamatters24/CaroleNDVoice:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf datamatters24/CaroleNDVoice:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf datamatters24/CaroleNDVoice:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf datamatters24/CaroleNDVoice:Q4_K_M
Use Docker
docker model run hf.co/datamatters24/CaroleNDVoice:Q4_K_M
Quick Links

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

CaroleNDVoice

A fine-tuned Llama 3.1 8B Instruct chatbot designed for neurodivergent users โ€” particularly those who experience rejection-sensitive dysphoria (RSD), the visceral spike that comes with criticism or perceived rejection.

Built with Llama.

Carole is a portfolio / educational project. She is named after the author's wife, who has a way of holding hard conversations: validate first, redirect with a question. The model was trained to mirror that pattern.

The live demo is at meetcarole.com (gated).


What this is

  • A QLoRA fine-tune of meta-llama/Meta-Llama-3.1-8B-Instruct
  • Trained on ~1,500 synthetic conversations seeded from 50 hand-written golden examples
  • Quantized to Q4_K_M GGUF (~4.6GB) for inference via llama.cpp
  • Deployed end-to-end with a RAG layer (ChromaDB, all-MiniLM-L6-v2 embeddings, 1,732 chunks)

The defining behavior is validate, then redirect โ€” not as a softener for sycophancy but as a way to deliver pushback without triggering RSD.

What this is not

  • Not therapy. Not medical advice. Not a substitute for a clinician.
  • Not a transformers checkpoint. This repo ships GGUF + LoRA adapter artifacts, not merged safetensors weights. Do not deploy as a standard Text Generation / transformers endpoint.

Files in this repo

File Purpose
carole-q4_k_m.gguf Production inference artifact (llama.cpp / HF GGUF endpoint)
final_adapter/ LoRA adapter (reproduce merge or fine-tune further)
logs/ Training logs

There is no root config.json โ€” that is expected for a GGUF repo.

Quickstart (llama.cpp)

llama-server \
  --hf-repo datamatters24/CaroleNDVoice \
  --hf-file carole-q4_k_m.gguf \
  --host 127.0.0.1 --port 8085 \
  --ctx-size 4096

Then POST to http://127.0.0.1:8085/v1/chat/completions with an OpenAI-compatible payload.

Hugging Face Inference Endpoint

Do not create a transformers Text Generation endpoint โ€” it will fail looking for config.json and weight shards.

Instead:

  1. Open Inference Endpoints
  2. New endpoint โ†’ repo datamatters24/CaroleNDVoice
  3. Engine: GGUF / llama.cpp (auto-selected when a .gguf file is present)
  4. GGUF file: carole-q4_k_m.gguf
  5. Hardware: GPU recommended (e.g. L4 / T4 โ€” ~6GB+ VRAM for Q4 8B)
  6. Deploy โ†’ OpenAI-compatible URL at /v1/chat/completions

Use that URL with Vercel AI SDK (@ai-sdk/huggingface) or the Vercel Connect integration.

Training

Setting Value
Base meta-llama/Meta-Llama-3.1-8B-Instruct
Method QLoRA (4-bit NF4 + LoRA) via TRL SFTTrainer
LoRA rank / alpha 64 / 128
Learning rate 2e-4, cosine schedule
Epochs 3, best checkpoint epoch 2 (eval_loss = 1.40)
Hardware 1ร— A100 80GB on RunPod

Intended use

Educational / portfolio demonstrations of a non-sycophantic, neurodivergence-aware conversational pattern with RAG.

Out of scope

  • Crisis intervention or clinical mental-health use
  • Medical / legal / financial advice
  • Unrestricted public deployment without rate limiting and disclaimer

License

Derivative of Meta Llama 3.1 8B Instruct under the Llama 3.1 Community License.

Built with Llama.

Downloads last month
30
GGUF
Model size
8B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for datamatters24/CaroleNDVoice

Quantized
(904)
this model