You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

CaroleNDVoice

A fine-tuned Llama 3.1 8B Instruct chatbot designed for neurodivergent users โ€” particularly those who experience rejection-sensitive dysphoria (RSD), the visceral spike that comes with criticism or perceived rejection.

Built with Llama.

Carole is a portfolio / educational project. She is named after the author's wife, who has a way of holding hard conversations: validate first, redirect with a question. The model was trained to mirror that pattern.

The live demo is at meetcarole.com (gated).


What this is

  • A QLoRA fine-tune of meta-llama/Meta-Llama-3.1-8B-Instruct
  • Trained on ~1,500 synthetic conversations seeded from 50 hand-written golden examples
  • Quantized to Q4_K_M GGUF (~4.6GB) for inference via llama.cpp
  • Deployed end-to-end with a RAG layer (ChromaDB, all-MiniLM-L6-v2 embeddings, 1,732 chunks)

The defining behavior is validate, then redirect โ€” not as a softener for sycophancy but as a way to deliver pushback without triggering RSD.

What this is not

  • Not therapy. Not medical advice. Not a substitute for a clinician.
  • Not a transformers checkpoint. This repo ships GGUF + LoRA adapter artifacts, not merged safetensors weights. Do not deploy as a standard Text Generation / transformers endpoint.

Files in this repo

File Purpose
carole-q4_k_m.gguf Production inference artifact (llama.cpp / HF GGUF endpoint)
final_adapter/ LoRA adapter (reproduce merge or fine-tune further)
logs/ Training logs

There is no root config.json โ€” that is expected for a GGUF repo.

Quickstart (llama.cpp)

llama-server \
  --hf-repo datamatters24/CaroleNDVoice \
  --hf-file carole-q4_k_m.gguf \
  --host 127.0.0.1 --port 8085 \
  --ctx-size 4096

Then POST to http://127.0.0.1:8085/v1/chat/completions with an OpenAI-compatible payload.

Hugging Face Inference Endpoint

Do not create a transformers Text Generation endpoint โ€” it will fail looking for config.json and weight shards.

Instead:

  1. Open Inference Endpoints
  2. New endpoint โ†’ repo datamatters24/CaroleNDVoice
  3. Engine: GGUF / llama.cpp (auto-selected when a .gguf file is present)
  4. GGUF file: carole-q4_k_m.gguf
  5. Hardware: GPU recommended (e.g. L4 / T4 โ€” ~6GB+ VRAM for Q4 8B)
  6. Deploy โ†’ OpenAI-compatible URL at /v1/chat/completions

Use that URL with Vercel AI SDK (@ai-sdk/huggingface) or the Vercel Connect integration.

Training

Setting Value
Base meta-llama/Meta-Llama-3.1-8B-Instruct
Method QLoRA (4-bit NF4 + LoRA) via TRL SFTTrainer
LoRA rank / alpha 64 / 128
Learning rate 2e-4, cosine schedule
Epochs 3, best checkpoint epoch 2 (eval_loss = 1.40)
Hardware 1ร— A100 80GB on RunPod

Intended use

Educational / portfolio demonstrations of a non-sycophantic, neurodivergence-aware conversational pattern with RAG.

Out of scope

  • Crisis intervention or clinical mental-health use
  • Medical / legal / financial advice
  • Unrestricted public deployment without rate limiting and disclaimer

License

Derivative of Meta Llama 3.1 8B Instruct under the Llama 3.1 Community License.

Built with Llama.

Downloads last month
30
GGUF
Model size
8B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for datamatters24/CaroleNDVoice

Quantized
(904)
this model