tabby-tavern-stack / DEVLOG.md
jpanasuk's picture
Remove MCP server (never got working), add SearXNG Redis, update README/DEVLOG, clean start-stack.sh
4ef91a7
|
Raw
History Blame Contribute Delete
1.83 kB

License: MIT text---

DEVLOG.md

# Tabby-Tavern Development Log

Engineering history, architectural decisions, and troubleshooting notes for the Tabby-Tavern stack.

---

## Week 1 — Core Integration & Production Hardening

### 1. Hardware & Environment Baseline

- **GPU:** NVIDIA GeForce RTX 4070
- **Environment:** Fully containerized Linux stack via Docker Compose with GPU passthrough (`capabilities: [gpu]`)
- **Primary model format:** EXL3 / ExLlamaV3

### 2. Backend Engine — EXL3 Adoption

- Upgraded TabbyAPI to native ExLlamaV3 (EXL3) support for better memory mapping and higher token throughput.
- Standardized on EXL3 model repositories for fast VRAM loading and clean integration with SillyTavern + Open WebUI.

### 3. Master Orchestration (`docker-compose.yml`)

- Consolidated all services into a single `docker-compose.yml`:
  - TabbyAPI
  - SillyTavern
  - Open WebUI
  - Ollama
  - SearXNG
  - Redis
- Added persistent volume mounts for configs, models, and workspace data.

### 4. Critical Bug Fix — TabbyAPI Whitelist / Auth Error

**Problem:**  
TabbyAPI was rejecting connections from other containers due to strict default whitelist/authorization settings.

**Investigation:**  
Reviewed upstream config options and security switches controlling endpoint access.

**Resolution:**  
Explicitly set admin credentials to `admin` / `admin` in the configuration. This unlocked full API access and restored communication between all services in the stack.

---

## Week 2 — Model Deployment

### 1. Model Loading & Stack Reset

- Downloaded `Llama-3.1-8B-Instruct-6.0bpw-exl3` into `./tabby_models/`
- Performed clean stack restart:
  ```bash
  docker compose down && docker compose up -d

Confirmed healthy startup via:Bashdocker compose logs -f tabbyapi


Last updated: August 2026