Spaces:
Running on Zero
Running on Zero
File size: 24,056 Bytes
d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f 2bfecd4 d3f3e8f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 | # MLOL — MultiDomain LLM Optimisation Lab
## Master Specification v1.1 — 2026-07-19
**Status:** approved blueprint for implementation.
**v1.1:** hardened per reliability review — positioning statement, graceful
degradation, P7/P8, config schema validation, untrusted-input handling,
runtime-aware routing, per-experiment manifests + state machine, paired
statistics, environment-stamped certificates, dependency separation, lazy
loading, mocked-backend-first delivery.
**Supersedes:** FinLLM Foundry prototype (this repo's v0; code is reused where it fits).
**Owner:** finpy1789 · Contact: finpy07@gmail.com
---
## 1. Vision
A Hugging Face Space that is a complete, research-grade **LLM fine-tuning and
optimisation laboratory**: choose a base model → upload a dataset → validate and
clean it → get a hardware recommendation → run a baseline evaluation → fine-tune
(on the right backend) → re-evaluate → compare → generate an optimisation report
and a Model Performance Certificate → publish or deploy the adapter. A
platform-wide AI Research Assistant acts as a copilot across every step.
The Space is the **control plane**. Training runs on the **training plane**
(ZeroGPU for demos, HF Jobs for serious runs, Colab export for the free route).
All state lives on the **persistence plane** (Hub repos / storage bucket).
The Space itself is stateless and restart-safe.
```
┌──────────────────────── CONTROL PLANE (this Space) ────────────────────────┐
│ upload/validation · baseline & post evals (sampled) · hardware estimation │
│ training config · job submission · status/log monitoring · reports/certs │
└──────────────┬─────────────────────────────────────────────┬──────────────┘
TRAINING PLANE PERSISTENCE PLANE
├ ZeroGPU — demo runs ≤1.5B ├ private dataset repo — experiments DB
├ HF Jobs — managed serious training ├ model repos — adapters/merged models
├ Colab export — free user-controlled └ storage bucket — checkpoints/logs/
└ external GPU — advanced route uploads (fallback: dataset repo)
```
**Positioning.** MLOL is an **orchestration and evaluation platform, not a
persistent GPU training server**. ZeroGPU provides bounded demonstrations;
large-model training is delegated to Colab, HF Jobs, or external compute.
**Graceful degradation** is an architectural objective: dataset preparation,
experiment configuration, report viewing, and Colab export must remain fully
usable when ZeroGPU is unavailable, queued, or rate-limited.
## 2. Design principles (binding)
P1. **Configuration over knowledge.** Every model, provider, domain, benchmark,
and hardware profile lives in `configs/`. The application never validates,
rejects, or special-cases a model name from code or from the assistant's
internal knowledge. If it is in the config, it is valid and the UI renders
it. Newly released models (e.g. Gemma 4, GPT-5.6) are added by editing
config only — zero code changes. Configuration files ARE validated for
structure at startup — required fields, types, supported capability
combinations, and `schema_version` — via Pydantic models; a config that
fails schema validation is reported precisely and skipped, never guessed at.
P2. **Never show an impossible button.** The routing engine decides what can run
where; ineligible options render disabled with the reason and an alternative.
P3. **Statelessness.** Any Space restart loses nothing: all experiment state is
on the Hub. The UI always renders from persisted state.
P4. **Statistical honesty.** Every sampled metric carries n, seed, and 95% CI.
Certificates never imply a sample is the full benchmark.
P5. **Credential separation.** The owner token only touches demo resources and
the platform's own experiment repo. User-owned training uses the user's
scoped token (HF OAuth preferred). Raw tokens are never persisted or logged.
P6. **Working increments.** Each delivery increment leaves the deployed Space
fully functional.
P7. **Hardware-aware execution.** CPU operations never acquire ZeroGPU. GPU-
decorated functions are small, independently callable, bounded, and release
GPU resources immediately after their operation completes.
P8. **Resumable and idempotent execution.** Every training and evaluation stage
records its state before and after execution; retrying a completed
operation must not create duplicate jobs or corrupt experiment artifacts.
## 3. Module map
```
MLOL
├── Home — pipeline overview, recent experiments, quick start
├── Tier 1 — General Fine-Tuning Lab
├── Tier 2 — Domain Foundry (premium: access code / OAuth allowlist)
├── Evaluation Lab
├── Reports — optimisation reports, certificates, research dashboard
├── Adapter Library
├── Hardware Advisor
├── Documentation
└── AI Research Assistant — docked bottom-right overlay, available on every page
```
Top-level navigation = Gradio Tabs. The assistant is a collapsible CSS-overlay
panel (bottom-right), not a tab, and is present regardless of active page.
## 4. Tier 1 — General Fine-Tuning Lab
### 4.1 Base model selection
- Dropdown fed by `configs/models.yaml`. Seed catalogue: TinyLlama, SmolLM,
Phi-4 Mini, Qwen2.5, Qwen3, Gemma, Mistral, Llama families (editable freely).
- Each entry displays: parameters, context length, license, quantization
support, gated flag — all from config metadata (P1).
- Schema per entry:
```yaml
- name: str # display name
repo: str # HF repo id
params_b: float # billions
context: int
license: str
gated: bool
quant: [4bit, 8bit, bf16] # supported load modes
chat_template: str # LLaMA-Factory template name (for export)
notes: str # optional, shown verbatim
```
### 4.2 Dataset upload & preparation
- Accepted: CSV, JSON, JSONL, TXT, PDF, DOCX (pypdf / python-docx for extraction).
- **All preparation runs on CPU** (P7): parsing, validation, cleaning, dedupe,
and approximate token estimation never reserve ZeroGPU. The selected model's
tokenizer is loaded lazily and cached; a fast generic tokenizer provides
estimates until then.
- **Untrusted input handling:** PDF/DOCX parsed as untrusted content;
configurable limits (`configs/limits.yaml`) on upload size, extracted text
size, sample count, and token count; encrypted, malformed, oversized, or
decompression-heavy (zip-bomb-like) files are rejected with a clear reason.
- Pipeline: parse → column/field mapping UI (auto-guess instruction/input/output
or prompt/response or raw text) → validate → clean (dedupe, strip empties,
length filter, encoding fixes) → convert to chat `messages` format →
persist to `uploaded_datasets/<run_id>/` (bucket or repo).
- Dataset Summary panel: samples, token count (tokenizer of selected model),
average/percentile lengths, missing values, duplicates removed, language
guess, **estimated training time per backend**.
- Validation for evaluation: if reference answers exist → reference-based
metrics enabled; else those metrics are greyed with the reason (P2).
### 4.3 Training configuration
- **Basic mode:** epochs, batch size, learning rate, LoRA rank, alpha, dropout.
- **Advanced mode:** + scheduler, optimizer, gradient accumulation, warmup,
weight decay, seed, method (LoRA/QLoRA/DoRA), max length, packing.
- "Recommend hyperparameters" button → rule-based suggestion from dataset size
+ model size (assistant can also do this via tool call, §9).
### 4.4 Training routing engine
```python
def choose_backend(model_params_b, sample_count, est_tokens, user_ctx) -> Route:
# zerogpu_demo: params_b <= 1.5 and sample_count <= 5_000 and est_minutes <= window
# hf_job: user authenticated + jobs eligibility verified
# colab_export: always available
# external: documented, config listed
```
- Routing decision panel shows: selected model, estimated GPU-hours, ZeroGPU
eligibility (with reason), recommended backend, alternatives.
- **Routing is capability-based and runtime-aware:** eligibility (ZeroGPU
availability, estimated duration, model memory needs, dataset size, user
authentication, backend access) is re-checked **immediately before launch**,
not only when the form renders. Each backend implements one
`TrainingBackend` capability contract (also satisfied by a mock backend used
in tests and Increment 1).
- **ZeroGPU demo path** ("Quick Demo Training — limited dataset and model size"):
single-window LoRA run, live stage display (Preparing Dataset → Loading Base
Model → Training → Saving Adapter → Evaluation), progress bar, loss curve
(live-updated plot), ETA; GPU usage where the runtime exposes it.
The demo profile is **enforced by configuration** (`configs/limits.yaml`):
max model size, token count, sequence length, epochs, estimated duration,
checkpoint frequency, and output size — hard limits, not UI guidance.
- **Progress without live streams:** the UI never depends on uninterrupted log
streaming from ZeroGPU or Colab. Stage transitions and periodic summaries are
persisted (§11), and the interface reconstructs progress from persisted state
after any Space refresh or restart.
- **HF Jobs path:** generate config → `HfApi.run_job(...)` with user token →
status (queued/running/completed/failed), log streaming/polling, cancel,
adapter retrieval, auto-trigger post-eval. Eligibility check before showing.
- **Colab export path:** downloadable package
`mlol_experiment_<run_id>/ {training_notebook.ipynb, train_config.yaml,
dataset_manifest.json, requirements.txt, run_training.py, run_evaluation.py,
README.md}`. The package is **pinned and self-contained**: exact dependency
versions, dataset revision, model revision, configuration hash, resume
instructions, and a completion-upload step. Notebook: HF auth → pull dataset
→ train (resume-capable) → push adapter → write completion metadata back to
the experiment repo.
- **Completion verification:** a Colab/Jobs run is accepted as complete only
after the Space verifies the expected adapter, config, and evaluation
artifacts exist and match the original `run_id` and configuration hash;
mismatches mark the run `failed` with the discrepancy listed.
## 5. Tier 2 — Domain Foundry (premium)
- Domains (config-driven, `configs/domains/*.yaml`): Finance, Law, Medical,
Regulatory, Accounting, Insurance, Computer Science, Programming,
Artificial Intelligence, Literature, General Science.
- Each domain config provides: curated dataset references, prompt/format
templates, benchmark definitions, evaluation metric set, recommended
hyperparameters, system prompt, disclaimer text.
- Execution: same Tier-1 machinery; domain config pre-fills everything.
Serious runs route to HF Jobs / external; sampled evals in-Space; full evals
via exported Job/Colab script; domain performance certificate.
- Access control: `PREMIUM_ACCESS_CODES` secret (Space-set codes) and/or HF
OAuth username allowlist. Locked UI states name the unlock path
(email finpy07@gmail.com).
## 6. Evaluation Lab
### 6.1 Levels
| Level | Items | Purpose | Execution |
|----------|----------|----------------------------|------------------------------------|
| Quick | 25–50 | UI sanity check | in-Space (CPU or one GPU window) |
| Standard | 100–200 | certificate comparison | in-Space, bounded batches (each batch a small GPU call, P7; resumable, P8) |
| Full | complete benchmark | publication-quality | exported Job/Colab script only — never in-Space |
Per-backend execution limits (batch size, max items per window, timeouts) live
in `configs/limits.yaml`.
- Identical item IDs, seed, and generation settings for baseline vs fine-tuned
(stored in the experiment record; re-used automatically).
- Every sampled metric reports: value, n, seed, 95% CI (Wilson for proportions,
bootstrap for continuous). Certificates print "Full benchmark: Not executed"
where applicable (P4).
### 6.2 Metrics
- Reference-based: accuracy (exact/normalized match), BLEU, ROUGE-L, BERTScore.
- Intrinsic: perplexity (held-out).
- Operational: latency (p50/p95), tokens/sec, peak memory, response length.
- **Hallucination — component estimate, never one opaque number:**
factual consistency vs references · unsupported-claim rate ·
citation/reference agreement (where applicable) · judge-model flagged rate →
composite % labeled "Estimated hallucination risk — not a direct measurement
of truthfulness", with judge model, prompt version, sample size, and
human-verification flag disclosed.
### 6.3 Flow
Baseline eval (pre-training, auto) → `baseline.json` → training → identical
post eval → `post_training.json` → comparison dashboard (metric | baseline |
fine-tuned | Δ | paired significance) → classification.
- **Paired statistics:** because baseline and post-training use identical items,
significance uses paired tests — paired bootstrap / permutation for continuous
metrics, McNemar-style paired proportion test for accuracy-type metrics. CIs
remain displayed on every metric (P4); CI overlap is never the significance rule.
- **Item-level persistence:** per-item outputs (or stable item hashes + scores)
are persisted, so interrupted evaluations resume without repeating completed
items (P8) and any comparison can be exactly reproduced later.
**Improved / Neutral / Degraded** (per-metric and overall) → Trial & Error
panel: data-driven diagnostics (dataset too small, LR too high, overfitting
signal from train/eval loss divergence, low-quality dataset signals, too few
epochs, catastrophic-forgetting probe results) with evidence for each claim.
## 7. Reports & Certificate
### 7.1 Optimisation report
Auto-generated: training summary, dataset summary, hyperparameters, evaluation
tables, performance deltas, charts (loss, metric comparisons), recommendations.
Download: **PDF** (reportlab), **CSV** (flat metrics), **JSON** (full record).
### 7.2 Model Performance Certificate (9 sections)
1. Identity: model, base, adapter, date, training time, dataset fingerprint.
2. Performance: accuracy, BLEU, BERTScore, latency, memory, hallucination
estimate — each with n and CI.
3. Overall result: ✓ Improved / ⚠ Neutral / ✗ Degraded (CI-aware rule:
improved = significant gains on primary metrics without significant
regressions; degraded = any significant regression on a primary metric).
4. Confidence rating ★1–5 — reflects **evaluation comprehensiveness** (sample
sizes, metric coverage, seed control, full-vs-sampled), never model quality.
Rubric printed on the certificate.
5. Strengths — generated from significant positive deltas.
6. Weaknesses — generated from significant negative deltas + diagnostics.
7. Deployment recommendation: Ready / Needs More Training / Needs Better
Dataset / Do Not Deploy (rule-based from 3+6).
8. Hardware recommendation: per-target table (Recommended / Minimum / Not
Recommended) from `configs/hardware.yaml` profiles.
9. Research summary — templated natural-language digest of the deltas.
Every certificate (and report) embeds an **execution environment block**:
backend, accelerator (where exposed), quantization mode, model revision,
dataset fingerprint, dependency lock fingerprint (captured at runtime),
evaluation seed and sample sizes, and a demo-run vs full-run marker.
### 7.3 Research Dashboard
Experiment history (from `experiments/index.jsonl`), per-experiment graphs
(loss, BLEU, BERTScore, latency, hallucination, accuracy), compare any two
experiments side-by-side.
## 8. Adapter Library
Registry of every trained adapter: name, version, domain, base model,
performance snapshot, date, links — Download (HF repo), Deploy (§10),
Load-in-playground (inference engine hot-swap). Backed by model repos +
`evaluation_summary.json` in each.
## 9. AI Research Assistant
- **Placement:** docked bottom-right collapsible overlay, on every page.
- **Provider layer** (`configs/providers.yaml`, P1): each entry = display name,
API type (hf-inference | openai-compatible | anthropic), model id, auth
source (user HF OAuth token / user-supplied API key). Seed entries include
Gemma 4, GPT-5.6, Claude, DeepSeek, Qwen — list is config, edit freely.
User keys held in session memory only (P5).
- **Modes** (auto-selected by context + explicit switcher):
1. **General LLM Assistant** — concept Q&A (LoRA, QLoRA, ORPO, forgetting…),
grounded in `Documentation/` content.
2. **Experiment Assistant** — context injection: current experiment's config,
training logs, dataset stats, eval JSONs → data-grounded diagnosis.
3. **Hardware Advisor** — deterministic estimator computes VRAM/time (the
numbers), assistant explains and recommends (the words). Numbers always
come from the estimator, not the LLM.
4. **Report Interpreter** — certificate/report JSON injected; explains each
metric, why the rating was assigned, highest-leverage improvements.
- **No GPU for explanations:** the assistant never acquires ZeroGPU to explain
an experiment (P7). Default inference = HF Inference Providers or the user's
supplied provider; local GPU inference only as an explicit option for small
config-listed models.
- **Untrusted context:** uploaded datasets, extracted documents, logs, and
model-generated text injected into the assistant are treated as untrusted
data — they must not override the assistant's system instructions, and text
inside them can never directly trigger platform actions; only the user's own
chat turns can invoke tools.
- **Tool awareness:** function-calling against a platform action registry:
`open_comparison(exp_a, exp_b)`, `regenerate_report(run_id)`,
`export_certificate_pdf(run_id)`, `suggest_hyperparameters(run_id)` (pre-fills
the training form), `navigate(module)`. Actions that mutate state are echoed
in the chat ("Pre-filled the training form — review before launching");
destructive actions are out of scope for the assistant.
## 10. Deployment targets
Per adapter: push to HF Hub (adapter or merged) · safetensors download ·
GGUF conversion (exported script; on-Space only for small models) · Ollama
Modelfile · Dockerfile (vLLM/TGI serving) · the Space's own inference API.
All generated artifacts, no owner-credential publishing on behalf of users (P5).
## 11. Persistence schema
```
Private dataset repo (mlol-experiments)
experiments/<run_id>/manifest.json # AUTHORITATIVE per-experiment record
experiments/<run_id>/config.yaml
experiments/<run_id>/dataset_summary.json
experiments/<run_id>/baseline.json
experiments/<run_id>/post_training.json
experiments/<run_id>/certificate.json
experiments/<run_id>/report.pdf
experiments/index.jsonl # rebuildable dashboard cache ONLY
Storage bucket (fallback: same dataset repo)
uploaded_datasets/<run_id>/ checkpoints/<run_id>/ logs/<run_id>/ tmp/<run_id>/
Model repos (per adapter)
adapter_config.json adapter_model.safetensors README.md evaluation_summary.json
```
- **Source of truth = one manifest per experiment.** `index.jsonl` is a derived
dashboard cache, rebuildable at any time by scanning manifests — concurrent
runs never contend on a shared mutable file.
- **State machine** (persisted in each manifest, transitions appended with
timestamps): `draft → data-ready → baseline-running → baseline-complete →
training-submitted → training-running → training-complete →
post-evaluation-running → complete`, plus `failed` and `cancelled` from any
state. The UI renders exclusively from the recorded state (P3, P8).
- **Artifact envelope:** every persisted artifact carries `schema_version`,
`created_at`, `updated_at`, `run_id`, config hash, model revision, dataset
fingerprint, backend, and producer version. Small migration functions read
older schema versions forward.
- Storage access goes through one `StorageBackend` abstraction (bucket | repo).
## 12. Tech stack
Gradio 5.x (SDK) · huggingface_hub (Jobs, repos, OAuth) · transformers/peft/trl
(demo training + inference) · datasets · evaluate + bert-score + sacrebleu ·
**pydantic** (config + experiment schema validation) · reportlab (PDF) ·
pypdf, python-docx (ingestion) · matplotlib (charts).
- **Pinned, split dependencies:** tested version ranges, not open-ended
`latest`. `requirements.txt` = lightweight Space/control-plane deps only;
`requirements-train.txt` = training-plane deps (used by ZeroGPU demo path,
HF Jobs image, and Colab package). The runtime lock fingerprint is captured
at startup for certificates (§7).
- **Service layer:** orchestration is framework-independent — Gradio callbacks
invoke plain Python services (`src/services/`); no training, persistence, or
routing logic lives inside UI callbacks.
- **Lazy everything:** transformers, torch, evaluation models, document
parsers, and plotting import lazily. The Space must render with no model
loaded and no GPU acquired (P7); first paint stays fast.
Existing FinLLM Foundry code reused: inference engine, EDGAR ingestion,
dataset mixing, training scripts (become the Jobs/Colab payload), guide logic
(absorbed into assistant mode 1).
## 13. Delivery plan (each increment ships working)
| # | Increment | Key acceptance test |
|---|-----------|--------------------|
| 1 | **Foundation**: nav skeleton, Pydantic config validation, experiment manifests + state machine, persistence recovery, `TrainingBackend` capability contract, **mocked training backend** | Space renders all modules; a mocked end-to-end run walks every state; experiments survive restart; invalid config reported precisely |
| 2a | Dataset preparation: upload→validate→clean→convert (CPU-only), limits, untrusted-input handling + **reproducible pinned Colab export** | malicious/oversized files rejected cleanly; export package runs in Colab and reports completion back |
| 2b | Bounded ZeroGPU demo training (config-enforced demo profile) | 0.5–1.5B LoRA run completes in one window; over-limit configs never reach the GPU |
| 3 | Evaluation Lab: adaptive sampling, CIs, paired tests, item-level resume, baseline/post/compare, diagnostics | same-seed before/after eval produces paired-significance dashboard; interrupted eval resumes |
| 4 | Reports + Certificate (PDF/CSV/JSON) with environment block | certificate downloads with all 9 sections + environment stamp |
| 5 | HF Jobs backend: submit, monitor, cancel, retrieve, completion verification, eligibility gating | 7B QLoRA job launched from UI with user token; tampered/mismatched artifacts rejected |
| 6 | Domain Foundry + domain benchmarks (Finance first, then the rest) | finance domain run yields domain certificate |
| 7 | AI Research Assistant (4 modes + tools, untrusted-context rules) + Adapter Library + Research Dashboard | assistant diagnoses a real experiment and pre-fills a config; injected instructions in logs are ignored |
## 14. Resolved decisions
- **Name:** MultiDomain LLM Optimisation Lab (MLOL).
- **Space:** evolve this repo/Space; rename to `finpy1789/mlol` (URL redirects).
- **Model/provider catalogues:** config-only; no code-side name validation (P1).
- **Assistant default provider:** open model via HF Inference (user OAuth);
Claude/GPT as bring-your-own-key.
- **Premium:** access codes now; OAuth allowlist when public.
|