auralynq-rag / docs /research /modelfit-gap-analysis.md
MHamdan's picture
Deploy Auralynq RAG (Llama-3.3-70B via HF Inference Providers)
8c1b9fe verified
|
Raw
History Blame Contribute Delete
4.51 kB
# ModelFit Index β€” Gap Analysis
## Framing the Gap
Existing tools for local model selection fail RAG practitioners in five ways:
### Gap 1: Hardware Fit Is Not Automated
Most tools require the user to manually look up VRAM requirements, compare to their GPU spec, and guess whether a model will fit. This:
- Wastes time (download a model β†’ OOM β†’ try again).
- Is error-prone (VRAM estimates online are often for different quantizations or outdated).
- Is impossible without knowing quantization overhead, KV cache cost, and framework overhead.
**ModelFit solution**: Hardware profiler probes CPU/GPU/VRAM/RAM locally. VRAM estimator computes fit before any download. Quantization recommender selects the best level for available hardware.
---
### Gap 2: Quantization Selection Is Unexplained
Users face an alphabet soup of quantizations: q4_0, q4_k, q5_k_m, q8_0, gguf, gptq, awq, exl2. No existing tool explains:
- What quality difference each level produces.
- Which level fits a given GPU.
- Whether KV cache overhead at long contexts will push them over the limit.
**ModelFit solution**: Explicit quantization estimator with fit levels (comfortable / tight / not_recommended / impossible). All estimates clearly labelled. Benchmark runner measures actual quality at each level.
---
### Gap 3: RAG Quality Is Not Measured for Local Models
Leaderboards (Open LLM, MTEB, Chatbot Arena) measure:
- Academic reasoning (MMLU, ARC, GSM8k)
- Embedding retrieval (MTEB)
- Human preference (Arena)
None measure:
- Citation faithfulness (does the model cite the actual retrieved passage?)
- Groundedness (is the answer supported by retrieved context?)
- Abstention accuracy (does the model correctly say "I don't know"?)
- Evidence coverage (are all relevant retrieved passages used?)
**ModelFit solution**: RAG-specific benchmark tasks (citation, groundedness, abstention). These use local documents and local models β€” no cloud judge required.
---
### Gap 4: Speed Data Is Unreliable or Fabricated
Online tok/s reports suffer from:
- Different hardware (A100 vs. RTX 3090 vs. M2 Pro).
- Different quantizations.
- Different prompt/generation lengths.
- Batch vs. single-request.
- Framework differences (llama.cpp vs. vLLM vs. Ollama).
**ModelFit solution**: Benchmarks run locally and produce hardware-attached results. Estimates are clearly labelled `is_estimate: true`. Measured results from benchmark runner are labelled `is_measured: true`. Community results are labelled with `verified_status`.
---
### Gap 5: Adapter Compatibility Is Undocumented for Inference
PEFT/QLoRA documentation focuses on training setup. For inference:
- Which quantized models accept LoRA adapters?
- Can a GGUF model load a LoRA adapter?
- Does an Ollama model support adapter merging?
No tool answers these in the context of RAG deployment.
**ModelFit solution**: `supports_adapters` flag in model metadata. Notes field captures adapter compatibility caveats.
---
## Summary Table
| Gap | Existing coverage | ModelFit coverage |
|---|---|---|
| Automated hardware profiling | None (LM Studio partial) | βœ“ Full local probe |
| VRAM estimation before download | None (manual TheBloke tables) | βœ“ Formula-based estimator |
| Quantization recommendation | None | βœ“ Fit-aware recommender |
| KV cache overhead estimation | None | βœ“ Context-length-aware |
| RAG groundedness measurement | None locally | βœ“ Local benchmark |
| Citation faithfulness | None locally | βœ“ Local benchmark |
| Abstention accuracy | None | βœ“ Local benchmark |
| Measured tok/s (not fabricated) | Community posts, unreproducible | βœ“ Local benchmark with hardware metadata |
| Adapter compatibility info | HF cards (inconsistent) | βœ“ Registry metadata field |
| Visual grounding model support | None indexed | βœ“ Vision model detection |
| Community benchmark index | None structured | βœ“ Validated JSON schema |
| RAG strategy ↔ model integration | None | βœ“ ModelFit in query response |
---
## Who Benefits
| User type | Pain point addressed |
|---|---|
| Consumer GPU user (8–16 GB VRAM) | Knows which models actually fit before pulling |
| RAG practitioner | Knows which model produces grounded citations |
| Research contributor | Can reproduce and share benchmark results |
| Adapter fine-tuner | Knows which base models accept their adapter |
| Visual document QA user | Knows which models support image grounding |
| Multilingual RAG user | Knows which models handle non-English documents |