AI_Chatbot / evals

Commit History

production: harden catalog tenants analytics and widget install
59bd3b5

Hamza-Naimat commited on

ingest product variants (size/colour: per-variant price + stock) + bot-probe audit mode
e07e661

Hamza-Naimat Claude Opus 4.8 commited on

coverage audit + scheduled daily re-ingest
68b2aca

Hamza-Naimat Claude Opus 4.8 commited on

catalog: sidebar-collection listing + reverse 'which category is X in'
2244f27

Hamza-Naimat Claude Opus 4.8 commited on

eval panel: fix custom-tests 500, judge-key UX, product-question generation
d707010

Hamza-Naimat Claude Opus 4.8 commited on

fix: 'I specialize in [] for store' + eval relevance metric + mashup topics
b847ba7

Hamza-Naimat Claude Fable 5 commited on

fix: eval grader accepts variant-price answers containing reference price
a37d0a0

Hamza-Naimat Claude Fable 5 commited on

feat: eval question diversity + OOS probes + score renormalization
e916e7c

Hamza-Naimat Claude Fable 5 commited on

fix(evals): add admin-only debug trace + evidence-based judge
16c7450

Hamza-Naimat commited on

fix(eval): correct product scoring and retrieval ranking
cbe2568

Hamza-Naimat commited on

fix(eval): product-catalog question templates + judge rate-limit fixes
e0cc848

Hamza-Naimat Claude Sonnet 4.6 commited on

fix(eval): exclude retail logistics terms from non-retail DB tenant filter
0aad53e

Hamza-Naimat Claude Sonnet 4.6 commited on

fix(eval): two correctness fixes
881ce50

Hamza-Naimat Claude Sonnet 4.6 commited on

fix(eval): use SQLite-first for document loading to avoid lock contention
b8a7c55

Hamza-Naimat Claude Sonnet 4.6 commited on

fix: preserve rich config.json during GitHub sync; fix eval OOS filter for retail DBs
292d60d

Hamza-Naimat Claude Sonnet 4.6 commited on

eval: derive product questions from chunks
ca4c7d9

Hamza-Naimat commited on

eval: tighten tenant scope matching
d4751ce

Hamza-Naimat commited on

eval: fix chat retrieval scoring and tenant scope fallback
ef96d4c

Hamza-Naimat commited on

eval: harden live run output and grounding filter
98dea1f

Hamza-Naimat commited on

eval: tighten tenant scoping and fallback diagnosis
0c0c95d

Hamza-Naimat commited on

eval: tighten tenant filter and judge guidance
a432b33

Hamza-Naimat commited on

eval: target judge calls at failing rows
4e1039f

Hamza-Naimat commited on

eval: confirm trace-backed judge verdicts locally
227b9ed

Hamza-Naimat commited on

eval: deepen judge trace certainty
a08fac6

Hamza-Naimat commited on

eval: add workflow-trace evidence to llm judge
e60a937

Hamza-Naimat commited on

eval: filter stale eval sets and diagnose idk rows
ee9d6cf

Hamza-Naimat commited on

eval: filter cross-tenant analytics seeds
5b31b76

Hamza-Naimat commited on

eval: add prompt-aware judge diagnostics
3e1e3bd

Hamza-Naimat commited on

eval: add separate Groq judge layer
5010826

Hamza-Naimat commited on

evals: add context recall and likely-cause hints
5e14059

Hamza-Naimat commited on

evals: improve live scoring and admin diagnostics
658b774

Hamza-Naimat commited on

evals: add ranked retrieval and grounded answer metrics
daddee6

Hamza-Naimat commited on

evals: harden kb question quality and grading consistency
237df70

Hamza-Naimat commited on

evals: remove cwd assumptions and expand coverage
f682d27

Hamza-Naimat commited on

evals: tighten grading and harden standalone suite
691d228

Hamza-Naimat commited on

eval: harden DB-grounded suite and rotation
c86a9a8

Hamza-Naimat commited on

ci: run pytest and add eval v1
20663d6

Hamza-Naimat commited on

security: harden tenant auth and add lite eval
ac7ecbc

Hamza-Naimat commited on