AI_Chatbot / tests /test_eval_v1.py

Commit History

tests: tenant-scoped analytics + CSAT-scale charts, eval sqlite-first rows
c40cb1d

Hamza-Naimat Claude Opus 4.8 commited on

fix(eval): correct product scoring and retrieval ranking
cbe2568

Hamza-Naimat commited on

eval: derive product questions from chunks
ca4c7d9

Hamza-Naimat commited on

eval: tighten tenant scope matching
d4751ce

Hamza-Naimat commited on

eval: fix chat retrieval scoring and tenant scope fallback
ef96d4c

Hamza-Naimat commited on

eval: tighten tenant scoping and fallback diagnosis
0c0c95d

Hamza-Naimat commited on

eval: filter stale eval sets and diagnose idk rows
ee9d6cf

Hamza-Naimat commited on

eval: filter cross-tenant analytics seeds
5b31b76

Hamza-Naimat commited on

eval: add prompt-aware judge diagnostics
3e1e3bd

Hamza-Naimat commited on

eval: add separate Groq judge layer
5010826

Hamza-Naimat commited on

evals: add context recall and likely-cause hints
5e14059

Hamza-Naimat commited on

evals: improve live scoring and admin diagnostics
658b774

Hamza-Naimat commited on

evals: add ranked retrieval and grounded answer metrics
daddee6

Hamza-Naimat commited on

evals: harden kb question quality and grading consistency
237df70

Hamza-Naimat commited on

evals: remove cwd assumptions and expand coverage
f682d27

Hamza-Naimat commited on

evals: tighten grading and harden standalone suite
691d228

Hamza-Naimat commited on

eval: harden DB-grounded suite and rotation
c86a9a8

Hamza-Naimat commited on