AI & ML interests

AI evaluation, LLM benchmarking, agent evaluation, reproducible eval workflows, model comparison, regression testing, failure analysis, eval datasets, and open-source developer tooling

Recent Activity

phranzia  updated a Space 3 days ago
quantiles/README
phranzia  updated a Space 10 days ago
quantiles/README
phranzia  published a Space 2 months ago
quantiles/README
View all activity

phranzia 
updated a Space 3 days ago
phranzia 
published a Space 2 months ago