Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
๐
In a Training Loop
Tarun Jain
lucifertrj
11
7
15
Follow
cuizhanming's profile picture
ai-everyday's profile picture
georgefreedland's profile picture
40 followers
ยท
12 following
https://youtube.com/@aiwithtarun
TRJ_0751
lucifertrj
AI & ML interests
Deep Learning, FPGA, and ML
Recent Activity
replied
to
their
post
5 days ago
You can now automate EDD (eval-driven development) with Coding Harness Agents > build a baseline LLM-based application > score every change with judge evals > keep what improves, reject what regresses I made a tutorial on what EDD is, how it works, and how to use eval scores across experiments to improve an LLM app. It builds on Jeffrey's (Confident AI) article on EDD and Eugene Yan's write-up on product evals. > Setup: a baseline RAG app using Qdrant and Gemini that every experiment starts from > Step 1: a binary-labelled dataset with critiques > Step 2: aligning the LLM-as-a-judge evaluator with Opik evals > Step 3: a harness loop that runs each experiment and scores it against the baseline. Tracing and experiment comparison then show what improved, what regressed, and what to tweak next. Source code is open source. Full guide (source code linked in the description): https://www.youtube.com/watch?v=e6akw_fKWPk
liked
a model
6 days ago
kaividlabs/LFM2.5-2.6B-litertlm
updated
a model
6 days ago
kaividlabs/LFM2.5-2.6B-litertlm
View all activity
Organizations
lucifertrj
's Spaces
4
Sort:ย Recently updated
Sleeping
Agents
AI Interview Assistant
๐ข
Agentic workflow for interview preparation platform
Runtime error
Joyride
๐ข
Runtime error
Agents
Gradio Blog Demo
โก
No application file
1
DCGAN
๐