-
Repro - MemEvolve: Meta-Evolution of Agent Memory Systems
🧬Collaborate with an AI agent to manage a shared experiment logbook
-
Repro - TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
🎯Explore experiment logs and sync findings with a coding agent
-
Repro - How to Correctly Report LLM-as-a-Judge Evaluations
🎯Log and share LLM evaluation findings in a collaborative notebook
-
Repro - Dependence-Aware Label Aggregation via Ising Models
🧲Explore and edit experiment logbooks with AI agent help
Kshitij Thakkar PRO
kshitijthakkar
AI & ML interests
Building the evaluation and observability layer for AI.
Creator of TraceVerse—turning real-world LLM interactions into datasets, benchmarks, and cost-efficient model insights.
Organizations
Trackio logbooks: claim-by-claim reproductions of ICML 2026 papers for the Agent Reproducibility Challenge.
DeepSeek V4 Replicas
Small-scale faithful replicas of the DeepSeek-V4 architecture for ablation and weight-transfer research.
-
kshitijthakkar/deepseek-v4-mini-300M-init
Text Generation • 0.3B • Updated • 13 -
kshitijthakkar/deepseek-v4-mini-1B-init
Text Generation • 1B • Updated • 129 • 1 -
kshitijthakkar/deepseek-v4-mini-3B-init
Text Generation • 3B • Updated • 55 • 1 -
kshitijthakkar/deepseek-v4-mini-6B-init
Text Generation • 8B • Updated • 52 • 4
Qwen3.5 Dense-to-MoE Weight Transfer
Qwen3.5 MoE models from dual-source weight transfer (dense backbone + 35B-A3B experts). Hybrid DeltaNet + GQA attention.
-
kshitijthakkar/qwen3.5-moe-0.87B-d0.8B
Image-Text-to-Text • 1B • Updated • 259 • 1 -
kshitijthakkar/qwen3.5-moe-2.3B-d2B
Image-Text-to-Text • 3B • Updated • 26 -
kshitijthakkar/qwen3.5-moe-4.7B-d4B
Image-Text-to-Text • 5B • Updated • 300 -
kshitijthakkar/qwen3.5-tiny-test
Image-Text-to-Text • 0.1B • Updated • 15
Mobile MoE Architecture Search
32 MoE models from 41 experiments exploring expert count, routing, and learning rates for mobile deployment.
TraceMind-AI
Collection of my submissions to MCP-1st-Birthday:- TraceMind Agent and MCP Server and smoltrace datasets generated for running evals using smoltrace.
- Sleeping10
TraceMind MCP Server
🤖10MCP server for agent evaluation with Gemini 2.5 Flash
- RunningAgents21
TraceMind AI
🧠21AI agent evaluation with MCP-powered intelligence
-
MCP-1st-Birthday/smoltrace-recruitment-tasks
Viewer • Updated • 101 • 26 -
MCP-1st-Birthday/smoltrace-smart-home-tasks
Viewer • Updated • 100 • 31
Kirigami: Zero-Shot Expert Carves of Qwen3.6
Qwen3.6-35B-A3B carved by expert importance to fit one consumer GPU. No training, no calibration. 800 tok/s on a laptop 5090.
-
kshitijthakkar/Kirigami-Qwen3.6-20B-A3B-NVFP4
Text Generation • 14B • Updated • 72 • 1 -
kshitijthakkar/Kirigami-Qwen3.6-24B-A3B-NVFP4
Text Generation • 16B • Updated • 285 -
kshitijthakkar/Kirigami-Qwen3.6-28B-A3B-NVFP4
Text Generation • 19B • Updated • 48 • 1 - Running
Kirigami Journey
🪷How we carved a 35B MoE to fit a 24GB GPU — zero training
mcp-server-bench
This is a collection of Benchmarking results between Gradio and FastMCP
Large MoE Architecture Search (1B-2B)
Systematic search for 1B-2B MoE models. Best: bs=1, ctx=2048 achieves 0.32 loss. Top-8 routing beats top-2.
-
kshitijthakkar/moe-1083m-781m-16x8-8L-large-moe-1.3b-bs4-ctx1024
1B • Updated -
kshitijthakkar/moe-1083m-781m-16x8-8L-large-moe-1.3b-bs2-ctx2048
1B • Updated -
kshitijthakkar/moe-1083m-781m-16x8-8L-large-moe-1.3b-bs2-ctx1024
1B • Updated -
kshitijthakkar/moe-1083m-781m-16x8-8L-large-moe-1.3b-bs1-ctx2048
1B • Updated
OutageOdyssey
My submission to Agents-MCP-Hackathon
Loggenix-MOE
Collection of Loggenix Models, Eval Dataset, Demo Playground. Soon will add the training dataset.
- Build error2
Loggenix Moe 0.3B A0.1B Demo
🏢2Demo Space for my model loggenix-moe-0.3B-A0.1B
-
kshitijthakkar/loggenix-synthetic-ai-tasks-eval-with-outputs
Viewer • Updated • 28 • 16 -
kshitijthakkar/loggenix-synthetic-ai-tasks-eval_v6-with-outputs
Viewer • Updated • 170 • 25 -
kshitijthakkar/loggenix-synthetic-ai-tasks-eval_v5-with-outputs-v7-sft-v1
Viewer • Updated • 170 • 21
ICML 2026 Reproductions - agent-repro challenge
Trackio logbooks: claim-by-claim reproductions of ICML 2026 papers for the Agent Reproducibility Challenge.
- Running
Repro - MemEvolve: Meta-Evolution of Agent Memory Systems
🧬Collaborate with an AI agent to manage a shared experiment logbook
- Running
Repro - TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
🎯Explore experiment logs and sync findings with a coding agent
- Running
Repro - How to Correctly Report LLM-as-a-Judge Evaluations
🎯Log and share LLM evaluation findings in a collaborative notebook
- Running
Repro - Dependence-Aware Label Aggregation via Ising Models
🧲Explore and edit experiment logbooks with AI agent help
Kirigami: Zero-Shot Expert Carves of Qwen3.6
Qwen3.6-35B-A3B carved by expert importance to fit one consumer GPU. No training, no calibration. 800 tok/s on a laptop 5090.
-
kshitijthakkar/Kirigami-Qwen3.6-20B-A3B-NVFP4
Text Generation • 14B • Updated • 72 • 1 -
kshitijthakkar/Kirigami-Qwen3.6-24B-A3B-NVFP4
Text Generation • 16B • Updated • 285 -
kshitijthakkar/Kirigami-Qwen3.6-28B-A3B-NVFP4
Text Generation • 19B • Updated • 48 • 1 - Running
Kirigami Journey
🪷How we carved a 35B MoE to fit a 24GB GPU — zero training
DeepSeek V4 Replicas
Small-scale faithful replicas of the DeepSeek-V4 architecture for ablation and weight-transfer research.
-
kshitijthakkar/deepseek-v4-mini-300M-init
Text Generation • 0.3B • Updated • 13 -
kshitijthakkar/deepseek-v4-mini-1B-init
Text Generation • 1B • Updated • 129 • 1 -
kshitijthakkar/deepseek-v4-mini-3B-init
Text Generation • 3B • Updated • 55 • 1 -
kshitijthakkar/deepseek-v4-mini-6B-init
Text Generation • 8B • Updated • 52 • 4
mcp-server-bench
This is a collection of Benchmarking results between Gradio and FastMCP
Qwen3.5 Dense-to-MoE Weight Transfer
Qwen3.5 MoE models from dual-source weight transfer (dense backbone + 35B-A3B experts). Hybrid DeltaNet + GQA attention.
-
kshitijthakkar/qwen3.5-moe-0.87B-d0.8B
Image-Text-to-Text • 1B • Updated • 259 • 1 -
kshitijthakkar/qwen3.5-moe-2.3B-d2B
Image-Text-to-Text • 3B • Updated • 26 -
kshitijthakkar/qwen3.5-moe-4.7B-d4B
Image-Text-to-Text • 5B • Updated • 300 -
kshitijthakkar/qwen3.5-tiny-test
Image-Text-to-Text • 0.1B • Updated • 15
Large MoE Architecture Search (1B-2B)
Systematic search for 1B-2B MoE models. Best: bs=1, ctx=2048 achieves 0.32 loss. Top-8 routing beats top-2.
-
kshitijthakkar/moe-1083m-781m-16x8-8L-large-moe-1.3b-bs4-ctx1024
1B • Updated -
kshitijthakkar/moe-1083m-781m-16x8-8L-large-moe-1.3b-bs2-ctx2048
1B • Updated -
kshitijthakkar/moe-1083m-781m-16x8-8L-large-moe-1.3b-bs2-ctx1024
1B • Updated -
kshitijthakkar/moe-1083m-781m-16x8-8L-large-moe-1.3b-bs1-ctx2048
1B • Updated
Mobile MoE Architecture Search
32 MoE models from 41 experiments exploring expert count, routing, and learning rates for mobile deployment.
OutageOdyssey
My submission to Agents-MCP-Hackathon
TraceMind-AI
Collection of my submissions to MCP-1st-Birthday:- TraceMind Agent and MCP Server and smoltrace datasets generated for running evals using smoltrace.
- Sleeping10
TraceMind MCP Server
🤖10MCP server for agent evaluation with Gemini 2.5 Flash
- RunningAgents21
TraceMind AI
🧠21AI agent evaluation with MCP-powered intelligence
-
MCP-1st-Birthday/smoltrace-recruitment-tasks
Viewer • Updated • 101 • 26 -
MCP-1st-Birthday/smoltrace-smart-home-tasks
Viewer • Updated • 100 • 31
Loggenix-MOE
Collection of Loggenix Models, Eval Dataset, Demo Playground. Soon will add the training dataset.
- Build error2
Loggenix Moe 0.3B A0.1B Demo
🏢2Demo Space for my model loggenix-moe-0.3B-A0.1B
-
kshitijthakkar/loggenix-synthetic-ai-tasks-eval-with-outputs
Viewer • Updated • 28 • 16 -
kshitijthakkar/loggenix-synthetic-ai-tasks-eval_v6-with-outputs
Viewer • Updated • 170 • 25 -
kshitijthakkar/loggenix-synthetic-ai-tasks-eval_v5-with-outputs-v7-sft-v1
Viewer • Updated • 170 • 21