How you find out whether an agent is any good: held-out splits, trajectory-level scoring, and leaderboards that publish their method.
👋 Open to Work
Yauhen Bichel
YauhenBichel
·
AI & ML interests
Create new models, create harnesses, use projects with AI agentic apps, platform engineering
Recent Activity
updated a Space about 15 hours ago
YauhenBichel/privacy-gate-llm-demo updated a model about 15 hours ago
YauhenBichel/privacy-gate-llm updated a dataset 5 days ago
YauhenBichel/py-harness-intentsOrganizations
Medical imaging CNNs
Dermoscopic lesion classification — public data and models behind the CNN work I did at MoleCare (NHS App Library). Research, not clinical tools.
-
marmal88/skin_cancer
Viewer • Updated • 13.4k • 1.59k • 46 -
surajbijjahalli/ISIC2018
Viewer • Updated • 3.69k • 1.2k • 3 -
DimiT/MelanoMaven_ISIC_2017-2020
Viewer • Updated • 2.65k • 111 • 3 -
Anwarkh1/Skin_Cancer-Image_Classification
Image Classification • 85.8M • Updated • 1.2k • • 44
Small models for on-prem inference
Models small enough to run on infrastructure you control — the practical constraint when data cannot leave the estate.
-
Qwen/Qwen3-4B-Instruct-2507
Text Generation • 4B • Updated • 3.65M • • 996 -
Qwen/Qwen3-4B-Thinking-2507
Text Generation • 4B • Updated • 604k • • 619 -
meta-llama/Llama-3.2-3B-Instruct
Text Generation • 3B • Updated • 1.55M • • 2.76k -
bartowski/Llama-3.2-3B-Instruct-GGUF
Text Generation • 3B • Updated • 157k • 244
Evaluation & benchmarks
How you find out whether an agent is any good: held-out splits, trajectory-level scoring, and leaderboards that publish their method.
Medical imaging CNNs
Dermoscopic lesion classification — public data and models behind the CNN work I did at MoleCare (NHS App Library). Research, not clinical tools.
-
marmal88/skin_cancer
Viewer • Updated • 13.4k • 1.59k • 46 -
surajbijjahalli/ISIC2018
Viewer • Updated • 3.69k • 1.2k • 3 -
DimiT/MelanoMaven_ISIC_2017-2020
Viewer • Updated • 2.65k • 111 • 3 -
Anwarkh1/Skin_Cancer-Image_Classification
Image Classification • 85.8M • Updated • 1.2k • • 44
AI SRE & agentic ops
Agents that diagnose production incidents — the corpora, tooling and tool-calling models. Built around my work on OpenSRE's evaluation environment.
Small models for on-prem inference
Models small enough to run on infrastructure you control — the practical constraint when data cannot leave the estate.
-
Qwen/Qwen3-4B-Instruct-2507
Text Generation • 4B • Updated • 3.65M • • 996 -
Qwen/Qwen3-4B-Thinking-2507
Text Generation • 4B • Updated • 604k • • 619 -
meta-llama/Llama-3.2-3B-Instruct
Text Generation • 3B • Updated • 1.55M • • 2.76k -
bartowski/Llama-3.2-3B-Instruct-GGUF
Text Generation • 3B • Updated • 157k • 244