Replace mermaid block in README.md with ASCII diagram (renders in HF blob viewer) 576b083 verified yashash045 commited on Apr 26
Replace mermaid block in BLOG.md with ASCII diagram (renders in HF blob viewer) 659254f verified yashash045 commited on Apr 26
Update model_cards/sft_adapter_README.md: humanise voice, narrative, captions f9f862c verified yashash045 commited on Apr 26
Update gradio_app.py: humanise voice, narrative, captions 918032b verified yashash045 commited on Apr 26
Update docs/architecture.md: humanise voice, narrative, captions 83e5379 verified yashash045 commited on Apr 26
Update devops_pipeline_gym_colab.ipynb: humanise voice, narrative, captions d9d5941 verified yashash045 commited on Apr 26
BLOG: add header rows to TL;DR + SFT-setup tables (was rendering empty bars) 1d23f7f verified yashash045 commited on Apr 26
frontier_results: add 5 frontier baseline JSONs (evidence behind comparison numbers) 2ee41ce verified yashash045 commited on Apr 26
Notebook: strip PREVIOUS HANDOFF NOTES line; remove broken Kaggle badge 094e300 verified yashash045 commited on Apr 26
grpo_train: strip PREVIOUS HANDOFF NOTES from prompt template c603d39 verified yashash045 commited on Apr 26
README: inline architecture mermaid; +200 steps in GRPO; drop broken Kaggle badge af9094b verified yashash045 commited on Apr 26
BLOG: rebuild as full artifact (mermaid arch, tables, results, GRPO chart, code, artifacts list) 79f7353 verified yashash045 commited on Apr 26
Honest reframe: baseline is Qwen2.5-7B (not Qwen3-1.7B); trained 1.7B beats every untrained baseline tested 4083da2 verified yashash045 commited on Apr 26
Honest reframe: baseline is Qwen2.5-7B (not Qwen3-1.7B); trained 1.7B beats every untrained baseline tested f33d61d verified yashash045 commited on Apr 26
Colab: drop Track-IO link, add concrete +1.156 number in wrap-up da8ba22 verified yashash045 commited on Apr 26
Colab notebook with executed output cells visible to judges d287b59 verified yashash045 commited on Apr 26
Lock in real numbers: +1.156 same-model delta, beats all 70B+ frontier (BLOG.md) 3c01c86 verified yashash045 commited on Apr 26
Lock in real numbers: +1.156 same-model delta, beats all 70B+ frontier (README.md) caa23ec verified yashash045 commited on Apr 26
Colab: align baseline+trained temp to 0.3 (apples-to-apples) 814072d verified yashash045 commited on Apr 26
Colab: add Kaggle badge, honest GRPO cell, cleaner wrap-up 32e0543 verified yashash045 commited on Apr 26
Hackathon submission: new README (3-5 min read), BLOG.md narrative, frontier baselines, design-principles framing 40de84e verified yashash045 commited on Apr 26
README: add Colab/Kaggle badges + judge re-run instructions + adapter links 20bea6e yashash04 commited on Apr 25
Submission deliverables: Colab notebook + per-step CSV + trainer log upload cc203b9 yashash04 commited on Apr 25
Phase M Stage B: align fully with kube-sre-gym (sid-rp) patterns + A100 prep 9f7a631 yashash04 commited on Apr 25
Phase M Option 3 v3: 3 fixes lifted from kube-sre-gym (sid-rp) GRPO patterns a598d78 yashash04 commited on Apr 25
Phase M Option 3 retry: vLLM debug logging + reduce gpu_memory + max-steps 30 f2a191c yashash04 commited on Apr 25
Phase M debug: log raw parse_completion data to see config_edits shape 1866fe2 yashash04 commited on Apr 25
Phase M Stage A v2: fix 3 compound bugs preventing GRPO learning c6c5f7f yashash04 commited on Apr 25
Phase M Stage A: revert wrapper + trainer to 4940a15 (last known-working non-vLLM state) bf1c711 yashash04 commited on Apr 25
Phase M debug Approach 2.5: pin peft>=0.18.0,<0.19 for SFT adapter loading 8d77eb7 yashash04 commited on Apr 25
Phase M Stage A vLLM Approach 1: pin vllm==0.6.3 (avoid v1-engine graph bug) 2599fd4 yashash04 commited on Apr 25
Phase M Stage A vLLM: enable use_vllm + fast_inference for ~3x speedup 22623bb yashash04 commited on Apr 25