{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "# GRPO SQL Optimizer — Colab Quickstart\n", "\n", "This notebook runs a **small, reproducible GRPO training run** on the **SQL Query Optimization Environment** (DuckDB-verifiable rewards).\n", "\n", "- Repo: `OfficialAbhinavSingh/SQL-Query-Optimization-Environment-`\n", "- Goal: give judges a one-click way to rerun training and see reward/loss curves.\n", "\n", "> Tip: For a quick demo run, keep episodes small (e.g. 40–80). For a longer run, increase episodes and/or group size." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# --- 1) Clone repo ---\n", "%cd /content\n", "!rm -rf /content/SQL-Query-Optimization-Environment-\n", "!git clone https://github.com/OfficialAbhinavSingh/SQL-Query-Optimization-Environment-.git\n", "%cd /content/SQL-Query-Optimization-Environment-" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# --- 2) Install deps ---\n", "!pip -q install -r requirements.txt\n", "\n", "# sanity (optional)\n", "!openenv validate ." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# --- 3) Run a SHORT training run (judge-friendly) ---\n", "# We run train.py via import so we can override config without editing the repo.\n", "\n", "import os\n", "import train\n", "\n", "# Tune these for speed / quality\n", "train.cfg.num_episodes = 60\n", "train.cfg.group_size = 4\n", "train.cfg.output_dir = \"./checkpoints_colab\"\n", "\n", "# Optional: reduce tokens for faster iterations\n", "train.cfg.max_new_tokens = 768\n", "\n", "history = train.train()\n", "history[\"best_reward\"], len(history[\"episode_rewards\"])" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# --- 4) View curves and key outputs ---\n", "from pathlib import Path\n", "\n", "out = Path(\"./checkpoints_colab\")\n", "print(\"Outputs:\")\n", "for p in [out / \"training_curves.png\", out / \"training_history.json\"]:\n", " print(\" -\", p, \"exists=\", p.exists())\n", "\n", "display(Image(filename=str(out / \"training_curves.png\")))" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "# --- 5) Optional: generate the environment-only before/after artifact ---\n", "!python training/eval_before_after.py --save-dir results\n", "from PIL import Image\n", "display(Image.open(\"results/before_after_chart.png\"))" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "name": "python", "version": "3.10" } }, "nbformat": 4, "nbformat_minor": 5 }