doodle-duel / README.md
AwsheshNath's picture
README: add Team section with HF username (AwsheshNath)
c394916 verified
|
Raw
History Blame Contribute Delete
8.72 kB
---
title: Doodle Duel
emoji: ๐ŸŽจ
colorFrom: pink
colorTo: blue
sdk: gradio
sdk_version: 6.16.0
app_file: app.py
pinned: false
license: mit
short_description: Draw a word, a small vision model guesses it in real time
tags:
- build-small-hackathon
- thousand-token-wood
- track:wood
- achievement:offbrand
- achievement:llama
---
# ๐ŸŽจ Doodle Duel
![Doodle Duel โ€” robot correctly guessing "helicopter" from a simple drawing](screenshot.png)
Pictionary against an AI vision model โ€” built for the Hugging Face **Build Small Hackathon** (Chapter Two: *An Adventure in Thousand Token Wood*), track: **Thousand Token Wood**.
You pick a difficulty, choose a word, draw it on the canvas, and a small vision model (โ‰ค32B) watches your canvas. Every few seconds it shouts its top guesses and shows its reasoning. The moment it recognises your drawing, you score โ€” faster means more points. Stump it until time runs out and *you* win the round.
**The AI's vision is the whole game.**
## โ–ถ๏ธ Demo & play
- **Watch the demo:** [youtube.com/watch?v=2icnvsOzWGM](https://www.youtube.com/watch?v=2icnvsOzWGM)
- **Play it live:** [build-small-hackathon/doodle-duel](https://huggingface.co/spaces/build-small-hackathon/doodle-duel)
- **Post on X:** [x.com/AwsheshNath](https://x.com/AwsheshNath/status/2066533146330538328)
## ๐Ÿ‘ฅ Team
- **[AwsheshNath](https://huggingface.co/AwsheshNath)** โ€” solo submission. Registered for the hackathon and a member of the `build-small-hackathon` org.
## โณ First time here? Warm up the AI (please read)
The vision model runs on a Hugging Face **Inference Endpoint that scales to zero when idle**. So if nobody has played for a while, the **first guess triggers a cold start (~1โ€“2 minutes)** while the model spins back up. This is normal โ€” it only happens on the first request after inactivity; after that, guesses come every few seconds.
**How to tell if the AI is ready** โ€” watch the status pill right under the title:
| Pill | Meaning |
|------|---------|
| ๐ŸŸก *Connecting to the AIโ€ฆ* / *Waking the AIโ€ฆ (cold start, up to ~2 min)* | Still warming up โ€” hang on, don't worry |
| ๐ŸŸข *AI ready* | Good to go โ€” the robot will start guessing within seconds |
| ๐Ÿ”ด *AI offline โ€” retryingโ€ฆ* | Temporary endpoint hiccup โ€” give it a moment |
**Tip:** you can start drawing while it warms up. The instant the pill turns ๐ŸŸข *AI ready*, the robot reads your current canvas and begins guessing โ€” no need to redraw.
## How it plays
- Pick a difficulty: **๐ŸŸข Easy** (everyday objects), **๐ŸŸก Medium** (everything), **๐Ÿ”ด Hard** (brands, famous people, characters, landmarks)
- Choose one of three secret words
- **Draw** on the canvas. A 2-minute countdown ring ticks down
- **Drawing tools** (canvas toolbar): adjustable **brush size and colour**, **eraser**, **undo / redo**, **clear**, and zoom โ€” so you can sketch rough or detailed
- The robot looks at your canvas every few seconds and guesses โ€” even if you've stopped drawing, so it keeps trying fresh words it hasn't guessed yet
- **Right-hand panel** shows the robot's live reasoning ("I see a round shape with legsโ€ฆ")
- **Guess matches โ†’ you score** (time-scaled + streak bonus)
- **Timer runs out โ†’ you stumped the robot.** Pick a difficulty again to keep going
- **Hint button** reveals letters of the answer to the robot (costs points)
- Streak and best-score tracked across rounds
## ๐ŸŽฎ Game modes โ€” and why Hard mode is the best part
- **๐ŸŸข Easy** โ€” everyday, clearly drawable things (animals, food, vehicles, household objects).
- **๐ŸŸก Medium** โ€” the full word bank, easy and hard mixed.
- **๐Ÿ”ด Hard** โ€” **the highlight of the game.** The answers are brands, famous people, fictional characters, and landmarks โ€” things you can't draw literally. The model has to **think laterally from the clues you give it**: jersey numbers, initials, logos, symbols, silhouettes, even song-title hints. It's not shape-matching any more; it's reasoning.
**Amazing examples (some shown in the demo video):**
- โœ๏ธ Sketch a cue for **"Rap God"** โ†’ the model reasons its way to **Eminem** ๐ŸŽค *(demonstrated in the [YouTube demo](https://www.youtube.com/watch?v=2icnvsOzWGM))*
- โœ๏ธ A swoosh โ†’ **Nike**
- โœ๏ธ A web and a mask โ†’ **Spider-Man**
- โœ๏ธ Jersey **#7** with **"CR"** โ†’ **Ronaldo**
Watching the right-hand "Robot's thoughts" panel reason through these out-of-the-box clues is the most fun part of Doodle Duel.
## ๐Ÿง  Under the hood (technical)
- **Model:** **Qwen3.6-27B** (vision-language), 4-bit quantized โ€” served file `Qwen3.6-27B-Q4_K_M.gguf` plus an F16 multimodal projector (`mmproj-F16.gguf`) for image input. ~27B parameters, comfortably within the hackathon's โ‰ค32B "small model" rule.
- **Serving:** runs on **llama.cpp** (CUDA `server`) behind an **OpenAI-compatible Hugging Face Inference Endpoint** (scale-to-zero). The endpoint runs the official **`ghcr.io/ggml-org/llama.cpp:server-cuda`** container, serving model id `Qwen3.6-27B-Q4_K_M.gguf` (auto-detected from `/v1/models`).
- **Per guess:** the canvas is sent as a base64 PNG together with a Pictionary prompt that includes the answer's **letter count, word count, and a revealed-letter dash pattern** (so guesses are constrained to the right shape). The model replies in exactly two lines โ€” `THINK:` (the reasoning you see live) and `GUESS:` (top 3, most likely first).
- **Token budget:** native chain-of-thought is disabled for speed and the response is capped at 200 output tokens, so each guess runs on roughly **~500 tokens total** โ€” the entire point of the *Thousand Token Wood* track.
- **Sequential, rate-limited guessing:** at most one in-flight request per session and โ‰ฅ4 s between guesses, so the Space never floods (or 429s) the endpoint.
## ๐ŸŽ–๏ธ Bonus quests earned
- ๐Ÿฆ™ **Llama Champion** โ€” the model runs through the **llama.cpp** runtime. The Hugging Face Inference Endpoint uses the official **`ghcr.io/ggml-org/llama.cpp:server-cuda`** container, serving `Qwen3.6-27B-Q4_K_M.gguf`.
- ๐ŸŽจ **Off-Brand (custom UI)** โ€” the entire interface is a hand-built dark "arcade" theme (custom `style.css`, custom HTML topbar, countdown ring, robot reasoning panel and chip bubble) that pushes well past the default Gradio look.
## Model configuration / self-hosting
Needs an OpenAI-compatible vision endpoint. Set these Space secrets:
| Secret | Example | Notes |
|--------|---------|-------|
| `MODEL_BASE_URL` | `https://your-endpoint` | vLLM / llama.cpp server |
| `MODEL_NAME` | `auto` | auto-detects from `/v1/models` |
| `MODEL_API_KEY` | `your-key` | optional, for authenticated endpoints |
No `MODEL_BASE_URL` โ†’ runs in **mock mode** (random guesses, for UI testing).
Recommended model: any vision-capable โ‰ค32B model. **This Space runs `Qwen3.6-27B-Q4_K_M.gguf`** (Qwen3.6-27B, 4-bit `Q4_K_M` quant) on llama.cpp.
## Tuning
| Var | Default | Meaning |
|-----|---------|---------|
| `ROUND_SECONDS` | 120 | Seconds per round |
| `GUESS_MIN_INTERVAL` | 4.0 | Min seconds between AI guesses |
| `POLL_INTERVAL` | 5.0 | Background guess cadence |
| `BRUSH_SIZE` | 4 | Default brush width |
## Files
- `app.py` โ€” Gradio UI: canvas, countdown timer, word choices, controls, status pill
- `game.py` โ€” round/clock/guess/scoring/hint logic
- `vision_client.py` โ€” sends canvas image to model, parses THINK/GUESS response
- `visuals.py` โ€” timer ring, robot speech bubble, reasoning panel, win/lose
- `config.py` โ€” word bank (1000+ words, 18 categories, easy/hard tagged) + timing config
- `style.css` โ€” bright playful theme
---
Been a while since I posted, but I built something fun and wanted to share it.
I made a little game called **Doodle Duel** for the Hugging Face Build Small Hackathon. The idea is simple: you draw something on a canvas, and a small vision model (Qwen3.6-27B, under 32B parameters) watches what you're drawing and tries to guess it in real time โ€” like reverse Pictionary against an AI.
What I find cool about it is that you can actually see the model's reasoning as it watches your strokes. It'll say things like "I see a round shape... maybe a wheel? a ball?" and keep guessing even if you stop drawing, trying words it hasn't said yet. There are easy, medium, and hard word modes, a hint system, and a 2-minute timer โ€” and in Hard mode it'll reason its way from a "Rap God" doodle all the way to **Eminem**.
The whole thing runs on roughly ~500 tokens per guess, which was a genuinely fun constraint to design around.
Built for the **[Build Small Hackathon](https://huggingface.co/spaces/build-small-hackathon/registration)** โ€” Track: Thousand Token Wood.