Piccaso-0.1

A 102M-parameter model that answers a text prompt with 361 brush strokes instead of pixels.

Piccaso started as a curiosity question: is it easier for a small model to paint a picture than to make one? Each painting is 361 quadratic Bezier strokes (position, curve, width, colour, on/off), rendered by a fixed renderer to PNG or exported as a real SVG.

Early research preview. Trained on 232k images for about 9.2M picture-views (~8 GPU-hours on RTX 5090s). It gets colour, light and layout right and objects some of the time. It is not a production image generator, and it is nowhere near pixel models trained on 100M+ images. The learning curves were still rising; more data and training should help a lot. Read the limitations below.

samples

What it does well, and what it doesn't

  • Good: colour (97% right on our prompt benchmark), light, mood and layout; landscapes, sunsets, interiors, food, vehicles, portraits as painted heads; single main subjects.
  • Weak: two separate objects in one scene (4%), exact shapes, identity, text, fine detail; doodle and emoji styles (out of domain).
  • By design: a painterly, quick-oil-sketch look. The 361-stroke format itself caps realism (see "ceiling" below).

Results

Caption retrieval with CLIP ViT-B/32: does the painting match its own caption better than the other captions of the set?

Top-1 among 200 Real photo Fitted 361 strokes (ceiling) Piccaso-0.1
Seen (training captions) 96.5% 83.5% 28.5%
Unseen (held-out PixelProse) 98.0% 87.5% 31.0%
DOCCI (different photo source, human captions) 80.0% n/a 9.5%

Chance is 0.5%. Seen and unseen are the same: the model generalises rather than memorising.

StrokeBench, 200 fixed prompts Score Chance
Right object, top-1 / top-5 of 80 30.3% / 56.3% 1.3% / 6.3%
Right colour (of 10) 96.9% 10%
Both objects in two-object prompts 4.4% 0.4%
Right style (of 5) 45.6% 20%

unseen Unseen captions. Columns: real photo, its fitted 361 strokes (what the format can show), two Piccaso paintings from the caption alone.

Usage

git clone https://huggingface.co/shing-dev/Piccaso-0.1 && cd Piccaso-0.1/code
pip install torch open_clip_torch ftfy regex safetensors huggingface_hub pillow numpy
python paint.py "a lighthouse on a cliff at sunset, oil painting" --n 4 --model .. --out paintings

Each painting is saved as a 512 px PNG and an SVG with 361 <path> strokes. The Long-CLIP-B text encoder (BeichenZhang/LongCLIP-B, ~600 MB) is downloaded on first run. Speed: about 0.11 s per painting on an RTX 5090 (batch 50), about 90 s on a laptop CPU (25 steps). Defaults: 25 DDIM steps, guidance 3.

Model

Output 361 strokes x 11 numbers: 165 base strokes on 4x4 / 7x7 / 10x10 grids + 196 detail strokes on a 14x14 grid
Network set diffusion transformer (DiT-style, adaLN), width 640, 11 layers, 10 heads, 102M parameters
Text Long-CLIP-B, frozen: pooled vector + cross-attention to up to 248 caption tokens
Training v-prediction, cosine schedule, self-conditioning, 25% reference-image conditioning; 8k steps at batch 256 + 7k at batch 1,024
Data 232,134 PixelProse images (clean, aesthetic >= 5), converted to strokes by gradient-descent fitting, filtered by how well the stroke render still matches the caption
Files model.safetensors (EMA weights + stroke normalisation, anchors, caption standardisation), config.json, code/

How it compares to a real image model

MobileDiffusion (Google, 2023) trains on 150M images with weeks of TPU time and runs in ~0.2 s on a phone. Piccaso-0.1 saw 232k images (about 650x fewer) for ~8 GPU-hours, and the whole project cost about $65 of rented GPUs. The gap is mostly data and compute, not the stroke format. Few-step distillation and more data are the obvious next steps.

Limitations and responsible use

  • Research preview; outputs are often wrong or abstract. Do not use it where accuracy matters.
  • Trained on web images with machine-written captions (PixelProse); it inherits their biases and gaps.
  • It cannot reproduce identities, logos or text in any recognisable way at this scale.

Links and credits

Downloads last month
12
Safetensors
Model size
0.1B params
Tensor type
F32
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train shing-dev/Piccaso-0.1