File size: 2,983 Bytes
0445247
 
fc127f3
0445247
 
 
fc127f3
 
0445247
fc127f3
 
0445247
fc127f3
 
 
0445247
 
fc127f3
0445247
fc127f3
 
0445247
fc127f3
 
0445247
 
 
fc127f3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0445247
 
 
fc127f3
 
 
 
 
0445247
 
 
fc127f3
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
---
license: apache-2.0
library_name: pytorch
tags:
  - gan
  - image-generation
  - dcgan
  - logo
  - from-scratch
  - small-model
  - pytorch
metrics:
  - mode-collapse
  - spatial-coherence
model_type: dcgan
---

# logo-gan

A small **DCGAN** trained **from scratch** to generate 64×64 company-logo-style
images. Trained on 1,500 real logos resized to 64×64×3.

This is the deliverable for [model-requests #1](https://huggingface.co/spaces/Compactbot/model-requests/discussions/1)
("a GAN that learns to make company logos").

## What it is

- **Architecture**: DCGAN. Generator = linear latent→256×8×8, then 3×
  ConvTranspose2d (256→128→64→3, Tanh out). Discriminator = 3× Conv2d
  (3→64→128→256) + AdaptiveAvgPool + Linear→1.
- **Params** (learnable): **generator 2,805,123 + discriminator 659,585 = 3,464,708**.
  (The saved checkpoint also carries BatchNorm running-stat buffers, so a raw
  numel count over all tensors reads 3,498,756 — the extra ~34k are non-learnable
  running mean/var, not parameters.)
- **Latent**: 128-dim. **Output**: 64×64×3, [-1, 1].
- **Training**: 8,000 steps, batch 16, Adam (lr 2e-4, β=(0.5, 0.999)),
  non-saturating GAN objective, seeded 0. Trained on an RTX 5090 in ~64s.

## Data

1,500 logos (64×64×3, float 0–1), assembled from public logo datasets on the Hub
and cached to `logos_big.npy`.

## Quality — measured, not asserted

Generated 64 samples (seed 42) from `final.pt` and measured:

| Check | Value | Reading |
|---|---|---|
| Min pairwise L2 (64 samples) | 51.7 | **No mode collapse** (0.0% of pairs < 0.01) |
| Mean pairwise L2 | 103.9 | Samples are diverse |
| Adjacent-pixel mean \|diff\| | 0.109 | Structured, not noise (real data 0.057, pure noise ~0.4–0.6) |
| Per-channel std | 0.85 | Full dynamic range used |

So the generator is **not** collapsed and **not** producing noise — it makes
diverse, spatially-coherent, logo-shaped color fields.

## What it is NOT

This is a 3.5M-param DCGAN on 1,500 images. It produces **logo-shaped blobs and
color fields**, not crisp, legible, trademark-accurate logos. At this scale and
data budget, expect abstract logo-likes, not usable brand marks. That is the
honest ceiling for this recipe; a real logo pipeline needs a diffusion model on
a much larger, cleaner dataset.

## Files

- `final.pt` — generator + discriminator state dicts (`g`, `d`), plus `step`, `zdim`.
  SHA256 `114765c79dc23099655d9e7477648c5a8c2b90fda03b7f3dbd4714f45f27b95f`.
- `grid_final.png` — 64 generated samples (8×8 grid).
  SHA256 `e9eee93950397a9f29028384b34809df432d0dfcbdeb4b1cce30328c4504bf5b`.
- `train_logo_gan_v2.py` — the exact training script (seeded, reproducible).

## Reproduce

```python
import torch
from train_logo_gan_v2 import G
ck = torch.load("final.pt", map_location="cpu", weights_only=False)
g = G(ck["zdim"]); g.load_state_dict(ck["g"]); g.eval()
with torch.no_grad():
    imgs = g(torch.randn(64, 128))   # (64,3,64,64) in [-1,1]
```