Spaces:
Running on Zero
Running on Zero
File size: 9,833 Bytes
4978d0d 587e1f0 4978d0d 587e1f0 4978d0d 587e1f0 d53cfbe 587e1f0 d53cfbe 587e1f0 d53cfbe 587e1f0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 | ---
title: Facade
emoji: π
colorFrom: blue
colorTo: red
sdk: gradio
app_file: app.py
pinned: false
license: mit
short_description: Photograph a building, find its architectural style
---
# Facade
**Point a camera at a building. Find out what you are looking at.**
Upload a photograph of a facade and the app returns the three closest
architectural styles from a 1,000-plate synthetic reference corpus, a reading
written for *your* photograph at query time, the measurements behind the match,
and real named buildings near you in the same style.
- **Dataset:** [Jonathandav/facade-styles](https://huggingface.co/datasets/Jonathandav/facade-styles)
- **Corpus:** 1,000 generated plates Β· 20 styles Β· 50 each
- **Retrieval model:** `laion/CLIP-ViT-B-32-laion2B-s34B-b79K`
- **Reading model:** `Qwen/Qwen2.5-1.5B-Instruct`
---
## The idea
Architectural style is a family resemblance rather than a fact about a
building. That makes it a good target for *synthetic* reference data: a
generated Bauhaus facade **is** a Bauhaus facade, because there is no external
specimen it can contradict.
It also makes style unusually measurable. Brutalism is low-saturation and
heavy-massed; Gothic Revival is vertical-edge dominant; Bauhaus is horizontal;
Ottoman Revival is curve-rich. Every one of those has a pixel-level correlate,
so "did the generator do what it was told?" is a computed quantity rather than
an opinion.
The whole project is built on that: a corpus specified attribute by attribute,
audited by measurement, and finally tested against real photographs.
---
## How it works
```
style_seed.csv 20 styles, each with massing, material, window
β rhythm, roofline, palette, and two EXPECTED
β measurements (saturation band, edge orientation)
βΌ
sample_plates() 50 specs per style; view / crop / light / condition
β CYCLED, not sampled β so they cannot correlate
βΌ with style
SDXL-Turbo style name + period + key features lead the prompt
β acceptance gate measures every render and retries
βΌ with a fresh seed until it passes
1,000 plates βββΊ LAION-CLIP embeddings βββΊ index_embeddings.npy
β
βΌ
Qwen2.5-1.5B style readings, barred from naming architects,
asserting dates, or claiming heritage status
```
At query time: embed the photograph β cosine similarity against the index β
optionally rerank with an OpenStreetMap prior over what actually stands nearby
β generate a reading from the matched style, the runner-up, and the
measurements taken off the photograph.
---
## What the corpus can identify
Bauhaus International Β· Brutalist Β· Art Deco Β· Ottoman Revival Β· Mandate
Eclectic Β· Neoclassical Β· Gothic Revival Β· Venetian Gothic Β· Georgian Β·
Victorian Terrace Β· Mid-century Modern Β· Postmodern Β· Deconstructivist Β·
Contemporary Curtain Wall Β· Mediterranean Vernacular Β· Adobe Pueblo Β· Soviet
Constructivist Β· Scandinavian Functionalist Β· Industrial Warehouse Β· Japanese
Metabolist
Anything outside these twenty is still forced onto the nearest of them. The app
says so in the interface, because a user should know that before trusting a
result.
---
## Did the generator obey?

Each style declares an expected saturation level. That is an **ordinal** claim β
brutalism less saturated than Art Deco β not a value on a calibrated scale.
**Spearman Ο = 0.67, p β 1.5 Γ 10β»ΒΉΒ³Β², n = 1000.** Declared rank predicts
measured saturation strongly.
The first version of this test scored **8%** compliance, because the bands were
absolute thresholds chosen before any image existed: measured saturation ran
0.30β0.85 while "medium" had been guessed at 0.22β0.50. Recalibrating the bands
to corpus quantiles and reporting rank agreement is the correct test, and it is
a better methodological story than thresholds that happened to work first time.
---
## Where will it get confused?

Plotting each style's centroid in saturation Γ edge-orientation space produces
a **prediction, before any embedding model runs**, of which pairs will be
conflated:
| pair | centroid distance |
|---|---|
| Neoclassical β Postmodern | 0.41 |
| Japanese Metabolist β Scandinavian Functionalist | 0.62 |
| Mandate Eclectic β Neoclassical | 0.65 |
| Brutalist β Japanese Metabolist | 0.65 |
| Georgian β Victorian Terrace | 0.75 |
These are architecturally meaningful, not arbitrary. Postmodernism is *defined*
by ironic quotation of classical form. Mandate Eclectic borrowed classical
proportion directly. Brutalism and Metabolism are contemporaneous raw-concrete
movements. A purely pixel-level metric recovered relationships an architectural
historian would recognise.
**Georgian β Victorian Terrace was confirmed** by the CLIP confusion matrix at
0.16 β the top confusion by a wide margin. The others were not, and CLIP's
actual confusions cluster differently: Bauhaus International acts as an
*attractor* for the modernist styles, with Brutalist, Soviet Constructivist and
Japanese Metabolist all leaking toward it. That is a semantic effect the pixel
metric cannot see.
One prediction confirmed and four refuted with a coherent explanation is a
better result than five vague hits.
---
## Does style dominate the nuisance factors?

Style explains **43β66%** of variance in every measured feature. Crop, view and
condition sit at essentially **zero** β which validates the sampler: those axes
are cycled rather than drawn independently, so they cannot correlate with
style. Had they done so, an embedding model could have scored well by learning
framing instead of architecture.
Light is the only meaningful contaminant, and only on the photometric features,
which is physically unsurprising.
---
## Retrieval
Image-to-image over the corpus, style as the label, chance β 0.049.
| model | Recall@1 | Recall@3 | MRR | lift |
|---|---|---|---|---|
| SigLIP-B/16 | 0.972 | 0.994 | 0.983 | 19.8Γ |
| LAION-CLIP-B/32 | 0.970 | 0.991 | 0.981 | 19.8Γ |
| CLIP-B/32 | 0.955 | 0.988 | 0.972 | 19.5Γ |
**These are at ceiling and the comparison cannot resolve.** Plates of the same
style share nearly identical prompts, differing only in angle, crop, light and
condition, so they are near-siblings. Every slice falls between 0.975 and
1.000. The differences between models are inside noise.
---
## Sim-to-real β the result that matters
17 hand-labelled photographs of real buildings in Tel Aviv, Jaffa, Paris and
the Greek islands, retrieved against the wholly synthetic index.
| model | real R@3 | synthetic R@3 | gap |
|---|---|---|---|
| **LAION-CLIP-B/32** | **0.94** (16/17) | 0.991 | **β0.05** |
| SigLIP-B/16 | 0.88 (15/17) | 0.994 | β0.11 |
| CLIP-B/32 | 0.76 (13/17) | 0.988 | β0.22 |
The synthetic benchmark ranked SigLIP first by 0.002 on Recall@1. On real
photographs LAION wins and has less than half the domain gap. **Selecting on
the saturated benchmark would have shipped the weaker model.** A benchmark at
ceiling does not merely fail to discriminate β it discriminates wrongly.
Both failure cases are architecturally sensible: a Bauhaus building retrieved
as Soviet Constructivist (contemporaneous interwar modernism, strip windows,
white render), and a night photograph of a Mandate building retrieved as
Mediterranean Vernacular (both cream render with shutters).
---
## Generator failures found along the way
Four, plus one measurement bug. **Three of the four were silent** β no
exception, no warning, and plausible-looking output.
| # | failure | how it was caught | silent |
|---|---|---|---|
| 1 | SDXL's pooled embedding overrode every style | Bauhaus, Brutalist and Gothic all rendered as the same beige office block | yes |
| 2 | detail crops destroyed the class signal | style unrecoverable from a cropped window grid | yes |
| 3 | `key_features` never reached the image prompt | code review after failure 1 | no |
| 4 | prompts truncated at 77 tokens | validator token-budget check | yes |
The fifth was a **metric** bug: the orientation classifier could only emit three
of the five declared labels, so every style declaring `curved` or `diagonal`
scored zero and looked like a total generator failure. Fixing the measurement
raised orientation compliance from 60% to 83%.
A broken metric can invent a failure as easily as a broken generator can hide
one. Both argue for the same thing: instrumented acceptance criteria, not
visual inspection.
---
## Limitations
- Visual-similarity search, **not** authoritative identification.
- No plate depicts a real building; none is attributable to any architect.
- The real-photograph evaluation is small (n = 17) and unevenly distributed
across styles.
- Style boundaries blur where visual grammars genuinely overlap.
- Location reranking depends on OpenStreetMap tag density, which varies
enormously by city β Paris returns far more than Tel Aviv.
- The app makes no claim about any building's date, architect, or heritage
status, and both models are explicitly prompted not to.
## Using it
Include the whole building where you can. Style lives in massing, roofline and
silhouette, and a cropped window grid discards all three.
For the location features, type a place with its country β "Rothschild
Boulevard, Tel Aviv, Israel" β and check the resolved name before running.
Geocoders will happily match a street name on the wrong continent. |