vision-space / README.md
techresearchspace's picture
Upload 2 files
b8bc0f0 verified
|
Raw
History Blame Contribute Delete
3.36 kB
metadata
title: Vision Lab  Multimodal AI In-Browser
emoji: 🔬
colorFrom: blue
colorTo: indigo
sdk: static
pinned: false
license: mit
models:
  - Xenova/vit-gpt2-image-captioning
  - Xenova/owlvit-base-patch32
  - Xenova/clip-vit-base-patch32
  - Xenova/trocr-small-printed

Vision Lab

A multimodal computer-vision + language playground that runs entirely in the browser — no backend, no API keys, no server costs, and nothing you upload ever leaves your machine. Built on 🤗 Transformers.js, which executes real transformer models client-side via WebAssembly, with WebGPU acceleration when the browser supports it.

Capabilities

Tab Task Model
Caption Image → natural-language description Xenova/vit-gpt2-image-captioning
Detect Objects Draws bounding boxes for any labels you define, even ones the model wasn't explicitly trained to detect Xenova/owlvit-base-patch32
Classify Zero-shot classification against any labels you define Xenova/clip-vit-base-patch32
Read Text OCR on printed text in an image Xenova/trocr-small-printed
Semantic Search Rank a gallery of images against a natural-language query Xenova/clip-vit-base-patch32 (shared)

How it works

  • Lazy loading. Each model downloads and initializes only the first time its tab is used — not all four load up front. Progress is streamed live to the system log at the bottom of the page.
  • WebGPU / WASM fallback. On load, the app probes for a WebGPU adapter. If present, inference runs hardware-accelerated; otherwise it falls back to WASM (CPU) automatically.
  • Semantic search without embeddings math. Rather than manually computing and comparing CLIP image/text embeddings, the search tab reuses the zero-shot classification pipeline with a single candidate label (the query) per gallery image — the resulting confidence score is a valid relevance signal for ranking.
  • Browser caching. Once a model has been downloaded, the browser's own cache keeps it available on repeat visits without re-downloading.

Why this shape

This is deliberately built as a static Hugging Face Space (sdk: static) rather than a Gradio Space, so it can be hosted for free with no compute plan required — all inference happens on the visitor's own device.

Running locally

No build step. Just open index.html in a browser, or serve the folder:

python -m http.server 8000

Extending it

Ideas for going further: add a webcam capture option alongside file upload, swap in a larger captioning model for better quality (trading off download size), or add an "export gallery as JSON" button to persist search results across sessions.

A note on transformers.js task coverage

Not every task in the Python transformers library has a browser equivalent. Visual question answering, for instance, isn't currently a supported transformers.js pipeline — only its listed NLP, vision, audio, and multimodal tasks (embeddings, zero-shot classification/detection) run client-side. Object detection was used here instead of VQA for exactly that reason: it's a confirmed, supported multimodal task that still combines free-form text labels with an image.