splasher / README.md
SmaugC137's picture
hf: Docker image + Hugging Face Space config (demo, api extra)
8f40b86
|
Raw
History Blame Contribute Delete
8.49 kB
metadata
title: Splasher
emoji: 💦
colorFrom: indigo
colorTo: blue
sdk: docker
app_port: 7860
pinned: false
license: mit
short_description: Label synchronized 3D/image channels into a 2D BEV grid

Splasher

Live demo — this Hugging Face Space runs splasher demo --serve (the synthetic source) inside a Docker container. The 3D cloud, camera and BEV grid are all interactive. Note: the demo serves a single shared session, so concurrent visitors edit the same grid. See the GitHub repo to run it locally.

A labeling tool with a generic core: you give it a synchronous dataset — at each timestamp, a pack of named channels (3D point cloud, camera image, pose, …) — and you label either a top-down 2D grid (BEV), the 3D points directly, or both.

First use case: traversability. But nothing is hard-wired: no imposed class schema, no imposed world semantics, and no mandatory dependency on any dataset format. apairo is just one optional input adapter.

Idea

  • Several synchronized channels, shown as references: you roam freely in the 3D cloud and look at the camera images.
  • You choose which channels to display (the Channels dock: show/hide each cloud or camera available in the source — multiple cameras and clouds supported).
  • You draw the grid of cells (top-down view): its extent and cell size, created explicitly via "New grid". Undo is per frame.
  • You drag a rectangle with the mouse on the top-down view. Depending on the target:
    • Grid: fills the covered cells with the active class (output = raster of IDs).
    • Points: assigns the class to the 3D points in the rectangle (output = per-point labels).
  • Interaction: left-drag = apply, right-drag = erase / deselect, Shift-drag or middle-button = pan, wheel = zoom. An ego marker (X = forward, Y = left) is drawn at the origin in both the BEV and the 3D views, using the same axis convention so the two read consistently.
  • Selection mode (desktop-style): left-drag adds cells to the selection, right-drag removes them (non-contiguous selections allowed), then you apply the class to the whole selection at once. Changing the grid asks for confirmation if a labeling already exists.
  • Accumulation: you can accumulate ±N frames registered by their poses into the current frame's frame of reference (a denser cloud helps labeling). The grid and labels stay per frame: a brush stroke on the accumulated cloud is de-accumulated back to each source frame. (Requires a POSE channel.)

Architecture (swappable GUI)

Layers from generic to specific — each depends only on the previous one:

splasher/
  core/      pure numpy model (grid, label targets, BEV projection, accumulation…)
  engine/    headless Session: all the state + operations, with no UI dependency.
             Returns a *semantic* ViewState (points + per-point labels + channel,
             BEV field, grid raster, selection, images) — not pixels.
  server/    FastAPI backend on the same Session + serves the web front; desktop app.
  web/       web front (vanilla, zero build) — the ONLY front (packaged in the wheel).
    vendor/  Three.js hosted locally (offline).

One front, one engine. The web front is served by the backend; the desktop app (splasher without --serve) opens that same front in a native webview (pywebview, Spotify/Electron style) — so the desktop is the web front, in a window.

The Session (splasher.engine.Session) is the flexibility point: it draws nothing, it exposes view_state() + commands (paint_rect, select_rect, apply_selection, set_frame, set_accum_radius, commit_grid, save/load, …). Plugging in another front = consuming this Session (the web does it over HTTP). Colorization stays on the front side: each front draws the primitives its own way (reference helpers in splasher.engine.render, mirrored in JS under splasher/web/src).

Installation

cd ~/dev/splasher
uv sync                 # core + engine only (numpy)
uv sync --extra app     # + desktop app (FastAPI + uvicorn + pywebview)
uv sync --extra api     # + headless web/API server (FastAPI + uvicorn)
uv sync --extra apairo  # + apairo adapter (optional)

Demo (zero external data)

uv run python examples/demo_arraysource.py   # desktop app (native window, `app` extra)
splasher                                      # empty: file-viewer mode (browse + open files)
splasher demo                                 # synthetic dataset
splasher demo --serve                         # web server → http://127.0.0.1:8077 (`api` extra)

File viewer: launched empty, Open file… browses the filesystem and opens individual point clouds (.npy/.bin/.pcd) and images into resizable views, with a clear error for unreadable / unsupported files. Independent from the dataset/labeling mode.

Dark brutalist black-&-blue design. Labelable top-down (BEV) view (underlay mode: height / density / intensity) + 3D/camera panels you can add, resize, and each bind to a channel (3D color-by: height / intensity). Sensor placements (from ChannelSpec.placement) are drawn as markers in 3D. Classes are fully editable (⚙ in the Classes panel). Three.js is hosted locally (splasher/web/vendor) → works offline.

Native desktop window: the app extra ships a Qt WebEngine (Chromium) backend, so uv sync --extra app && splasher demo opens a real native window out of the box. If no backend can start, it falls back to the default browser and keeps serving.

Input

The core consumes a Source: __len__, __getitem__(i) -> Frame, channels(). ArraySource builds one from in-memory numpy arrays. ApairoSource (apairo extra) wraps any synchronous apairo dataset.

apairo

uv sync --extra apairo --extra app                     # adapter + desktop
splasher /path/to/dataset --adapter apairo             # all channels
splasher /path/to/dataset --adapter apairo --channels lidar,cam_front,pose   # only these
splasher /path/to/dataset --adapter apairo --reference lidar --tolerance 0.05  # sync an async dataset

Two ways to pick what you work on: load everything (a synchronous dataset) and choose what to display in the UI (the Clouds (BEV) toggles + Add view per channel), or select channels at load time with --channels a,b,c. An asynchronous dataset needs --reference <channel> (and optionally --tolerance) to be synchronized first.

API (back/front mode)

splasher demo --serve [--host H --port P] starts a FastAPI server driven by the same Session as the desktop app. Main endpoints (each command returns the updated ViewState so a front renders in a single round-trip):

Method Route Purpose
GET /api/session ~static description (channels, classes, n_frames)
GET /api/view current render state
POST /api/frame, /api/class, /api/tool, /api/targets, /api/accum, /api/visibility settings
POST /api/paint, /api/select, /api/selection/apply, /api/selection/clear, /api/clear, /api/undo labeling
POST /api/grid, /api/save, /api/load grid & I/O

numpy arrays travel as {dtype, shape, data(base64)} (splasher.server.protocol), decodable directly into a TypedArray on the JavaScript side. Interactive docs at /docs.

File-viewer mode

Launched empty (splasher), Open file… browses the filesystem (type a path, Tab to complete) and opens point clouds (.npy/.bin/.pcd) and images (.png/.jpg/… or .npy HxWxC) into resizable views. The Clouds (BEV) panel selects which open clouds feed the BEV (multiple = combined), with a color mode (height/intensity/normal). The BEV grid and its labels are independent of the displayed cloud — switching/combining clouds never wipes them. Export writes the grid raster to a single .npy (default name <cloud>_bev.npy); Save writes a full session folder.

Development

uv sync --extra api          # core + engine + server (for the test suite)
uv run --extra api pytest -q # run the tests

The core/engine import without any UI dependency; tests are pure-numpy + FastAPI's TestClient. CI runs the suite on Python 3.11 and 3.12 (.github/workflows/ci.yml).