Safetensors
GeoText1652_model / EXAMPLES.md
mhassanch's picture
Support JPEG and PNG image inputs
ac44754
|
Raw History Blame Contribute Delete
6.84 kB

GeoText-1652 request examples

All image routes require a public RGB GeoTIFF/COG, JPEG, or PNG URL. GeoTIFFs return CRS and bounds; JPEG and PNG inputs return null for those fields. Use this image in the examples below:

https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif

Inputs and outcomes at a glance

Workflow Inputs Primary returned fields What to do next
Embed text queries text_embeddings Search global or patch vector collections.
Embed one image image_url image_embedding Store in an image index for image-to-image retrieval.
Embed image patches image_url, include_patch_embeddings patch_embeddings.patches Search each footprint for text-guided localization.
Rank text descriptions image_url, queries ranked_queries Determine which description best matches the image.
Rank a large raster image, queries, tile_size ranked_queries with tile_index Find the best tile for each query.
Propose matching regions image, queries, include_bboxes text_conditioned_boxes / stitched_text_conditioned_boxes Inspect semantic candidate regions.
Persist patch index data image, store_patch_embeddings patch_embeddings_storage Download Parquet and run vector similarity search.

1. Health check

curl https://YOUR-ENDPOINT.endpoints.huggingface.cloud/health

2. Text-only embeddings

POST /embed/text

{
  "queries": ["a parking lot", "a large building", "a sports field"]
}

Returns one normalized 256-dimensional embedding per query. This is the route to use when searching a previously stored image or patch-vector collection.

3. Global image embedding

POST /embed/image

{
  "image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
  "include_global_embedding": true
}

JPEG or PNG input

The same image routes accept ordinary RGB imagery. This public OpenDroneMap drone image is a PNG, so its response has null CRS and bounds metadata:

{
  "image_url": "https://raw.githubusercontent.com/pierotofy/dataset_banana/master/banana.png",
  "include_global_embedding": true,
  "include_patch_embeddings": true
}

4. Image patch embeddings

POST /embed/image

{
  "image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
  "include_global_embedding": false,
  "include_patch_embeddings": true
}

Each patch includes source_pixel_xyxy and a normalized 256-dimensional embedding. The model returns a 12 by 12 patch grid for one 384 px model input.

5. Tiled global and patch image embeddings

POST /embed/image

{
  "image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
  "tile_size": 384,
  "tile_overlap": 64,
  "include_global_embedding": true,
  "include_patch_embeddings": true,
  "return_tile_results": false
}

The global vector is an area-weighted, L2-normalized mean of tile vectors. Patch vectors keep their tile_index and original-image pixel footprints.

6. Store tiled patch embeddings in the Hub

POST /embed/image

{
  "image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
  "tile_size": 384,
  "tile_overlap": 64,
  "include_global_embedding": true,
  "store_patch_embeddings": true,
  "output_prefix": "geotext/demo-image"
}

Endpoint secrets/configuration required:

HF_TOKEN=<write-capable token>
HF_BUCKET=your-namespace/your-embedding-dataset

This stores a compressed Parquet file but leaves the large patch collection out of the response. Add "include_patch_embeddings": true to return it too.

7. Image-level text ranking

POST /infer (or POST / in the Inference Endpoints Playground)

{
  "image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
  "queries": ["a parking lot", "a building", "a sports field"]
}

Returns ranked_queries, sorted by descending cosine similarity. This answers questions such as: which of these descriptions best matches this image?

8. Search with global image and text vectors

POST /infer

{
  "image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
  "queries": ["a parking lot", "a building"],
  "include_embeddings": true
}

9. Large-raster ranking, localization, and stitched region proposals

POST /infer

{
  "image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
  "queries": ["a parking lot", "a building"],
  "tile_size": 384,
  "tile_overlap": 64,
  "include_embeddings": true,
  "include_bboxes": true,
  "return_tile_results": false
}

Per-query similarity is the maximum matching tile score. Region proposals are converted to original-image pixels and deduplicated with non-maximum suppression.

The tile_index in ranked_queries identifies the strongest tile for each query. Add return_tile_results: true to receive all tile scores and inspect the full ranking surface.

10. Inspect every tile

POST /infer

{
  "image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
  "queries": ["a parking lot"],
  "tile_size": 384,
  "tile_overlap": 64,
  "include_embeddings": true,
  "include_patch_embeddings": true,
  "include_bboxes": true,
  "return_tile_results": true
}

This is for debugging only: it returns all tile results and can be very large.

11. Text-to-patch localization from stored embeddings

This is a two-step workflow. First create or retrieve a text vector:

POST /embed/text

{
  "queries": ["a parking lot"]
}

Then compute dot-product similarity between that 256-D vector and the embedding column in patch_embeddings.parquet. Sort descending and use each result's source_pixel_xyxy field to draw the matching locations on the original image. This is the scalable text-guided localization workflow.

12. Image-to-image retrieval

Create a global vector for a query image:

POST /embed/image

{
  "image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
  "include_global_embedding": true
}

Store image_embedding in a vector database or local matrix with global vectors from other images. Rank candidates by dot product to find visually and semantically similar images.