Safetensors
GeoText1652_model / EXAMPLES.md
mhassanch's picture
Support JPEG and PNG image inputs
ac44754
|
Raw History Blame Contribute Delete
6.84 kB
# GeoText-1652 request examples
All image routes require a public RGB GeoTIFF/COG, JPEG, or PNG URL. GeoTIFFs
return CRS and bounds; JPEG and PNG inputs return `null` for those fields. Use
this image in the examples below:
```text
https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif
```
## Inputs and outcomes at a glance
| Workflow | Inputs | Primary returned fields | What to do next |
| --- | --- | --- | --- |
| Embed text | `queries` | `text_embeddings` | Search global or patch vector collections. |
| Embed one image | `image_url` | `image_embedding` | Store in an image index for image-to-image retrieval. |
| Embed image patches | `image_url`, `include_patch_embeddings` | `patch_embeddings.patches` | Search each footprint for text-guided localization. |
| Rank text descriptions | `image_url`, `queries` | `ranked_queries` | Determine which description best matches the image. |
| Rank a large raster | image, queries, `tile_size` | `ranked_queries` with `tile_index` | Find the best tile for each query. |
| Propose matching regions | image, queries, `include_bboxes` | `text_conditioned_boxes` / `stitched_text_conditioned_boxes` | Inspect semantic candidate regions. |
| Persist patch index data | image, `store_patch_embeddings` | `patch_embeddings_storage` | Download Parquet and run vector similarity search. |
## 1. Health check
```bash
curl https://YOUR-ENDPOINT.endpoints.huggingface.cloud/health
```
## 2. Text-only embeddings
`POST /embed/text`
```json
{
"queries": ["a parking lot", "a large building", "a sports field"]
}
```
Returns one normalized 256-dimensional embedding per query. This is the route
to use when searching a previously stored image or patch-vector collection.
## 3. Global image embedding
`POST /embed/image`
```json
{
"image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
"include_global_embedding": true
}
```
### JPEG or PNG input
The same image routes accept ordinary RGB imagery. This public OpenDroneMap
drone image is a PNG, so its response has `null` CRS and bounds metadata:
```json
{
"image_url": "https://raw.githubusercontent.com/pierotofy/dataset_banana/master/banana.png",
"include_global_embedding": true,
"include_patch_embeddings": true
}
```
## 4. Image patch embeddings
`POST /embed/image`
```json
{
"image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
"include_global_embedding": false,
"include_patch_embeddings": true
}
```
Each patch includes `source_pixel_xyxy` and a normalized 256-dimensional
embedding. The model returns a 12 by 12 patch grid for one 384 px model input.
## 5. Tiled global and patch image embeddings
`POST /embed/image`
```json
{
"image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
"tile_size": 384,
"tile_overlap": 64,
"include_global_embedding": true,
"include_patch_embeddings": true,
"return_tile_results": false
}
```
The global vector is an area-weighted, L2-normalized mean of tile vectors.
Patch vectors keep their `tile_index` and original-image pixel footprints.
## 6. Store tiled patch embeddings in the Hub
`POST /embed/image`
```json
{
"image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
"tile_size": 384,
"tile_overlap": 64,
"include_global_embedding": true,
"store_patch_embeddings": true,
"output_prefix": "geotext/demo-image"
}
```
Endpoint secrets/configuration required:
```text
HF_TOKEN=<write-capable token>
HF_BUCKET=your-namespace/your-embedding-dataset
```
This stores a compressed Parquet file but leaves the large patch collection out
of the response. Add `"include_patch_embeddings": true` to return it too.
## 7. Image-level text ranking
`POST /infer` (or `POST /` in the Inference Endpoints Playground)
```json
{
"image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
"queries": ["a parking lot", "a building", "a sports field"]
}
```
Returns `ranked_queries`, sorted by descending cosine similarity. This answers
questions such as: *which of these descriptions best matches this image?*
## 8. Search with global image and text vectors
`POST /infer`
```json
{
"image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
"queries": ["a parking lot", "a building"],
"include_embeddings": true
}
```
## 9. Large-raster ranking, localization, and stitched region proposals
`POST /infer`
```json
{
"image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
"queries": ["a parking lot", "a building"],
"tile_size": 384,
"tile_overlap": 64,
"include_embeddings": true,
"include_bboxes": true,
"return_tile_results": false
}
```
Per-query similarity is the maximum matching tile score. Region proposals are
converted to original-image pixels and deduplicated with non-maximum
suppression.
The `tile_index` in `ranked_queries` identifies the strongest tile for each
query. Add `return_tile_results: true` to receive all tile scores and inspect
the full ranking surface.
## 10. Inspect every tile
`POST /infer`
```json
{
"image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
"queries": ["a parking lot"],
"tile_size": 384,
"tile_overlap": 64,
"include_embeddings": true,
"include_patch_embeddings": true,
"include_bboxes": true,
"return_tile_results": true
}
```
This is for debugging only: it returns all tile results and can be very large.
## 11. Text-to-patch localization from stored embeddings
This is a two-step workflow. First create or retrieve a text vector:
`POST /embed/text`
```json
{
"queries": ["a parking lot"]
}
```
Then compute dot-product similarity between that 256-D vector and the
`embedding` column in `patch_embeddings.parquet`. Sort descending and use each
result's `source_pixel_xyxy` field to draw the matching locations on the
original image. This is the scalable text-guided localization workflow.
## 12. Image-to-image retrieval
Create a global vector for a query image:
`POST /embed/image`
```json
{
"image_url": "https://huggingface.co/datasets/geobase/geoai-cogs/resolve/main/geoembeddings-demo/building-detection_sm.tif",
"include_global_embedding": true
}
```
Store `image_embedding` in a vector database or local matrix with global
vectors from other images. Rank candidates by dot product to find visually and
semantically similar images.