locateanything_3b / README.md
IMvision12's picture
Fix Collection badge link to the current zeromodels collection slug
626a37d verified
|
Raw
History Blame Contribute Delete
4.27 kB
---
pipeline_tag: image-text-to-text
license: other
license_name: nvidia-license
license_link: https://huggingface.co/nvidia/LocateAnything-3B/blob/main/LICENSE
base_model: nvidia/LocateAnything-3B
library_name: zeromodels
tags:
- keras
- zeromodels
- locateanything
- object-detection
- grounding
- vision
- arxiv:2605.27365
- pytorch
- jax
- tf
---
## ***See [our collection](https://huggingface.co/collections/zeromodels/locateanything-6a8eaf39bae20e37dd3f7fc3) for all versions of LocateAnything.***
# Run LocateAnything with Keras 3: JAX, PyTorch, or TensorFlow
[![GitHub](https://img.shields.io/badge/GitHub-ZeroModels-black?logo=github)](https://github.com/IMvision12/ZeroModels) [![Docs](https://img.shields.io/badge/Docs-LocateAnything-blue)](https://imvision12.github.io/ZeroModels/locateanything/) [![Collection](https://img.shields.io/badge/HF-LocateAnything%20collection-yellow)](https://huggingface.co/collections/zeromodels/locateanything-6a8eaf39bae20e37dd3f7fc3)
# zeromodels/locateanything_3b
Paper: [LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding (arXiv:2605.27365)](https://arxiv.org/abs/2605.27365) · [HF Papers](https://huggingface.co/papers/2605.27365)
LocateAnything is a vision-language grounding model for **detection, referring, pointing, layout, GUI/text grounding, and OCR**. Build the instruction with `locate_prompt(task, text)`, then parse boxes / points / grounding from the generated token ids.
For more details on the model, please go to the upstream [model card](https://huggingface.co/nvidia/LocateAnything-3B).
Pure-**Keras 3** conversion of [`nvidia/LocateAnything-3B`](https://huggingface.co/nvidia/LocateAnything-3B) for [zeromodels](https://github.com/IMvision12/ZeroModels). One implementation runs unmodified on **TensorFlow / Torch / JAX**.
This is a **grounding VLM** checkpoint (`LocateAnythingConditionalGenerate`). Prefer `load_dtype="bfloat16"`.
## ✨ Quick start (open-vocabulary detection)
```python
import os
os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow"
import keras
import numpy as np
from PIL import Image
from zeromodels.models.locateanything import (
LocateAnythingConditionalGenerate,
LocateAnythingProcessor,
locate_prompt,
)
model = LocateAnythingConditionalGenerate.from_weights(
"zeromodels/locateanything_3b", load_dtype="bfloat16"
)
processor = LocateAnythingProcessor.from_weights("zeromodels/locateanything_3b")
image = Image.open("your_image.jpg").convert("RGB")
# Tasks: detection | referring | phrase_grounding | pointing |
# layout | text_grounding | OCR
prompt = locate_prompt("detection", "zebra")
inputs = processor(
conversation=[
{
"role": "user",
"content": [
{"type": "image", "image": image},
{"type": "text", "text": prompt},
],
}
]
)
out = model.generate(
**inputs, max_new_tokens=192, tokenizer=processor.tokenizer
)
ids = np.asarray(keras.ops.convert_to_numpy(out))[0].tolist()
boxes = processor.tokenizer.parse_boxes(ids) # [0, 1000] grid
print(len(boxes), boxes[:2])
```
Load any LocateAnything variant the same way with `from_weights("zeromodels/<variant>")`:
| Variant | Hub |
|---|---|
| `locateanything_3b` | [`zeromodels/locateanything_3b`](https://huggingface.co/zeromodels/locateanything_3b) |
## Tips
- Set `KERAS_BACKEND` **before** importing Keras / zeromodels.
- Use `parse_boxes` for detection, `parse_points` for pointing, `parse_grounding` for referring / layout / text / OCR.
- Boxes and points are on a `[0, 1000]` grid; scale to pixels yourself.
- See [LocateAnything docs](https://imvision12.github.io/ZeroModels/locateanything/) and [Loading Weights](https://imvision12.github.io/ZeroModels/loading_weights/).
- Community / upstream safetensors still work via the `hf:` prefix, e.g. `LocateAnythingConditionalGenerate.from_weights("hf:nvidia/LocateAnything-3B")`.
## Special Thanks
A huge thank you to the NVIDIA LocateAnything authors for creating and releasing these models.
License: [NVIDIA License](https://huggingface.co/nvidia/LocateAnything-3B/blob/main/LICENSE) (non-commercial / research). See the upstream card for component licenses (Qwen2.5, MoonViT, etc.).