IcecreamDetect / README.md
ibsocr1's picture
Upload 6 files
dbb4a08 verified
|
Raw
History Blame Contribute Delete
5.4 kB
---
title: Ice Cream Dataset + Counter
emoji: 🍦
colorFrom: yellow
colorTo: red
sdk: gradio
sdk_version: 6.5.1
app_file: app.py
pinned: false
---
# 🍦 Ice Cream Dataset + Counter — Gradio Space
This is the **Gradio + ZeroGPU-compatible** version of the uploaded Ice Cream Counter project. It does **not** use Docker, FastAPI, Uvicorn, or a custom HTML frontend.
## What it does
1. Upload training freezer photos.
2. Define your product classes.
3. Annotate each ice cream with bounding boxes.
4. Train an RT-DETR object detector.
5. Upload one new freezer image.
6. Get the total count, per-product counts, confidence scores, and an annotated result image.
The model is still:
`PekingU/rtdetr_r50vd`
No YOLO is used.
## Create the Hugging Face Space
Create a new Space and choose:
- **SDK:** Gradio
- **Hardware:** ZeroGPU or a dedicated GPU is recommended for training
The app uses `@spaces.GPU` for training and counting, so it also boots correctly when the Space hardware is **ZeroGPU**.
Then upload these files/folders:
```text
app.py
requirements.txt
README.md
training/
```
You do **not** need:
```text
Dockerfile
FastAPI
Uvicorn
static/index.html
app/main.py
```
## Persistent dataset and model
The app stores:
```text
images/
dataset.json
model/
generated_dataset/
```
When `/data` is available and writable, the app automatically uses:
```text
/data/icecream_counter/
```
You can also explicitly set:
```text
DATA_DIR=/data/icecream_counter
```
in the Space variables.
### Important Hugging Face storage point
A normal Space filesystem is not permanent storage across every rebuild/restart. If you need the dataset and trained model to survive Space restarts/rebuilds, attach **persistent storage** to the Space or move the data/model to an external persistent service.
The Gradio conversion itself does not change this storage rule.
## Annotation workflow
In the **Annotate** tab the actual uploaded training image is displayed directly using Gradio's Image component. This avoids browser canvas/JavaScript issues that can make the preview appear black.
1. Select a training image.
2. The real image appears in the preview.
3. Read the image dimensions shown below it.
4. Select the product class.
5. Enter the bounding box in original-image pixel coordinates:
- X (left)
- Y (top)
- Width
- Height
6. Click **Save Box**.
7. Saved boxes are drawn in red on the real image.
8. Repeat for every ice cream.
9. Use **Delete Box** or **Clear All Boxes** when needed.
This version prioritizes a reliable visible image over the previous JavaScript canvas approach.
## Training
The app creates an 80/20 COCO train/validation split from the annotated images and starts the existing RT-DETR training script.
Default settings:
```text
Epochs: 30
Batch size: 2
Learning rate: 1e-5
```
For an initial test on a small dataset, use fewer epochs such as 2–5. Once everything works, increase the epochs.
A GPU Space is strongly recommended.
## Training fix
This release includes a CUDA-device fix for RT-DETR's contrastive-denoising training path. On some
Transformers/PyTorch combinations, the denoising class-index tensor can remain on CPU while the
RT-DETR class embedding is on CUDA, producing:
`RuntimeError: Expected all tensors to be on the same device ... cpu ... cuda:0`
The training script now moves nested target tensors explicitly and patches the RT-DETR denoising
helper so its target tensors follow the class-embedding device. The `num_labels=10` vs. checkpoint
`80` message is expected when fine-tuning the COCO-pretrained checkpoint for 10 custom classes;
`ignore_mismatched_sizes=True` intentionally reinitializes the classification heads.
## Counting
After training, open the **Count** tab and upload one image.
The result contains:
```json
{
"total": 31,
"counts": {
"cone": 6,
"correto": 7,
"cornetto": 12,
"magnum": 6
}
}
```
The result image also shows the detected bounding boxes and confidence values.
## Default classes
```text
cornetto
magnum
correto
cone
cup
sandwich
stick
other
```
You can change them from the Dataset tab.
## Local test
```bash
pip install -r requirements.txt
python app.py
```
Then open:
```text
http://localhost:7860
```
## Recommended Space setup
For the first deployment:
1. Create the Space as **Gradio**.
2. Upload `app.py`, `requirements.txt`, `README.md`, and `training/`.
3. Wait for dependencies to install.
4. Open the Dataset tab.
5. Upload 2+ training images.
6. Annotate them.
7. Start with 2–5 epochs to verify training.
8. After the model finishes, test the Count tab.
9. For serious training, attach a GPU and persistent storage.
### RT-DETR training compatibility
The trainer disables RT-DETR contrastive denoising by default because some Transformers releases can create CPU class-index tensors while the embedding is on CUDA. The detector's normal supervised detection loss remains enabled. A device-safe denoising patch is also included for future re-enablement.
## Fixed22 training patch
This version fixes the RT-DETR CPU/CUDA denoising crash by wrapping the denoising class embedding as a real `torch.nn.Module` and moving its index tensor to the embedding weight device before `nn.Embedding` is called. A plain Python function wrapper is intentionally not used because Transformers expects the embedding to remain a module.