Repository Info:

This repository contains a mixed precision version of lodestones/taggerine.
The backbone weights are converted to bf16 and the projection head weights are kept in fp32.

Why? – Because the original ReadMe states that the backbone was trained in bf16, so I'm thinking keeping it in fp32 is waste of storage space and bandwidth depending on situation.
There is also a full bf16 and full fp8 version here, but note that scores may be a very tiny bit less accurate.


Original ReadMe content

DINOv3 ViT-H/16+ Booru Tagger

A multi-label image tagger trained on e621 and Danbooru annotations, using a DINOv3 ViT-H/16+ backbone fine-tuned end-to-end with a single linear projection head.

Model Details

Property Value
Backbone facebook/dinov3-vith16plus-pretrain-lvd1689m
Architecture ViT-H/16+ Β· 32 layers Β· hidden dim 1280 Β· 20 heads Β· SwiGLU MLP Β· RoPE Β· 4 register tokens
Head Linear((1 + 4) Γ— 1280 β†’ 74 625) β€” CLS + 4 register tokens concatenated
Vocabulary 74 625 tags (min frequency β‰₯ 50 across training set)
Input resolution Any multiple of 16 px β€” trained at 512 px, generalises to higher resolutions
Input normalisation ImageNet mean/std [0.485, 0.456, 0.406] / [0.229, 0.224, 0.225]
Output Raw logits β€” apply sigmoid for per-tag probabilities
Parameters ~632 M (backbone) + ~480 M (head)

Training

Hyperparameter Value
Training data e621 + Danbooru (parquet)
Batch size 32
Learning rate 1e-6
Warmup steps 50
Loss BCEWithLogitsLoss with per-tag pos_weight = (neg/pos)^(1/T), cap 100
Optimiser AdamW (β₁=0.9, Ξ²β‚‚=0.999, wd=0.01)
Precision bfloat16 (backbone) / float32 (projection + loss)
Hardware 2Γ— GPU, ThreadPoolExecutor + NCCL all-reduce

eval_viz

Usage

1. Install dependencies

pip install -r requirements.txt

Or manually:

pip install torch torchvision safetensors Pillow requests \
            python-multipart fastapi uvicorn jinja2 aiofiles

2. Download model files

huggingface-cli download lodestones/taggerine \
    tagger_proto.safetensors \
    tagger_vocab_with_categories_and_alias_updated.json \
    tagger_ui_server.py \
    inference_tagger_standalone.py \
    --local-dir .

Note: tagger_proto.safetensors is ~5.3 GB. Make sure you have enough disk space.

3. Download the tagger_ui/ templates folder

The server requires the tagger_ui/templates/ directory to be present alongside tagger_ui_server.py:

huggingface-cli download lodestones/taggerine \
    --include "tagger_ui/**" \
    --local-dir .

4. Run the Web UI

python tagger_ui_server.py \
    --checkpoint tagger_proto.safetensors \
    --vocab tagger_vocab_with_categories_and_alias_updated.json \
    --port 7860
# β†’ open http://localhost:7860

CPU-only machine? Add --device cpu (inference will be slower):

python tagger_ui_server.py \
    --checkpoint tagger_proto.safetensors \
    --vocab tagger_vocab_with_categories_and_alias_updated.json \
    --device cpu \
    --port 7860

Standalone CLI inference (no server)

python inference_tagger_standalone.py \
    --checkpoint tagger_proto.safetensors \
    --vocab tagger_vocab_with_categories_and_alias_updated.json \
    --images photo.jpg \
    --topk 30

Files

File Description
tagger_proto.safetensors Model weights (bfloat16)
tagger_vocab_with_categories_and_alias_updated.json {"idx2tag": [...], "tag2category": {...}} β€” 74 625 tags with category metadata
tagger_vocab_with_categories.json Same without alias data
tagger_vocab.json Minimal vocab β€” {"idx2tag": [...]} only
inference_tagger_standalone.py Self-contained CLI inference script (no transformers dep)
tagger_ui_server.py FastAPI + Jinja2 web UI server
requirements.txt Python dependencies

Tag Vocabulary

Tags are sourced from e621 and Danbooru annotations and cover:

  • Subject β€” species, character count, gender (solo, duo, anthro, 1girl, male, …)
  • Body β€” anatomy, fur/scale/skin markings, body parts
  • Action / pose β€” looking at viewer, sitting, …
  • Scene β€” background, lighting, setting
  • Style β€” digital art, hi res, sketch, watercolor, …
  • Rating β€” explicit content tags are included; filter as needed for your use case

Minimum tag frequency threshold: 50 occurrences across the combined dataset.

Limitations

  • Evaluated on booru-style illustrations and furry art; performance on photographic images or other art works to some extend.
  • The vocabulary reflects the biases of e621 and Danbooru annotation practices.

License

Apache 2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for DraconicDragon/taggerine-mixed-bf16

Quantized
(2)
this model