Anima Tagger

Multi-label anime image tagger. Given an image it emits a booru-style caption in exactly the format the Anima diffusion model was trained on:

rating, count, characters, copyrights, general tags

It is the tagger behind anime_tools (dataset autotagging, position-clause captions, the curation GUI) and the anima_lora training pipeline, and works standalone.

What is in this repo

Only our half of the checkpoint β€” a few MB. The backbone weights are fetched from their own repo at load time (see the license note below).

File Role
config.json Backend descriptor: backbone repo / arch / input size
vocab.json The 2,532-tag Anima vocabulary with categories, emit order and 60 tag groups
rules.yaml Caption normalization: replacements, tag aliases, always-remove, clothing dedup
groups.yaml Softmax / sentinel tag groups (eye color, hair color, …)
thresholds.safetensors Per-tag inference thresholds
sidecar.safetensors + sidecar.json The linear sidecar head and its row map
sidecar_metrics.json Held-out metrics of the shipped head

Architecture

  1. Backbone β€” animetimm/caformer_b36.dbv4-full (134 M params, 384 px, 12,476 danbooru tags, timm). Supplies rating, characters and general tags.
  2. Vocab projection. dbv4's snake_case names are joined onto the Anima vocab β€” 2,182 of 2,532 tags match, with rules.yaml renames recovered on the way. Per-tag thresholds come from the dbv4 card's best_threshold (median 0.32). Of the 350 tags dbv4 cannot express, 223 are covered by the sidecar and 127 are hard-disabled with a never-fire threshold.
  3. Sidecar head (~3 MB, trained by us). A linear head on the backbone's 3,072-d MLP hidden feature that emits only what dbv4 has no labels for: 118 copyright tags, 23 characters dbv4 lacks, 82 renamed generals, plus an 8-way people-count bucket.

Everything downstream of the score vector is shared post-processing: threshold gating, softmax groups ("at most one" eye color / hair color β€” the winner must still clear its own threshold), count-tag dedupe and character cap, a character-confidence floor with original fallback, top-1 copyright collapse, and slot ordering with underscores turned into spaces.

Two kinds of tag are left out of the head on purpose:

  • Artist OCs β€” a character whose trailing qualifier is an artist handle (shiro (mignon)). Such a name means nothing outside the dataset it came from, and a head trained on a few dozen positives fires it on any look-alike. The 15 rows left out are listed in sidecar.json["dropped_artist_oc"]; franchise characters dbv4 lacks stay.
  • Retired booru names (silver hair, light brown hair, black footwear, …). Their positives look exactly like the live tag's, so a head over them measured macro-F1 0.18. rules.yaml aliases: folds each onto its live name instead (silver hair β†’ grey hair).

Backbone weights are gated GPL-3.0 and are never bundled here

animetimm/*.dbv4-full is GPL-3.0-licensed and access-gated. Nothing in this repo vendors, mirrors or redistributes it: the loader calls hf_hub_download under your token, and accepting the upstream repo's terms is what grants the download. Run hf auth login and click through once on the backbone's page before first use. The MIT license on this repo covers our part only β€” vocab, rules, groups, thresholds and the sidecar head.

Numbers

Threshold-free mAP on the in-house 791-image held-out split, intersection vocab, against the previous in-house PE-backbone tagger (2026-08-26):

caformer_b36 dbv4 previous in-house head
mAP (all tags) 0.633 0.297
Characters 0.964 0.619
Tail tags (freq < 200) 0.630 0.285
Rating (4-way acc) 0.905 0.833

The gap is the backbone β€” a frozen-feature linear probe against a model fine-tuned on all of danbooru β€” not the head.

Shipped sidecar head on the same split (2026-09-13): copyright macro-F1 0.81, characters 0.86, renamed generals 0.42. For people count the caption count-tag rule is authoritative (0.943) over the sidecar's own softmax (0.927), which is exposed as people_count_scores only.

Usage

Install anime_tools β€” one line, no checkout β€” or uv sync in an anima_lora checkout, which depends on it. Both halves of the checkpoint (this repo and the gated backbone) auto-download on first use.

curl -fsSL https://github.com/sorryhyun/anime_tools/releases/latest/download/install.sh | sh
from PIL import Image
from anime_tools.tagger import AnimaTagger   # anima_lora.captioning.AnimaTagger is the same class

tagger = AnimaTagger()   # defaults to models/captioners/anima-tagger-dbv4/
print(tagger.predict_caption(Image.open("image.png")))
# "nsfw, 1girl, blue archive, animal ears, black hair, blush, fox girl, halo, long hair, ..."

predict() returns the structured form instead β€” rating / rating_scores, people_count (+ people_count_source), scores and thresholds over the whole vocab, kept (the emitted positives) and groups ({group_name: winner_or_None}). predict_batch / predict_caption_batch run one backbone forward over a list. The constructor takes device / dtype (defaults: CUDA if available, bf16; float32 on CPU and MPS) and character_floor.

From the command line:

python -m anime_tools.tagger.cli.main --mode predict --image image.png --show_scores

Batch-tagging a dataset is the autotag stage (make caption-autotag in anima_lora, or the GUI's Autotag button). In ComfyUI, the anima_tagger nodes shipped in anime_tools/comfyui/ give AnimaTaggerLoader β†’ AnimaTaggerCaption β†’ STRING.

Rebuilding the checkpoint

Every step is a python -m module in anime_tools.tagger.cli; the full recipe is in docs/anima_tagger.md.

python -m anime_tools.tagger.cli.main --mode build_vocab --min_freq 5   # vocab + split + groups
python -m anime_tools.tagger.cli.build_dbv4_ckpt                        # checkpoint dir + card thresholds
python -m anime_tools.tagger.cli.train_sidecar --ckpt_dir models/captioners/anima-tagger-dbv4

train_sidecar caches one backbone forward per image, trains the head (BCE + CE, best epoch by val mAP), F1-calibrates the sidecar rows' thresholds and writes the head. --keep_artist_oc trains the OC rows too; --categories picks which vocab categories the head covers.

Limitations

  • No @artist tags. dbv4 has no artist category and the sidecar deliberately does not model one, so the @artists caption slot is always empty. The 92 artist tags in the vocab are hard-disabled.
  • No artist OCs, by the rule above β€” a dataset's recurring original character comes out as original.
  • Rating bands are danbooru's, mapped onto Anima's four (general β†’ safe, questionable β†’ nsfw).
  • Thresholds for backbone tags are the upstream card's calibrated values, not re-fit on our split; re-fitting on 791 images overfits.
  • The vocabulary, normalization rules and emit order are tuned for the Anima caption distribution. For a general-purpose booru tagger use the upstream model directly.
Downloads last month
278
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for sorryhyun/anima-tagger