File size: 7,705 Bytes
aa3d7f9 0a6a675 aa3d7f9 0a6a675 aa3d7f9 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 | ---
library_name: pytorch
pipeline_tag: feature-extraction
license: other
license_name: circlestone-labs-non-commercial-license
license_link: LICENSE.md
base_model:
- circlestone-labs/Anima
- timm/vit_so400m_patch16_siglip_gap_512.v2_webli
tags:
- anime
- style-embedding
- image-retrieval
- reference-image
- siglip2
- anima
- pytorch
---
# Anima Style Embedding v1.0
**A multi-reference style encoder for person-centric anime and illustration images.**
Anima Style Embedding converts one to eight reference images into a normalized 512-dimensional style vector. It combines multi-level SigLIP2 features from the full image and face with internal Anima features, then grounds the result across the complete reference set. The embedding is designed for open-set style retrieval and reference-driven style transfer systems rather than a fixed bank of known artists.
| Model size | References | Embedding | Modalities |
|---:|---:|---:|---|
| 17.49M parameters | 1–8 images | 512-D, L2-normalized | Full image + face + Anima internals |

## Highlights
- **Open-set style representation:** unseen artists can be embedded directly from reference images.
- **Native multi-reference grounding:** Reference Consensus Re-reading lets each image be interpreted in the context of the complete 1–8 image set.
- **Full-image and face awareness:** the encoder combines global composition and rendering cues with face-specific style information.
- **Anima-aware features:** three internal Anima blocks complement SigLIP2 visual features with representations aligned to the target generator.
- **Strong retrieval gains:** unseen synthetic-style MRR improves from 32.0% with the frozen SigLIP2 Full+Face baseline to 60.8% with the learned encoder on the 4,500-way validation gallery.
## Performance
Retrieval uses cosine similarity between normalized query embeddings and five-image style prototypes.
### 4,500-way validation — unseen identities
The unseen subset contains 500 identities retrieved against a gallery of 4,500 styles.
| Source | Top-1 | Top-5 | MRR |
|---|---:|---:|---:|
| Synthetic | 51.8% | 71.7% | 60.8% |
| Human | 56.4% | 75.3% | 65.0% |
### 500-way sealed test — unseen identities
| Source | Top-1 | Top-5 | MRR |
|---|---:|---:|---:|
| Synthetic | 73.8% | 90.3% | 81.1% |
| Human | 76.1% | 92.0% | 83.0% |
The sealed test uses a smaller gallery than validation, so its absolute scores should be read independently.

### Reference count
More references consistently improve human-style grounding. Synthetic K=5 uses the locked controlled/matched/cross-content protocol and is shown separately from the generic K diagnostics.

## Architecture
The encoder projects full-image, face, and Anima tokens into a shared 512-wide space and reads them with four eight-head mixer blocks. Functional Style Factorization introduces four auxiliary style queries during training. Reference Consensus Re-reading computes a consensus across the reference set, re-reads every image under that context, and forms an agreement-weighted prototype.
| Input | Shape per batch |
|---|---|
| Full-image SigLIP2 features | `[B, R, 30, 1152]` |
| Face SigLIP2 features | `[B, R, 30, 1152]` |
| Face validity mask | `[B, R]` |
| Anima internal features | `[B, R, 3, 4096]` |
| Reference mask | `[B, R]` |
SigLIP2 features use [`timm/vit_so400m_patch16_siglip_gap_512.v2_webli`](https://huggingface.co/timm/vit_so400m_patch16_siglip_gap_512.v2_webli) at 512×512 and three backbone depths. Anima features use 768×768 inputs and blocks 8, 18, and 26 at sigma 0.1.
## Quick start
```python
import json
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from modeling_anima_style_embedding import UnifiedStyleEncoder
repo_id = "Baragi-AI/Anima-Style-Embedding"
config = json.load(open(hf_hub_download(repo_id, "config.json"), encoding="utf-8"))
cfg = config["architecture"]
encoder = UnifiedStyleEncoder(
width=cfg["width"],
blocks=cfg["blocks"],
heads=cfg["heads"],
embedding_dim=cfg["embedding_dim"],
dropout=cfg["dropout"],
use_anima=True,
reference_grounding=cfg["reference_grounding"],
functional_factors=cfg["functional_factors"],
)
encoder.load_state_dict(
load_file(hf_hub_download(repo_id, "model.safetensors")),
strict=True,
)
encoder.eval()
```
```python
with torch.inference_mode():
style_embedding, per_image_embeddings = encoder(
full_features,
reference_mask,
face_features=face_features,
face_mask=face_mask,
anima_features=anima_features,
)
assert style_embedding.shape == (full_features.shape[0], 512)
```
`modeling_anima_style_embedding.py` contains the matching encoder definition. The checkpoint expects precomputed SigLIP2 and Anima features; RGB and Anima feature extraction are supplied by downstream integrations.
## Training
The model was trained on 320,000 person-centric style records: 160,000 Anima-generated images and 160,000 human-created illustrations across 8,000 training identities. Synthetic and human examples are mixed within each batch. Training uses four RTX 5090 GPUs, BF16 DistributedDataParallel, 16-way episodes, two queries per prototype, variable 1–8 image support, global negatives, and weak hard-negative sampling. The 4,500-step checkpoint produced the best full-validation score.

The training record is available in [Weights & Biases](https://wandb.ai/1wndrla17-kyung-hee-university/anima-style-embedding/runs/5c9ae843a14f46529fe9aaa604a78d3b). Machine-readable evaluation results are included under `metrics/`.
## Applications
- reference-driven style transfer for Anima;
- artist-style retrieval and clustering;
- aggregation of multiple references into a single style prototype;
- open-set style recognition for previously unseen artists;
- similarity scoring and reference-set curation.
## Files
| File | Contents |
|---|---|
| `model.safetensors` | Encoder weights |
| `config.json` | Architecture and feature contract |
| `modeling_anima_style_embedding.py` | PyTorch model definition |
| `metrics/` | Training curves and retrieval results |
| `SHA256SUMS.json` | Release checksums |
## Limitations
- The model is optimized for person-centric anime and illustration images; other domains are not evaluated.
- Retrieval metrics measure representation quality, not generated-image fidelity.
- Style and artist-associated content preferences are not completely disentangled.
- Anima features must follow the documented model, block, resolution, and sigma configuration.
## License
The model weights are distributed under the [CircleStone Labs Non-Commercial License v1.2](LICENSE.md). Commercial or production use requires appropriate authorization from CircleStone Labs.
The original source code supplied with this release is available under the [Apache License 2.0](CODE_LICENSE). Third-party components remain subject to their respective licenses; attribution details are listed in [NOTICE](NOTICE) and [THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md).
Built on NVIDIA Cosmos.
## Citation
```bibtex
@misc{baragi2026animastyleembedding,
title = {Anima Style Embedding v1.0},
author = {{Baragi AI}},
year = {2026},
howpublished = {Hugging Face model repository},
url = {https://huggingface.co/Baragi-AI/Anima-Style-Embedding}
}
```
|