File size: 9,896 Bytes
bfc210e 4db1886 bfc210e e173dbc bfc210e ba98af4 bfc210e 4db1886 6915460 ba98af4 6915460 ba98af4 6915460 bfc210e f062949 bfc210e 4326f2e bfc210e 4326f2e bfc210e 4326f2e bfc210e 4db1886 bfc210e 4326f2e bfc210e 4326f2e 8dca2bd bfc210e 4326f2e bfc210e 4326f2e bfc210e 4326f2e bfc210e 4326f2e 32f0301 4326f2e bfc210e 4db1886 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 | ---
license: other
license_name: circlestone-labs-non-commercial-license
license_link: https://huggingface.co/circlestone-labs/Anima/blob/main/LICENSE.md
base_model: circlestone-labs/Anima
pipeline_tag: text-to-image
tags:
- anime
- lokr
- finetune
- Anima
---
# Anima-Telescopa
ALL PREVIEW IMAGES HAVE COMFYUI METADATA.
<p align="center">
<img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/hero.png" width="80%" />
</p>
<table>
<tr>
<td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample1.png" width="100%"/></td>
<td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample2.png" width="100%"/></td>
<td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample3.png" width="100%"/></td>
<td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample4.png" width="100%"/></td>
</tr>
<tr>
<td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample5.png" width="100%"/></td>
<td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample6.png" width="100%"/></td>
<td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample7.png" width="100%"/></td>
<td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample8.png" width="100%"/></td>
</tr>
</table>
---
## I. Introduction
A **full matrix LoKr** of [Anima Base v1.0](https://huggingface.co/circlestone-labs/Anima/blob/main/split_files/diffusion_models/anima-base-v1.0.safetensors) by [CircleStone Labs](https://huggingface.co/circlestone-labs), trained using the [LyCoRIS](https://github.com/kohakublueleaf/lycoris) full matrix LoKr method ([paper](https://arxiv.org/abs/2309.14859)).
This was my first attempt at fine-tuning an image model on a dataset above 10k samples. Alongside general practice, the run was also used to:
- Test whether a carefully hand-curated aesthetic dataset could improve Anima's background generation quality
- Compare full matrix LoKr training against a true full finetune on the same settings
Compared to other Anima fine-tunes such as Tdrussel's Aes B, Motimalu's KirazuriV4, duongve's AnimaYume, and the Silvermoon mixes, Telescopa is perhaps a lil worse on fine detail, stability, and knowledge/style retention β but it shows slightly stronger background composition and detail than the base model and most contemporaries. For a first larger-scale training run, I'm happy with the result.
This model is produced independently, as a hobbyist project, with no external funding.
This repo contain 3 model variant, the ideal epoch 10 LoKR, a earlier and more unstable epoch 3 that also doesnt have as much style bias or potential knowledge forgetting, and a failed epoch 10 full finetune.
---
## II. Model Details
| | |
| ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| **Base model** | [circlestone-labs/Anima](https://huggingface.co/circlestone-labs/Anima) (v1.0) |
| **Method** | Full Matrix LoKr ([LyCoRIS](https://github.com/kohakublueleaf/lycoris)) |
| **Trainer** | [Bluvoll diffusion-pipe fork](https://github.com/bluvoll/diffusion-pipe) (originally by Anima's creator, Tdrussel) |
| **Hardware** | 2x NVIDIA RTX A4000 16GB, courtesy of [Astromahdi's gpu.garden](https://gpu.garden/) |
| **Total training time** | ~3 days (~68 hours) |
| **Total samples seen (unbatched steps)** | ~100,000 |
| **Training resolutions** | 1024Β², 1280Β² |
### Training run (initial 3-epoch run β extended 7-epoch run)
Full config: [TelescopaLOKR.toml](https://github.com/RicemanT/model-training-configs/blob/main/diffusion-pipe/Anima/configs/TelescopaLOKR.toml)
- **Samples seen:** ~30,000 + ~70,000
- **Learning rate:** 3e-6
- **LR scheduler:** Constant with warmup
- **LLM adaptor LR:** Disabled
- **Precision:** Full BF16
- **Optimizer:** AdamW8bit with Kahan summation (to offset BF16 precision roundoff)
- **Weight decay:** 0.01
- **Timestep sampling:** Logit-Normal, shift 4, sigmoid scale 1.3
### LoKr configuration
- Dim / rank: 16
- Alpha: 16
- Dtype: bfloat16
- Dropout: 0.0
- Factor: 2 (~800MB file size)
### Additional training features
- Tag dropout: 10%
- Caption dropout: 5%
- Mixed captions at a 25/25/25/25 ratio
- Tag shuffle
---
## III. Recommended Settings
These are the settings used for the sample images above (ComfyUI, `ModelSamplingAuraFlow` node):
- **Shift:** 5.0
- **Steps:** 28
- **CFG:** 4.0
- **Sampler:** `er_sde`
- **Scheduler:** simple
These are just my usual settings β feel free to experiment and really go buck wild with it, 'euler a' at same steps and `sa_solver_pece` at a lower 15β20 steps also recommended.
---
## IV. Dataset
[Full dataset](https://huggingface.co/datasets/RicemanT/Anime-Background-Finetuning-V1.1)
Sourced from Danbooru (curated by artist background quality) and [Akanyan's personal anime screencap collection scattered throughout his reddit account](https://www.reddit.com/r/anime/comments/4lhd5l/complete_monogatari_series_background_art/). Dataset tooling lives in the [`utils` folder](https://github.com/RicemanT/model-training-configs/tree/main/diffusion-pipe/Anima/utils) of the training repo.
**Curation:** ~43k Danbooru images plus ~7-8k screencaps (50k+ raw) were hand-reviewed for background quality β about two weeks of manual review, landing on a final **10,143-image** dataset.
**Processing:** PNGs converted to lossless WebP, resolutions capped at 2000px (cv2 `INTER_AREA`), deduplicated, and filtered for unusual aspect ratios. A YOLO β IOPaint watermark/logo removal pass was planned but skipped due to time constraints and unreliable open-source detection models.
**Tagging:** [convnextv2_huge.dbv4-full](https://huggingface.co/animetimm/convnextv2_huge.dbv4-full) by animetimm/DeepGHS, run on a free Colab T4.
**Captioning:** [Qwen 3.6 27B FP8](https://huggingface.co/Qwen/Qwen3.6-27B-FP8) via VLLM on an H100 (~3h for ~9k images, non-thinking mode), with the remaining ~1k finished on Qwen 3.6 27B Q6_K via llama.cpp on 2x A4000. Captioning [system prompt](https://github.com/RicemanT/model-training-configs/blob/main/diffusion-pipe/Anima/utils/Anima-Telescopa-captioning-sysprompt.md).
---
## V. Failed Full-Finetune Comparison
A full finetune run with matching settings was also attempted on 4x A4000, finishing in 1.5 days. Training was stable, but the resulting model lost a significant amount of character and artist-style knowledge, making it uncompetitive as an Anima tune β likely because a learning rate of 3e-6 was still too high for full finetuning at the batch size available. That config is included in the training repo for reference. Model file is included here.
---
## VI. Notes from the Training Diary
- **Always back up your dataset before deleting anything locally.** During transfer to the HF dataset repo, an upload silently failed to commit, and ~4k already-"transferred" screencap images (much of it curated Monogatari content) were lost after being deleted locally. The dataset was rebuilt with a more diverse series selection β which ended up improving overall balance.
- WebP isn't always smaller than JPEG β worth checking before batch-converting a dataset.
Full training diary (unedited, written without LLM assistance) (yes I used Claude to help partially write this model page because lazy, sorry lol): [TelescopaLOKR-diary.md](https://github.com/RicemanT/model-training-configs/blob/main/diffusion-pipe/Anima/notes/TelescopaLOKR-diary.md)
---
## VII. License
This model is a Derivative of Anima and is distributed under the same **[CircleStone Labs Non-Commercial License v1.1](https://huggingface.co/circlestone-labs/Anima/blob/main/LICENSE.md)** as the base model, with no additional restrictions. Non-commercial use only β see the license for full terms.
---
## VIII. Acknowledgments
Beeg thanks to:
- [CircleStone Labs](https://huggingface.co/circlestone-labs) β for Anima and the original diffusion-pipe trainer
- [Motimalu](https://civitai.red/user/motimalu) β Inspiration behind finetuning practices and configs, his diary was such a huge help for a completely beginner to get an idea behind how finetuning a model looks like
- [Bluvoll](https://github.com/bluvoll/diffusion-pipe) β diffusion-pipe fork used for training, and general training advice
- [Anzhc](https://huggingface.co/Anzhc) β general training advice
- [Akanyan](https://www.reddit.com/user/akanyan/) β for the goated collection of anime screencap from 9 years ago
- [Astromahdi](https://gpu.garden/) β compute via gpu.garden
- [KohakuBlueleaf](https://github.com/kohakublueleaf/lycoris) β LyCoRIS / LoKr method
- [animetimm/DeepGHS](https://huggingface.co/animetimm/convnextv2_huge.dbv4-full) β tagging model
- Format inspired by [Motimalu's Kirazuri diary](https://github.com/motimalu/diffusion-training-configs/blob/main/diffusion-pipe/anima/notes/kirazuri4.0-notes%20.md) |