| --- |
| license: other |
| license_name: circlestone-labs-non-commercial-license |
| license_link: https://huggingface.co/circlestone-labs/Anima/blob/main/LICENSE.md |
| base_model: circlestone-labs/Anima |
| pipeline_tag: text-to-image |
| tags: |
| - anime |
| - lokr |
| - finetune |
| - Anima |
| --- |
| |
| # Anima-Telescopa |
| ALL PREVIEW IMAGES HAVE COMFYUI METADATA. |
| <p align="center"> |
| <img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/hero.png" width="80%" /> |
| </p> |
|
|
| <table> |
| <tr> |
| <td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample1.png" width="100%"/></td> |
| <td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample2.png" width="100%"/></td> |
| <td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample3.png" width="100%"/></td> |
| <td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample4.png" width="100%"/></td> |
| </tr> |
| <tr> |
| <td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample5.png" width="100%"/></td> |
| <td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample6.png" width="100%"/></td> |
| <td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample7.png" width="100%"/></td> |
| <td width="25%"><img src="https://huggingface.co/RicemanT/Anima-Telescopa/resolve/main/images/sample8.png" width="100%"/></td> |
| </tr> |
| </table> |
|
|
| --- |
|
|
| ## I. Introduction |
|
|
| A **full matrix LoKr** of [Anima Base v1.0](https://huggingface.co/circlestone-labs/Anima/blob/main/split_files/diffusion_models/anima-base-v1.0.safetensors) by [CircleStone Labs](https://huggingface.co/circlestone-labs), trained using the [LyCoRIS](https://github.com/kohakublueleaf/lycoris) full matrix LoKr method ([paper](https://arxiv.org/abs/2309.14859)). |
|
|
| This was my first attempt at fine-tuning an image model on a dataset above 10k samples. Alongside general practice, the run was also used to: |
|
|
| - Test whether a carefully hand-curated aesthetic dataset could improve Anima's background generation quality |
| - Compare full matrix LoKr training against a true full finetune on the same settings |
|
|
| Compared to other Anima fine-tunes such as Tdrussel's Aes B, Motimalu's KirazuriV4, duongve's AnimaYume, and the Silvermoon mixes, Telescopa is perhaps a lil worse on fine detail, stability, and knowledge/style retention β but it shows slightly stronger background composition and detail than the base model and most contemporaries. For a first larger-scale training run, I'm happy with the result. |
|
|
| This model is produced independently, as a hobbyist project, with no external funding. |
| This repo contain 3 model variant, the ideal epoch 10 LoKR, a earlier and more unstable epoch 3 that also doesnt have as much style bias or potential knowledge forgetting, and a failed epoch 10 full finetune. |
|
|
| --- |
|
|
| ## II. Model Details |
|
|
| | | | |
| | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | |
| | **Base model** | [circlestone-labs/Anima](https://huggingface.co/circlestone-labs/Anima) (v1.0) | |
| | **Method** | Full Matrix LoKr ([LyCoRIS](https://github.com/kohakublueleaf/lycoris)) | |
| | **Trainer** | [Bluvoll diffusion-pipe fork](https://github.com/bluvoll/diffusion-pipe) (originally by Anima's creator, Tdrussel) | |
| | **Hardware** | 2x NVIDIA RTX A4000 16GB, courtesy of [Astromahdi's gpu.garden](https://gpu.garden/) | |
| | **Total training time** | ~3 days (~68 hours) | |
| | **Total samples seen (unbatched steps)** | ~100,000 | |
| | **Training resolutions** | 1024Β², 1280Β² | |
|
|
| ### Training run (initial 3-epoch run β extended 7-epoch run) |
|
|
| Full config: [TelescopaLOKR.toml](https://github.com/RicemanT/model-training-configs/blob/main/diffusion-pipe/Anima/configs/TelescopaLOKR.toml) |
|
|
| - **Samples seen:** ~30,000 + ~70,000 |
| - **Learning rate:** 3e-6 |
| - **LR scheduler:** Constant with warmup |
| - **LLM adaptor LR:** Disabled |
| - **Precision:** Full BF16 |
| - **Optimizer:** AdamW8bit with Kahan summation (to offset BF16 precision roundoff) |
| - **Weight decay:** 0.01 |
| - **Timestep sampling:** Logit-Normal, shift 4, sigmoid scale 1.3 |
|
|
| ### LoKr configuration |
|
|
| - Dim / rank: 16 |
| - Alpha: 16 |
| - Dtype: bfloat16 |
| - Dropout: 0.0 |
| - Factor: 2 (~800MB file size) |
|
|
| ### Additional training features |
|
|
| - Tag dropout: 10% |
| - Caption dropout: 5% |
| - Mixed captions at a 25/25/25/25 ratio |
| - Tag shuffle |
|
|
| --- |
|
|
| ## III. Recommended Settings |
|
|
| These are the settings used for the sample images above (ComfyUI, `ModelSamplingAuraFlow` node): |
|
|
| - **Shift:** 5.0 |
| - **Steps:** 28 |
| - **CFG:** 4.0 |
| - **Sampler:** `er_sde` |
| - **Scheduler:** simple |
|
|
| These are just my usual settings β feel free to experiment and really go buck wild with it, 'euler a' at same steps and `sa_solver_pece` at a lower 15β20 steps also recommended. |
|
|
| --- |
|
|
| ## IV. Dataset |
|
|
| [Full dataset](https://huggingface.co/datasets/RicemanT/Anime-Background-Finetuning-V1.1) |
|
|
| Sourced from Danbooru (curated by artist background quality) and [Akanyan's personal anime screencap collection scattered throughout his reddit account](https://www.reddit.com/r/anime/comments/4lhd5l/complete_monogatari_series_background_art/). Dataset tooling lives in the [`utils` folder](https://github.com/RicemanT/model-training-configs/tree/main/diffusion-pipe/Anima/utils) of the training repo. |
|
|
| **Curation:** ~43k Danbooru images plus ~7-8k screencaps (50k+ raw) were hand-reviewed for background quality β about two weeks of manual review, landing on a final **10,143-image** dataset. |
|
|
| **Processing:** PNGs converted to lossless WebP, resolutions capped at 2000px (cv2 `INTER_AREA`), deduplicated, and filtered for unusual aspect ratios. A YOLO β IOPaint watermark/logo removal pass was planned but skipped due to time constraints and unreliable open-source detection models. |
|
|
| **Tagging:** [convnextv2_huge.dbv4-full](https://huggingface.co/animetimm/convnextv2_huge.dbv4-full) by animetimm/DeepGHS, run on a free Colab T4. |
|
|
| **Captioning:** [Qwen 3.6 27B FP8](https://huggingface.co/Qwen/Qwen3.6-27B-FP8) via VLLM on an H100 (~3h for ~9k images, non-thinking mode), with the remaining ~1k finished on Qwen 3.6 27B Q6_K via llama.cpp on 2x A4000. Captioning [system prompt](https://github.com/RicemanT/model-training-configs/blob/main/diffusion-pipe/Anima/utils/Anima-Telescopa-captioning-sysprompt.md). |
| |
| --- |
| |
| ## V. Failed Full-Finetune Comparison |
| |
| A full finetune run with matching settings was also attempted on 4x A4000, finishing in 1.5 days. Training was stable, but the resulting model lost a significant amount of character and artist-style knowledge, making it uncompetitive as an Anima tune β likely because a learning rate of 3e-6 was still too high for full finetuning at the batch size available. That config is included in the training repo for reference. Model file is included here. |
| |
| --- |
| |
| ## VI. Notes from the Training Diary |
| |
| - **Always back up your dataset before deleting anything locally.** During transfer to the HF dataset repo, an upload silently failed to commit, and ~4k already-"transferred" screencap images (much of it curated Monogatari content) were lost after being deleted locally. The dataset was rebuilt with a more diverse series selection β which ended up improving overall balance. |
| - WebP isn't always smaller than JPEG β worth checking before batch-converting a dataset. |
| |
| Full training diary (unedited, written without LLM assistance) (yes I used Claude to help partially write this model page because lazy, sorry lol): [TelescopaLOKR-diary.md](https://github.com/RicemanT/model-training-configs/blob/main/diffusion-pipe/Anima/notes/TelescopaLOKR-diary.md) |
| |
| --- |
| |
| ## VII. License |
| |
| This model is a Derivative of Anima and is distributed under the same **[CircleStone Labs Non-Commercial License v1.1](https://huggingface.co/circlestone-labs/Anima/blob/main/LICENSE.md)** as the base model, with no additional restrictions. Non-commercial use only β see the license for full terms. |
| |
| --- |
| |
| ## VIII. Acknowledgments |
| |
| Beeg thanks to: |
| |
| - [CircleStone Labs](https://huggingface.co/circlestone-labs) β for Anima and the original diffusion-pipe trainer |
| - [Motimalu](https://civitai.red/user/motimalu) β Inspiration behind finetuning practices and configs, his diary was such a huge help for a completely beginner to get an idea behind how finetuning a model looks like |
| - [Bluvoll](https://github.com/bluvoll/diffusion-pipe) β diffusion-pipe fork used for training, and general training advice |
| - [Anzhc](https://huggingface.co/Anzhc) β general training advice |
| - [Akanyan](https://www.reddit.com/user/akanyan/) β for the goated collection of anime screencap from 9 years ago |
| - [Astromahdi](https://gpu.garden/) β compute via gpu.garden |
| - [KohakuBlueleaf](https://github.com/kohakublueleaf/lycoris) β LyCoRIS / LoKr method |
| - [animetimm/DeepGHS](https://huggingface.co/animetimm/convnextv2_huge.dbv4-full) β tagging model |
| - Format inspired by [Motimalu's Kirazuri diary](https://github.com/motimalu/diffusion-training-configs/blob/main/diffusion-pipe/anima/notes/kirazuri4.0-notes%20.md) |