Text-to-Image
anima
illustration
diffusion
anime

Status: Training in progress.

Next step: Pretraining on a general 10M samples on various concepts to improve prompt understanding, while also expand significantly on the main anime/illustration dataset (I will update the knowledge as latest as possible)

If you'd like to support me or to support the training progress :

Ko-fi PayPal

Vast.ai: thangquay347@gmail.com

Every bit of support helps expand the model's scope and capability even further!

You will need to install ComfyUI-Anima-2.9B to the custom node folder. Plug and play, there is no workflow node needed. Sometimes may not work with other custom nodes

collage

Overview

Anima-2.9B is a fine-tune and layer-expansion of circlestone-labs/Anima. The base Anima model targets anime, illustration, and non-photorealistic art; this release continues training on that foundation with an expanded architecture. The model is trained on an additional 1.7M anime/illustration samples, with knowledge cutoff in July 2026, making Anima-2.9B one of the most up-to-date anime/illustration model at release.

Versions

  • Anima-2.9B-preview-v1: inital release

Training/Dataset

  • Trained using Muon optimizer on a 8x 5080s cluster, with earlier steps trained locally on my PC

  • As of preview v1, only new layers have been trained, with roughly 70% of the compute spent on 1024px

  • Knowledge cutoff in July 2026, training data included both new and old samples prior to September 2025

  • Mixed captioning, including both tags and natural languages, using a mix of Gemini 3.1 Flash-Lite, Gemini 3.5 Flash-Lite, and Claude Sonnet 5

  • NO scoring

Architecture

  • Transformer depth expansion: expanded from 28 transformers layers to 40, growing the model to ~2.9B parameters. Each new layer is added by deep-copying its neighboring layer's weights, using interleaved insertion with zeroed-out output projections, making the new model functionally identical to Anima-base at initialization.

Prompting tips :

Follow Anima prompting tips: quality tags, year/period tags, @artist tags, character count (1girl, 1boy), character tags (follow Danbooru and Gelbooru tags), series/copyrights, base appearance.

Character name/tags should be follow with series/copyrights tags or else the model might confuse.

For multi-character images, attribute the character and names with their respective tags/appearance.

The model does improve the base art style slightly, but I'd still recommend using artist tags.

The dataset does not include score in its captions, however, you can still use them.

(IMPORTANT) THE MORE DETAILED THE PROMPT, THE BETTER, short prompt will often generate a bland simple background, and may not able to produce the desire results

Generation (Recommendation)

  • Sampler: Euler/Res-multistep/Er-sde
  • Scheduler: sgm-uniform/beta/beta57/linear-quadratic
  • Resolution: 812x1216, 1152x1536, 1536x1536 (iffy)
  • Steps: 28-50
  • CFG: 3.5-5

My personal usage are euler + sgm-uniform, which has a good balance between composition and fine details. Additionally res-multistep + linear-quadratic spend more time at high noise steps, which does lead to visibly better composition. My recommendation for the highest quality is 50 steps, there are some images where 3.5 CFG do better than 5 CFG and vice versa. Experiment yourself!

To-do:

Lora Training will be supported via my Anima Standalone Trainer in a few days.

License

Model weights are released under the CircleStone Labs Non-Commercial License, falling under derivative model category.

Acknowledgements

Built on nvidia/Cosmos-Predict2-2B-Text2Image and circlestone-labs/Anima.

LLaMA Pro: Progressive LLaMA with Block Expansion.

Training infrastructure built on sd-scripts.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Gazingstars123/Anima-2.9B

Finetuned
(82)
this model

Space using Gazingstars123/Anima-2.9B 1

Paper for Gazingstars123/Anima-2.9B