Anima-Lightning

Anima-Lightning is the public release name for a four-step distilled text-to-image model derived from circlestone-labs/Anima.

It keeps the original Anima architecture and Diffusers components while replacing the diffusion transformer weights with four-step distilled weights. The text encoder, tokenizers, text conditioner, VAE, and scheduler follow the upstream Anima pipeline.

This repository contains the public BF16 Diffusers checkpoint. The name “Anima-Lightning” refers to this release and is not the name of an upstream architecture or Diffusers component.

Anima-Lightning is intended for anime, illustration, and other non-photorealistic artwork. It is not optimized for photorealism.

This repository is not an official CircleStone Labs or NVIDIA release.

Architecture

Anima-Lightning uses the upstream Anima modular architecture:

  • Pipeline: AnimaModularPipeline
  • Transformer: CosmosTransformer3DModel
  • Text encoder: Qwen3Model
  • Tokenizer: Qwen2Tokenizer
  • T5 tokenizer: T5Tokenizer
  • Text conditioner: AnimaTextConditioner
  • VAE: AutoencoderKLQwenImage
  • Scheduler: FlowMatchEulerDiscreteScheduler

Required four-step inference

This checkpoint must be sampled with the distilled four-step trajectory:

  • Model evaluations: 4
  • Sigma schedule: [1.0, 0.75, 0.5, 0.25]
  • Terminal sigma: 0.0 (appended by the scheduler)
  • Runtime CFG: 1.0
  • Scheduler: FlowMatchEulerDiscreteScheduler
  • Scheduler shift: 1.0
  • Training timesteps: 1000
  • Recommended resolution: 1024 x 1024
  • Recommended precision: bfloat16
  • Maximum sequence length: 512

This is not the normal 30-50-step upstream Anima runtime. Running this checkpoint with settings such as 40 steps and CFG 5.0 will produce incorrect or degraded output.

distilled_generation_config.json documents this runtime contract, but Diffusers does not automatically apply that file during inference. Pass the four-step sigma schedule and use CFG 1.0 as shown below.

Installation

Anima currently requires a recent Diffusers build containing the Anima modular pipeline and Cosmos text-to-image components.

pip install -U git+https://github.com/huggingface/diffusers.git
pip install -U transformers accelerate safetensors sentencepiece protobuf

A recent CUDA-capable PyTorch installation is recommended.

Generation

import torch
from diffusers import AnimaModularPipeline


repo_id = "aina-tech/Anima-Lightning"

pipe = AnimaModularPipeline.from_pretrained(
    repo_id,
    torch_dtype=torch.bfloat16,
)

# Load all original Anima components and the distilled transformer from this
# repository. Passing the repository explicitly also avoids stale local paths
# embedded by older Modular Diffusers save formats.
pipe.load_components(
    [
        "text_encoder",
        "tokenizer",
        "t5_tokenizer",
        "text_conditioner",
        "transformer",
        "scheduler",
        "vae",
    ],
    pretrained_model_name_or_path=repo_id,
    torch_dtype=torch.bfloat16,
)

pipe.to("cuda")

# The distilled runtime uses one conditional transformer pass per step.
pipe.guider.guidance_scale = 1.0

prompt = (
    "masterpiece, best quality, score_8, safe, anime illustration, "
    "solo shrine maiden standing in a rainlit courtyard, "
    "wide shot, reflections, evening light, detailed background"
)

generator = torch.Generator(device="cpu").manual_seed(0)

image = pipe(
    prompt=prompt,
    negative_prompt=None,
    height=1024,
    width=1024,
    num_inference_steps=4,
    sigmas=[1.0, 0.75, 0.5, 0.25],
    max_sequence_length=512,
    generator=generator,
    output="images",
)[0]

image.save("anima_lightning.png")

Why the sigmas are explicit

The model was distilled for exactly four transitions:

1.00 -> 0.75
0.75 -> 0.50
0.50 -> 0.25
0.25 -> 0.00

Passing the sigmas explicitly prevents an accidental fallback to the normal full-step Anima schedule. The included scheduler performs the Euler updates and appends the terminal 0.0 sigma.

Do not replace the included scheduler with a scheduler intended for another model family.

Guidance and negative prompts

The recommended runtime uses:

pipe.guider.guidance_scale = 1.0

At CFG 1.0, inference uses one conditional transformer evaluation per diffusion step. A negative-prompt branch is not used:

negative_prompt = None

Using CFG greater than 1.0 changes the runtime from the distilled trajectory, increases computation, and may degrade output.

Lower-memory loading

If the complete pipeline does not fit in GPU memory, use CPU offloading instead of pipe.to("cuda"):

pipe.enable_model_cpu_offload()
pipe.vae.enable_tiling()
pipe.vae.enable_slicing()

CPU offloading reduces peak GPU memory use but makes generation slower.

Prompting

Anima-Lightning accepts Danbooru-style tags, natural-language descriptions, or a mixture of both.

A useful prompt structure is:

[quality and safety tags], [subject], [character], [series or artist], [scene and details]

Tips:

  • Put quality, style, and safety tags near the beginning.
  • Use lowercase tags.
  • Use spaces instead of underscores, except for score tags such as score_8.
  • Prefix artist tags with @.
  • Describe appearance, pose, composition, lighting, and background explicitly.
  • For multiple characters, describe each character separately.
  • Very short prompts may provide weak subject control.

Incorrect usage

Do not use upstream full-step settings:

# Incorrect for Anima-Lightning
image = pipe(
    prompt=prompt,
    num_inference_steps=40,
    guidance_scale=5.0,
).images[0]

Do not rely on the pipeline's default step count:

# Incorrect: this may select the normal full-step default.
image = pipe(prompt)

Use the explicit distilled runtime:

pipe.guider.guidance_scale = 1.0

image = pipe(
    prompt=prompt,
    negative_prompt=None,
    height=1024,
    width=1024,
    num_inference_steps=4,
    sigmas=[1.0, 0.75, 0.5, 0.25],
    max_sequence_length=512,
    output="images",
)[0]

Reproducibility

For deterministic seed handling, use a CPU generator:

generator = torch.Generator(device="cpu").manual_seed(42)

Exact output can still vary between PyTorch, CUDA, Diffusers, and GPU versions.

Limitations

  • The model is designed for anime, illustration, and stylized artwork.
  • It is not optimized for photorealism.
  • It requires the four-step distilled runtime described above.
  • Text rendering may be unreliable, especially for long phrases.
  • Hands, complex anatomy, and dense multi-character scenes may contain errors.
  • Short or underspecified prompts may produce weak subject control.
  • Results differ from upstream Anima because the transformer weights are distilled.
  • Alternative schedulers, higher CFG, and full-step sampling are not validated.
  • This model has not been optimized for the Hugging Face Inference API.

License

This model is derived from circlestone-labs/Anima and is distributed under the CircleStone Labs Non-Commercial License.

Anima is itself derived from nvidia/Cosmos-Predict2-2B-Text2Image, so NVIDIA's Open Model License Agreement may also apply.

Review the repository's LICENSE.md and NOTICE.md before using, redistributing, or publishing the model. This summary is not legal advice.

Credits

Anima-Lightning builds on:

Thanks to the upstream authors and open-source contributors.

Downloads last month
2
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aina-tech/Anima-Lightning

Finetuned
(106)
this model