geolip-beatrix-anima
Experiments on steering Anima, a 2B illustration model built on NVIDIA Cosmos-Predict2, with a second conditioning source: Beatrix, a byte-level language model from the geolip line. Anima reads its prompt through a small language model (Qwen3 0.6B) and a light adapter, so an added signal is not drowned by a very large text encoder; that makes it a bed for testing how far a second source can steer the image without a full diffusion training run. The first experiments measure the bed itself: how its conditioning responds to mood words and to a mood direction added to it, and what LoRAs trained on the model's own mood images do.
Each experiment has its own folder under experiments/ with a README (the question, the recipe, the rule fixed before
the run, the result), meta.json, and its configuration, weights, evaluation and logs where it has them.
Experiments
| folder | date | what | result |
|---|---|---|---|
e001_anima_flavor_test |
2026-10-03 | The stock model: mood words, a mood direction added to the conditioning, and the conditioning norms | upbeat words +2.59 (UPBEAT WORDS MOVE IT), downbeat words -1.39 (DOWNBEAT WORDS MOVE IT); dial per unit alpha: source +0.03 (NO EFFECT), context +0.31 (A DIAL) |
e002_lora_mood_up |
2026-10-03 | Mood LoRA: upbeat images, neutral captions | NO EFFECT: mood effect +0.35 +- 0.24 at scale 1, 56% of 32 held-out cells the expected way |
e003_lora_neutral_control |
2026-10-03 | Control LoRA: the model's own neutral images, neutral captions | control: mood effect -0.03 +- 0.18 at scale 1; CONTROL QUIET against e002 |
e004_lora_mood_down |
2026-10-03 | Mood LoRA, mirrored: downbeat images, neutral captions | MIXED: mood effect -0.69 +- 0.23 at scale 1, 69% of 32 held-out cells the expected way |
e005_lora_mood_up_reseed |
2026-10-03 | Mood LoRA, replicate: a fresh draw of upbeat images | FLAVOR LORA: mood effect +1.31 +- 0.28 at scale 1, 84% of 32 held-out cells the expected way |
e006_lora_mood_up_lr_1e-5 |
2026-10-03 | Mood LoRA at half the learning rate | FLAVOR LORA: mood effect +1.07 +- 0.24 at scale 1, 75% of 32 held-out cells the expected way |
e007_lora_mood_up_lr_4e-5 |
2026-10-03 | Mood LoRA at double the learning rate | FLAVOR LORA: mood effect +1.00 +- 0.27 at scale 1, 81% of 32 held-out cells the expected way |
e008_lora_neutral_control_draw2 |
2026-10-03 | Control LoRA, second draw: the model's own neutral images | control: mood effect +0.12 +- 0.19 at scale 1; CONTROL QUIET against e005 |
e009_lora_mood_down_draw2 |
2026-10-03 | Mood LoRA, mirrored, second draw: downbeat images | NO EFFECT: mood effect -0.39 +- 0.21 at scale 1, 62% of 32 held-out cells the expected way |
e010_lora_mood_up_lr_1e-5_draw2 |
2026-10-03 | Mood LoRA at half the learning rate, second draw | FLAVOR LORA: mood effect +1.18 +- 0.19 at scale 1, 84% of 32 held-out cells the expected way |
e011_lora_mood_up_lr_4e-5_draw2 |
2026-10-03 | Mood LoRA at double the learning rate, second draw | FLAVOR LORA: mood effect +1.40 +- 0.26 at scale 1, 75% of 32 held-out cells the expected way |
e012_anima_attribute_screen |
2026-10-03 | The stock model: attribute words, attribute sliders after the adapter, and their cross-talk | hair_length WORDS ONLY; hair_colour SLIDER WITH CROSS-TALK; eye_colour WORDS ONLY; proportions NO HANDLE; age WORDS ONLY; style WORDS ONLY |
e013_beatrix_mood_connector |
2026-10-03 | Beatrix's mood phrases steer the image through a learned push | NOT LEARNED (upbeat +0.23, downbeat +0.21); NOT LEARNED (upbeat +0.29, downbeat +0.13); NEUTRAL MOVES (+0.29) |
e014_beatrix_random_trunk_connector |
2026-10-03 | Control: the same connector on an untrained Beatrix of the same shape | NOT LEARNED (upbeat -0.02, downbeat -0.07); NOT LEARNED (upbeat -0.01, downbeat -0.05); NEUTRAL MOVES (-0.05); against e013: NOT READABLE (Beatrix's connector did not learn) |
e015_free_vector_connector |
2026-10-03 | Capacity reference: one free learned vector per mood class, no encoder | NOT LEARNED (upbeat +0.43, downbeat -1.16); NEUTRAL QUIET (+0.03) |
e016_beatrix_mood_connector_whitened |
2026-10-03 | Beatrix's mood phrases steer the image through a learned push, whitened input | NOT LEARNED (upbeat +0.31, downbeat -0.36); NOT LEARNED (upbeat +0.14, downbeat -0.15); NEUTRAL MOVES (+0.13) |
e017_beatrix_random_trunk_connector_whitened |
2026-10-03 | Control: the same whitened connector on an untrained Beatrix of the same shape | NOT LEARNED (upbeat +0.19, downbeat -0.12); NOT LEARNED (upbeat +0.17, downbeat -0.06); NEUTRAL QUIET (-0.03); against e016: NOT READABLE (Beatrix's connector did not learn) |
e018_beatrix_mood_slider |
2026-10-03 | Beatrix's reading of a phrase's mood as a slider value | NOT LEARNED (upbeat +0.78, downbeat +0.14); NOT LEARNED (upbeat +0.46, downbeat -0.01); NEUTRAL QUIET (+0.22) |
e019_beatrix_random_trunk_slider |
2026-10-03 | Control: the same slider on an untrained Beatrix of the same shape | NOT LEARNED (upbeat +0.39, downbeat -0.46); NOT LEARNED (upbeat +0.10, downbeat -0.09); NEUTRAL QUIET (-0.03); against e018: NOT READABLE (Beatrix's connector did not learn) |
e020_anima_route_split |
2026-10-04 | The stock model: which of the adapter's two readings of a caption carries the mood words, Qwen3's states or the T5 word ids | the words +2.59 / -1.39; Qwen3's states only -0.04 / -0.12 (CARRIES NOTHING); the T5 ids only +1.94 / -0.43 (ONE WAY) |
e021_anima_route_split_appended |
2026-10-04 | The stock model: the mood words appended after the scene, through the adapter's query half, its source half, or both | the words +2.62 / -2.65; Qwen3's states only -0.08 / -0.11 (CARRIES NOTHING); the T5 ids only +2.53 / -0.52 (ONE WAY); the source half beside the query half +0.09 / -2.13 (downbeat: THE SOURCE HALF ADDS TO THE QUERY HALF) |
e022_anima_query_dial |
2026-10-04 | The stock model: a mood direction added to the adapter's queries, alone and paired with a source token a query can look up | query dial +0.057/unit (MIXED); downbeat side, a source token beside it: direction +0.01 (NO EFFECT), state +0.05 (NO EFFECT); uniform control +0.05 (NO EFFECT) |
e023_beatrix_mood_slider_two_sided |
2026-10-03 | Beatrix's reading of a phrase's mood as a two-sided slider | NOT LEARNED (upbeat +0.49, downbeat -0.14); NOT LEARNED (upbeat +0.17, downbeat -0.33); NEUTRAL QUIET (-0.07) |
e024_beatrix_random_trunk_slider_two_sided |
2026-10-03 | Control: the same two-sided slider on an untrained Beatrix of the same shape | NOT LEARNED (upbeat +0.20, downbeat -0.42); NOT LEARNED (upbeat +0.08, downbeat -0.04); NEUTRAL QUIET (+0.04); against e023: NOT READABLE (Beatrix's connector did not learn) |
e025_beatrix_mood_slider_smooth |
2026-10-03 | Beatrix's reading of a phrase's mood as a smooth two-sided slider | NOT LEARNED (upbeat +2.33, downbeat +0.08); NOT LEARNED (upbeat +0.98, downbeat +0.07); NEUTRAL QUIET (+0.35) |
e026_anima_word_split |
2026-10-04 | The stock model: single mood words through the adapter's query half alone or through both halves, whole T5 tokens against shattered ones | BY TOKENIZATION; the query half's share: cheerful whole 1.03, cheerful shattered 0.03, gloomy whole 1.33, gloomy shattered 0.37 |
e027_anima_slot_pair |
2026-10-04 | The stock model: a word-sized mood push at one word's position, on the adapter's query side, its source side, both, or the source side under a content-free question | per word of push: the query at the slot +0.709 (A DIAL), the pair +0.642 (A DIAL), the answer alone +0.140 (MIXED), the answer under a content-free question +0.122 (MIXED); downbeat side, the matched source beside the query +0.04 (NO EFFECT) |
e028_anima_word_swap |
2026-10-04 | The stock model: a mood word cut into pieces, read with another word's Qwen3 states in its place; does the picture follow the answer's mood? | THE ANSWER CARRIES THE MOOD: a word's own pieces with another word's reading, cheerful answers minus gloomy ones +2.050 +- 0.370 (91% of cells), 0.64 of the words' own axis; the neutral carriers: 'workaday' NO EFFECT (0.07 of the answer's own effect), 'quotidian' NO EFFECT (0.12) |
Reads across the LoRA sequence (rules fixed before the runs)
Each draw is one independent set of training images (seeds N to N+7); its upbeat LoRA at learning rate 2e-05 is the reference the control, the mirror and the learning rates are read against.
| read | draw 1000-1007 | draw 2000-2007 |
|---|---|---|
| control LoRA (own neutral images) against the upbeat LoRA | CONTROL QUIET | CONTROL QUIET |
| upbeat minus control, cell by cell | +0.381 +- 0.192, 53% positive: NOT SHOWN | +1.188 +- 0.289, 84% positive: THE MOOD COMES FROM THE IMAGES |
| downbeat LoRA | -0.692 (upbeat +0.347): MIXED downward | -0.390 (upbeat +1.313): NO EFFECT downward |
- The upbeat LoRA on the two draws: DOES NOT REPLICATE.
- the control: SETTLED across the two draws.
- upbeat minus control: UNSETTLED across the two draws.
- the mirror: UNSETTLED across the two draws.
| learning rate | draw 1000-1007: final effect, first epoch beyond 3 SE | draw 2000-2007: final effect, first epoch beyond 3 SE |
|---|---|---|
| 1e-05 | +1.073 +- 0.238, 4 | +1.183 +- 0.192, 4 |
| 2e-05 | +0.347 +- 0.242, 4 | +1.313 +- 0.278, 2 |
| 4e-05 | +0.999 +- 0.266, 2 | +1.397 +- 0.259, 2 |
How the mood experiments are measured
The mood judge. Each image is scored on its pixels only, with CLIP ViT-L/14: 100 x (the mean cosine similarity to three upbeat phrases, "a cheerful, upbeat image", "a happy, joyful scene", "a bright, uplifting photo", minus the same for three downbeat phrases, "a gloomy, downbeat image", "a sad, melancholy scene", "a dark, depressing photo"); the same judge as the Sana experiments, so the two beds read on one scale. How far mood words in the prompt move this score on the stock model is experiment e001. Content kept is the CLIP image cosine between an image and the no-LoRA image of the same prompt and seed.
The LoRA experiments train on 24 everyday scenes and are scored on 8 scenes they never saw (4 seeds each, 32 paired cells): each cell compares the same prompt and seed with and without the LoRA.
Tools
Training: diffusion-pipe (the AbstractEyes fork, model type anima; plain Adam,
no weight decay, fp32 master weights over the bf16 LoRA; the LLM adapter frozen), driven by
anima-trainer (notebooks/anima_colab_experiments.ipynb). The images are rendered in the notebook
process with the fork's own Anima model code (Euler flow sampler, shift 3, 30 steps, guidance 4.5, the model card's quality
prefix and negative prompt). The LoRAs are in ComfyUI format.
References
- CircleStone Labs and Comfy Org, "Anima" (model card, 2026). https://huggingface.co/circlestone-labs/Anima
- NVIDIA et al., "Cosmos World Foundation Model Platform for Physical AI" (2025). https://arxiv.org/abs/2501.03575
- Yang, Li, Yang, Zhang, Hui, Zheng et al., "Qwen3 Technical Report" (2025). https://arxiv.org/abs/2505.09388
- Hu, Shen, Wallis, Allen-Zhu, Li, Wang et al., "LoRA: Low-Rank Adaptation of Large Language Models" (2021). https://arxiv.org/abs/2106.09685
- Liu, Gong, Liu, "Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow" (2022). https://arxiv.org/abs/2209.03003
- Esser, Kulal, Blattmann, Entezari, Muller, Saini et al., "Scaling Rectified Flow Transformers for High-Resolution Image Synthesis" (2024). https://arxiv.org/abs/2403.03206
- Gandikota, Materzynska, Zhou, Torralba, Bau, "Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models" (2023). https://arxiv.org/abs/2311.12092
- Turner, Thiergart, Udell, Leech, Mini, MacDiarmid, "Activation Addition: Steering Language Models Without Optimization" (2023). https://arxiv.org/abs/2308.10248
- Radford, Kim, Hallacy et al., "Learning Transferable Visual Models From Natural Language Supervision" (2021). https://arxiv.org/abs/2103.00020
Licences
Anima's weights are under the CircleStone Labs Non-Commercial License (Anima is a derivative of NVIDIA Cosmos-Predict2-2B, under the NVIDIA Open Model License); the LoRAs here are derivatives under the same non-commercial terms. Generated images are not restricted by that licence.
Model tree for AbstractPhil/geolip-beatrix-anima
Base model
nvidia/Cosmos-Predict2-2B-Text2Image