geolip-beatrix-anima

Experiments on steering Anima, a 2B illustration model built on NVIDIA Cosmos-Predict2, with a second conditioning source: Beatrix, a byte-level language model from the geolip line. Anima reads its prompt through a small language model (Qwen3 0.6B) and a light adapter, so an added signal is not drowned by a very large text encoder; that makes it a bed for testing how far a second source can steer the image without a full diffusion training run. The first experiments measure the bed itself: how its conditioning responds to mood words and to a mood direction added to it, and what LoRAs trained on the model's own mood images do.

Each experiment has its own folder under experiments/ with a README (the question, the recipe, the rule fixed before the run, the result), meta.json, and its configuration, weights, evaluation and logs where it has them.

Experiments

folder date what result
e001_anima_flavor_test 2026-10-03 The stock model: mood words, a mood direction added to the conditioning, and the conditioning norms upbeat words +2.59 (UPBEAT WORDS MOVE IT), downbeat words -1.39 (DOWNBEAT WORDS MOVE IT); dial per unit alpha: source +0.03 (NO EFFECT), context +0.31 (A DIAL)
e002_lora_mood_up 2026-10-03 Mood LoRA: upbeat images, neutral captions NO EFFECT: mood effect +0.35 +- 0.24 at scale 1, 56% of 32 held-out cells the expected way
e003_lora_neutral_control 2026-10-03 Control LoRA: the model's own neutral images, neutral captions control: mood effect -0.03 +- 0.18 at scale 1; CONTROL QUIET against e002
e004_lora_mood_down 2026-10-03 Mood LoRA, mirrored: downbeat images, neutral captions MIXED: mood effect -0.69 +- 0.23 at scale 1, 69% of 32 held-out cells the expected way
e005_lora_mood_up_reseed 2026-10-03 Mood LoRA, replicate: a fresh draw of upbeat images FLAVOR LORA: mood effect +1.31 +- 0.28 at scale 1, 84% of 32 held-out cells the expected way
e006_lora_mood_up_lr_1e-5 2026-10-03 Mood LoRA at half the learning rate FLAVOR LORA: mood effect +1.07 +- 0.24 at scale 1, 75% of 32 held-out cells the expected way
e007_lora_mood_up_lr_4e-5 2026-10-03 Mood LoRA at double the learning rate FLAVOR LORA: mood effect +1.00 +- 0.27 at scale 1, 81% of 32 held-out cells the expected way
e008_lora_neutral_control_draw2 2026-10-03 Control LoRA, second draw: the model's own neutral images control: mood effect +0.12 +- 0.19 at scale 1; CONTROL QUIET against e005
e009_lora_mood_down_draw2 2026-10-03 Mood LoRA, mirrored, second draw: downbeat images NO EFFECT: mood effect -0.39 +- 0.21 at scale 1, 62% of 32 held-out cells the expected way
e010_lora_mood_up_lr_1e-5_draw2 2026-10-03 Mood LoRA at half the learning rate, second draw FLAVOR LORA: mood effect +1.18 +- 0.19 at scale 1, 84% of 32 held-out cells the expected way
e011_lora_mood_up_lr_4e-5_draw2 2026-10-03 Mood LoRA at double the learning rate, second draw FLAVOR LORA: mood effect +1.40 +- 0.26 at scale 1, 75% of 32 held-out cells the expected way
e012_anima_attribute_screen 2026-10-03 The stock model: attribute words, attribute sliders after the adapter, and their cross-talk hair_length WORDS ONLY; hair_colour SLIDER WITH CROSS-TALK; eye_colour WORDS ONLY; proportions NO HANDLE; age WORDS ONLY; style WORDS ONLY
e013_beatrix_mood_connector 2026-10-03 Beatrix's mood phrases steer the image through a learned push NOT LEARNED (upbeat +0.23, downbeat +0.21); NOT LEARNED (upbeat +0.29, downbeat +0.13); NEUTRAL MOVES (+0.29)
e014_beatrix_random_trunk_connector 2026-10-03 Control: the same connector on an untrained Beatrix of the same shape NOT LEARNED (upbeat -0.02, downbeat -0.07); NOT LEARNED (upbeat -0.01, downbeat -0.05); NEUTRAL MOVES (-0.05); against e013: NOT READABLE (Beatrix's connector did not learn)
e015_free_vector_connector 2026-10-03 Capacity reference: one free learned vector per mood class, no encoder NOT LEARNED (upbeat +0.43, downbeat -1.16); NEUTRAL QUIET (+0.03)
e016_beatrix_mood_connector_whitened 2026-10-03 Beatrix's mood phrases steer the image through a learned push, whitened input NOT LEARNED (upbeat +0.31, downbeat -0.36); NOT LEARNED (upbeat +0.14, downbeat -0.15); NEUTRAL MOVES (+0.13)
e017_beatrix_random_trunk_connector_whitened 2026-10-03 Control: the same whitened connector on an untrained Beatrix of the same shape NOT LEARNED (upbeat +0.19, downbeat -0.12); NOT LEARNED (upbeat +0.17, downbeat -0.06); NEUTRAL QUIET (-0.03); against e016: NOT READABLE (Beatrix's connector did not learn)
e018_beatrix_mood_slider 2026-10-03 Beatrix's reading of a phrase's mood as a slider value NOT LEARNED (upbeat +0.78, downbeat +0.14); NOT LEARNED (upbeat +0.46, downbeat -0.01); NEUTRAL QUIET (+0.22)
e019_beatrix_random_trunk_slider 2026-10-03 Control: the same slider on an untrained Beatrix of the same shape NOT LEARNED (upbeat +0.39, downbeat -0.46); NOT LEARNED (upbeat +0.10, downbeat -0.09); NEUTRAL QUIET (-0.03); against e018: NOT READABLE (Beatrix's connector did not learn)
e020_anima_route_split 2026-10-04 The stock model: which of the adapter's two readings of a caption carries the mood words, Qwen3's states or the T5 word ids the words +2.59 / -1.39; Qwen3's states only -0.04 / -0.12 (CARRIES NOTHING); the T5 ids only +1.94 / -0.43 (ONE WAY)
e021_anima_route_split_appended 2026-10-04 The stock model: the mood words appended after the scene, through the adapter's query half, its source half, or both the words +2.62 / -2.65; Qwen3's states only -0.08 / -0.11 (CARRIES NOTHING); the T5 ids only +2.53 / -0.52 (ONE WAY); the source half beside the query half +0.09 / -2.13 (downbeat: THE SOURCE HALF ADDS TO THE QUERY HALF)
e022_anima_query_dial 2026-10-04 The stock model: a mood direction added to the adapter's queries, alone and paired with a source token a query can look up query dial +0.057/unit (MIXED); downbeat side, a source token beside it: direction +0.01 (NO EFFECT), state +0.05 (NO EFFECT); uniform control +0.05 (NO EFFECT)
e023_beatrix_mood_slider_two_sided 2026-10-03 Beatrix's reading of a phrase's mood as a two-sided slider NOT LEARNED (upbeat +0.49, downbeat -0.14); NOT LEARNED (upbeat +0.17, downbeat -0.33); NEUTRAL QUIET (-0.07)
e024_beatrix_random_trunk_slider_two_sided 2026-10-03 Control: the same two-sided slider on an untrained Beatrix of the same shape NOT LEARNED (upbeat +0.20, downbeat -0.42); NOT LEARNED (upbeat +0.08, downbeat -0.04); NEUTRAL QUIET (+0.04); against e023: NOT READABLE (Beatrix's connector did not learn)
e025_beatrix_mood_slider_smooth 2026-10-03 Beatrix's reading of a phrase's mood as a smooth two-sided slider NOT LEARNED (upbeat +2.33, downbeat +0.08); NOT LEARNED (upbeat +0.98, downbeat +0.07); NEUTRAL QUIET (+0.35)
e026_anima_word_split 2026-10-04 The stock model: single mood words through the adapter's query half alone or through both halves, whole T5 tokens against shattered ones BY TOKENIZATION; the query half's share: cheerful whole 1.03, cheerful shattered 0.03, gloomy whole 1.33, gloomy shattered 0.37
e027_anima_slot_pair 2026-10-04 The stock model: a word-sized mood push at one word's position, on the adapter's query side, its source side, both, or the source side under a content-free question per word of push: the query at the slot +0.709 (A DIAL), the pair +0.642 (A DIAL), the answer alone +0.140 (MIXED), the answer under a content-free question +0.122 (MIXED); downbeat side, the matched source beside the query +0.04 (NO EFFECT)
e028_anima_word_swap 2026-10-04 The stock model: a mood word cut into pieces, read with another word's Qwen3 states in its place; does the picture follow the answer's mood? THE ANSWER CARRIES THE MOOD: a word's own pieces with another word's reading, cheerful answers minus gloomy ones +2.050 +- 0.370 (91% of cells), 0.64 of the words' own axis; the neutral carriers: 'workaday' NO EFFECT (0.07 of the answer's own effect), 'quotidian' NO EFFECT (0.12)

Reads across the LoRA sequence (rules fixed before the runs)

Each draw is one independent set of training images (seeds N to N+7); its upbeat LoRA at learning rate 2e-05 is the reference the control, the mirror and the learning rates are read against.

read draw 1000-1007 draw 2000-2007
control LoRA (own neutral images) against the upbeat LoRA CONTROL QUIET CONTROL QUIET
upbeat minus control, cell by cell +0.381 +- 0.192, 53% positive: NOT SHOWN +1.188 +- 0.289, 84% positive: THE MOOD COMES FROM THE IMAGES
downbeat LoRA -0.692 (upbeat +0.347): MIXED downward -0.390 (upbeat +1.313): NO EFFECT downward
  • The upbeat LoRA on the two draws: DOES NOT REPLICATE.
  • the control: SETTLED across the two draws.
  • upbeat minus control: UNSETTLED across the two draws.
  • the mirror: UNSETTLED across the two draws.
learning rate draw 1000-1007: final effect, first epoch beyond 3 SE draw 2000-2007: final effect, first epoch beyond 3 SE
1e-05 +1.073 +- 0.238, 4 +1.183 +- 0.192, 4
2e-05 +0.347 +- 0.242, 4 +1.313 +- 0.278, 2
4e-05 +0.999 +- 0.266, 2 +1.397 +- 0.259, 2

How the mood experiments are measured

The mood judge. Each image is scored on its pixels only, with CLIP ViT-L/14: 100 x (the mean cosine similarity to three upbeat phrases, "a cheerful, upbeat image", "a happy, joyful scene", "a bright, uplifting photo", minus the same for three downbeat phrases, "a gloomy, downbeat image", "a sad, melancholy scene", "a dark, depressing photo"); the same judge as the Sana experiments, so the two beds read on one scale. How far mood words in the prompt move this score on the stock model is experiment e001. Content kept is the CLIP image cosine between an image and the no-LoRA image of the same prompt and seed.

The LoRA experiments train on 24 everyday scenes and are scored on 8 scenes they never saw (4 seeds each, 32 paired cells): each cell compares the same prompt and seed with and without the LoRA.

Tools

Training: diffusion-pipe (the AbstractEyes fork, model type anima; plain Adam, no weight decay, fp32 master weights over the bf16 LoRA; the LLM adapter frozen), driven by anima-trainer (notebooks/anima_colab_experiments.ipynb). The images are rendered in the notebook process with the fork's own Anima model code (Euler flow sampler, shift 3, 30 steps, guidance 4.5, the model card's quality prefix and negative prompt). The LoRAs are in ComfyUI format.

References

Licences

Anima's weights are under the CircleStone Labs Non-Commercial License (Anima is a derivative of NVIDIA Cosmos-Predict2-2B, under the NVIDIA Open Model License); the LoRAs here are derivatives under the same non-commercial terms. Generated images are not restricted by that licence.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AbstractPhil/geolip-beatrix-anima

Adapter
(129)
this model

Papers for AbstractPhil/geolip-beatrix-anima