Florence-Chan-2.1

A Florence-2-large-derived anime caption model, baked into a standalone BF16 safetensors checkpoint. No adapter is required at inference time.

What changed

  • Based on the v2 cumulative teacher line and its follow-up SFT continuation.
  • Extended with the latest tranche plus an equal-sized v2 retention mix.
  • Kept the standard Florence 2 large baseline behavior close while improving the validated follow-up row set.
  • The training lineage used GLM-4.6V and Grok-4.3 teacher captions.

Training data

  • Broader cumulative teacher lineage: 44,647 rows.
  • Follow-up SFT used for this checkpoint: 980 rows.
  • Follow-up SFT composition: 490 latest-tranche rows + 490 retention rows.

Package

  • model: Florence-Chan-2.1
  • parent model: Florence 2 large
  • follow-up SFT size: 980 image-caption rows
  • cumulative teacher lineage size: 44,647 rows
  • base model: microsoft/Florence-2-large
  • adapter source: /data/anime-captioner/checkpoints/florence2_cumulative_teacher_v2_extension_late_stage_v1/best
  • merged checkpoint: /data/anime-captioner/hf.co/florence2-cumulative-teacher-v2p1
  • safetensors: /data/anime-captioner/hf.co/florence2-cumulative-teacher-v2p1/model.safetensors

Safe comparison set

The table below compares the baked model against the standard Florence 2 large baseline and the raw WD14 tags on eight safe, diverse images.

dataset/danbooru_images/danbooru_8428374.jpg

dataset/danbooru_images/danbooru_8428374.jpg

View Content
WD14 tags weapon (0.88)
long hair (0.88)
sword (0.85)
animal ears (0.85)
holding weapon (0.68)
fox ears (0.67)
holding (0.67)
holding sword (0.65)
katana (0.61)
skirt (0.59)
Florence-2 large Short: An illustration of a woman with a sword in a city.
Detail: The image shows an animated person holding a sword in front of a cityscape, with buildings in the background and a watermark on the image.
Florence-Chan-2.1 Short: Dark-haired horned warrior in dark armor wielding a large sword amid blue energy effects in a ruined city.
Detail: A dark-haired figure with horns and a flowing black cloak stands in a dynamic pose, gripping a large sword with both hands, surrounded by swirling blue energy effects and debris in a ruined cityscape with broken buildings and a pink flower-like structure floating in the air.

dataset/danbooru_images/danbooru_5341516.jpg

dataset/danbooru_images/danbooru_5341516.jpg

View Content
WD14 tags mostima (arknights) (0.97)
multiple girls (0.91)
2girls (0.91)
doughnut (0.90)
food (0.88)
? (0.87)
tail (0.86)
horns (0.86)
halo (0.85)
hug from behind (0.83)
Florence-2 large Short: A couple of anime girls sitting next to each other.
Detail: The image shows two anime girls sitting next to each other on a couch, with one of them holding a donut in her hand. The image is animated, giving it a lively and vibrant feel.
Florence-Chan-2.1 Short: Blue-haired cat-eared girl holds a donut beside a red-haired horned girl in a hoodie.
Detail: A blue-haired girl with cat ears and a halo holds a donut in her right hand while resting her left hand on the lap of a pink-haired horned girl in a black hoodie, their arms wrapped around each other in a close embrace.

dataset/danbooru_images/danbooru_5294973.jpg

dataset/danbooru_images/danbooru_5294973.jpg

View Content
WD14 tags guts (berserk) (0.97)
armor (0.92)
multiple boys (0.92)
2boys (0.92)
weapon (0.88)
long hair (0.83)
sword (0.78)
blue skin (0.74)
colored skin (0.73)
wavy hair (0.70)
Florence-2 large Short: A group of anime characters standing next to each other.
Detail: The image shows a group of four people standing next to each other in front of a white background. The person in the front is wearing a white dress and is holding a sword, while the person to the left is wearing an unknown outfit. The image is animated, giving it a dynamic feel.
Florence-Chan-2.1 Short: White-haired woman in silver armor flanked by two dark-haired men in red robes.
Detail: A woman with long white curly hair and blue eyes stands in the foreground wearing a white and silver armor with a glowing blue orb at her chest. To her left is a young man with short black hair and a red jacket, and to her right is a man with spiky dark hair wearing a red cape and holding a sword.

dataset/danbooru_images/danbooru_7019581.jpg

dataset/danbooru_images/danbooru_7019581.jpg

View Content
WD14 tags 1girl (0.94)
scarf (0.91)
solo (0.89)
lying (0.85)
animal ears (0.85)
halo (0.85)
one eye closed (0.82)
on back (0.82)
winter clothes (0.73)
smile (0.68)
Florence-2 large Short: A girl in a white coat laying on a pillow.
Detail: The image shows a girl in a white coat laying on top of a bed next to a cat, with a pillow beneath her head. She appears to be sleeping peacefully, with her eyes closed and her arms tucked in close to her body. The background is a soft, muted color, giving the image a dreamy, ethereal feel.
Florence-Chan-2.1 Short: Cat-eared girl in a white coat lies on a bed with a fluffy tail.
Detail: A cat-eared girl with brown and white hair lies on her back across white sheets, eyes closed and mouth open, wearing a white puffy coat, pink scarf, black gloves, and brown boots. A large fluffy tail rests between her legs, and a pink circular emblem floats above her head.

dataset/danbooru_images/danbooru_8654121.jpg

dataset/danbooru_images/danbooru_8654121.jpg

View Content
WD14 tags 1girl (0.99)
japanese clothes (0.96)
kimono (0.95)
umbrella (0.94)
solo (0.89)
animal ears (0.88)
looking back (0.80)
sword (0.79)
weapon (0.79)
from behind (0.78)
Florence-2 large Short: A woman in a kimono holding an umbrella and a cat.
Detail: The image shows a woman in a kimono sitting on a bench, holding an umbrella in one hand and a stick in the other. She is surrounded by trees and a shed in the background, creating a peaceful atmosphere. The image is animated, giving it a dynamic feel.
Florence-Chan-2.1 Short: Cat-eared girl in a blue floral kimono holding a red umbrella.
Detail: A young woman with black hair in a high bun and fox ears wears a teal kimono patterned with white and gold floral designs while seated on a wooden surface. She holds a small red object in one hand and rests the other on her lap, with a red umbrella visible behind her and a sword resting on the ground in front of her.

dataset/danbooru_images/danbooru_5221484.jpg

dataset/danbooru_images/danbooru_5221484.jpg

View Content
WD14 tags hina (blue archive) (0.99)
1girl (0.97)
hina (swimsuit) (blue archive) (0.96)
whistle (0.95)
innertube (0.93)
horns (0.92)
purple eyes (0.85)
ahoge (0.85)
long hair (0.84)
hair ornament (0.82)
Florence-2 large Short: A cartoon of a girl floating in a pool with a duck.
Detail: The image shows a cartoon of a girl floating in the water with a crown on her head, wearing a life jacket and holding an object in her hand. The background is white and there is some text at the bottom of the image.
Florence-Chan-2.1 Short: White-haired horned girl in an inflatable ring with a white duck floating in blue water.
Detail: A white-haired girl with spiky hair and a red bandage floats in blue water, her face flushed with eyes closed and mouth open, while a small white creature with a pink bow floats beside her and a purple crown floats above her head.

dataset/danbooru_images/danbooru_8643115.jpg

dataset/danbooru_images/danbooru_8643115.jpg

View Content
WD14 tags multiple girls (0.98)
2girls (0.96)
kimono (0.96)
japanese clothes (0.95)
wings (0.93)
seiza (0.91)
hina (blue archive) (0.91)
sash (0.89)
halo (0.87)
sitting (0.85)
Florence-2 large Short: A couple of girls sitting on top of a tatami mat.
Detail: The image shows three anime girls in kimono sitting on the floor in front of a window, with a houseplant on the left side and text on the image. Through the window, we can see trees and snow outside.
Florence-Chan-2.1 Short: Pink-haired girl in a white kimono and white-haired dragon-eared girl in purple kimonos sit on a tatami mat.
Detail: A pink-haired girl in a white kimono with a pink bow and floral headpiece sits on the left, holding a pink bowl, while a white-haired dragon-like girl with purple eyes kneels on the right in a purple kimonos with black wings and a white headpiece, both positioned on yellow tatami mats in front of large windows with snow-covered trees visible outside.

dataset/danbooru_images/danbooru_6621403.jpg

dataset/danbooru_images/danbooru_6621403.jpg

View Content
WD14 tags 1girl (0.99)
solo (0.97)
animal ears (0.95)
sword (0.95)
weapon (0.94)
holding sword (0.93)
holding (0.93)
holding weapon (0.92)
long hair (0.86)
bear ears (0.86)
Florence-2 large Short: A girl in a black dress holding a sword in a field.
Detail: The image shows a girl in a black dress holding a sword in a field surrounded by plants with flowers, trees, hills, and a starry sky.
Florence-Chan-2.1 Short: Purple-haired girl with bear ears holds a sword in a field at night.
Detail: A young woman with long purple hair in twin pigtails tied with purple ribbons and teddy bear ears sits on a grassy field at night, wearing a black kimono-style outfit with purple trim and a plaid skirt, holding a tall silver sword in her right hand with her left arm raised.

Notes

  • Images are included at 512px for table display.
  • Baseline is the standard Florence-2-large captioning path.
  • WD14 tags are shown as the raw local tagger output for the same images.
  • The checkpoint is merged and serialized as safetensors, so there is no adapter dependency.

Run with Transformers

Known-good runtime: transformers==4.46.3. The standard Florence-2 processor/model API works directly against the baked checkpoint.

from PIL import Image
import torch
from transformers import AutoModelForCausalLM, AutoProcessor

model_id = "/data/anime-captioner/hf.co/florence2-cumulative-teacher-v2p1"
image_path = "/data/anime-captioner/dataset/danbooru_images/danbooru_8428374.jpg"

processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
).eval()

image = Image.open(image_path).convert("RGB")

for task in ["<CAPTION>", "<DETAILED_CAPTION>"]:
    inputs = processor(text=task, images=image, return_tensors="pt")
    inputs = {k: v.to(model.device) for k, v in inputs.items()}

    with torch.no_grad():
        generated_ids = model.generate(
            input_ids=inputs["input_ids"],
            pixel_values=inputs["pixel_values"],
            max_new_tokens=128,
            do_sample=False,
            num_beams=1,
        )

    text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
    print(task, text)
Downloads last month
262
Safetensors
Model size
0.8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Mitchins/Florence-Chan-2.1

Adapter
(10)
this model