--- license: mit tags: - executorch - xnnpack - pte - on-device - image-text-to-text - image-to-text - object-detection base_model: - microsoft/Florence-2-large --- # Florence-2-large — ExecuTorch (vision + text encoder + text decoder) One set of weights that captions, detects, reads text and grounds phrases, with the task chosen by the prompt you send. 0.77 B parameters — the same interface as [Florence-2-base](https://huggingface.co/mlboydaisuke/Florence-2-base-ExecuTorch) at roughly three times the size, and it notices more: on the same photograph base says *"A black and white photo of a person playing a piano"* where large reads the brand off the instrument. Three `.pte` files, split the way Whisper is on this shelf and for the same reason: the vision tower and the text encoder run once per image, the decoder runs once per generated token. ``` vision : pixel_values (1,3,768,768) -> image_features (1,577,1024) encoder: (image_features, input_ids (1,32), mask (1,32)) -> hidden (1,609,1024) decoder: (hidden, mask (1,32), decoder_input_ids (1,128)) -> logits (1,128,51328) ``` Every file takes and returns fp32 tensors, so a precision is a file swap and the three parts can be mixed. | part | build | file | MB | corr vs fp32 eager | Mac ms* | |---|---|---|---|---|---| | vision | XNNPACK fp32 | `florence2_large_vision_xnnpack_fp32.pte` | 1452.6 | 1.000000 | 803.9 | | vision | Core ML (iOS) | `florence2_large_vision_coreml_all.pte` | 729.5 | 0.999953 | 238.2 | | encoder | XNNPACK fp32 | `florence2_large_encoder_xnnpack_fp32.pte` | 831.8 | 1.000000 | 144.4 | | encoder | Core ML (iOS) | `florence2_large_encoder_coreml_all.pte` | 409.2 | 0.999944 | 43.1 | | decoder | XNNPACK fp32 | `florence2_large_decoder_xnnpack_fp32.pte` | 1243.8 | 1.000000 | 85.8 | | decoder | Core ML (iOS) | `florence2_large_decoder_coreml_all.pte` | 509.5 | 0.999980 | 13.8 | Two sets: **3528 MB** all-fp32, **1648 MB** Core ML. Torch eager fp32 on the same machine: vision 1288.2 ms, encoder 96.8 ms, decoder 60.6 ms. \*Mac arm64, single process, median of 5 — a reference point for relative cost, not a device number. A caption of *n* tokens costs one vision pass, one encoder pass and *n* decoder passes: about 2.2 s for a 15-token caption on the fp32 set and 0.49 s on Core ML, on this Mac. The decoder graph is a fixed 128-token window, so every step costs the same whether it is the first token or the fiftieth. ## Running it Identical to the base model's contract except for the width (1024 instead of 768). **1. The image.** RGB, divide by 255, ImageNet normalise (mean .485/.456/.406, std .229/.224/.225), bicubic resize to 768×768. No crop. Run the vision `.pte`. **2. The prompt.** Florence-2's task tokens are shorthand the processor expands into a sentence before tokenising — the model never sees ``: | task | the sentence that is actually tokenised | |---|---| | `` | What does the image describe? | | `` | Describe in detail what is shown in the image. | | `` | Describe with a paragraph what is shown in the image. | | `` | Locate the objects with category name in the image. | | `` | Locate the objects in the image, with their descriptions. | | `` | Locate the region proposals in the image. | | `` | What is the text in the image? | | `` | What is the text in the image, with regions? | | `` | Locate the phrases in the caption: {your caption} | | `` | Locate {your phrase} in the image. | Tokenise that sentence as ` sentence ` with the repo's tokenizer, right-pad to 32 with the pad id (1), and build an `attention_mask` that is 1 on the real tokens and 0 on the padding. Run the encoder `.pte` with `(image_features, input_ids, mask)`. The 577 image tokens are handled inside the graphs. The prompt sequence the original model sees is `×577 + prompt `, and because the image tokens are a contiguous prefix, the encoder here concatenates the vision features in front of the text embeddings instead of scattering them into placeholder positions. **3. Greedy decoding, and the rule you cannot skip.** Fill a `(1,128)` int64 window with the pad id, write the decoder start token (2) at position 0, then for step `t`: ``` logits = decoder(hidden, mask, window) # mask is the same one the encoder took ban every token that would repeat a 3-gram already in the output # <- see below next = argmax(logits[0, t]) if next == 2: stop # window[0, t + 1] = next ``` **`no_repeat_ngram_size: 3` from the model's `generation_config.json` is load-bearing on this size.** Large's decoder returns `` as its argmax three times in a row on most images. The ban on repeating that 3-gram is the only thing that moves it on to the caption — a plain argmax loop emits `` forever and returns an empty string. On five test photographs it did so every time, before and after conversion, in eager PyTorch as well as through the `.pte`. Base tolerates the omission and large does not, so implement the rule. **4. Reading a detection answer.** Grounded tasks answer with `` tokens, `N` in 0..999. Four in a row are a box, and each coordinate is `(N + 0.5) × side / 1000` in the original image's pixels — `side` being the image's width for x and its height for y, not 768. ## Verification The three wrappers reproduce `Florence2ForConditionalGeneration` **exactly**: composition `max_abs_diff` 0.000e+00 against the full model's logits on the same image and prompt. End to end through the three `.pte` files, greedy `` on five photographs against the same decoding in eager PyTorch: **fp32 5/5 and Core ML 5/5 captions identical**, character for character. ``` A person's hands playing a piano with the words Lauberger and Gloss written on it. A long wooden pier stretching out into the ocean on a sunny day. A forest of dead trees in the middle of a forest. A road in the middle of a pine forest lined with tall trees. A couple of wooden benches sitting on top of a park bench covered in leaves. ``` ```bash python convert/check_florence2.py large fp32 # or coreml_all ``` ## Not shipped, and why **int8 converts and is not published.** Dynamic int8 brings the set from 3528 MB to 1304 MB at correlations of 0.999, 0.9995 and 0.9893 — numbers that would pass any gate on this shelf. The captions do not: **2 of 5 match eager**, and the other three are plausible but different sentences, one of them reworded from the first word. The same recipe on Florence-2-base keeps all five, so this is a property of the larger decoder rather than of the recipe. On iOS the Core ML set is the small one (1648 MB, 5/5); on Android it is fp32 or nothing until a better recipe is found. **fp16 for the encoder and the decoder does not export.** BART clamps its activations when, and only when, they are half precision: ```python if hidden_states.dtype == torch.float16 and not torch.isfinite(hidden_states).all(): ``` In fp32 that line short-circuits and never reaches the graph. Halve the model and it becomes a question about values `torch.export` cannot answer, and export stops with `GuardOnDataDependentSymNode`. ## Conversion notes Converted from [florence-community/Florence-2-large](https://huggingface.co/florence-community/Florence-2-large), the transformers-format mirror of [microsoft/Florence-2-large](https://huggingface.co/microsoft/Florence-2-large) — same MIT weights. The original repo predates the in-tree implementation and its weight names do not match it: loading it into `Florence2ForConditionalGeneration` prints a load report where every key is unexpected and hands back a randomly initialised model without raising. The checkpoint declares `torch_dtype: float16` and transformers honours it, so `from_pretrained` must be given `dtype=torch.float32` explicitly. - **Source**: microsoft/Florence-2-large (via florence-community mirror) - **License**: MIT torch.export -> to_edge_transform_and_lower(partitioner) -> .pte (conversion scripts: [executorch-models](https://github.com/john-rocky/executorch-models))