File size: 10,060 Bytes
3cec3d6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
656bf90
3cec3d6
084b099
3cec3d6
 
 
656bf90
 
3cec3d6
 
 
a0fb2f8
 
 
3cec3d6
24da67a
 
3cec3d6
 
 
 
084b099
3cec3d6
084b099
3cec3d6
084b099
3cec3d6
084b099
 
 
 
 
 
 
3cec3d6
084b099
3cec3d6
084b099
 
 
 
3cec3d6
084b099
 
 
 
 
 
3cec3d6
084b099
 
3cec3d6
084b099
 
 
 
 
3cec3d6
084b099
 
 
 
 
 
 
 
 
 
3cec3d6
084b099
 
 
 
3cec3d6
084b099
 
 
 
 
 
3cec3d6
084b099
3cec3d6
084b099
 
 
3cec3d6
084b099
 
 
 
 
 
3cec3d6
084b099
3cec3d6
084b099
 
3cec3d6
084b099
3cec3d6
084b099
 
 
 
 
3cec3d6
084b099
 
 
 
3cec3d6
084b099
3cec3d6
084b099
3cec3d6
084b099
 
 
 
 
 
 
 
 
3cec3d6
084b099
 
3cec3d6
084b099
3cec3d6
084b099
 
 
 
 
 
 
 
 
 
 
 
3cec3d6
084b099
3cec3d6
084b099
3cec3d6
084b099
 
 
3cec3d6
084b099
3cec3d6
c047c57
 
 
 
 
 
 
 
 
 
 
982e411
c047c57
084b099
3cec3d6
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
---
license: cc-by-nc-4.0
pipeline_tag: feature-extraction
base_model: EximiusLabs/fusion-embedding-2-2b-preview
tags:
- embeddings
- retrieval
- multimodal
- thermal
- infrared
- adapters
---

<p align="center">
  <img src="assets/ember-banner.png" alt="Ember — the thermal sense for Fusion Embedding 2 (2B-Preview) — Eximius Labs" width="100%">
</p>

# Ember — the thermal sense for fusion-embedding-2

<div align="center">

[![Python](https://img.shields.io/badge/python-3.11+-blue.svg)](https://github.com/Eximius-Labs/fusion-embedding) [![PyTorch](https://img.shields.io/badge/PyTorch-2.x-ee4c2c.svg)](https://github.com/Eximius-Labs/fusion-embedding) [![Weights](https://img.shields.io/badge/weights-CC--BY--NC--4.0-green.svg)](#license) [![Status](https://img.shields.io/badge/status-research%20preview%20v0.1-orange.svg)](#) [![Code](https://img.shields.io/badge/code-GitHub-black.svg)](https://github.com/Eximius-Labs/fusion-embedding)

[![Deploy on RunPod](https://api.runpod.io/badge/Eximius-Labs/ember)](https://www.runpod.io/console/hub/Eximius-Labs/ember)

</div>

Ember is the first sense pack for [fusion-embedding-2](https://huggingface.co/EximiusLabs/fusion-embedding-2-2b-preview): it teaches the
model to embed thermal infrared images in the same vector space as its text,
image, video, and audio embeddings. Packs are named for the physical trace their
sensor reads; Ember reads heat, and its sibling pack
[fusion-embedding-2-tactus](https://huggingface.co/EximiusLabs/fusion-embedding-2-tactus)
reads touch (32x32 pressure/taxel arrays).

**The family.** Each sense is a separately loadable pack over the same frozen base: [Tactus](https://huggingface.co/EximiusLabs/fusion-embedding-2-tactus) reads touch from a 32x32 pressure glove, [Tactus Mat](https://huggingface.co/EximiusLabs/fusion-embedding-2-tactus-mat) reads a 64x32 body pressure mat, [Ember](https://huggingface.co/EximiusLabs/fusion-embedding-2-ember) reads heat, and [Tremor](https://huggingface.co/EximiusLabs/fusion-embedding-2-tremor) reads motion, with a [Unitree-G1 head](https://huggingface.co/EximiusLabs/fusion-embedding-2-tremor-g1). Because the base is never modified, adding a sense costs a small trained head and an afternoon of compute rather than a new foundation model.

Ember is strictly additive. Technically it is a 44M-parameter gated adapter pack
that attaches to the frozen decoder behind a thermal-only gate: when the gate is
closed (every non-thermal input), the model's outputs are bit-for-bit identical
to the model without the pack. This is verified, not aspirational; see
Correctness below.

![Ember architecture overview](assets/fe2_ember_overview.png)

## What it does

- Thermal image to text retrieval: R@10 0.785 on a held-out 2,000-caption gallery
  (frozen base: 0.224).
- Thermal zero-shot classification is preserved: LLVIP person/background 94.3
  (calibrated ensemble harness; frozen base reference 95.4).
- Cross-domain thermal-to-visible retrieval: R@10 0.348 on LLVIP registered pairs,
  above the frozen baseline 0.165, without training on any LLVIP data.
- Text, RGB image, video, and audio embeddings unchanged, bit-for-bit.

## Usage

Ember loads as an adapter pack through the multi-gate adapter registry in the
[fusion-embedding GitHub repository](https://github.com/Eximius-Labs/fusion-embedding)
(`fusion_embedding/adapters.py`). Thermal images are single-channel; replicate to
three channels and encode through the ordinary image path with the thermal scope open.

```python
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from fusion_embedding.adapters import AdapterPacks
from inference import FusionEmbedder          # fe2_release/inference.py (GitHub repo)

emb = FusionEmbedder.from_pretrained(
    "EximiusLabs/fusion-embedding-2-2b-preview", revision="v0.2-preview", device="cuda")

packs = AdapterPacks()
adapters, gate = packs.add_pack("thermal", emb.model.base_lm, 2048, rank=384)
adapters.load_state_dict(load_file(hf_hub_download(
    "EximiusLabs/fusion-embedding-2-ember", "model.safetensors")))
packs.to("cuda")

# thermal encode: thermal readout template, thermal scope open (the scope
# spans forward and backward)
import torch.nn.functional as F
text = ("<|im_start|>system\nRepresent this thermal infrared image.<|im_end|>\n"
        "<|im_start|>user\n<|vision_start|><|image_pad|><|vision_end|><|im_end|>\n"
        "<|im_start|>assistant\n")
inp = emb.proc(text=[text], images=[thermal_image_3ch], return_tensors="pt").to("cuda")
with torch.no_grad(), packs.scope("thermal"):
    h = emb.full(**inp).last_hidden_state
thermal_vec = F.normalize(h[0, inp["attention_mask"][0].sum() - 1].float(), dim=-1)

# everything else: leave the scope closed; outputs equal the pack-free model exactly
text_vec = emb.embed_text("a person crossing a dark road")
audio_vec = emb.embed_audio(wav, sr=16000)    # audio pack co-loaded, unaffected
```

The pack attaches equally to the raw base (`AutoModel.from_pretrained("Qwen/Qwen3-VL-Embedding-2B")`,
attach at `.language_model`), which is the exact configuration it was trained in;
both loading paths are verified bitwise in the release smoke. The readout is the
standard fusion-embedding protocol: chat template, last non-pad token pooling,
L2 normalization (see `config.json` for the exact templates and the trained
temperature).

## Evaluation

Training: contrastive thermal-to-caption alignment on IR-TD (61,320 pairs after
FLIR exclusion and eval dedup), 3,900 steps, batch 16, 1,024 bank negatives,
frozen base, bf16 base precision with fp32 adapters. Three seeds; seed 2 shipped.

| release run (61K corpus) | holdout t2t R@10 | delta vs frozen | LLVIP-ZS | LLVIP twin R@10 |
|---|---|---|---|---|
| frozen base | 0.224 | - | 95.4 | 0.165 |
| seed 1 | 0.783 | +0.560 | 94.1 | 0.333 |
| **seed 2 (shipped)** | **0.785** | **+0.561** | **94.3** | **0.348** |
| seed 3 | 0.777 | +0.554 | 91.8 | 0.341 |

Seed 2 holdout detail: R@1 0.412, R@5 0.692, R@10 0.785 over a 2,000-item gallery.

Text to thermal retrieval on the release holdout (queries shortened for display;
retrieval used the full captions):

![Ember retrieval gallery](assets/fe2_ember_retrieval_gallery.png)

A caption-style ablation on the pre-exclusion corpus (82K pairs, 5,000 steps)
found that caption richness is a generalization lever, not just an in-domain fit
lever: training on full descriptive captions reached holdout R@10 0.843 and LLVIP
twin 0.343, while first-sentence captions reached 0.594 and collapsed cross-domain
transfer to 0.089, below the frozen baseline. Ember ships the full-caption arm.

Domain note: IR-TD spans 63 source collections but is still a finite domain mix.
The LLVIP numbers above are cross-domain signal (night pedestrian scenes never
seen in training), not a claim of parity with in-domain retrieval. Expect the gap
to vary with distance from the training domains.

## Correctness

The bit-for-bit preservation claim is tested at three levels:

1. Unit suite (GitHub repo, `tests/test_thermal_adapters.py`): closed-gate
   forwards equal the base exactly; gradients reach only the open pack; the gate
   must span forward and backward under gradient checkpointing.
2. Release checks with the trained weights: RGB-image and text forwards
   bit-for-bit equal to the pack-free base with the thermal gate closed, per seed.
3. Composability matrix (audio pack + thermal pack co-loaded on the same frozen
   decoder): audio through the registry vs the shipped single-gate path, audio
   co-loaded vs audio-only, thermal co-loaded vs thermal-only, and text / RGB /
   video vs the raw base, all bitwise; retrieval scores identical under co-load.

Mixed inputs that would open two gates in a single forward are outside the
guarantee and are not tested.

## Provenance

- Training corpus: IR-TD early access (IRGPT, ICCV 2025,
  [arXiv:2507.14449](https://arxiv.org/abs/2507.14449),
  [repository](https://github.com/WheatCao/ICCV2025-IRGPT)); 84,284 real thermal
  images with LLM-generated descriptive captions; academic research use only.
- FLIR exclusion: IR-TD includes FLIR-derived sources whose terms restrict
  redistribution of trained weights. All 20,964 images matching the FLIR capture
  signature (640x512) were excluded from training and holdout. This is a
  size-based heuristic, not an author-provided source mapping; the exclusion list
  ships in this repository (`release_strip_640x512.json`, sha256 `65a870df5e4fe0008fd9bacd9fa81bd0c47f919779a28e398fea8d1bcabad09f`).
- LLVIP is used for evaluation only (zero-shot gate and cross-domain retrieval).
- Eval hygiene: perceptual-hash dedup between the training set and the LLVIP test
  set found 0 collisions (hamming distance <= 4).

## Deploy on RunPod

One-click from the [RunPod Hub](https://www.runpod.io/console/hub/Eximius-Labs/ember).

```bash
curl -s https://api.runpod.ai/v2/<ENDPOINT_ID>/runsync \n  -H "Authorization: Bearer $RUNPOD_API_KEY" \n  -H "Content-Type: application/json" \n  -d '{"input": {"thermal": "<https url | data-uri | base64>"}}'
```

Use `text` instead of `thermal` to embed a caption. Returns 2048-d vectors, thermal and text in one space.

## Engram

This pack is one of the modalities [Engram](https://github.com/Eximius-Labs/engram) searches. Engram is
the open cross-modal memory layer for physical AI: it indexes a robot's video, audio, and motion into
one embedding space and answers questions about it in plain language, including temporal reasoning that
retrieval alone cannot.

```bash
pip install engram-robomem
```

Repo: https://github.com/Eximius-Labs/engram  &middot;  PyPI: https://pypi.org/project/engram-robomem  &middot;  Playground: https://www.eximiuslabs.com/playground

## License

The Ember weights in this repository are released under CC-BY-NC-4.0 for research
use, reflecting the academic-use terms of the training corpus. The core
fusion-embedding-2 model is a separate artifact under its own license; this pack
is optional and separable, and does not modify the core model's weights.