jvogan
Improve Hugging Face discovery metadata
6e715ed
|
Raw
History Blame Contribute Delete
14.5 kB
---
license: other
license_name: qwen-research-license
license_link: https://huggingface.co/JacobMolBio/vibethinker-3b-jlens-model/blob/main/LICENSES/QWEN-RESEARCH.txt
thumbnail: https://huggingface.co/JacobMolBio/vibethinker-3b-jlens-model/resolve/main/assets/jlens-model-banner.png
tags:
- vibethinker-3b
- jacobian-lens
- mechanistic-interpretability
- interpretability
- llm-interpretability
- activation-analysis
- model-visualization
- qwen
- qwen2
- safetensors
---
![Stacked VibeThinker-3B and J Lens title beside a folded black-layer sculpture divided by a cyan plane. The subtitle reads: Watch the model’s answer take shape, layer by layer. Weights, traces, code, and a viewer.](assets/jlens-model-banner.png)
# VibeThinker-3B J Lens
The J Lens is a Jacobian lens fitted to
[`WeiboAI/VibeThinker-3B`](https://huggingface.co/WeiboAI/VibeThinker-3B/tree/77bd2cced09193c8b9a59a32bd8577bbd1f3e01c).
Pick one of the 18 fitted source layers and a token position, and the lens
decodes that residual-stream activation into a ranked list of vocabulary
tokens. Following one position across layers shows how the decoded ranking
changes on the way to the model's final output.
The lens contains 18 matrices, one for each even-numbered source layer from 0
through 34. Each matrix maps its source-layer residual-stream activation into
layer-35 coordinates. VibeThinker-3B's final normalization and vocabulary
projection produce the ranked token scores.
This repository contains the fitted lens in two Safetensors precisions.
`model.safetensors` is the FP16 lens used to capture the released traces.
`evaluation.safetensors` is the FP32 lens used for the recorded readout
evaluation. Casting each FP32 matrix to FP16 reproduces the FP16 file exactly;
both are included so every published result stays paired with the tensor
values that produced it.
The companion trace repository contains six saved traces. Each trace pairs a
prompt and its generated response with the top 12 decoded tokens at each
captured position for the 18 source layers and the final model layer.
## Choose a path
| Goal | Use |
|---|---|
| Browse saved traces | Open the [static viewer](https://jvogan.github.io/vibethinker-3b-jlens). It reads released trace JSON in the browser; it does not download model weights or run inference. |
| Run a new prompt locally | Use `scripts/render_slice.py` in the [source repository](https://github.com/jvogan/vibethinker-3b-jlens). It uses the pinned VibeThinker-3B base-model weights and `model.safetensors`, the FP16 J Lens. |
| Load the released lens in Python | Use `model.safetensors`, the FP16 trace lens. |
| Recalculate the recorded readout statistics | Use `evaluation.safetensors`, the FP32 evaluation lens, with the released rank rows and `scripts/recalculate_readout.py` in the trace repository. |
## Companion repositories
| Release component | Location |
|---|---|
| Source code and Pages source | [https://github.com/jvogan/vibethinker-3b-jlens](https://github.com/jvogan/vibethinker-3b-jlens) |
| Captured trace dataset | [https://huggingface.co/datasets/JacobMolBio/vibethinker-3b-jlens-traces](https://huggingface.co/datasets/JacobMolBio/vibethinker-3b-jlens-traces) |
| Static Pages site | [https://jvogan.github.io/vibethinker-3b-jlens](https://jvogan.github.io/vibethinker-3b-jlens) |
| Hugging Face model repository ID | [JacobMolBio/vibethinker-3b-jlens-model](https://huggingface.co/JacobMolBio/vibethinker-3b-jlens-model) |
## Artifact identity
| Field | Trace artifact | Evaluation artifact |
|---|---|---|
| File | `model.safetensors` | `evaluation.safetensors` |
| SHA-256 | `089d776979408f23e5377539c15aa8025d171633718ccdce709bcd3372e7942c` | `0cc184eb65d273ac8bfee5450a141c8cdf8dd6ec8caa68d7d47260ef621777d1` |
| Size | 150,996,824 bytes | 301,991,904 bytes |
| Dtype | FP16 | FP32 |
| Role | Captured trace readouts | Recorded readout evaluation |
[![Use model.safetensors for FP16 trace work and evaluation.safetensors for the recorded FP32 evaluation.](assets/two-lens-files.png)](assets/two-lens-files.svg)
Both artifacts contain 18 matrices with shape 2,048 × 2,048. Their source
layers are every even layer from 0 through 34, and their target layer is 35.
The lens was fit on 1,000 prompts for model revision
`77bd2cced09193c8b9a59a32bd8577bbd1f3e01c`. The revision binding was
reconstructed from the Hub head after the fit;
[`provenance.json`](provenance.json) records that binding and its limits.
## Files
- `model.safetensors` — the 18 losslessly reserialized FP16 lens matrices.
- `evaluation.safetensors` — the 18 losslessly reserialized FP32 evaluation matrices.
- `lens_config.json` — model binding, layer layout, fit parameters, and artifact ID.
- `tensor_manifest.json` — shape, dtype, size, and raw-byte SHA-256 for every tensor.
- `evaluation_tensor_manifest.json` — the corresponding FP32 tensor manifest.
- `evaluation_provenance.json` — derivation and source-hash bindings.
- `evaluation_compatibility.json` — the FP32-to-FP16 comparison for every layer.
- `evaluation_validation.json` — the exact FP32 conversion checks.
- `provenance.json` — fit, corpus, software, and conversion provenance.
- `validation.json` — exhaustive source-to-safetensors comparison result.
- `evaluation.json` — recorded readout-only evaluation and its scope.
- `SHA256SUMS` — checksums for both Safetensors files.
- `requirements.txt` — pinned packages for loading and validation.
- `scripts/validate_artifact.py` — validates the artifact without pickle.
- `scripts/convert_checkpoint.py` — reproducible checkpoint conversion for a user-provided source file.
- `scripts/convert_evaluation_checkpoint.py` — reproducible FP32 conversion.
- `LICENSES/QWEN-RESEARCH.txt` — complete distribution terms for the artifacts.
- `NOTICE` and `THIRD_PARTY_NOTICES.md` — required attribution and license context.
## Tensor layout
The keys are `J.<source_layer>`:
```text
J.0, J.2, J.4, J.6, J.8, J.10, J.12, J.14, J.16,
J.18, J.20, J.22, J.24, J.26, J.28, J.30, J.32, J.34
```
Each matrix maps the post-transformer-block residual at its source layer into
the layer-35 residual basis. With row-major batches of residual vectors, the
transport used by the reference implementation is:
```python
transported = residual.float() @ J.float().T
logits = unembed(transported)
```
Both Safetensors headers contain the model ID and revision, source and target
layers, width, prompt count, source-checkpoint hash, and key pattern.
`requirements.txt` pins the runtime used to create and validate
`evaluation.safetensors`. `provenance.json` records the earlier runtime used
to convert `model.safetensors`.
## Install and load
The `jlens` Python package comes from the companion source repository. This
repository contains the fitted tensors and the dependencies needed to inspect
and validate them.
From this repository checkout, install both parts:
```bash
export JLENS_CODE_REPO_URL="https://github.com/jvogan/vibethinker-3b-jlens"
git clone "$JLENS_CODE_REPO_URL" ../vibethinker-3b-jlens
python3 -m pip install -r requirements.txt
python3 -m pip install ../vibethinker-3b-jlens
```
Load the local artifact:
```python
from jlens import JacobianLens
lens = JacobianLens.load("model.safetensors")
print(lens.source_layers)
```
Or load an immutable Hugging Face revision, setting `JLENS_MODEL_REVISION`
to the model-repository commit recorded in the source repository's
`release-manifest.json`:
```bash
export JLENS_MODEL_REPO_ID="JacobMolBio/vibethinker-3b-jlens-model"
export JLENS_MODEL_REVISION="<commit from the source repository's release-manifest.json>"
```
```python
import os
from jlens import JacobianLens
lens = JacobianLens.from_pretrained(
os.environ["JLENS_MODEL_REPO_ID"],
filename="model.safetensors",
revision=os.environ["JLENS_MODEL_REVISION"],
)
```
To inspect tensors without the companion package, install `requirements.txt`
and use Safetensors directly:
```python
from safetensors import safe_open
with safe_open("model.safetensors", framework="pt", device="cpu") as handle:
metadata = handle.metadata()
jacobians = {
int(key.split(".", 1)[1]): handle.get_tensor(key).clone()
for key in handle.keys()
}
assert metadata["model_revision"] == (
"77bd2cced09193c8b9a59a32bd8577bbd1f3e01c"
)
assert sorted(jacobians) == list(range(0, 36, 2))
```
Load VibeThinker itself from the upstream repository at the exact bound
revision; this repository does not include VibeThinker or Qwen weights:
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "WeiboAI/VibeThinker-3B"
revision = "77bd2cced09193c8b9a59a32bd8577bbd1f3e01c"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForCausalLM.from_pretrained(
model_id,
revision=revision,
dtype="auto",
)
```
Review and accept the upstream terms before downloading or using the model.
## Viewer and new-prompt compute
The companion Pages site is a static viewer for captured trace JSON. It
parses, filters, and renders those records in the browser; loading a local
JSON file uses the browser's file reader and does not upload the file. The
static site cannot analyze a new prompt. That requires the companion source
package, this lens artifact, the pinned VibeThinker weights, and local Python
compute.
## Fit method
The lens was fit with the Apache-2.0 Anthropic Jacobian Lens reference
implementation at commit
[`581d398613e5602a5af361e1c34d3a92ea82ba8e`](https://github.com/anthropics/jacobian-lens/tree/581d398613e5602a5af361e1c34d3a92ea82ba8e).
For each fitted source layer, the estimator sums cotangents over valid causal
targets at or after a source position, averages over valid source positions,
and gives each prompt equal weight.
- Hook point: post-transformer-block output residual
- Source layers: every even layer from 0 through 34
- Target layer: 35
- Sequence cap: 128 tokens
- Leading positions skipped: 16
- Final position excluded from fitting
- Dimension batch: 8
- Fit runtime dtype: BF16
- Trace artifact dtype: FP16
- Evaluation artifact dtype: FP32
- Fit manifest: 1,000 examples from WikiText-103 raw train, minimum 600
characters
[`provenance.json`](provenance.json) records the fit manifest, a later local
rematerialization of the prompt set, and the hash checks that bind them.
Prompt text is not included in this repository.
## Validation
Validation checked the following for every tensor in both artifacts:
- exact key set;
- exact shape and each artifact's declared FP16 or FP32 dtype;
- exact numerical equality with the recorded source checkpoint;
- exact FP16 or FP32 bit-pattern equality;
- exact raw tensor-byte SHA-256;
- contiguity and full finiteness; and
- a Safetensors header free of private source paths.
All 18 tensors in each artifact passed, and zero values changed. The
FP32-to-FP16 cast also matches `model.safetensors` exactly for all 18 layers.
Recheck both artifacts with:
```bash
python3 scripts/validate_artifact.py
```
The validator verifies both file checksums, cross-file model and artifact
bindings, header metadata, tensor shapes and hashes, license and notice files,
and FP32-to-FP16 compatibility.
## Evaluation
The recorded evaluation measures ranked-token readouts: for each prompt in a
551-item suite, does a known target term rank highly in the decoded tokens?
The paired token-target and shuffled-layer mapping checks passed for the
selected band at layers 24, 26, 28, 30, 32, and 34. The comparison with the
ordinary logit lens did not establish improvement.
[`evaluation.json`](evaluation.json) records this readout-only scope as
`validated_readout_only`.
The 551 items cover association, multihop, multilingual,
order-of-operations, poetry, and typo prompts, each pairing a prompt with one
or more target terms. A deterministic split assigned 377 eligible items to
the test set; the band was selected on development items. The companion trace
repository contains all evaluation inputs, the 50,050 readout rows, the
aggregate metrics, the bootstrap intervals, and the exact method, and its
`scripts/recalculate_readout.py` recomputes every recorded statistic from the
released rows. Regenerating the rows themselves requires the pinned base
model, `evaluation.safetensors`, and a separate evaluator implementation.
The recorded metrics bind to `evaluation.safetensors`
(SHA-256 `0cc184eb65d273ac8bfee5450a141c8cdf8dd6ec8caa68d7d47260ef621777d1`,
from source FP32 checkpoint SHA-256
`8f752032a26a5196c1cb447ef63f01e8a29820ff0c57178dd80c6e26d32b12a9`);
they do not evaluate `model.safetensors`. The evaluation does not measure
general model accuracy or test causal steering or free-generation behavior.
The lens is stride-2, was fit on 1,000 prompts, and is model-revision
specific. Single-cell readouts can be noisy; compare patterns across nearby
layers and positions before interpreting an isolated token.
## License and attribution
The Qwen Research License applies to this distribution and limits use to
non-commercial research and evaluation. Commercial use requires a separate
license from the upstream rights holder. Redistribution must include the
complete [Qwen Research License](LICENSES/QWEN-RESEARCH.txt) and retain the
required notice in [`NOTICE`](NOTICE). Built with Qwen. No VibeThinker or
Qwen base-model weights are redistributed here.
The pinned VibeThinker model-card metadata declares `license: mit` and names
`Qwen/Qwen2.5-Coder-3B` as its base model; the pinned tree contains no
standalone license file.
[`LICENSES/VIBETHINKER-LICENSE-NOTE.txt`](LICENSES/VIBETHINKER-LICENSE-NOTE.txt)
records that metadata and the pinned revision URL.
The fitting implementation is Anthropic's Apache-2.0
[`jacobian-lens`](https://github.com/anthropics/jacobian-lens). WikiText-103
is available under CC BY-SA 4.0; no source text is redistributed here. See
[`THIRD_PARTY_NOTICES.md`](THIRD_PARTY_NOTICES.md) for details and links.
## References
- Anthropic, [*Verbalizable Representations Form a Global Workspace in Language Models*](https://transformer-circuits.pub/2026/workspace/index.html)
- Anthropic, [`jacobian-lens`](https://github.com/anthropics/jacobian-lens)
- [`WeiboAI/VibeThinker-3B`](https://huggingface.co/WeiboAI/VibeThinker-3B/tree/77bd2cced09193c8b9a59a32bd8577bbd1f3e01c)
- [`Qwen/Qwen2.5-Coder-3B`](https://huggingface.co/Qwen/Qwen2.5-Coder-3B)
- [`Salesforce/wikitext`](https://huggingface.co/datasets/Salesforce/wikitext)