RynnLAM β motion-grounded latent action model
A Latent Action Model that compresses the dynamics between two RGB frames into a compact, embodiment-agnostic latent action. Trained self-supervised to predict future visual features, regress dynamic 3D flow, and decouple camera motion from ego motion, so the latent describes physical motion rather than 2D appearance change.
This is the checkpoint used to densely label the Stage-1 pretraining corpus behind RynnVLA-Latent. It is published so that the latent actions in a downstream VLA run are reproducible from a fixed, verifiable labeler rather than from a description of one.
What it produces
The default representation is ktoken_zcam, 608 floats per frame pair:
| slice | dims | content |
|---|---|---|
k_tokens.flatten(1) |
512 | 8 K-tokens x 64, from the K-token compressor |
z |
64 | latent action |
camera_pose_latent |
32 | camera ego-motion, decoupled from scene motion |
Six other representations are available from the same forward pass: ktoken, z, zcam,
hints_pool, full, encoder.
Architecture
397.3 M parameters, 684 tensors, float32.
| module | tensors | params | role |
|---|---|---|---|
encoder |
407 | 304.4 M | DA3-Large / DINOv2-family backbone over frame pairs |
flow_decoder |
104 | 40.7 M | dynamic 3D flow regression |
latent_encoder |
99 | 20.3 M | self-attention latent encoder |
hint_compressor |
9 | 16.9 M | compresses motion hints into the 8 K-tokens |
recon_decoder_ktoken |
44 | 14.7 M | K-token reconstruction decoder |
cam_dec / cam_dec_adversarial |
20 | 0.2 M | camera-pose decoder |
Key config: model_version=v5, flow_decoder_version=v5, latent_encoder_type=selfattn,
use_k_token_recon=True, k_token_source=features, num_k_tokens=8,
k_token_bottleneck_dim=64, latent_dim=64, camera_pose_latent_dim=32,
motion_hint_dim=128, embed_dim=1024, patch_size=14, target_hw=[238, 322].
The artifact is self-contained. encoder_checkpoint_path is None and the 407 encoder
tensors are inside the file, so loading it downloads nothing and needs no separate encoder
weights. The loader pins encoder_finetune_mode="freeze" because there is nothing left to
initialize.
Usage
Install RynnVLA-Latent first; this
checkpoint is loaded by its rynnlam package.
Label a corpus from a metadata manifest:
python scripts/label_latent.py \
--checkpoint RynnLAM.pt \
--metadata <manifest.json> \
--output-dir <out> \
--device cuda --precision bf16 --representation ktoken_zcam \
--gap 4 --pair-stride 4 --normalization imagenet --bucket auto
Each episode yields a latent.npz with latent_action float16 [N, 608], pair_indices
int32 [N, 2], and a meta JSON string recording schema_version, code_shape, and the
checkpoint_sha256 of the labeler that produced it. Those gap/pair_stride/normalization/
bucket values are the ones the Stage-1 corpus was labeled with; changing them produces a
different, incompatible corpus.
Or encode frame pairs directly:
from rynnlam.inference import RynnLAMEncoder
enc = RynnLAMEncoder("RynnLAM.pt", device="cuda", precision="bf16")
latents = enc(images, representation="ktoken_zcam") # [batch, 608]
images is [batch, 2, height, width, 3], float32 in [0, 1], RGB. Along the pair axis,
index 0 is the earlier frame t and index 1 is the later frame t + gap; the latent describes
the motion from 0 to 1. Out-of-range or non-finite input raises rather than being clamped.
Frames are resized to target_hw = [238, 322] internally.
Loading never requires --trust-checkpoint: the file reads under torch.load(..., weights_only=True) and then strict-loads with 0 trainable parameters.
Provenance and integrity
sha256 (this file) 0ab843b0b308357db48b50d3b27ce5c39eac134fa9d5f7fe5172c3320aded424
sha256 (internal) d066f871fe6e05525ab7ad48c706cd3f564cc8546f9a67af2d597426d560ed56
size 1,589,652,281 bytes
Verify a download:
sha256sum RynnLAM.pt
Both digests are published deliberately. The second is the hash of the file as it was used
internally, and it is the value written into the meta of every latent .npz in the Stage-1
corpus. This file has the same weights β all 684 tensors are bit-identical β and differs only
in the config block, which was sanitized for release. So a corpus labeled internally records the
internal digest, and a corpus you label with this file records the public one. If you are
checking whether your latents came from these weights, compare against the internal digest; if
you are checking whether your download is intact, compare against the public one.
The equivalence is not asserted, it was measured: labeling the same videos with each file
produced float16 [24, 608] arrays that are bit-identical, maxabsdiff = 0.0.
What was removed from the config, and why
The internal checkpoint carried 143 config entries; this one carries 134. The weights were not touched. Removed:
| key | why |
|---|---|
encoder_checkpoint_path |
absolute path on an internal mount; also dead at load time, the encoder is in the file |
output_root, resume |
internal training-run locations |
safetensors_root |
internal location of the LAM training corpus |
dataset_sampling_weights, dataset_sampling_temperature, dataset_role_sampling_weights |
the training corpus composition and its sampling recipe are not part of this release |
val_scene_id |
an identifier into that corpus |
wandb_project, oss_mode |
internal experiment-tracking and storage-backend flags |
None of these reach the model: rynnlam.model.build_model filters the config by the RynnLAM
constructor signature, so training-only keys are dropped at load time regardless. They were
removed because they disclosed internal infrastructure, not because they affected behaviour β
which is why the bit-identical-latents check above is the meaningful guarantee.
The remaining 134 entries are architecture and loss hyperparameters, and are what
build_model consumes.
Requirements
Verified against torch 2.7.1+cu126 on an NVIDIA H20, and under weights_only=True on CPU.
precision="bf16" needs a GPU with bfloat16 support; use precision="fp32" otherwise. The
encoder is DA3-Large-class, so budget roughly 1.6 GB of disk and a few GB of host memory to
load, plus GPU memory for the pair forward pass.
Limitations
- The latent is a relative action between two frames. It has no metric scale and no absolute pose; converting it to a robot action space is what RynnVLA-Latent's Stage-2 post-training is for.
target_hw = [238, 322]is fixed by training. Input at other aspect ratios is bucketed and resized by the labeler, which is a distribution shift the model was not trained to handle.- No optimizer, scheduler, or EMA state is included. This is an inference artifact and cannot resume the training run that produced it.
License
Apache-2.0. Note that the RynnVLA-Latent source license covers code, not weights or datasets; this file is released separately under the same license. Training-data provenance is not redistributed here β obtaining this checkpoint grants no rights to any corpus it was trained on.