RynnLAM β€” motion-grounded latent action model

A Latent Action Model that compresses the dynamics between two RGB frames into a compact, embodiment-agnostic latent action. Trained self-supervised to predict future visual features, regress dynamic 3D flow, and decouple camera motion from ego motion, so the latent describes physical motion rather than 2D appearance change.

This is the checkpoint used to densely label the Stage-1 pretraining corpus behind RynnVLA-Latent. It is published so that the latent actions in a downstream VLA run are reproducible from a fixed, verifiable labeler rather than from a description of one.

What it produces

The default representation is ktoken_zcam, 608 floats per frame pair:

slice dims content
k_tokens.flatten(1) 512 8 K-tokens x 64, from the K-token compressor
z 64 latent action
camera_pose_latent 32 camera ego-motion, decoupled from scene motion

Six other representations are available from the same forward pass: ktoken, z, zcam, hints_pool, full, encoder.

Architecture

397.3 M parameters, 684 tensors, float32.

module tensors params role
encoder 407 304.4 M DA3-Large / DINOv2-family backbone over frame pairs
flow_decoder 104 40.7 M dynamic 3D flow regression
latent_encoder 99 20.3 M self-attention latent encoder
hint_compressor 9 16.9 M compresses motion hints into the 8 K-tokens
recon_decoder_ktoken 44 14.7 M K-token reconstruction decoder
cam_dec / cam_dec_adversarial 20 0.2 M camera-pose decoder

Key config: model_version=v5, flow_decoder_version=v5, latent_encoder_type=selfattn, use_k_token_recon=True, k_token_source=features, num_k_tokens=8, k_token_bottleneck_dim=64, latent_dim=64, camera_pose_latent_dim=32, motion_hint_dim=128, embed_dim=1024, patch_size=14, target_hw=[238, 322].

The artifact is self-contained. encoder_checkpoint_path is None and the 407 encoder tensors are inside the file, so loading it downloads nothing and needs no separate encoder weights. The loader pins encoder_finetune_mode="freeze" because there is nothing left to initialize.

Usage

Install RynnVLA-Latent first; this checkpoint is loaded by its rynnlam package.

Label a corpus from a metadata manifest:

python scripts/label_latent.py \
    --checkpoint RynnLAM.pt \
    --metadata <manifest.json> \
    --output-dir <out> \
    --device cuda --precision bf16 --representation ktoken_zcam \
    --gap 4 --pair-stride 4 --normalization imagenet --bucket auto

Each episode yields a latent.npz with latent_action float16 [N, 608], pair_indices int32 [N, 2], and a meta JSON string recording schema_version, code_shape, and the checkpoint_sha256 of the labeler that produced it. Those gap/pair_stride/normalization/ bucket values are the ones the Stage-1 corpus was labeled with; changing them produces a different, incompatible corpus.

Or encode frame pairs directly:

from rynnlam.inference import RynnLAMEncoder

enc = RynnLAMEncoder("RynnLAM.pt", device="cuda", precision="bf16")
latents = enc(images, representation="ktoken_zcam")   # [batch, 608]

images is [batch, 2, height, width, 3], float32 in [0, 1], RGB. Along the pair axis, index 0 is the earlier frame t and index 1 is the later frame t + gap; the latent describes the motion from 0 to 1. Out-of-range or non-finite input raises rather than being clamped. Frames are resized to target_hw = [238, 322] internally.

Loading never requires --trust-checkpoint: the file reads under torch.load(..., weights_only=True) and then strict-loads with 0 trainable parameters.

Provenance and integrity

sha256 (this file)  0ab843b0b308357db48b50d3b27ce5c39eac134fa9d5f7fe5172c3320aded424
sha256 (internal)   d066f871fe6e05525ab7ad48c706cd3f564cc8546f9a67af2d597426d560ed56
size                1,589,652,281 bytes

Verify a download:

sha256sum RynnLAM.pt

Both digests are published deliberately. The second is the hash of the file as it was used internally, and it is the value written into the meta of every latent .npz in the Stage-1 corpus. This file has the same weights β€” all 684 tensors are bit-identical β€” and differs only in the config block, which was sanitized for release. So a corpus labeled internally records the internal digest, and a corpus you label with this file records the public one. If you are checking whether your latents came from these weights, compare against the internal digest; if you are checking whether your download is intact, compare against the public one.

The equivalence is not asserted, it was measured: labeling the same videos with each file produced float16 [24, 608] arrays that are bit-identical, maxabsdiff = 0.0.

What was removed from the config, and why

The internal checkpoint carried 143 config entries; this one carries 134. The weights were not touched. Removed:

key why
encoder_checkpoint_path absolute path on an internal mount; also dead at load time, the encoder is in the file
output_root, resume internal training-run locations
safetensors_root internal location of the LAM training corpus
dataset_sampling_weights, dataset_sampling_temperature, dataset_role_sampling_weights the training corpus composition and its sampling recipe are not part of this release
val_scene_id an identifier into that corpus
wandb_project, oss_mode internal experiment-tracking and storage-backend flags

None of these reach the model: rynnlam.model.build_model filters the config by the RynnLAM constructor signature, so training-only keys are dropped at load time regardless. They were removed because they disclosed internal infrastructure, not because they affected behaviour β€” which is why the bit-identical-latents check above is the meaningful guarantee.

The remaining 134 entries are architecture and loss hyperparameters, and are what build_model consumes.

Requirements

Verified against torch 2.7.1+cu126 on an NVIDIA H20, and under weights_only=True on CPU. precision="bf16" needs a GPU with bfloat16 support; use precision="fp32" otherwise. The encoder is DA3-Large-class, so budget roughly 1.6 GB of disk and a few GB of host memory to load, plus GPU memory for the pair forward pass.

Limitations

  • The latent is a relative action between two frames. It has no metric scale and no absolute pose; converting it to a robot action space is what RynnVLA-Latent's Stage-2 post-training is for.
  • target_hw = [238, 322] is fixed by training. Input at other aspect ratios is bucketed and resized by the labeler, which is a distribution shift the model was not trained to handle.
  • No optimizer, scheduler, or EMA state is included. This is an inference artifact and cannot resume the training run that produced it.

License

Apache-2.0. Note that the RynnVLA-Latent source license covers code, not weights or datasets; this file is released separately under the same license. Training-data provenance is not redistributed here β€” obtaining this checkpoint grants no rights to any corpus it was trained on.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Collection including Alibaba-DAMO-Academy/RynnLAM