Papers
arxiv:2610.04211

CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model

Published on Oct 3
· Submitted by
mingyi
on Oct 6
Authors:
,
,
,

Abstract

Skeletal motion is stored as every joint's transform at every frame, yet most of it is implied by the body rather than by what the motion is about. Compression is one way to ask what a motion must still say once the body is known, and a production codec must answer it for any skeleton with a stated error bound. Our earlier codec, CurveCodec, matched the mean error of ACL, the production library of modern game engines, with a learned prior over sparse anchors, but not ACL's worst case, and it counted its payload as floats rather than bits. Here we ask where the redundancy of skeletal motion lies and which part of a codec a learned model should take over. Measurements give three answers. At production precision the largest saving comes from predicting each quantized curve from its own past, the second from choosing per joint, in closed loop through the hierarchy, which samples not to code. On the gaps such an encoder leaves, a nearest-neighbour oracle over millions of training samples is no better than linear interpolation, and no learned in-betweener we tried paid for itself. What a network does learn is the distribution of the residuals the codec must send. CurveCodec 2 codes every sub-track as a curve in the log map, quantized in closed loop and thinned to rate-distortion-selected keys, with residuals entropy-coded under a small learned model whose integer inference is bit-exact across platforms. Two contracts are verified on every decoded clip: ACL's own worst case per joint within a stated tolerance, or ACL's mean error per clip. On a held-out test side of 4,472 clips from 33 datasets, CurveCodec 2 needs 0.37x ACL's bytes at ACL's default precision of 0.01 cm under the worst-case contract and 0.22x at 0.1 cm under the mean contract, decodes on one CPU core, and transfers without retraining to a species absent from training. Project page: https://rubbly.cn/publications/curvecodec/

Community

CurveCodec v0.2.0

Skeleton-agnostic animation compression with a learned entropy model

SIGGRAPH Asia 2026

[arXiv] [homepage] [code] [demo]

Mingyi Shi1 · Huancheng Lin1 · Xuelin Chen2,* · Taku Komura1,*

1The University of Hong Kong · 2Adobe Research · *Co-corresponding authors

Grid of twelve characters in motion (a walking dog, turtle, chicken and leopard, a flying buzzard, bat and pteranodon, a swimming shark, a striking anaconda, a Unitree Go2 robot driven by dog motion capture, and two humans punching and kicking), each animated from its CurveCodec stream at 0.1 cm and labelled with its size before and after compression, 56x to 154x smaller than the original float32 clip

Original float32 clip → CurveCodec at 0.1 cm (mean error 0.1–0.4 mm). Across our held-out test set, the average at this precision is about 100× smaller than float32.

Why study compression at first?

  • Motion data is highly redundant. A clip is stored as every joint's local transform at every frame. But when a person performs an action, they rarely attend to how each joint gets from one place to the next: the trajectories are largely produced by a strong prior, the body itself, rather than being what the motion is about. Spending more effort and computation on generating these curves does little for understanding behaviour or action.

  • Compression is a principled way to learn what matters. Embodied intelligence works the same way: an agent decides what to do, and its body, its morphology and its dynamics, decide most of how. A model of motion intelligence should spend its capacity on the decisions, not on the kinematics the embodiment already implies. Compression separates the two: what a good codec must still send is the decision; what it can drop is the body.

  • We want a representation general enough to model the dynamics of motion. If the body is only the prior, the representation should not be tied to one body: it has to generalize across motions and across skeletons. Today's hierarchical skeleton representations bake a specific topology into the data, so every rig needs its own model and little of what is learned on one body carries over to another; they do not scale. CurveCodec 2 treats one model serves any rig, and transfers without retraining to a species it has never seen.

What the library does

Input Any skeleton, any rig: per-joint rotations and translations (BVH). One model for humans, hands, animals, robots and game rigs, with no per-rig training.
Precision 0.01 to 1 cm, set per clip. Every decoded clip is verified against its error bound, and a clip that misses is re-encoded with tighter margins.
Entropy model A 108 K-parameter causal transformer (870 KB). Its integer inference is bit-exact across platforms, and it changes only the bits, never the decoded motion.
Footprint About 2 bits per joint sample at 0.1 cm. Our training corpus of about 900 h of motion fits in a few GB.
Decode 2.1 s per million joint samples on one CPU core, or 0.76 s on four threads. An optional CUDA path is available.
Training The codec dumps its own residual streams, so you can retrain the model on your data. The released model took about 1.5 h on one RTX 4090.

Compared with ACL

Bar chart of bits per joint sample on our held-out test set: ACL vs CurveCodec at precisions 0.01 to 1 cm; CurveCodec needs 0.37x, 0.27x, 0.22x, 0.14x and 0.07x of ACL's bytes

On our held-out test set of about 4,500 clips (about 20 h), every clip keeps ACL's mean error at the same p,
with per-joint caps and an anti-pop guard. Each clip that misses this contract is counted with its fallback stream.
On that set, CurveCodec needs 0.37× ACL's bytes at 0.01 cm, 0.27× at 0.05 cm, 0.22× at 0.1 cm, 0.14×
at 0.3 cm and 0.07× at 1 cm.

A flying dragon rendered three times: the original, ACL 2.1 at 0.01 cm (169.8 KB) and CurveCodec (79.0 KB), both with 0.028 mm mean error

The two codecs serve different goals, and CurveCodec is not a replacement for ACL. ACL is built for runtime: it is
stateless, keeps the clip compressed in memory and decompresses only the poses a frame needs, and samples any pose at
random with minimal memory traffic. CurveCodec corrects
its encoder with error feedback through forward kinematics. Its stream is entropy-coded and decoded once per clip. It
is meant for storing and streaming whole clips, and for asking what motion data really contains.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.04211
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.04211 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.04211 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.04211 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.