Reka Inverse Dynamics Model (RIDM)
Two models predict camera motion (W/A/S/D/Shift, yaw, pitch) from video. Both trained on game renders.
| flow | pixel | |
|---|---|---|
| Input | Optical flow (RAFT-small) | Raw frames |
| Parameters | 1,795,337 trained, plus 990,162 in frozen RAFT-small (2,785,499 in total) | 9,836,063 |
| Output | One prediction per 17-frame window | One prediction per frame |
| Real-video turn accuracy | 91.5 | 51.1 |
| Gaming footage turn accuracy | 84.9 | 70.9 |
| Dependencies | torch, torchvision, av, opencv, safetensors | torch, av, opencv, safetensors |
Repository contents:
Reka-Inverse-Dynamics-Model/
README.md LICENSE NOTICE requirements.txt results.json
flow/ model.safetensors config.json config_walking.json inference.py README.md
pixel/ model.safetensors config.json inference.py README.md
assets/ demo and example prediction videos
examples/ 10 s sample clip and expected_flow_output.json
tests/ test_smoke.py
Demo
The flow model on a walking tour of Bristol, with frame gap 12. Blue bars show the five keys. The dials show yaw and pitch in degrees.
More clips are in assets/: walk-empty-street.mp4, turn-right.mp4 and the failure case failure-walk-key-with-no-walk.mp4.
In the failure clip the camera stands still and pans. The model reports the walk key (Shift) at 0.72 to 0.98 in every window.
Which model to use
The flow model has better generalization capabilities to real videos. The pixel model might have a high potential for gaming footage especially if trained further.
Use
pip install -r requirements.txt
hf download RekaAI/Reka-Inverse-Dynamics-Model --local-dir ridm
python ridm/flow/inference.py clip.mp4 --stride 17
- The script prints JSON. Each window has
key_probabilities,keys_pressed,yaw_degandpitch_deg. - Default settings match game footage. For walking video, run
python ridm/flow/inference.py clip.mp4 --config ridm/flow/config_walking.json --stride 49. - To check your install, run
pytest tests/test_smoke.pyfrom the repository root. It runs the sample clip and compares withexamples/expected_flow_output.json. - Each folder runs alone. See
flow/README.mdandpixel/README.mdfor details. - Tested with Python 3.12 and the versions in
requirements.txt. Newer versions are untested.
Status: preliminary release. The scores marked † are unconfirmed.
Training data
- Source: 4,071 recorded gaming footage plays, captured at 1,280 x 720 and 48 frames per second.
- Labels: the game engine gives the keys, the mouse movement and the camera angles for every frame. Nobody annotates.
- Flow model: 7,071,037 frame pairs. Pixel model: about 104,000 gameplay windows.
- Both models trained on one NVIDIA L4 GPU.
License
- The weights and the code in this repository use the Apache License 2.0. See
LICENSE. - RAFT-small weights use the BSD-3 license. They come from torchvision and are not in this repository. See
NOTICE. - The models trained on gameplay captures. This repository holds no game assets, no gameplay clips and no training data.
- The videos in
assets/andexamples/are not Apache 2.0. They come from walking tours by POPtravel on Wikimedia Commons, licensed CC BY 3.0, and carry the model overlay or are shortened excerpts. Sources: Bristol, Ingolstadt, Bamberg.
Contact and citation
- Questions and problems: open a discussion on this repository.
Citation
@misc{RIDM,
title = {Reka Inverse Dynamics Model for Interactive World Models},
author = {Reka AI},
year = {2026},
url = {https://reka.ai/labs/research/reka-inverse-dynamics-model-for-interactive-world-models}
}
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Evaluation results
- Accuracy (percent) on Real video, 78 clipsself-reported85.900
- Accuracy (percent) on Real video, 47 turn clipsself-reported91.500
- Accuracy (percent)self-reported84.500
- Accuracy (percent)self-reported84.900