--- license: apache-2.0 library_name: pytorch pipeline_tag: video-classification tags: - inverse-dynamics - camera-motion - optical-flow model-index: - name: RIDM flow model results: - task: type: video-classification name: Forward, turn left or turn right dataset: name: Real video, 78 clips type: real-video metrics: - type: accuracy value: 85.9 name: Accuracy (percent) - task: type: video-classification name: Turn direction dataset: name: Real video, 47 turn clips type: real-video metrics: - type: accuracy value: 91.5 name: Accuracy (percent) - task: type: video-classification name: Forward, turn left or turn right metrics: - type: accuracy value: 84.5 name: Accuracy (percent) - task: type: video-classification name: Turn direction metrics: - type: accuracy value: 84.9 name: Accuracy (percent) --- # Reka Inverse Dynamics Model (RIDM) Two models predict camera motion (W/A/S/D/Shift, yaw, pitch) from video. Both trained on game renders. | | [flow](flow/) | [pixel](pixel/) | |---|---|---| | Input | Optical flow (RAFT-small) | Raw frames | | Parameters | 1,795,337 trained, plus 990,162 in frozen RAFT-small (2,785,499 in total) | 9,836,063 | | Output | One prediction per 17-frame window | One prediction per frame | | Real-video turn accuracy | 91.5 | 51.1 | | Gaming footage turn accuracy | 84.9 | 70.9 | | Dependencies | torch, torchvision, av, opencv, safetensors | torch, av, opencv, safetensors | Repository contents: ```text Reka-Inverse-Dynamics-Model/ README.md LICENSE NOTICE requirements.txt results.json flow/ model.safetensors config.json config_walking.json inference.py README.md pixel/ model.safetensors config.json inference.py README.md assets/ demo and example prediction videos examples/ 10 s sample clip and expected_flow_output.json tests/ test_smoke.py ``` ## Demo The flow model on a walking tour of Bristol, with frame gap 12. Blue bars show the five keys. The dials show yaw and pitch in degrees. More clips are in `assets/`: `walk-empty-street.mp4`, `turn-right.mp4` and the failure case `failure-walk-key-with-no-walk.mp4`. In the failure clip the camera stands still and pans. The model reports the walk key (Shift) at 0.72 to 0.98 in every window. ## Which model to use The flow model has better generalization capabilities to real videos. The pixel model might have a high potential for gaming footage especially if trained further. ## Use ```bash pip install -r requirements.txt hf download RekaAI/Reka-Inverse-Dynamics-Model --local-dir ridm python ridm/flow/inference.py clip.mp4 --stride 17 ``` - The script prints JSON. Each window has `key_probabilities`, `keys_pressed`, `yaw_deg` and `pitch_deg`. - Default settings match game footage. For walking video, run `python ridm/flow/inference.py clip.mp4 --config ridm/flow/config_walking.json --stride 49`. - To check your install, run `pytest tests/test_smoke.py` from the repository root. It runs the sample clip and compares with `examples/expected_flow_output.json`. - Each folder runs alone. See `flow/README.md` and `pixel/README.md` for details. - Tested with Python 3.12 and the versions in `requirements.txt`. Newer versions are untested. Status: preliminary release. The scores marked † are unconfirmed. ## Training data - Source: 4,071 recorded gaming footage plays, captured at 1,280 x 720 and 48 frames per second. - Labels: the game engine gives the keys, the mouse movement and the camera angles for every frame. Nobody annotates. - Flow model: 7,071,037 frame pairs. Pixel model: about 104,000 gameplay windows. - Both models trained on one NVIDIA L4 GPU. ## License - The weights and the code in this repository use the Apache License 2.0. See `LICENSE`. - RAFT-small weights use the BSD-3 license. They come from torchvision and are not in this repository. See `NOTICE`. - The models trained on gameplay captures. This repository holds no game assets, no gameplay clips and no training data. - The videos in `assets/` and `examples/` are not Apache 2.0. They come from walking tours by POPtravel on Wikimedia Commons, licensed CC BY 3.0, and carry the model overlay or are shortened excerpts. Sources: [Bristol](https://commons.wikimedia.org/wiki/File:Walking_in_BRISTOL_-_UK_-_4K_60fps_(UHD).webm), [Ingolstadt](https://commons.wikimedia.org/wiki/File:Walking_in_INGOLSTADT_-_Germany_-_4K_60fps_(UHD).webm), [Bamberg](https://commons.wikimedia.org/wiki/File:Walking_in_BAMBERG_-_Germany_-_4K_60fps_(UHD).webm). ## Contact and citation - Questions and problems: open a discussion on this repository. ## Citation ```bibtex @misc{RIDM, title = {Reka Inverse Dynamics Model for Interactive World Models}, author = {Reka AI}, year = {2026}, url = {https://reka.ai/labs/research/reka-inverse-dynamics-model-for-interactive-world-models} } ```