|
Download README.md from RekaAI/Reka-Inverse-Dynamics-Model: direct link, hf CLI and curl.
- Browser
- Download file 5.16 kB
-
https://huggingface.co/RekaAI/Reka-Inverse-Dynamics-Model/resolve/main/README.md
- Command line
-
hf download hf://RekaAI/Reka-Inverse-Dynamics-Model/README.md
-
curl -L -o README.md https://huggingface.co/RekaAI/Reka-Inverse-Dynamics-Model/resolve/main/README.md
5.16 kB
| license: apache-2.0 | |
| library_name: pytorch | |
| pipeline_tag: video-classification | |
| tags: | |
| - inverse-dynamics | |
| - camera-motion | |
| - optical-flow | |
| model-index: | |
| - name: RIDM flow model | |
| results: | |
| - task: | |
| type: video-classification | |
| name: Forward, turn left or turn right | |
| dataset: | |
| name: Real video, 78 clips | |
| type: real-video | |
| metrics: | |
| - type: accuracy | |
| value: 85.9 | |
| name: Accuracy (percent) | |
| - task: | |
| type: video-classification | |
| name: Turn direction | |
| dataset: | |
| name: Real video, 47 turn clips | |
| type: real-video | |
| metrics: | |
| - type: accuracy | |
| value: 91.5 | |
| name: Accuracy (percent) | |
| - task: | |
| type: video-classification | |
| name: Forward, turn left or turn right | |
| metrics: | |
| - type: accuracy | |
| value: 84.5 | |
| name: Accuracy (percent) | |
| - task: | |
| type: video-classification | |
| name: Turn direction | |
| metrics: | |
| - type: accuracy | |
| value: 84.9 | |
| name: Accuracy (percent) | |
| # Reka Inverse Dynamics Model (RIDM) | |
| Two models predict camera motion (W/A/S/D/Shift, yaw, pitch) from video. | |
| Both trained on game renders. | |
| | | [flow](flow/) | [pixel](pixel/) | | |
| |---|---|---| | |
| | Input | Optical flow (RAFT-small) | Raw frames | | |
| | Parameters | 1,795,337 trained, plus 990,162 in frozen RAFT-small (2,785,499 in total) | 9,836,063 | | |
| | Output | One prediction per 17-frame window | One prediction per frame | | |
| | Real-video turn accuracy | 91.5 | 51.1 | | |
| | Gaming footage turn accuracy | 84.9 | 70.9 | | |
| | Dependencies | torch, torchvision, av, opencv, safetensors | torch, av, opencv, safetensors | | |
| Repository contents: | |
| ```text | |
| Reka-Inverse-Dynamics-Model/ | |
| README.md LICENSE NOTICE requirements.txt results.json | |
| flow/ model.safetensors config.json config_walking.json inference.py README.md | |
| pixel/ model.safetensors config.json inference.py README.md | |
| assets/ demo and example prediction videos | |
| examples/ 10 s sample clip and expected_flow_output.json | |
| tests/ test_smoke.py | |
| ``` | |
| ## Demo | |
| <video src="https://huggingface.co/RekaAI/Reka-Inverse-Dynamics-Model/resolve/main/assets/demo-walk-crowd-forward.mp4" controls muted loop width="100%"></video> | |
| The flow model on a walking tour of Bristol, with frame gap 12. Blue bars show the five keys. The dials show yaw and pitch in degrees. | |
| More clips are in `assets/`: `walk-empty-street.mp4`, `turn-right.mp4` and the failure case `failure-walk-key-with-no-walk.mp4`. | |
| In the failure clip the camera stands still and pans. The model reports the walk key (Shift) at 0.72 to 0.98 in every window. | |
| ## Which model to use | |
| The flow model has better generalization capabilities to real videos. | |
| The pixel model might have a high potential for gaming footage especially if trained further. | |
| ## Use | |
| ```bash | |
| pip install -r requirements.txt | |
| hf download RekaAI/Reka-Inverse-Dynamics-Model --local-dir ridm | |
| python ridm/flow/inference.py clip.mp4 --stride 17 | |
| ``` | |
| - The script prints JSON. Each window has `key_probabilities`, `keys_pressed`, `yaw_deg` and `pitch_deg`. | |
| - Default settings match game footage. For walking video, run | |
| `python ridm/flow/inference.py clip.mp4 --config ridm/flow/config_walking.json --stride 49`. | |
| - To check your install, run `pytest tests/test_smoke.py` from the repository root. It runs the sample clip and compares with `examples/expected_flow_output.json`. | |
| - Each folder runs alone. See `flow/README.md` and `pixel/README.md` for details. | |
| - Tested with Python 3.12 and the versions in `requirements.txt`. Newer versions are untested. | |
| Status: preliminary release. The scores marked † are unconfirmed. | |
| ## Training data | |
| - Source: 4,071 recorded gaming footage plays, captured at 1,280 x 720 and 48 frames per second. | |
| - Labels: the game engine gives the keys, the mouse movement and the camera angles for every frame. Nobody annotates. | |
| - Flow model: 7,071,037 frame pairs. Pixel model: about 104,000 gameplay windows. | |
| - Both models trained on one NVIDIA L4 GPU. | |
| ## License | |
| - The weights and the code in this repository use the Apache License 2.0. See `LICENSE`. | |
| - RAFT-small weights use the BSD-3 license. They come from torchvision and are not in this repository. See `NOTICE`. | |
| - The models trained on gameplay captures. This repository holds no game assets, no gameplay clips and no training data. | |
| - The videos in `assets/` and `examples/` are not Apache 2.0. They come from walking tours by POPtravel on Wikimedia Commons, licensed CC BY 3.0, and carry the model overlay or are shortened excerpts. | |
| Sources: [Bristol](https://commons.wikimedia.org/wiki/File:Walking_in_BRISTOL_-_UK_-_4K_60fps_(UHD).webm), | |
| [Ingolstadt](https://commons.wikimedia.org/wiki/File:Walking_in_INGOLSTADT_-_Germany_-_4K_60fps_(UHD).webm), | |
| [Bamberg](https://commons.wikimedia.org/wiki/File:Walking_in_BAMBERG_-_Germany_-_4K_60fps_(UHD).webm). | |
| ## Contact and citation | |
| - Questions and problems: open a discussion on this repository. | |
| ## Citation | |
| ```bibtex | |
| @misc{RIDM, | |
| title = {Reka Inverse Dynamics Model for Interactive World Models}, | |
| author = {Reka AI}, | |
| year = {2026}, | |
| url = {https://reka.ai/labs/research/reka-inverse-dynamics-model-for-interactive-world-models} | |
| } | |
| ``` | |