File size: 5,160 Bytes
1d7ff26 97a0336 1d7ff26 3838ef5 1d7ff26 94070e8 1d7ff26 94070e8 1d7ff26 838d84d 1d7ff26 3838ef5 97a0336 3838ef5 1d7ff26 97a0336 1d7ff26 97a0336 177adcd 97a0336 3838ef5 94070e8 3838ef5 97a0336 3838ef5 1d7ff26 97a0336 94070e8 97a0336 1d7ff26 94070e8 97a0336 3838ef5 5170853 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 | ---
license: apache-2.0
library_name: pytorch
pipeline_tag: video-classification
tags:
- inverse-dynamics
- camera-motion
- optical-flow
model-index:
- name: RIDM flow model
results:
- task:
type: video-classification
name: Forward, turn left or turn right
dataset:
name: Real video, 78 clips
type: real-video
metrics:
- type: accuracy
value: 85.9
name: Accuracy (percent)
- task:
type: video-classification
name: Turn direction
dataset:
name: Real video, 47 turn clips
type: real-video
metrics:
- type: accuracy
value: 91.5
name: Accuracy (percent)
- task:
type: video-classification
name: Forward, turn left or turn right
metrics:
- type: accuracy
value: 84.5
name: Accuracy (percent)
- task:
type: video-classification
name: Turn direction
metrics:
- type: accuracy
value: 84.9
name: Accuracy (percent)
---
# Reka Inverse Dynamics Model (RIDM)
Two models predict camera motion (W/A/S/D/Shift, yaw, pitch) from video.
Both trained on game renders.
| | [flow](flow/) | [pixel](pixel/) |
|---|---|---|
| Input | Optical flow (RAFT-small) | Raw frames |
| Parameters | 1,795,337 trained, plus 990,162 in frozen RAFT-small (2,785,499 in total) | 9,836,063 |
| Output | One prediction per 17-frame window | One prediction per frame |
| Real-video turn accuracy | 91.5 | 51.1 |
| Gaming footage turn accuracy | 84.9 | 70.9 |
| Dependencies | torch, torchvision, av, opencv, safetensors | torch, av, opencv, safetensors |
Repository contents:
```text
Reka-Inverse-Dynamics-Model/
README.md LICENSE NOTICE requirements.txt results.json
flow/ model.safetensors config.json config_walking.json inference.py README.md
pixel/ model.safetensors config.json inference.py README.md
assets/ demo and example prediction videos
examples/ 10 s sample clip and expected_flow_output.json
tests/ test_smoke.py
```
## Demo
<video src="https://huggingface.co/RekaAI/Reka-Inverse-Dynamics-Model/resolve/main/assets/demo-walk-crowd-forward.mp4" controls muted loop width="100%"></video>
The flow model on a walking tour of Bristol, with frame gap 12. Blue bars show the five keys. The dials show yaw and pitch in degrees.
More clips are in `assets/`: `walk-empty-street.mp4`, `turn-right.mp4` and the failure case `failure-walk-key-with-no-walk.mp4`.
In the failure clip the camera stands still and pans. The model reports the walk key (Shift) at 0.72 to 0.98 in every window.
## Which model to use
The flow model has better generalization capabilities to real videos.
The pixel model might have a high potential for gaming footage especially if trained further.
## Use
```bash
pip install -r requirements.txt
hf download RekaAI/Reka-Inverse-Dynamics-Model --local-dir ridm
python ridm/flow/inference.py clip.mp4 --stride 17
```
- The script prints JSON. Each window has `key_probabilities`, `keys_pressed`, `yaw_deg` and `pitch_deg`.
- Default settings match game footage. For walking video, run
`python ridm/flow/inference.py clip.mp4 --config ridm/flow/config_walking.json --stride 49`.
- To check your install, run `pytest tests/test_smoke.py` from the repository root. It runs the sample clip and compares with `examples/expected_flow_output.json`.
- Each folder runs alone. See `flow/README.md` and `pixel/README.md` for details.
- Tested with Python 3.12 and the versions in `requirements.txt`. Newer versions are untested.
Status: preliminary release. The scores marked † are unconfirmed.
## Training data
- Source: 4,071 recorded gaming footage plays, captured at 1,280 x 720 and 48 frames per second.
- Labels: the game engine gives the keys, the mouse movement and the camera angles for every frame. Nobody annotates.
- Flow model: 7,071,037 frame pairs. Pixel model: about 104,000 gameplay windows.
- Both models trained on one NVIDIA L4 GPU.
## License
- The weights and the code in this repository use the Apache License 2.0. See `LICENSE`.
- RAFT-small weights use the BSD-3 license. They come from torchvision and are not in this repository. See `NOTICE`.
- The models trained on gameplay captures. This repository holds no game assets, no gameplay clips and no training data.
- The videos in `assets/` and `examples/` are not Apache 2.0. They come from walking tours by POPtravel on Wikimedia Commons, licensed CC BY 3.0, and carry the model overlay or are shortened excerpts.
Sources: [Bristol](https://commons.wikimedia.org/wiki/File:Walking_in_BRISTOL_-_UK_-_4K_60fps_(UHD).webm),
[Ingolstadt](https://commons.wikimedia.org/wiki/File:Walking_in_INGOLSTADT_-_Germany_-_4K_60fps_(UHD).webm),
[Bamberg](https://commons.wikimedia.org/wiki/File:Walking_in_BAMBERG_-_Germany_-_4K_60fps_(UHD).webm).
## Contact and citation
- Questions and problems: open a discussion on this repository.
## Citation
```bibtex
@misc{RIDM,
title = {Reka Inverse Dynamics Model for Interactive World Models},
author = {Reka AI},
year = {2026},
url = {https://reka.ai/labs/research/reka-inverse-dynamics-model-for-interactive-world-models}
}
```
|