File size: 5,160 Bytes
1d7ff26
 
 
 
 
 
 
 
97a0336
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1d7ff26
 
3838ef5
1d7ff26
 
94070e8
1d7ff26
 
 
 
 
 
94070e8
 
1d7ff26
 
838d84d
1d7ff26
 
3838ef5
97a0336
3838ef5
1d7ff26
97a0336
 
 
1d7ff26
 
97a0336
 
177adcd
97a0336
 
 
 
 
3838ef5
 
94070e8
 
3838ef5
 
 
 
 
 
 
 
 
 
 
 
97a0336
3838ef5
 
1d7ff26
 
 
97a0336
 
94070e8
97a0336
 
 
 
1d7ff26
 
 
 
94070e8
97a0336
 
 
 
3838ef5
 
 
 
5170853
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
---
license: apache-2.0
library_name: pytorch
pipeline_tag: video-classification
tags:
- inverse-dynamics
- camera-motion
- optical-flow
model-index:
- name: RIDM flow model
  results:
  - task:
      type: video-classification
      name: Forward, turn left or turn right
    dataset:
      name: Real video, 78 clips
      type: real-video
    metrics:
    - type: accuracy
      value: 85.9
      name: Accuracy (percent)
  - task:
      type: video-classification
      name: Turn direction
    dataset:
      name: Real video, 47 turn clips
      type: real-video
    metrics:
    - type: accuracy
      value: 91.5
      name: Accuracy (percent)
  - task:
      type: video-classification
      name: Forward, turn left or turn right
    metrics:
    - type: accuracy
      value: 84.5
      name: Accuracy (percent)
  - task:
      type: video-classification
      name: Turn direction
    metrics:
    - type: accuracy
      value: 84.9
      name: Accuracy (percent)
---

# Reka Inverse Dynamics Model (RIDM)

Two models predict camera motion (W/A/S/D/Shift, yaw, pitch) from video.
Both trained on game renders.

| | [flow](flow/) | [pixel](pixel/) |
|---|---|---|
| Input | Optical flow (RAFT-small) | Raw frames |
| Parameters | 1,795,337 trained, plus 990,162 in frozen RAFT-small (2,785,499 in total) | 9,836,063 |
| Output | One prediction per 17-frame window | One prediction per frame |
| Real-video turn accuracy | 91.5 | 51.1 |
| Gaming footage turn accuracy | 84.9 | 70.9 |
| Dependencies | torch, torchvision, av, opencv, safetensors | torch, av, opencv, safetensors |

Repository contents:

```text
Reka-Inverse-Dynamics-Model/
  README.md  LICENSE  NOTICE  requirements.txt  results.json
  flow/   model.safetensors  config.json  config_walking.json  inference.py  README.md
  pixel/  model.safetensors  config.json  inference.py  README.md
  assets/  demo and example prediction videos
  examples/  10 s sample clip and expected_flow_output.json
  tests/  test_smoke.py
```

## Demo

<video src="https://huggingface.co/RekaAI/Reka-Inverse-Dynamics-Model/resolve/main/assets/demo-walk-crowd-forward.mp4" controls muted loop width="100%"></video>

The flow model on a walking tour of Bristol, with frame gap 12. Blue bars show the five keys. The dials show yaw and pitch in degrees.
More clips are in `assets/`: `walk-empty-street.mp4`, `turn-right.mp4` and the failure case `failure-walk-key-with-no-walk.mp4`.
In the failure clip the camera stands still and pans. The model reports the walk key (Shift) at 0.72 to 0.98 in every window.

## Which model to use

The flow model has better generalization capabilities to real videos.
The pixel model might have a high potential for gaming footage especially if trained further.

## Use

```bash
pip install -r requirements.txt
hf download RekaAI/Reka-Inverse-Dynamics-Model --local-dir ridm
python ridm/flow/inference.py clip.mp4 --stride 17
```

- The script prints JSON. Each window has `key_probabilities`, `keys_pressed`, `yaw_deg` and `pitch_deg`.
- Default settings match game footage. For walking video, run
  `python ridm/flow/inference.py clip.mp4 --config ridm/flow/config_walking.json --stride 49`.
- To check your install, run `pytest tests/test_smoke.py` from the repository root. It runs the sample clip and compares with `examples/expected_flow_output.json`.
- Each folder runs alone. See `flow/README.md` and `pixel/README.md` for details.
- Tested with Python 3.12 and the versions in `requirements.txt`. Newer versions are untested.

Status: preliminary release. The scores marked † are unconfirmed.

## Training data

- Source: 4,071 recorded gaming footage plays, captured at 1,280 x 720 and 48 frames per second.
- Labels: the game engine gives the keys, the mouse movement and the camera angles for every frame. Nobody annotates.
- Flow model: 7,071,037 frame pairs. Pixel model: about 104,000 gameplay windows.
- Both models trained on one NVIDIA L4 GPU.

## License

- The weights and the code in this repository use the Apache License 2.0. See `LICENSE`.
- RAFT-small weights use the BSD-3 license. They come from torchvision and are not in this repository. See `NOTICE`.
- The models trained on gameplay captures. This repository holds no game assets, no gameplay clips and no training data.
- The videos in `assets/` and `examples/` are not Apache 2.0. They come from walking tours by POPtravel on Wikimedia Commons, licensed CC BY 3.0, and carry the model overlay or are shortened excerpts.
  Sources: [Bristol](https://commons.wikimedia.org/wiki/File:Walking_in_BRISTOL_-_UK_-_4K_60fps_(UHD).webm),
  [Ingolstadt](https://commons.wikimedia.org/wiki/File:Walking_in_INGOLSTADT_-_Germany_-_4K_60fps_(UHD).webm),
  [Bamberg](https://commons.wikimedia.org/wiki/File:Walking_in_BAMBERG_-_Germany_-_4K_60fps_(UHD).webm).

## Contact and citation

- Questions and problems: open a discussion on this repository.

## Citation

```bibtex
@misc{RIDM,
  title  = {Reka Inverse Dynamics Model for Interactive World Models},
  author = {Reka AI},
  year   = {2026},
  url    = {https://reka.ai/labs/research/reka-inverse-dynamics-model-for-interactive-world-models}
}
```