tijayantML's picture
Update README.md
5170853 verified
|
Raw History Blame Contribute Delete
5.16 kB
---
license: apache-2.0
library_name: pytorch
pipeline_tag: video-classification
tags:
- inverse-dynamics
- camera-motion
- optical-flow
model-index:
- name: RIDM flow model
results:
- task:
type: video-classification
name: Forward, turn left or turn right
dataset:
name: Real video, 78 clips
type: real-video
metrics:
- type: accuracy
value: 85.9
name: Accuracy (percent)
- task:
type: video-classification
name: Turn direction
dataset:
name: Real video, 47 turn clips
type: real-video
metrics:
- type: accuracy
value: 91.5
name: Accuracy (percent)
- task:
type: video-classification
name: Forward, turn left or turn right
metrics:
- type: accuracy
value: 84.5
name: Accuracy (percent)
- task:
type: video-classification
name: Turn direction
metrics:
- type: accuracy
value: 84.9
name: Accuracy (percent)
---
# Reka Inverse Dynamics Model (RIDM)
Two models predict camera motion (W/A/S/D/Shift, yaw, pitch) from video.
Both trained on game renders.
| | [flow](flow/) | [pixel](pixel/) |
|---|---|---|
| Input | Optical flow (RAFT-small) | Raw frames |
| Parameters | 1,795,337 trained, plus 990,162 in frozen RAFT-small (2,785,499 in total) | 9,836,063 |
| Output | One prediction per 17-frame window | One prediction per frame |
| Real-video turn accuracy | 91.5 | 51.1 |
| Gaming footage turn accuracy | 84.9 | 70.9 |
| Dependencies | torch, torchvision, av, opencv, safetensors | torch, av, opencv, safetensors |
Repository contents:
```text
Reka-Inverse-Dynamics-Model/
README.md LICENSE NOTICE requirements.txt results.json
flow/ model.safetensors config.json config_walking.json inference.py README.md
pixel/ model.safetensors config.json inference.py README.md
assets/ demo and example prediction videos
examples/ 10 s sample clip and expected_flow_output.json
tests/ test_smoke.py
```
## Demo
<video src="https://huggingface.co/RekaAI/Reka-Inverse-Dynamics-Model/resolve/main/assets/demo-walk-crowd-forward.mp4" controls muted loop width="100%"></video>
The flow model on a walking tour of Bristol, with frame gap 12. Blue bars show the five keys. The dials show yaw and pitch in degrees.
More clips are in `assets/`: `walk-empty-street.mp4`, `turn-right.mp4` and the failure case `failure-walk-key-with-no-walk.mp4`.
In the failure clip the camera stands still and pans. The model reports the walk key (Shift) at 0.72 to 0.98 in every window.
## Which model to use
The flow model has better generalization capabilities to real videos.
The pixel model might have a high potential for gaming footage especially if trained further.
## Use
```bash
pip install -r requirements.txt
hf download RekaAI/Reka-Inverse-Dynamics-Model --local-dir ridm
python ridm/flow/inference.py clip.mp4 --stride 17
```
- The script prints JSON. Each window has `key_probabilities`, `keys_pressed`, `yaw_deg` and `pitch_deg`.
- Default settings match game footage. For walking video, run
`python ridm/flow/inference.py clip.mp4 --config ridm/flow/config_walking.json --stride 49`.
- To check your install, run `pytest tests/test_smoke.py` from the repository root. It runs the sample clip and compares with `examples/expected_flow_output.json`.
- Each folder runs alone. See `flow/README.md` and `pixel/README.md` for details.
- Tested with Python 3.12 and the versions in `requirements.txt`. Newer versions are untested.
Status: preliminary release. The scores marked † are unconfirmed.
## Training data
- Source: 4,071 recorded gaming footage plays, captured at 1,280 x 720 and 48 frames per second.
- Labels: the game engine gives the keys, the mouse movement and the camera angles for every frame. Nobody annotates.
- Flow model: 7,071,037 frame pairs. Pixel model: about 104,000 gameplay windows.
- Both models trained on one NVIDIA L4 GPU.
## License
- The weights and the code in this repository use the Apache License 2.0. See `LICENSE`.
- RAFT-small weights use the BSD-3 license. They come from torchvision and are not in this repository. See `NOTICE`.
- The models trained on gameplay captures. This repository holds no game assets, no gameplay clips and no training data.
- The videos in `assets/` and `examples/` are not Apache 2.0. They come from walking tours by POPtravel on Wikimedia Commons, licensed CC BY 3.0, and carry the model overlay or are shortened excerpts.
Sources: [Bristol](https://commons.wikimedia.org/wiki/File:Walking_in_BRISTOL_-_UK_-_4K_60fps_(UHD).webm),
[Ingolstadt](https://commons.wikimedia.org/wiki/File:Walking_in_INGOLSTADT_-_Germany_-_4K_60fps_(UHD).webm),
[Bamberg](https://commons.wikimedia.org/wiki/File:Walking_in_BAMBERG_-_Germany_-_4K_60fps_(UHD).webm).
## Contact and citation
- Questions and problems: open a discussion on this repository.
## Citation
```bibtex
@misc{RIDM,
title = {Reka Inverse Dynamics Model for Interactive World Models},
author = {Reka AI},
year = {2026},
url = {https://reka.ai/labs/research/reka-inverse-dynamics-model-for-interactive-world-models}
}
```