330 GB
208 files
Updated 3 months ago
Name
Size
.gitattributes2.46 kB
xet
BIAS.md529 Bytes
xet
EXPLAINABILITY.md2.43 kB
xet
LICENSE4.73 kB
xet
PRIVACY.md1.57 kB
xet
README.md5.88 kB
xet
SAFETY.md945 Bytes
xet
SHARD_0000.tar.gz1.74 GB
xet
SHARD_0001.tar.gz1.56 GB
xet
SHARD_0002.tar.gz1.69 GB
xet
SHARD_0003.tar.gz1.43 GB
xet
SHARD_0004.tar.gz1.63 GB
xet
SHARD_0005.tar.gz1.52 GB
xet
SHARD_0006.tar.gz1.8 GB
xet
SHARD_0007.tar.gz1.8 GB
xet
SHARD_0008.tar.gz1.58 GB
xet
SHARD_0009.tar.gz1.65 GB
xet
SHARD_0010.tar.gz1.63 GB
xet
SHARD_0011.tar.gz1.59 GB
xet
SHARD_0012.tar.gz1.85 GB
xet
SHARD_0013.tar.gz1.44 GB
xet
SHARD_0014.tar.gz1.6 GB
xet
SHARD_0015.tar.gz1.61 GB
xet
SHARD_0016.tar.gz1.61 GB
xet
SHARD_0017.tar.gz1.77 GB
xet
SHARD_0018.tar.gz1.9 GB
xet
SHARD_0019.tar.gz1.8 GB
xet
SHARD_0020.tar.gz1.92 GB
xet
SHARD_0021.tar.gz1.76 GB
xet
SHARD_0022.tar.gz1.82 GB
xet
SHARD_0023.tar.gz1.66 GB
xet
SHARD_0024.tar.gz1.8 GB
xet
SHARD_0025.tar.gz1.77 GB
xet
SHARD_0026.tar.gz1.72 GB
xet
SHARD_0027.tar.gz1.62 GB
xet
SHARD_0028.tar.gz1.59 GB
xet
SHARD_0029.tar.gz1.49 GB
xet
SHARD_0030.tar.gz1.5 GB
xet
SHARD_0031.tar.gz1.76 GB
xet
SHARD_0032.tar.gz1.73 GB
xet
SHARD_0033.tar.gz1.5 GB
xet
SHARD_0034.tar.gz1.33 GB
xet
SHARD_0035.tar.gz1.3 GB
xet
SHARD_0036.tar.gz1.37 GB
xet
SHARD_0037.tar.gz1.75 GB
xet
SHARD_0038.tar.gz1.77 GB
xet
SHARD_0039.tar.gz1.66 GB
xet
SHARD_0040.tar.gz1.61 GB
xet
SHARD_0041.tar.gz1.85 GB
xet
SHARD_0042.tar.gz1.5 GB
xet
SHARD_0043.tar.gz1.62 GB
xet
SHARD_0044.tar.gz1.66 GB
xet
SHARD_0045.tar.gz1.31 GB
xet
SHARD_0046.tar.gz1.61 GB
xet
SHARD_0047.tar.gz1.71 GB
xet
SHARD_0048.tar.gz1.51 GB
xet
SHARD_0049.tar.gz1.6 GB
xet
SHARD_0050.tar.gz1.65 GB
xet
SHARD_0051.tar.gz1.53 GB
xet
SHARD_0052.tar.gz1.52 GB
xet
SHARD_0053.tar.gz1.6 GB
xet
SHARD_0054.tar.gz1.56 GB
xet
SHARD_0055.tar.gz1.78 GB
xet
SHARD_0056.tar.gz1.64 GB
xet
SHARD_0057.tar.gz1.74 GB
xet
SHARD_0058.tar.gz1.53 GB
xet
SHARD_0059.tar.gz1.57 GB
xet
SHARD_0060.tar.gz1.71 GB
xet
SHARD_0061.tar.gz1.55 GB
xet
SHARD_0062.tar.gz1.57 GB
xet
SHARD_0063.tar.gz1.59 GB
xet
SHARD_0064.tar.gz1.83 GB
xet
SHARD_0065.tar.gz1.78 GB
xet
SHARD_0066.tar.gz1.51 GB
xet
SHARD_0067.tar.gz1.6 GB
xet
SHARD_0068.tar.gz1.51 GB
xet
SHARD_0069.tar.gz1.62 GB
xet
SHARD_0070.tar.gz1.48 GB
xet
SHARD_0071.tar.gz1.56 GB
xet
SHARD_0072.tar.gz1.7 GB
xet
SHARD_0073.tar.gz1.64 GB
xet
SHARD_0074.tar.gz1.66 GB
xet
SHARD_0075.tar.gz1.6 GB
xet
SHARD_0076.tar.gz1.76 GB
xet
SHARD_0077.tar.gz1.65 GB
xet
SHARD_0078.tar.gz1.68 GB
xet
SHARD_0079.tar.gz1.48 GB
xet
SHARD_0080.tar.gz1.67 GB
xet
SHARD_0081.tar.gz1.52 GB
xet
SHARD_0082.tar.gz1.44 GB
xet
SHARD_0083.tar.gz1.56 GB
xet
SHARD_0084.tar.gz1.87 GB
xet
SHARD_0085.tar.gz1.88 GB
xet
SHARD_0086.tar.gz1.88 GB
xet
SHARD_0087.tar.gz1.75 GB
xet
SHARD_0088.tar.gz1.75 GB
xet
SHARD_0089.tar.gz1.58 GB
xet
SHARD_0090.tar.gz1.66 GB
xet
SHARD_0091.tar.gz1.64 GB
xet
SHARD_0092.tar.gz1.52 GB
xet
README.md

NitroGen Dataset

Dataset Description:

The NitroGen dataset contains action annotations for publicly available gameplay videos. Specifically, we used an in-house model to annotate each video frame with gamepad actions. Note that reproducing results from the NitroGen paper requires additional filtering, such as IDLE frame filtering.

This repository is structured as follows:

├── actions
│   ├── SHARD_0000
│   │   ├── <video_id>
│   │   │   ├── <video_id>_chunk_0000
│   │   │   │   ├── actions_processed.parquet
│   │   │   │   ├── actions_raw.parquet
│   │   │   │   └── metadata.json
│   │   │   ├── <video_id>_chunk_0001
│   │   │   │   ├── actions_processed.parquet
│   │   │   │   ├── actions_raw.parquet
│   │   │   │   └── metadata.json
│   │   │   ├── ...
│   ├── SHARD_0001
│   │   ├── ...
│   ├── ...

Annotations for each video are split into 20-second chunks. Each chunk directory contains the following files:

  • actions_raw.parquet: this is a table that stores per-frame gamepad actions
  • metadata.json: contains all metadata related to the chunk, such as timestamps, length or url
  • actions_processed.parquet (optional): same format as actions_raw.parquet but with quality filtering and remapping applied

metadata.json contains the following:

{
    "uuid": "<video_id>_chunk_<chunk_number>_actions",
    "chunk_id": "<chunk_number>",
    "chunk_size": int, # number of frames in the chunk
    "original_video": {
        "resolution": [1080, 1920],
        "video_id": "<video_id>",
        "source": str,
        "url": str,

        # chunk start and end timestamps
        "start_time": float, # in seconds
        "end_time": float,
        "duration": float,

        "start_frame": int,
        "end_frame": int,
    },
    "game": str,
    "controller_type": str,
    
    # bbox to mask the on-screen controller in pixel space, relative to resolution above
    "bbox_controller_overlay": [xtl, ytl, w, h],

    # optional, only if the gameplay is not full screen in the video, relative coordinates in [0, 1] 
    "bbox_game_area": {
        "xtl": float,
        "ytl": float,
        "xbr": float,
        "ybr": float
    },

    # optional, list of bounding boxes for elements that are not gameplay
    "bbox_others": [
        {
            "xtl": float,
            "ytl": float,
            "xbr": float,
            "ybr": float
        },
        ...
    ]
}

actions_raw.parquet and actions_processed.parquet are tables containing gamepad actions, one row corresponds to a gamepad state for one frame from the original video. Each row follows a standard gamepad layout, with $17$ boolean columns for buttons and $2$ columns for each joystick, containing pairs of $[-1,1]$ values.

Button columns are the following:

[
  "dpad_down",
  "dpad_left",
  "dpad_right",
  "dpad_up",
  "left_shoulder",
  "left_thumb",
  "left_trigger",
  "right_shoulder",
  "right_thumb",
  "right_trigger",
  "south",
  "west",
  "east",
  "north",
  "back",
  "start",
  "guide",
]

Joystick columns are j_left and j_right. They contain $x,y$ coordinates in $[-1, 1]$. Note that $(-1,-1)$ is the top-left as is standard for joystick axes.

This dataset only includes the gamepad action labels. This dataset is for research and development only.

Dataset Owner(s):

NVIDIA Corporation

Dataset Creation Date:

2025-12-19

License/Terms of Use:

CC BY-NC 4.0

Intended Usage:

This dataset is intended for training behavior cloning policies (video to actions) and world models (actions to video)

Dataset Characterization

** Data Collection Method
Automated

** Labeling Method
Synthetic

Dataset Format

Tabular, parquet files

Dataset Quantification

Annotated videos: 30k Total number of frames annotated: ~15B

Citation

If you find this work useful, please consider citing:

@misc{magne2026nitrogen,
      title={NitroGen: An Open Foundation Model for Generalist Gaming Agents}, 
      author={Loïc Magne and Anas Awadalla and Guanzhi Wang and Yinzhen Xu and Joshua Belofsky and Fengyuan Hu and Joohwan Kim and Ludwig Schmidt and Georgia Gkioxari and Jan Kautz and Yisong Yue and Yejin Choi and Yuke Zhu and Linxi "Jim" Fan},
      year={2026},
      eprint={2601.02427},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2601.02427}, 
}

Ethical Considerations:

NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.

Total size
330 GB
Files
208
Last updated
Jun 10
Pre-warmed CDN
US EU US EU

Contributors