---
language:
- en
- zh
license: other
license_name: tencent-hy-world-2.0-community
license_link: https://github.com/Tencent-Hunyuan/HY-World-2.0/blob/main/License.txt
pipeline_tag: image-to-3d
library_name: hy-world-2
tags:
- worldmodel
- 3d
- hy-world
extra_gated_eu_disallowed: true
---
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
[English](README.md) | [简体中文](README_zh.md)
"What Is Now Proved Was Once Only Imagined"
## 🎥 Video
https://github.com/user-attachments/assets/b56f4750-25c9-48fb-83ff-d58526711463
## 🔥 News
- **[May 18, 2026]**: 🤗 Open-source World Generation inference code and WorldStereo 2.0 model weights!
- **[May 11, 2026]**: 🤗 Open-source HY-Pano 2.0 inference code and model weights!
- **[April 16, 2026]**: 🚀 Release HY-World 2.0 technical report & partial codes!
- **[April 16, 2026]**: 🤗 Open-source WorldMirror 2.0 inference code and model weights!
## 📋 Table of Contents
- [📖 Introduction](#-introduction)
- [✨ Highlights](#-highlights)
- [🧩 Architecture](#-architecture)
- [📝 Open-Source Plan](#-open-source-plan)
- [🎁 Model Zoo](#-model-zoo)
- [🤗 Get Started](#-get-started)
- [🔮 Performance](#-performance)
- [🎬 More Examples](#-more-examples)
- [📚 Citation](#-citation)
## 📖 Introduction
**HY-World 2.0** is a multi-modal world model framework for **world generation** and **world reconstruction**. It accepts diverse input modalities — text, single-view images, multi-view images, and videos — and produces 3D world representations (meshes / Gaussian Splattings). It offers two core capabilities:
- **World Generation** (text / single image → 3D world): syntheses high-fidelity, navigable 3D scenes through a four-stage method —— a)  with HY-Pano 2.0, b)  with WorldNav, c)  with WorldStereo 2.0, and d)  with WorldMirror 2.0 & 3DGS learning.
- **World Reconstruction** (multi-view images / video → 3D): Powered by WorldMirror 2.0, a unified feed-forward model that simultaneously predicts depth, surface normals, camera parameters, 3D point clouds, and 3DGS attributes in a single forward pass.
HY-World 2.0 is an **open-source state-of-the-art** world model. We released all model weights, code, and technical details to facilitate reproducibility and advance research in this field.
### Why 3D World Models?
Existing world models, such as Genie 3, Cosmos, and HY-World 1.5 (WorldPlay+WorldCompass), generate pixel-level videos — essentially "watching a movie" that vanishes once playback ends. **HY-World 2.0 takes a fundamentally different approach**: it directly produces editable, persistent 3D assets (meshes / 3DGS) that can be imported into game engines like Blender/Unity/Unreal Engine/Isaac Sim — more like "building a playable game" than recording a clip. This paradigm shift natively resolves many long-standing pain points of video world models:
| | Video World Models | 3D World Model (HY-World 2.0) |
|--|---|---|
| **Output** | Pixel videos (non-editable) | Real 3D assets — meshes / 3DGS (fully editable) |
| **Playable Duration** | Limited (typically 1 min) | Unlimited — assets persist permanently |
| **3D Consistency** | No (flickering, artifacts across views) | Native — inherently consistent in 3D |
| **Real-Time Rendering** | Requires per-frame inference; high latency | Consumer GPUs can render in real time |
| **Controllability** | Weak (imprecise character control, no real physics) | Precise — zero-error control, real physics collision, accurate lighting |
| **Inference Cost** | Accumulates with every interaction | One-time generation; rendering cost ≈ 0 |
| **Engine Compatibility** | ✗ Video files only | ✓ Directly importable into Blender / UE / Isaac Engine |
| | $\color{IndianRed}{\textsf{Watch a video, then it's gone}}$ | $\color{RoyalBlue}{\textbf{Build a world, keep it forever}}$ |
All above are real 3D assets (not generated videos) and entirely created by HY-World 2.0 -- captured from live real-time interaction.
## ✨ Highlights
- **Real 3D Worlds, Not Just Videos**
Unlike video-only world models (e.g., Genie 3, HY World 1.5), HY-World 2.0 generates **real 3D assets** — 3DGS, meshes, and point clouds — that are freely explorable, editable, and directly importable into **Unity / Unreal Engine / Isaac**. From a single text prompt or image, create navigable 3D worlds with diverse styles: realistic, cartoon, game, and more.
- **Instant 3D Reconstruction from Photos & Videos**
Powered by **WorldMirror 2.0**, a unified feed-forward model that predicts dense point clouds, depth maps, surface normals, camera parameters, and 3DGS from multi-view images or casual videos in a single forward pass. Supports flexible-resolution inference (50K–500K pixels) with SOTA accuracy. Capture a video, get a digital twin.
- **Interactive Character Exploration**
Go beyond viewing — **play inside your generated worlds**. HY-World 2.0 supports first-person navigation and third-person character mode, enabling users to freely explore AI-generated streets, buildings, and landscapes with physics-based collision. Go to [our product page](https://3d.hunyuan.tencent.com/sceneTo3D) for free try ().
## 🧩 Architecture
- **Refer to our tech report for more details**
A systematic pipeline of HY-World 2.0 — *Panorama Generation* (HY-Pano-2.0) → *Trajectory Planning* (WorldNav) → *World Expansion* (WorldStereo 2.0) → *World Composition* (WorldMirror 2.0 + Splattings Learning) — that automatically transforms text or a single image into a high-fidelity, navigable 3D world (3DGS/mesh outputs).
## 📝 Open-Source Plan
- [x] Technical Report
- [x] WorldMirror 2.0 Code & Model Checkpoints
- [x] Full Inference Code for World Generation (WorldNav + WorldStereo + World Composition)
- [x] Panorama Generation (HY-Pano 2.0) Model & Code
- [x] World Expansion (WorldStereo 2.0) Model & Code
## 🎁 Model Zoo
### World Reconstruction — WorldMirror Series
| Model | Description | Params | Date | Hugging Face |
|-------|-------------|--------|------|--------------|
| WorldMirror-2 [new] | Multi-view / video → 3D reconstruction | ~1.2B | 2026 | [Download](https://huggingface.co/tencent/HY-World-2.0/tree/main/HY-WorldMirror-2.0) |
| WorldMirror-1 | Multi-view / video → 3D reconstruction (legacy) | ~1.2B | 2025 | [Download](https://huggingface.co/tencent/HunyuanWorld-Mirror/tree/main) |
### Panorama Generation — HY-Pano Series
| Model | Description | Params | Date | Hugging Face |
|-------|-------------|--------|------|--------------|
| HY-Pano-2 [new] | Text / image → 360° panorama | ~80B | 2026 | [Download](https://huggingface.co/tencent/HY-World-2.0/tree/main/HY-Pano-2.0) |
| HY-Pano-2-Qwen [new] | Text / image → 360° panorama | ~425M | 2026 | [Download](https://huggingface.co/tencent/HY-World-2.0/blob/main/HY-Pano-2.0/pytorch_lora_weights.safetensors) |
### World Expansion — WorldStereo Series
| Model | Description | Params | Date | Hugging Face |
|-----------------|-------------|-----|------|--------------|
| WorldStereo-2 [new] | Panorama → 3DGS world | ~17B | 2026 | [Download](https://huggingface.co/hanshanxue/WorldStereo/tree/main) |
We recommend referring to our previous works, [WorldStereo](https://github.com/FuchengSu/WorldStereo) and [WorldMirror](https://github.com/Tencent-Hunyuan/HunyuanWorld-Mirror), for background knowledge on 3D world generation and reconstruction.
## 🤗 Get Started
### Install Requirements
We recommend **CUDA 12.8** and **Python 3.11+**. The easiest path is to prepare one shared environment, first make **World Reconstruction (WorldMirror 2.0)** work, and then install the extra components required by **World Generation**.
#### 1. Create the shared environment
```bash
git clone https://github.com/Tencent-Hunyuan/HY-World-2.0
cd HY-World-2.0
conda create -n hyworld2 python=3.11.15
conda activate hyworld2
```
#### 2. Install World Reconstruction dependencies
After this step, the environment is ready for **worldrecon / WorldMirror 2.0**.
```bash
# Base dependencies shared by worldrecon and worldgen
pip install -r requirements.txt
# Recommended: install the custom gsplat variant once for both worldrecon and worldgen
cd hyworld2/worldgen/third_party/gsplat_maskgaussian
pip install -e . --no-build-isolation
cd ../../../../
```
If you only need **worldrecon** and want a simpler fallback, official `gsplat` is also supported:
```bash
pip install git+https://github.com/nerfstudio-project/gsplat.git
```
Install **one** FlashAttention backend:
```bash
# Recommended for Hopper GPUs: FlashAttention-3
git clone https://github.com/Dao-AILab/flash-attention.git
cd flash-attention/hopper
python setup.py install
cd ../../
rm -rf flash-attention
```
```bash
# Simpler alternative: FlashAttention-2
pip install flash-attn --no-build-isolation
```
#### 3. Add extra World Generation dependencies
Run the following extra steps only if you need **worldgen**. These commands assume the shared `hyworld2` environment above is already active.
```bash
# Git-based dependencies require torch/CUDA to be installed first
pip install --no-build-isolation -r requirements_git.txt
# recastnavigation is managed as a git submodule
git submodule update --init --recursive
# Recast navmesh extension for trajectory planning
cd hyworld2/worldgen/third_party/navmesh
pip install . --no-build-isolation
cd ../../../../
```
For **HY-Pano-2** installation, please refer to **[hyworld2/panogen/README.md](hyworld2/panogen/README.md)**.
### Code Usage — Panorama Generation (HY-Pano-2)
For full documentation and CLI reference, see **[hyworld2/panogen/README.md](hyworld2/panogen/README.md)**.
We provide a `diffusers`-like Python API for HY-Pano 2.0. Model weights are automatically downloaded from Hugging Face on first run.
```python
from pipeline import HunyuanPanoPipeline
pipeline = HunyuanPanoPipeline.from_pretrained('tencent/HY-World-2.0')
output = pipeline('input.png')
output.save('output_panorama.png')
```
### Code Usage — World Generation (WorldNav, WorldStereo-2, and 3DGS)
The world Generation pipeline turns a panorama scene into a navigable 3D world through five stages:
| Stage | Script | Description |
|-------|--------|-------------|
| 1. Trajectory Planning | `traj_generate.py` | VLM-guided camera trajectory planning with obstacle-aware navigation |
| 2. Trajectory Rendering | `traj_render.py` | Multi-GPU point-cloud rendering along planned trajectories |
| 3. World Expansion | `video_gen.py` | WorldStereo-2 keyframe generation with memory-guided consistency |
| 4. GS Data Preparation | `gen_gs_data.py` | Extract frames, aligned depth, normals, and cameras for 3DGS training |
| 5. 3DGS Training | `world_gs_trainer.py` | Optimize and export the final Gaussian Splatting world |
For full documentation, prerequisites, and CLI arguments, see **[hyworld2/worldgen/README.md](hyworld2/worldgen/README.md)**.
### Code Usage — WorldMirror 2.0
WorldMirror 2.0 supports the following usage modes:
- [Code Usage](#code-usage--worldmirror-20)
- [Gradio App](#gradio-app--worldmirror-20)
We provide a `diffusers`-like Python API for WorldMirror 2.0. Model weights are automatically downloaded from Hugging Face on first run.
```python
from hyworld2.worldrecon.pipeline import WorldMirrorPipeline
pipeline = WorldMirrorPipeline.from_pretrained('tencent/HY-World-2.0')
result = pipeline('path/to/images')
```
**With Prior Injection (Camera & Depth):**
```python
result = pipeline(
'path/to/images',
prior_cam_path='path/to/prior_camera.json',
prior_depth_path='path/to/prior_depth/',
)
```
> For the detailed structure of camera/depth priors and how to prepare them, see [Prior Preparation Guide](DOCUMENTATION.md#prior-injection).
**CLI:**
```bash
# Single GPU
python -m hyworld2.worldrecon.pipeline --input_path path/to/images
# Multi-GPU
torchrun --nproc_per_node=2 -m hyworld2.worldrecon.pipeline \
--input_path path/to/images \
--use_fsdp --enable_bf16
```
> **Important:** In multi-GPU mode, the number of input images must be **>= the number of GPUs**. For example, with `--nproc_per_node=8`, provide at least 8 images.
### Gradio App — WorldMirror 2.0
We provide an interactive [Gradio](https://www.gradio.app/) web demo for WorldMirror 2.0. Upload images or videos and visualize 3DGS, point clouds, depth maps, normal maps, and camera parameters in your browser.
```bash
# Single GPU
python -m hyworld2.worldrecon.gradio_app
# Multi-GPU
torchrun --nproc_per_node=2 -m hyworld2.worldrecon.gradio_app \
--use_fsdp --enable_bf16
```
For the full list of Gradio app arguments (port, share, local checkpoints, etc.), see [DOCUMENTATION.md](DOCUMENTATION.md#gradio-app).
## 🔮 Performance
For full benchmark results, please refer to the [technical report](https://3d-models.hunyuan.tencent.com/world/).
### WorldStereo 2.0 — Camera Control
| Methods |
Camera Metrics |
Visual Quality |
| RotErr ↓ | TransErr ↓ | ATE ↓ |
Q-Align ↑ | CLIP-IQA+ ↑ | Laion-Aes ↑ | CLIP-I ↑ |
| SEVA | 1.690 | 1.578 | 2.879 | 3.232 | 0.479 | 4.623 | 77.16 |
| Gen3C | 0.944 | 1.580 | 2.789 | 3.353 | 0.489 | 4.863 | 82.33 |
| WorldStereo | 0.762 | 1.245 | 2.141 | 4.149 | 0.547 | 5.257 | 89.05 |
| WorldStereo 2.0 | 0.492 | 0.968 | 1.768 | 4.205 | 0.544 | 5.266 | 89.43 |
### WorldStereo 2.0 — Single-View-Generated Reconstruction
| Methods |
Tanks-and-Temples |
MipNeRF360 |
| Precision ↑ |
Recall ↑ |
F1-Score ↑ |
AUC ↑ |
Precision ↑ |
Recall ↑ |
F1-Score ↑ |
AUC ↑ |
| SEVA |
33.59 |
35.34 |
36.73 |
51.03 |
22.38 |
55.63 |
28.75 |
46.81 |
| Gen3C |
46.73 |
25.51 |
31.24 |
42.44 |
23.28 |
75.37 |
35.26 |
52.10 |
| Lyra |
50.38 |
28.67 |
32.54 |
43.05 |
30.02 |
58.60 |
36.05 |
49.89 |
| FlashWorld |
26.58 |
20.72 |
22.29 |
30.45 |
35.97 |
53.77 |
42.60 |
53.86 |
| WorldStereo 2.0 |
43.62 |
41.02 |
41.43 |
58.19 |
43.19 |
65.32 |
51.27 |
65.79 |
| WorldStereo 2.0 (DMD) |
40.41 |
44.41 |
43.16 |
60.09 |
42.34 |
64.83 |
50.52 |
65.64 |
### WorldMirror 2.0 — Point Map Reconstruction
**Point Map Reconstruction on 7-Scenes, NRGBD, and DTU.** We report the mean Accuracy and Completeness of WorldMirror under different input configurations. **Bold** results are best. "L / M / H" denote low / medium / high inference resolution. "+ all priors" denotes injection of camera extrinsics, camera intrinsics, and depth priors.
| Method |
7-Scenes (scene) |
NRGBD (scene) |
DTU (object) |
| Acc. ↓ | Comp. ↓ |
Acc. ↓ | Comp. ↓ |
Acc. ↓ | Comp. ↓ |
| WorldMirror 1.0 |
| L | 0.043 | 0.055 | 0.046 | 0.049 | 1.476 | 1.768 |
| L + all priors | 0.021 | 0.026 | 0.022 | 0.020 | 1.347 | 1.392 |
| M | 0.043 | 0.049 | 0.041 | 0.045 | 1.017 | 1.780 |
| M + all priors | 0.018 | 0.023 | 0.016 | 0.014 | 0.735 | 0.935 |
| H | 0.079 | 0.087 | 0.077 | 0.093 | 2.271 | 2.113 |
| H + all priors | 0.042 | 0.041 | 0.078 | 0.082 | 1.773 | 1.478 |
|
| WorldMirror 2.0 |
| L | 0.041 | 0.052 | 0.047 | 0.058 | 1.352 | 2.009 |
| L + all priors | 0.019 | 0.024 | 0.017 | 0.015 | 1.100 | 1.201 |
| M | 0.033 | 0.046 | 0.039 | 0.047 | 1.005 | 1.892 |
| M + all priors | 0.013 | 0.017 | 0.013 | 0.013 | 0.690 | 0.876 |
| H | 0.037 | 0.040 | 0.046 | 0.053 | 0.845 | 1.904 |
| H + all priors | 0.012 | 0.016 | 0.015 | 0.016 | 0.554 | 0.771 |
### WorldMirror 2.0 — Prior Comparison
**Comparison with Pow3R and MapAnything under Different Prior Conditions.** Results are averaged on 7-Scenes, NRGBD, and DTU datasets. Pow3R (pro) refers to the original Pow3R with Procrustes alignment.
## 🎬 More Examples
## 📖 Documentation
For detailed usage guides, parameter references, output format specifications, and prior injection instructions, see **[DOCUMENTATION.md](DOCUMENTATION.md)**.
## 📚 Citation
If you find HunyuanWorld 2.0 useful for your research, please cite:
```bibtex
@article{hyworld22026,
title={HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds},
author={Team HY-World},
journal={arXiv preprint arXiv:2604.14268},
year={2026}
}
@article{hunyuanworld2025tencent,
title={HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels},
author={Team HunyuanWorld},
year={2025},
journal={arXiv preprint}
}
```
## 📧 Contact
Please send emails to tengfeiwang12@gmail.com for questions or feedback.
## 🙏 Acknowledgements
We would like to thank [HunyuanWorld 1.0](https://github.com/Tencent-Hunyuan/HunyuanWorld-1.0), [WorldMirror](https://github.com/Tencent-Hunyuan/HunyuanWorld-Mirror), [WorldPlay](https://github.com/Tencent-Hunyuan/HY-WorldPlay), [WorldStereo](https://github.com/FuchengSu/WorldStereo), [HunyuanImage](https://github.com/Tencent-Hunyuan/HunyuanImage-3.0) for their great work.