Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

VISIONx @ NYU

university
https://www.sainingxie.com/
Activity Feed

AI & ML interests

None defined yet.

Recent Activity

craigwu  submitted a paper 20 days ago
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System
sayakpaul  authored a paper 25 days ago
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
sayakpaul  submitted a paper 25 days ago
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
View all activity

Papers

Benchmarking Visual State Tracking in Multimodal Video Understanding

PaintBench: Deterministic Evaluation of Precise Visual Editing

View all Papers

Ellis Brown's profile picture Peter Tong's profile picture Manoj Middepogu's profile picture Sai Charitha Akula's profile picture Penghao Wu's profile picture Jihan Yang's profile picture Saining Xie's profile picture Bingda Tang's profile picture BoYang Zheng's profile picture Sayak Paul's profile picture Shusheng Yang's profile picture Chenyu, Li's profile picture Anjali W Gupta's profile picture Xichen Pan's profile picture Pinzhi Huang's profile picture Nanye Ma's profile picture Jaskirat Singh's profile picture Ziteng Wang's profile picture Junwan Kim's profile picture Georgy Savva's profile picture Daohan Lu's profile picture Sihyun Yu's profile picture Zifan Zhao's profile picture
nyu-visionx 's papers 7
Submitted by
Pinzhi Huang
52

Benchmarking Visual State Tracking in Multimodal Video Understanding

nyu-visionx VISIONx @ NYU
46 1
Submitted by
Ellis Brown
4

PaintBench: Deterministic Evaluation of Precise Visual Editing

nyu-visionx VISIONx @ NYU
4 3
Submitted by
taesiri
31

Solaris: Building a Multiplayer Video World Model in Minecraft

nyu-visionx VISIONx @ NYU
229 3
Submitted by
BoYang Zheng
55

Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders

nyu-visionx VISIONx @ NYU
263 2
Submitted by
Ellis Brown
6

SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding

nyu-visionx VISIONx @ NYU
12 2
Submitted by
Jihan Yang
10

Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts

nyu-visionx VISIONx @ NYU
2
Submitted by
Peter Tong
170

Diffusion Transformers with Representation Autoencoders

nyu-visionx VISIONx @ NYU
2.02k 6
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs