Submitted by Haiwen Diao 66 From Pixels to Words -- Towards Native Vision-Language Primitives at Scale SenseTime 450 2