microsoft/VibeVoice-ASR-Streaming-7B Automatic Speech Recognition • 9B • Updated 21 days ago • 6.55k • 247
microsoft/VibeVoice-ASR-Streaming-1.5B Automatic Speech Recognition • 3B • Updated 21 days ago • 11.5k • 57
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models Paper • 2605.21573 • Published May 20 • 109
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding Paper • 2604.05015 • Published Apr 6 • 232
VibeVoice Collection Frontier Text-to-Speech Models https://microsoft.github.io/VibeVoice/ • 11 items • Updated 22 days ago • 264
Kosmos-G: Generating Images in Context with Multimodal Large Language Models Paper • 2310.02992 • Published Oct 4, 2023 • 4
Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object Detection Paper • 2205.09613 • Published May 19, 2022
BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers Paper • 2208.06366 • Published Aug 12, 2022
Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks Paper • 2208.10442 • Published Aug 22, 2022