embedl/Cosmos-Reason2-8B-W4A16-FlashHead Image-Text-to-Text β’ 3B β’ Updated 1 day ago β’ 124 β’ 1
embedl/Cosmos-Reason2-32B-W4A16-FlashHead Image-Text-to-Text β’ 6B β’ Updated 1 day ago β’ 33 β’ 1
embedl/Cosmos-Reason2-2B-W4A16-Edge2-FlashHead Image-Text-to-Text β’ 2B β’ Updated 1 day ago β’ 961 β’ 9
view post Post 3142 π The reasoning backbone quadruples from 8B to 32B , while the action expert remains at 2.3B! π We took a closer look at the architectural evolution from nvidia/Alpamayo-1.5-10B to nvidia/Alpamayo2-Super .Read the analysis here:https://huggingface.co/blog/JonnaMat/alpamayo2-superOur analysis explores some implications of this design choice, especially from a distillation perspective where keeping the expert compact could be key for efficient deployment. π§ See translation 3 replies Β· π₯ 7 7 + Reply