video-SALMONN 2 Pro tsinghua-ee/video-SALMONN-2-Pro-4B 5B • Updated 26 days ago • 240 • 1 tsinghua-ee/video-SALMONN-2-Pro-8B 10B • Updated 26 days ago • 118 • 2 tsinghua-ee/video-SALMONN-2-Pro-32B 34B • Updated 26 days ago • 66 • 3
Spatial Audio & Visual Spatial Audio & Visual LLMs JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Paper • 2602.18527 • Published Feb 20 • 3 tsinghua-ee/JAEGER Updated Jun 23 • 14 • 4
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Paper • 2602.18527 • Published Feb 20 • 3
General Time Series SciTS: Scientific Time Series Understanding and Generation with LLMs Paper • 2510.03255 • Published Sep 26, 2025 OpenTSLab/SciTS Preview • Updated Mar 19 • 5.57k • 4
SciTS: Scientific Time Series Understanding and Generation with LLMs Paper • 2510.03255 • Published Sep 26, 2025
Brain Signals BrainOmni: A Brain Foundation Model for Unified EEG and MEG Signals Paper • 2505.18185 • Published May 18, 2025 • 1 OpenTSLab/BrainOmni Updated Oct 15, 2025 • 5
BrainOmni: A Brain Foundation Model for Unified EEG and MEG Signals Paper • 2505.18185 • Published May 18, 2025 • 1
SPEAR SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations Paper • 2510.25955 • Published Oct 29, 2025 • 1 marcoyang/spear-xlarge-speech-audio Feature Extraction • 0.6B • Updated Jun 22 • 517 • 7
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations Paper • 2510.25955 • Published Oct 29, 2025 • 1
video-SALMONN 2 video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions. tsinghua-ee/video-SALMONN-2_plus_72B Updated Sep 28, 2025 • 18 • 2 tsinghua-ee/video_SALMONN2plus_72B_audioAlign 74B • Updated Jan 28 • 10 tsinghua-ee/video-SALMONN2_plus_7B_full 9B • Updated Feb 23 • 879 • 2 tsinghua-ee/video-SALMONN-2_plus_7B Updated Sep 28, 2025 • 24 • 7
video-SALMONN 2 Pro tsinghua-ee/video-SALMONN-2-Pro-4B 5B • Updated 26 days ago • 240 • 1 tsinghua-ee/video-SALMONN-2-Pro-8B 10B • Updated 26 days ago • 118 • 2 tsinghua-ee/video-SALMONN-2-Pro-32B 34B • Updated 26 days ago • 66 • 3
Spatial Audio & Visual Spatial Audio & Visual LLMs JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Paper • 2602.18527 • Published Feb 20 • 3 tsinghua-ee/JAEGER Updated Jun 23 • 14 • 4
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Paper • 2602.18527 • Published Feb 20 • 3
SPEAR SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations Paper • 2510.25955 • Published Oct 29, 2025 • 1 marcoyang/spear-xlarge-speech-audio Feature Extraction • 0.6B • Updated Jun 22 • 517 • 7
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations Paper • 2510.25955 • Published Oct 29, 2025 • 1
General Time Series SciTS: Scientific Time Series Understanding and Generation with LLMs Paper • 2510.03255 • Published Sep 26, 2025 OpenTSLab/SciTS Preview • Updated Mar 19 • 5.57k • 4
SciTS: Scientific Time Series Understanding and Generation with LLMs Paper • 2510.03255 • Published Sep 26, 2025
video-SALMONN 2 video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions. tsinghua-ee/video-SALMONN-2_plus_72B Updated Sep 28, 2025 • 18 • 2 tsinghua-ee/video_SALMONN2plus_72B_audioAlign 74B • Updated Jan 28 • 10 tsinghua-ee/video-SALMONN2_plus_7B_full 9B • Updated Feb 23 • 879 • 2 tsinghua-ee/video-SALMONN-2_plus_7B Updated Sep 28, 2025 • 24 • 7
Brain Signals BrainOmni: A Brain Foundation Model for Unified EEG and MEG Signals Paper • 2505.18185 • Published May 18, 2025 • 1 OpenTSLab/BrainOmni Updated Oct 15, 2025 • 5
BrainOmni: A Brain Foundation Model for Unified EEG and MEG Signals Paper • 2505.18185 • Published May 18, 2025 • 1