Model information (WIP)

The SpeechLMM 2.0 collection of multimodal and multilingual large language models is a collection of instruction-tuned generative models in 4 different sizes: S (1.7B), M (3B), L (7B) and XL (30B), supporting text, audio and video as input and text and audio as output. The SpeechLMM 2.0 models are optimized for various X-to-X generation tasks, namely:

  • Audio Chaptering (ACHAP)
  • Automatic Speech Recognition (ASR)
  • Audiovisual Speaker Diarization (AVSPEAKD)
  • Lip Reading (LIPREAD)
  • Machine Translation (MT)
  • Speech-to-speech Translation (S2ST)
  • Spoken Language Understanding - Intent (SLU-I)
  • Spoken Question Answering - Abstractive (SQA-A)
  • Spoken Question Answering - Extractive (SQA-E)
  • Spoken Question Answering - Multiple Choice (SQA-M)
  • Speech Summarization (SSUM)
  • Speech Translation (ST)
  • Textual Question Answering - Abstractive (TQA-A)
  • Textual Question Answering - Extractive (TQA-E)
  • Textual Question Answering - Multiple Choice (TQA-M)
  • Text Summarization (TSUM)
  • Text-to-Speech (TTS)

Model Developer: Meetween consortium

Supported Languages: Czech, German, English, Spanish, French, Hungarian, Italian, Dutch, Portuguese, Swedish are officially supported (for a subset of the supported tasks). The Qwen Omni backbone have been originally trained on a broader collection of languages than these 10 supported ones, therefore, the model might exhibit good performance on other languages too.

Model Release Date: July 31, 2026

License: Apache 2.0

Model Architecture (WIP)

SpeechLMM 2.0 is an auto-regressive multimodal language model based on a Qwen-X-Omni backbone (X varies with the model size) and incorporates a dedicated visual speech recognition pathway.

The architecture consists of:

  • Qwen3-Omni-30B-A3B-Instruct as the multimodal foundation model backbone;
  • The Qwen3 Thinker module for multimodal reasoning;
  • The Qwen3 Talker module for speech generation;
  • The Audio Transformer (AuT) encoder inherited from Qwen3-Omni for speech processing;
  • A SigLIP2-So400M-based [2] vision encoder;
  • An AutoAVSR-based [3] lip-reading encoder for extracting visual speech representations;
  • Trainable adapter layers connecting video, lip-reading and audio encoders outputs to the Qwen3 Thinker representation space.
Model Backbone Params Input modalities Output modalities Context Length
SpeechLMM 2.0 S --- 1.7B Multilingual text and audio, English video Multilingual Text ---
SpeechLMM 2.0 M Qwen-2.5-Omni-3B 3B Multilingual text and audio, English video Multilingual Text ---
SpeechLMM 2.0 L Qwen-2.5-Omni-7B 7B Multilingual text and audio, English video Multilingual Text ---
SpeechLMM 2.0 XL Qwen-3.0-Omni-30B-A3B-Instruct 30B Multilingual text and audio, English video Multilingual Text ---

Training Details (WIP)

How to use (WIP)

Refer to the instructions in our codebase:

Training Data (WIP)

TASK Dataset Language License
ACHAP AMI Meeting Corpus en → en -
YTSeg en → en -
ASR FLEURS cs → cs, de → de, en → en, es → es, fr → fr, hu → hu, it → it, nl → nl, pt → pt, sv → sv -
Multilingual Spoken Topical-Chat cs → cs, de → de, en → en, es → es, fr → fr, hu → hu, it → it, nl → nl, pt → pt, sv → sv -
Spoken DGT-TM cs → cs, de → de, en → en, es → es, fr → fr, hu → hu, it → it, nl → nl, pt → pt, sv → sv -
VoxPopuli cs → cs, de → de, en → en, es → es, fr → fr, hu → hu, it → it, nl → nl, pt → pt, sv → sv -
EuroSpeech cs → cs, de → de, en → en, es → es, fr → fr, hu → hu, it → it, nl → nl, pt → pt, sv → sv -
AVSPEAKD AMI Meeting Corpus en → en -
LIPREAD LipCrops en → en -
MT Europarl-ST {de, en, es, fr, it, nl, pt} → {de, en, es, fr, it, nl, pt} -
Multilingual Spoken Topical-Chat {cs, de, en, es, fr, hu, it, nl, pt, sv} → {cs, de, en, es, fr, hu, it, nl, pt, sv} -
Spoken DGT-TM {cs, de, en, es, fr, hu, it, nl, pt, sv} → {cs, de, en, es, fr, hu, it, nl, pt, sv} -
SLU-I SLURP en → en -
Speech-MASSIVE de → de, fr → fr -
SQA-E Multilingual Spoken SQuAD cs → cs, de → de, en → en, es → es, fr → fr, hu → hu, it → it, nl → nl, pt → pt, sv → sv -
SSUM ICSI en → en -
ST Europarl-ST {de, en, es, fr, it, nl, pt} → {de, en, es, fr, it, nl, pt} -
Multilingual Spoken Topical-Chat {cs, de, en, es, fr, hu, it, nl, pt, sv} → {cs, de, en, es, fr, hu, it, nl, pt, sv} -
Spoken DGT-TM {cs, de, en, es, fr, hu, it, nl, pt, sv} → {cs, de, en, es, fr, hu, it, nl, pt, sv} -
TQA-E Multilingual Spoken SQuAD cs → cs, de → de, en → en, es → es, fr → fr, hu → hu, it → it, nl → nl, pt → pt, sv → sv -
TQA-M Multilingual Text LibriSQA cs → cs, de → de, en → en, es → es, fr → fr, hu → hu, it → it, nl → nl, pt → pt, sv → sv -
TSUM AMI Meeting Corpus en → en -
ELITR Minuting Corpus cs → cs, en → en -

Evaluation Results (WIP)

The following results specifically refer to the L model.

ASR Metrics (WIP)

SLU Metrics (WIP)

SQA Metrics (WIP)

SSUM Metrics (WIP)

ST Metrics (WIP)

LIPREAD Metrics (WIP)

MT Metrics (WIP)

TSUM Metrics (WIP)

Framework versions (WIP)

Downloads last month
63
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for meetween/SpeechLMM-v2.0-XL-30B

Finetuned
(28)
this model