caT-MDC / README.md
cabbagel's picture
Add caT-MDC model card
5b5c57c verified
|
Raw
History Blame Contribute Delete
4.6 kB
---
library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE
pipeline_tag: image-text-to-text
base_model:
- Qwen/Qwen3.5-9B
tags:
- qwen3.5
- vision-language
- multimodal
- video
- reinforcement-learning
- gspo
- mars2
- mdc
language:
- en
- zh
---
# caT-MDC
`caT-MDC` is the model submitted by team **caT** to the **MDC track of the
MARS2 2026 Challenge**.
The model uses the Qwen3.5-9B multimodal architecture and was post-trained with
the team's cold-start and group-based reinforcement-learning pipeline. This
repository contains the complete merged model in Hugging Face Transformers
format rather than a LoRA adapter.
## Model Details
| Item | Description |
|---|---|
| Team | caT |
| Challenge | MARS2 2026 |
| Track | MDC |
| Backbone | Qwen3.5-9B |
| Architecture | `Qwen3_5ForConditionalGeneration` |
| Model type | Multimodal vision-language model |
| Weight format | Safetensors |
| Precision | BFloat16 |
| Training stage | Cold-start post-training followed by GSPO-stage reinforcement learning |
| Release format | Complete merged model |
## Training Summary
The released checkpoint is the selected MDC submission model. According to the
archived configuration in `args.json`, its reinforcement-learning stage used:
- learning rate: `1e-5`
- epochs: `1`
- per-device batch size: `4`
- gradient accumulation steps: `2`
- rollout generations per prompt: `8`
- maximum completion length: `8048`
- precision: BFloat16
- optimizer: fused AdamW
- learning-rate schedule: cosine
- experiment tracking: SwanLab and TensorBoard
The competition training dataset is not redistributed in this model repository.
## Repository Contents
- `config.json`: model architecture and configuration
- `generation_config.json`: default generation configuration
- `model-*.safetensors`: sharded model weights
- `model.safetensors.index.json`: weight index
- `preprocessor_config.json`: multimodal preprocessing configuration
- `processor_config.json`: processor configuration
- `tokenizer.json`: tokenizer
- `tokenizer_config.json`: tokenizer configuration
- `chat_template.jinja`: conversation template
- `args.json`: archived training arguments
## Installation
```bash
pip install -U transformers accelerate pillow
```
Qwen3.5 requires a recent Transformers version. Refer to the official
[Qwen3.5-9B model card](https://huggingface.co/Qwen/Qwen3.5-9B) for current
compatibility guidance.
## Loading the Model
```python
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
model_id = "cabbagel/caT-MDC"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
model_id,
dtype="auto",
device_map="auto",
)
print(model.__class__.__name__)
```
Expected model class:
```text
Qwen3_5ForConditionalGeneration
```
## Basic Text Inference
```python
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
model_id = "cabbagel/caT-MDC"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
model_id,
dtype="auto",
device_map="auto",
)
messages = [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Briefly describe your multimodal reasoning capabilities.",
}
],
}
]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
generated_ids = model.generate(**inputs, max_new_tokens=256)
output_ids = generated_ids[:, inputs["input_ids"].shape[1]:]
response = processor.batch_decode(
output_ids,
skip_special_tokens=True,
)[0]
print(response)
```
For image and video inputs, follow the multimodal message format documented in
the official Qwen3.5 model card.
## Intended Use
This model is released for:
- reproduction and verification of the caT MDC submission;
- research on multimodal understanding and reasoning;
- evaluation within the MARS2 MDC task setting.
## Limitations
- The model was optimized for the MDC competition setting and may not generalize
to unrelated tasks.
- The model may produce inaccurate or unsupported responses.
- No claim is made that the model is suitable for safety-critical or high-stakes
applications.
- Users should independently verify model outputs.
## Acknowledgements
This work builds on [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B). We
thank the Qwen team and the MARS2 2026 organizers.