Irfanuruchi's picture
Update README.md
f48b46b verified
|
Raw
History Blame Contribute Delete
4.29 kB
---
language:
- en
license: apache-2.0
base_model:
- Qwen/Qwen3-4B
pipeline_tag: text-generation
tags:
- qwen3
- computer-science
- software-engineering
- programming
- awq
- compressed-tensors
- int4
- w4a16
- code
- python
---
# Qwen3-4B-Computer-Science-AWQ
## Overview
Qwen3-4B-Computer-Science-AWQ is the AWQ-calibrated quantized release of **Qwen3-4B-Computer-Science**.
Weights are stored using the **Compressed-Tensors** format with **4-bit asymmetric group-wise quantization (W4A16)**. The checkpoint was produced using Activation-aware Weight Quantization (AWQ) calibration and validated by successful quantization, checksum verification, and CPU inference.
---
## Model Information
| Property | Value |
|-----------|-------|
| Base Model | Qwen/Qwen3-4B |
| Model Family | Qwen3-4B-Computer-Science |
| Quantization | AWQ |
| Storage Format | Compressed-Tensors |
| Weight Precision | INT4 |
| Activation Precision | FP16 / BF16 |
| Quantization Scheme | W4A16 |
| Group Size | 128 |
| Weight Quantization | Asymmetric |
| lm_head | Excluded from Quantization |
| Language | English |
| License | Apache-2.0 |
---
## Training Data
The base model was instruction tuned using permissively licensed datasets.
| Dataset | Configuration | License |
|---------|---------------|---------|
| HuggingFaceTB/smoltalk | smol-magpie-ultra | Apache-2.0 |
| agentica-org/DeepCoder-Preview-Dataset | primeintellect | MIT |
### Dataset Size
| Split | Samples |
|-------|---------:|
| Training | 60,989 |
| Evaluation | 512 |
---
## Intended Use
This model is intended for:
- Software engineering
- Programming
- Code generation
- Debugging
- Code review
- Algorithm implementation
- Computer science education
- General technical reasoning
---
## Quantization
This release was generated using Activation-aware Weight Quantization (AWQ).
The resulting checkpoint stores weights using packed 4-bit group-wise asymmetric quantization.
| Parameter | Value |
|-----------|-------|
| Weight Format | Packed INT4 |
| Group Size | 128 |
| Symmetric | No |
| Observer | memoryless_minmax |
| Compression Format | Compressed-Tensors |
---
## Runtime Compatibility
This checkpoint uses the **Compressed-Tensors** format.
It is intended for runtimes that support Compressed-Tensors models.
Validation performed for this release:
- Successful AWQ calibration
- Successful model serialization
- CPU inference
- SHA256 verification of release artifacts
Loading this checkpoint with standard Transformers may decompress weights during execution depending on the runtime and available hardware.
---
## Usage
### Transformers
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"Irfanuruchi/Qwen3-4B-Computer-Science-AWQ",
device_map="auto",
dtype="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
"Irfanuruchi/Qwen3-4B-Computer-Science-AWQ"
)
```
---
## Release Artifacts
```
model.safetensors
config.json
generation_config.json
recipe.yaml
tokenizer.json
tokenizer_config.json
chat_template.jinja
SHA256SUMS
LICENSE
README.md
```
---
## Integrity Verification
Every release artifact includes a SHA256 checksum.
Verify downloaded files:
```bash
sha256sum -c SHA256SUMS
```
---
## Validation
The published checkpoint was verified before release.
Completed validation:
- AWQ calibration completed successfully
- Quantized checkpoint generated successfully
- CPU inference completed successfully
- SHA256 checksums verified
---
## Limitations
- Quantization may affect output quality compared to the BF16 checkpoint.
- Runtime support depends on the inference engine.
- GPU memory requirements depend on whether the runtime executes directly on compressed weights or decompresses them during inference.
---
## License
Base model:
- Apache-2.0
Training datasets:
- Apache-2.0
- MIT
This repository is distributed under the Apache-2.0 License.
---
## Acknowledgements
- Alibaba Qwen Team
- Hugging Face
- vLLM Project
- LLM Compressor Project
- SmolTalk Contributors
- DeepCoder Contributors
---
## Citation
```bibtex
@software{uruci2026qwen3csawq,
title={Qwen3-4B-Computer-Science-AWQ},
author={Irfan Uruçi},
year={2026},
publisher={Hugging Face}
}
```