Instructions to use optimum-intel-internal-testing/tiny-random-qwen3-dflash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use optimum-intel-internal-testing/tiny-random-qwen3-dflash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="optimum-intel-internal-testing/tiny-random-qwen3-dflash", trust_remote_code=True)# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("optimum-intel-internal-testing/tiny-random-qwen3-dflash", trust_remote_code=True) model = AutoModel.from_pretrained("optimum-intel-internal-testing/tiny-random-qwen3-dflash", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| library_name: transformers | |
| tags: | |
| - qwen3 | |
| - dflash | |
| - tiny-random | |
| # Tiny Random Qwen3 DFlash b16 | |
| Tiny random DFlash draft model compatible with `optimum-intel-internal-testing/tiny-random-qwen3`. | |
| - Source remote-code template: `z-lab/Qwen3-4B-DFlash-b16` | |
| - Target hidden size: 32 | |
| - Target hidden layers: 2 | |
| - DFlash block size: 16 | |
| - DFlash target layer ids: `[0, 1]` | |