File size: 5,881 Bytes
8ce2c21
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
---
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
language:
  - zh
  - en
tags:
  - qwen3
  - conversational
  - text-generation
  - safetensors
  - tensorplay
  - tensormind
  - preview
inference: true
---

<div align="center">
  <img src="./assets/poster.png" alt="TensorMind 1.5 Preview" width="100%" />
</div>

<h1 align="center">TensorMind 1.5 Preview</h1>

<p align="center">
  A compact, open-weight bilingual language model from <strong>TensorPlay AI</strong>.
</p>

<p align="center">
  <code>536.9M parameters</code> · <code>BF16</code> · <code>Qwen3 architecture</code> · <code>Apache-2.0</code>
</p>

## Overview

TensorMind 1.5 Preview is a lightweight Chinese/English conversational model for local experimentation, instruction-following research, and continued fine-tuning. It ships as a standard Transformers package with Safetensors weights, tokenizer, generation config, and chat template—no custom model code is required.

### Highlights

- **Compact deployment:** 536,941,568 parameters in a 1.00 GiB BF16 weight file.
- **Standard runtime:** native `Qwen3ForCausalLM` support in Transformers.
- **Chat ready:** bundled bilingual chat template with optional non-thinking generation.
- **Reproducible evaluation:** exact score data, vector figures, and rendering source are included.
- **Open release:** Apache-2.0 licensed weights for research and application prototyping.

## Quick start

```bash
pip install "transformers>=4.51.0" accelerate torch
```

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "AATensorPlay/TensorMind-1.5-preview"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "用三句话解释什么是强化学习。"},
]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.inference_mode():
    output = model.generate(
        **inputs,
        max_new_tokens=256,
        do_sample=True,
        temperature=0.7,
        top_p=0.8,
    )

new_tokens = output[0, inputs.input_ids.shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
```

For deterministic decoding, set `do_sample=False` and omit `temperature` and `top_p`.

## Model details

| Item | Value |
|---|---|
| Architecture | `Qwen3ForCausalLM` |
| Parameters | 536,941,568 |
| Precision | BF16 |
| Layers | 32 |
| Hidden size | 1,024 |
| Intermediate size | 4,096 |
| Attention heads / KV heads | 16 / 8 |
| Vocabulary | 32,768 |
| Position configuration | 10,240 positions¹ |
| Attention | Full attention; no sliding window |
| Weight tying | Enabled |
| Weight format | Safetensors |

¹ `max_position_embeddings=10240` is the exported configuration ceiling. This preview has not been evaluated on Needle, RULER, LongBench, or another effective long-context benchmark, so it should not be interpreted as a validated long-context quality claim.

## Evaluation

<div align="center">
  <img src="./assets/benchmark-suite.png" alt="TensorMind 1.5 Preview benchmark suite" width="100%" />
</div>

Scores use one matched protocol: `lm-eval 0.4.12`, SGLang `0.5.14`, zero-shot evaluation, full datasets, batch size 48, and fixed seeds.

| Model | CMMLU | AGIEval-CN | A-CLUE | C-Eval | TMMLU+ | Macro |
|---|---:|---:|---:|---:|---:|---:|
| **TensorMind 1.5 Preview** | 24.8834 | 32.3822 | 24.7282 | 23.2541 | 24.7272 | **25.9950** |

These are standalone release scores. Compare them with other models only when the harness, prompts, datasets, and decoding protocol are matched. The exact values are available in [`benchmark-results.json`](./assets/benchmark-results.json), with an additional [benchmark scorecard](./assets/benchmark-matrix.png).

## Training snapshot

| Item | Value |
|---|---:|
| Final alignment packing length | 512 tokens |
| Packed train / eval sequences | 8,429 / 422 |
| Optimization steps | 50 |
| Per-device train / eval batch | 28 / 28 |
| Peak learning rate | 5e-7 |
| Precision | BF16 |
| Train loss | 2.2760 |
| Eval loss, start → finish | 2.16935 → 2.16928 |

## Intended use

Good fits include:

- local and edge-oriented language-model experiments;
- Chinese/English chat prototypes;
- instruction-tuning and alignment research;
- reproducible inference, quantization, and serving studies.

This preview is **not** a safety-tuned production assistant, a high-stakes decision system, or a substitute for domain experts. Evaluate it for your task before deployment.

## Limitations

- At 0.5B scale, the model can hallucinate, miss instructions, and produce incorrect reasoning or arithmetic.
- Benchmark scores are protocol-specific and do not establish superiority over other releases.
- Effective long-context quality has not been measured; the position setting is a configuration value only.
- Safety, multilingual breadth, tool use, and production robustness have not received comprehensive evaluation.

## Checksums

| File | SHA-256 |
|---|---|
| `model.safetensors` | `b99f6c0448831a6cb2012b93793dd72c9e64e94deee15fd8c4d8206c056817d9` |
| `tokenizer.json` | `2b31db3618982e4b4425535b563b4f12f8b62f82114abf3ef67e363a310bc44b` |
| `chat_template.jinja` | `ed2c59f23b9eb551f51d870b3a2aa7f9e431d424f8e34f27e4012049e85ba814` |
| `config.json` | `bf6d70f9520550eec8aef7961b3f240a06b69a46e8c2bd273b8cb18d61a494f1` |

## License

TensorMind 1.5 Preview is released under the [Apache License 2.0](./LICENSE).

---

<div align="center">
  <img src="./assets/tensorplay-ai-logo.png" alt="TensorPlay AI" width="140" />
  <br />
  Built by TensorPlay AI
</div>