Text Generation
Transformers
Safetensors
Arabic
llama
arabic
reasoning
chain-of-thought
math
gsm8k
small-language-model
slm
sft
conversational
text-generation-inference
File size: 5,757 Bytes
867d0f3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
"""
Push the Arabic GSM8K reasoning dataset to the Hub as a PRIVATE dataset repo.

Usage: python push_dataset.py [--repo oddadmix/gsm8k-reasoning-ar] [--dry-run]
"""
import argparse
import json
from pathlib import Path

import pyarrow.parquet as pq
from huggingface_hub import HfApi

OUT = Path("out_gsm")
PARQUET = OUT / "gsm8k_reasoning_ar.parquet"
SOURCE = "Ajhesh7/gsm8k-reasoning-SFT-datas"
MT_MODEL = "ByteDance-Seed/Seed-X-PPO-7B"

CARD = """---
license: apache-2.0
language:
- ar
- en
task_categories:
- text-generation
tags:
- arabic
- reasoning
- chain-of-thought
- math
- gsm8k
- machine-translated
size_categories:
- 100K<n<1M
dataset_info:
  features:
  - name: text
    dtype: string
  - name: question
    dtype: string
  - name: thinking
    dtype: string
  - name: answer
    dtype: string
  - name: question_en
    dtype: string
  - name: thinking_en
    dtype: string
  - name: source_index
    dtype: int64
  splits:
  - name: train
    num_examples: {rows}
configs:
- config_name: default
  data_files:
  - split: train
    path: data/train-*
---

# GSM8K Reasoning — Arabic (مترجم آليًا)

**{rows:,}** grade-school math reasoning items translated from English into Arabic with
[`{mt}`](https://huggingface.co/{mt}), a 7B translation model.

Source: [`{source}`](https://huggingface.co/datasets/{source}) (600,000 rows).

> **بالعربية:** مجموعة بيانات للاستدلال الرياضي بالعربية، مترجمة آليًا من الإنجليزية.
> كل مثال يحتوي على سؤال، وخطوات التفكير، والإجابة النهائية.

## Format

`text` keeps the source's tag layout, with Arabic content:

```
<question>يجمع فريا 168 صندوقًا وجمعت هانا 19 صندوقًا...</question> <thinking>دعونا نفكر خطوة بخطوة...</thinking> <answer>187</answer>
```

The parts are also available as separate columns — `question`, `thinking`, `answer` (Arabic;
`answer` is the untouched numeral) — with `question_en` / `thinking_en` carrying the English
source so every row is auditable, and `source_index` pointing back into the source dataset.

## How it was built

1. **Sampling.** {selected:,} of the 600,000 source rows. The corpus is generated from only
   **2,814** underlying question patterns (numbers and names masked), so the sample is stratified
   by pattern with a floor of {floor} rows per pattern — every pattern is represented rather than
   over-weighting the common ones.
2. **Translation.** Question and reasoning translated separately, each as its own sentence, with
   `Translate the following English sentence into Arabic:\\n{{text}} <ar>` and greedy decoding.
   Numbers and names were left in place rather than masked, so Arabic gender agreement follows the
   actual name (`اشترت` for Aisha) and number agreement follows the actual quantity. The final
   `answer` numeral is never sent to the translator.
3. **Validation.** A row is kept only if, for **both** segments, the numbers in the Arabic exactly
   match the English (order-insensitive), the output is non-empty Arabic script, has no degenerate
   repetition loop, and has no significant Latin-script residue. **{kept_pct:.2f}%** of translated
   rows passed.

Rejection breakdown: `{rejects}`

## Limitations

This is **machine translation**, not human-verified Arabic. It inherits the source's synthetic,
templated phrasing — {selected:,} rows expand from 2,814 patterns, so linguistic diversity is far
lower than the row count suggests.

**Gender agreement.** The English source pairs names with pronouns arbitrarily ("This week Emil
did chores and earned $76. **She** bought a bottle…"), which English mostly hides but Arabic does
not: a row can read `قام جورج …` and then `اشترت …` for the same person. The translator rendered
the source faithfully; the disagreement is upstream, and it is visible throughout. The arithmetic itself is copied from the source and was not
re-verified; in the source, the reasoning's final number agrees with the `answer` field ~96.6% of
the time, so a small fraction of items are internally inconsistent. Suitable for SFT on reasoning
*format* and basic Arabic math phrasing; not a benchmark.
"""


def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--repo", default="oddadmix/gsm8k-reasoning-ar")
    ap.add_argument("--dry-run", action="store_true")
    args = ap.parse_args()

    stats = json.loads((OUT / "build_stats.json").read_text(encoding="utf-8"))
    rows = pq.ParquetFile(PARQUET).metadata.num_rows
    floor = max(1, stats["selected"] // (2814 * 4))

    card = CARD.format(
        rows=rows, mt=MT_MODEL, source=SOURCE, selected=stats["selected"],
        kept_pct=stats["kept_pct"], rejects=stats["rejects"], floor=floor,
    )
    (OUT / "README.md").write_text(card, encoding="utf-8")
    print(f"[+] wrote card ({len(card)} chars), {rows} rows")

    if args.dry_run:
        print("[dry-run] not pushing")
        return

    api = HfApi()
    api.create_repo(args.repo, repo_type="dataset", private=True, exist_ok=True)
    api.upload_file(path_or_fileobj=str(PARQUET), path_in_repo="data/train-00000-of-00001.parquet",
                    repo_id=args.repo, repo_type="dataset")
    api.upload_file(path_or_fileobj=str(OUT / "README.md"), path_in_repo="README.md",
                    repo_id=args.repo, repo_type="dataset")
    for script in ("gsm_common.py", "translate_gsm.py", "build_dataset.py"):
        api.upload_file(path_or_fileobj=script, path_in_repo=f"scripts/{script}",
                        repo_id=args.repo, repo_type="dataset")
    print(f"[+] https://huggingface.co/datasets/{args.repo}")


if __name__ == "__main__":
    main()