Clemylia commited on Jan 9

Commit

99f6792

verified ·

1 Parent(s): 771ccd8

Upload folder using huggingface_hub

Browse files

Files changed (18) hide show

README.md +64 -0
added_tokens.json +3 -0
config.json +35 -0
generation_config.json +9 -0
merges.txt +0 -0
onnx/model.onnx +3 -0
onnx/model_bnb4.onnx +3 -0
onnx/model_fp16.onnx +3 -0
onnx/model_int8.onnx +3 -0
onnx/model_q4.onnx +3 -0
onnx/model_q4f16.onnx +3 -0
onnx/model_quantized.onnx +3 -0
onnx/model_uint8.onnx +3 -0
quantize_config.json +18 -0
special_tokens_map.json +30 -0
tokenizer.json +0 -0
tokenizer_config.json +37 -0
vocab.json +0 -0

README.md ADDED Viewed

	@@ -0,0 +1,64 @@

+---
+library_name: transformers.js
+license: other
+language:
+- fr
+pipeline_tag: text-generation
+tags:
+- concept of code
+- Codeur
+- Basique Codeur creatif
+- SLM
+base_model:
+- NaA-IA/Qsana-coder-base
+---
+# Qsana-coder-base (ONNX)
+This is an ONNX version of [NaA-IA/Qsana-coder-base](https://huggingface.co/NaA-IA/Qsana-coder-base). It was automatically converted and uploaded using [this Hugging Face Space](https://huggingface.co/spaces/onnx-community/convert-to-onnx).
+## Usage with Transformers.js
+See the pipeline documentation for `text-generation`: https://huggingface.co/docs/transformers.js/api/pipelines#module_pipelines.TextGenerationPipeline
+---
+# 🤖 Qsana-coder-base : L'Assistant de Pseudo-Code Éducatif
+![Qsana](http://www.image-heberg.fr/files/17637511953035484323.png)
+## 🌟 Mission & Positionnement
+**Qsana-coder-base** est un *Small Language Model* (SLM) conçu pour la **créativité conceptuelle** autour des bases du codage (Python, pseudocode). Il ne vise **PAS** à produire du code exécutable en production, mais à générer des **fragments de logique codée** pour des contextes éducatifs et de prototypage rapide.
+> 💡 **Le but n'est pas la validité syntaxique à 100%, mais la stimulation de la pensée logique et la visualisation des concepts de codage (variables, boucles, conditions) pour les débutants.**
+## 🎯 Cas d'Usage Principaux
+| Emojis | Cas d'Usage | Description |
+| :--- | :--- | :--- |
+| 🧑‍🏫 | **Outil Pédagogique** | Générer des exemples de code courts et thématiques pour les jeunes apprenants (enfants, collégiens) qui illustrent la *logique* d'une fonction, même si la syntaxe est "créative". |
+| 🧪 | **Prototypage Conceptuel** | Pour les développeurs qui veulent rapidement coucher sur le papier la *structure* d'une idée sans se soucier des détails syntaxiques stricts. |
+| ✍️ | **Génération de Pseudo-Code** | Produire des fragments de code qui se rapprochent du langage naturel et qui sont faciles à expliquer sans nécessiter un environnement de développement complet. |
+## ⚙️ Détails Techniques
+  * **Modèle de Base :** `lam-4-zero-f` (51M, Fine-Tuned)
+  * **Langage Principal :** Français / Pseudo-Code Python
+  * **Précision Syntaxique :** **Intentionalité Créative (Non Rigide)**. Les erreurs de syntaxe font partie du comportement attendu pour illustrer le concept de "code prototype".
+  * **Poids / Efficacité :** Optimisé pour une exécution locale rapide (SLM).
+## 🛑 Limitations et Comportement Attendu
+Veuillez noter le comportement intentionnel suivant de **Qsana-coder-base** :
+1.  **Non-Exécutable :** Le code généré **n'est pas destiné à être copié/collé et exécuté** sans correction.
+2.  **Créativité Lexicale :** Le modèle mélange parfois les opérateurs (`<=`, `!=`) et les mots-clés (`continue`, `print`) d'une manière qui n'est pas standard en Python. **Ceci est le résultat du *fine-tuning* visant la créativité.**
+3.  **Utilisation du Chat Template :** Pour obtenir les résultats les plus cohérents, il est **fortement recommandé** d'utiliser le Chat Template fourni.

added_tokens.json ADDED Viewed

	@@ -0,0 +1,3 @@

+{
+  "[PAD]": 50257
+}

config.json ADDED Viewed

	@@ -0,0 +1,35 @@

+{
+  "_attn_implementation_autoset": true,
+  "_name_or_path": "NaA-IA/Qsana-coder-base",
+  "activation_function": "gelu_new",
+  "architectures": [
+    "GPT2LMHeadModel"
+  ],
+  "attn_pdrop": 0.1,
+  "bos_token_id": 50256,
+  "dtype": "float32",
+  "embd_pdrop": 0.1,
+  "eos_token_id": 50256,
+  "initializer_range": 0.02,
+  "layer_norm_epsilon": 1e-05,
+  "model_type": "gpt2",
+  "n_embd": 512,
+  "n_head": 8,
+  "n_inner": null,
+  "n_layer": 8,
+  "n_positions": 128,
+  "pad_token_id": 50257,
+  "reorder_and_upcast_attn": false,
+  "resid_pdrop": 0.1,
+  "scale_attn_by_inverse_layer_idx": false,
+  "scale_attn_weights": true,
+  "summary_activation": null,
+  "summary_first_dropout": 0.1,
+  "summary_proj_to_labels": true,
+  "summary_type": "cls_index",
+  "summary_use_proj": true,
+  "torch_dtype": "float32",
+  "transformers_version": "4.49.0",
+  "use_cache": true,
+  "vocab_size": 50258
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,9 @@

+{
+  "_from_model_config": true,
+  "bos_token_id": 50256,
+  "eos_token_id": [
+    50256
+  ],
+  "pad_token_id": 50257,
+  "transformers_version": "4.49.0"
+}

merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

onnx/model.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:aa355ca084d8940198e8db78ca77247de56f98651450486f7d9dadede60d1300
+size 204300051

onnx/model_bnb4.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:fc8819f30f71433c3b222daddf9760e771e9095fa1a8e8ea5621cfc05d5a983b
+size 204300070

onnx/model_fp16.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:57e8fc701837895077f13337760116990820a89ae7476bb86f8900d075588971
+size 102268734

onnx/model_int8.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:0ed38ba0b636417c115a15ab56a6c10438f2c98083de3d644f516edbd93b27a6
+size 154386872

onnx/model_q4.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:fc8819f30f71433c3b222daddf9760e771e9095fa1a8e8ea5621cfc05d5a983b
+size 204300070

onnx/model_q4f16.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:c48c8d2e2da07b532327294a67639183be489cd7145439756ed473ad7191c774
+size 102268753

onnx/model_quantized.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:0ed38ba0b636417c115a15ab56a6c10438f2c98083de3d644f516edbd93b27a6
+size 154386872

onnx/model_uint8.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:c27daf1c2d0db82e76612248ca764c52a0c7471fa834958bbc7b0193afb15879
+size 154386888

quantize_config.json ADDED Viewed

	@@ -0,0 +1,18 @@

+{
+    "modes": [
+        "fp16",
+        "q8",
+        "int8",
+        "uint8",
+        "q4",
+        "q4f16",
+        "bnb4"
+    ],
+    "per_channel": false,
+    "reduce_range": false,
+    "block_size": null,
+    "is_symmetric": true,
+    "accuracy_level": null,
+    "quant_type": 1,
+    "op_block_list": null
+}

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,30 @@

+{
+  "bos_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": {
+    "content": "[PAD]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "unk_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,37 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "50256": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "50257": {
+      "content": "[PAD]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": "<|endoftext|>",
+  "chat_template": "{%- for message in messages -%}\n    {{ '### Instruction:\\n' + message.content + '\\n\\n' }}\n{%- endfor -%}\n{{ '### Response:\\n' }}",
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|endoftext|>",
+  "extra_special_tokens": {},
+  "max_length": 128,
+  "model_max_length": 1024,
+  "pad_to_multiple_of": null,
+  "pad_token": "[PAD]",
+  "pad_token_type_id": 0,
+  "padding_side": "right",
+  "stride": 0,
+  "tokenizer_class": "GPT2Tokenizer",
+  "truncation_side": "right",
+  "truncation_strategy": "longest_first",
+  "unk_token": "<|endoftext|>"
+}

vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff