End of training

Browse files

Files changed (8) hide show

README.md +66 -0
adapter_config.json +17 -0
adapter_model.safetensors +3 -0
chat_template.jinja +45 -0
special_tokens_map.json +41 -0
tokenizer.json +0 -0
tokenizer_config.json +0 -0
training_args.bin +3 -0

README.md ADDED Viewed

	@@ -0,0 +1,66 @@

+---
+library_name: peft
+license: other
+base_model: tiiuae/Falcon3-1B-Instruct
+tags:
+- base_model:adapter:tiiuae/Falcon3-1B-Instruct
+- transformers
+pipeline_tag: text-generation
+model-index:
+- name: Falcon3-1B-Instruct-tuned
+  results: []
+---
+<!-- This model card has been generated automatically according to the information the Trainer had access to. You
+should probably proofread and complete it, then remove this comment. -->
+# Falcon3-1B-Instruct-tuned
+This model is a fine-tuned version of [tiiuae/Falcon3-1B-Instruct](https://huggingface.co/tiiuae/Falcon3-1B-Instruct) on an unknown dataset.
+It achieves the following results on the evaluation set:
+- Loss: 1.7228
+- Perplexity: 5.6001
+## Model description
+More information needed
+## Intended uses & limitations
+More information needed
+## Training and evaluation data
+More information needed
+## Training procedure
+### Training hyperparameters
+The following hyperparameters were used during training:
+- learning_rate: 0.003
+- train_batch_size: 12
+- eval_batch_size: 12
+- seed: 42
+- optimizer: Use OptimizerNames.ADAMW_TORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
+- lr_scheduler_type: linear
+- num_epochs: 3
+### Training results
+| Training Loss | Epoch  | Step | Validation Loss | Perplexity |
+|:-------------:|:------:|:----:|:---------------:|:----------:|
+| No log        | 0      | 0    | 3.5178          | 33.7116    |
+| No log        | 0.6011 | 333  | 1.8889          | 6.6122     |
+| 1.9189        | 1.2022 | 666  | 1.8092          | 6.1057     |
+| 1.9189        | 1.8032 | 999  | 1.7672          | 5.8541     |
+| 1.7617        | 2.4043 | 1332 | 1.7440          | 5.7203     |
+### Framework versions
+- PEFT 0.16.0
+- Transformers 4.54.1
+- Pytorch 2.7.1+cu128
+- Datasets 4.0.0
+- Tokenizers 0.21.4

adapter_config.json ADDED Viewed

	@@ -0,0 +1,17 @@

+{
+  "auto_mapping": null,
+  "base_model_name_or_path": "tiiuae/Falcon3-1B-Instruct",
+  "encoder_dropout": 0.0,
+  "encoder_hidden_size": 128,
+  "encoder_num_layers": 2,
+  "encoder_reparameterization_type": "MLP",
+  "inference_mode": true,
+  "num_attention_heads": 8,
+  "num_layers": 18,
+  "num_transformer_submodules": 1,
+  "num_virtual_tokens": 20,
+  "peft_type": "P_TUNING",
+  "revision": null,
+  "task_type": "CAUSAL_LM",
+  "token_dim": 2048
+}

adapter_model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:c947cf37af31eaafbb0fc2939cd2361a1bdebcf3228c8cd6d301abcf17dfe710
+size 163960

chat_template.jinja ADDED Viewed

	@@ -0,0 +1,45 @@

+{%- if tools %}
+{{- '<|system|>\n' }}
+{%- if messages[0]['role'] == 'system' %}
+{{- messages[0]['content'] }}
+{%- set remaining_messages = messages[1:] %}
+{%- else %}
+{%- set remaining_messages = messages %}
+{%- endif %}
+{{- 'You are a Falcon assistant skilled in function calling. You are helpful, respectful, and concise.\n\n# Tools\n\nYou have access to the following functions. You MUST use them to answer questions when needed. For each function call, you MUST return a JSON object inside <tool_call></tool_call> tags.\n\n<tools>' + tools|tojson(indent=2) + '</tools>\n\n# Output Format\n\nYour response MUST follow this format when making function calls:\n<tool_call>\n[\n  {"name": "function_name", "arguments": {"arg1": "value1", "arg2": "value2"}},\n  {"name": "another_function", "arguments": {"arg": "value"}}\n]\n</tool_call>\nIf no function calls are needed, respond normally without the tool_call tags.\n' }}
+{%- for message in remaining_messages %}
+{%- if message['role'] == 'user' %}
+{{- '<|user|>\n' + message['content'] + '\n' }}
+{%- elif message['role'] == 'assistant' %}
+{%- if message.content %}
+{{- '<|assistant|>\n' + message['content'] }}
+{%- endif %}
+{%- if message.tool_calls %}
+{{- '\n<tool_call>\n' }}
+{{- message.tool_calls|tojson(indent=2) }}
+{{- '\n</tool_call>' }}
+{%- endif %}
+{{- eos_token + '\n' }}
+{%- elif message['role'] == 'tool' %}
+{{- '<|assistant|>\n<tool_response>\n' + message['content'] + '\n</tool_response>\n' }}
+{%- endif %}
+{%- endfor %}
+{{- '<|assistant|>\n' if add_generation_prompt }}
+{%- else %}
+{%- for message in messages %}
+{%- if message['role'] == 'system' %}
+{{- '<|system|>\n' + message['content'] + '\n' }}
+{%- elif message['role'] == 'user' %}
+{{- '<|user|>\n' + message['content'] + '\n' }}
+{%- elif message['role'] == 'assistant' %}
+{%- if not loop.last %}
+{{- '<|assistant|>\n' + message['content'] + eos_token + '\n' }}
+{%- else %}
+{{- '<|assistant|>\n' + message['content'] + eos_token }}
+{%- endif %}
+{%- endif %}
+{%- if loop.last and add_generation_prompt %}
+{{- '<|assistant|>\n' }}
+{%- endif %}
+{%- endfor %}
+{%- endif %}

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,41 @@

+{
+  "additional_special_tokens": [
+    ">>TITLE<<",
+    ">>ABSTRACT<<",
+    ">>INTRODUCTION<<",
+    ">>SUMMARY<<",
+    ">>COMMENT<<",
+    ">>ANSWER<<",
+    ">>QUESTION<<",
+    ">>DOMAIN<<",
+    ">>EMAIL_ADDRESS<<",
+    ">>IP_ADDRESS<<",
+    "<|startoftext|>",
+    ">>IP_ADDRESS_0<<",
+    ">>IP_ADDRESS_1<<",
+    ">>IP_ADDRESS_2<<",
+    ">>IP_ADDRESS_3<<",
+    ">>IP_ADDRESS_4<<",
+    ">>IP_ADDRESS_5<<",
+    ">>IP_ADDRESS_6<<",
+    ">>IP_ADDRESS_7<<",
+    ">>IP_ADDRESS_8<<",
+    ">>IP_ADDRESS_9<<",
+    ">>PASSWORD<<",
+    ">>KEY<<"
+  ],
+  "eos_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": {
+    "content": "<|pad|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

The diff for this file is too large to render. See raw diff

training_args.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6d98584e0461e502f549b7bde34f31699e91dd101d049334a5151592f036cf72
+size 5841