Tharun007 commited on Sep 19, 2025

Commit

d6e6ec4

verified ·

1 Parent(s): 8684641

Upload folder using huggingface_hub

Browse files

Files changed (44) hide show

.gitattributes +3 -0
qwen2-7b-commitpackft-lora-final/README.md +207 -0
qwen2-7b-commitpackft-lora-final/adapter_config.json +37 -0
qwen2-7b-commitpackft-lora-final/adapter_model.safetensors +3 -0
qwen2-7b-commitpackft-lora-final/added_tokens.json +5 -0
qwen2-7b-commitpackft-lora-final/chat_template.jinja +6 -0
qwen2-7b-commitpackft-lora-final/merges.txt +0 -0
qwen2-7b-commitpackft-lora-final/special_tokens_map.json +14 -0
qwen2-7b-commitpackft-lora-final/tokenizer.json +3 -0
qwen2-7b-commitpackft-lora-final/tokenizer_config.json +43 -0
qwen2-7b-commitpackft-lora-final/vocab.json +0 -0
qwen2-7b-commitpackft-lora/checkpoint-352/README.md +207 -0
qwen2-7b-commitpackft-lora/checkpoint-352/adapter_config.json +37 -0
qwen2-7b-commitpackft-lora/checkpoint-352/adapter_model.safetensors +3 -0
qwen2-7b-commitpackft-lora/checkpoint-352/added_tokens.json +5 -0
qwen2-7b-commitpackft-lora/checkpoint-352/chat_template.jinja +6 -0
qwen2-7b-commitpackft-lora/checkpoint-352/merges.txt +0 -0
qwen2-7b-commitpackft-lora/checkpoint-352/optimizer.pt +3 -0
qwen2-7b-commitpackft-lora/checkpoint-352/rng_state.pth +3 -0
qwen2-7b-commitpackft-lora/checkpoint-352/scaler.pt +3 -0
qwen2-7b-commitpackft-lora/checkpoint-352/scheduler.pt +3 -0
qwen2-7b-commitpackft-lora/checkpoint-352/special_tokens_map.json +14 -0
qwen2-7b-commitpackft-lora/checkpoint-352/tokenizer.json +3 -0
qwen2-7b-commitpackft-lora/checkpoint-352/tokenizer_config.json +43 -0
qwen2-7b-commitpackft-lora/checkpoint-352/trainer_state.json +83 -0
qwen2-7b-commitpackft-lora/checkpoint-352/training_args.bin +3 -0
qwen2-7b-commitpackft-lora/checkpoint-352/vocab.json +0 -0
qwen2-7b-commitpackft-lora/checkpoint-528/README.md +207 -0
qwen2-7b-commitpackft-lora/checkpoint-528/adapter_config.json +37 -0
qwen2-7b-commitpackft-lora/checkpoint-528/adapter_model.safetensors +3 -0
qwen2-7b-commitpackft-lora/checkpoint-528/added_tokens.json +5 -0
qwen2-7b-commitpackft-lora/checkpoint-528/chat_template.jinja +6 -0
qwen2-7b-commitpackft-lora/checkpoint-528/merges.txt +0 -0
qwen2-7b-commitpackft-lora/checkpoint-528/optimizer.pt +3 -0
qwen2-7b-commitpackft-lora/checkpoint-528/rng_state.pth +3 -0
qwen2-7b-commitpackft-lora/checkpoint-528/scaler.pt +3 -0
qwen2-7b-commitpackft-lora/checkpoint-528/scheduler.pt +3 -0
qwen2-7b-commitpackft-lora/checkpoint-528/special_tokens_map.json +14 -0
qwen2-7b-commitpackft-lora/checkpoint-528/tokenizer.json +3 -0
qwen2-7b-commitpackft-lora/checkpoint-528/tokenizer_config.json +43 -0
qwen2-7b-commitpackft-lora/checkpoint-528/trainer_state.json +104 -0
qwen2-7b-commitpackft-lora/checkpoint-528/training_args.bin +3 -0
qwen2-7b-commitpackft-lora/checkpoint-528/vocab.json +0 -0
qwen2-finetune.ipynb +522 -0

.gitattributes CHANGED Viewed

@@ -33,3 +33,6 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+qwen2-7b-commitpackft-lora/checkpoint-352/tokenizer.json filter=lfs diff=lfs merge=lfs -text
+qwen2-7b-commitpackft-lora/checkpoint-528/tokenizer.json filter=lfs diff=lfs merge=lfs -text
+qwen2-7b-commitpackft-lora-final/tokenizer.json filter=lfs diff=lfs merge=lfs -text

qwen2-7b-commitpackft-lora-final/README.md ADDED Viewed

	@@ -0,0 +1,207 @@

+---
+base_model: Qwen/Qwen2-7B-Instruct
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen2-7B-Instruct
+- lora
+- transformers
+---
+# Model Card for Model ID
+<!-- Provide a quick summary of what the model is/does. -->
+## Model Details
+### Model Description
+<!-- Provide a longer summary of what this model is. -->
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+### Model Sources [optional]
+<!-- Provide the basic links for the model. -->
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+## Uses
+<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
+### Direct Use
+<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
+[More Information Needed]
+### Downstream Use [optional]
+<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
+[More Information Needed]
+### Out-of-Scope Use
+<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
+[More Information Needed]
+## Bias, Risks, and Limitations
+<!-- This section is meant to convey both technical and sociotechnical limitations. -->
+[More Information Needed]
+### Recommendations
+<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+## How to Get Started with the Model
+Use the code below to get started with the model.
+[More Information Needed]
+## Training Details
+### Training Data
+<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
+[More Information Needed]
+### Training Procedure
+<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
+#### Preprocessing [optional]
+[More Information Needed]
+#### Training Hyperparameters
+- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
+#### Speeds, Sizes, Times [optional]
+<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
+[More Information Needed]
+## Evaluation
+<!-- This section describes the evaluation protocols and provides the results. -->
+### Testing Data, Factors & Metrics
+#### Testing Data
+<!-- This should link to a Dataset Card if possible. -->
+[More Information Needed]
+#### Factors
+<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
+[More Information Needed]
+#### Metrics
+<!-- These are the evaluation metrics being used, ideally with a description of why. -->
+[More Information Needed]
+### Results
+[More Information Needed]
+#### Summary
+## Model Examination [optional]
+<!-- Relevant interpretability work for the model goes here -->
+[More Information Needed]
+## Environmental Impact
+<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+## Technical Specifications [optional]
+### Model Architecture and Objective
+[More Information Needed]
+### Compute Infrastructure
+[More Information Needed]
+#### Hardware
+[More Information Needed]
+#### Software
+[More Information Needed]
+## Citation [optional]
+<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
+**BibTeX:**
+[More Information Needed]
+**APA:**
+[More Information Needed]
+## Glossary [optional]
+<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
+[More Information Needed]
+## More Information [optional]
+[More Information Needed]
+## Model Card Authors [optional]
+[More Information Needed]
+## Model Card Contact
+[More Information Needed]
+### Framework versions
+- PEFT 0.17.1

qwen2-7b-commitpackft-lora-final/adapter_config.json ADDED Viewed

	@@ -0,0 +1,37 @@

+{
+  "alpha_pattern": {},
+  "auto_mapping": null,
+  "base_model_name_or_path": "Qwen/Qwen2-7B-Instruct",
+  "bias": "none",
+  "corda_config": null,
+  "eva_config": null,
+  "exclude_modules": null,
+  "fan_in_fan_out": false,
+  "inference_mode": true,
+  "init_lora_weights": true,
+  "layer_replication": null,
+  "layers_pattern": null,
+  "layers_to_transform": null,
+  "loftq_config": {},
+  "lora_alpha": 32,
+  "lora_bias": false,
+  "lora_dropout": 0.05,
+  "megatron_config": null,
+  "megatron_core": "megatron.core",
+  "modules_to_save": null,
+  "peft_type": "LORA",
+  "qalora_group_size": 16,
+  "r": 16,
+  "rank_pattern": {},
+  "revision": null,
+  "target_modules": [
+    "q_proj",
+    "v_proj"
+  ],
+  "target_parameters": null,
+  "task_type": "CAUSAL_LM",
+  "trainable_token_indices": null,
+  "use_dora": false,
+  "use_qalora": false,
+  "use_rslora": false
+}

qwen2-7b-commitpackft-lora-final/adapter_model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:73daf2e5d58f6bc6c45949042c4d30ae952c1c941427e5d9f4e25d750c0ae45e
+size 20200056

qwen2-7b-commitpackft-lora-final/added_tokens.json ADDED Viewed

	@@ -0,0 +1,5 @@

+{
+  "<|endoftext|>": 151643,
+  "<|im_end|>": 151645,
+  "<|im_start|>": 151644
+}

qwen2-7b-commitpackft-lora-final/chat_template.jinja ADDED Viewed

	@@ -0,0 +1,6 @@

+{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system
+You are a helpful assistant.<|im_end|>
+' }}{% endif %}{{'<|im_start|>' + message['role'] + '
+' + message['content'] + '<|im_end|>' + '
+'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant
+' }}{% endif %}

qwen2-7b-commitpackft-lora-final/merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

qwen2-7b-commitpackft-lora-final/special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,14 @@

+{
+  "additional_special_tokens": [
+    "<|im_start|>",
+    "<|im_end|>"
+  ],
+  "eos_token": {
+    "content": "<|im_end|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": "<|im_end|>"
+}

qwen2-7b-commitpackft-lora-final/tokenizer.json ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:bcfe42da0a4497e8b2b172c1f9f4ec423a46dc12907f4349c55025f670422ba9
+size 11418266

qwen2-7b-commitpackft-lora-final/tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,43 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "151643": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151644": {
+      "content": "<|im_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151645": {
+      "content": "<|im_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "additional_special_tokens": [
+    "<|im_start|>",
+    "<|im_end|>"
+  ],
+  "bos_token": null,
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|im_end|>",
+  "errors": "replace",
+  "extra_special_tokens": {},
+  "model_max_length": 131072,
+  "pad_token": "<|im_end|>",
+  "split_special_tokens": false,
+  "tokenizer_class": "Qwen2Tokenizer",
+  "unk_token": null
+}

qwen2-7b-commitpackft-lora-final/vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff

qwen2-7b-commitpackft-lora/checkpoint-352/README.md ADDED Viewed

	@@ -0,0 +1,207 @@

+---
+base_model: Qwen/Qwen2-7B-Instruct
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen2-7B-Instruct
+- lora
+- transformers
+---
+# Model Card for Model ID
+<!-- Provide a quick summary of what the model is/does. -->
+## Model Details
+### Model Description
+<!-- Provide a longer summary of what this model is. -->
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+### Model Sources [optional]
+<!-- Provide the basic links for the model. -->
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+## Uses
+<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
+### Direct Use
+<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
+[More Information Needed]
+### Downstream Use [optional]
+<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
+[More Information Needed]
+### Out-of-Scope Use
+<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
+[More Information Needed]
+## Bias, Risks, and Limitations
+<!-- This section is meant to convey both technical and sociotechnical limitations. -->
+[More Information Needed]
+### Recommendations
+<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+## How to Get Started with the Model
+Use the code below to get started with the model.
+[More Information Needed]
+## Training Details
+### Training Data
+<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
+[More Information Needed]
+### Training Procedure
+<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
+#### Preprocessing [optional]
+[More Information Needed]
+#### Training Hyperparameters
+- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
+#### Speeds, Sizes, Times [optional]
+<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
+[More Information Needed]
+## Evaluation
+<!-- This section describes the evaluation protocols and provides the results. -->
+### Testing Data, Factors & Metrics
+#### Testing Data
+<!-- This should link to a Dataset Card if possible. -->
+[More Information Needed]
+#### Factors
+<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
+[More Information Needed]
+#### Metrics
+<!-- These are the evaluation metrics being used, ideally with a description of why. -->
+[More Information Needed]
+### Results
+[More Information Needed]
+#### Summary
+## Model Examination [optional]
+<!-- Relevant interpretability work for the model goes here -->
+[More Information Needed]
+## Environmental Impact
+<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+## Technical Specifications [optional]
+### Model Architecture and Objective
+[More Information Needed]
+### Compute Infrastructure
+[More Information Needed]
+#### Hardware
+[More Information Needed]
+#### Software
+[More Information Needed]
+## Citation [optional]
+<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
+**BibTeX:**
+[More Information Needed]
+**APA:**
+[More Information Needed]
+## Glossary [optional]
+<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
+[More Information Needed]
+## More Information [optional]
+[More Information Needed]
+## Model Card Authors [optional]
+[More Information Needed]
+## Model Card Contact
+[More Information Needed]
+### Framework versions
+- PEFT 0.17.1

qwen2-7b-commitpackft-lora/checkpoint-352/adapter_config.json ADDED Viewed

	@@ -0,0 +1,37 @@

+{
+  "alpha_pattern": {},
+  "auto_mapping": null,
+  "base_model_name_or_path": "Qwen/Qwen2-7B-Instruct",
+  "bias": "none",
+  "corda_config": null,
+  "eva_config": null,
+  "exclude_modules": null,
+  "fan_in_fan_out": false,
+  "inference_mode": true,
+  "init_lora_weights": true,
+  "layer_replication": null,
+  "layers_pattern": null,
+  "layers_to_transform": null,
+  "loftq_config": {},
+  "lora_alpha": 32,
+  "lora_bias": false,
+  "lora_dropout": 0.05,
+  "megatron_config": null,
+  "megatron_core": "megatron.core",
+  "modules_to_save": null,
+  "peft_type": "LORA",
+  "qalora_group_size": 16,
+  "r": 16,
+  "rank_pattern": {},
+  "revision": null,
+  "target_modules": [
+    "q_proj",
+    "v_proj"
+  ],
+  "target_parameters": null,
+  "task_type": "CAUSAL_LM",
+  "trainable_token_indices": null,
+  "use_dora": false,
+  "use_qalora": false,
+  "use_rslora": false
+}

qwen2-7b-commitpackft-lora/checkpoint-352/adapter_model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:16c3c11b05de2b97aaefdf7a68e10cf46accf46aa85264f140c9195dd57f15a8
+size 20200056

qwen2-7b-commitpackft-lora/checkpoint-352/added_tokens.json ADDED Viewed

	@@ -0,0 +1,5 @@

+{
+  "<|endoftext|>": 151643,
+  "<|im_end|>": 151645,
+  "<|im_start|>": 151644
+}

qwen2-7b-commitpackft-lora/checkpoint-352/chat_template.jinja ADDED Viewed

	@@ -0,0 +1,6 @@

+{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system
+You are a helpful assistant.<|im_end|>
+' }}{% endif %}{{'<|im_start|>' + message['role'] + '
+' + message['content'] + '<|im_end|>' + '
+'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant
+' }}{% endif %}

qwen2-7b-commitpackft-lora/checkpoint-352/merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

qwen2-7b-commitpackft-lora/checkpoint-352/optimizer.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:835152a734a83b9683275fa5ba1907304e3fd50d16a35cb503d5db7158b9e434
+size 40466443

qwen2-7b-commitpackft-lora/checkpoint-352/rng_state.pth ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:a1308f6b2fade250ef62416ffe329f2383295ca36bacfb10a7f84d8acd690afa
+size 14645

qwen2-7b-commitpackft-lora/checkpoint-352/scaler.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:edfe6f2786141265307b771141da9e10c3bd8b16f7cf5e238280fdd25f38e919
+size 1383

qwen2-7b-commitpackft-lora/checkpoint-352/scheduler.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:27f6d97bd79641ac48a4eca561192eeec39003cb5168dc16074a35d937944926
+size 1465

qwen2-7b-commitpackft-lora/checkpoint-352/special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,14 @@

+{
+  "additional_special_tokens": [
+    "<|im_start|>",
+    "<|im_end|>"
+  ],
+  "eos_token": {
+    "content": "<|im_end|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": "<|im_end|>"
+}

qwen2-7b-commitpackft-lora/checkpoint-352/tokenizer.json ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:bcfe42da0a4497e8b2b172c1f9f4ec423a46dc12907f4349c55025f670422ba9
+size 11418266

qwen2-7b-commitpackft-lora/checkpoint-352/tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,43 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "151643": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151644": {
+      "content": "<|im_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151645": {
+      "content": "<|im_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "additional_special_tokens": [
+    "<|im_start|>",
+    "<|im_end|>"
+  ],
+  "bos_token": null,
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|im_end|>",
+  "errors": "replace",
+  "extra_special_tokens": {},
+  "model_max_length": 131072,
+  "pad_token": "<|im_end|>",
+  "split_special_tokens": false,
+  "tokenizer_class": "Qwen2Tokenizer",
+  "unk_token": null
+}

qwen2-7b-commitpackft-lora/checkpoint-352/trainer_state.json ADDED Viewed

	@@ -0,0 +1,83 @@

+{
+  "best_global_step": null,
+  "best_metric": null,
+  "best_model_checkpoint": null,
+  "epoch": 2.0,
+  "eval_steps": 500,
+  "global_step": 352,
+  "is_hyper_param_search": false,
+  "is_local_process_zero": true,
+  "is_world_process_zero": true,
+  "log_history": [
+    {
+      "epoch": 0.28551034975017847,
+      "grad_norm": 0.1307421177625656,
+      "learning_rate": 0.00018143939393939395,
+      "loss": 0.6947,
+      "step": 50
+    },
+    {
+      "epoch": 0.5710206995003569,
+      "grad_norm": 0.13502074778079987,
+      "learning_rate": 0.00016250000000000002,
+      "loss": 0.5291,
+      "step": 100
+    },
+    {
+      "epoch": 0.8565310492505354,
+      "grad_norm": 0.11538226157426834,
+      "learning_rate": 0.00014356060606060607,
+      "loss": 0.5294,
+      "step": 150
+    },
+    {
+      "epoch": 1.1370449678800856,
+      "grad_norm": 0.12343299388885498,
+      "learning_rate": 0.00012462121212121211,
+      "loss": 0.5218,
+      "step": 200
+    },
+    {
+      "epoch": 1.422555317630264,
+      "grad_norm": 0.10965840518474579,
+      "learning_rate": 0.00010568181818181819,
+      "loss": 0.5029,
+      "step": 250
+    },
+    {
+      "epoch": 1.7080656673804424,
+      "grad_norm": 0.11586486548185349,
+      "learning_rate": 8.674242424242425e-05,
+      "loss": 0.5026,
+      "step": 300
+    },
+    {
+      "epoch": 1.993576017130621,
+      "grad_norm": 0.1560799926519394,
+      "learning_rate": 6.78030303030303e-05,
+      "loss": 0.5231,
+      "step": 350
+    }
+  ],
+  "logging_steps": 50,
+  "max_steps": 528,
+  "num_input_tokens_seen": 0,
+  "num_train_epochs": 3,
+  "save_steps": 500,
+  "stateful_callbacks": {
+    "TrainerControl": {
+      "args": {
+        "should_epoch_stop": false,
+        "should_evaluate": false,
+        "should_log": false,
+        "should_save": true,
+        "should_training_stop": false
+      },
+      "attributes": {}
+    }
+  },
+  "total_flos": 1.2176756003517235e+17,
+  "train_batch_size": 2,
+  "trial_name": null,
+  "trial_params": null
+}

qwen2-7b-commitpackft-lora/checkpoint-352/training_args.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:67d35dafe077b66520bf707bd5b750e11840108d94f0dbc9a06e3580c3a40c2a
+size 5777

qwen2-7b-commitpackft-lora/checkpoint-352/vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff

qwen2-7b-commitpackft-lora/checkpoint-528/README.md ADDED Viewed

	@@ -0,0 +1,207 @@

+---
+base_model: Qwen/Qwen2-7B-Instruct
+library_name: peft
+pipeline_tag: text-generation
+tags:
+- base_model:adapter:Qwen/Qwen2-7B-Instruct
+- lora
+- transformers
+---
+# Model Card for Model ID
+<!-- Provide a quick summary of what the model is/does. -->
+## Model Details
+### Model Description
+<!-- Provide a longer summary of what this model is. -->
+- **Developed by:** [More Information Needed]
+- **Funded by [optional]:** [More Information Needed]
+- **Shared by [optional]:** [More Information Needed]
+- **Model type:** [More Information Needed]
+- **Language(s) (NLP):** [More Information Needed]
+- **License:** [More Information Needed]
+- **Finetuned from model [optional]:** [More Information Needed]
+### Model Sources [optional]
+<!-- Provide the basic links for the model. -->
+- **Repository:** [More Information Needed]
+- **Paper [optional]:** [More Information Needed]
+- **Demo [optional]:** [More Information Needed]
+## Uses
+<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
+### Direct Use
+<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
+[More Information Needed]
+### Downstream Use [optional]
+<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
+[More Information Needed]
+### Out-of-Scope Use
+<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
+[More Information Needed]
+## Bias, Risks, and Limitations
+<!-- This section is meant to convey both technical and sociotechnical limitations. -->
+[More Information Needed]
+### Recommendations
+<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+## How to Get Started with the Model
+Use the code below to get started with the model.
+[More Information Needed]
+## Training Details
+### Training Data
+<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
+[More Information Needed]
+### Training Procedure
+<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
+#### Preprocessing [optional]
+[More Information Needed]
+#### Training Hyperparameters
+- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
+#### Speeds, Sizes, Times [optional]
+<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
+[More Information Needed]
+## Evaluation
+<!-- This section describes the evaluation protocols and provides the results. -->
+### Testing Data, Factors & Metrics
+#### Testing Data
+<!-- This should link to a Dataset Card if possible. -->
+[More Information Needed]
+#### Factors
+<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
+[More Information Needed]
+#### Metrics
+<!-- These are the evaluation metrics being used, ideally with a description of why. -->
+[More Information Needed]
+### Results
+[More Information Needed]
+#### Summary
+## Model Examination [optional]
+<!-- Relevant interpretability work for the model goes here -->
+[More Information Needed]
+## Environmental Impact
+<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
+Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
+- **Hardware Type:** [More Information Needed]
+- **Hours used:** [More Information Needed]
+- **Cloud Provider:** [More Information Needed]
+- **Compute Region:** [More Information Needed]
+- **Carbon Emitted:** [More Information Needed]
+## Technical Specifications [optional]
+### Model Architecture and Objective
+[More Information Needed]
+### Compute Infrastructure
+[More Information Needed]
+#### Hardware
+[More Information Needed]
+#### Software
+[More Information Needed]
+## Citation [optional]
+<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
+**BibTeX:**
+[More Information Needed]
+**APA:**
+[More Information Needed]
+## Glossary [optional]
+<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
+[More Information Needed]
+## More Information [optional]
+[More Information Needed]
+## Model Card Authors [optional]
+[More Information Needed]
+## Model Card Contact
+[More Information Needed]
+### Framework versions
+- PEFT 0.17.1

qwen2-7b-commitpackft-lora/checkpoint-528/adapter_config.json ADDED Viewed

	@@ -0,0 +1,37 @@

+{
+  "alpha_pattern": {},
+  "auto_mapping": null,
+  "base_model_name_or_path": "Qwen/Qwen2-7B-Instruct",
+  "bias": "none",
+  "corda_config": null,
+  "eva_config": null,
+  "exclude_modules": null,
+  "fan_in_fan_out": false,
+  "inference_mode": true,
+  "init_lora_weights": true,
+  "layer_replication": null,
+  "layers_pattern": null,
+  "layers_to_transform": null,
+  "loftq_config": {},
+  "lora_alpha": 32,
+  "lora_bias": false,
+  "lora_dropout": 0.05,
+  "megatron_config": null,
+  "megatron_core": "megatron.core",
+  "modules_to_save": null,
+  "peft_type": "LORA",
+  "qalora_group_size": 16,
+  "r": 16,
+  "rank_pattern": {},
+  "revision": null,
+  "target_modules": [
+    "q_proj",
+    "v_proj"
+  ],
+  "target_parameters": null,
+  "task_type": "CAUSAL_LM",
+  "trainable_token_indices": null,
+  "use_dora": false,
+  "use_qalora": false,
+  "use_rslora": false
+}

qwen2-7b-commitpackft-lora/checkpoint-528/adapter_model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:73daf2e5d58f6bc6c45949042c4d30ae952c1c941427e5d9f4e25d750c0ae45e
+size 20200056

qwen2-7b-commitpackft-lora/checkpoint-528/added_tokens.json ADDED Viewed

	@@ -0,0 +1,5 @@

+{
+  "<|endoftext|>": 151643,
+  "<|im_end|>": 151645,
+  "<|im_start|>": 151644
+}

qwen2-7b-commitpackft-lora/checkpoint-528/chat_template.jinja ADDED Viewed

	@@ -0,0 +1,6 @@

+{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system
+You are a helpful assistant.<|im_end|>
+' }}{% endif %}{{'<|im_start|>' + message['role'] + '
+' + message['content'] + '<|im_end|>' + '
+'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant
+' }}{% endif %}

qwen2-7b-commitpackft-lora/checkpoint-528/merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

qwen2-7b-commitpackft-lora/checkpoint-528/optimizer.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:142af714adc3574c15975a453aa271337ac75e7672b0f4ef5eb26181a50c66f9
+size 40466443

qwen2-7b-commitpackft-lora/checkpoint-528/rng_state.pth ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d423fa4d00cbc6a351787764a0cbce8eb5c10e46c3382024ce3dcd8648a5f641
+size 14645

qwen2-7b-commitpackft-lora/checkpoint-528/scaler.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f5fdd1a36f3bbbcea6b8062f359c0175b0022085d16d0e16e66eae10443c4cb3
+size 1383

qwen2-7b-commitpackft-lora/checkpoint-528/scheduler.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:2e49b91544c4a1784c4e83826d64bde15557e0f648672231532efef1b8eef895
+size 1465

qwen2-7b-commitpackft-lora/checkpoint-528/special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,14 @@

+{
+  "additional_special_tokens": [
+    "<|im_start|>",
+    "<|im_end|>"
+  ],
+  "eos_token": {
+    "content": "<|im_end|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": "<|im_end|>"
+}

qwen2-7b-commitpackft-lora/checkpoint-528/tokenizer.json ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:bcfe42da0a4497e8b2b172c1f9f4ec423a46dc12907f4349c55025f670422ba9
+size 11418266

qwen2-7b-commitpackft-lora/checkpoint-528/tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,43 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "151643": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151644": {
+      "content": "<|im_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151645": {
+      "content": "<|im_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "additional_special_tokens": [
+    "<|im_start|>",
+    "<|im_end|>"
+  ],
+  "bos_token": null,
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|im_end|>",
+  "errors": "replace",
+  "extra_special_tokens": {},
+  "model_max_length": 131072,
+  "pad_token": "<|im_end|>",
+  "split_special_tokens": false,
+  "tokenizer_class": "Qwen2Tokenizer",
+  "unk_token": null
+}

qwen2-7b-commitpackft-lora/checkpoint-528/trainer_state.json ADDED Viewed

	@@ -0,0 +1,104 @@

+{
+  "best_global_step": null,
+  "best_metric": null,
+  "best_model_checkpoint": null,
+  "epoch": 3.0,
+  "eval_steps": 500,
+  "global_step": 528,
+  "is_hyper_param_search": false,
+  "is_local_process_zero": true,
+  "is_world_process_zero": true,
+  "log_history": [
+    {
+      "epoch": 0.28551034975017847,
+      "grad_norm": 0.1307421177625656,
+      "learning_rate": 0.00018143939393939395,
+      "loss": 0.6947,
+      "step": 50
+    },
+    {
+      "epoch": 0.5710206995003569,
+      "grad_norm": 0.13502074778079987,
+      "learning_rate": 0.00016250000000000002,
+      "loss": 0.5291,
+      "step": 100
+    },
+    {
+      "epoch": 0.8565310492505354,
+      "grad_norm": 0.11538226157426834,
+      "learning_rate": 0.00014356060606060607,
+      "loss": 0.5294,
+      "step": 150
+    },
+    {
+      "epoch": 1.1370449678800856,
+      "grad_norm": 0.12343299388885498,
+      "learning_rate": 0.00012462121212121211,
+      "loss": 0.5218,
+      "step": 200
+    },
+    {
+      "epoch": 1.422555317630264,
+      "grad_norm": 0.10965840518474579,
+      "learning_rate": 0.00010568181818181819,
+      "loss": 0.5029,
+      "step": 250
+    },
+    {
+      "epoch": 1.7080656673804424,
+      "grad_norm": 0.11586486548185349,
+      "learning_rate": 8.674242424242425e-05,
+      "loss": 0.5026,
+      "step": 300
+    },
+    {
+      "epoch": 1.993576017130621,
+      "grad_norm": 0.1560799926519394,
+      "learning_rate": 6.78030303030303e-05,
+      "loss": 0.5231,
+      "step": 350
+    },
+    {
+      "epoch": 2.274089935760171,
+      "grad_norm": 0.1398245394229889,
+      "learning_rate": 4.886363636363637e-05,
+      "loss": 0.5195,
+      "step": 400
+    },
+    {
+      "epoch": 2.5596002855103497,
+      "grad_norm": 0.1260460615158081,
+      "learning_rate": 2.9924242424242427e-05,
+      "loss": 0.5049,
+      "step": 450
+    },
+    {
+      "epoch": 2.845110635260528,
+      "grad_norm": 0.11539369821548462,
+      "learning_rate": 1.0984848484848486e-05,
+      "loss": 0.4878,
+      "step": 500
+    }
+  ],
+  "logging_steps": 50,
+  "max_steps": 528,
+  "num_input_tokens_seen": 0,
+  "num_train_epochs": 3,
+  "save_steps": 500,
+  "stateful_callbacks": {
+    "TrainerControl": {
+      "args": {
+        "should_epoch_stop": false,
+        "should_evaluate": false,
+        "should_log": false,
+        "should_save": true,
+        "should_training_stop": true
+      },
+      "attributes": {}
+    }
+  },
+  "total_flos": 1.8265134005275853e+17,
+  "train_batch_size": 2,
+  "trial_name": null,
+  "trial_params": null
+}

qwen2-7b-commitpackft-lora/checkpoint-528/training_args.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:67d35dafe077b66520bf707bd5b750e11840108d94f0dbc9a06e3580c3a40c2a
+size 5777

qwen2-7b-commitpackft-lora/checkpoint-528/vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff

qwen2-finetune.ipynb ADDED Viewed

	@@ -0,0 +1,522 @@

+{
+ "cells": [
+  {
+   "cell_type": "markdown",
+   "id": "666ac2b7",
+   "metadata": {},
+   "source": [
+    "# Qwen2-7B-Instruct LoRA Fine-tuning with bigcode/commitpackft\n",
+    "\n",
+    "## Install required libraries"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 1,
+   "id": "457fec89",
+   "metadata": {},
+   "outputs": [
+    {
+     "name": "stdout",
+     "output_type": "stream",
+     "text": [
+      "Note: you may need to restart the kernel to use updated packages.\n"
+     ]
+    }
+   ],
+   "source": [
+    "%pip install -q transformers accelerate peft datasets bitsandbytes trl"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "a1fc5166",
+   "metadata": {},
+   "source": [
+    "## Imports"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 2,
+   "id": "aa2eba59",
+   "metadata": {},
+   "outputs": [
+    {
+     "name": "stderr",
+     "output_type": "stream",
+     "text": [
+      "c:\\Users\\Admin\\AppData\\Local\\Programs\\Python\\Python311\\Lib\\site-packages\\tqdm\\auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html\n",
+      "  from .autonotebook import tqdm as notebook_tqdm\n"
+     ]
+    },
+    {
+     "name": "stdout",
+     "output_type": "stream",
+     "text": [
+      "WARNING:tensorflow:From c:\\Users\\Admin\\AppData\\Local\\Programs\\Python\\Python311\\Lib\\site-packages\\keras\\src\\losses.py:2976: The name tf.losses.sparse_softmax_cross_entropy is deprecated. Please use tf.compat.v1.losses.sparse_softmax_cross_entropy instead.\n",
+      "\n"
+     ]
+    },
+    {
+     "name": "stderr",
+     "output_type": "stream",
+     "text": [
+      "W0919 09:55:13.094000 10452 site-packages\\torch\\distributed\\elastic\\multiprocessing\\redirects.py:29] NOTE: Redirects are currently not supported in Windows or MacOs.\n"
+     ]
+    }
+   ],
+   "source": [
+    "import torch\n",
+    "from datasets import load_dataset\n",
+    "from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, Trainer\n",
+    "from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "580fb870",
+   "metadata": {},
+   "source": [
+    "## 1. Load dataset"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 3,
+   "id": "3ed2eec5",
+   "metadata": {},
+   "outputs": [
+    {
+     "name": "stdout",
+     "output_type": "stream",
+     "text": [
+      "Dataset columns: ['commit', 'old_file', 'new_file', 'old_contents', 'new_contents', 'subject', 'message', 'lang', 'license', 'repos']\n",
+      "First example:\n",
+      "commit: e905334869af72025592de586b81650cb3468b8a\n",
+      "old_file: sentry/queue/client.py\n",
+      "new_file: sentry/queue/client.py\n",
+      "old_contents: \"\"\"\n",
+      "sentry.queue.client\n",
+      "~~~~~~~~~~~~~~~~~~~\n",
+      "\n",
+      ":copyright: (c) 2010 by the Sentry Team, see AUTHORS fo...\n",
+      "new_contents: \"\"\"\n",
+      "sentry.queue.client\n",
+      "~~~~~~~~~~~~~~~~~~~\n",
+      "\n",
+      ":copyright: (c) 2010 by the Sentry Team, see AUTHORS fo...\n",
+      "subject: Declare queues when broker is instantiated\n",
+      "message: Declare queues when broker is instantiated\n",
+      "\n",
+      "lang: Python\n",
+      "license: bsd-3-clause\n",
+      "repos: imankulov/sentry,BuildingLink/sentry,zenefits/sentry,korealerts1/sentry,kevinastone/sentry,fotinakis...\n"
+     ]
+    }
+   ],
+   "source": [
+    "# Load dataset with python config (you can choose another language if preferred)\n",
+    "dataset = load_dataset(\"bigcode/commitpackft\", \"python\", split=\"train[:5%]\")  # Using 5% of data to keep training time reasonable\n",
+    "\n",
+    "# Let's examine the dataset structure\n",
+    "print(\"Dataset columns:\", dataset.column_names)\n",
+    "print(\"First example:\")\n",
+    "for key, value in dataset[0].items():\n",
+    "    if isinstance(value, str) and len(value) > 100:\n",
+    "        print(f\"{key}: {value[:100]}...\")\n",
+    "    else:\n",
+    "        print(f\"{key}: {value}\")"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "aa276c12",
+   "metadata": {},
+   "source": [
+    "## 2. Load tokenizer & model (Qwen3-4B)"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 4,
+   "id": "7204f957",
+   "metadata": {},
+   "outputs": [],
+   "source": [
+    "# The correct model name for Qwen models\n",
+    "model_name = \"Qwen/Qwen2-7B-Instruct\"  # Using Qwen2 7B Instruct model\n",
+    "tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)\n",
+    "tokenizer.pad_token = tokenizer.eos_token"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "e29a46e0",
+   "metadata": {},
+   "source": [
+    "## Load in 4-bit quantized mode (saves VRAM)"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 5,
+   "id": "e0d06509",
+   "metadata": {},
+   "outputs": [
+    {
+     "name": "stderr",
+     "output_type": "stream",
+     "text": [
+      "The `load_in_4bit` and `load_in_8bit` arguments are deprecated and will be removed in the future versions. Please, pass a `BitsAndBytesConfig` object in `quantization_config` argument instead.\n",
+      "Loading checkpoint shards: 100%|██████████| 4/4 [00:12<00:00,  3.20s/it]\n",
+      "\n"
+     ]
+    }
+   ],
+   "source": [
+    "# Load model in 4-bit quantized mode to save VRAM\n",
+    "model = AutoModelForCausalLM.from_pretrained(\n",
+    "    model_name,\n",
+    "    device_map=\"auto\",\n",
+    "    load_in_4bit=True,\n",
+    "    trust_remote_code=True\n",
+    ")"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "b061b83c",
+   "metadata": {},
+   "source": [
+    "## 3. Prepare model for LoRA training"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 6,
+   "id": "35b3b2cd",
+   "metadata": {},
+   "outputs": [],
+   "source": [
+    "model = prepare_model_for_kbit_training(model)\n",
+    "\n",
+    "lora_config = LoraConfig(\n",
+    "    r=16,              # Rank\n",
+    "    lora_alpha=32,     \n",
+    "    target_modules=[\"q_proj\", \"v_proj\"],  # LoRA on attention layers\n",
+    "    lora_dropout=0.05,\n",
+    "    bias=\"none\",\n",
+    "    task_type=\"CAUSAL_LM\"\n",
+    ")\n",
+    "\n",
+    "model = get_peft_model(model, lora_config)"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "6d17bd85",
+   "metadata": {},
+   "source": [
+    "## 4. Tokenize dataset"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 7,
+   "id": "174e630b",
+   "metadata": {},
+   "outputs": [],
+   "source": [
+    "def tokenize_function(examples):\n",
+    "    # Create instruction-based prompts for code fixing\n",
+    "    # Use the correct column names based on the dataset structure\n",
+    "    # Commonly used names in commitpackft are \"old_contents\" and \"new_contents\"\n",
+    "    \n",
+    "    prompts = [\n",
+    "        f\"### Instruction:\\nFix the following buggy code:\\n{before}\\n\\n### Response:\\n{after}\"\n",
+    "        for before, after in zip(examples[\"old_contents\"], examples[\"new_contents\"])\n",
+    "    ]\n",
+    "    \n",
+    "    # Tokenize the prompts\n",
+    "    tokenized = tokenizer(\n",
+    "        prompts,\n",
+    "        padding=\"max_length\", \n",
+    "        truncation=True, \n",
+    "        max_length=512,\n",
+    "        return_tensors=\"pt\"\n",
+    "    )\n",
+    "    \n",
+    "    # For causal language modeling, labels are the input_ids\n",
+    "    tokenized[\"labels\"] = tokenized[\"input_ids\"].clone()\n",
+    "    \n",
+    "    return tokenized\n",
+    "\n",
+    "# After examining the dataset structure, apply the tokenization\n",
+    "tokenized_dataset = dataset.map(tokenize_function, batched=True, remove_columns=dataset.column_names)"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "2701bd75",
+   "metadata": {},
+   "source": [
+    "## 5. Training arguments"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 8,
+   "id": "3f0218bd",
+   "metadata": {},
+   "outputs": [
+    {
+     "name": "stdout",
+     "output_type": "stream",
+     "text": [
+      "Training configuration:\n",
+      "- Output directory: ./qwen2-7b-commitpackft-lora\n",
+      "- Batch size: 2 (x8 grad accum)\n",
+      "- Learning rate: 0.0002\n",
+      "- Epochs: 3\n",
+      "- FP16: True\n"
+     ]
+    }
+   ],
+   "source": [
+    "training_args = TrainingArguments(\n",
+    "    output_dir=\"./qwen2-7b-commitpackft-lora\",\n",
+    "    per_device_train_batch_size=2,\n",
+    "    gradient_accumulation_steps=8,\n",
+    "    num_train_epochs=3,\n",
+    "    learning_rate=2e-4,\n",
+    "    fp16=True,\n",
+    "    logging_steps=50,\n",
+    "    save_strategy=\"epoch\",\n",
+    "    # Removed evaluation_strategy parameter as it's not supported in this version\n",
+    "    save_total_limit=2,\n",
+    "    push_to_hub=False,\n",
+    "    report_to=\"none\"\n",
+    ")\n",
+    "\n",
+    "# Print training configuration for verification\n",
+    "print(f\"Training configuration:\")\n",
+    "print(f\"- Output directory: {training_args.output_dir}\")\n",
+    "print(f\"- Batch size: {training_args.per_device_train_batch_size} (x{training_args.gradient_accumulation_steps} grad accum)\")\n",
+    "print(f\"- Learning rate: {training_args.learning_rate}\")\n",
+    "print(f\"- Epochs: {training_args.num_train_epochs}\")\n",
+    "print(f\"- FP16: {training_args.fp16}\")"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "ff6fa420",
+   "metadata": {},
+   "source": [
+    "## 6. Trainer"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 9,
+   "id": "20744ef0",
+   "metadata": {},
+   "outputs": [
+    {
+     "name": "stderr",
+     "output_type": "stream",
+     "text": [
+      "C:\\Users\\Admin\\AppData\\Local\\Temp\\ipykernel_10452\\3424097219.py:1: FutureWarning: `tokenizer` is deprecated and will be removed in version 5.0.0 for `Trainer.__init__`. Use `processing_class` instead.\n",
+      "  trainer = Trainer(\n"
+     ]
+    }
+   ],
+   "source": [
+    "trainer = Trainer(\n",
+    "    model=model,\n",
+    "    args=training_args,\n",
+    "    train_dataset=tokenized_dataset,\n",
+    "    tokenizer=tokenizer\n",
+    ")"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "dd392bb1",
+   "metadata": {},
+   "source": [
+    "## 7. Start training"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 10,
+   "id": "32152c46",
+   "metadata": {},
+   "outputs": [
+    {
+     "name": "stderr",
+     "output_type": "stream",
+     "text": [
+      "The tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config. The model config and generation config were aligned accordingly, being updated with the tokenizer's values. Updated tokens: {'bos_token_id': None, 'pad_token_id': 151645}.\n",
+      "`use_cache=True` is incompatible with gradient checkpointing. Setting `use_cache=False`.\n",
+      "`use_cache=True` is incompatible with gradient checkpointing. Setting `use_cache=False`.\n",
+      "c:\\Users\\Admin\\AppData\\Local\\Programs\\Python\\Python311\\Lib\\site-packages\\torch\\_dynamo\\eval_frame.py:929: UserWarning: torch.utils.checkpoint: the use_reentrant parameter should be passed explicitly. In version 2.5 we will raise an exception if use_reentrant is not passed. use_reentrant=False is recommended, but if you need to preserve the current default behavior, you can pass use_reentrant=True. Refer to docs for more details on the differences between the two variants.\n",
+      "  return fn(*args, **kwargs)\n",
+      "c:\\Users\\Admin\\AppData\\Local\\Programs\\Python\\Python311\\Lib\\site-packages\\torch\\_dynamo\\eval_frame.py:929: UserWarning: torch.utils.checkpoint: the use_reentrant parameter should be passed explicitly. In version 2.5 we will raise an exception if use_reentrant is not passed. use_reentrant=False is recommended, but if you need to preserve the current default behavior, you can pass use_reentrant=True. Refer to docs for more details on the differences between the two variants.\n",
+      "  return fn(*args, **kwargs)\n"
+     ]
+    },
+    {
+     "data": {
+      "text/html": [
+       "\n",
+       "    <div>\n",
+       "      \n",
+       "      <progress value='528' max='528' style='width:300px; height:20px; vertical-align: middle;'></progress>\n",
+       "      [528/528 1:10:15, Epoch 3/3]\n",
+       "    </div>\n",
+       "    <table border=\"1\" class=\"dataframe\">\n",
+       "  <thead>\n",
+       " <tr style=\"text-align: left;\">\n",
+       "      <th>Step</th>\n",
+       "      <th>Training Loss</th>\n",
+       "    </tr>\n",
+       "  </thead>\n",
+       "  <tbody>\n",
+       "    <tr>\n",
+       "      <td>50</td>\n",
+       "      <td>0.694700</td>\n",
+       "    </tr>\n",
+       "    <tr>\n",
+       "      <td>100</td>\n",
+       "      <td>0.529100</td>\n",
+       "    </tr>\n",
+       "    <tr>\n",
+       "      <td>150</td>\n",
+       "      <td>0.529400</td>\n",
+       "    </tr>\n",
+       "    <tr>\n",
+       "      <td>200</td>\n",
+       "      <td>0.521800</td>\n",
+       "    </tr>\n",
+       "    <tr>\n",
+       "      <td>250</td>\n",
+       "      <td>0.502900</td>\n",
+       "    </tr>\n",
+       "    <tr>\n",
+       "      <td>300</td>\n",
+       "      <td>0.502600</td>\n",
+       "    </tr>\n",
+       "    <tr>\n",
+       "      <td>350</td>\n",
+       "      <td>0.523100</td>\n",
+       "    </tr>\n",
+       "    <tr>\n",
+       "      <td>400</td>\n",
+       "      <td>0.519500</td>\n",
+       "    </tr>\n",
+       "    <tr>\n",
+       "      <td>450</td>\n",
+       "      <td>0.504900</td>\n",
+       "    </tr>\n",
+       "    <tr>\n",
+       "      <td>500</td>\n",
+       "      <td>0.487800</td>\n",
+       "    </tr>\n",
+       "  </tbody>\n",
+       "</table><p>"
+      ],
+      "text/plain": [
+       "<IPython.core.display.HTML object>"
+      ]
+     },
+     "metadata": {},
+     "output_type": "display_data"
+    },
+    {
+     "name": "stderr",
+     "output_type": "stream",
+     "text": [
+      "c:\\Users\\Admin\\AppData\\Local\\Programs\\Python\\Python311\\Lib\\site-packages\\torch\\_dynamo\\eval_frame.py:929: UserWarning: torch.utils.checkpoint: the use_reentrant parameter should be passed explicitly. In version 2.5 we will raise an exception if use_reentrant is not passed. use_reentrant=False is recommended, but if you need to preserve the current default behavior, you can pass use_reentrant=True. Refer to docs for more details on the differences between the two variants.\n",
+      "  return fn(*args, **kwargs)\n",
+      "c:\\Users\\Admin\\AppData\\Local\\Programs\\Python\\Python311\\Lib\\site-packages\\torch\\_dynamo\\eval_frame.py:929: UserWarning: torch.utils.checkpoint: the use_reentrant parameter should be passed explicitly. In version 2.5 we will raise an exception if use_reentrant is not passed. use_reentrant=False is recommended, but if you need to preserve the current default behavior, you can pass use_reentrant=True. Refer to docs for more details on the differences between the two variants.\n",
+      "  return fn(*args, **kwargs)\n",
+      "c:\\Users\\Admin\\AppData\\Local\\Programs\\Python\\Python311\\Lib\\site-packages\\torch\\_dynamo\\eval_frame.py:929: UserWarning: torch.utils.checkpoint: the use_reentrant parameter should be passed explicitly. In version 2.5 we will raise an exception if use_reentrant is not passed. use_reentrant=False is recommended, but if you need to preserve the current default behavior, you can pass use_reentrant=True. Refer to docs for more details on the differences between the two variants.\n",
+      "  return fn(*args, **kwargs)\n"
+     ]
+    },
+    {
+     "data": {
+      "text/plain": [
+       "TrainOutput(global_step=528, training_loss=0.5297952763962023, metrics={'train_runtime': 4224.0462, 'train_samples_per_second': 1.989, 'train_steps_per_second': 0.125, 'total_flos': 1.8265134005275853e+17, 'train_loss': 0.5297952763962023, 'epoch': 3.0})"
+      ]
+     },
+     "execution_count": 10,
+     "metadata": {},
+     "output_type": "execute_result"
+    }
+   ],
+   "source": [
+    "trainer.train()"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "44e0c3df",
+   "metadata": {},
+   "source": [
+    "## 8. Save final LoRA adapter"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 11,
+   "id": "d9f0c7e2",
+   "metadata": {},
+   "outputs": [
+    {
+     "data": {
+      "text/plain": [
+       "('./qwen2-7b-commitpackft-lora-final\\\\tokenizer_config.json',\n",
+       " './qwen2-7b-commitpackft-lora-final\\\\special_tokens_map.json',\n",
+       " './qwen2-7b-commitpackft-lora-final\\\\chat_template.jinja',\n",
+       " './qwen2-7b-commitpackft-lora-final\\\\vocab.json',\n",
+       " './qwen2-7b-commitpackft-lora-final\\\\merges.txt',\n",
+       " './qwen2-7b-commitpackft-lora-final\\\\added_tokens.json',\n",
+       " './qwen2-7b-commitpackft-lora-final\\\\tokenizer.json')"
+      ]
+     },
+     "execution_count": 11,
+     "metadata": {},
+     "output_type": "execute_result"
+    }
+   ],
+   "source": [
+    "model.save_pretrained(\"./qwen2-7b-commitpackft-lora-final\")\n",
+    "tokenizer.save_pretrained(\"./qwen2-7b-commitpackft-lora-final\")"
+   ]
+  }
+ ],
+ "metadata": {
+  "kernelspec": {
+   "display_name": "Python 3",
+   "language": "python",
+   "name": "python3"
+  },
+  "language_info": {
+   "codemirror_mode": {
+    "name": "ipython",
+    "version": 3
+   },
+   "file_extension": ".py",
+   "mimetype": "text/x-python",
+   "name": "python",
+   "nbconvert_exporter": "python",
+   "pygments_lexer": "ipython3",
+   "version": "3.11.5"
+  }
+ },
+ "nbformat": 4,
+ "nbformat_minor": 5
+}