File size: 4,116 Bytes
26738d0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
---
library_name: transformers
tags:
- generated_from_trainer
datasets:
- /home/athuser/modelC_train/sft_modelC.jsonl
model-index:
- name: dev/shm/modelC_e2
  results: []
---

<!-- This model card has been generated automatically according to the information the Trainer had access to. You
should probably proofread and complete it, then remove this comment. -->

[<img src="https://raw.githubusercontent.com/axolotl-ai-cloud/axolotl/main/image/axolotl-badge-web.png" alt="Built with Axolotl" width="200" height="32"/>](https://github.com/axolotl-ai-cloud/axolotl)
<details><summary>See axolotl config</summary>

axolotl version: `0.12.2`
```yaml
# Model C epoch 2 (continuation from e1; constant LR makes this equivalent to a
# continuous 2-epoch run). PAD FIX: distinct pad token so EOS gets real labels —
# e1 never learned to end documents (pad==eos masked EOS from loss).
# Fresh prepared-dataset path so tokenization redoes with the new pad.
# Derived from /models/axolot/llama3_70b_fsdp.yaml (the out_FFT_E precedent);
# dataset swapped to Model C keepers in completion format (full-doc LM loss,
# both speakers, 15% header dropout baked into the jsonl by export_sft.py).
base_model: /models/modelC_out/70B_fft_e1
model_type: LlamaForCausalLM
tokenizer_type: AutoTokenizer
load_in_8bit: false
load_in_4bit: false
datasets:
  - path: /home/athuser/modelC_train/sft_modelC.jsonl
    type: completion
    field: text
dataset_prepared_path: /home/athuser/modelC_train/prepared_e2_padfix
val_set_size: 0.02
output_dir: /dev/shm/modelC_e2
sequence_len: 4096
sample_packing: true
tf32: true
gradient_accumulation_steps: 4
micro_batch_size: 1
num_epochs: 1
optimizer: adamw_torch_fused
lr_scheduler: constant_with_warmup
learning_rate: 2.0e-05
bf16: true
resume_from_checkpoint:
logging_steps: 1
flash_attention: true
warmup_ratio: 0.03
evals_per_epoch: 4
saves_per_epoch: 1
save_only_model: true
weight_decay: 0.0
ddp_backend: nccl
fsdp_version: 2
fsdp_config:
  offload_params: false
  cpu_ram_efficient_loading: true
  auto_wrap_policy: TRANSFORMER_BASED_WRAP
  transformer_layer_cls_to_wrap: LlamaDecoderLayer
  state_dict_type: FULL_STATE_DICT
  reshard_after_forward: true
  activation_checkpointing: true
special_tokens:
  pad_token: <|finetune_right_pad_id|>

```

</details><br>

# dev/shm/modelC_e2

This model was trained from scratch on the /home/athuser/modelC_train/sft_modelC.jsonl dataset.
It achieves the following results on the evaluation set:
- Loss: 1.4188
- Memory/max Mem Active(gib): 89.15
- Memory/max Mem Allocated(gib): 89.15
- Memory/device Mem Reserved(gib): 94.15

## Model description

More information needed

## Intended uses & limitations

More information needed

## Training and evaluation data

More information needed

## Training procedure

### Training hyperparameters

The following hyperparameters were used during training:
- learning_rate: 2e-05
- train_batch_size: 1
- eval_batch_size: 1
- seed: 42
- distributed_type: multi-GPU
- num_devices: 8
- gradient_accumulation_steps: 4
- total_train_batch_size: 32
- total_eval_batch_size: 8
- optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
- lr_scheduler_type: constant_with_warmup
- lr_scheduler_warmup_steps: 3
- training_steps: 113

### Training results

| Training Loss | Epoch  | Step | Validation Loss | Mem Active(gib) | Mem Allocated(gib) | Mem Reserved(gib) |
|:-------------:|:------:|:----:|:---------------:|:---------------:|:------------------:|:-----------------:|
| No log        | 0      | 0    | 1.4115          | 27.73           | 27.73              | 31.33             |
| 1.448         | 0.2549 | 29   | 1.4173          | 89.15           | 89.15              | 94.15             |
| 1.4109        | 0.5099 | 58   | 1.4175          | 89.15           | 89.15              | 94.15             |
| 1.4639        | 0.7648 | 87   | 1.4188          | 89.15           | 89.15              | 94.15             |


### Framework versions

- Transformers 4.55.2
- Pytorch 2.7.0+cu128
- Datasets 4.0.0
- Tokenizers 0.21.2