| nohup: ignoring input |
| 当前指标:ZD |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:09<00:27, 9.21s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:18<00:18, 9.20s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:27<00:09, 9.24s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:30<00:00, 6.57s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:30<00:00, 7.54s/it] |
| Setting `pad_token_id` to `eos_token_id`:128001 for open-end generation. |
| Once upon a time, a group of 3,000 people lived in a small village. The people were all very happy and had a peaceful life. One day, a visitor came to the village. He was a very handsome man, and the people of the village were very impressed by him. They wanted to know more about him, so they asked him to tell them about his life. |
| The visitor told the people of the village that he was from a very powerful and wealthy family. He said that he had everything |
| LlamaForCausalLM( |
| (model): LlamaModel( |
| (embed_tokens): Embedding(128256, 4096) |
| (layers): ModuleList( |
| (0-31): 32 x LlamaDecoderLayer( |
| (self_attn): LlamaAttention( |
| (q_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| (k_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (v_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (o_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| ) |
| (mlp): LlamaMLP( |
| (gate_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (up_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (down_proj): Linear(in_features=14336, out_features=4096, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| (post_attention_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| ) |
| ) |
| (norm): LlamaRMSNorm((4096,), eps=1e-05) |
| (rotary_emb): LlamaRotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=4096, out_features=128256, bias=False) |
| ) |
| config: |
| LlamaConfig { |
| "architectures": [ |
| "LlamaForCausalLM" |
| ], |
| "attention_bias": false, |
| "attention_dropout": 0.0, |
| "bos_token_id": 128000, |
| "dtype": "float16", |
| "eos_token_id": 128001, |
| "head_dim": 128, |
| "hidden_act": "silu", |
| "hidden_size": 4096, |
| "initializer_range": 0.02, |
| "intermediate_size": 14336, |
| "max_position_embeddings": 131072, |
| "mlp_bias": false, |
| "model_type": "llama", |
| "num_attention_heads": 32, |
| "num_hidden_layers": 32, |
| "num_key_value_heads": 8, |
| "pretraining_tp": 1, |
| "rms_norm_eps": 1e-05, |
| "rope_scaling": { |
| "factor": 8.0, |
| "high_freq_factor": 4.0, |
| "low_freq_factor": 1.0, |
| "original_max_position_embeddings": 8192, |
| "rope_type": "llama3" |
| }, |
| "rope_theta": 500000.0, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "vocab_size": 128256 |
| } |
|
|
| Processing layer 0 |
| ZD value of layer 0 |
| Processing layer 1 |
| ZD value of layer 1 |
| Processing layer 2 |
| ZD value of layer 2 |
| Processing layer 3 |
| ZD value of layer 3 |
| Processing layer 4 |
| ZD value of layer 4 |
| Processing layer 5 |
| ZD value of layer 5 |
| Processing layer 6 |
| ZD value of layer 6 |
| Processing layer 7 |
| ZD value of layer 7 |
| Processing layer 8 |
| ZD value of layer 8 |
| Processing layer 9 |
| ZD value of layer 9 |
| Processing layer 10 |
| ZD value of layer 10 |
| Processing layer 11 |
| ZD value of layer 11 |
| Processing layer 12 |
| ZD value of layer 12 |
| Processing layer 13 |
| ZD value of layer 13 |
| Processing layer 14 |
| ZD value of layer 14 |
| Processing layer 15 |
| ZD value of layer 15 |
| Processing layer 16 |
| ZD value of layer 16 |
| Processing layer 17 |
| ZD value of layer 17 |
| Processing layer 18 |
| ZD value of layer 18 |
| Processing layer 19 |
| ZD value of layer 19 |
| Processing layer 20 |
| ZD value of layer 20 |
| Processing layer 21 |
| ZD value of layer 21 |
| Processing layer 22 |
| ZD value of layer 22 |
| Processing layer 23 |
| ZD value of layer 23 |
| Processing layer 24 |
| ZD value of layer 24 |
| Processing layer 25 |
| ZD value of layer 25 |
| Processing layer 26 |
| ZD value of layer 26 |
| Processing layer 27 |
| ZD value of layer 27 |
| Processing layer 28 |
| ZD value of layer 28 |
| Processing layer 29 |
| ZD value of layer 29 |
| Processing layer 30 |
| ZD value of layer 30 |
| Processing layer 31 |
| ZD value of layer 31 |
| [(0.15277639031410217, 23), (0.15186314284801483, 22), (0.15127883851528168, 25), (0.15086320042610168, 28), (0.15078411996364594, 21), (0.15071263909339905, 24), (0.1505168378353119, 26), (0.15024136006832123, 29), (0.14989669620990753, 12), (0.14976008236408234, 20), (0.14974866807460785, 19), (0.14963041245937347, 18), (0.14953136444091797, 27), (0.14945906400680542, 3), (0.1493656039237976, 6), (0.14915773272514343, 7), (0.14816699922084808, 4), (0.14805446565151215, 16), (0.14798566699028015, 5), (0.14793498814105988, 11), (0.14778870344161987, 17), (0.1470077931880951, 13), (0.14679034054279327, 30), (0.14632324874401093, 14), (0.14624223113059998, 15), (0.14602778851985931, 2), (0.14589068293571472, 8), (0.14568227529525757, 10), (0.14557501673698425, 31), (0.14533302187919617, 9), (0.14315810799598694, 1), (0.1303202509880066, 0)] |
| metric_name ZD: [23, 22, 25, 28, 21, 24, 26, 29, 12, 20, 19, 18, 27, 3, 6, 7, 4, 16, 5, 11, 17, 13, 30, 14, 15, 2, 8, 10, 31, 9, 1, 0] |
| 当前指标:alpha |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:09<00:27, 9.33s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:18<00:18, 9.02s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:27<00:09, 9.03s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:29<00:00, 6.29s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:29<00:00, 7.32s/it] |
| Setting `pad_token_id` to `eos_token_id`:128001 for open-end generation. |
| Once upon a time there was a guy. He was a nice guy. He was a smart guy. He was a guy who was not very good at taking care of himself. He had a wife who loved him. He had a family who loved him. But he was not very good at taking care of himself. He was not very good at being a good husband. He was not very good at being a good father. He was not very good at being a good son. He was not very good at being |
| LlamaForCausalLM( |
| (model): LlamaModel( |
| (embed_tokens): Embedding(128256, 4096) |
| (layers): ModuleList( |
| (0-31): 32 x LlamaDecoderLayer( |
| (self_attn): LlamaAttention( |
| (q_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| (k_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (v_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (o_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| ) |
| (mlp): LlamaMLP( |
| (gate_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (up_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (down_proj): Linear(in_features=14336, out_features=4096, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| (post_attention_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| ) |
| ) |
| (norm): LlamaRMSNorm((4096,), eps=1e-05) |
| (rotary_emb): LlamaRotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=4096, out_features=128256, bias=False) |
| ) |
| config: |
| LlamaConfig { |
| "architectures": [ |
| "LlamaForCausalLM" |
| ], |
| "attention_bias": false, |
| "attention_dropout": 0.0, |
| "bos_token_id": 128000, |
| "dtype": "float16", |
| "eos_token_id": 128001, |
| "head_dim": 128, |
| "hidden_act": "silu", |
| "hidden_size": 4096, |
| "initializer_range": 0.02, |
| "intermediate_size": 14336, |
| "max_position_embeddings": 131072, |
| "mlp_bias": false, |
| "model_type": "llama", |
| "num_attention_heads": 32, |
| "num_hidden_layers": 32, |
| "num_key_value_heads": 8, |
| "pretraining_tp": 1, |
| "rms_norm_eps": 1e-05, |
| "rope_scaling": { |
| "factor": 8.0, |
| "high_freq_factor": 4.0, |
| "low_freq_factor": 1.0, |
| "original_max_position_embeddings": 8192, |
| "rope_type": "llama3" |
| }, |
| "rope_theta": 500000.0, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "vocab_size": 128256 |
| } |
|
|
| Processing layer 0 |
| alpha value of layer 0 |
| Processing layer 1 |
| alpha value of layer 1 |
| Processing layer 2 |
| alpha value of layer 2 |
| Processing layer 3 |
| alpha value of layer 3 |
| Processing layer 4 |
| alpha value of layer 4 |
| Processing layer 5 |
| alpha value of layer 5 |
| Processing layer 6 |
| alpha value of layer 6 |
| Processing layer 7 |
| alpha value of layer 7 |
| Processing layer 8 |
| alpha value of layer 8 |
| Processing layer 9 |
| alpha value of layer 9 |
| Processing layer 10 |
| alpha value of layer 10 |
| Processing layer 11 |
| alpha value of layer 11 |
| Processing layer 12 |
| alpha value of layer 12 |
| Processing layer 13 |
| alpha value of layer 13 |
| Processing layer 14 |
| alpha value of layer 14 |
| Processing layer 15 |
| alpha value of layer 15 |
| Processing layer 16 |
| alpha value of layer 16 |
| Processing layer 17 |
| alpha value of layer 17 |
| Processing layer 18 |
| alpha value of layer 18 |
| Processing layer 19 |
| alpha value of layer 19 |
| Processing layer 20 |
| alpha value of layer 20 |
| Processing layer 21 |
| alpha value of layer 21 |
| Processing layer 22 |
| alpha value of layer 22 |
| Processing layer 23 |
| alpha value of layer 23 |
| Processing layer 24 |
| alpha value of layer 24 |
| Processing layer 25 |
| alpha value of layer 25 |
| Processing layer 26 |
| alpha value of layer 26 |
| Processing layer 27 |
| alpha value of layer 27 |
| Processing layer 28 |
| alpha value of layer 28 |
| Processing layer 29 |
| alpha value of layer 29 |
| Processing layer 30 |
| alpha value of layer 30 |
| Processing layer 31 |
| alpha value of layer 31 |
| [(5.11770486831665, 28), (4.464050769805908, 29), (4.385020732879639, 23), (4.369636058807373, 2), (4.368300914764404, 22), (4.3416218757629395, 24), (4.305949687957764, 21), (4.280331611633301, 19), (4.195464134216309, 20), (4.1130595207214355, 16), (4.0612688064575195, 25), (4.042068004608154, 27), (4.032166957855225, 26), (4.020874977111816, 4), (3.978748083114624, 18), (3.8516407012939453, 17), (3.767652988433838, 5), (3.7508127689361572, 6), (3.678138017654419, 15), (3.645509958267212, 3), (3.6096770763397217, 1), (3.456982374191284, 30), (3.420032501220703, 7), (3.2777371406555176, 13), (3.2595152854919434, 31), (3.2005908489227295, 14), (3.189865827560425, 12), (3.048171281814575, 11), (3.0236618518829346, 8), (2.9705560207366943, 10), (2.91736102104187, 0), (2.772552013397217, 9)] |
| metric_name alpha: [28, 29, 23, 2, 22, 24, 21, 19, 20, 16, 25, 27, 26, 4, 18, 17, 5, 6, 15, 3, 1, 30, 7, 13, 31, 14, 12, 11, 8, 10, 0, 9] |
| 当前指标:alpha_hat |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:09<00:27, 9.12s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:17<00:17, 8.94s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:26<00:08, 8.85s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 6.20s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 7.21s/it] |
| Setting `pad_token_id` to `eos_token_id`:128001 for open-end generation. |
| Once upon a time, in the world of marketing, there were two mighty forces that ruled the digital landscape: Google and Facebook. These two giants held the key to success for businesses looking to make their mark online. |
| But then, something extraordinary happened. A third contender emerged, armed with an innovative approach and a passion for transforming the way businesses engage with their audience. It was the dawn of Instagram marketing, and it was set to change the game forever. |
| The Power of Visual Storytelling |
| One of the most captivating |
| LlamaForCausalLM( |
| (model): LlamaModel( |
| (embed_tokens): Embedding(128256, 4096) |
| (layers): ModuleList( |
| (0-31): 32 x LlamaDecoderLayer( |
| (self_attn): LlamaAttention( |
| (q_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| (k_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (v_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (o_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| ) |
| (mlp): LlamaMLP( |
| (gate_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (up_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (down_proj): Linear(in_features=14336, out_features=4096, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| (post_attention_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| ) |
| ) |
| (norm): LlamaRMSNorm((4096,), eps=1e-05) |
| (rotary_emb): LlamaRotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=4096, out_features=128256, bias=False) |
| ) |
| config: |
| LlamaConfig { |
| "architectures": [ |
| "LlamaForCausalLM" |
| ], |
| "attention_bias": false, |
| "attention_dropout": 0.0, |
| "bos_token_id": 128000, |
| "dtype": "float16", |
| "eos_token_id": 128001, |
| "head_dim": 128, |
| "hidden_act": "silu", |
| "hidden_size": 4096, |
| "initializer_range": 0.02, |
| "intermediate_size": 14336, |
| "max_position_embeddings": 131072, |
| "mlp_bias": false, |
| "model_type": "llama", |
| "num_attention_heads": 32, |
| "num_hidden_layers": 32, |
| "num_key_value_heads": 8, |
| "pretraining_tp": 1, |
| "rms_norm_eps": 1e-05, |
| "rope_scaling": { |
| "factor": 8.0, |
| "high_freq_factor": 4.0, |
| "low_freq_factor": 1.0, |
| "original_max_position_embeddings": 8192, |
| "rope_type": "llama3" |
| }, |
| "rope_theta": 500000.0, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "vocab_size": 128256 |
| } |
|
|
| Processing layer 0 |
| alpha_hat value of layer 0 |
| Processing layer 1 |
| alpha_hat value of layer 1 |
| Processing layer 2 |
| alpha_hat value of layer 2 |
| Processing layer 3 |
| alpha_hat value of layer 3 |
| Processing layer 4 |
| alpha_hat value of layer 4 |
| Processing layer 5 |
| alpha_hat value of layer 5 |
| Processing layer 6 |
| alpha_hat value of layer 6 |
| Processing layer 7 |
| alpha_hat value of layer 7 |
| Processing layer 8 |
| alpha_hat value of layer 8 |
| Processing layer 9 |
| alpha_hat value of layer 9 |
| Processing layer 10 |
| alpha_hat value of layer 10 |
| Processing layer 11 |
| alpha_hat value of layer 11 |
| Processing layer 12 |
| alpha_hat value of layer 12 |
| Processing layer 13 |
| alpha_hat value of layer 13 |
| Processing layer 14 |
| alpha_hat value of layer 14 |
| Processing layer 15 |
| alpha_hat value of layer 15 |
| Processing layer 16 |
| alpha_hat value of layer 16 |
| Processing layer 17 |
| alpha_hat value of layer 17 |
| Processing layer 18 |
| alpha_hat value of layer 18 |
| Processing layer 19 |
| alpha_hat value of layer 19 |
| Processing layer 20 |
| alpha_hat value of layer 20 |
| Processing layer 21 |
| alpha_hat value of layer 21 |
| Processing layer 22 |
| alpha_hat value of layer 22 |
| Processing layer 23 |
| alpha_hat value of layer 23 |
| Processing layer 24 |
| alpha_hat value of layer 24 |
| Processing layer 25 |
| alpha_hat value of layer 25 |
| Processing layer 26 |
| alpha_hat value of layer 26 |
| Processing layer 27 |
| alpha_hat value of layer 27 |
| Processing layer 28 |
| alpha_hat value of layer 28 |
| Processing layer 29 |
| alpha_hat value of layer 29 |
| Processing layer 30 |
| alpha_hat value of layer 30 |
| Processing layer 31 |
| alpha_hat value of layer 31 |
| [(16.43989372253418, 28), (14.465442657470703, 29), (14.207252502441406, 27), (13.2962064743042, 31), (13.146225929260254, 23), (13.113412857055664, 20), (13.112133026123047, 22), (13.104711532592773, 21), (12.81999683380127, 24), (12.806440353393555, 26), (12.788350105285645, 30), (12.653292655944824, 25), (12.637839317321777, 2), (12.502681732177734, 19), (12.156525611877441, 18), (11.846827507019043, 17), (11.685830116271973, 4), (11.275540351867676, 16), (11.184784889221191, 6), (10.987555503845215, 5), (10.852300643920898, 0), (10.767301559448242, 15), (10.484479904174805, 7), (10.44900131225586, 3), (10.145772933959961, 1), (9.989272117614746, 12), (9.591448783874512, 13), (9.408217430114746, 8), (9.372644424438477, 14), (9.301414489746094, 10), (9.262565612792969, 11), (8.966535568237305, 9)] |
| metric_name alpha_hat: [28, 29, 27, 31, 23, 20, 22, 21, 24, 26, 30, 25, 2, 19, 18, 17, 4, 16, 6, 5, 0, 15, 7, 3, 1, 12, 13, 8, 14, 10, 11, 9] |
| 当前指标:stable_rank |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:09<00:27, 9.05s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:18<00:18, 9.13s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:27<00:09, 9.05s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:29<00:00, 6.31s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:29<00:00, 7.32s/it] |
| Setting `pad_token_id` to `eos_token_id`:128001 for open-end generation. |
| Once upon a time, there was a little girl who was very, very naughty. She was always making a mess and being a bother to her parents. One day, her parents decided that they had had enough and they sent her to live with her Aunt for a while. The girl was very sad to leave her home, but she knew that she had to go. |
| When she arrived at her Aunt’s house, she was surprised to find that it was much different than her own home. The girl had never seen such |
| LlamaForCausalLM( |
| (model): LlamaModel( |
| (embed_tokens): Embedding(128256, 4096) |
| (layers): ModuleList( |
| (0-31): 32 x LlamaDecoderLayer( |
| (self_attn): LlamaAttention( |
| (q_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| (k_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (v_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (o_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| ) |
| (mlp): LlamaMLP( |
| (gate_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (up_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (down_proj): Linear(in_features=14336, out_features=4096, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| (post_attention_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| ) |
| ) |
| (norm): LlamaRMSNorm((4096,), eps=1e-05) |
| (rotary_emb): LlamaRotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=4096, out_features=128256, bias=False) |
| ) |
| config: |
| LlamaConfig { |
| "architectures": [ |
| "LlamaForCausalLM" |
| ], |
| "attention_bias": false, |
| "attention_dropout": 0.0, |
| "bos_token_id": 128000, |
| "dtype": "float16", |
| "eos_token_id": 128001, |
| "head_dim": 128, |
| "hidden_act": "silu", |
| "hidden_size": 4096, |
| "initializer_range": 0.02, |
| "intermediate_size": 14336, |
| "max_position_embeddings": 131072, |
| "mlp_bias": false, |
| "model_type": "llama", |
| "num_attention_heads": 32, |
| "num_hidden_layers": 32, |
| "num_key_value_heads": 8, |
| "pretraining_tp": 1, |
| "rms_norm_eps": 1e-05, |
| "rope_scaling": { |
| "factor": 8.0, |
| "high_freq_factor": 4.0, |
| "low_freq_factor": 1.0, |
| "original_max_position_embeddings": 8192, |
| "rope_type": "llama3" |
| }, |
| "rope_theta": 500000.0, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "vocab_size": 128256 |
| } |
|
|
| Processing layer 0 |
| frobenius_norm tensor(76.8856, device='cuda:6') |
| spectral_norm tensor(32.1192, device='cuda:6') |
| frobenius_norm tensor(55.4457, device='cuda:6') |
| spectral_norm tensor(18.8411, device='cuda:6') |
| frobenius_norm tensor(14.8614, device='cuda:6') |
| spectral_norm tensor(1.9605, device='cuda:6') |
| frobenius_norm tensor(34.3062, device='cuda:6') |
| spectral_norm tensor(4.9343, device='cuda:6') |
| frobenius_norm tensor(98.5077, device='cuda:6') |
| spectral_norm tensor(10.0864, device='cuda:6') |
| frobenius_norm tensor(90.2663, device='cuda:6') |
| spectral_norm tensor(5.6548, device='cuda:6') |
| frobenius_norm tensor(90.4429, device='cuda:6') |
| spectral_norm tensor(6.7945, device='cuda:6') |
| stable_rank value of layer 0 |
| Processing layer 1 |
| frobenius_norm tensor(79.7186, device='cuda:6') |
| spectral_norm tensor(12.8741, device='cuda:6') |
| frobenius_norm tensor(56.4302, device='cuda:6') |
| spectral_norm tensor(9.5045, device='cuda:6') |
| frobenius_norm tensor(18.3755, device='cuda:6') |
| spectral_norm tensor(1.3126, device='cuda:6') |
| frobenius_norm tensor(39.3655, device='cuda:6') |
| spectral_norm tensor(5.1106, device='cuda:6') |
| frobenius_norm tensor(99.7901, device='cuda:6') |
| spectral_norm tensor(7.4131, device='cuda:6') |
| frobenius_norm tensor(91.8738, device='cuda:6') |
| spectral_norm tensor(3.5568, device='cuda:6') |
| frobenius_norm tensor(91.7333, device='cuda:6') |
| spectral_norm tensor(5.1422, device='cuda:6') |
| stable_rank value of layer 1 |
| Processing layer 2 |
| frobenius_norm tensor(76.9517, device='cuda:6') |
| spectral_norm tensor(9.4978, device='cuda:6') |
| frobenius_norm tensor(57.6473, device='cuda:6') |
| spectral_norm tensor(7.2243, device='cuda:6') |
| frobenius_norm tensor(15.2013, device='cuda:6') |
| spectral_norm tensor(1.0522, device='cuda:6') |
| frobenius_norm tensor(36.0830, device='cuda:6') |
| spectral_norm tensor(4.5946, device='cuda:6') |
| frobenius_norm tensor(102.2242, device='cuda:6') |
| spectral_norm tensor(8.5718, device='cuda:6') |
| frobenius_norm tensor(91.4413, device='cuda:6') |
| spectral_norm tensor(3.3492, device='cuda:6') |
| frobenius_norm tensor(91.8334, device='cuda:6') |
| spectral_norm tensor(5.1150, device='cuda:6') |
| stable_rank value of layer 2 |
| Processing layer 3 |
| frobenius_norm tensor(77.2064, device='cuda:6') |
| spectral_norm tensor(10.1484, device='cuda:6') |
| frobenius_norm tensor(57.4959, device='cuda:6') |
| spectral_norm tensor(7.5375, device='cuda:6') |
| frobenius_norm tensor(17.8962, device='cuda:6') |
| spectral_norm tensor(1.0208, device='cuda:6') |
| frobenius_norm tensor(41.1775, device='cuda:6') |
| spectral_norm tensor(4.3115, device='cuda:6') |
| frobenius_norm tensor(107.6994, device='cuda:6') |
| spectral_norm tensor(9.9873, device='cuda:6') |
| frobenius_norm tensor(90.0023, device='cuda:6') |
| spectral_norm tensor(3.0939, device='cuda:6') |
| frobenius_norm tensor(89.8130, device='cuda:6') |
| spectral_norm tensor(4.4840, device='cuda:6') |
| stable_rank value of layer 3 |
| Processing layer 4 |
| frobenius_norm tensor(76.3690, device='cuda:6') |
| spectral_norm tensor(9.9006, device='cuda:6') |
| frobenius_norm tensor(56.6955, device='cuda:6') |
| spectral_norm tensor(7.6416, device='cuda:6') |
| frobenius_norm tensor(19.1979, device='cuda:6') |
| spectral_norm tensor(1.1709, device='cuda:6') |
| frobenius_norm tensor(41.9950, device='cuda:6') |
| spectral_norm tensor(3.6805, device='cuda:6') |
| frobenius_norm tensor(112.9379, device='cuda:6') |
| spectral_norm tensor(11.9836, device='cuda:6') |
| frobenius_norm tensor(88.0967, device='cuda:6') |
| spectral_norm tensor(3.2062, device='cuda:6') |
| frobenius_norm tensor(87.9027, device='cuda:6') |
| spectral_norm tensor(4.8592, device='cuda:6') |
| stable_rank value of layer 4 |
| Processing layer 5 |
| frobenius_norm tensor(76.0779, device='cuda:6') |
| spectral_norm tensor(9.3366, device='cuda:6') |
| frobenius_norm tensor(56.8678, device='cuda:6') |
| spectral_norm tensor(7.5451, device='cuda:6') |
| frobenius_norm tensor(16.0436, device='cuda:6') |
| spectral_norm tensor(1.0232, device='cuda:6') |
| frobenius_norm tensor(38.9231, device='cuda:6') |
| spectral_norm tensor(2.9655, device='cuda:6') |
| frobenius_norm tensor(112.7281, device='cuda:6') |
| spectral_norm tensor(12.0050, device='cuda:6') |
| frobenius_norm tensor(88.6070, device='cuda:6') |
| spectral_norm tensor(3.5774, device='cuda:6') |
| frobenius_norm tensor(88.3036, device='cuda:6') |
| spectral_norm tensor(4.7503, device='cuda:6') |
| stable_rank value of layer 5 |
| Processing layer 6 |
| frobenius_norm tensor(77.5564, device='cuda:6') |
| spectral_norm tensor(9.4346, device='cuda:6') |
| frobenius_norm tensor(56.8842, device='cuda:6') |
| spectral_norm tensor(8.2064, device='cuda:6') |
| frobenius_norm tensor(17.3765, device='cuda:6') |
| spectral_norm tensor(1.1139, device='cuda:6') |
| frobenius_norm tensor(40.5982, device='cuda:6') |
| spectral_norm tensor(3.0531, device='cuda:6') |
| frobenius_norm tensor(113.1005, device='cuda:6') |
| spectral_norm tensor(12.2282, device='cuda:6') |
| frobenius_norm tensor(88.5713, device='cuda:6') |
| spectral_norm tensor(3.6645, device='cuda:6') |
| frobenius_norm tensor(88.2304, device='cuda:6') |
| spectral_norm tensor(5.4514, device='cuda:6') |
| stable_rank value of layer 6 |
| Processing layer 7 |
| frobenius_norm tensor(74.5972, device='cuda:6') |
| spectral_norm tensor(9.6876, device='cuda:6') |
| frobenius_norm tensor(57.3617, device='cuda:6') |
| spectral_norm tensor(8.4875, device='cuda:6') |
| frobenius_norm tensor(17.1551, device='cuda:6') |
| spectral_norm tensor(1.1882, device='cuda:6') |
| frobenius_norm tensor(41.2336, device='cuda:6') |
| spectral_norm tensor(2.8647, device='cuda:6') |
| frobenius_norm tensor(111.0432, device='cuda:6') |
| spectral_norm tensor(11.8029, device='cuda:6') |
| frobenius_norm tensor(89.6853, device='cuda:6') |
| spectral_norm tensor(3.8356, device='cuda:6') |
| frobenius_norm tensor(89.3886, device='cuda:6') |
| spectral_norm tensor(6.4686, device='cuda:6') |
| stable_rank value of layer 7 |
| Processing layer 8 |
| frobenius_norm tensor(74.6059, device='cuda:6') |
| spectral_norm tensor(8.0167, device='cuda:6') |
| frobenius_norm tensor(56.4533, device='cuda:6') |
| spectral_norm tensor(7.4413, device='cuda:6') |
| frobenius_norm tensor(18.0929, device='cuda:6') |
| spectral_norm tensor(1.1291, device='cuda:6') |
| frobenius_norm tensor(42.1143, device='cuda:6') |
| spectral_norm tensor(3.3892, device='cuda:6') |
| frobenius_norm tensor(111.3565, device='cuda:6') |
| spectral_norm tensor(11.8704, device='cuda:6') |
| frobenius_norm tensor(89.3811, device='cuda:6') |
| spectral_norm tensor(3.7573, device='cuda:6') |
| frobenius_norm tensor(89.2145, device='cuda:6') |
| spectral_norm tensor(6.2894, device='cuda:6') |
| stable_rank value of layer 8 |
| Processing layer 9 |
| frobenius_norm tensor(74.5075, device='cuda:6') |
| spectral_norm tensor(8.3014, device='cuda:6') |
| frobenius_norm tensor(56.1731, device='cuda:6') |
| spectral_norm tensor(7.0673, device='cuda:6') |
| frobenius_norm tensor(20.9612, device='cuda:6') |
| spectral_norm tensor(1.3668, device='cuda:6') |
| frobenius_norm tensor(45.0585, device='cuda:6') |
| spectral_norm tensor(3.3602, device='cuda:6') |
| frobenius_norm tensor(112.6965, device='cuda:6') |
| spectral_norm tensor(12.7243, device='cuda:6') |
| frobenius_norm tensor(90.2687, device='cuda:6') |
| spectral_norm tensor(3.7826, device='cuda:6') |
| frobenius_norm tensor(89.5318, device='cuda:6') |
| spectral_norm tensor(6.5008, device='cuda:6') |
| stable_rank value of layer 9 |
| Processing layer 10 |
| frobenius_norm tensor(75.8125, device='cuda:6') |
| spectral_norm tensor(8.0557, device='cuda:6') |
| frobenius_norm tensor(56.8728, device='cuda:6') |
| spectral_norm tensor(6.8490, device='cuda:6') |
| frobenius_norm tensor(18.1161, device='cuda:6') |
| spectral_norm tensor(1.2821, device='cuda:6') |
| frobenius_norm tensor(41.8618, device='cuda:6') |
| spectral_norm tensor(3.2646, device='cuda:6') |
| frobenius_norm tensor(110.1133, device='cuda:6') |
| spectral_norm tensor(12.3665, device='cuda:6') |
| frobenius_norm tensor(91.5807, device='cuda:6') |
| spectral_norm tensor(3.7899, device='cuda:6') |
| frobenius_norm tensor(90.7305, device='cuda:6') |
| spectral_norm tensor(6.3435, device='cuda:6') |
| stable_rank value of layer 10 |
| Processing layer 11 |
| frobenius_norm tensor(72.1934, device='cuda:6') |
| spectral_norm tensor(8.1887, device='cuda:6') |
| frobenius_norm tensor(56.2865, device='cuda:6') |
| spectral_norm tensor(6.8306, device='cuda:6') |
| frobenius_norm tensor(17.7975, device='cuda:6') |
| spectral_norm tensor(1.2927, device='cuda:6') |
| frobenius_norm tensor(42.6298, device='cuda:6') |
| spectral_norm tensor(3.1474, device='cuda:6') |
| frobenius_norm tensor(108.6994, device='cuda:6') |
| spectral_norm tensor(12.0712, device='cuda:6') |
| frobenius_norm tensor(92.1148, device='cuda:6') |
| spectral_norm tensor(3.8644, device='cuda:6') |
| frobenius_norm tensor(91.3552, device='cuda:6') |
| spectral_norm tensor(7.2518, device='cuda:6') |
| stable_rank value of layer 11 |
| Processing layer 12 |
| frobenius_norm tensor(72.9665, device='cuda:6') |
| spectral_norm tensor(10.0784, device='cuda:6') |
| frobenius_norm tensor(54.8966, device='cuda:6') |
| spectral_norm tensor(7.9653, device='cuda:6') |
| frobenius_norm tensor(20.8604, device='cuda:6') |
| spectral_norm tensor(1.2892, device='cuda:6') |
| frobenius_norm tensor(44.9508, device='cuda:6') |
| spectral_norm tensor(3.1262, device='cuda:6') |
| frobenius_norm tensor(107.3952, device='cuda:6') |
| spectral_norm tensor(12.6191, device='cuda:6') |
| frobenius_norm tensor(93.8434, device='cuda:6') |
| spectral_norm tensor(4.2624, device='cuda:6') |
| frobenius_norm tensor(92.6804, device='cuda:6') |
| spectral_norm tensor(6.7675, device='cuda:6') |
| stable_rank value of layer 12 |
| Processing layer 13 |
| frobenius_norm tensor(73.6418, device='cuda:6') |
| spectral_norm tensor(8.6689, device='cuda:6') |
| frobenius_norm tensor(56.4655, device='cuda:6') |
| spectral_norm tensor(7.3055, device='cuda:6') |
| frobenius_norm tensor(19.4327, device='cuda:6') |
| spectral_norm tensor(1.1633, device='cuda:6') |
| frobenius_norm tensor(44.0239, device='cuda:6') |
| spectral_norm tensor(3.2448, device='cuda:6') |
| frobenius_norm tensor(107.4902, device='cuda:6') |
| spectral_norm tensor(13.0488, device='cuda:6') |
| frobenius_norm tensor(94.2244, device='cuda:6') |
| spectral_norm tensor(4.0377, device='cuda:6') |
| frobenius_norm tensor(92.7343, device='cuda:6') |
| spectral_norm tensor(5.5242, device='cuda:6') |
| stable_rank value of layer 13 |
| Processing layer 14 |
| frobenius_norm tensor(72.2095, device='cuda:6') |
| spectral_norm tensor(8.4329, device='cuda:6') |
| frobenius_norm tensor(55.9640, device='cuda:6') |
| spectral_norm tensor(7.2572, device='cuda:6') |
| frobenius_norm tensor(19.2428, device='cuda:6') |
| spectral_norm tensor(1.1132, device='cuda:6') |
| frobenius_norm tensor(43.8044, device='cuda:6') |
| spectral_norm tensor(2.6452, device='cuda:6') |
| frobenius_norm tensor(109.8463, device='cuda:6') |
| spectral_norm tensor(13.1032, device='cuda:6') |
| frobenius_norm tensor(93.9992, device='cuda:6') |
| spectral_norm tensor(3.9854, device='cuda:6') |
| frobenius_norm tensor(92.6320, device='cuda:6') |
| spectral_norm tensor(5.2650, device='cuda:6') |
| stable_rank value of layer 14 |
| Processing layer 15 |
| frobenius_norm tensor(78.9840, device='cuda:6') |
| spectral_norm tensor(9.9497, device='cuda:6') |
| frobenius_norm tensor(54.9467, device='cuda:6') |
| spectral_norm tensor(7.7205, device='cuda:6') |
| frobenius_norm tensor(20.9664, device='cuda:6') |
| spectral_norm tensor(1.2816, device='cuda:6') |
| frobenius_norm tensor(45.5146, device='cuda:6') |
| spectral_norm tensor(3.4394, device='cuda:6') |
| frobenius_norm tensor(112.3823, device='cuda:6') |
| spectral_norm tensor(14.2334, device='cuda:6') |
| frobenius_norm tensor(93.4551, device='cuda:6') |
| spectral_norm tensor(3.9940, device='cuda:6') |
| frobenius_norm tensor(92.3903, device='cuda:6') |
| spectral_norm tensor(5.1075, device='cuda:6') |
| stable_rank value of layer 15 |
| Processing layer 16 |
| frobenius_norm tensor(76.3742, device='cuda:6') |
| spectral_norm tensor(9.9213, device='cuda:6') |
| frobenius_norm tensor(55.5387, device='cuda:6') |
| spectral_norm tensor(8.2110, device='cuda:6') |
| frobenius_norm tensor(19.7775, device='cuda:6') |
| spectral_norm tensor(1.1561, device='cuda:6') |
| frobenius_norm tensor(44.4344, device='cuda:6') |
| spectral_norm tensor(3.4351, device='cuda:6') |
| frobenius_norm tensor(114.8397, device='cuda:6') |
| spectral_norm tensor(14.0994, device='cuda:6') |
| frobenius_norm tensor(92.3233, device='cuda:6') |
| spectral_norm tensor(4.2256, device='cuda:6') |
| frobenius_norm tensor(91.3091, device='cuda:6') |
| spectral_norm tensor(4.5477, device='cuda:6') |
| stable_rank value of layer 16 |
| Processing layer 17 |
| frobenius_norm tensor(76.7515, device='cuda:6') |
| spectral_norm tensor(10.3508, device='cuda:6') |
| frobenius_norm tensor(55.5269, device='cuda:6') |
| spectral_norm tensor(7.5316, device='cuda:6') |
| frobenius_norm tensor(21.6980, device='cuda:6') |
| spectral_norm tensor(1.2998, device='cuda:6') |
| frobenius_norm tensor(46.0819, device='cuda:6') |
| spectral_norm tensor(3.8492, device='cuda:6') |
| frobenius_norm tensor(115.8902, device='cuda:6') |
| spectral_norm tensor(13.8304, device='cuda:6') |
| frobenius_norm tensor(92.0595, device='cuda:6') |
| spectral_norm tensor(4.2038, device='cuda:6') |
| frobenius_norm tensor(91.2759, device='cuda:6') |
| spectral_norm tensor(4.6559, device='cuda:6') |
| stable_rank value of layer 17 |
| Processing layer 18 |
| frobenius_norm tensor(75.6062, device='cuda:6') |
| spectral_norm tensor(9.9070, device='cuda:6') |
| frobenius_norm tensor(56.8777, device='cuda:6') |
| spectral_norm tensor(7.8718, device='cuda:6') |
| frobenius_norm tensor(19.7982, device='cuda:6') |
| spectral_norm tensor(1.1998, device='cuda:6') |
| frobenius_norm tensor(44.9770, device='cuda:6') |
| spectral_norm tensor(3.6232, device='cuda:6') |
| frobenius_norm tensor(116.1832, device='cuda:6') |
| spectral_norm tensor(13.2450, device='cuda:6') |
| frobenius_norm tensor(91.6447, device='cuda:6') |
| spectral_norm tensor(4.1425, device='cuda:6') |
| frobenius_norm tensor(91.0653, device='cuda:6') |
| spectral_norm tensor(4.6787, device='cuda:6') |
| stable_rank value of layer 18 |
| Processing layer 19 |
| frobenius_norm tensor(75.5331, device='cuda:6') |
| spectral_norm tensor(9.8490, device='cuda:6') |
| frobenius_norm tensor(54.4802, device='cuda:6') |
| spectral_norm tensor(7.2052, device='cuda:6') |
| frobenius_norm tensor(20.8951, device='cuda:6') |
| spectral_norm tensor(1.2835, device='cuda:6') |
| frobenius_norm tensor(45.5737, device='cuda:6') |
| spectral_norm tensor(3.6876, device='cuda:6') |
| frobenius_norm tensor(116.8122, device='cuda:6') |
| spectral_norm tensor(12.4213, device='cuda:6') |
| frobenius_norm tensor(91.3425, device='cuda:6') |
| spectral_norm tensor(3.9625, device='cuda:6') |
| frobenius_norm tensor(90.9342, device='cuda:6') |
| spectral_norm tensor(4.0861, device='cuda:6') |
| stable_rank value of layer 19 |
| Processing layer 20 |
| frobenius_norm tensor(74.6753, device='cuda:6') |
| spectral_norm tensor(9.4898, device='cuda:6') |
| frobenius_norm tensor(53.6751, device='cuda:6') |
| spectral_norm tensor(7.2770, device='cuda:6') |
| frobenius_norm tensor(22.0046, device='cuda:6') |
| spectral_norm tensor(1.3160, device='cuda:6') |
| frobenius_norm tensor(45.1627, device='cuda:6') |
| spectral_norm tensor(4.0906, device='cuda:6') |
| frobenius_norm tensor(116.7308, device='cuda:6') |
| spectral_norm tensor(12.2498, device='cuda:6') |
| frobenius_norm tensor(91.6521, device='cuda:6') |
| spectral_norm tensor(4.0443, device='cuda:6') |
| frobenius_norm tensor(91.2816, device='cuda:6') |
| spectral_norm tensor(4.0475, device='cuda:6') |
| stable_rank value of layer 20 |
| Processing layer 21 |
| frobenius_norm tensor(73.7490, device='cuda:6') |
| spectral_norm tensor(9.8504, device='cuda:6') |
| frobenius_norm tensor(54.0573, device='cuda:6') |
| spectral_norm tensor(7.0908, device='cuda:6') |
| frobenius_norm tensor(22.6461, device='cuda:6') |
| spectral_norm tensor(1.3368, device='cuda:6') |
| frobenius_norm tensor(46.2102, device='cuda:6') |
| spectral_norm tensor(3.1179, device='cuda:6') |
| frobenius_norm tensor(117.5385, device='cuda:6') |
| spectral_norm tensor(11.7937, device='cuda:6') |
| frobenius_norm tensor(91.9506, device='cuda:6') |
| spectral_norm tensor(4.1673, device='cuda:6') |
| frobenius_norm tensor(91.5696, device='cuda:6') |
| spectral_norm tensor(3.9444, device='cuda:6') |
| stable_rank value of layer 21 |
| Processing layer 22 |
| frobenius_norm tensor(71.8484, device='cuda:6') |
| spectral_norm tensor(9.1078, device='cuda:6') |
| frobenius_norm tensor(52.7300, device='cuda:6') |
| spectral_norm tensor(6.7125, device='cuda:6') |
| frobenius_norm tensor(23.8695, device='cuda:6') |
| spectral_norm tensor(1.3795, device='cuda:6') |
| frobenius_norm tensor(47.0670, device='cuda:6') |
| spectral_norm tensor(3.6429, device='cuda:6') |
| frobenius_norm tensor(117.4744, device='cuda:6') |
| spectral_norm tensor(11.7691, device='cuda:6') |
| frobenius_norm tensor(92.2259, device='cuda:6') |
| spectral_norm tensor(4.0605, device='cuda:6') |
| frobenius_norm tensor(91.8678, device='cuda:6') |
| spectral_norm tensor(3.5128, device='cuda:6') |
| stable_rank value of layer 22 |
| Processing layer 23 |
| frobenius_norm tensor(71.9286, device='cuda:6') |
| spectral_norm tensor(9.4720, device='cuda:6') |
| frobenius_norm tensor(52.4818, device='cuda:6') |
| spectral_norm tensor(6.7927, device='cuda:6') |
| frobenius_norm tensor(25.1328, device='cuda:6') |
| spectral_norm tensor(1.6420, device='cuda:6') |
| frobenius_norm tensor(47.9436, device='cuda:6') |
| spectral_norm tensor(3.2829, device='cuda:6') |
| frobenius_norm tensor(117.6067, device='cuda:6') |
| spectral_norm tensor(11.1983, device='cuda:6') |
| frobenius_norm tensor(92.5378, device='cuda:6') |
| spectral_norm tensor(3.8395, device='cuda:6') |
| frobenius_norm tensor(92.2068, device='cuda:6') |
| spectral_norm tensor(3.6470, device='cuda:6') |
| stable_rank value of layer 23 |
| Processing layer 24 |
| frobenius_norm tensor(70.7268, device='cuda:6') |
| spectral_norm tensor(8.9255, device='cuda:6') |
| frobenius_norm tensor(49.9702, device='cuda:6') |
| spectral_norm tensor(6.5083, device='cuda:6') |
| frobenius_norm tensor(27.4158, device='cuda:6') |
| spectral_norm tensor(1.5623, device='cuda:6') |
| frobenius_norm tensor(50.0108, device='cuda:6') |
| spectral_norm tensor(2.8417, device='cuda:6') |
| frobenius_norm tensor(118.0258, device='cuda:6') |
| spectral_norm tensor(10.3830, device='cuda:6') |
| frobenius_norm tensor(92.8487, device='cuda:6') |
| spectral_norm tensor(3.5258, device='cuda:6') |
| frobenius_norm tensor(92.5677, device='cuda:6') |
| spectral_norm tensor(4.2689, device='cuda:6') |
| stable_rank value of layer 24 |
| Processing layer 25 |
| frobenius_norm tensor(69.7075, device='cuda:6') |
| spectral_norm tensor(9.2668, device='cuda:6') |
| frobenius_norm tensor(49.8228, device='cuda:6') |
| spectral_norm tensor(6.4607, device='cuda:6') |
| frobenius_norm tensor(27.5467, device='cuda:6') |
| spectral_norm tensor(1.7434, device='cuda:6') |
| frobenius_norm tensor(50.2014, device='cuda:6') |
| spectral_norm tensor(3.0538, device='cuda:6') |
| frobenius_norm tensor(118.7821, device='cuda:6') |
| spectral_norm tensor(10.1132, device='cuda:6') |
| frobenius_norm tensor(93.4874, device='cuda:6') |
| spectral_norm tensor(3.6431, device='cuda:6') |
| frobenius_norm tensor(93.2013, device='cuda:6') |
| spectral_norm tensor(4.6891, device='cuda:6') |
| stable_rank value of layer 25 |
| Processing layer 26 |
| frobenius_norm tensor(70.0908, device='cuda:6') |
| spectral_norm tensor(9.0741, device='cuda:6') |
| frobenius_norm tensor(51.2286, device='cuda:6') |
| spectral_norm tensor(7.0144, device='cuda:6') |
| frobenius_norm tensor(28.7508, device='cuda:6') |
| spectral_norm tensor(1.6545, device='cuda:6') |
| frobenius_norm tensor(51.0959, device='cuda:6') |
| spectral_norm tensor(3.6542, device='cuda:6') |
| frobenius_norm tensor(119.5556, device='cuda:6') |
| spectral_norm tensor(11.0651, device='cuda:6') |
| frobenius_norm tensor(94.1382, device='cuda:6') |
| spectral_norm tensor(4.1865, device='cuda:6') |
| frobenius_norm tensor(93.8620, device='cuda:6') |
| spectral_norm tensor(4.2123, device='cuda:6') |
| stable_rank value of layer 26 |
| Processing layer 27 |
| frobenius_norm tensor(68.9026, device='cuda:6') |
| spectral_norm tensor(9.8716, device='cuda:6') |
| frobenius_norm tensor(50.5827, device='cuda:6') |
| spectral_norm tensor(7.1128, device='cuda:6') |
| frobenius_norm tensor(30.7386, device='cuda:6') |
| spectral_norm tensor(1.9273, device='cuda:6') |
| frobenius_norm tensor(52.3594, device='cuda:6') |
| spectral_norm tensor(4.7373, device='cuda:6') |
| frobenius_norm tensor(120.4429, device='cuda:6') |
| spectral_norm tensor(12.6819, device='cuda:6') |
| frobenius_norm tensor(95.1424, device='cuda:6') |
| spectral_norm tensor(4.8977, device='cuda:6') |
| frobenius_norm tensor(94.6858, device='cuda:6') |
| spectral_norm tensor(4.1269, device='cuda:6') |
| stable_rank value of layer 27 |
| Processing layer 28 |
| frobenius_norm tensor(68.7252, device='cuda:6') |
| spectral_norm tensor(10.2899, device='cuda:6') |
| frobenius_norm tensor(48.1675, device='cuda:6') |
| spectral_norm tensor(7.2302, device='cuda:6') |
| frobenius_norm tensor(31.7899, device='cuda:6') |
| spectral_norm tensor(1.9608, device='cuda:6') |
| frobenius_norm tensor(53.7083, device='cuda:6') |
| spectral_norm tensor(5.2134, device='cuda:6') |
| frobenius_norm tensor(119.7390, device='cuda:6') |
| spectral_norm tensor(12.8574, device='cuda:6') |
| frobenius_norm tensor(96.8375, device='cuda:6') |
| spectral_norm tensor(5.8801, device='cuda:6') |
| frobenius_norm tensor(96.0067, device='cuda:6') |
| spectral_norm tensor(4.0887, device='cuda:6') |
| stable_rank value of layer 28 |
| Processing layer 29 |
| frobenius_norm tensor(68.5741, device='cuda:6') |
| spectral_norm tensor(9.4460, device='cuda:6') |
| frobenius_norm tensor(50.4763, device='cuda:6') |
| spectral_norm tensor(7.7513, device='cuda:6') |
| frobenius_norm tensor(32.9435, device='cuda:6') |
| spectral_norm tensor(2.0426, device='cuda:6') |
| frobenius_norm tensor(55.2837, device='cuda:6') |
| spectral_norm tensor(4.3637, device='cuda:6') |
| frobenius_norm tensor(119.4804, device='cuda:6') |
| spectral_norm tensor(13.5208, device='cuda:6') |
| frobenius_norm tensor(99.0342, device='cuda:6') |
| spectral_norm tensor(7.1489, device='cuda:6') |
| frobenius_norm tensor(97.3141, device='cuda:6') |
| spectral_norm tensor(3.2229, device='cuda:6') |
| stable_rank value of layer 29 |
| Processing layer 30 |
| frobenius_norm tensor(65.2094, device='cuda:6') |
| spectral_norm tensor(11.1630, device='cuda:6') |
| frobenius_norm tensor(44.0073, device='cuda:6') |
| spectral_norm tensor(7.2343, device='cuda:6') |
| frobenius_norm tensor(39.1436, device='cuda:6') |
| spectral_norm tensor(2.5942, device='cuda:6') |
| frobenius_norm tensor(58.8802, device='cuda:6') |
| spectral_norm tensor(4.7271, device='cuda:6') |
| frobenius_norm tensor(123.5870, device='cuda:6') |
| spectral_norm tensor(15.2486, device='cuda:6') |
| frobenius_norm tensor(101.0766, device='cuda:6') |
| spectral_norm tensor(10.7552, device='cuda:6') |
| frobenius_norm tensor(97.5836, device='cuda:6') |
| spectral_norm tensor(3.7103, device='cuda:6') |
| stable_rank value of layer 30 |
| Processing layer 31 |
| frobenius_norm tensor(71.5148, device='cuda:6') |
| spectral_norm tensor(12.3063, device='cuda:6') |
| frobenius_norm tensor(48.6604, device='cuda:6') |
| spectral_norm tensor(8.2152, device='cuda:6') |
| frobenius_norm tensor(34.5621, device='cuda:6') |
| spectral_norm tensor(2.7818, device='cuda:6') |
| frobenius_norm tensor(57.8552, device='cuda:6') |
| spectral_norm tensor(3.7047, device='cuda:6') |
| frobenius_norm tensor(140.3628, device='cuda:6') |
| spectral_norm tensor(23.9930, device='cuda:6') |
| frobenius_norm tensor(115.5870, device='cuda:6') |
| spectral_norm tensor(19.0393, device='cuda:6') |
| frobenius_norm tensor(99.7985, device='cuda:6') |
| spectral_norm tensor(4.2950, device='cuda:6') |
| stable_rank value of layer 31 |
| [(290.3305358886719, 24), (270.76318359375, 23), (269.9619445800781, 22), (268.34088134765625, 3), (261.071044921875, 25), (249.41912841796875, 21), (247.0523681640625, 26), (242.49368286132812, 29), (240.6619415283203, 4), (235.54725646972656, 19), (234.5662841796875, 14), (232.97064208984375, 20), (229.96060180664062, 2), (226.9339141845703, 5), (215.97666931152344, 16), (213.65187072753906, 1), (213.3599090576172, 15), (212.75192260742188, 13), (211.72381591796875, 18), (209.96926879882812, 27), (209.63941955566406, 6), (209.34996032714844, 17), (201.4922332763672, 8), (199.4788360595703, 9), (198.48497009277344, 10), (195.3269500732422, 28), (192.40182495117188, 7), (189.51531982421875, 11), (187.5979766845703, 12), (185.66880798339844, 30), (154.01223754882812, 31), (92.50985717773438, 0)] |
| metric_name stable_rank: [24, 23, 22, 3, 25, 21, 26, 29, 4, 19, 14, 20, 2, 5, 16, 1, 15, 13, 18, 27, 6, 17, 8, 9, 10, 28, 7, 11, 12, 30, 31, 0] |
| 当前指标:effective_rank |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:09<00:27, 9.27s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:18<00:18, 9.32s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:27<00:09, 9.18s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:29<00:00, 6.41s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:29<00:00, 7.45s/it] |
| Setting `pad_token_id` to `eos_token_id`:128001 for open-end generation. |
| Once upon a time, I loved the sound of a dog’s nails clicking on the pavement. I loved hearing a dog’s tail thumping against the floor when he got excited. Now, I hate those sounds. I have two dogs that are so much in love with each other that when they’re playing, it sounds like the floor is falling apart. When they’re running down the hallway, they’re making the same noise. When they run outside to go potty, the noise is deaf |
| LlamaForCausalLM( |
| (model): LlamaModel( |
| (embed_tokens): Embedding(128256, 4096) |
| (layers): ModuleList( |
| (0-31): 32 x LlamaDecoderLayer( |
| (self_attn): LlamaAttention( |
| (q_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| (k_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (v_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (o_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| ) |
| (mlp): LlamaMLP( |
| (gate_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (up_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (down_proj): Linear(in_features=14336, out_features=4096, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| (post_attention_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| ) |
| ) |
| (norm): LlamaRMSNorm((4096,), eps=1e-05) |
| (rotary_emb): LlamaRotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=4096, out_features=128256, bias=False) |
| ) |
| config: |
| LlamaConfig { |
| "architectures": [ |
| "LlamaForCausalLM" |
| ], |
| "attention_bias": false, |
| "attention_dropout": 0.0, |
| "bos_token_id": 128000, |
| "dtype": "float16", |
| "eos_token_id": 128001, |
| "head_dim": 128, |
| "hidden_act": "silu", |
| "hidden_size": 4096, |
| "initializer_range": 0.02, |
| "intermediate_size": 14336, |
| "max_position_embeddings": 131072, |
| "mlp_bias": false, |
| "model_type": "llama", |
| "num_attention_heads": 32, |
| "num_hidden_layers": 32, |
| "num_key_value_heads": 8, |
| "pretraining_tp": 1, |
| "rms_norm_eps": 1e-05, |
| "rope_scaling": { |
| "factor": 8.0, |
| "high_freq_factor": 4.0, |
| "low_freq_factor": 1.0, |
| "original_max_position_embeddings": 8192, |
| "rope_type": "llama3" |
| }, |
| "rope_theta": 500000.0, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "vocab_size": 128256 |
| } |
|
|
| Processing layer 0 |
| effective_rank value of layer 0 |
| Processing layer 1 |
| effective_rank value of layer 1 |
| Processing layer 2 |
| effective_rank value of layer 2 |
| Processing layer 3 |
| effective_rank value of layer 3 |
| Processing layer 4 |
| effective_rank value of layer 4 |
| Processing layer 5 |
| effective_rank value of layer 5 |
| Processing layer 6 |
| effective_rank value of layer 6 |
| Processing layer 7 |
| effective_rank value of layer 7 |
| Processing layer 8 |
| effective_rank value of layer 8 |
| Processing layer 9 |
| effective_rank value of layer 9 |
| Processing layer 10 |
| effective_rank value of layer 10 |
| Processing layer 11 |
| effective_rank value of layer 11 |
| Processing layer 12 |
| effective_rank value of layer 12 |
| Processing layer 13 |
| effective_rank value of layer 13 |
| Processing layer 14 |
| effective_rank value of layer 14 |
| Processing layer 15 |
| effective_rank value of layer 15 |
| Processing layer 16 |
| effective_rank value of layer 16 |
| Processing layer 17 |
| effective_rank value of layer 17 |
| Processing layer 18 |
| effective_rank value of layer 18 |
| Processing layer 19 |
| effective_rank value of layer 19 |
| Processing layer 20 |
| effective_rank value of layer 20 |
| Processing layer 21 |
| effective_rank value of layer 21 |
| Processing layer 22 |
| effective_rank value of layer 22 |
| Processing layer 23 |
| effective_rank value of layer 23 |
| Processing layer 24 |
| effective_rank value of layer 24 |
| Processing layer 25 |
| effective_rank value of layer 25 |
| Processing layer 26 |
| effective_rank value of layer 26 |
| Processing layer 27 |
| effective_rank value of layer 27 |
| Processing layer 28 |
| effective_rank value of layer 28 |
| Processing layer 29 |
| effective_rank value of layer 29 |
| Processing layer 30 |
| effective_rank value of layer 30 |
| Processing layer 31 |
| effective_rank value of layer 31 |
| [(2794.071533203125, 23), (2792.9208984375, 29), (2782.5615234375, 25), (2781.080810546875, 21), (2780.562744140625, 28), (2779.262939453125, 26), (2777.256591796875, 22), (2777.125732421875, 24), (2768.924072265625, 20), (2767.594970703125, 27), (2763.34814453125, 19), (2752.18798828125, 3), (2743.938720703125, 18), (2734.289306640625, 4), (2732.52294921875, 31), (2724.106201171875, 6), (2722.3828125, 17), (2714.739501953125, 2), (2712.323486328125, 30), (2709.111328125, 16), (2703.84765625, 5), (2702.1064453125, 7), (2690.130126953125, 15), (2689.672607421875, 12), (2673.497314453125, 14), (2669.786376953125, 8), (2663.650146484375, 10), (2662.1806640625, 9), (2661.160888671875, 11), (2650.41259765625, 13), (2606.497802734375, 1), (2424.426513671875, 0)] |
| metric_name effective_rank: [23, 29, 25, 21, 28, 26, 22, 24, 20, 27, 19, 3, 18, 4, 31, 6, 17, 2, 30, 16, 5, 7, 15, 12, 14, 8, 10, 9, 11, 13, 1, 0] |
| 当前指标:head_diversity |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:08<00:26, 8.98s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:17<00:17, 8.96s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:26<00:08, 8.75s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 6.09s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 7.11s/it] |
| Setting `pad_token_id` to `eos_token_id`:128001 for open-end generation. |
| Once upon a time, I was very sick. I was in the hospital for weeks and I was very, very tired. I was tired because my body was in constant pain, and I was tired because I was having to spend all my energy just to breathe. I was tired because I was being fed through a tube, and I was tired because I was so weak. I was tired because I was always alone and I was tired because I was scared. I was tired because I was losing hope. I was tired |
| LlamaForCausalLM( |
| (model): LlamaModel( |
| (embed_tokens): Embedding(128256, 4096) |
| (layers): ModuleList( |
| (0-31): 32 x LlamaDecoderLayer( |
| (self_attn): LlamaAttention( |
| (q_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| (k_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (v_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (o_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| ) |
| (mlp): LlamaMLP( |
| (gate_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (up_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (down_proj): Linear(in_features=14336, out_features=4096, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| (post_attention_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| ) |
| ) |
| (norm): LlamaRMSNorm((4096,), eps=1e-05) |
| (rotary_emb): LlamaRotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=4096, out_features=128256, bias=False) |
| ) |
| config: |
| LlamaConfig { |
| "architectures": [ |
| "LlamaForCausalLM" |
| ], |
| "attention_bias": false, |
| "attention_dropout": 0.0, |
| "bos_token_id": 128000, |
| "dtype": "float16", |
| "eos_token_id": 128001, |
| "head_dim": 128, |
| "hidden_act": "silu", |
| "hidden_size": 4096, |
| "initializer_range": 0.02, |
| "intermediate_size": 14336, |
| "max_position_embeddings": 131072, |
| "mlp_bias": false, |
| "model_type": "llama", |
| "num_attention_heads": 32, |
| "num_hidden_layers": 32, |
| "num_key_value_heads": 8, |
| "pretraining_tp": 1, |
| "rms_norm_eps": 1e-05, |
| "rope_scaling": { |
| "factor": 8.0, |
| "high_freq_factor": 4.0, |
| "low_freq_factor": 1.0, |
| "original_max_position_embeddings": 8192, |
| "rope_type": "llama3" |
| }, |
| "rope_theta": 500000.0, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "vocab_size": 128256 |
| } |
|
|
| Processing layer 0 |
| head_diversity value of layer 0 |
| Processing layer 1 |
| head_diversity value of layer 1 |
| Processing layer 2 |
| head_diversity value of layer 2 |
| Processing layer 3 |
| head_diversity value of layer 3 |
| Processing layer 4 |
| head_diversity value of layer 4 |
| Processing layer 5 |
| head_diversity value of layer 5 |
| Processing layer 6 |
| head_diversity value of layer 6 |
| Processing layer 7 |
| head_diversity value of layer 7 |
| Processing layer 8 |
| head_diversity value of layer 8 |
| Processing layer 9 |
| head_diversity value of layer 9 |
| Processing layer 10 |
| head_diversity value of layer 10 |
| Processing layer 11 |
| head_diversity value of layer 11 |
| Processing layer 12 |
| head_diversity value of layer 12 |
| Processing layer 13 |
| head_diversity value of layer 13 |
| Processing layer 14 |
| head_diversity value of layer 14 |
| Processing layer 15 |
| head_diversity value of layer 15 |
| Processing layer 16 |
| head_diversity value of layer 16 |
| Processing layer 17 |
| head_diversity value of layer 17 |
| Processing layer 18 |
| head_diversity value of layer 18 |
| Processing layer 19 |
| head_diversity value of layer 19 |
| Processing layer 20 |
| head_diversity value of layer 20 |
| Processing layer 21 |
| head_diversity value of layer 21 |
| Processing layer 22 |
| head_diversity value of layer 22 |
| Processing layer 23 |
| head_diversity value of layer 23 |
| Processing layer 24 |
| head_diversity value of layer 24 |
| Processing layer 25 |
| head_diversity value of layer 25 |
| Processing layer 26 |
| head_diversity value of layer 26 |
| Processing layer 27 |
| head_diversity value of layer 27 |
| Processing layer 28 |
| head_diversity value of layer 28 |
| Processing layer 29 |
| head_diversity value of layer 29 |
| Processing layer 30 |
| head_diversity value of layer 30 |
| Processing layer 31 |
| head_diversity value of layer 31 |
| [(0.9939982891082764, 23), (0.9938094615936279, 3), (0.9935836791992188, 24), (0.9935479164123535, 28), (0.993516206741333, 2), (0.9933868646621704, 19), (0.9932371377944946, 1), (0.9932280778884888, 21), (0.9932112693786621, 31), (0.993040919303894, 20), (0.9929364919662476, 4), (0.9929299354553223, 26), (0.9929077625274658, 6), (0.9928981065750122, 27), (0.9927928447723389, 22), (0.9927812218666077, 25), (0.9926327466964722, 29), (0.9925468564033508, 17), (0.992446780204773, 18), (0.9918708801269531, 7), (0.9917168617248535, 30), (0.9916648864746094, 16), (0.9915013313293457, 5), (0.9912642240524292, 15), (0.990802526473999, 8), (0.9907873868942261, 9), (0.9907500743865967, 12), (0.9904842376708984, 11), (0.9903575778007507, 10), (0.9888362288475037, 14), (0.988161563873291, 13), (0.9860825538635254, 0)] |
| metric_name head_diversity: [23, 3, 24, 28, 2, 19, 1, 21, 31, 20, 4, 26, 6, 27, 22, 25, 29, 17, 18, 7, 30, 16, 5, 15, 8, 9, 12, 11, 10, 14, 13, 0] |
| 当前指标:coherence |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:08<00:26, 8.69s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:17<00:17, 8.78s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:26<00:08, 8.66s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 6.03s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 7.01s/it] |
| Setting `pad_token_id` to `eos_token_id`:128001 for open-end generation. |
| Once upon a time, the words “sustainable agriculture” conjured up images of hippies, granola, and Birkenstocks. Today, the concept is gaining acceptance across the political spectrum as an effective way to address some of the most pressing problems facing America. A growing number of businesses, cities, and communities are embracing sustainable agriculture as a way to reduce our carbon footprint, improve our health, and support local economies. Sustainable agriculture is a holistic approach to food production that emphasizes ecological balance and social equity. It |
| LlamaForCausalLM( |
| (model): LlamaModel( |
| (embed_tokens): Embedding(128256, 4096) |
| (layers): ModuleList( |
| (0-31): 32 x LlamaDecoderLayer( |
| (self_attn): LlamaAttention( |
| (q_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| (k_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (v_proj): Linear(in_features=4096, out_features=1024, bias=False) |
| (o_proj): Linear(in_features=4096, out_features=4096, bias=False) |
| ) |
| (mlp): LlamaMLP( |
| (gate_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (up_proj): Linear(in_features=4096, out_features=14336, bias=False) |
| (down_proj): Linear(in_features=14336, out_features=4096, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| (post_attention_layernorm): LlamaRMSNorm((4096,), eps=1e-05) |
| ) |
| ) |
| (norm): LlamaRMSNorm((4096,), eps=1e-05) |
| (rotary_emb): LlamaRotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=4096, out_features=128256, bias=False) |
| ) |
| config: |
| LlamaConfig { |
| "architectures": [ |
| "LlamaForCausalLM" |
| ], |
| "attention_bias": false, |
| "attention_dropout": 0.0, |
| "bos_token_id": 128000, |
| "dtype": "float16", |
| "eos_token_id": 128001, |
| "head_dim": 128, |
| "hidden_act": "silu", |
| "hidden_size": 4096, |
| "initializer_range": 0.02, |
| "intermediate_size": 14336, |
| "max_position_embeddings": 131072, |
| "mlp_bias": false, |
| "model_type": "llama", |
| "num_attention_heads": 32, |
| "num_hidden_layers": 32, |
| "num_key_value_heads": 8, |
| "pretraining_tp": 1, |
| "rms_norm_eps": 1e-05, |
| "rope_scaling": { |
| "factor": 8.0, |
| "high_freq_factor": 4.0, |
| "low_freq_factor": 1.0, |
| "original_max_position_embeddings": 8192, |
| "rope_type": "llama3" |
| }, |
| "rope_theta": 500000.0, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "vocab_size": 128256 |
| } |
|
|
| Processing layer 0 |
| coherence value of layer 0 |
| Processing layer 1 |
| coherence value of layer 1 |
| Processing layer 2 |
| coherence value of layer 2 |
| Processing layer 3 |
| coherence value of layer 3 |
| Processing layer 4 |
| coherence value of layer 4 |
| Processing layer 5 |
| coherence value of layer 5 |
| Processing layer 6 |
| coherence value of layer 6 |
| Processing layer 7 |
| coherence value of layer 7 |
| Processing layer 8 |
| coherence value of layer 8 |
| Processing layer 9 |
| coherence value of layer 9 |
| Processing layer 10 |
| coherence value of layer 10 |
| Processing layer 11 |
| coherence value of layer 11 |
| Processing layer 12 |
| coherence value of layer 12 |
| Processing layer 13 |
| coherence value of layer 13 |
| Processing layer 14 |
| coherence value of layer 14 |
| Processing layer 15 |
| coherence value of layer 15 |
| Processing layer 16 |
| coherence value of layer 16 |
| Processing layer 17 |
| coherence value of layer 17 |
| Processing layer 18 |
| coherence value of layer 18 |
| Processing layer 19 |
| coherence value of layer 19 |
| Processing layer 20 |
| coherence value of layer 20 |
| Processing layer 21 |
| coherence value of layer 21 |
| Processing layer 22 |
| coherence value of layer 22 |
| Processing layer 23 |
| coherence value of layer 23 |
| Processing layer 24 |
| coherence value of layer 24 |
| Processing layer 25 |
| coherence value of layer 25 |
| Processing layer 26 |
| coherence value of layer 26 |
| Processing layer 27 |
| coherence value of layer 27 |
| Processing layer 28 |
| coherence value of layer 28 |
| Processing layer 29 |
| coherence value of layer 29 |
| Processing layer 30 |
| coherence value of layer 30 |
| Processing layer 31 |
| coherence value of layer 31 |
| [(0.04617950692772865, 0), (0.025356026366353035, 1), (0.019824203103780746, 31), (0.019280431792140007, 13), (0.01895361766219139, 12), (0.018941296264529228, 11), (0.01878860592842102, 7), (0.018680352717638016, 30), (0.01830293983221054, 9), (0.01820538565516472, 8), (0.01808197796344757, 10), (0.0178972240537405, 4), (0.017885111272335052, 14), (0.01778978481888771, 5), (0.017596283927559853, 2), (0.017451299354434013, 6), (0.017325662076473236, 3), (0.017044615000486374, 18), (0.016964342445135117, 15), (0.01687796786427498, 16), (0.016494933515787125, 17), (0.0163591206073761, 27), (0.016093522310256958, 19), (0.016048356890678406, 26), (0.01604018732905388, 28), (0.015906821936368942, 20), (0.015847351402044296, 24), (0.01560671441257, 25), (0.015601491555571556, 22), (0.015557424165308475, 21), (0.015549903735518456, 29), (0.015353316441178322, 23)] |
| metric_name coherence: [0, 1, 31, 13, 12, 11, 7, 30, 9, 8, 10, 4, 14, 5, 2, 6, 3, 18, 15, 16, 17, 27, 19, 26, 28, 20, 24, 25, 22, 21, 29, 23] |
| 当前指标:ZD |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:07<00:21, 7.14s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:14<00:14, 7.06s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:20<00:06, 6.87s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:27<00:00, 6.72s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:27<00:00, 6.82s/it] |
| Setting `pad_token_id` to `eos_token_id`:151643 for open-end generation. |
| Once upon a time, I took an advanced placement class in high school, a class that promised to teach us the “basics” of calculus. It was an interesting class in its own right, but I was struck with the fact that we were learning calculus as a high schooler. The teacher was a nice guy, but I remember thinking, “If I can learn this in high school, why can’t high school students learn more math in high school?”\nI’m not saying that we should be teaching |
| Qwen2ForCausalLM( |
| (model): Qwen2Model( |
| (embed_tokens): Embedding(152064, 3584) |
| (layers): ModuleList( |
| (0-27): 28 x Qwen2DecoderLayer( |
| (self_attn): Qwen2Attention( |
| (q_proj): Linear(in_features=3584, out_features=3584, bias=True) |
| (k_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (v_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (o_proj): Linear(in_features=3584, out_features=3584, bias=False) |
| ) |
| (mlp): Qwen2MLP( |
| (gate_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (up_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (down_proj): Linear(in_features=18944, out_features=3584, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (post_attention_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| ) |
| ) |
| (norm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (rotary_emb): Qwen2RotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=3584, out_features=152064, bias=False) |
| ) |
| config: |
| Qwen2Config { |
| "architectures": [ |
| "Qwen2ForCausalLM" |
| ], |
| "attention_dropout": 0.0, |
| "bos_token_id": 151643, |
| "dtype": "float16", |
| "eos_token_id": 151643, |
| "hidden_act": "silu", |
| "hidden_size": 3584, |
| "initializer_range": 0.02, |
| "intermediate_size": 18944, |
| "layer_types": [ |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention" |
| ], |
| "max_position_embeddings": 131072, |
| "max_window_layers": 28, |
| "model_type": "qwen2", |
| "num_attention_heads": 28, |
| "num_hidden_layers": 28, |
| "num_key_value_heads": 4, |
| "rms_norm_eps": 1e-06, |
| "rope_scaling": null, |
| "rope_theta": 1000000.0, |
| "sliding_window": null, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "use_mrope": false, |
| "use_sliding_window": false, |
| "vocab_size": 152064 |
| } |
|
|
| Processing layer 0 |
| ZD value of layer 0 |
| Processing layer 1 |
| ZD value of layer 1 |
| Processing layer 2 |
| ZD value of layer 2 |
| Processing layer 3 |
| ZD value of layer 3 |
| Processing layer 4 |
| ZD value of layer 4 |
| Processing layer 5 |
| ZD value of layer 5 |
| Processing layer 6 |
| ZD value of layer 6 |
| Processing layer 7 |
| ZD value of layer 7 |
| Processing layer 8 |
| ZD value of layer 8 |
| Processing layer 9 |
| ZD value of layer 9 |
| Processing layer 10 |
| ZD value of layer 10 |
| Processing layer 11 |
| ZD value of layer 11 |
| Processing layer 12 |
| ZD value of layer 12 |
| Processing layer 13 |
| ZD value of layer 13 |
| Processing layer 14 |
| ZD value of layer 14 |
| Processing layer 15 |
| ZD value of layer 15 |
| Processing layer 16 |
| ZD value of layer 16 |
| Processing layer 17 |
| ZD value of layer 17 |
| Processing layer 18 |
| ZD value of layer 18 |
| Processing layer 19 |
| ZD value of layer 19 |
| Processing layer 20 |
| ZD value of layer 20 |
| Processing layer 21 |
| ZD value of layer 21 |
| Processing layer 22 |
| ZD value of layer 22 |
| Processing layer 23 |
| ZD value of layer 23 |
| Processing layer 24 |
| ZD value of layer 24 |
| Processing layer 25 |
| ZD value of layer 25 |
| Processing layer 26 |
| ZD value of layer 26 |
| Processing layer 27 |
| ZD value of layer 27 |
| [(0.14779578149318695, 24), (0.14769357442855835, 6), (0.1476735919713974, 10), (0.14743219316005707, 3), (0.14733701944351196, 7), (0.14689013361930847, 5), (0.14676602184772491, 11), (0.14655494689941406, 4), (0.14630182087421417, 8), (0.1456553190946579, 20), (0.14544296264648438, 25), (0.14524872601032257, 9), (0.14517414569854736, 13), (0.14493051171302795, 12), (0.14477485418319702, 16), (0.14454184472560883, 23), (0.14420966804027557, 19), (0.1438734531402588, 18), (0.1436464488506317, 17), (0.14361177384853363, 2), (0.14331096410751343, 15), (0.14328131079673767, 26), (0.14289206266403198, 21), (0.1421867161989212, 22), (0.14147840440273285, 27), (0.14123240113258362, 0), (0.14055506885051727, 14), (0.1311480700969696, 1)] |
| metric_name ZD: [24, 6, 10, 3, 7, 5, 11, 4, 8, 20, 25, 9, 13, 12, 16, 23, 19, 18, 17, 2, 15, 26, 21, 22, 27, 0, 14, 1] |
| 当前指标:alpha |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:07<00:21, 7.07s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:13<00:13, 6.98s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:20<00:06, 6.87s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:27<00:00, 6.74s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:27<00:00, 6.81s/it] |
| Setting `pad_token_id` to `eos_token_id`:151643 for open-end generation. |
| Once upon a time, a king had a very wise minister who was loved by everyone. One day, the king called for the minister and said, "I want to test your wisdom. I have three bags of gold coins, but two of them contain fake coins while the other one contains real coins. You need to identify the bag with real coins within just one weighing on my scales." The minister thought for a moment and then replied, "Yes, Your Majesty. I can do that." How did the minister manage |
| Qwen2ForCausalLM( |
| (model): Qwen2Model( |
| (embed_tokens): Embedding(152064, 3584) |
| (layers): ModuleList( |
| (0-27): 28 x Qwen2DecoderLayer( |
| (self_attn): Qwen2Attention( |
| (q_proj): Linear(in_features=3584, out_features=3584, bias=True) |
| (k_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (v_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (o_proj): Linear(in_features=3584, out_features=3584, bias=False) |
| ) |
| (mlp): Qwen2MLP( |
| (gate_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (up_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (down_proj): Linear(in_features=18944, out_features=3584, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (post_attention_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| ) |
| ) |
| (norm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (rotary_emb): Qwen2RotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=3584, out_features=152064, bias=False) |
| ) |
| config: |
| Qwen2Config { |
| "architectures": [ |
| "Qwen2ForCausalLM" |
| ], |
| "attention_dropout": 0.0, |
| "bos_token_id": 151643, |
| "dtype": "float16", |
| "eos_token_id": 151643, |
| "hidden_act": "silu", |
| "hidden_size": 3584, |
| "initializer_range": 0.02, |
| "intermediate_size": 18944, |
| "layer_types": [ |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention" |
| ], |
| "max_position_embeddings": 131072, |
| "max_window_layers": 28, |
| "model_type": "qwen2", |
| "num_attention_heads": 28, |
| "num_hidden_layers": 28, |
| "num_key_value_heads": 4, |
| "rms_norm_eps": 1e-06, |
| "rope_scaling": null, |
| "rope_theta": 1000000.0, |
| "sliding_window": null, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "use_mrope": false, |
| "use_sliding_window": false, |
| "vocab_size": 152064 |
| } |
|
|
| Processing layer 0 |
| alpha value of layer 0 |
| Processing layer 1 |
| alpha value of layer 1 |
| Processing layer 2 |
| alpha value of layer 2 |
| Processing layer 3 |
| alpha value of layer 3 |
| Processing layer 4 |
| alpha value of layer 4 |
| Processing layer 5 |
| alpha value of layer 5 |
| Processing layer 6 |
| alpha value of layer 6 |
| Processing layer 7 |
| alpha value of layer 7 |
| Processing layer 8 |
| alpha value of layer 8 |
| Processing layer 9 |
| alpha value of layer 9 |
| Processing layer 10 |
| alpha value of layer 10 |
| Processing layer 11 |
| alpha value of layer 11 |
| Processing layer 12 |
| alpha value of layer 12 |
| Processing layer 13 |
| alpha value of layer 13 |
| Processing layer 14 |
| alpha value of layer 14 |
| Processing layer 15 |
| alpha value of layer 15 |
| Processing layer 16 |
| alpha value of layer 16 |
| Processing layer 17 |
| alpha value of layer 17 |
| Processing layer 18 |
| alpha value of layer 18 |
| Processing layer 19 |
| alpha value of layer 19 |
| Processing layer 20 |
| alpha value of layer 20 |
| Processing layer 21 |
| alpha value of layer 21 |
| Processing layer 22 |
| alpha value of layer 22 |
| Processing layer 23 |
| alpha value of layer 23 |
| Processing layer 24 |
| alpha value of layer 24 |
| Processing layer 25 |
| alpha value of layer 25 |
| Processing layer 26 |
| alpha value of layer 26 |
| Processing layer 27 |
| alpha value of layer 27 |
| [(6.430380821228027, 6), (6.222294330596924, 23), (6.176848411560059, 13), (5.387862682342529, 21), (5.161396503448486, 8), (5.114707946777344, 25), (4.984698295593262, 18), (4.962588310241699, 9), (4.807549953460693, 24), (4.767049312591553, 22), (4.559838771820068, 7), (4.515918731689453, 0), (4.414379596710205, 10), (4.310111999511719, 11), (4.209693431854248, 16), (4.1970930099487305, 4), (4.061600685119629, 12), (4.0341081619262695, 26), (4.019566535949707, 5), (3.9200167655944824, 14), (3.8895316123962402, 1), (3.783604383468628, 3), (3.776597261428833, 17), (3.751697301864624, 19), (3.6569418907165527, 15), (3.6451375484466553, 20), (3.5805065631866455, 2), (3.413137912750244, 27)] |
| metric_name alpha: [6, 23, 13, 21, 8, 25, 18, 9, 24, 22, 7, 0, 10, 11, 16, 4, 12, 26, 5, 14, 1, 3, 17, 19, 15, 20, 2, 27] |
| 当前指标:alpha_hat |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:07<00:21, 7.06s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:14<00:14, 7.13s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:21<00:07, 7.05s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:27<00:00, 6.79s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:27<00:00, 6.90s/it] |
| Setting `pad_token_id` to `eos_token_id`:151643 for open-end generation. |
| Once upon a time, there was a young man named Alex who had a great passion for traveling and exploring new places. However, his parents were strict and never allowed him to leave their house. One day, Alex's curiosity got the better of him and he decided to sneak out of the house and explore the neighborhood. After wandering around for a few hours, he stumbled upon an old, abandoned building that caught his attention. He cautiously entered the building and found himself in a dimly lit room with a locked door. |
| Qwen2ForCausalLM( |
| (model): Qwen2Model( |
| (embed_tokens): Embedding(152064, 3584) |
| (layers): ModuleList( |
| (0-27): 28 x Qwen2DecoderLayer( |
| (self_attn): Qwen2Attention( |
| (q_proj): Linear(in_features=3584, out_features=3584, bias=True) |
| (k_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (v_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (o_proj): Linear(in_features=3584, out_features=3584, bias=False) |
| ) |
| (mlp): Qwen2MLP( |
| (gate_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (up_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (down_proj): Linear(in_features=18944, out_features=3584, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (post_attention_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| ) |
| ) |
| (norm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (rotary_emb): Qwen2RotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=3584, out_features=152064, bias=False) |
| ) |
| config: |
| Qwen2Config { |
| "architectures": [ |
| "Qwen2ForCausalLM" |
| ], |
| "attention_dropout": 0.0, |
| "bos_token_id": 151643, |
| "dtype": "float16", |
| "eos_token_id": 151643, |
| "hidden_act": "silu", |
| "hidden_size": 3584, |
| "initializer_range": 0.02, |
| "intermediate_size": 18944, |
| "layer_types": [ |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention" |
| ], |
| "max_position_embeddings": 131072, |
| "max_window_layers": 28, |
| "model_type": "qwen2", |
| "num_attention_heads": 28, |
| "num_hidden_layers": 28, |
| "num_key_value_heads": 4, |
| "rms_norm_eps": 1e-06, |
| "rope_scaling": null, |
| "rope_theta": 1000000.0, |
| "sliding_window": null, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "use_mrope": false, |
| "use_sliding_window": false, |
| "vocab_size": 152064 |
| } |
| |
| Processing layer 0--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 0 ---17.878414154052734 |
| Processing layer 1--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 1 ---10.724651336669922 |
| Processing layer 2--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 2 ---11.09423542022705 |
| Processing layer 3--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 3 ---13.750286102294922 |
| Processing layer 4--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 4 ---12.229127883911133 |
| Processing layer 5--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 5 ---13.23713493347168 |
| Processing layer 6--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 6 ---15.239927291870117 |
| Processing layer 7--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 7 ---13.34253215789795 |
| Processing layer 8--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 8 ---12.03944206237793 |
| Processing layer 9--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 9 ---14.739602088928223 |
| Processing layer 10--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 10 ---12.660069465637207 |
| Processing layer 11--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 11 ---11.740548133850098 |
| Processing layer 12--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 12 ---12.344042778015137 |
| Processing layer 13--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 13 ---12.779866218566895 |
| Processing layer 14--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 14 ---10.932424545288086 |
| Processing layer 15--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 15 ---10.816814422607422 |
| Processing layer 16--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 16 ---12.136594772338867 |
| Processing layer 17--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 17 ---10.876593589782715 |
| Processing layer 18--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 18 ---12.142037391662598 |
| Processing layer 19--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 19 ---11.50020694732666 |
| Processing layer 20--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 20 ---11.685124397277832 |
| Processing layer 21--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 21 ---14.077293395996094 |
| Processing layer 22--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 22 ---14.83903980255127 |
| Processing layer 23--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 23 ---17.10747718811035 |
| Processing layer 24--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 24 ---14.79826831817627 |
| Processing layer 25--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 25 ---16.314620971679688 |
| Processing layer 26--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 26 ---14.33527946472168 |
| Processing layer 27--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| alpha_hat value of layer 27 ---15.28555679321289 |
| [(17.878414154052734, 0), (17.10747718811035, 23), (16.314620971679688, 25), (15.28555679321289, 27), (15.239927291870117, 6), (14.83903980255127, 22), (14.79826831817627, 24), (14.739602088928223, 9), (14.33527946472168, 26), (14.077293395996094, 21), (13.750286102294922, 3), (13.34253215789795, 7), (13.23713493347168, 5), (12.779866218566895, 13), (12.660069465637207, 10), (12.344042778015137, 12), (12.229127883911133, 4), (12.142037391662598, 18), (12.136594772338867, 16), (12.03944206237793, 8), (11.740548133850098, 11), (11.685124397277832, 20), (11.50020694732666, 19), (11.09423542022705, 2), (10.932424545288086, 14), (10.876593589782715, 17), (10.816814422607422, 15), (10.724651336669922, 1)] |
| metric_name alpha_hat: [0, 23, 25, 27, 6, 22, 24, 9, 26, 21, 3, 7, 5, 13, 10, 12, 4, 18, 16, 8, 11, 20, 19, 2, 14, 17, 15, 1] |
| 当前指标:stable_rank |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:07<00:22, 7.37s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:14<00:14, 7.21s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:21<00:07, 7.26s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 7.04s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 7.12s/it] |
| Setting `pad_token_id` to `eos_token_id`:151643 for open-end generation. |
| Once upon a time a poor man had a large field, but he was only able to work on it for three hours a day. He was very sad because of his poor life. One day a rich man said to him, "I want you to work for me for one year. I will pay you with rice. One heap of rice will be given to you for the first day, and two heaps for the second day, and so on." "How much will I get on the last day?" the |
| Qwen2ForCausalLM( |
| (model): Qwen2Model( |
| (embed_tokens): Embedding(152064, 3584) |
| (layers): ModuleList( |
| (0-27): 28 x Qwen2DecoderLayer( |
| (self_attn): Qwen2Attention( |
| (q_proj): Linear(in_features=3584, out_features=3584, bias=True) |
| (k_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (v_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (o_proj): Linear(in_features=3584, out_features=3584, bias=False) |
| ) |
| (mlp): Qwen2MLP( |
| (gate_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (up_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (down_proj): Linear(in_features=18944, out_features=3584, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (post_attention_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| ) |
| ) |
| (norm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (rotary_emb): Qwen2RotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=3584, out_features=152064, bias=False) |
| ) |
| config: |
| Qwen2Config { |
| "architectures": [ |
| "Qwen2ForCausalLM" |
| ], |
| "attention_dropout": 0.0, |
| "bos_token_id": 151643, |
| "dtype": "float16", |
| "eos_token_id": 151643, |
| "hidden_act": "silu", |
| "hidden_size": 3584, |
| "initializer_range": 0.02, |
| "intermediate_size": 18944, |
| "layer_types": [ |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention" |
| ], |
| "max_position_embeddings": 131072, |
| "max_window_layers": 28, |
| "model_type": "qwen2", |
| "num_attention_heads": 28, |
| "num_hidden_layers": 28, |
| "num_key_value_heads": 4, |
| "rms_norm_eps": 1e-06, |
| "rope_scaling": null, |
| "rope_theta": 1000000.0, |
| "sliding_window": null, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "use_mrope": false, |
| "use_sliding_window": false, |
| "vocab_size": 152064 |
| } |
| |
| Processing layer 0--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(74.6100, device='cuda:6') |
| spectral_norm tensor(22.8962, device='cuda:6') |
| frobenius_norm tensor(37.2340, device='cuda:6') |
| spectral_norm tensor(4.1839, device='cuda:6') |
| frobenius_norm tensor(12.2050, device='cuda:6') |
| spectral_norm tensor(1.2453, device='cuda:6') |
| frobenius_norm tensor(44.6291, device='cuda:6') |
| spectral_norm tensor(6.0001, device='cuda:6') |
| frobenius_norm tensor(126.6843, device='cuda:6') |
| spectral_norm tensor(37.7735, device='cuda:6') |
| frobenius_norm tensor(108.8807, device='cuda:6') |
| spectral_norm tensor(7.8392, device='cuda:6') |
| frobenius_norm tensor(115.3682, device='cuda:6') |
| spectral_norm tensor(6.5322, device='cuda:6') |
| stable_rank value of layer 0 ---108.18235778808594 |
| Processing layer 1--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(56.7970, device='cuda:6') |
| spectral_norm tensor(10.4783, device='cuda:6') |
| frobenius_norm tensor(29.9237, device='cuda:6') |
| spectral_norm tensor(4.7115, device='cuda:6') |
| frobenius_norm tensor(18.3318, device='cuda:6') |
| spectral_norm tensor(1.4635, device='cuda:6') |
| frobenius_norm tensor(51.8109, device='cuda:6') |
| spectral_norm tensor(3.7208, device='cuda:6') |
| frobenius_norm tensor(123.2531, device='cuda:6') |
| spectral_norm tensor(20.7839, device='cuda:6') |
| frobenius_norm tensor(101.4097, device='cuda:6') |
| spectral_norm tensor(8.8381, device='cuda:6') |
| frobenius_norm tensor(101.5777, device='cuda:6') |
| spectral_norm tensor(13.5321, device='cuda:6') |
| stable_rank value of layer 1 ---91.95645141601562 |
| Processing layer 2--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(65.4367, device='cuda:6') |
| spectral_norm tensor(5.5540, device='cuda:6') |
| frobenius_norm tensor(31.8038, device='cuda:6') |
| spectral_norm tensor(3.0331, device='cuda:6') |
| frobenius_norm tensor(14.7804, device='cuda:6') |
| spectral_norm tensor(1.2128, device='cuda:6') |
| frobenius_norm tensor(50.6930, device='cuda:6') |
| spectral_norm tensor(3.9768, device='cuda:6') |
| frobenius_norm tensor(138.8778, device='cuda:6') |
| spectral_norm tensor(23.4270, device='cuda:6') |
| frobenius_norm tensor(112.7775, device='cuda:6') |
| spectral_norm tensor(6.5452, device='cuda:6') |
| frobenius_norm tensor(115.6860, device='cuda:6') |
| spectral_norm tensor(7.9797, device='cuda:6') |
| stable_rank value of layer 2 ---157.4268035888672 |
| Processing layer 3--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(67.9445, device='cuda:6') |
| spectral_norm tensor(7.3888, device='cuda:6') |
| frobenius_norm tensor(32.4597, device='cuda:6') |
| spectral_norm tensor(3.6898, device='cuda:6') |
| frobenius_norm tensor(17.2702, device='cuda:6') |
| spectral_norm tensor(1.2119, device='cuda:6') |
| frobenius_norm tensor(53.1248, device='cuda:6') |
| spectral_norm tensor(3.7626, device='cuda:6') |
| frobenius_norm tensor(153.6270, device='cuda:6') |
| spectral_norm tensor(22.2753, device='cuda:6') |
| frobenius_norm tensor(132.3810, device='cuda:6') |
| spectral_norm tensor(7.2662, device='cuda:6') |
| frobenius_norm tensor(131.2425, device='cuda:6') |
| spectral_norm tensor(17.0355, device='cuda:6') |
| stable_rank value of layer 3 ---143.3182830810547 |
| Processing layer 4--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(64.7004, device='cuda:6') |
| spectral_norm tensor(6.6270, device='cuda:6') |
| frobenius_norm tensor(29.1071, device='cuda:6') |
| spectral_norm tensor(3.3637, device='cuda:6') |
| frobenius_norm tensor(20.8693, device='cuda:6') |
| spectral_norm tensor(1.5107, device='cuda:6') |
| frobenius_norm tensor(53.3331, device='cuda:6') |
| spectral_norm tensor(3.7430, device='cuda:6') |
| frobenius_norm tensor(156.7403, device='cuda:6') |
| spectral_norm tensor(22.8238, device='cuda:6') |
| frobenius_norm tensor(129.6389, device='cuda:6') |
| spectral_norm tensor(5.4833, device='cuda:6') |
| frobenius_norm tensor(128.7504, device='cuda:6') |
| spectral_norm tensor(8.4121, device='cuda:6') |
| stable_rank value of layer 4 ---200.63299560546875 |
| Processing layer 5--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(63.2364, device='cuda:6') |
| spectral_norm tensor(5.9598, device='cuda:6') |
| frobenius_norm tensor(26.6552, device='cuda:6') |
| spectral_norm tensor(2.6609, device='cuda:6') |
| frobenius_norm tensor(19.8491, device='cuda:6') |
| spectral_norm tensor(1.5480, device='cuda:6') |
| frobenius_norm tensor(53.4192, device='cuda:6') |
| spectral_norm tensor(4.0186, device='cuda:6') |
| frobenius_norm tensor(149.4161, device='cuda:6') |
| spectral_norm tensor(21.1506, device='cuda:6') |
| frobenius_norm tensor(133.9768, device='cuda:6') |
| spectral_norm tensor(6.2399, device='cuda:6') |
| frobenius_norm tensor(132.0446, device='cuda:6') |
| spectral_norm tensor(8.4305, device='cuda:6') |
| stable_rank value of layer 5 ---187.18458557128906 |
| Processing layer 6--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(64.2853, device='cuda:6') |
| spectral_norm tensor(5.5097, device='cuda:6') |
| frobenius_norm tensor(28.5616, device='cuda:6') |
| spectral_norm tensor(2.9881, device='cuda:6') |
| frobenius_norm tensor(20.6937, device='cuda:6') |
| spectral_norm tensor(1.4517, device='cuda:6') |
| frobenius_norm tensor(54.9730, device='cuda:6') |
| spectral_norm tensor(5.0745, device='cuda:6') |
| frobenius_norm tensor(151.2180, device='cuda:6') |
| spectral_norm tensor(14.3238, device='cuda:6') |
| frobenius_norm tensor(129.6978, device='cuda:6') |
| spectral_norm tensor(4.1238, device='cuda:6') |
| frobenius_norm tensor(127.6298, device='cuda:6') |
| spectral_norm tensor(7.3585, device='cuda:6') |
| stable_rank value of layer 6 ---278.4970703125 |
| Processing layer 7--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(60.9057, device='cuda:6') |
| spectral_norm tensor(4.5698, device='cuda:6') |
| frobenius_norm tensor(23.8630, device='cuda:6') |
| spectral_norm tensor(2.5382, device='cuda:6') |
| frobenius_norm tensor(24.6983, device='cuda:6') |
| spectral_norm tensor(1.7165, device='cuda:6') |
| frobenius_norm tensor(60.3206, device='cuda:6') |
| spectral_norm tensor(3.5918, device='cuda:6') |
| frobenius_norm tensor(141.6723, device='cuda:6') |
| spectral_norm tensor(13.0737, device='cuda:6') |
| frobenius_norm tensor(133.2031, device='cuda:6') |
| spectral_norm tensor(5.1650, device='cuda:6') |
| frobenius_norm tensor(132.6786, device='cuda:6') |
| spectral_norm tensor(7.6574, device='cuda:6') |
| stable_rank value of layer 7 ---262.54864501953125 |
| Processing layer 8--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(63.3943, device='cuda:6') |
| spectral_norm tensor(3.9526, device='cuda:6') |
| frobenius_norm tensor(27.2992, device='cuda:6') |
| spectral_norm tensor(2.5236, device='cuda:6') |
| frobenius_norm tensor(20.6254, device='cuda:6') |
| spectral_norm tensor(1.4007, device='cuda:6') |
| frobenius_norm tensor(55.2465, device='cuda:6') |
| spectral_norm tensor(3.4821, device='cuda:6') |
| frobenius_norm tensor(139.1350, device='cuda:6') |
| spectral_norm tensor(12.2201, device='cuda:6') |
| frobenius_norm tensor(135.6841, device='cuda:6') |
| spectral_norm tensor(4.7790, device='cuda:6') |
| frobenius_norm tensor(134.1169, device='cuda:6') |
| spectral_norm tensor(8.0838, device='cuda:6') |
| stable_rank value of layer 8 ---293.3992919921875 |
| Processing layer 9--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(58.6742, device='cuda:6') |
| spectral_norm tensor(4.6933, device='cuda:6') |
| frobenius_norm tensor(23.2825, device='cuda:6') |
| spectral_norm tensor(2.6721, device='cuda:6') |
| frobenius_norm tensor(24.7824, device='cuda:6') |
| spectral_norm tensor(1.5406, device='cuda:6') |
| frobenius_norm tensor(60.4736, device='cuda:6') |
| spectral_norm tensor(6.3921, device='cuda:6') |
| frobenius_norm tensor(152.4723, device='cuda:6') |
| spectral_norm tensor(28.9447, device='cuda:6') |
| frobenius_norm tensor(124.9508, device='cuda:6') |
| spectral_norm tensor(5.0401, device='cuda:6') |
| frobenius_norm tensor(123.2868, device='cuda:6') |
| spectral_norm tensor(7.9764, device='cuda:6') |
| stable_rank value of layer 9 ---208.82086181640625 |
| Processing layer 10--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(63.5265, device='cuda:6') |
| spectral_norm tensor(4.3337, device='cuda:6') |
| frobenius_norm tensor(27.1029, device='cuda:6') |
| spectral_norm tensor(2.6480, device='cuda:6') |
| frobenius_norm tensor(22.3527, device='cuda:6') |
| spectral_norm tensor(1.3822, device='cuda:6') |
| frobenius_norm tensor(57.0854, device='cuda:6') |
| spectral_norm tensor(4.1727, device='cuda:6') |
| frobenius_norm tensor(141.8071, device='cuda:6') |
| spectral_norm tensor(14.3111, device='cuda:6') |
| frobenius_norm tensor(133.1904, device='cuda:6') |
| spectral_norm tensor(5.2121, device='cuda:6') |
| frobenius_norm tensor(132.4129, device='cuda:6') |
| spectral_norm tensor(9.0536, device='cuda:6') |
| stable_rank value of layer 10 ---247.63009643554688 |
| Processing layer 11--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(64.2904, device='cuda:6') |
| spectral_norm tensor(4.0832, device='cuda:6') |
| frobenius_norm tensor(28.0532, device='cuda:6') |
| spectral_norm tensor(2.6312, device='cuda:6') |
| frobenius_norm tensor(19.8152, device='cuda:6') |
| spectral_norm tensor(1.4528, device='cuda:6') |
| frobenius_norm tensor(55.3427, device='cuda:6') |
| spectral_norm tensor(4.3578, device='cuda:6') |
| frobenius_norm tensor(139.0032, device='cuda:6') |
| spectral_norm tensor(12.7478, device='cuda:6') |
| frobenius_norm tensor(135.5063, device='cuda:6') |
| spectral_norm tensor(5.4040, device='cuda:6') |
| frobenius_norm tensor(134.1321, device='cuda:6') |
| spectral_norm tensor(9.8842, device='cuda:6') |
| stable_rank value of layer 11 ---234.3900146484375 |
| Processing layer 12--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(62.3252, device='cuda:6') |
| spectral_norm tensor(3.9938, device='cuda:6') |
| frobenius_norm tensor(26.8499, device='cuda:6') |
| spectral_norm tensor(2.5529, device='cuda:6') |
| frobenius_norm tensor(20.6057, device='cuda:6') |
| spectral_norm tensor(1.3954, device='cuda:6') |
| frobenius_norm tensor(55.6693, device='cuda:6') |
| spectral_norm tensor(4.6503, device='cuda:6') |
| frobenius_norm tensor(136.5591, device='cuda:6') |
| spectral_norm tensor(12.8626, device='cuda:6') |
| frobenius_norm tensor(136.9887, device='cuda:6') |
| spectral_norm tensor(5.2525, device='cuda:6') |
| frobenius_norm tensor(135.8750, device='cuda:6') |
| spectral_norm tensor(10.9823, device='cuda:6') |
| stable_rank value of layer 12 ---237.3602294921875 |
| Processing layer 13--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(61.0881, device='cuda:6') |
| spectral_norm tensor(4.2847, device='cuda:6') |
| frobenius_norm tensor(25.6297, device='cuda:6') |
| spectral_norm tensor(2.9570, device='cuda:6') |
| frobenius_norm tensor(22.4836, device='cuda:6') |
| spectral_norm tensor(1.3594, device='cuda:6') |
| frobenius_norm tensor(57.5732, device='cuda:6') |
| spectral_norm tensor(5.3631, device='cuda:6') |
| frobenius_norm tensor(139.4029, device='cuda:6') |
| spectral_norm tensor(12.9032, device='cuda:6') |
| frobenius_norm tensor(135.3871, device='cuda:6') |
| spectral_norm tensor(5.2956, device='cuda:6') |
| frobenius_norm tensor(133.7299, device='cuda:6') |
| spectral_norm tensor(10.5605, device='cuda:6') |
| stable_rank value of layer 13 ---228.27247619628906 |
| Processing layer 14--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(59.9385, device='cuda:6') |
| spectral_norm tensor(4.1989, device='cuda:6') |
| frobenius_norm tensor(25.3315, device='cuda:6') |
| spectral_norm tensor(2.3592, device='cuda:6') |
| frobenius_norm tensor(19.7791, device='cuda:6') |
| spectral_norm tensor(1.4382, device='cuda:6') |
| frobenius_norm tensor(54.6680, device='cuda:6') |
| spectral_norm tensor(4.8688, device='cuda:6') |
| frobenius_norm tensor(136.9790, device='cuda:6') |
| spectral_norm tensor(12.1678, device='cuda:6') |
| frobenius_norm tensor(136.4235, device='cuda:6') |
| spectral_norm tensor(5.1185, device='cuda:6') |
| frobenius_norm tensor(134.9485, device='cuda:6') |
| spectral_norm tensor(9.6906, device='cuda:6') |
| stable_rank value of layer 14 ---237.89956665039062 |
| Processing layer 15--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(62.2632, device='cuda:6') |
| spectral_norm tensor(4.4386, device='cuda:6') |
| frobenius_norm tensor(26.5240, device='cuda:6') |
| spectral_norm tensor(2.5572, device='cuda:6') |
| frobenius_norm tensor(20.8667, device='cuda:6') |
| spectral_norm tensor(1.5160, device='cuda:6') |
| frobenius_norm tensor(55.3834, device='cuda:6') |
| spectral_norm tensor(4.4431, device='cuda:6') |
| frobenius_norm tensor(135.3356, device='cuda:6') |
| spectral_norm tensor(11.3039, device='cuda:6') |
| frobenius_norm tensor(137.9369, device='cuda:6') |
| spectral_norm tensor(5.1105, device='cuda:6') |
| frobenius_norm tensor(135.9716, device='cuda:6') |
| spectral_norm tensor(10.2615, device='cuda:6') |
| stable_rank value of layer 15 ---242.37692260742188 |
| Processing layer 16--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(59.5504, device='cuda:6') |
| spectral_norm tensor(4.1705, device='cuda:6') |
| frobenius_norm tensor(23.7914, device='cuda:6') |
| spectral_norm tensor(2.3827, device='cuda:6') |
| frobenius_norm tensor(22.4299, device='cuda:6') |
| spectral_norm tensor(1.8284, device='cuda:6') |
| frobenius_norm tensor(57.2100, device='cuda:6') |
| spectral_norm tensor(5.3112, device='cuda:6') |
| frobenius_norm tensor(135.1444, device='cuda:6') |
| spectral_norm tensor(11.1679, device='cuda:6') |
| frobenius_norm tensor(137.6841, device='cuda:6') |
| spectral_norm tensor(5.1499, device='cuda:6') |
| frobenius_norm tensor(135.1051, device='cuda:6') |
| spectral_norm tensor(10.5823, device='cuda:6') |
| stable_rank value of layer 16 ---227.75900268554688 |
| Processing layer 17--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(61.2952, device='cuda:6') |
| spectral_norm tensor(3.9247, device='cuda:6') |
| frobenius_norm tensor(24.2607, device='cuda:6') |
| spectral_norm tensor(2.3428, device='cuda:6') |
| frobenius_norm tensor(22.0993, device='cuda:6') |
| spectral_norm tensor(1.6527, device='cuda:6') |
| frobenius_norm tensor(56.9147, device='cuda:6') |
| spectral_norm tensor(4.7962, device='cuda:6') |
| frobenius_norm tensor(133.7680, device='cuda:6') |
| spectral_norm tensor(11.1640, device='cuda:6') |
| frobenius_norm tensor(138.4800, device='cuda:6') |
| spectral_norm tensor(5.2358, device='cuda:6') |
| frobenius_norm tensor(135.3299, device='cuda:6') |
| spectral_norm tensor(8.6134, device='cuda:6') |
| stable_rank value of layer 17 ---251.5305633544922 |
| Processing layer 18--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(57.8086, device='cuda:6') |
| spectral_norm tensor(4.1784, device='cuda:6') |
| frobenius_norm tensor(23.1346, device='cuda:6') |
| spectral_norm tensor(2.7649, device='cuda:6') |
| frobenius_norm tensor(24.9173, device='cuda:6') |
| spectral_norm tensor(1.5424, device='cuda:6') |
| frobenius_norm tensor(60.8592, device='cuda:6') |
| spectral_norm tensor(5.2139, device='cuda:6') |
| frobenius_norm tensor(133.3184, device='cuda:6') |
| spectral_norm tensor(10.9756, device='cuda:6') |
| frobenius_norm tensor(140.8582, device='cuda:6') |
| spectral_norm tensor(5.6010, device='cuda:6') |
| frobenius_norm tensor(137.9481, device='cuda:6') |
| spectral_norm tensor(7.5009, device='cuda:6') |
| stable_rank value of layer 18 ---253.84030151367188 |
| Processing layer 19--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(57.4449, device='cuda:6') |
| spectral_norm tensor(4.5447, device='cuda:6') |
| frobenius_norm tensor(21.3000, device='cuda:6') |
| spectral_norm tensor(2.4431, device='cuda:6') |
| frobenius_norm tensor(25.0779, device='cuda:6') |
| spectral_norm tensor(1.6237, device='cuda:6') |
| frobenius_norm tensor(59.6389, device='cuda:6') |
| spectral_norm tensor(4.9505, device='cuda:6') |
| frobenius_norm tensor(135.1317, device='cuda:6') |
| spectral_norm tensor(12.0810, device='cuda:6') |
| frobenius_norm tensor(140.5321, device='cuda:6') |
| spectral_norm tensor(5.5793, device='cuda:6') |
| frobenius_norm tensor(137.1833, device='cuda:6') |
| spectral_norm tensor(6.9554, device='cuda:6') |
| stable_rank value of layer 19 ---252.57142639160156 |
| Processing layer 20--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(58.7482, device='cuda:6') |
| spectral_norm tensor(4.0858, device='cuda:6') |
| frobenius_norm tensor(22.3411, device='cuda:6') |
| spectral_norm tensor(2.3521, device='cuda:6') |
| frobenius_norm tensor(25.9143, device='cuda:6') |
| spectral_norm tensor(1.7104, device='cuda:6') |
| frobenius_norm tensor(60.9286, device='cuda:6') |
| spectral_norm tensor(4.8709, device='cuda:6') |
| frobenius_norm tensor(134.6000, device='cuda:6') |
| spectral_norm tensor(11.1297, device='cuda:6') |
| frobenius_norm tensor(141.2537, device='cuda:6') |
| spectral_norm tensor(5.9169, device='cuda:6') |
| frobenius_norm tensor(138.0343, device='cuda:6') |
| spectral_norm tensor(6.5506, device='cuda:6') |
| stable_rank value of layer 20 ---263.31256103515625 |
| Processing layer 21--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(57.1627, device='cuda:6') |
| spectral_norm tensor(3.6813, device='cuda:6') |
| frobenius_norm tensor(19.9702, device='cuda:6') |
| spectral_norm tensor(2.2489, device='cuda:6') |
| frobenius_norm tensor(27.7738, device='cuda:6') |
| spectral_norm tensor(1.6762, device='cuda:6') |
| frobenius_norm tensor(63.2683, device='cuda:6') |
| spectral_norm tensor(5.2737, device='cuda:6') |
| frobenius_norm tensor(136.7462, device='cuda:6') |
| spectral_norm tensor(11.6961, device='cuda:6') |
| frobenius_norm tensor(141.2342, device='cuda:6') |
| spectral_norm tensor(5.6582, device='cuda:6') |
| frobenius_norm tensor(137.9897, device='cuda:6') |
| spectral_norm tensor(5.3100, device='cuda:6') |
| stable_rank value of layer 21 ---310.4990234375 |
| Processing layer 22--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(58.1757, device='cuda:6') |
| spectral_norm tensor(4.0771, device='cuda:6') |
| frobenius_norm tensor(19.4699, device='cuda:6') |
| spectral_norm tensor(1.9257, device='cuda:6') |
| frobenius_norm tensor(27.1717, device='cuda:6') |
| spectral_norm tensor(2.0305, device='cuda:6') |
| frobenius_norm tensor(63.5510, device='cuda:6') |
| spectral_norm tensor(4.3551, device='cuda:6') |
| frobenius_norm tensor(136.9238, device='cuda:6') |
| spectral_norm tensor(12.4543, device='cuda:6') |
| frobenius_norm tensor(141.9774, device='cuda:6') |
| spectral_norm tensor(6.5433, device='cuda:6') |
| frobenius_norm tensor(139.0448, device='cuda:6') |
| spectral_norm tensor(5.1260, device='cuda:6') |
| stable_rank value of layer 22 ---289.32818603515625 |
| Processing layer 23--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(59.5436, device='cuda:6') |
| spectral_norm tensor(4.0933, device='cuda:6') |
| frobenius_norm tensor(20.5584, device='cuda:6') |
| spectral_norm tensor(2.1754, device='cuda:6') |
| frobenius_norm tensor(27.5827, device='cuda:6') |
| spectral_norm tensor(1.8698, device='cuda:6') |
| frobenius_norm tensor(64.9988, device='cuda:6') |
| spectral_norm tensor(5.3918, device='cuda:6') |
| frobenius_norm tensor(138.1573, device='cuda:6') |
| spectral_norm tensor(14.3250, device='cuda:6') |
| frobenius_norm tensor(141.6219, device='cuda:6') |
| spectral_norm tensor(6.6344, device='cuda:6') |
| frobenius_norm tensor(138.9087, device='cuda:6') |
| spectral_norm tensor(5.2948, device='cuda:6') |
| stable_rank value of layer 23 ---271.5490417480469 |
| Processing layer 24--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(56.8169, device='cuda:6') |
| spectral_norm tensor(3.8277, device='cuda:6') |
| frobenius_norm tensor(19.4813, device='cuda:6') |
| spectral_norm tensor(1.8291, device='cuda:6') |
| frobenius_norm tensor(31.0836, device='cuda:6') |
| spectral_norm tensor(2.3168, device='cuda:6') |
| frobenius_norm tensor(65.3215, device='cuda:6') |
| spectral_norm tensor(6.2412, device='cuda:6') |
| frobenius_norm tensor(135.7100, device='cuda:6') |
| spectral_norm tensor(12.7437, device='cuda:6') |
| frobenius_norm tensor(143.0019, device='cuda:6') |
| spectral_norm tensor(6.6948, device='cuda:6') |
| frobenius_norm tensor(141.1184, device='cuda:6') |
| spectral_norm tensor(5.4309, device='cuda:6') |
| stable_rank value of layer 24 ---266.88092041015625 |
| Processing layer 25--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(55.3445, device='cuda:6') |
| spectral_norm tensor(3.8164, device='cuda:6') |
| frobenius_norm tensor(17.7185, device='cuda:6') |
| spectral_norm tensor(1.8025, device='cuda:6') |
| frobenius_norm tensor(33.9660, device='cuda:6') |
| spectral_norm tensor(3.0984, device='cuda:6') |
| frobenius_norm tensor(68.8981, device='cuda:6') |
| spectral_norm tensor(5.3780, device='cuda:6') |
| frobenius_norm tensor(134.2979, device='cuda:6') |
| spectral_norm tensor(12.3842, device='cuda:6') |
| frobenius_norm tensor(144.5099, device='cuda:6') |
| spectral_norm tensor(7.7998, device='cuda:6') |
| frobenius_norm tensor(143.8913, device='cuda:6') |
| spectral_norm tensor(6.7086, device='cuda:6') |
| stable_rank value of layer 25 ---216.0190887451172 |
| Processing layer 26--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(51.7868, device='cuda:6') |
| spectral_norm tensor(3.9548, device='cuda:6') |
| frobenius_norm tensor(16.9575, device='cuda:6') |
| spectral_norm tensor(1.9295, device='cuda:6') |
| frobenius_norm tensor(40.6153, device='cuda:6') |
| spectral_norm tensor(3.1825, device='cuda:6') |
| frobenius_norm tensor(71.6628, device='cuda:6') |
| spectral_norm tensor(6.2830, device='cuda:6') |
| frobenius_norm tensor(134.3241, device='cuda:6') |
| spectral_norm tensor(10.3940, device='cuda:6') |
| frobenius_norm tensor(145.9831, device='cuda:6') |
| spectral_norm tensor(10.3400, device='cuda:6') |
| frobenius_norm tensor(143.8145, device='cuda:6') |
| spectral_norm tensor(5.9343, device='cuda:6') |
| stable_rank value of layer 26 ---213.61611938476562 |
| Processing layer 27--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| frobenius_norm tensor(56.6802, device='cuda:6') |
| spectral_norm tensor(7.8725, device='cuda:6') |
| frobenius_norm tensor(18.0258, device='cuda:6') |
| spectral_norm tensor(2.1181, device='cuda:6') |
| frobenius_norm tensor(36.7123, device='cuda:6') |
| spectral_norm tensor(4.5755, device='cuda:6') |
| frobenius_norm tensor(66.9187, device='cuda:6') |
| spectral_norm tensor(10.3449, device='cuda:6') |
| frobenius_norm tensor(139.7623, device='cuda:6') |
| spectral_norm tensor(14.5234, device='cuda:6') |
| frobenius_norm tensor(145.0968, device='cuda:6') |
| spectral_norm tensor(16.8470, device='cuda:6') |
| frobenius_norm tensor(133.4643, device='cuda:6') |
| spectral_norm tensor(7.9846, device='cuda:6') |
| stable_rank value of layer 27 ---96.66703033447266 |
| [(310.4990234375, 21), (293.3992919921875, 8), (289.32818603515625, 22), (278.4970703125, 6), (271.5490417480469, 23), (266.88092041015625, 24), (263.31256103515625, 20), (262.54864501953125, 7), (253.84030151367188, 18), (252.57142639160156, 19), (251.5305633544922, 17), (247.63009643554688, 10), (242.37692260742188, 15), (237.89956665039062, 14), (237.3602294921875, 12), (234.3900146484375, 11), (228.27247619628906, 13), (227.75900268554688, 16), (216.0190887451172, 25), (213.61611938476562, 26), (208.82086181640625, 9), (200.63299560546875, 4), (187.18458557128906, 5), (157.4268035888672, 2), (143.3182830810547, 3), (108.18235778808594, 0), (96.66703033447266, 27), (91.95645141601562, 1)] |
| metric_name stable_rank: [21, 8, 22, 6, 23, 24, 20, 7, 18, 19, 17, 10, 15, 14, 12, 11, 13, 16, 25, 26, 9, 4, 5, 2, 3, 0, 27, 1] |
| 当前指标:effective_rank |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:07<00:21, 7.09s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:13<00:13, 6.97s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:21<00:07, 7.01s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:27<00:00, 6.70s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:27<00:00, 6.82s/it] |
| Setting `pad_token_id` to `eos_token_id`:151643 for open-end generation. |
| Once upon a time, there was a small town called Harmonyville. The town had a population of 2,000 people, and 1,500 of them had access to the internet. A new survey found that 1,350 people in Harmonyville used the internet for social media, while 1,000 people used it for online shopping. How many people in Harmonyville used the internet for both social media and online shopping? |
| |
| To determine how many people in Harmonyville |
| Qwen2ForCausalLM( |
| (model): Qwen2Model( |
| (embed_tokens): Embedding(152064, 3584) |
| (layers): ModuleList( |
| (0-27): 28 x Qwen2DecoderLayer( |
| (self_attn): Qwen2Attention( |
| (q_proj): Linear(in_features=3584, out_features=3584, bias=True) |
| (k_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (v_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (o_proj): Linear(in_features=3584, out_features=3584, bias=False) |
| ) |
| (mlp): Qwen2MLP( |
| (gate_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (up_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (down_proj): Linear(in_features=18944, out_features=3584, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (post_attention_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| ) |
| ) |
| (norm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (rotary_emb): Qwen2RotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=3584, out_features=152064, bias=False) |
| ) |
| config: |
| Qwen2Config { |
| "architectures": [ |
| "Qwen2ForCausalLM" |
| ], |
| "attention_dropout": 0.0, |
| "bos_token_id": 151643, |
| "dtype": "float16", |
| "eos_token_id": 151643, |
| "hidden_act": "silu", |
| "hidden_size": 3584, |
| "initializer_range": 0.02, |
| "intermediate_size": 18944, |
| "layer_types": [ |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention" |
| ], |
| "max_position_embeddings": 131072, |
| "max_window_layers": 28, |
| "model_type": "qwen2", |
| "num_attention_heads": 28, |
| "num_hidden_layers": 28, |
| "num_key_value_heads": 4, |
| "rms_norm_eps": 1e-06, |
| "rope_scaling": null, |
| "rope_theta": 1000000.0, |
| "sliding_window": null, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "use_mrope": false, |
| "use_sliding_window": false, |
| "vocab_size": 152064 |
| } |
| |
| Processing layer 0--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 0 ---2094.5888671875 |
| Processing layer 1--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 1 ---2108.23583984375 |
| Processing layer 2--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 2 ---2268.78173828125 |
| Processing layer 3--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 3 ---2322.7314453125 |
| Processing layer 4--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 4 ---2320.885009765625 |
| Processing layer 5--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 5 ---2335.18017578125 |
| Processing layer 6--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 6 ---2337.6826171875 |
| Processing layer 7--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 7 ---2348.204833984375 |
| Processing layer 8--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 8 ---2334.47705078125 |
| Processing layer 9--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 9 ---2323.93115234375 |
| Processing layer 10--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 10 ---2351.43994140625 |
| Processing layer 11--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 11 ---2331.3955078125 |
| Processing layer 12--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 12 ---2333.96142578125 |
| Processing layer 13--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 13 ---2327.25927734375 |
| Processing layer 14--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 14 ---2299.207763671875 |
| Processing layer 15--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 15 ---2312.1357421875 |
| Processing layer 16--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 16 ---2319.4013671875 |
| Processing layer 17--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 17 ---2322.21826171875 |
| Processing layer 18--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 18 ---2311.623779296875 |
| Processing layer 19--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 19 ---2325.7373046875 |
| Processing layer 20--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 20 ---2334.143798828125 |
| Processing layer 21--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 21 ---2338.0810546875 |
| Processing layer 22--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 22 ---2328.93994140625 |
| Processing layer 23--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 23 ---2365.697021484375 |
| Processing layer 24--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 24 ---2355.1064453125 |
| Processing layer 25--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 25 ---2361.042236328125 |
| Processing layer 26--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 26 ---2345.724365234375 |
| Processing layer 27--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| effective_rank value of layer 27 ---2320.07958984375 |
| [(2365.697021484375, 23), (2361.042236328125, 25), (2355.1064453125, 24), (2351.43994140625, 10), (2348.204833984375, 7), (2345.724365234375, 26), (2338.0810546875, 21), (2337.6826171875, 6), (2335.18017578125, 5), (2334.47705078125, 8), (2334.143798828125, 20), (2333.96142578125, 12), (2331.3955078125, 11), (2328.93994140625, 22), (2327.25927734375, 13), (2325.7373046875, 19), (2323.93115234375, 9), (2322.7314453125, 3), (2322.21826171875, 17), (2320.885009765625, 4), (2320.07958984375, 27), (2319.4013671875, 16), (2312.1357421875, 15), (2311.623779296875, 18), (2299.207763671875, 14), (2268.78173828125, 2), (2108.23583984375, 1), (2094.5888671875, 0)] |
| metric_name effective_rank: [23, 25, 24, 10, 7, 26, 21, 6, 5, 8, 20, 12, 11, 22, 13, 19, 9, 3, 17, 4, 27, 16, 15, 18, 14, 2, 1, 0] |
| 当前指标:head_diversity |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:07<00:21, 7.32s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:14<00:14, 7.20s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:21<00:07, 7.17s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 6.93s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 7.03s/it] |
| Setting `pad_token_id` to `eos_token_id`:151643 for open-end generation. |
| Once upon a time, in a small town, there lived a young girl named Lily. She was known for her love of music and her beautiful voice. One day, she decided to audition for the local symphony orchestra. She was nervous but excited, and after a few weeks, she received a call. She had been accepted! |
| |
| Lily was thrilled and worked hard to prepare for her first performance. On the day of the concert, she was surrounded by all the other musicians, each playing their instruments with precision and |
| Qwen2ForCausalLM( |
| (model): Qwen2Model( |
| (embed_tokens): Embedding(152064, 3584) |
| (layers): ModuleList( |
| (0-27): 28 x Qwen2DecoderLayer( |
| (self_attn): Qwen2Attention( |
| (q_proj): Linear(in_features=3584, out_features=3584, bias=True) |
| (k_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (v_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (o_proj): Linear(in_features=3584, out_features=3584, bias=False) |
| ) |
| (mlp): Qwen2MLP( |
| (gate_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (up_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (down_proj): Linear(in_features=18944, out_features=3584, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (post_attention_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| ) |
| ) |
| (norm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (rotary_emb): Qwen2RotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=3584, out_features=152064, bias=False) |
| ) |
| config: |
| Qwen2Config { |
| "architectures": [ |
| "Qwen2ForCausalLM" |
| ], |
| "attention_dropout": 0.0, |
| "bos_token_id": 151643, |
| "dtype": "float16", |
| "eos_token_id": 151643, |
| "hidden_act": "silu", |
| "hidden_size": 3584, |
| "initializer_range": 0.02, |
| "intermediate_size": 18944, |
| "layer_types": [ |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention" |
| ], |
| "max_position_embeddings": 131072, |
| "max_window_layers": 28, |
| "model_type": "qwen2", |
| "num_attention_heads": 28, |
| "num_hidden_layers": 28, |
| "num_key_value_heads": 4, |
| "rms_norm_eps": 1e-06, |
| "rope_scaling": null, |
| "rope_theta": 1000000.0, |
| "sliding_window": null, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "use_mrope": false, |
| "use_sliding_window": false, |
| "vocab_size": 152064 |
| } |
| |
| Processing layer 0--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 0 ---0.9926437139511108 |
| Processing layer 1--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 1 ---0.990115761756897 |
| Processing layer 2--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 2 ---0.9928579330444336 |
| Processing layer 3--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 3 ---0.9908387064933777 |
| Processing layer 4--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 4 ---0.9908434748649597 |
| Processing layer 5--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 5 ---0.9908813834190369 |
| Processing layer 6--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 6 ---0.9935818314552307 |
| Processing layer 7--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 7 ---0.9881868362426758 |
| Processing layer 8--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 8 ---0.9893801212310791 |
| Processing layer 9--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 9 ---0.9851914644241333 |
| Processing layer 10--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 10 ---0.9917221069335938 |
| Processing layer 11--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 11 ---0.988156795501709 |
| Processing layer 12--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 12 ---0.9904806017875671 |
| Processing layer 13--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 13 ---0.9887334108352661 |
| Processing layer 14--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 14 ---0.984466552734375 |
| Processing layer 15--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 15 ---0.9898031949996948 |
| Processing layer 16--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 16 ---0.9885467290878296 |
| Processing layer 17--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 17 ---0.9883378744125366 |
| Processing layer 18--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 18 ---0.987558901309967 |
| Processing layer 19--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 19 ---0.9857799410820007 |
| Processing layer 20--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 20 ---0.9886536598205566 |
| Processing layer 21--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 21 ---0.988256573677063 |
| Processing layer 22--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 22 ---0.9854484796524048 |
| Processing layer 23--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 23 ---0.9908662438392639 |
| Processing layer 24--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 24 ---0.9898425936698914 |
| Processing layer 25--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 25 ---0.9891295433044434 |
| Processing layer 26--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 26 ---0.9852111339569092 |
| Processing layer 27--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| head_diversity value of layer 27 ---0.9841814637184143 |
| [(0.9935818314552307, 6), (0.9928579330444336, 2), (0.9926437139511108, 0), (0.9917221069335938, 10), (0.9908813834190369, 5), (0.9908662438392639, 23), (0.9908434748649597, 4), (0.9908387064933777, 3), (0.9904806017875671, 12), (0.990115761756897, 1), (0.9898425936698914, 24), (0.9898031949996948, 15), (0.9893801212310791, 8), (0.9891295433044434, 25), (0.9887334108352661, 13), (0.9886536598205566, 20), (0.9885467290878296, 16), (0.9883378744125366, 17), (0.988256573677063, 21), (0.9881868362426758, 7), (0.988156795501709, 11), (0.987558901309967, 18), (0.9857799410820007, 19), (0.9854484796524048, 22), (0.9852111339569092, 26), (0.9851914644241333, 9), (0.984466552734375, 14), (0.9841814637184143, 27)] |
| metric_name head_diversity: [6, 2, 0, 10, 5, 23, 4, 3, 12, 1, 24, 15, 8, 25, 13, 20, 16, 17, 21, 7, 11, 18, 19, 22, 26, 9, 14, 27] |
| 当前指标:coherence |
| `torch_dtype` is deprecated! Use `dtype` instead! |
|
Loading checkpoint shards: 0%| | 0/4 [00:00<?, ?it/s]
Loading checkpoint shards: 25%|██▌ | 1/4 [00:07<00:21, 7.31s/it]
Loading checkpoint shards: 50%|█████ | 2/4 [00:14<00:14, 7.20s/it]
Loading checkpoint shards: 75%|███████▌ | 3/4 [00:21<00:07, 7.12s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 6.93s/it]
Loading checkpoint shards: 100%|██████████| 4/4 [00:28<00:00, 7.03s/it] |
| Setting `pad_token_id` to `eos_token_id`:151643 for open-end generation. |
| Once upon a time,there lived a king.He had three daughters.One day,the king asked his daughters,"Which of you loves me most?" The two elder daughters answered at once. "I love you most,Father," they said.But the youngest daughter was quiet for a while.Then she said, "I love you more than they do,Father." The king was not happy with her answer.So he gave each of them a coin and said, "Go to the street and spend this coin in any |
| Qwen2ForCausalLM( |
| (model): Qwen2Model( |
| (embed_tokens): Embedding(152064, 3584) |
| (layers): ModuleList( |
| (0-27): 28 x Qwen2DecoderLayer( |
| (self_attn): Qwen2Attention( |
| (q_proj): Linear(in_features=3584, out_features=3584, bias=True) |
| (k_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (v_proj): Linear(in_features=3584, out_features=512, bias=True) |
| (o_proj): Linear(in_features=3584, out_features=3584, bias=False) |
| ) |
| (mlp): Qwen2MLP( |
| (gate_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (up_proj): Linear(in_features=3584, out_features=18944, bias=False) |
| (down_proj): Linear(in_features=18944, out_features=3584, bias=False) |
| (act_fn): SiLUActivation() |
| ) |
| (input_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (post_attention_layernorm): Qwen2RMSNorm((3584,), eps=1e-06) |
| ) |
| ) |
| (norm): Qwen2RMSNorm((3584,), eps=1e-06) |
| (rotary_emb): Qwen2RotaryEmbedding() |
| ) |
| (lm_head): Linear(in_features=3584, out_features=152064, bias=False) |
| ) |
| config: |
| Qwen2Config { |
| "architectures": [ |
| "Qwen2ForCausalLM" |
| ], |
| "attention_dropout": 0.0, |
| "bos_token_id": 151643, |
| "dtype": "float16", |
| "eos_token_id": 151643, |
| "hidden_act": "silu", |
| "hidden_size": 3584, |
| "initializer_range": 0.02, |
| "intermediate_size": 18944, |
| "layer_types": [ |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention", |
| "full_attention" |
| ], |
| "max_position_embeddings": 131072, |
| "max_window_layers": 28, |
| "model_type": "qwen2", |
| "num_attention_heads": 28, |
| "num_hidden_layers": 28, |
| "num_key_value_heads": 4, |
| "rms_norm_eps": 1e-06, |
| "rope_scaling": null, |
| "rope_theta": 1000000.0, |
| "sliding_window": null, |
| "tie_word_embeddings": false, |
| "transformers_version": "4.57.3", |
| "use_cache": true, |
| "use_mrope": false, |
| "use_sliding_window": false, |
| "vocab_size": 152064 |
| } |
| |
| Processing layer 0--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 0 ---0.02408885583281517 |
| Processing layer 1--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 1 ---0.056700460612773895 |
| Processing layer 2--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 2 ---0.02794474922120571 |
| Processing layer 3--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 3 ---0.018478330224752426 |
| Processing layer 4--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 4 ---0.0186083372682333 |
| Processing layer 5--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 5 ---0.017582347616553307 |
| Processing layer 6--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 6 ---0.017218835651874542 |
| Processing layer 7--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 7 ---0.01599242351949215 |
| Processing layer 8--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 8 ---0.01598191447556019 |
| Processing layer 9--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 9 ---0.01899208500981331 |
| Processing layer 10--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 10 ---0.015675809234380722 |
| Processing layer 11--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 11 ---0.01689709722995758 |
| Processing layer 12--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 12 ---0.016846707090735435 |
| Processing layer 13--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 13 ---0.01713874563574791 |
| Processing layer 14--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 14 ---0.016889071092009544 |
| Processing layer 15--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 15 ---0.017316650599241257 |
| Processing layer 16--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 16 ---0.0172375850379467 |
| Processing layer 17--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 17 ---0.016330666840076447 |
| Processing layer 18--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 18 ---0.017193637788295746 |
| Processing layer 19--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 19 ---0.01639077439904213 |
| Processing layer 20--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 20 ---0.016309073194861412 |
| Processing layer 21--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 21 ---0.0156699325889349 |
| Processing layer 22--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 22 ---0.016703739762306213 |
| Processing layer 23--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 23 ---0.016167931258678436 |
| Processing layer 24--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 24 ---0.01581120304763317 |
| Processing layer 25--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 25 ---0.016396205872297287 |
| Processing layer 26--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 26 ---0.017451899126172066 |
| Processing layer 27--subset--{'self_attn.q_proj': Linear(in_features=3584, out_features=3584, bias=True), 'self_attn.k_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.v_proj': Linear(in_features=3584, out_features=512, bias=True), 'self_attn.o_proj': Linear(in_features=3584, out_features=3584, bias=False), 'mlp.gate_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.up_proj': Linear(in_features=3584, out_features=18944, bias=False), 'mlp.down_proj': Linear(in_features=18944, out_features=3584, bias=False)} |
| coherence value of layer 27 ---0.02082100510597229 |
| [(0.056700460612773895, 1), (0.02794474922120571, 2), (0.02408885583281517, 0), (0.02082100510597229, 27), (0.01899208500981331, 9), (0.0186083372682333, 4), (0.018478330224752426, 3), (0.017582347616553307, 5), (0.017451899126172066, 26), (0.017316650599241257, 15), (0.0172375850379467, 16), (0.017218835651874542, 6), (0.017193637788295746, 18), (0.01713874563574791, 13), (0.01689709722995758, 11), (0.016889071092009544, 14), (0.016846707090735435, 12), (0.016703739762306213, 22), (0.016396205872297287, 25), (0.01639077439904213, 19), (0.016330666840076447, 17), (0.016309073194861412, 20), (0.016167931258678436, 23), (0.01599242351949215, 7), (0.01598191447556019, 8), (0.01581120304763317, 24), (0.015675809234380722, 10), (0.0156699325889349, 21)] |
| metric_name coherence: [1, 2, 0, 27, 9, 4, 3, 5, 26, 15, 16, 6, 18, 13, 11, 14, 12, 22, 25, 19, 17, 20, 23, 7, 8, 24, 10, 21] |
| |