PowerMachine commited on
Commit
64ae428
·
verified ·
1 Parent(s): 1b4e2d4

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +200 -0
README.md ADDED
@@ -0,0 +1,200 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # BiGRU_T_version — Refatoração formal do GRU-RING v13.9.2
2
+
3
+ Refatoração matemática do modelo [PowerMachine/gru-ring-v13-9-2](https://huggingface.co/PowerMachine/gru-ring-v13-9-2) em torno de **4 lemas formais** que tornam explícitos os mecanismos de desacoplamento, cirurgia de gradiente, cancelamento de ruído de quantização e auto-configuração.
4
+
5
+ ## Lemas
6
+
7
+ ### Lema 1 — Desacoplamento via atenção hierárquica
8
+ Seleção dos módulos `u8cell_T` por softmax com temperatura controlada (α). Durante o treino, a temperatura é meta-ajustada e a entropia dos pesos é minimizada indiretamente pela regularização de agudeza, forçando especialização e reduzindo o produto interno dos gradientes (disputa).
9
+
10
+ **Implementação**: `src/bigru_t/model/module_selector.py` — `ModuleSelector`
11
+ - `module_logits: nn.Parameter(torch.zeros(K))` — logits aprendíveis
12
+ - `alpha = F.softmax(module_logits / T, dim=0)` no forward
13
+ - `entropy_reg = -lambda_ent * (alpha * log(alpha + eps)).sum()`
14
+
15
+ ### Lema 2 — Gradiente cirúrgico
16
+ `apply_gradient_surgery` implementa a projeção ortogonal quando os gradientes da tarefa principal e da hipótese são conflitantes (produto interno negativo). Garante que ambas as perdas possam ser reduzidas sem interferência destrutiva.
17
+
18
+ **Implementação**: `src/bigru_t/training/gradient_surgery.py`
19
+ - `orthogonalize_gradient(g_main, g_hyp)`: projeção vetorial
20
+ - `apply_gradient_surgery(model, loss_main, loss_hyp)`: calcula gradientes independentes via `torch.autograd.grad`, aplica projeção e atribui `p.grad = g_main + g_hyp_perp`
21
+
22
+ ### Lema 3 — Cancelamento de ruído de quantização
23
+ A hipótese (`hyp_T`) é treinada com a perda sobre `y_final = y_hat + delta`, aprendendo a corrigir o viés e variância introduzidos pela quantização W8A8 (simulada no forward pelas camadas `QuantizedLinear`). `stop_grad_hyp=True` isola o sinal de correção do gradiente principal, e a cirurgia gerencia os parâmetros compartilhados.
24
+
25
+ **Implementação**: `src/bigru_t/quantization/quantized_linear.py` + `src/bigru_t/model/hyp_t.py`
26
+ - `QuantizedLinear(nn.Linear)`: fake quant W8A8 no forward (STE backward)
27
+ - `HypT.forward(o, stop_grad=True)`: aplica `o.detach()` quando stop_grad
28
+ - Ativação condicional: hipótese só é calculada quando `loss_main > tau`
29
+
30
+ ### Lema 4 — Auto-configuração e suavização
31
+ `MetaConfigurator` ajusta temperatura e limiar τ, e penaliza a agudeza da perda (norma do gradiente), empurrando o modelo para mínimos planos onde o ruído de gradiente é menor. A quantidade de módulos ativos é implicitamente controlada pela softmax com temperatura.
32
+
33
+ **Implementação**: `src/bigru_t/training/meta_configurator.py`
34
+ - `log_temperature: nn.Parameter` — `T = exp(log_temperature)` (positivo)
35
+ - `log_tau: nn.Parameter` — `τ = exp(log_tau)` (positivo)
36
+ - `forward_with_meta(x_val, y_val)`: meta-loss = `L_val + λ_s * ||∇_θ L_val||²`
37
+
38
+ ## Arquitetura
39
+
40
+ ```
41
+ x (batch, T, input_dim)
42
+ │
43
+ ├──→ u8cell_T_1 ─→ h_1 ─┐
44
+ ├──→ u8cell_T_2 ─→ h_2 ─┤ Lema 1: alpha = softmax(logits / T)
45
+ ├──→ ... ├─→ H = stack([alpha_k * h_k]) → OrqCell
46
+ └──→ u8cell_T_K ─→ h_K ─┘ │
47
+ ↓
48
+ o (batch, d_cache)
49
+ │
50
+ ┌──────────────┴──────────────┐
51
+ ↓ ↓
52
+ TrainT(o) HypT(o, stop_grad=True)
53
+ │ │
54
+ ↓ ↓
55
+ y_hat delta
56
+ │ │
57
+ └──────→ y_final = y_hat + delta * use_hyp ←── Lema 3
58
+
59
+ Lema 2: gradient_surgery(g_main, g_hyp) → projeção ortogonal
60
+ Lema 4: MetaConfigurator ajusta T (temperatura) e τ (limiar de ativação da hipótese)
61
+ ```
62
+
63
+ ### Componentes
64
+ | Componente | Arquivo | Descrição |
65
+ |---|---|---|
66
+ | `BiGRU4` | `model/bigru4.py` | 4 camadas BiGRU sequenciais |
67
+ | `TransformerUnit` | `model/transformer_unit.py` | Self-attention + FFN para 8 representações |
68
+ | `u8cell_T` | `model/u8cell_t.py` | 8 BiGRU4 paralelas + TransformerUnit |
69
+ | `OrqCell` | `model/orq_cell.py` | Cache aprendível + cross-attention |
70
+ | `TrainT` | `model/train_t.py` | Cabeça de predição principal |
71
+ | `HypT` | `model/hyp_t.py` | Cabeça de hipótese (correção delta) |
72
+ | `ModuleSelector` | `model/module_selector.py` | Lema 1: softmax + entropia |
73
+ | `UnifiedModel` | `model/unified_model.py` | Orquestra todos os componentes |
74
+ | `QuantizedLinear` | `quantization/quantized_linear.py` | Lema 3: W8A8 fake quant |
75
+ | `apply_gradient_surgery` | `training/gradient_surgery.py` | Lema 2: projeção ortogonal |
76
+ | `MetaConfigurator` | `training/meta_configurator.py` | Lema 4: auto-configuração |
77
+ | `KillSwitch` | `training/kill_switch.py` | Monitor RAM/disk/loss + kill automático |
78
+ | `BiGRU_T_Trainer` | `training/trainer.py` | Loop de treino (2 épocas) |
79
+
80
+ ## Reaproveitamento do source (PowerMachine/gru-ring-v13-9-2)
81
+
82
+ | Componente | Source | Target |
83
+ |---|---|---|
84
+ | BBPE Tokenizer | `flexnet/bbpe_tokenizer.py` | `tokenizer/bbpe_tokenizer.py` |
85
+ | Streaming datasets (9 base + 3 PT-BR) | `scripts/streaming_datasets_v13_9.py` | `data/streaming_datasets.py` |
86
+ | W8A8 QOperator (inferência) | `flexnet/w8a8_qoperator.py` | `quantization/w8a8_qoperator.py` |
87
+ | Hardware detector | `flexnet/hardware_detector.py` | `utils/hardware_detector.py` |
88
+ | Xeon runtime | `flexnet/xeon_runtime.py` | `utils/xeon_runtime.py` |
89
+ | OOM guard | `flexnet/oom_guard.py` | `utils/oom_guard.py` |
90
+ | Memory monitor | `flexnet/memory_monitor.py` | `utils/memory_monitor.py` |
91
+ | Tensor ops | `xavante/utils/tensor_ops.py` | `utils/tensor_ops.py` |
92
+ | Validators | `xavante/utils/validators.py` | `utils/validators.py` |
93
+ | Logging utils | `xavante/utils/logging_utils.py` | `utils/logging_utils.py` |
94
+ | Hamiltonian-Wasserstein optimizer | `flexnet/hamiltonian_wasserstein_optimizer.py` | `optim/hamiltonian_wasserstein.py` |
95
+ | Multimodal: text/image/audio/video encoders | `xavante/multimodal/*` | `multimodal/*` |
96
+
97
+ ## Datasets (12, PT-BR)
98
+
99
+ **9 base (v13.9.2)**:
100
+ 1. `CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1`
101
+ 2. `Madras1/corpus-ptbr-v2`
102
+ 3. `rhaymison/multmodal_175k_portuguese`
103
+ 4. `TucanoBR/GigaVerbo`
104
+ 5. `nvidia/OpenMathReasoning`
105
+ 6. `MathLLMs/MathVision`
106
+ 7. `nvidia/OpenMathInstruct-2`
107
+ 8. `dominguesm/restore-punctuation-ptbr-dataset`
108
+ 9. `carolina-c4ai/corpus-carolina`
109
+
110
+ **3 PT-BR finetune**:
111
+ 10. `orion-research/translations-en_US-pt_BR` (format: `### Instruction:/### Response:`)
112
+ 11. `cnmoro/Instruct-PTBR-10M` (format: `### Instruction:/### Response:`)
113
+ 12. `strak2005/corpus-ptbr-v1` (plain text)
114
+
115
+ ## Instalação
116
+
117
+ ```bash
118
+ cd BiGRU_T_version
119
+ pip install -r requirements.txt
120
+ pip install -e .
121
+ ```
122
+
123
+ ## Uso — Treino de bug-detection (2 épocas)
124
+
125
+ ```bash
126
+ export HF_TOKEN="hf_xxx" # para datasets públicos, opcional
127
+ python scripts/train.py \
128
+ --datasets CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1,Madras1/corpus-ptbr-v2 \
129
+ --max-samples 30 \
130
+ --epochs 2 \
131
+ --output-dir model_final
132
+ ```
133
+
134
+ ## Uso — Inferência
135
+
136
+ ```python
137
+ import torch
138
+ from bigru_t import create_unified_model
139
+ from tokenizers import Tokenizer
140
+
141
+ model, config = create_unified_model()
142
+ state = torch.load("model_final/pytorch_model.bin", map_location="cpu")
143
+ model.load_state_dict(state, strict=False)
144
+ model.eval()
145
+
146
+ tokenizer = Tokenizer.from_file("model_final/tokenizer/tokenizer.json")
147
+ input_ids = torch.tensor([tokenizer.encode("O presidente anunciou que").ids])
148
+ y_hat, delta = model(input_ids, temperature=1.0, use_hypothesis=False)
149
+ next_token = y_hat.argmax(dim=-1)
150
+ print(tokenizer.decode([next_token.item()]))
151
+ ```
152
+
153
+ ## Configuração Xeon
154
+
155
+ O ambiente é automaticamente configurado para Xeon com AVX512/VNNI/AMX:
156
+ - `OMP_NUM_THREADS=2`, `MKL_NUM_THREADS=2`
157
+ - `torch.set_num_threads(2)`
158
+ - `optimize_xeon_environment()` (de `utils/xeon_runtime.py`) chamado no startup
159
+
160
+ ## Estrutura
161
+
162
+ ```
163
+ BiGRU_T_version/
164
+ ├── README.md
165
+ ├── requirements.txt
166
+ ├── src/
167
+ │ ├── setup.py
168
+ │ └── bigru_t/
169
+ │ ├── __init__.py
170
+ │ ├── model/ # BiGRU4, TransformerUnit, u8cell_T, OrqCell, TrainT, HypT, ModuleSelector, UnifiedModel
171
+ │ ├── quantization/ # QuantizedLinear (Lema 3), w8a8_qoperator (reaproveitado)
172
+ │ ├── training/ # gradient_surgery (Lema 2), meta_configurator (Lema 4), kill_switch, trainer
173
+ │ ├── data/ # streaming_datasets (reaproveitado, 12 datasets)
174
+ │ ├── multimodal/ # text/image/audio/video encoders (reaproveitado)
175
+ │ ├── optim/ # hamiltonian_wasserstein (reaproveitado)
176
+ │ ├── utils/ # hardware_detector, xeon_runtime, oom_guard, memory_monitor, tensor_ops, validators, logging_utils
177
+ │ └── tokenizer/ # bbpe_tokenizer (reaproveitado)
178
+ ├── scripts/
179
+ │ ├── train.py # Treino de bug-detection (2 épocas)
180
+ │ ├── smoke_test.py # Smoke test do pipeline
181
+ │ └── upload_to_hf.py # Upload para HF
182
+ ├── tests/
183
+ │ ├── test_model.py
184
+ │ ├── test_gradient_surgery.py
185
+ │ ├── test_meta_configurator.py
186
+ │ └── test_w8a8.py
187
+ ├── docs/
188
+ │ ├── analysis.md # Análise matemática dos 4 lemas
189
+ │ └── architecture.md
190
+ ├── configs/
191
+ │ └── default.yaml
192
+ └── model_final/ # Artefatos treinados
193
+ ├── config.json
194
+ ├── pytorch_model.bin
195
+ └── tokenizer/
196
+ ```
197
+
198
+ ## Licença
199
+
200
+ MIT (herdado do source PowerMachine/gru-ring-v13-9-2).