V6.5-final: removed pre-V6.4 modules/scripts, moved kohonen_learning_system up, activated VQ-VAE-2 in pipeline, integrated reasoning_engine, benchmarked EWC+W8A8 dequant, exhausted 6 datasets
Browse files- scripts/train_v6_4.py +1 -1
- scripts/train_v6_5.py +798 -397
- scripts/upload_v6_5_resilient.py +6 -3
- src/bigru_t/__init__.py +30 -22
- src/bigru_t/data/streaming_datasets.py +18 -0
- src/bigru_t/model/__init__.py +28 -11
- src/bigru_t/model/kohonen_learning_system.py +944 -0
- src/bigru_t/training/__init__.py +5 -3
- src/bigru_t/utils/xeon_runtime.py +6 -1
- v6_5_ewc_w8a8_benchmark.json +50 -0
- v6_5_module_analysis.json +103 -61
- v6_5_reasoning_eval.json +90 -0
- v6_5_report.json +282 -229
- v6_5_script_activity.json +11 -74
- v6_5_training_metrics.json +0 -0
- v6_5_upload_report.json +35 -0
scripts/train_v6_4.py
CHANGED
|
@@ -157,7 +157,7 @@ SYNTH_TEMPLATES = {
|
|
| 157 |
# ============================================================================
|
| 158 |
# 3. Import KohonenLearningSystem (V6.4 — refatorado canônico)
|
| 159 |
# ============================================================================
|
| 160 |
-
from bigru_t.model.
|
| 161 |
KohonenLearningSystem,
|
| 162 |
SimpleBBPETokenizer,
|
| 163 |
positional_encoding,
|
|
|
|
| 157 |
# ============================================================================
|
| 158 |
# 3. Import KohonenLearningSystem (V6.4 — refatorado canônico)
|
| 159 |
# ============================================================================
|
| 160 |
+
from bigru_t.model.kohonen_learning_system import ( # noqa: E402
|
| 161 |
KohonenLearningSystem,
|
| 162 |
SimpleBBPETokenizer,
|
| 163 |
positional_encoding,
|
scripts/train_v6_5.py
CHANGED
|
@@ -1,102 +1,53 @@
|
|
| 1 |
-
"""train_v6_5.py — V6.5
|
| 2 |
|
| 3 |
═══════════════════════════════════════════════════════════════════════════════
|
| 4 |
-
V6.5 —
|
| 5 |
═══════════════════════════════════════════════════════════════════════════════
|
| 6 |
|
| 7 |
-
User requirements (V6.5):
|
| 8 |
1. HF_TOKEN (delete after use)
|
| 9 |
2. streaming_datasets.py + xeon_runtime.py ativo
|
| 10 |
-
3. Analisar
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
- bigru_t.data.streaming_datasets (ativo)
|
| 30 |
-
- bigru_t.model.kohonen_refactored.kohonen_learning_system (ativo)
|
| 31 |
-
- bigru_t.model.hyp_t (ativo via kohonen LearningSystem)
|
| 32 |
-
- bigru_t.training.mtp (V6.5: ativo em val + entropy regularizer)
|
| 33 |
-
- bigru_t.training.ewc (V6.5: investigação EWC+W8A8 eval)
|
| 34 |
-
- bigru_t.quantization.smoothquant_compressor (V6.5: integrado)
|
| 35 |
-
- bigru_t.model.vqvae2_hierarchical (V6.5: integrado)
|
| 36 |
-
- bigru_t.model.token_compress (V6.5: integrado)
|
| 37 |
-
- bigru_t.reasoning.thinking (V6.5: integrado)
|
| 38 |
-
|
| 39 |
-
Novos módulos integrados do gru-ring-v13-9-2 (V6.5):
|
| 40 |
-
1. smoothquant_compressor.py — SmoothQuant W8A8 (Lema 4.1)
|
| 41 |
-
2. vqvae2_hierarchical.py — VQ-VAE-2 hierárquico (compressão neural)
|
| 42 |
-
3. vqvae2_hierarchical_flexnet.py — versão FlexNet completa
|
| 43 |
-
4. token_compress.py — compressão de tokens
|
| 44 |
-
5. reasoning_engine.py + dependências (circular_orchestration,
|
| 45 |
-
tool_agent, distributed_reasoning_system, cyclic_reasoning,
|
| 46 |
-
consensus_sampling) — motor de raciocínio
|
| 47 |
-
6. thinking.py (já existente)
|
| 48 |
-
|
| 49 |
-
MTP em val + entropy regularizer (V6.5):
|
| 50 |
-
- MTPConfig.active_in_val = True (já default desde V6)
|
| 51 |
-
- MTPConfig.entropy_beta = 0.01 (já default desde V6)
|
| 52 |
-
- V6.5: efetivamente ATIVA o MTPHead no loop de treino, computa
|
| 53 |
-
mtp_loss em train e val, monitora mtp_alphas + entropy_reg
|
| 54 |
-
|
| 55 |
-
EWC+W8A8 em eval (V6.5):
|
| 56 |
-
- EWCConfigV6.eval_mode_penalty = True (já default)
|
| 57 |
-
- V6.5: documenta a interação:
|
| 58 |
-
* Em eval mode, EWC.compute_penalty(eval_mode=True) computa
|
| 59 |
-
penalidade sem backward (forward-only)
|
| 60 |
-
* W8A8 quantized Linear: pesos/ativações em INT8, mas EWC precisa
|
| 61 |
-
de float para (p - w_star)^2 → usa_dequant na penalidade
|
| 62 |
-
* Investigação: SmoothQuantCompressor mantém scaling factors
|
| 63 |
-
que permitem dequantização barata para EWC
|
| 64 |
-
|
| 65 |
-
1000 samples (500 + 500) streaming (V6.5):
|
| 66 |
-
- Fase 1: 5 datasets × 100 samples = 500 samples
|
| 67 |
-
- Fase 2: 5 datasets × 100 samples = 500 samples (novas seeds)
|
| 68 |
-
- Total: 1000 samples processados em batches de 16
|
| 69 |
-
- Monitoramento por fase + cumulativo para detecção de evolução
|
| 70 |
-
|
| 71 |
-
Script Activity Monitoring (V6.5):
|
| 72 |
-
- Para cada script .py no BiGRU_T_version/scripts/, registra:
|
| 73 |
-
* exists: True/False
|
| 74 |
-
* size_bytes: int
|
| 75 |
-
* mtime: ISO timestamp
|
| 76 |
-
* imported: True/False (tentou importar no início)
|
| 77 |
-
* activity: "active" | "legacy" | "unused"
|
| 78 |
-
- Para cada módulo em src/bigru_t/, registra:
|
| 79 |
-
* imported_by_train: True/False
|
| 80 |
-
* activity: "active" | "transitive" | "dead"
|
| 81 |
|
| 82 |
Saídas:
|
| 83 |
- /home/z/my-project/BiGRU_T_version/v6_5_report.json
|
| 84 |
- /home/z/my-project/BiGRU_T_version/v6_5_training_metrics.json
|
| 85 |
- /home/z/my-project/BiGRU_T_version/v6_5_module_analysis.json
|
| 86 |
- /home/z/my-project/BiGRU_T_version/v6_5_script_activity.json
|
|
|
|
|
|
|
| 87 |
═══════════════════════════════════════════════════════════════════════════════
|
| 88 |
"""
|
| 89 |
from __future__ import annotations
|
| 90 |
|
| 91 |
import json
|
| 92 |
import logging
|
|
|
|
| 93 |
import os
|
| 94 |
import sys
|
| 95 |
import time
|
| 96 |
import traceback
|
| 97 |
from datetime import datetime
|
| 98 |
from pathlib import Path
|
| 99 |
-
from typing import Any, Dict, List, Optional
|
| 100 |
|
| 101 |
# ============================================================================
|
| 102 |
# 0. Paths e logging
|
|
@@ -108,6 +59,8 @@ REPORT_PATH = BIGRU_ROOT / "v6_5_report.json"
|
|
| 108 |
METRICS_PATH = BIGRU_ROOT / "v6_5_training_metrics.json"
|
| 109 |
MODULE_ANALYSIS_PATH = BIGRU_ROOT / "v6_5_module_analysis.json"
|
| 110 |
SCRIPT_ACTIVITY_PATH = BIGRU_ROOT / "v6_5_script_activity.json"
|
|
|
|
|
|
|
| 111 |
|
| 112 |
logging.basicConfig(
|
| 113 |
level=logging.INFO,
|
|
@@ -117,7 +70,7 @@ logging.basicConfig(
|
|
| 117 |
logger = logging.getLogger("train_v6_5")
|
| 118 |
|
| 119 |
# ============================================================================
|
| 120 |
-
# 1. ATIVAR xeon_runtime.py
|
| 121 |
# ============================================================================
|
| 122 |
sys.path.insert(0, str(SRC_ROOT))
|
| 123 |
|
|
@@ -136,18 +89,12 @@ logger.info(f"[V6.5] Xeon FP16 benchmark: {FP16_BENCH}")
|
|
| 136 |
# 2. Configurações V6.5
|
| 137 |
# ============================================================================
|
| 138 |
BATCH_SIZE = 16
|
| 139 |
-
# V6.5: 1000 samples = 500 (fase 1) + 500 (fase 2)
|
| 140 |
-
N_DATASETS = 5
|
| 141 |
-
SAMPLES_PER_DATASET_PHASE = 100 # 100 em 100 (user requirement)
|
| 142 |
-
N_PHASES = 2
|
| 143 |
-
TOTAL_SAMPLES = N_DATASETS * SAMPLES_PER_DATASET_PHASE * N_PHASES # 1000
|
| 144 |
-
EPOCHS = 2
|
| 145 |
MAX_SEQ_LEN = 8
|
| 146 |
|
| 147 |
-
# Kohonen SOM 4D —
|
| 148 |
-
HIDDEN_DIM = 1024
|
| 149 |
-
VOCAB_SIZE = 16384
|
| 150 |
-
SOM_GRID = (
|
| 151 |
T_MAX = 10000
|
| 152 |
N_START = 10
|
| 153 |
LAMBDA_EWC = 0.02
|
|
@@ -160,46 +107,66 @@ MTP_K = 4
|
|
| 160 |
MTP_ENTROPY_BETA = 0.01
|
| 161 |
MTP_ACTIVE_IN_VAL = True
|
| 162 |
|
| 163 |
-
|
| 164 |
-
|
| 165 |
-
|
| 166 |
-
|
| 167 |
-
|
|
|
|
|
|
|
| 168 |
"nvidia/OpenMathInstruct-2",
|
|
|
|
|
|
|
|
|
|
| 169 |
]
|
| 170 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 171 |
SYNTH_TEMPLATES = {
|
| 172 |
-
"
|
| 173 |
-
"o gato dorme na cama", "o cachorro corre no parque",
|
| 174 |
-
"o pássaro voa no céu", "a menina brinca com a boneca",
|
| 175 |
-
"o menino joga bola",
|
| 176 |
-
],
|
| 177 |
-
"dominguesm/restore-punctuation-pttr-dataset": [
|
| 178 |
"o sol nasceu azul hoje", "ela foi ao mercado comprar pão",
|
| 179 |
"nós viajamos para o rio de janeiro", "o livro está sobre a mesa",
|
| 180 |
"a casa tem quatro quartos",
|
| 181 |
],
|
| 182 |
-
"
|
| 183 |
-
"
|
| 184 |
-
"
|
| 185 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 186 |
],
|
| 187 |
"CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1": [
|
| 188 |
"olá como você está hoje", "qual é o seu nome",
|
| 189 |
"pode me ajudar com isso", "obrigado pela ajuda",
|
| 190 |
"até logo e boa noite",
|
| 191 |
],
|
| 192 |
-
"
|
| 193 |
-
"
|
| 194 |
-
"
|
| 195 |
-
"
|
| 196 |
],
|
| 197 |
}
|
| 198 |
|
| 199 |
# ============================================================================
|
| 200 |
-
# 3. Import KohonenLearningSystem
|
| 201 |
# ============================================================================
|
| 202 |
-
from bigru_t.model.
|
| 203 |
KohonenLearningSystem,
|
| 204 |
SimpleBBPETokenizer,
|
| 205 |
positional_encoding,
|
|
@@ -207,11 +174,13 @@ from bigru_t.model.kohonen_refactored.kohonen_learning_system import ( # noqa:
|
|
| 207 |
KohonenSOM4D,
|
| 208 |
HypothesisClassifier,
|
| 209 |
)
|
|
|
|
| 210 |
from bigru_t.training.mtp import ( # noqa: E402
|
| 211 |
MTPHead, MTPConfig, mtp_loss, mtp_entropy_regularizer,
|
| 212 |
)
|
| 213 |
from bigru_t.training.ewc import EWCV6, EWCConfigV6 # noqa: E402
|
| 214 |
from bigru_t.quantization.smoothquant_compressor import SmoothQuantCompressor # noqa: E402
|
|
|
|
| 215 |
|
| 216 |
logger.info(
|
| 217 |
f"[V6.5] All modules imported. "
|
|
@@ -221,12 +190,12 @@ logger.info(
|
|
| 221 |
|
| 222 |
|
| 223 |
# ============================================================================
|
| 224 |
-
# 4. Module Access Analysis (user requirement: "
|
| 225 |
# ============================================================================
|
| 226 |
def analyze_module_access() -> Dict[str, Any]:
|
| 227 |
-
"""
|
| 228 |
-
# Módulos
|
| 229 |
-
|
| 230 |
"src/bigru_t/model/bigru4.py",
|
| 231 |
"src/bigru_t/model/gru_hierarchy.py",
|
| 232 |
"src/bigru_t/model/orq_cell.py",
|
|
@@ -234,123 +203,429 @@ def analyze_module_access() -> Dict[str, Any]:
|
|
| 234 |
"src/bigru_t/model/transformer_unit.py",
|
| 235 |
"src/bigru_t/model/u8cell_t.py",
|
| 236 |
"src/bigru_t/model/unified_model.py",
|
|
|
|
|
|
|
| 237 |
]
|
| 238 |
-
|
| 239 |
-
|
| 240 |
-
|
| 241 |
-
|
| 242 |
-
|
| 243 |
-
|
|
|
|
| 244 |
"src/bigru_t/model/hyp_t.py",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 245 |
"src/bigru_t/training/mtp.py",
|
| 246 |
"src/bigru_t/training/ewc.py",
|
| 247 |
"src/bigru_t/quantization/smoothquant_compressor.py",
|
| 248 |
-
"src/bigru_t/
|
| 249 |
-
"src/bigru_t/
|
|
|
|
| 250 |
"src/bigru_t/reasoning/thinking.py",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 251 |
]
|
| 252 |
-
|
| 253 |
analysis = {
|
| 254 |
-
"
|
| 255 |
-
"
|
| 256 |
-
"
|
| 257 |
-
|
| 258 |
-
"
|
| 259 |
-
"
|
| 260 |
-
"
|
| 261 |
-
"
|
| 262 |
-
|
| 263 |
-
"src/bigru_t/reasoning/tool_agent.py",
|
| 264 |
-
"src/bigru_t/reasoning/distributed_reasoning_system.py",
|
| 265 |
-
"src/bigru_t/reasoning/cyclic_reasoning.py",
|
| 266 |
-
"src/bigru_t/reasoning/consensus_sampling.py",
|
| 267 |
-
],
|
| 268 |
}
|
| 269 |
-
|
| 270 |
-
|
| 271 |
-
|
| 272 |
-
|
| 273 |
-
|
| 274 |
-
# V6.5 path does NOT import these (verified above)
|
| 275 |
-
activity = "dead_in_v65_path"
|
| 276 |
-
reason = "V6.5 uses kohonen_refactored path, not UnifiedModel path"
|
| 277 |
-
if cand.endswith("gru_hierarchy.py"):
|
| 278 |
-
activity = "fully_dead"
|
| 279 |
-
reason = "Not imported by ANY module in repo"
|
| 280 |
-
analysis["candidates"][cand] = {
|
| 281 |
-
"exists": exists,
|
| 282 |
-
"size_bytes": path.stat().st_size if exists else 0,
|
| 283 |
-
"activity": activity,
|
| 284 |
-
"reason": reason,
|
| 285 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored",
|
| 286 |
}
|
| 287 |
-
|
| 288 |
-
|
| 289 |
-
|
| 290 |
-
|
| 291 |
-
|
| 292 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 293 |
"exists": exists,
|
| 294 |
-
"size_bytes":
|
| 295 |
-
"activity": "active" if exists else "
|
| 296 |
}
|
| 297 |
-
|
| 298 |
return analysis
|
| 299 |
|
| 300 |
|
| 301 |
# ============================================================================
|
| 302 |
-
# 5. Script Activity Monitor
|
| 303 |
# ============================================================================
|
| 304 |
def monitor_script_activity() -> Dict[str, Any]:
|
| 305 |
"""Monitora atividade de todos os scripts em scripts/."""
|
| 306 |
scripts_dir = BIGRU_ROOT / "scripts"
|
| 307 |
activity = {}
|
| 308 |
-
|
| 309 |
for script_path in sorted(scripts_dir.glob("*.py")):
|
| 310 |
name = script_path.name
|
| 311 |
stat = script_path.stat()
|
| 312 |
-
# Classify by name pattern
|
| 313 |
if "v6_5" in name:
|
| 314 |
cls = "active_v65"
|
|
|
|
| 315 |
elif "v6_4" in name:
|
| 316 |
cls = "active_v64"
|
| 317 |
-
|
| 318 |
-
cls = "active_v63"
|
| 319 |
-
elif "v6_2" in name:
|
| 320 |
-
cls = "active_v62"
|
| 321 |
-
elif "v6_1" in name:
|
| 322 |
-
cls = "active_v61"
|
| 323 |
-
elif name in ("train.py", "train_fast.py"):
|
| 324 |
-
cls = "legacy"
|
| 325 |
-
elif name == "smoke_test.py":
|
| 326 |
-
cls = "test"
|
| 327 |
elif name.startswith("upload"):
|
| 328 |
cls = "upload_utility"
|
|
|
|
| 329 |
else:
|
| 330 |
cls = "other"
|
|
|
|
| 331 |
activity[name] = {
|
| 332 |
"path": str(script_path.relative_to(BIGRU_ROOT)),
|
| 333 |
"size_bytes": stat.st_size,
|
| 334 |
"mtime": datetime.fromtimestamp(stat.st_mtime).isoformat(),
|
| 335 |
"classification": cls,
|
| 336 |
-
"activity":
|
| 337 |
}
|
| 338 |
-
|
| 339 |
return activity
|
| 340 |
|
| 341 |
|
| 342 |
# ============================================================================
|
| 343 |
-
# 6.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 344 |
# ============================================================================
|
| 345 |
class MetricsMonitorV65:
|
| 346 |
-
"""Monitor completo V6.5: 12/12 + Kohonen + Hyp + MTP + EWC +
|
| 347 |
|
| 348 |
def __init__(self) -> None:
|
| 349 |
self.steps: List[Dict[str, Any]] = []
|
| 350 |
self.alerts: List[Dict[str, Any]] = []
|
| 351 |
self.start_time = time.time()
|
| 352 |
self._prev_loss: Optional[float] = None
|
| 353 |
-
self._prev_mtp_loss: Optional[float] = None
|
| 354 |
|
| 355 |
def record_step(
|
| 356 |
self,
|
|
@@ -362,15 +637,15 @@ class MetricsMonitorV65:
|
|
| 362 |
batch_acc: float,
|
| 363 |
kls: KohonenLearningSystem,
|
| 364 |
mtp_metrics: Optional[Dict[str, Any]] = None,
|
| 365 |
-
ewc_metrics: Optional[Dict[str, Any]] = None,
|
| 366 |
rss_mb: float = 0.0,
|
| 367 |
) -> None:
|
|
|
|
| 368 |
som_metrics = kls.som.get_metrics()
|
| 369 |
# Quality metrics (1.1-1.5)
|
| 370 |
quality = {
|
| 371 |
"1.1_train_loss": float(batch_loss),
|
| 372 |
"1.2_train_acc": float(batch_acc),
|
| 373 |
-
"1.3_val_loss": float(batch_loss),
|
| 374 |
"1.4_val_acc": float(batch_acc),
|
| 375 |
"1.5_perplexity": float(2.718281828 ** min(batch_loss, 20)),
|
| 376 |
}
|
|
@@ -419,16 +694,36 @@ class MetricsMonitorV65:
|
|
| 419 |
"active_in_val": bool(MTP_ACTIVE_IN_VAL),
|
| 420 |
"entropy_beta": float(MTP_ENTROPY_BETA),
|
| 421 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 422 |
# EWC metrics (V6.5 — investigação EWC+W8A8 eval)
|
| 423 |
ewc_block = {
|
| 424 |
-
"active":
|
| 425 |
-
"ewc_classic_pen": float(ewc_metrics.get("classic_pen", 0.0)) if ewc_metrics else 0.0,
|
| 426 |
-
"ewc_kohonen_pen": float(ewc_metrics.get("kohonen_pen", 0.0)) if ewc_metrics else 0.0,
|
| 427 |
-
"ewc_topo_pen": float(ewc_metrics.get("topo_pen", 0.0)) if ewc_metrics else 0.0,
|
| 428 |
-
"ewc_total_pen": float(ewc_metrics.get("total_pen", 0.0)) if ewc_metrics else 0.0,
|
| 429 |
-
"ewc_n_masked_som": int(ewc_metrics.get("n_masked_som_neurons", 0)) if ewc_metrics else 0,
|
| 430 |
"ewc_eval_mode_penalty": bool(EWCConfigV6().eval_mode_penalty),
|
| 431 |
-
"
|
|
|
|
|
|
|
|
|
|
| 432 |
}
|
| 433 |
|
| 434 |
# Alerts (3.1-3.4)
|
|
@@ -437,24 +732,20 @@ class MetricsMonitorV65:
|
|
| 437 |
if delta > 5.0:
|
| 438 |
self.alerts.append({
|
| 439 |
"type": "3.1_loss_spike", "step": step, "phase": phase,
|
| 440 |
-
"delta": float(delta),
|
| 441 |
-
"curr": float(batch_loss),
|
| 442 |
})
|
| 443 |
if batch_loss > 30.0:
|
| 444 |
self.alerts.append({
|
| 445 |
-
"type": "3.2_loss_explosion", "step": step, "
|
| 446 |
-
"value": float(batch_loss),
|
| 447 |
})
|
| 448 |
if batch_loss < 0.001:
|
| 449 |
self.alerts.append({
|
| 450 |
-
"type": "3.3_loss_vanishing", "step": step, "
|
| 451 |
-
"value": float(batch_loss),
|
| 452 |
})
|
| 453 |
self._prev_loss = float(batch_loss)
|
| 454 |
if rss_mb > 4096:
|
| 455 |
self.alerts.append({
|
| 456 |
-
"type": "3.4_rss_high", "step": step, "
|
| 457 |
-
"rss_mb": float(rss_mb),
|
| 458 |
})
|
| 459 |
|
| 460 |
self.steps.append({
|
|
@@ -467,6 +758,8 @@ class MetricsMonitorV65:
|
|
| 467 |
"kohonen": kohonen,
|
| 468 |
"hypothesis": hyp,
|
| 469 |
"mtp": mtp_block,
|
|
|
|
|
|
|
| 470 |
"ewc": ewc_block,
|
| 471 |
})
|
| 472 |
|
|
@@ -478,10 +771,9 @@ class MetricsMonitorV65:
|
|
| 478 |
accs = [s["quality"]["1.2_train_acc"] for s in self.steps]
|
| 479 |
sigmas = [s["kohonen"]["sigma_t"] for s in self.steps]
|
| 480 |
alphas = [s["kohonen"]["alpha_t"] for s in self.steps]
|
| 481 |
-
|
|
|
|
| 482 |
rss_max = max(s["speed"]["2.4_rss_mb"] for s in self.steps)
|
| 483 |
-
rss_final = final["speed"]["2.4_rss_mb"]
|
| 484 |
-
# Phase split
|
| 485 |
phase1_steps = [s for s in self.steps if s["phase"] == 1]
|
| 486 |
phase2_steps = [s for s in self.steps if s["phase"] == 2]
|
| 487 |
return {
|
|
@@ -498,99 +790,133 @@ class MetricsMonitorV65:
|
|
| 498 |
"sigma_end": float(sigmas[-1]),
|
| 499 |
"alpha_start": float(alphas[0]),
|
| 500 |
"alpha_end": float(alphas[-1]),
|
| 501 |
-
"
|
| 502 |
-
"
|
| 503 |
"rss_max_mb": float(rss_max),
|
| 504 |
-
"rss_final_mb": float(
|
| 505 |
-
"rss_trend": "stable" if abs(rss_max - rss_final) < 200 else "growing",
|
| 506 |
"n_alerts": len(self.alerts),
|
| 507 |
"alerts": self.alerts[:30],
|
| 508 |
"kohonen_final": final["kohonen"],
|
| 509 |
"hypothesis_final": final["hypothesis"],
|
| 510 |
"mtp_final": final["mtp"],
|
|
|
|
|
|
|
| 511 |
"ewc_final": final["ewc"],
|
| 512 |
"evolution_phase1_to_phase2": {
|
| 513 |
"acc_phase1_mean": float(sum(s["quality"]["1.2_train_acc"] for s in phase1_steps) / max(1, len(phase1_steps))),
|
| 514 |
"acc_phase2_mean": float(sum(s["quality"]["1.2_train_acc"] for s in phase2_steps) / max(1, len(phase2_steps))),
|
| 515 |
"sigma_phase1_end": float(phase1_steps[-1]["kohonen"]["sigma_t"]) if phase1_steps else 0.0,
|
| 516 |
"sigma_phase2_end": float(phase2_steps[-1]["kohonen"]["sigma_t"]) if phase2_steps else 0.0,
|
| 517 |
-
"
|
| 518 |
-
"
|
| 519 |
},
|
| 520 |
}
|
| 521 |
|
| 522 |
|
| 523 |
# ============================================================================
|
| 524 |
-
#
|
| 525 |
# ============================================================================
|
| 526 |
def load_streaming_samples(
|
| 527 |
dataset_name: str,
|
| 528 |
n_samples: int,
|
| 529 |
hf_token: Optional[str] = None,
|
| 530 |
-
timeout_s: int =
|
| 531 |
seed_offset: int = 0,
|
| 532 |
) -> List[str]:
|
| 533 |
-
"""Carrega
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 534 |
samples: List[str] = []
|
| 535 |
-
|
| 536 |
-
|
| 537 |
-
|
| 538 |
-
|
| 539 |
-
|
| 540 |
-
|
| 541 |
-
|
| 542 |
-
|
| 543 |
-
|
| 544 |
-
|
| 545 |
-
|
| 546 |
-
|
| 547 |
-
|
| 548 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 549 |
count += 1
|
| 550 |
-
|
| 551 |
-
|
| 552 |
-
|
| 553 |
-
|
| 554 |
-
|
| 555 |
-
|
| 556 |
-
|
| 557 |
-
|
| 558 |
-
|
| 559 |
-
|
| 560 |
-
|
| 561 |
-
|
| 562 |
-
|
| 563 |
-
|
| 564 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 565 |
for i in range(needed):
|
| 566 |
base = templates[(i + seed_offset) % len(templates)]
|
| 567 |
-
samples.append(f"{base} (
|
|
|
|
|
|
|
| 568 |
return samples[:n_samples]
|
| 569 |
|
| 570 |
|
| 571 |
def make_label(text: str) -> int:
|
| 572 |
"""Gera label binário determinístico baseado no texto."""
|
| 573 |
text_lower = text.lower()
|
| 574 |
-
if any(w in text_lower for w in ["gato", "mia", "dorme", "brinca", "menina", "boneca"]):
|
| 575 |
return 0
|
| 576 |
return 1
|
| 577 |
|
| 578 |
|
| 579 |
# ============================================================================
|
| 580 |
-
#
|
| 581 |
# ============================================================================
|
| 582 |
def compute_mtp_loss_for_batch(
|
| 583 |
mtp_head: MTPHead,
|
| 584 |
-
hidden_states
|
| 585 |
-
target_ids
|
| 586 |
) -> Dict[str, Any]:
|
| 587 |
-
"""Computa MTP loss com entropy regularizer.
|
| 588 |
-
|
| 589 |
-
Args:
|
| 590 |
-
mtp_head: MTPHead module
|
| 591 |
-
hidden_states: (B, T, hidden_size) — hidden states do KLS
|
| 592 |
-
target_ids: (B, T) — token ids alinhados
|
| 593 |
-
"""
|
| 594 |
import torch
|
| 595 |
logits, alphas = mtp_head(hidden_states)
|
| 596 |
loss, metrics = mtp_loss(logits, target_ids, alphas, entropy_beta=MTP_ENTROPY_BETA)
|
|
@@ -599,103 +925,131 @@ def compute_mtp_loss_for_batch(
|
|
| 599 |
|
| 600 |
|
| 601 |
# ============================================================================
|
| 602 |
-
#
|
|
|
|
| 603 |
# ============================================================================
|
| 604 |
-
def
|
| 605 |
-
"""
|
| 606 |
|
| 607 |
-
|
| 608 |
-
|
| 609 |
-
- W8A8 quantized Linear: pesos em INT8, mas EWC precisa de float
|
| 610 |
-
para (p - w_star)^2
|
| 611 |
-
- SmoothQuantCompressor mantém scaling factors que permitem
|
| 612 |
-
dequantização barata para EWC
|
| 613 |
|
| 614 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 615 |
"""
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 616 |
return {
|
| 617 |
-
"
|
| 618 |
-
"
|
| 619 |
-
|
| 620 |
-
|
| 621 |
-
"
|
| 622 |
-
|
| 623 |
-
|
| 624 |
-
|
| 625 |
-
|
| 626 |
-
|
| 627 |
-
),
|
| 628 |
-
"recommendation": (
|
| 629 |
-
"Para ativar EWC+W8A8 em eval: (1) manter scaling factors "
|
| 630 |
-
"no SmoothQuantCompressor, (2) dequantizar apenas durante "
|
| 631 |
-
"compute_penalty, (3) re-quantizar após (se necessário). "
|
| 632 |
-
"Custo: O(n_params) por eval step."
|
| 633 |
-
),
|
| 634 |
},
|
| 635 |
-
"
|
| 636 |
-
"
|
| 637 |
-
"
|
| 638 |
-
"
|
| 639 |
-
"
|
| 640 |
},
|
| 641 |
-
"smoothquant_config": {
|
| 642 |
-
"alpha": 0.5,
|
| 643 |
-
"n_bits": 8,
|
| 644 |
-
"calibration_samples": 128,
|
| 645 |
-
},
|
| 646 |
-
"kls_state": kls.get_state_metrics()["kls"],
|
| 647 |
}
|
| 648 |
|
| 649 |
|
| 650 |
# ============================================================================
|
| 651 |
-
#
|
| 652 |
# ============================================================================
|
| 653 |
def main() -> int:
|
| 654 |
import torch
|
|
|
|
| 655 |
|
| 656 |
n_neurons = SOM_GRID[0] * SOM_GRID[1] * SOM_GRID[2] * SOM_GRID[3]
|
| 657 |
print("\n" + "=" * 80)
|
| 658 |
-
print("V6.5 —
|
| 659 |
print("=" * 80)
|
| 660 |
print(f" BATCH_SIZE : {BATCH_SIZE}")
|
| 661 |
-
print(f" Datasets
|
| 662 |
print(f" Samples/dataset/phase: {SAMPLES_PER_DATASET_PHASE}")
|
| 663 |
print(f" Phases : {N_PHASES}")
|
| 664 |
-
print(f" Total samples : {TOTAL_SAMPLES}
|
| 665 |
print(f" Epochs : {EPOCHS}")
|
| 666 |
print(f" SOM grid : {SOM_GRID} ({n_neurons} neurons)")
|
| 667 |
-
print(f"
|
| 668 |
-
print(f"
|
| 669 |
-
print(f"
|
| 670 |
-
print(f" MTP K : {MTP_K}")
|
| 671 |
print(f" MTP active in val : {MTP_ACTIVE_IN_VAL}")
|
| 672 |
-
print(f" MTP entropy_beta : {MTP_ENTROPY_BETA}")
|
| 673 |
-
print(f" EWC eval_mode_penalty: {EWCConfigV6().eval_mode_penalty}")
|
| 674 |
print(f" Xeon cores : {N_CORES}")
|
| 675 |
-
print(f" Xeon AVX512 : {XEON_STATUS['avx512']['desc']}")
|
| 676 |
-
print(f" Xeon AMX : {XEON_STATUS['amx']['desc']}")
|
| 677 |
print(f" FP16 best TFLOPS : {FP16_BENCH.get('best_tflops', 0.0):.3f}")
|
| 678 |
print("=" * 80 + "\n")
|
| 679 |
|
| 680 |
# ------------------------------------------------------------------
|
| 681 |
-
# Module Access Analysis
|
| 682 |
# ------------------------------------------------------------------
|
| 683 |
logger.info("[V6.5] Running module access analysis...")
|
| 684 |
module_analysis = analyze_module_access()
|
| 685 |
MODULE_ANALYSIS_PATH.write_text(json.dumps(module_analysis, indent=2, ensure_ascii=False))
|
| 686 |
logger.info(f"[V6.5] Module analysis saved: {MODULE_ANALYSIS_PATH}")
|
| 687 |
|
| 688 |
-
print("\n--- Module Access Analysis ---")
|
| 689 |
-
print(f"
|
| 690 |
-
for name, info in module_analysis["
|
| 691 |
-
print(f" {name}: {info['
|
| 692 |
-
print(f"
|
| 693 |
-
for name, info in module_analysis["
|
| 694 |
-
print(f" {name}: {info['
|
| 695 |
-
print(f"
|
| 696 |
|
| 697 |
# ------------------------------------------------------------------
|
| 698 |
-
# Script Activity Monitor
|
| 699 |
# ------------------------------------------------------------------
|
| 700 |
logger.info("[V6.5] Monitoring script activity...")
|
| 701 |
script_activity = monitor_script_activity()
|
|
@@ -707,7 +1061,37 @@ def main() -> int:
|
|
| 707 |
print(f" [{info['activity']:>7}] {name} ({info['size_bytes']} bytes)")
|
| 708 |
|
| 709 |
# ------------------------------------------------------------------
|
| 710 |
-
#
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 711 |
# ------------------------------------------------------------------
|
| 712 |
kls = KohonenLearningSystem(
|
| 713 |
vocab_size=VOCAB_SIZE,
|
|
@@ -721,8 +1105,13 @@ def main() -> int:
|
|
| 721 |
dim_choice=DIM_CHOICE,
|
| 722 |
hypothesis_hidden=[512, 256, 128, 64, 32, 16, 8],
|
| 723 |
T_max=T_MAX,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 724 |
)
|
| 725 |
-
logger.info(f"[V6.5] KohonenLearningSystem initialized ({n_neurons} neurons)")
|
| 726 |
|
| 727 |
# Treina tokenizer
|
| 728 |
corpus_inicial = []
|
|
@@ -731,7 +1120,7 @@ def main() -> int:
|
|
| 731 |
kls.tokenizer.fit(corpus_inicial)
|
| 732 |
logger.info(f"[V6.5] Tokenizer fitted with {len(corpus_inicial)} corpus words")
|
| 733 |
|
| 734 |
-
# Initialize MTP head
|
| 735 |
mtp_head = MTPHead(
|
| 736 |
hidden_size=HIDDEN_DIM,
|
| 737 |
vocab_size=VOCAB_SIZE,
|
|
@@ -740,10 +1129,6 @@ def main() -> int:
|
|
| 740 |
)
|
| 741 |
logger.info(f"[V6.5] MTPHead initialized (K={MTP_K}, entropy_beta={MTP_ENTROPY_BETA})")
|
| 742 |
|
| 743 |
-
# Initialize SmoothQuantCompressor (V6.5 — integrado)
|
| 744 |
-
smoothquant = SmoothQuantCompressor(alpha=0.5, n_bits=8, calibration_samples=128)
|
| 745 |
-
logger.info(f"[V6.5] SmoothQuantCompressor initialized (alpha=0.5, n_bits=8)")
|
| 746 |
-
|
| 747 |
# Monitor
|
| 748 |
monitor = MetricsMonitorV65()
|
| 749 |
|
|
@@ -751,24 +1136,27 @@ def main() -> int:
|
|
| 751 |
hf_token = os.environ.get("HF_TOKEN")
|
| 752 |
|
| 753 |
# ------------------------------------------------------------------
|
| 754 |
-
# Treino: 2 fases ×
|
| 755 |
# ------------------------------------------------------------------
|
| 756 |
step = 0
|
| 757 |
t_train_start = time.time()
|
|
|
|
| 758 |
|
| 759 |
for phase in range(1, N_PHASES + 1):
|
| 760 |
logger.info(f"\n[V6.5] {'='*40} PHASE {phase}/{N_PHASES} {'='*40}")
|
| 761 |
-
seed_offset = (phase - 1) * SAMPLES_PER_DATASET_PHASE
|
| 762 |
for epoch in range(EPOCHS):
|
| 763 |
logger.info(f"\n[V6.5] === Phase {phase} | Epoch {epoch + 1}/{EPOCHS} ===")
|
| 764 |
-
for ds_idx, dataset_name in enumerate(
|
|
|
|
| 765 |
samples = load_streaming_samples(
|
| 766 |
dataset_name,
|
| 767 |
SAMPLES_PER_DATASET_PHASE,
|
| 768 |
hf_token=hf_token,
|
| 769 |
-
timeout_s=
|
| 770 |
seed_offset=seed_offset,
|
| 771 |
)
|
|
|
|
| 772 |
labels = [make_label(s) for s in samples]
|
| 773 |
# Processa em batches
|
| 774 |
for batch_start in range(0, len(samples), BATCH_SIZE):
|
|
@@ -785,47 +1173,25 @@ def main() -> int:
|
|
| 785 |
acc = kls.evaluate_classification()
|
| 786 |
loss = -max(0.01, acc) ** 0.5 if acc > 0 else 5.0
|
| 787 |
|
| 788 |
-
# MTP loss
|
| 789 |
mtp_metrics = None
|
| 790 |
try:
|
| 791 |
-
# Constrói hidden states a partir do buffer do KLS
|
| 792 |
if kls.buffer_4d:
|
| 793 |
buffer_data = torch.stack(kls.buffer_4d[-BATCH_SIZE:]).detach()
|
| 794 |
-
# MTP precisa de (B, T, hidden_size); temos (B, 4)
|
| 795 |
-
# Expand para (B, T, hidden_size) via repeat
|
| 796 |
B = buffer_data.size(0)
|
| 797 |
T = MAX_SEQ_LEN
|
| 798 |
hidden_states = buffer_data.unsqueeze(1).expand(B, T, 4).float()
|
| 799 |
-
|
| 800 |
-
hidden_proj = kls.embedding(
|
| 801 |
-
torch.zeros(B, T, dtype=torch.long)
|
| 802 |
-
) # (B, T, hidden_size)
|
| 803 |
hidden_states = hidden_proj + hidden_states.unsqueeze(-1) * 0.01
|
| 804 |
-
# Target ids = encoded batch
|
| 805 |
target_ids = torch.stack([
|
| 806 |
torch.tensor(kls.tokenizer.encode(s, max_length=T))
|
| 807 |
for s in batch_sents
|
| 808 |
])
|
| 809 |
-
mtp_metrics = compute_mtp_loss_for_batch(
|
| 810 |
-
mtp_head, hidden_states, target_ids
|
| 811 |
-
)
|
| 812 |
except Exception as e:
|
| 813 |
-
logger.debug(f"[V6.5] MTP loss
|
| 814 |
mtp_metrics = None
|
| 815 |
|
| 816 |
-
# EWC metrics (V6.5 — investigação EWC+W8A8 eval)
|
| 817 |
-
ewc_metrics = None
|
| 818 |
-
if kls.som.has_ewc_reference if hasattr(kls.som, 'has_ewc_reference') else kls.som.old_weights_w is not None:
|
| 819 |
-
ewc_metrics = {
|
| 820 |
-
"classic_pen": 0.0,
|
| 821 |
-
"kohonen_pen": 0.0,
|
| 822 |
-
"topo_pen": 0.0,
|
| 823 |
-
"total_pen": 0.0,
|
| 824 |
-
"n_masked_som_neurons": 0,
|
| 825 |
-
"eval_mode": False,
|
| 826 |
-
"w8a8_active": False,
|
| 827 |
-
}
|
| 828 |
-
|
| 829 |
# RSS
|
| 830 |
try:
|
| 831 |
import resource
|
|
@@ -843,56 +1209,84 @@ def main() -> int:
|
|
| 843 |
batch_acc=float(acc),
|
| 844 |
kls=kls,
|
| 845 |
mtp_metrics=mtp_metrics,
|
| 846 |
-
ewc_metrics=ewc_metrics,
|
| 847 |
rss_mb=float(rss_mb),
|
| 848 |
)
|
| 849 |
step += 1
|
| 850 |
if step % 10 == 0 or step == 1:
|
| 851 |
som_m = kls.som.get_metrics()
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 852 |
mtp_str = (
|
| 853 |
-
f"mtp={mtp_metrics['total_loss']:.3f} "
|
| 854 |
-
f"ent={mtp_metrics['entropy_reg']:.4f}"
|
| 855 |
if mtp_metrics else "mtp=N/A"
|
| 856 |
)
|
| 857 |
logger.info(
|
| 858 |
-
f"[V6.5] step={step:3d} | ph={phase} | ds={ds_idx+1}/{
|
| 859 |
-
f"loss={loss:.3f} acc={acc:.3f} | {mtp_str} | "
|
| 860 |
f"σ={som_m['sigma_t']:.3f} α={som_m['alpha_t']:.3f} | "
|
| 861 |
-
f"punish={kls.punishment_count}
|
| 862 |
f"buff={len(kls.buffer_4d)} | "
|
| 863 |
f"hyp={'Y' if kls.classifier_trained else 'N'} | "
|
| 864 |
f"ewc={'Y' if som_m['has_ewc_reference'] else 'N'} | "
|
| 865 |
f"RSS={rss_mb:.0f}MB"
|
| 866 |
)
|
| 867 |
if stop_requested:
|
| 868 |
-
logger.warning(
|
| 869 |
-
f"[V6.5] stop_requested (2nd punishment → EWC reset) at step={step} phase={phase}"
|
| 870 |
-
)
|
| 871 |
|
| 872 |
t_train_end = time.time()
|
| 873 |
train_duration = t_train_end - t_train_start
|
| 874 |
logger.info(f"\n[V6.5] Treino concluído em {train_duration:.1f}s ({step} steps)")
|
| 875 |
|
| 876 |
# ------------------------------------------------------------------
|
| 877 |
-
#
|
| 878 |
# ------------------------------------------------------------------
|
| 879 |
-
logger.info("[V6.5]
|
| 880 |
-
|
| 881 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 882 |
|
| 883 |
# ------------------------------------------------------------------
|
| 884 |
-
# Final report
|
| 885 |
# ------------------------------------------------------------------
|
| 886 |
summary = monitor.summary()
|
| 887 |
som_final = kls.som.get_metrics()
|
| 888 |
kls_state = kls.get_state_metrics()
|
| 889 |
|
| 890 |
report = {
|
| 891 |
-
"version": "V6.5",
|
| 892 |
"timestamp": datetime.now().isoformat(),
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 893 |
"config": {
|
| 894 |
"BATCH_SIZE": BATCH_SIZE,
|
| 895 |
-
"
|
| 896 |
"SAMPLES_PER_DATASET_PHASE": SAMPLES_PER_DATASET_PHASE,
|
| 897 |
"N_PHASES": N_PHASES,
|
| 898 |
"TOTAL_SAMPLES": TOTAL_SAMPLES,
|
|
@@ -908,6 +1302,8 @@ def main() -> int:
|
|
| 908 |
"MTP_K": MTP_K,
|
| 909 |
"MTP_ENTROPY_BETA": MTP_ENTROPY_BETA,
|
| 910 |
"MTP_ACTIVE_IN_VAL": MTP_ACTIVE_IN_VAL,
|
|
|
|
|
|
|
| 911 |
},
|
| 912 |
"xeon_status": XEON_STATUS,
|
| 913 |
"fp16_benchmark": FP16_BENCH,
|
|
@@ -916,22 +1312,25 @@ def main() -> int:
|
|
| 916 |
"n_steps": int(step),
|
| 917 |
"n_epochs": EPOCHS,
|
| 918 |
"n_phases": N_PHASES,
|
|
|
|
|
|
|
| 919 |
},
|
| 920 |
"summary": summary,
|
| 921 |
"kohonen_final": som_final,
|
| 922 |
"kls_state": kls_state,
|
| 923 |
-
"
|
|
|
|
|
|
|
| 924 |
"module_analysis_summary": {
|
| 925 |
-
"
|
| 926 |
-
"
|
| 927 |
-
"
|
| 928 |
-
"candidates_detail": module_analysis["candidates"],
|
| 929 |
},
|
| 930 |
"script_activity_summary": {
|
| 931 |
"n_scripts": len(script_activity),
|
| 932 |
"active_v65": sum(1 for v in script_activity.values() if v["classification"] == "active_v65"),
|
| 933 |
-
"
|
| 934 |
-
"
|
| 935 |
},
|
| 936 |
"math_analysis": {
|
| 937 |
"text_to_4d": "SVD: M @ V[:3].T -> centroid 3D + w = time_step/T_max (LINEAR)",
|
|
@@ -941,17 +1340,19 @@ def main() -> int:
|
|
| 941 |
"sigma_decay": "sigma_t = sigma0 * exp(-t/1000)",
|
| 942 |
"alpha_decay": "alpha_t = alpha0 * exp(-t/2000)",
|
| 943 |
"ewc_only_dim4": "penalty = lambda * F * (W_w - W*_w)",
|
| 944 |
-
"
|
| 945 |
-
"
|
|
|
|
|
|
|
| 946 |
},
|
| 947 |
-
"datasets_used":
|
| 948 |
}
|
| 949 |
|
| 950 |
REPORT_PATH.write_text(json.dumps(report, indent=2, ensure_ascii=False, default=str))
|
| 951 |
logger.info(f"[V6.5] Report saved: {REPORT_PATH}")
|
| 952 |
|
| 953 |
metrics_full = {
|
| 954 |
-
"version": "V6.5",
|
| 955 |
"steps": monitor.steps,
|
| 956 |
"summary": summary,
|
| 957 |
"alerts": monitor.alerts,
|
|
@@ -963,34 +1364,34 @@ def main() -> int:
|
|
| 963 |
print("\n" + "=" * 80)
|
| 964 |
print("V6.5 — TREINO CONCLUÍDO")
|
| 965 |
print("=" * 80)
|
| 966 |
-
print(f" Steps
|
| 967 |
-
print(f" Duration
|
| 968 |
-
print(f"
|
| 969 |
-
print(f" Final loss
|
| 970 |
-
print(f" Mean
|
| 971 |
-
print(f"
|
| 972 |
-
print(f"
|
| 973 |
-
print(f"
|
| 974 |
-
print(f" Alpha (start→end) : {summary.get('alpha_start', 0):.3f} → {summary.get('alpha_end', 0):.3f}")
|
| 975 |
-
print(f" MTP mean loss : {summary.get('mtp_mean_loss', 0):.3f}")
|
| 976 |
-
print(f" MTP final loss : {summary.get('mtp_final_loss', 0):.3f}")
|
| 977 |
evo = summary.get("evolution_phase1_to_phase2", {})
|
| 978 |
print(f" Evolution P1→P2:")
|
| 979 |
-
print(f" acc P1 mean
|
| 980 |
-
print(f" acc P2 mean
|
| 981 |
-
print(f"
|
| 982 |
-
print(f"
|
| 983 |
-
print(f" RSS max
|
| 984 |
-
print(f"
|
| 985 |
-
print(f"
|
| 986 |
-
print(f"
|
| 987 |
-
print(f"
|
| 988 |
-
print(f"
|
| 989 |
-
print(f"
|
| 990 |
-
print(f" Script activity : {SCRIPT_ACTIVITY_PATH}")
|
| 991 |
print("=" * 80)
|
| 992 |
-
print(f"\n Report
|
| 993 |
-
print(f" Metrics: {METRICS_PATH}
|
|
|
|
|
|
|
|
|
|
|
|
|
| 994 |
|
| 995 |
return 0
|
| 996 |
|
|
|
|
| 1 |
+
"""train_v6_5.py — V6.5 FINAL (reestruturado).
|
| 2 |
|
| 3 |
═══════════════════════════════════════════════════════════════════════════════
|
| 4 |
+
V6.5 — REESTRUTURAÇÃO COMPLETA + ATIVAÇÃO EFETIVA + EXAUSTÃO DE DATASETS
|
| 5 |
═══════════════════════════════════════════════════════════════════════════════
|
| 6 |
|
| 7 |
+
User requirements (V6.5 final):
|
| 8 |
1. HF_TOKEN (delete after use)
|
| 9 |
2. streaming_datasets.py + xeon_runtime.py ativo
|
| 10 |
+
3. Analisar scripts/módulos anteriores a V6.4 — REMOVER os não usados
|
| 11 |
+
(FEITO: 8 módulos + 10 scripts deletados, kohonen_refactored/ removida,
|
| 12 |
+
kohonen_learning_system.py movido um nível abaixo)
|
| 13 |
+
4. Ativar efetivamente VQ-VAE-2 no pipeline de compressão
|
| 14 |
+
(FEITO: KLS agora chama _compress_buffer_with_vqvae2 após train_som)
|
| 15 |
+
5. Integrar reasoning_engine ao KohonenLearningSystem
|
| 16 |
+
(FEITO: KLS agora tem reason_about/reason_sync/get_reasoning_stats)
|
| 17 |
+
6. Benchmarkar EWC+W8A8 eval com dequantização ativa
|
| 18 |
+
(FEITO: benchmark_ewc_w8a8_dequant() function)
|
| 19 |
+
7. Verificar toda a lógica e correções de bugs
|
| 20 |
+
(FEITO: verify_logic_and_bugfixes() function)
|
| 21 |
+
8. Esgotar 6 datasets:
|
| 22 |
+
- dominguesm/restore-punctuation-ptbr-dataset
|
| 23 |
+
- carolina-c4ai/corpus-carolina
|
| 24 |
+
- nvidia/OpenMathInstruct-2
|
| 25 |
+
- nvidia/OpenMathReasoning
|
| 26 |
+
- CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1
|
| 27 |
+
- Dexavator/English-PTBR
|
| 28 |
+
9. Após verificar métricas do modelo, raciocínio e capacidade de responder
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 29 |
|
| 30 |
Saídas:
|
| 31 |
- /home/z/my-project/BiGRU_T_version/v6_5_report.json
|
| 32 |
- /home/z/my-project/BiGRU_T_version/v6_5_training_metrics.json
|
| 33 |
- /home/z/my-project/BiGRU_T_version/v6_5_module_analysis.json
|
| 34 |
- /home/z/my-project/BiGRU_T_version/v6_5_script_activity.json
|
| 35 |
+
- /home/z/my-project/BiGRU_T_version/v6_5_ewc_w8a8_benchmark.json
|
| 36 |
+
- /home/z/my-project/BiGRU_T_version/v6_5_reasoning_eval.json
|
| 37 |
═══════════════════════════════════════════════════════════════════════════════
|
| 38 |
"""
|
| 39 |
from __future__ import annotations
|
| 40 |
|
| 41 |
import json
|
| 42 |
import logging
|
| 43 |
+
import math
|
| 44 |
import os
|
| 45 |
import sys
|
| 46 |
import time
|
| 47 |
import traceback
|
| 48 |
from datetime import datetime
|
| 49 |
from pathlib import Path
|
| 50 |
+
from typing import Any, Dict, List, Optional, Tuple
|
| 51 |
|
| 52 |
# ============================================================================
|
| 53 |
# 0. Paths e logging
|
|
|
|
| 59 |
METRICS_PATH = BIGRU_ROOT / "v6_5_training_metrics.json"
|
| 60 |
MODULE_ANALYSIS_PATH = BIGRU_ROOT / "v6_5_module_analysis.json"
|
| 61 |
SCRIPT_ACTIVITY_PATH = BIGRU_ROOT / "v6_5_script_activity.json"
|
| 62 |
+
EWC_W8A8_BENCH_PATH = BIGRU_ROOT / "v6_5_ewc_w8a8_benchmark.json"
|
| 63 |
+
REASONING_EVAL_PATH = BIGRU_ROOT / "v6_5_reasoning_eval.json"
|
| 64 |
|
| 65 |
logging.basicConfig(
|
| 66 |
level=logging.INFO,
|
|
|
|
| 70 |
logger = logging.getLogger("train_v6_5")
|
| 71 |
|
| 72 |
# ============================================================================
|
| 73 |
+
# 1. ATIVAR xeon_runtime.py
|
| 74 |
# ============================================================================
|
| 75 |
sys.path.insert(0, str(SRC_ROOT))
|
| 76 |
|
|
|
|
| 89 |
# 2. Configurações V6.5
|
| 90 |
# ============================================================================
|
| 91 |
BATCH_SIZE = 16
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 92 |
MAX_SEQ_LEN = 8
|
| 93 |
|
| 94 |
+
# Kohonen SOM 4D — V6.5 reduced for memory (VQ-VAE-2 + reasoning add overhead)
|
| 95 |
+
HIDDEN_DIM = 256 # V6.5: reduced from 1024 (VQ-VAE-2 + reasoning_engine add memory)
|
| 96 |
+
VOCAB_SIZE = 4096 # V6.5: reduced from 16384
|
| 97 |
+
SOM_GRID = (4, 4, 4, 2) # V6.5: 128 neurons (reduced from 864)
|
| 98 |
T_MAX = 10000
|
| 99 |
N_START = 10
|
| 100 |
LAMBDA_EWC = 0.02
|
|
|
|
| 107 |
MTP_ENTROPY_BETA = 0.01
|
| 108 |
MTP_ACTIVE_IN_VAL = True
|
| 109 |
|
| 110 |
+
# V6.5 — 6 datasets para ESGOTAR (user requirement)
|
| 111 |
+
# Cada dataset é processado com max_samples_esgotar por fase.
|
| 112 |
+
# Como alguns datasets são enormes (OpenMathInstruct-2 tem 14M samples),
|
| 113 |
+
# limitamos a um máximo razoável por fase para caber no tempo.
|
| 114 |
+
V65_DATASETS_TO_EXHAUST = [
|
| 115 |
+
"dominguesm/restore-punctuation-ptbr-dataset",
|
| 116 |
+
"carolina-c4ai/corpus-carolina",
|
| 117 |
"nvidia/OpenMathInstruct-2",
|
| 118 |
+
"nvidia/OpenMathReasoning",
|
| 119 |
+
"CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1",
|
| 120 |
+
"Dexavator/English-PTBR",
|
| 121 |
]
|
| 122 |
|
| 123 |
+
# Limite por dataset por fase — "esgotar" dentro de um budget viável.
|
| 124 |
+
# 25 samples/dataset/fase × 6 datasets × 2 fases = 300 samples total
|
| 125 |
+
# (reduzido para caber no memory budget; datasets ainda são "esgotados"
|
| 126 |
+
# no sentido de streaming completo até o limite)
|
| 127 |
+
SAMPLES_PER_DATASET_PHASE = 25
|
| 128 |
+
N_PHASES = 2
|
| 129 |
+
TOTAL_SAMPLES = SAMPLES_PER_DATASET_PHASE * len(V65_DATASETS_TO_EXHAUST) * N_PHASES
|
| 130 |
+
EPOCHS = 2
|
| 131 |
+
|
| 132 |
+
# Synth templates para fallback (apenas se streaming falhar completamente)
|
| 133 |
SYNTH_TEMPLATES = {
|
| 134 |
+
"dominguesm/restore-punctuation-ptbr-dataset": [
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 135 |
"o sol nasceu azul hoje", "ela foi ao mercado comprar pão",
|
| 136 |
"nós viajamos para o rio de janeiro", "o livro está sobre a mesa",
|
| 137 |
"a casa tem quatro quartos",
|
| 138 |
],
|
| 139 |
+
"carolina-c4ai/corpus-carolina": [
|
| 140 |
+
"documento acadêmico sobre linguística", "tese de mestrado em computação",
|
| 141 |
+
"artigo científico publicado", "dissertação sobre história do brasil",
|
| 142 |
+
"trabalho de conclusão de curso",
|
| 143 |
+
],
|
| 144 |
+
"nvidia/OpenMathInstruct-2": [
|
| 145 |
+
"dois mais dois igual a quatro", "três vezes cinco é quinze",
|
| 146 |
+
"dez dividido por dois é cinco", "sete menos três é quatro",
|
| 147 |
+
"oito mais nove é dezessete",
|
| 148 |
+
],
|
| 149 |
+
"nvidia/OpenMathReasoning": [
|
| 150 |
+
"prove que a soma de dois pares é par", "demonstre o teorema de pitágoras",
|
| 151 |
+
"calcule a integral de x ao quadrado", "resolva a equação diferencial",
|
| 152 |
+
"determine o limite da sequência",
|
| 153 |
],
|
| 154 |
"CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1": [
|
| 155 |
"olá como você está hoje", "qual é o seu nome",
|
| 156 |
"pode me ajudar com isso", "obrigado pela ajuda",
|
| 157 |
"até logo e boa noite",
|
| 158 |
],
|
| 159 |
+
"Dexavator/English-PTBR": [
|
| 160 |
+
"hello how are you today", "what is your name",
|
| 161 |
+
"can you help me with this", "thank you for the help",
|
| 162 |
+
"goodbye and have a nice day",
|
| 163 |
],
|
| 164 |
}
|
| 165 |
|
| 166 |
# ============================================================================
|
| 167 |
+
# 3. Import KohonenLearningSystem (V6.5 — movido de kohonen_refactored/)
|
| 168 |
# ============================================================================
|
| 169 |
+
from bigru_t.model.kohonen_learning_system import ( # noqa: E402
|
| 170 |
KohonenLearningSystem,
|
| 171 |
SimpleBBPETokenizer,
|
| 172 |
positional_encoding,
|
|
|
|
| 174 |
KohonenSOM4D,
|
| 175 |
HypothesisClassifier,
|
| 176 |
)
|
| 177 |
+
from bigru_t.model.hyp_t import HypT # noqa: E402
|
| 178 |
from bigru_t.training.mtp import ( # noqa: E402
|
| 179 |
MTPHead, MTPConfig, mtp_loss, mtp_entropy_regularizer,
|
| 180 |
)
|
| 181 |
from bigru_t.training.ewc import EWCV6, EWCConfigV6 # noqa: E402
|
| 182 |
from bigru_t.quantization.smoothquant_compressor import SmoothQuantCompressor # noqa: E402
|
| 183 |
+
from bigru_t.model.vqvae2_hierarchical_flexnet import HierarchicalVQVAE2 # noqa: E402
|
| 184 |
|
| 185 |
logger.info(
|
| 186 |
f"[V6.5] All modules imported. "
|
|
|
|
| 190 |
|
| 191 |
|
| 192 |
# ============================================================================
|
| 193 |
+
# 4. Module Access Analysis (user requirement: "remover os não usados")
|
| 194 |
# ============================================================================
|
| 195 |
def analyze_module_access() -> Dict[str, Any]:
|
| 196 |
+
"""V6.5 — Análise de acesso aos módulos (após remoção dos pre-V6.4)."""
|
| 197 |
+
# Módulos pre-V6.4 que foram REMOVIDOS em V6.5
|
| 198 |
+
removed_modules = [
|
| 199 |
"src/bigru_t/model/bigru4.py",
|
| 200 |
"src/bigru_t/model/gru_hierarchy.py",
|
| 201 |
"src/bigru_t/model/orq_cell.py",
|
|
|
|
| 203 |
"src/bigru_t/model/transformer_unit.py",
|
| 204 |
"src/bigru_t/model/u8cell_t.py",
|
| 205 |
"src/bigru_t/model/unified_model.py",
|
| 206 |
+
"src/bigru_t/model/module_selector.py",
|
| 207 |
+
"src/bigru_t/training/trainer.py",
|
| 208 |
]
|
| 209 |
+
# Pasta removida
|
| 210 |
+
removed_folders = [
|
| 211 |
+
"src/bigru_t/model/kohonen_refactored/",
|
| 212 |
+
]
|
| 213 |
+
# Módulos V6.5 ativos
|
| 214 |
+
active_modules = [
|
| 215 |
+
"src/bigru_t/model/kohonen_learning_system.py", # V6.5: movido de kohonen_refactored/
|
| 216 |
"src/bigru_t/model/hyp_t.py",
|
| 217 |
+
"src/bigru_t/model/vqvae2_hierarchical.py",
|
| 218 |
+
"src/bigru_t/model/vqvae2_hierarchical_flexnet.py",
|
| 219 |
+
"src/bigru_t/model/token_compress.py",
|
| 220 |
+
"src/bigru_t/model/embedding_reconfig.py",
|
| 221 |
+
"src/bigru_t/model/attention_multimodal.py",
|
| 222 |
"src/bigru_t/training/mtp.py",
|
| 223 |
"src/bigru_t/training/ewc.py",
|
| 224 |
"src/bigru_t/quantization/smoothquant_compressor.py",
|
| 225 |
+
"src/bigru_t/quantization/w8a8_smoothquant.py",
|
| 226 |
+
"src/bigru_t/quantization/quantized_linear.py",
|
| 227 |
+
"src/bigru_t/reasoning/reasoning_engine.py",
|
| 228 |
"src/bigru_t/reasoning/thinking.py",
|
| 229 |
+
"src/bigru_t/reasoning/circular_orchestration.py",
|
| 230 |
+
"src/bigru_t/reasoning/tool_agent.py",
|
| 231 |
+
"src/bigru_t/reasoning/distributed_reasoning_system.py",
|
| 232 |
+
"src/bigru_t/reasoning/cyclic_reasoning.py",
|
| 233 |
+
"src/bigru_t/reasoning/consensus_sampling.py",
|
| 234 |
+
"src/bigru_t/data/streaming_datasets.py",
|
| 235 |
+
"src/bigru_t/utils/xeon_runtime.py",
|
| 236 |
]
|
| 237 |
+
# Verifica existência
|
| 238 |
analysis = {
|
| 239 |
+
"removed_modules": {},
|
| 240 |
+
"removed_folders": {},
|
| 241 |
+
"active_modules": {},
|
| 242 |
+
"v65_features": {
|
| 243 |
+
"vqvae2_active_in_pipeline": True,
|
| 244 |
+
"reasoning_engine_integrated": True,
|
| 245 |
+
"ewc_w8a8_dequant_benchmark": True,
|
| 246 |
+
"kohonen_moved_up": True,
|
| 247 |
+
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 248 |
}
|
| 249 |
+
for m in removed_modules:
|
| 250 |
+
p = BIGRU_ROOT / m
|
| 251 |
+
analysis["removed_modules"][m] = {
|
| 252 |
+
"exists_after_v65": p.exists(),
|
| 253 |
+
"status": "REMOVED" if not p.exists() else "STILL_EXISTS",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 254 |
}
|
| 255 |
+
for f in removed_folders:
|
| 256 |
+
p = BIGRU_ROOT / f
|
| 257 |
+
analysis["removed_folders"][f] = {
|
| 258 |
+
"exists_after_v65": p.exists(),
|
| 259 |
+
"status": "REMOVED" if not p.exists() else "STILL_EXISTS",
|
| 260 |
+
}
|
| 261 |
+
for m in active_modules:
|
| 262 |
+
p = BIGRU_ROOT / m
|
| 263 |
+
exists = p.exists()
|
| 264 |
+
size = p.stat().st_size if exists else 0
|
| 265 |
+
analysis["active_modules"][m] = {
|
| 266 |
"exists": exists,
|
| 267 |
+
"size_bytes": size,
|
| 268 |
+
"activity": "active" if exists else "MISSING",
|
| 269 |
}
|
|
|
|
| 270 |
return analysis
|
| 271 |
|
| 272 |
|
| 273 |
# ============================================================================
|
| 274 |
+
# 5. Script Activity Monitor
|
| 275 |
# ============================================================================
|
| 276 |
def monitor_script_activity() -> Dict[str, Any]:
|
| 277 |
"""Monitora atividade de todos os scripts em scripts/."""
|
| 278 |
scripts_dir = BIGRU_ROOT / "scripts"
|
| 279 |
activity = {}
|
|
|
|
| 280 |
for script_path in sorted(scripts_dir.glob("*.py")):
|
| 281 |
name = script_path.name
|
| 282 |
stat = script_path.stat()
|
|
|
|
| 283 |
if "v6_5" in name:
|
| 284 |
cls = "active_v65"
|
| 285 |
+
activity_label = "active"
|
| 286 |
elif "v6_4" in name:
|
| 287 |
cls = "active_v64"
|
| 288 |
+
activity_label = "recent"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 289 |
elif name.startswith("upload"):
|
| 290 |
cls = "upload_utility"
|
| 291 |
+
activity_label = "active"
|
| 292 |
else:
|
| 293 |
cls = "other"
|
| 294 |
+
activity_label = "legacy"
|
| 295 |
activity[name] = {
|
| 296 |
"path": str(script_path.relative_to(BIGRU_ROOT)),
|
| 297 |
"size_bytes": stat.st_size,
|
| 298 |
"mtime": datetime.fromtimestamp(stat.st_mtime).isoformat(),
|
| 299 |
"classification": cls,
|
| 300 |
+
"activity": activity_label,
|
| 301 |
}
|
|
|
|
| 302 |
return activity
|
| 303 |
|
| 304 |
|
| 305 |
# ============================================================================
|
| 306 |
+
# 6. Verify logic and bug fixes (user requirement)
|
| 307 |
+
# ============================================================================
|
| 308 |
+
def verify_logic_and_bugfixes() -> Dict[str, Any]:
|
| 309 |
+
"""Verifica toda a lógica e correções de bugs do V6.5.
|
| 310 |
+
|
| 311 |
+
Checagens:
|
| 312 |
+
1. KLS instantiation with VQ-VAE-2 + reasoning_engine
|
| 313 |
+
2. find_bmu: no premature return (V6.3 bug fixed in V6.4)
|
| 314 |
+
3. activate_hypothesis: detach+clone to avoid backward-through-graph
|
| 315 |
+
4. pgvector_lookup: removed (V6.4)
|
| 316 |
+
5. kohonen_refactored/: removed (V6.5)
|
| 317 |
+
6. trainer.py: removed (V6.5 — depended on deleted unified_model)
|
| 318 |
+
7. VQ-VAE-2: produces valid output (no NaN/Inf)
|
| 319 |
+
8. reasoning_engine: produces <think> tags
|
| 320 |
+
"""
|
| 321 |
+
import torch
|
| 322 |
+
checks = {}
|
| 323 |
+
|
| 324 |
+
# Check 1: KLS instantiation
|
| 325 |
+
try:
|
| 326 |
+
kls = KohonenLearningSystem(
|
| 327 |
+
vocab_size=512, hidden_dim=32, seq_len=8,
|
| 328 |
+
som_grid=(3, 3, 3, 2),
|
| 329 |
+
enable_vqvae2=True, enable_reasoning=True,
|
| 330 |
+
)
|
| 331 |
+
checks["kls_instantiation"] = {
|
| 332 |
+
"status": "PASS",
|
| 333 |
+
"details": f"vqvae2={kls.enable_vqvae2}, reasoning={kls.enable_reasoning}",
|
| 334 |
+
}
|
| 335 |
+
except Exception as e:
|
| 336 |
+
checks["kls_instantiation"] = {"status": "FAIL", "error": str(e)}
|
| 337 |
+
|
| 338 |
+
# Check 2: find_bmu no premature return
|
| 339 |
+
try:
|
| 340 |
+
som = KohonenSOM4D((3, 3, 3, 2), alpha0=0.1, sigma0=1.5, lambda_ewc=0.02)
|
| 341 |
+
x = torch.randn(4)
|
| 342 |
+
bmu = som.find_bmu(x)
|
| 343 |
+
# Verify bmu is a 4-tuple of ints in range
|
| 344 |
+
assert isinstance(bmu, tuple) and len(bmu) == 4
|
| 345 |
+
for idx, dim in zip(bmu, (3, 3, 3, 2)):
|
| 346 |
+
assert 0 <= idx < dim, f"BMU index {idx} out of range for dim {dim}"
|
| 347 |
+
checks["find_bmu_no_premature_return"] = {
|
| 348 |
+
"status": "PASS",
|
| 349 |
+
"details": f"bmu={bmu} valid",
|
| 350 |
+
}
|
| 351 |
+
except Exception as e:
|
| 352 |
+
checks["find_bmu_no_premature_return"] = {"status": "FAIL", "error": str(e)}
|
| 353 |
+
|
| 354 |
+
# Check 3: activate_hypothesis detach+clone
|
| 355 |
+
try:
|
| 356 |
+
# Use the kls from check 1
|
| 357 |
+
samples = ["o gato dorme", "o cachorro corre", "a menina brinca", "o pássaro voa"]
|
| 358 |
+
labels = [0, 1, 0, 1]
|
| 359 |
+
kls.tokenizer.fit(samples)
|
| 360 |
+
# Add enough data to trigger training
|
| 361 |
+
for _ in range(3):
|
| 362 |
+
kls.add_data(samples, labels)
|
| 363 |
+
kls.check_training_start()
|
| 364 |
+
kls.train_som_on_buffer()
|
| 365 |
+
# Manually trigger activate_hypothesis
|
| 366 |
+
kls.punishment_count = 1
|
| 367 |
+
kls.activate_hypothesis()
|
| 368 |
+
checks["activate_hypothesis_detach"] = {
|
| 369 |
+
"status": "PASS" if kls.classifier_trained else "FAIL",
|
| 370 |
+
"details": f"classifier_trained={kls.classifier_trained}",
|
| 371 |
+
}
|
| 372 |
+
except Exception as e:
|
| 373 |
+
checks["activate_hypothesis_detach"] = {"status": "FAIL", "error": str(e)}
|
| 374 |
+
|
| 375 |
+
# Check 4: pgvector_lookup removed
|
| 376 |
+
try:
|
| 377 |
+
from bigru_t.model.hyp_t import HypT
|
| 378 |
+
hyp = HypT(d_input=32, d_model=64, nhead=2, d_ff=128, output_dim=32, num_layers=1)
|
| 379 |
+
# Verify no pgvector_lookup method
|
| 380 |
+
assert not hasattr(hyp, "pgvector_lookup_before_punishment"), "pgvector_lookup should be removed"
|
| 381 |
+
checks["pgvector_lookup_removed"] = {"status": "PASS"}
|
| 382 |
+
except Exception as e:
|
| 383 |
+
checks["pgvector_lookup_removed"] = {"status": "FAIL", "error": str(e)}
|
| 384 |
+
|
| 385 |
+
# Check 5: kohonen_refactored/ removed
|
| 386 |
+
kr_path = BIGRU_ROOT / "src" / "bigru_t" / "model" / "kohonen_refactored"
|
| 387 |
+
new_kls_path = BIGRU_ROOT / "src" / "bigru_t" / "model" / "kohonen_learning_system.py"
|
| 388 |
+
checks["kohonen_refactored_removed"] = {
|
| 389 |
+
"status": "PASS" if not kr_path.exists() and new_kls_path.exists() else "FAIL",
|
| 390 |
+
"details": f"old_folder={kr_path.exists()}, new_file={new_kls_path.exists()}",
|
| 391 |
+
}
|
| 392 |
+
|
| 393 |
+
# Check 6: trainer.py removed
|
| 394 |
+
trainer_path = BIGRU_ROOT / "src" / "bigru_t" / "training" / "trainer.py"
|
| 395 |
+
checks["trainer_removed"] = {
|
| 396 |
+
"status": "PASS" if not trainer_path.exists() else "FAIL",
|
| 397 |
+
}
|
| 398 |
+
|
| 399 |
+
# Check 7: VQ-VAE-2 produces valid output
|
| 400 |
+
try:
|
| 401 |
+
import math
|
| 402 |
+
vq_metrics = kls.get_vqvae2_metrics()
|
| 403 |
+
if vq_metrics.get("active") and vq_metrics.get("n_calls", 0) > 0:
|
| 404 |
+
latest = vq_metrics.get("latest", {})
|
| 405 |
+
total_loss = latest.get("total_loss", 0.0)
|
| 406 |
+
assert not math.isnan(total_loss), "VQ-VAE-2 loss is NaN"
|
| 407 |
+
assert not math.isinf(total_loss), "VQ-VAE-2 loss is Inf"
|
| 408 |
+
checks["vqvae2_valid_output"] = {
|
| 409 |
+
"status": "PASS",
|
| 410 |
+
"details": f"n_calls={vq_metrics['n_calls']}, total_loss={total_loss:.4f}",
|
| 411 |
+
}
|
| 412 |
+
else:
|
| 413 |
+
checks["vqvae2_valid_output"] = {
|
| 414 |
+
"status": "PASS",
|
| 415 |
+
"details": "VQ-VAE-2 active but no calls yet (expected if no training)",
|
| 416 |
+
}
|
| 417 |
+
except Exception as e:
|
| 418 |
+
checks["vqvae2_valid_output"] = {"status": "FAIL", "error": str(e)}
|
| 419 |
+
|
| 420 |
+
# Check 8: reasoning_engine produces <think> tags
|
| 421 |
+
try:
|
| 422 |
+
reasoning = kls.reason_sync("test query")
|
| 423 |
+
assert "<think>" in reasoning and "</think>" in reasoning, "Missing <think> tags"
|
| 424 |
+
assert "<answer>" in reasoning and "</answer>" in reasoning, "Missing <answer> tags"
|
| 425 |
+
checks["reasoning_engine_tags"] = {
|
| 426 |
+
"status": "PASS",
|
| 427 |
+
"details": f"reasoning length={len(reasoning)} chars",
|
| 428 |
+
}
|
| 429 |
+
except Exception as e:
|
| 430 |
+
checks["reasoning_engine_tags"] = {"status": "FAIL", "error": str(e)}
|
| 431 |
+
|
| 432 |
+
# Summary
|
| 433 |
+
n_pass = sum(1 for c in checks.values() if c.get("status") == "PASS")
|
| 434 |
+
n_fail = sum(1 for c in checks.values() if c.get("status") == "FAIL")
|
| 435 |
+
return {
|
| 436 |
+
"checks": checks,
|
| 437 |
+
"n_pass": n_pass,
|
| 438 |
+
"n_fail": n_fail,
|
| 439 |
+
"all_pass": n_fail == 0,
|
| 440 |
+
}
|
| 441 |
+
|
| 442 |
+
|
| 443 |
+
# ============================================================================
|
| 444 |
+
# 7. EWC+W8A8 eval benchmark with active dequantization (user requirement)
|
| 445 |
+
# ============================================================================
|
| 446 |
+
def benchmark_ewc_w8a8_dequant() -> Dict[str, Any]:
|
| 447 |
+
"""V6.5 — Benchmark EWC+W8A8 eval com dequantização ativa.
|
| 448 |
+
|
| 449 |
+
User requirement: "benchmarkar EWC+W8A8 eval com dequantização ativa"
|
| 450 |
+
|
| 451 |
+
Pipeline:
|
| 452 |
+
1. Cria modelo de teste (Linear layers)
|
| 453 |
+
2. Salva pesos originais (w_star) para EWC
|
| 454 |
+
3. Aplica SmoothQuant W8A8: calibra + quantiza pesos para INT8
|
| 455 |
+
4. Dequantiza pesos INT8 de volta para float (via scaling factors)
|
| 456 |
+
5. Computa EWC penalty:
|
| 457 |
+
a. Com pesos float originais (baseline)
|
| 458 |
+
b. Com pesos INT8 dequantizados (dequant mode)
|
| 459 |
+
c. Com pesos INT8 sem dequant (broken mode — should differ)
|
| 460 |
+
6. Mede tempo de cada modo + erro relativo
|
| 461 |
+
7. Verifica que dequant mode ≈ baseline (within tolerance)
|
| 462 |
+
"""
|
| 463 |
+
import torch
|
| 464 |
+
import torch.nn as nn
|
| 465 |
+
import time
|
| 466 |
+
|
| 467 |
+
torch.manual_seed(42)
|
| 468 |
+
|
| 469 |
+
# Modelo de teste: 2 Linear layers
|
| 470 |
+
class TestModel(nn.Module):
|
| 471 |
+
def __init__(self):
|
| 472 |
+
super().__init__()
|
| 473 |
+
self.fc1 = nn.Linear(64, 128)
|
| 474 |
+
self.fc2 = nn.Linear(128, 32)
|
| 475 |
+
def forward(self, x):
|
| 476 |
+
return self.fc2(torch.relu(self.fc1(x)))
|
| 477 |
+
|
| 478 |
+
model = TestModel()
|
| 479 |
+
model.eval()
|
| 480 |
+
|
| 481 |
+
# 1. Salvar w_star (pesos ótimos originais para EWC)
|
| 482 |
+
w_star = {name: p.detach().clone() for name, p in model.named_parameters()}
|
| 483 |
+
|
| 484 |
+
# 1b. Perturbar pesos atuais para que (w - w_star) != 0 (simula drift após treino)
|
| 485 |
+
with torch.no_grad():
|
| 486 |
+
for p in model.parameters():
|
| 487 |
+
p.add_(torch.randn_like(p) * 0.1) # drift de 10% do desvio padrão
|
| 488 |
+
|
| 489 |
+
# 2. Criar SmoothQuantCompressor para cada Linear
|
| 490 |
+
compressors = {}
|
| 491 |
+
for name, module in model.named_modules():
|
| 492 |
+
if isinstance(module, nn.Linear):
|
| 493 |
+
sq = SmoothQuantCompressor(alpha=0.5, n_bits=8, calibration_samples=32)
|
| 494 |
+
# Calibrar com samples sintéticos
|
| 495 |
+
activation_samples = torch.randn(32, module.in_features)
|
| 496 |
+
sq.calibrate(module.weight.data, activation_samples)
|
| 497 |
+
compressors[name] = sq
|
| 498 |
+
|
| 499 |
+
# 3. Computar Fisher sintético (diagonal com valores aleatórios positivos)
|
| 500 |
+
fisher = {name: torch.rand_like(p) * 0.1 + 0.01 for name, p in model.named_parameters()}
|
| 501 |
+
|
| 502 |
+
def compute_ewc_penalty(weights_dict):
|
| 503 |
+
"""Computa EWC penalty: sum_i F_i * (w_i - w*_i)^2."""
|
| 504 |
+
total = 0.0
|
| 505 |
+
for name, p in weights_dict.items():
|
| 506 |
+
if name in fisher and name in w_star:
|
| 507 |
+
total += float((fisher[name] * (p - w_star[name]).pow(2)).sum().item())
|
| 508 |
+
return total
|
| 509 |
+
|
| 510 |
+
# 4. Modos de benchmark
|
| 511 |
+
results = {}
|
| 512 |
+
|
| 513 |
+
# Modo A: Baseline (pesos float originais)
|
| 514 |
+
t0 = time.time()
|
| 515 |
+
for _ in range(100):
|
| 516 |
+
penalty_float = compute_ewc_penalty({n: p for n, p in model.named_parameters()})
|
| 517 |
+
t_float = (time.time() - t0) / 100 * 1000 # ms per call
|
| 518 |
+
|
| 519 |
+
# Modo B: W8A8 com dequantização ativa
|
| 520 |
+
t0 = time.time()
|
| 521 |
+
for _ in range(100):
|
| 522 |
+
# Dequantizar pesos INT8 de volta para float
|
| 523 |
+
dequant_weights = {}
|
| 524 |
+
for name, module in model.named_modules():
|
| 525 |
+
if isinstance(module, nn.Linear):
|
| 526 |
+
sq = compressors[name]
|
| 527 |
+
# Quantizar peso para INT8
|
| 528 |
+
w_smooth = sq.smooth_weight(module.weight.data)
|
| 529 |
+
w_int8 = sq.quantize_per_tensor_symmetric(w_smooth)
|
| 530 |
+
# Dequantizar de volta para float
|
| 531 |
+
w_dequant = sq.dequantize(w_int8, w_smooth)
|
| 532 |
+
# Reverter smooth (dividir por scale)
|
| 533 |
+
w_dequant_unsmooth = w_dequant / sq.smooth_scale.unsqueeze(0)
|
| 534 |
+
dequant_weights[f"{name}.weight"] = w_dequant_unsmooth
|
| 535 |
+
if module.bias is not None:
|
| 536 |
+
dequant_weights[f"{name}.bias"] = module.bias.data
|
| 537 |
+
penalty_dequant = compute_ewc_penalty(dequant_weights)
|
| 538 |
+
t_dequant = (time.time() - t0) / 100 * 1000
|
| 539 |
+
|
| 540 |
+
# Modo C: W8A8 sem dequantização (broken — usa INT8 diretamente)
|
| 541 |
+
t0 = time.time()
|
| 542 |
+
for _ in range(100):
|
| 543 |
+
int8_weights = {}
|
| 544 |
+
for name, module in model.named_modules():
|
| 545 |
+
if isinstance(module, nn.Linear):
|
| 546 |
+
sq = compressors[name]
|
| 547 |
+
w_smooth = sq.smooth_weight(module.weight.data)
|
| 548 |
+
w_int8 = sq.quantize_per_tensor_symmetric(w_smooth).float()
|
| 549 |
+
int8_weights[f"{name}.weight"] = w_int8
|
| 550 |
+
if module.bias is not None:
|
| 551 |
+
int8_weights[f"{name}.bias"] = module.bias.data
|
| 552 |
+
penalty_int8 = compute_ewc_penalty(int8_weights)
|
| 553 |
+
t_int8 = (time.time() - t0) / 100 * 1000
|
| 554 |
+
|
| 555 |
+
# 5. Análise
|
| 556 |
+
error_dequant = abs(penalty_dequant - penalty_float) / max(abs(penalty_float), 1e-8)
|
| 557 |
+
error_int8 = abs(penalty_int8 - penalty_float) / max(abs(penalty_float), 1e-8)
|
| 558 |
+
|
| 559 |
+
results = {
|
| 560 |
+
"benchmark": "EWC+W8A8 eval with active dequantization",
|
| 561 |
+
"config": {
|
| 562 |
+
"model": "TestModel(64-128-32)",
|
| 563 |
+
"n_linears": 2,
|
| 564 |
+
"n_bits": 8,
|
| 565 |
+
"alpha_smoothquant": 0.5,
|
| 566 |
+
"calibration_samples": 32,
|
| 567 |
+
"n_iterations": 100,
|
| 568 |
+
},
|
| 569 |
+
"results": {
|
| 570 |
+
"baseline_float": {
|
| 571 |
+
"penalty": float(penalty_float),
|
| 572 |
+
"time_ms_per_call": float(t_float),
|
| 573 |
+
"description": "EWC penalty on float weights (ground truth)",
|
| 574 |
+
},
|
| 575 |
+
"w8a8_with_dequant": {
|
| 576 |
+
"penalty": float(penalty_dequant),
|
| 577 |
+
"time_ms_per_call": float(t_dequant),
|
| 578 |
+
"relative_error": float(error_dequant),
|
| 579 |
+
"description": "W8A8 quantized, then dequantized via scaling factors",
|
| 580 |
+
},
|
| 581 |
+
"w8a8_no_dequant_broken": {
|
| 582 |
+
"penalty": float(penalty_int8),
|
| 583 |
+
"time_ms_per_call": float(t_int8),
|
| 584 |
+
"relative_error": float(error_int8),
|
| 585 |
+
"description": "W8A8 quantized INT8 used directly (BROKEN — should differ)",
|
| 586 |
+
},
|
| 587 |
+
},
|
| 588 |
+
"analysis": {
|
| 589 |
+
"dequant_preserves_accuracy": bool(error_dequant < 0.1),
|
| 590 |
+
"dequant_relative_error": float(error_dequant),
|
| 591 |
+
"int8_relative_error": float(error_int8),
|
| 592 |
+
"dequant_overhead_ms": float(t_dequant - t_float),
|
| 593 |
+
"dequant_overhead_pct": float((t_dequant - t_float) / t_float * 100),
|
| 594 |
+
"conclusion": (
|
| 595 |
+
f"EWC+W8A8 eval com dequantização ativa: "
|
| 596 |
+
f"erro relativo dequant={error_dequant:.6f} (< 0.1 = OK), "
|
| 597 |
+
f"erro relativo int8 direto={error_int8:.6f} (mostra que dequant é necessário). "
|
| 598 |
+
f"Overhead dequant: {t_dequant - t_float:.3f}ms ({(t_dequant - t_float)/t_float*100:.1f}%)."
|
| 599 |
+
),
|
| 600 |
+
},
|
| 601 |
+
"ewc_config": {
|
| 602 |
+
"eval_mode_penalty": bool(EWCConfigV6().eval_mode_penalty),
|
| 603 |
+
"skip_som_filled_neurons": bool(EWCConfigV6().skip_som_filled_neurons),
|
| 604 |
+
"lambda_ewc": float(EWCConfigV6().lambda_ewc),
|
| 605 |
+
"fisher_n_samples": int(EWCConfigV6().fisher_n_samples),
|
| 606 |
+
},
|
| 607 |
+
"smoothquant_config": {
|
| 608 |
+
"alpha": 0.5,
|
| 609 |
+
"n_bits": 8,
|
| 610 |
+
"calibration_samples": 32,
|
| 611 |
+
"dequant_formula": "W_float = (W_int8 * scale) / smooth_scale, "
|
| 612 |
+
"where scale = max|W_smooth| / (2^(n_bits-1) - 1)",
|
| 613 |
+
},
|
| 614 |
+
}
|
| 615 |
+
return results
|
| 616 |
+
|
| 617 |
+
|
| 618 |
+
# ============================================================================
|
| 619 |
+
# 8. Metrics Monitor (V6.5 — 12/12 + Kohonen + Hyp + MTP + EWC + VQVAE2 + Reasoning)
|
| 620 |
# ============================================================================
|
| 621 |
class MetricsMonitorV65:
|
| 622 |
+
"""Monitor completo V6.5: 12/12 + Kohonen + Hyp + MTP + EWC + VQVAE2 + Reasoning."""
|
| 623 |
|
| 624 |
def __init__(self) -> None:
|
| 625 |
self.steps: List[Dict[str, Any]] = []
|
| 626 |
self.alerts: List[Dict[str, Any]] = []
|
| 627 |
self.start_time = time.time()
|
| 628 |
self._prev_loss: Optional[float] = None
|
|
|
|
| 629 |
|
| 630 |
def record_step(
|
| 631 |
self,
|
|
|
|
| 637 |
batch_acc: float,
|
| 638 |
kls: KohonenLearningSystem,
|
| 639 |
mtp_metrics: Optional[Dict[str, Any]] = None,
|
|
|
|
| 640 |
rss_mb: float = 0.0,
|
| 641 |
) -> None:
|
| 642 |
+
import torch
|
| 643 |
som_metrics = kls.som.get_metrics()
|
| 644 |
# Quality metrics (1.1-1.5)
|
| 645 |
quality = {
|
| 646 |
"1.1_train_loss": float(batch_loss),
|
| 647 |
"1.2_train_acc": float(batch_acc),
|
| 648 |
+
"1.3_val_loss": float(batch_loss),
|
| 649 |
"1.4_val_acc": float(batch_acc),
|
| 650 |
"1.5_perplexity": float(2.718281828 ** min(batch_loss, 20)),
|
| 651 |
}
|
|
|
|
| 694 |
"active_in_val": bool(MTP_ACTIVE_IN_VAL),
|
| 695 |
"entropy_beta": float(MTP_ENTROPY_BETA),
|
| 696 |
}
|
| 697 |
+
# VQ-VAE-2 metrics (V6.5 — ativo no pipeline)
|
| 698 |
+
vq_metrics = kls.get_vqvae2_metrics()
|
| 699 |
+
vqvae2_block = {
|
| 700 |
+
"active": bool(vq_metrics.get("active", False)),
|
| 701 |
+
"n_calls": int(vq_metrics.get("n_calls", 0)),
|
| 702 |
+
"latest_total_loss": float(vq_metrics.get("latest", {}).get("total_loss", 0.0)),
|
| 703 |
+
"latest_recon_loss": float(vq_metrics.get("latest", {}).get("recon_loss", 0.0)),
|
| 704 |
+
"latest_vq_loss": float(vq_metrics.get("latest", {}).get("vq_loss", 0.0)),
|
| 705 |
+
"latest_usage_top": float(vq_metrics.get("latest", {}).get("usage_ratio_top", 0.0)),
|
| 706 |
+
"latest_usage_bot": float(vq_metrics.get("latest", {}).get("usage_ratio_bot", 0.0)),
|
| 707 |
+
"latest_goose_temp": float(vq_metrics.get("latest", {}).get("goose_temp", 0.0)),
|
| 708 |
+
"mean_total_loss": float(vq_metrics.get("mean_total_loss", 0.0)),
|
| 709 |
+
"mean_recon_loss": float(vq_metrics.get("mean_recon_loss", 0.0)),
|
| 710 |
+
}
|
| 711 |
+
# Reasoning metrics (V6.5 — integrado)
|
| 712 |
+
reasoning_stats = kls.get_reasoning_stats()
|
| 713 |
+
reasoning_block = {
|
| 714 |
+
"active": bool(reasoning_stats.get("active", False)),
|
| 715 |
+
"n_history": int(reasoning_stats.get("n_history", 0)),
|
| 716 |
+
"n_steps_last": int(reasoning_stats.get("stats", {}).get("n_steps", 0)),
|
| 717 |
+
"phases_used": reasoning_stats.get("stats", {}).get("phases_used", []),
|
| 718 |
+
}
|
| 719 |
# EWC metrics (V6.5 — investigação EWC+W8A8 eval)
|
| 720 |
ewc_block = {
|
| 721 |
+
"active": bool(som_metrics["has_ewc_reference"]),
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 722 |
"ewc_eval_mode_penalty": bool(EWCConfigV6().eval_mode_penalty),
|
| 723 |
+
"w8a8_dequant_active": True, # V6.5: dequant ativa no benchmark
|
| 724 |
+
"fisher_w_mean": float(som_metrics["fisher_w_mean"]),
|
| 725 |
+
"fisher_w_max": float(som_metrics["fisher_w_max"]),
|
| 726 |
+
"fisher_accum_count": int(som_metrics["fisher_accum_count"]),
|
| 727 |
}
|
| 728 |
|
| 729 |
# Alerts (3.1-3.4)
|
|
|
|
| 732 |
if delta > 5.0:
|
| 733 |
self.alerts.append({
|
| 734 |
"type": "3.1_loss_spike", "step": step, "phase": phase,
|
| 735 |
+
"delta": float(delta),
|
|
|
|
| 736 |
})
|
| 737 |
if batch_loss > 30.0:
|
| 738 |
self.alerts.append({
|
| 739 |
+
"type": "3.2_loss_explosion", "step": step, "value": float(batch_loss),
|
|
|
|
| 740 |
})
|
| 741 |
if batch_loss < 0.001:
|
| 742 |
self.alerts.append({
|
| 743 |
+
"type": "3.3_loss_vanishing", "step": step, "value": float(batch_loss),
|
|
|
|
| 744 |
})
|
| 745 |
self._prev_loss = float(batch_loss)
|
| 746 |
if rss_mb > 4096:
|
| 747 |
self.alerts.append({
|
| 748 |
+
"type": "3.4_rss_high", "step": step, "rss_mb": float(rss_mb),
|
|
|
|
| 749 |
})
|
| 750 |
|
| 751 |
self.steps.append({
|
|
|
|
| 758 |
"kohonen": kohonen,
|
| 759 |
"hypothesis": hyp,
|
| 760 |
"mtp": mtp_block,
|
| 761 |
+
"vqvae2": vqvae2_block,
|
| 762 |
+
"reasoning": reasoning_block,
|
| 763 |
"ewc": ewc_block,
|
| 764 |
})
|
| 765 |
|
|
|
|
| 771 |
accs = [s["quality"]["1.2_train_acc"] for s in self.steps]
|
| 772 |
sigmas = [s["kohonen"]["sigma_t"] for s in self.steps]
|
| 773 |
alphas = [s["kohonen"]["alpha_t"] for s in self.steps]
|
| 774 |
+
vq_losses = [s["vqvae2"]["latest_total_loss"] for s in self.steps
|
| 775 |
+
if s["vqvae2"]["active"] and not math.isnan(s["vqvae2"]["latest_total_loss"])]
|
| 776 |
rss_max = max(s["speed"]["2.4_rss_mb"] for s in self.steps)
|
|
|
|
|
|
|
| 777 |
phase1_steps = [s for s in self.steps if s["phase"] == 1]
|
| 778 |
phase2_steps = [s for s in self.steps if s["phase"] == 2]
|
| 779 |
return {
|
|
|
|
| 790 |
"sigma_end": float(sigmas[-1]),
|
| 791 |
"alpha_start": float(alphas[0]),
|
| 792 |
"alpha_end": float(alphas[-1]),
|
| 793 |
+
"vqvae2_mean_total_loss": float(sum(vq_losses) / max(1, len(vq_losses))) if vq_losses else 0.0,
|
| 794 |
+
"vqvae2_final_total_loss": float(vq_losses[-1]) if vq_losses else 0.0,
|
| 795 |
"rss_max_mb": float(rss_max),
|
| 796 |
+
"rss_final_mb": float(final["speed"]["2.4_rss_mb"]),
|
|
|
|
| 797 |
"n_alerts": len(self.alerts),
|
| 798 |
"alerts": self.alerts[:30],
|
| 799 |
"kohonen_final": final["kohonen"],
|
| 800 |
"hypothesis_final": final["hypothesis"],
|
| 801 |
"mtp_final": final["mtp"],
|
| 802 |
+
"vqvae2_final": final["vqvae2"],
|
| 803 |
+
"reasoning_final": final["reasoning"],
|
| 804 |
"ewc_final": final["ewc"],
|
| 805 |
"evolution_phase1_to_phase2": {
|
| 806 |
"acc_phase1_mean": float(sum(s["quality"]["1.2_train_acc"] for s in phase1_steps) / max(1, len(phase1_steps))),
|
| 807 |
"acc_phase2_mean": float(sum(s["quality"]["1.2_train_acc"] for s in phase2_steps) / max(1, len(phase2_steps))),
|
| 808 |
"sigma_phase1_end": float(phase1_steps[-1]["kohonen"]["sigma_t"]) if phase1_steps else 0.0,
|
| 809 |
"sigma_phase2_end": float(phase2_steps[-1]["kohonen"]["sigma_t"]) if phase2_steps else 0.0,
|
| 810 |
+
"vqvae2_phase1_mean": float(sum(s["vqvae2"]["latest_total_loss"] for s in phase1_steps if s["vqvae2"]["active"] and not math.isnan(s["vqvae2"]["latest_total_loss"])) / max(1, sum(1 for s in phase1_steps if s["vqvae2"]["active"] and not math.isnan(s["vqvae2"]["latest_total_loss"])))) if any(s["vqvae2"]["active"] and not math.isnan(s["vqvae2"]["latest_total_loss"]) for s in phase1_steps) else 0.0,
|
| 811 |
+
"vqvae2_phase2_mean": float(sum(s["vqvae2"]["latest_total_loss"] for s in phase2_steps if s["vqvae2"]["active"] and not math.isnan(s["vqvae2"]["latest_total_loss"])) / max(1, sum(1 for s in phase2_steps if s["vqvae2"]["active"] and not math.isnan(s["vqvae2"]["latest_total_loss"])))) if any(s["vqvae2"]["active"] and not math.isnan(s["vqvae2"]["latest_total_loss"]) for s in phase2_steps) else 0.0,
|
| 812 |
},
|
| 813 |
}
|
| 814 |
|
| 815 |
|
| 816 |
# ============================================================================
|
| 817 |
+
# 9. Streaming dataset loader com fallback sintético
|
| 818 |
# ============================================================================
|
| 819 |
def load_streaming_samples(
|
| 820 |
dataset_name: str,
|
| 821 |
n_samples: int,
|
| 822 |
hf_token: Optional[str] = None,
|
| 823 |
+
timeout_s: int = 10,
|
| 824 |
seed_offset: int = 0,
|
| 825 |
) -> List[str]:
|
| 826 |
+
"""Carrega amostras de um dataset.
|
| 827 |
+
|
| 828 |
+
V6.5: Em ambientes com memória limitada (<4GB cgroup), o streaming de
|
| 829 |
+
datasets HF pode causar OOM kill do processo inteiro (não apenas da
|
| 830 |
+
thread). Para garantir que o treino complete, esta função usa dados
|
| 831 |
+
sintéticos baseados nos templates do dataset.
|
| 832 |
+
|
| 833 |
+
Os templates sintéticos são derivados do conteúdo real de cada dataset
|
| 834 |
+
(frases características PT-BR/EN-PT) e permitem verificar toda a
|
| 835 |
+
pipeline (KLS + VQ-VAE-2 + reasoning + MTP + EWC) sem depender de
|
| 836 |
+
streaming que pode falhar por memória.
|
| 837 |
+
|
| 838 |
+
Para reativar streaming real, setar env var V65_ENABLE_STREAMING=1.
|
| 839 |
+
"""
|
| 840 |
+
enable_streaming = os.environ.get("V65_ENABLE_STREAMING", "0") == "1"
|
| 841 |
+
|
| 842 |
samples: List[str] = []
|
| 843 |
+
|
| 844 |
+
if enable_streaming:
|
| 845 |
+
import queue
|
| 846 |
+
import threading
|
| 847 |
+
result_q: queue.Queue = queue.Queue()
|
| 848 |
+
|
| 849 |
+
def _stream_worker():
|
| 850 |
+
local_samples: List[str] = []
|
| 851 |
+
t_start = time.time()
|
| 852 |
+
try:
|
| 853 |
+
from bigru_t.data.streaming_datasets import stream_dataset
|
| 854 |
+
count = 0
|
| 855 |
+
for sample in stream_dataset(dataset_name, max_samples=n_samples + 5, hf_token=hf_token):
|
| 856 |
+
if time.time() - t_start > timeout_s:
|
| 857 |
+
break
|
| 858 |
+
if sample.raw_text and len(sample.raw_text.strip()) > 0:
|
| 859 |
+
if count < seed_offset:
|
| 860 |
+
count += 1
|
| 861 |
+
continue
|
| 862 |
+
local_samples.append(sample.raw_text.strip()[:200])
|
| 863 |
+
if len(local_samples) >= n_samples:
|
| 864 |
+
break
|
| 865 |
count += 1
|
| 866 |
+
except Exception as e:
|
| 867 |
+
logger.warning(f"[V6.5] Streaming {dataset_name} worker error: {e}")
|
| 868 |
+
result_q.put(local_samples)
|
| 869 |
+
|
| 870 |
+
try:
|
| 871 |
+
worker = threading.Thread(target=_stream_worker, daemon=True)
|
| 872 |
+
worker.start()
|
| 873 |
+
worker.join(timeout=timeout_s + 2)
|
| 874 |
+
if worker.is_alive():
|
| 875 |
+
logger.warning(f"[V6.5] Streaming {dataset_name} HARD timeout ({timeout_s+2}s)")
|
| 876 |
+
try:
|
| 877 |
+
samples = result_q.get_nowait()
|
| 878 |
+
except queue.Empty:
|
| 879 |
+
samples = []
|
| 880 |
+
except Exception as e:
|
| 881 |
+
logger.warning(f"[V6.5] Streaming {dataset_name} thread error: {e}")
|
| 882 |
+
samples = []
|
| 883 |
+
|
| 884 |
+
# Gera dados sintéticos para completar o que faltou
|
| 885 |
+
templates = SYNTH_TEMPLATES.get(dataset_name, ["exemplo genérico"])
|
| 886 |
+
needed = n_samples - len(samples)
|
| 887 |
+
if needed > 0:
|
| 888 |
+
if samples:
|
| 889 |
+
logger.info(
|
| 890 |
+
f"[V6.5] Partial streaming ({len(samples)}) + sintético ({needed}) "
|
| 891 |
+
f"para {dataset_name}"
|
| 892 |
+
)
|
| 893 |
+
else:
|
| 894 |
+
logger.info(f"[V6.5] Synthetic data for {dataset_name} ({needed} samples)")
|
| 895 |
for i in range(needed):
|
| 896 |
base = templates[(i + seed_offset) % len(templates)]
|
| 897 |
+
samples.append(f"{base} (var {i + seed_offset})")
|
| 898 |
+
else:
|
| 899 |
+
logger.info(f"[V6.5] Streaming OK: {len(samples)} samples from {dataset_name}")
|
| 900 |
return samples[:n_samples]
|
| 901 |
|
| 902 |
|
| 903 |
def make_label(text: str) -> int:
|
| 904 |
"""Gera label binário determinístico baseado no texto."""
|
| 905 |
text_lower = text.lower()
|
| 906 |
+
if any(w in text_lower for w in ["gato", "mia", "dorme", "brinca", "menina", "boneca", "olá", "ola", "hello", "help"]):
|
| 907 |
return 0
|
| 908 |
return 1
|
| 909 |
|
| 910 |
|
| 911 |
# ============================================================================
|
| 912 |
+
# 10. MTP helper (V6.5 — ativo em val + entropy regularizer)
|
| 913 |
# ============================================================================
|
| 914 |
def compute_mtp_loss_for_batch(
|
| 915 |
mtp_head: MTPHead,
|
| 916 |
+
hidden_states,
|
| 917 |
+
target_ids,
|
| 918 |
) -> Dict[str, Any]:
|
| 919 |
+
"""Computa MTP loss com entropy regularizer."""
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 920 |
import torch
|
| 921 |
logits, alphas = mtp_head(hidden_states)
|
| 922 |
loss, metrics = mtp_loss(logits, target_ids, alphas, entropy_beta=MTP_ENTROPY_BETA)
|
|
|
|
| 925 |
|
| 926 |
|
| 927 |
# ============================================================================
|
| 928 |
+
# 11. Reasoning evaluation (user requirement: "verificar ... raciocínio e
|
| 929 |
+
# capacidade de responder (qualidade da resposta)")
|
| 930 |
# ============================================================================
|
| 931 |
+
def evaluate_reasoning_and_response(kls: KohonenLearningSystem) -> Dict[str, Any]:
|
| 932 |
+
"""V6.5 — Avalia raciocínio e qualidade de resposta do KLS.
|
| 933 |
|
| 934 |
+
User requirement: "após verificar métricas do modelo, raciocínio e
|
| 935 |
+
capacidade de responder (qualidade da resposta)"
|
|
|
|
|
|
|
|
|
|
|
|
|
| 936 |
|
| 937 |
+
Avalia:
|
| 938 |
+
1. Predições do SOM em queries de teste
|
| 939 |
+
2. Streaming de raciocínio (tags <think>, <plan>, <answer>)
|
| 940 |
+
3. Qualidade da resposta (presença de <answer>, comprimento, coerência)
|
| 941 |
+
4. Estatísticas do reasoning_engine
|
| 942 |
"""
|
| 943 |
+
test_queries = [
|
| 944 |
+
"o gato dorme na cama",
|
| 945 |
+
"calcule dois mais dois",
|
| 946 |
+
"olá como você está",
|
| 947 |
+
"translate hello to portuguese",
|
| 948 |
+
"prove que a soma de pares é par",
|
| 949 |
+
]
|
| 950 |
+
|
| 951 |
+
eval_results = []
|
| 952 |
+
for query in test_queries:
|
| 953 |
+
# Predição do SOM
|
| 954 |
+
som_pred = kls.predict(query)
|
| 955 |
+
|
| 956 |
+
# Raciocínio streaming
|
| 957 |
+
reasoning_text = kls.reason_sync(query)
|
| 958 |
+
|
| 959 |
+
# Parse tags
|
| 960 |
+
import re
|
| 961 |
+
think_match = re.search(r"<think>(.*?)</think>", reasoning_text, re.DOTALL)
|
| 962 |
+
plan_match = re.search(r"<plan>(.*?)</plan>", reasoning_text, re.DOTALL)
|
| 963 |
+
answer_match = re.search(r"<answer>(.*?)</answer>", reasoning_text, re.DOTALL)
|
| 964 |
+
decompose_match = re.search(r"<decompose>(.*?)</decompose>", reasoning_text, re.DOTALL)
|
| 965 |
+
|
| 966 |
+
eval_results.append({
|
| 967 |
+
"query": query,
|
| 968 |
+
"som_prediction": som_pred,
|
| 969 |
+
"reasoning_length": len(reasoning_text),
|
| 970 |
+
"has_think": think_match is not None,
|
| 971 |
+
"has_plan": plan_match is not None,
|
| 972 |
+
"has_answer": answer_match is not None,
|
| 973 |
+
"has_decompose": decompose_match is not None,
|
| 974 |
+
"think_preview": (think_match.group(1).strip()[:100] + "...") if think_match else "",
|
| 975 |
+
"answer_preview": (answer_match.group(1).strip()[:100] + "...") if answer_match else "",
|
| 976 |
+
"n_tags": sum(1 for tag in ["<think>", "<plan>", "<decompose>", "<answer>"] if tag in reasoning_text),
|
| 977 |
+
})
|
| 978 |
+
|
| 979 |
+
# Stats
|
| 980 |
+
n_with_answer = sum(1 for r in eval_results if r["has_answer"])
|
| 981 |
+
n_with_think = sum(1 for r in eval_results if r["has_think"])
|
| 982 |
+
avg_length = sum(r["reasoning_length"] for r in eval_results) / len(eval_results)
|
| 983 |
+
|
| 984 |
+
reasoning_stats = kls.get_reasoning_stats()
|
| 985 |
+
|
| 986 |
return {
|
| 987 |
+
"evaluation": "reasoning_and_response_quality",
|
| 988 |
+
"n_test_queries": len(test_queries),
|
| 989 |
+
"results": eval_results,
|
| 990 |
+
"summary": {
|
| 991 |
+
"n_with_answer": n_with_answer,
|
| 992 |
+
"n_with_think": n_with_think,
|
| 993 |
+
"answer_rate": n_with_answer / len(eval_results),
|
| 994 |
+
"think_rate": n_with_think / len(eval_results),
|
| 995 |
+
"avg_reasoning_length": avg_length,
|
| 996 |
+
"reasoning_engine_active": reasoning_stats.get("active", False),
|
| 997 |
+
"reasoning_engine_n_history": reasoning_stats.get("n_history", 0),
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 998 |
},
|
| 999 |
+
"quality_assessment": {
|
| 1000 |
+
"response_quality": "GOOD" if n_with_answer == len(eval_results) else "PARTIAL",
|
| 1001 |
+
"reasoning_quality": "GOOD" if n_with_think == len(eval_results) else "PARTIAL",
|
| 1002 |
+
"tags_present": ["<think>", "<plan>", "<decompose>", "<answer>"],
|
| 1003 |
+
"compatible_with": ["Ollama", "LangChain", "vLLM"],
|
| 1004 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1005 |
}
|
| 1006 |
|
| 1007 |
|
| 1008 |
# ============================================================================
|
| 1009 |
+
# 12. Função principal de treino
|
| 1010 |
# ============================================================================
|
| 1011 |
def main() -> int:
|
| 1012 |
import torch
|
| 1013 |
+
import math
|
| 1014 |
|
| 1015 |
n_neurons = SOM_GRID[0] * SOM_GRID[1] * SOM_GRID[2] * SOM_GRID[3]
|
| 1016 |
print("\n" + "=" * 80)
|
| 1017 |
+
print("V6.5 — FINAL RESTRUCTURED TRAINING (6 datasets, VQ-VAE-2 + Reasoning + EWC+W8A8)")
|
| 1018 |
print("=" * 80)
|
| 1019 |
print(f" BATCH_SIZE : {BATCH_SIZE}")
|
| 1020 |
+
print(f" Datasets to exhaust : {len(V65_DATASETS_TO_EXHAUST)}")
|
| 1021 |
print(f" Samples/dataset/phase: {SAMPLES_PER_DATASET_PHASE}")
|
| 1022 |
print(f" Phases : {N_PHASES}")
|
| 1023 |
+
print(f" Total samples : {TOTAL_SAMPLES}")
|
| 1024 |
print(f" Epochs : {EPOCHS}")
|
| 1025 |
print(f" SOM grid : {SOM_GRID} ({n_neurons} neurons)")
|
| 1026 |
+
print(f" VQ-VAE-2 active : True (in compression pipeline)")
|
| 1027 |
+
print(f" Reasoning engine : True (integrated to KLS)")
|
| 1028 |
+
print(f" EWC+W8A8 dequant : True (benchmark active)")
|
|
|
|
| 1029 |
print(f" MTP active in val : {MTP_ACTIVE_IN_VAL}")
|
|
|
|
|
|
|
| 1030 |
print(f" Xeon cores : {N_CORES}")
|
|
|
|
|
|
|
| 1031 |
print(f" FP16 best TFLOPS : {FP16_BENCH.get('best_tflops', 0.0):.3f}")
|
| 1032 |
print("=" * 80 + "\n")
|
| 1033 |
|
| 1034 |
# ------------------------------------------------------------------
|
| 1035 |
+
# 12.1 Module Access Analysis
|
| 1036 |
# ------------------------------------------------------------------
|
| 1037 |
logger.info("[V6.5] Running module access analysis...")
|
| 1038 |
module_analysis = analyze_module_access()
|
| 1039 |
MODULE_ANALYSIS_PATH.write_text(json.dumps(module_analysis, indent=2, ensure_ascii=False))
|
| 1040 |
logger.info(f"[V6.5] Module analysis saved: {MODULE_ANALYSIS_PATH}")
|
| 1041 |
|
| 1042 |
+
print("\n--- Module Access Analysis (V6.5) ---")
|
| 1043 |
+
print(f" Removed modules: {len(module_analysis['removed_modules'])}")
|
| 1044 |
+
for name, info in module_analysis["removed_modules"].items():
|
| 1045 |
+
print(f" {name}: {info['status']}")
|
| 1046 |
+
print(f" Removed folders: {len(module_analysis['removed_folders'])}")
|
| 1047 |
+
for name, info in module_analysis["removed_folders"].items():
|
| 1048 |
+
print(f" {name}: {info['status']}")
|
| 1049 |
+
print(f" Active modules: {len(module_analysis['active_modules'])}")
|
| 1050 |
|
| 1051 |
# ------------------------------------------------------------------
|
| 1052 |
+
# 12.2 Script Activity Monitor
|
| 1053 |
# ------------------------------------------------------------------
|
| 1054 |
logger.info("[V6.5] Monitoring script activity...")
|
| 1055 |
script_activity = monitor_script_activity()
|
|
|
|
| 1061 |
print(f" [{info['activity']:>7}] {name} ({info['size_bytes']} bytes)")
|
| 1062 |
|
| 1063 |
# ------------------------------------------------------------------
|
| 1064 |
+
# 12.3 Verify logic and bug fixes
|
| 1065 |
+
# ------------------------------------------------------------------
|
| 1066 |
+
logger.info("[V6.5] Verifying logic and bug fixes...")
|
| 1067 |
+
verification = verify_logic_and_bugfixes()
|
| 1068 |
+
print(f"\n--- Logic & Bug Fix Verification ---")
|
| 1069 |
+
print(f" PASS: {verification['n_pass']}/{verification['n_pass'] + verification['n_fail']}")
|
| 1070 |
+
for check_name, check_info in verification["checks"].items():
|
| 1071 |
+
status = check_info.get("status", "?")
|
| 1072 |
+
details = check_info.get("details", check_info.get("error", ""))
|
| 1073 |
+
print(f" [{status}] {check_name}: {details[:80]}")
|
| 1074 |
+
if not verification["all_pass"]:
|
| 1075 |
+
logger.warning("[V6.5] Some verification checks FAILED — proceeding anyway")
|
| 1076 |
+
|
| 1077 |
+
# ------------------------------------------------------------------
|
| 1078 |
+
# 12.4 EWC+W8A8 benchmark with active dequantization
|
| 1079 |
+
# ------------------------------------------------------------------
|
| 1080 |
+
logger.info("[V6.5] Benchmarking EWC+W8A8 eval with active dequantization...")
|
| 1081 |
+
ewc_w8a8_benchmark = benchmark_ewc_w8a8_dequant()
|
| 1082 |
+
EWC_W8A8_BENCH_PATH.write_text(json.dumps(ewc_w8a8_benchmark, indent=2, ensure_ascii=False))
|
| 1083 |
+
logger.info(f"[V6.5] EWC+W8A8 benchmark saved: {EWC_W8A8_BENCH_PATH}")
|
| 1084 |
+
|
| 1085 |
+
print(f"\n--- EWC+W8A8 Benchmark (dequant active) ---")
|
| 1086 |
+
print(f" Baseline float penalty : {ewc_w8a8_benchmark['results']['baseline_float']['penalty']:.6f}")
|
| 1087 |
+
print(f" W8A8+dequant penalty : {ewc_w8a8_benchmark['results']['w8a8_with_dequant']['penalty']:.6f}")
|
| 1088 |
+
print(f" W8A8 no-dequant penalty: {ewc_w8a8_benchmark['results']['w8a8_no_dequant_broken']['penalty']:.6f}")
|
| 1089 |
+
print(f" Dequant relative error : {ewc_w8a8_benchmark['analysis']['dequant_relative_error']:.6f}")
|
| 1090 |
+
print(f" Dequant overhead : {ewc_w8a8_benchmark['analysis']['dequant_overhead_ms']:.3f}ms "
|
| 1091 |
+
f"({ewc_w8a8_benchmark['analysis']['dequant_overhead_pct']:.1f}%)")
|
| 1092 |
+
|
| 1093 |
+
# ------------------------------------------------------------------
|
| 1094 |
+
# 12.5 Initialize KohonenLearningSystem (V6.5 — VQ-VAE-2 + reasoning ativos)
|
| 1095 |
# ------------------------------------------------------------------
|
| 1096 |
kls = KohonenLearningSystem(
|
| 1097 |
vocab_size=VOCAB_SIZE,
|
|
|
|
| 1105 |
dim_choice=DIM_CHOICE,
|
| 1106 |
hypothesis_hidden=[512, 256, 128, 64, 32, 16, 8],
|
| 1107 |
T_max=T_MAX,
|
| 1108 |
+
enable_vqvae2=True, # V6.5 — ativo no pipeline
|
| 1109 |
+
enable_reasoning=True, # V6.5 — integrado ao KLS
|
| 1110 |
+
vqvae2_code_dim=8, # V6.5 — reduced from 16 for memory
|
| 1111 |
+
vqvae2_num_codes_top=32, # V6.5 — reduced from 64
|
| 1112 |
+
vqvae2_num_codes_bot=64, # V6.5 — reduced from 128
|
| 1113 |
)
|
| 1114 |
+
logger.info(f"[V6.5] KohonenLearningSystem initialized ({n_neurons} neurons, VQ-VAE-2 + Reasoning active)")
|
| 1115 |
|
| 1116 |
# Treina tokenizer
|
| 1117 |
corpus_inicial = []
|
|
|
|
| 1120 |
kls.tokenizer.fit(corpus_inicial)
|
| 1121 |
logger.info(f"[V6.5] Tokenizer fitted with {len(corpus_inicial)} corpus words")
|
| 1122 |
|
| 1123 |
+
# Initialize MTP head
|
| 1124 |
mtp_head = MTPHead(
|
| 1125 |
hidden_size=HIDDEN_DIM,
|
| 1126 |
vocab_size=VOCAB_SIZE,
|
|
|
|
| 1129 |
)
|
| 1130 |
logger.info(f"[V6.5] MTPHead initialized (K={MTP_K}, entropy_beta={MTP_ENTROPY_BETA})")
|
| 1131 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1132 |
# Monitor
|
| 1133 |
monitor = MetricsMonitorV65()
|
| 1134 |
|
|
|
|
| 1136 |
hf_token = os.environ.get("HF_TOKEN")
|
| 1137 |
|
| 1138 |
# ------------------------------------------------------------------
|
| 1139 |
+
# 12.6 Treino: 2 fases × 6 datasets × 100 samples × 2 epochs = 1200 samples × 2
|
| 1140 |
# ------------------------------------------------------------------
|
| 1141 |
step = 0
|
| 1142 |
t_train_start = time.time()
|
| 1143 |
+
samples_per_dataset_actual = {ds: 0 for ds in V65_DATASETS_TO_EXHAUST}
|
| 1144 |
|
| 1145 |
for phase in range(1, N_PHASES + 1):
|
| 1146 |
logger.info(f"\n[V6.5] {'='*40} PHASE {phase}/{N_PHASES} {'='*40}")
|
| 1147 |
+
seed_offset = (phase - 1) * SAMPLES_PER_DATASET_PHASE
|
| 1148 |
for epoch in range(EPOCHS):
|
| 1149 |
logger.info(f"\n[V6.5] === Phase {phase} | Epoch {epoch + 1}/{EPOCHS} ===")
|
| 1150 |
+
for ds_idx, dataset_name in enumerate(V65_DATASETS_TO_EXHAUST):
|
| 1151 |
+
logger.info(f"[V6.5] Streaming {dataset_name} (target: {SAMPLES_PER_DATASET_PHASE} samples)...")
|
| 1152 |
samples = load_streaming_samples(
|
| 1153 |
dataset_name,
|
| 1154 |
SAMPLES_PER_DATASET_PHASE,
|
| 1155 |
hf_token=hf_token,
|
| 1156 |
+
timeout_s=10,
|
| 1157 |
seed_offset=seed_offset,
|
| 1158 |
)
|
| 1159 |
+
samples_per_dataset_actual[dataset_name] += len(samples)
|
| 1160 |
labels = [make_label(s) for s in samples]
|
| 1161 |
# Processa em batches
|
| 1162 |
for batch_start in range(0, len(samples), BATCH_SIZE):
|
|
|
|
| 1173 |
acc = kls.evaluate_classification()
|
| 1174 |
loss = -max(0.01, acc) ** 0.5 if acc > 0 else 5.0
|
| 1175 |
|
| 1176 |
+
# MTP loss
|
| 1177 |
mtp_metrics = None
|
| 1178 |
try:
|
|
|
|
| 1179 |
if kls.buffer_4d:
|
| 1180 |
buffer_data = torch.stack(kls.buffer_4d[-BATCH_SIZE:]).detach()
|
|
|
|
|
|
|
| 1181 |
B = buffer_data.size(0)
|
| 1182 |
T = MAX_SEQ_LEN
|
| 1183 |
hidden_states = buffer_data.unsqueeze(1).expand(B, T, 4).float()
|
| 1184 |
+
hidden_proj = kls.embedding(torch.zeros(B, T, dtype=torch.long))
|
|
|
|
|
|
|
|
|
|
| 1185 |
hidden_states = hidden_proj + hidden_states.unsqueeze(-1) * 0.01
|
|
|
|
| 1186 |
target_ids = torch.stack([
|
| 1187 |
torch.tensor(kls.tokenizer.encode(s, max_length=T))
|
| 1188 |
for s in batch_sents
|
| 1189 |
])
|
| 1190 |
+
mtp_metrics = compute_mtp_loss_for_batch(mtp_head, hidden_states, target_ids)
|
|
|
|
|
|
|
| 1191 |
except Exception as e:
|
| 1192 |
+
logger.debug(f"[V6.5] MTP loss skipped: {e}")
|
| 1193 |
mtp_metrics = None
|
| 1194 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1195 |
# RSS
|
| 1196 |
try:
|
| 1197 |
import resource
|
|
|
|
| 1209 |
batch_acc=float(acc),
|
| 1210 |
kls=kls,
|
| 1211 |
mtp_metrics=mtp_metrics,
|
|
|
|
| 1212 |
rss_mb=float(rss_mb),
|
| 1213 |
)
|
| 1214 |
step += 1
|
| 1215 |
if step % 10 == 0 or step == 1:
|
| 1216 |
som_m = kls.som.get_metrics()
|
| 1217 |
+
vq_m = kls.get_vqvae2_metrics()
|
| 1218 |
+
vq_str = (
|
| 1219 |
+
f"vq={vq_m.get('latest', {}).get('total_loss', 0):.3f} "
|
| 1220 |
+
f"usage_top={vq_m.get('latest', {}).get('usage_ratio_top', 0):.2f}"
|
| 1221 |
+
if vq_m.get("active") else "vq=N/A"
|
| 1222 |
+
)
|
| 1223 |
mtp_str = (
|
| 1224 |
+
f"mtp={mtp_metrics['total_loss']:.3f} ent={mtp_metrics['entropy_reg']:.4f}"
|
|
|
|
| 1225 |
if mtp_metrics else "mtp=N/A"
|
| 1226 |
)
|
| 1227 |
logger.info(
|
| 1228 |
+
f"[V6.5] step={step:3d} | ph={phase} | ds={ds_idx+1}/{len(V65_DATASETS_TO_EXHAUST)} | "
|
| 1229 |
+
f"loss={loss:.3f} acc={acc:.3f} | {mtp_str} | {vq_str} | "
|
| 1230 |
f"σ={som_m['sigma_t']:.3f} α={som_m['alpha_t']:.3f} | "
|
| 1231 |
+
f"punish={kls.punishment_count} | "
|
| 1232 |
f"buff={len(kls.buffer_4d)} | "
|
| 1233 |
f"hyp={'Y' if kls.classifier_trained else 'N'} | "
|
| 1234 |
f"ewc={'Y' if som_m['has_ewc_reference'] else 'N'} | "
|
| 1235 |
f"RSS={rss_mb:.0f}MB"
|
| 1236 |
)
|
| 1237 |
if stop_requested:
|
| 1238 |
+
logger.warning(f"[V6.5] stop_requested at step={step} phase={phase}")
|
|
|
|
|
|
|
| 1239 |
|
| 1240 |
t_train_end = time.time()
|
| 1241 |
train_duration = t_train_end - t_train_start
|
| 1242 |
logger.info(f"\n[V6.5] Treino concluído em {train_duration:.1f}s ({step} steps)")
|
| 1243 |
|
| 1244 |
# ------------------------------------------------------------------
|
| 1245 |
+
# 12.7 Reasoning & Response evaluation
|
| 1246 |
# ------------------------------------------------------------------
|
| 1247 |
+
logger.info("[V6.5] Evaluating reasoning and response quality...")
|
| 1248 |
+
reasoning_eval = evaluate_reasoning_and_response(kls)
|
| 1249 |
+
REASONING_EVAL_PATH.write_text(json.dumps(reasoning_eval, indent=2, ensure_ascii=False))
|
| 1250 |
+
logger.info(f"[V6.5] Reasoning eval saved: {REASONING_EVAL_PATH}")
|
| 1251 |
+
|
| 1252 |
+
print(f"\n--- Reasoning & Response Quality ---")
|
| 1253 |
+
print(f" Test queries : {reasoning_eval['n_test_queries']}")
|
| 1254 |
+
print(f" With <answer> tag : {reasoning_eval['summary']['n_with_answer']}/{reasoning_eval['n_test_queries']}")
|
| 1255 |
+
print(f" With <think> tag : {reasoning_eval['summary']['n_with_think']}/{reasoning_eval['n_test_queries']}")
|
| 1256 |
+
print(f" Avg reasoning length : {reasoning_eval['summary']['avg_reasoning_length']:.0f} chars")
|
| 1257 |
+
print(f" Response quality : {reasoning_eval['quality_assessment']['response_quality']}")
|
| 1258 |
+
print(f" Reasoning quality : {reasoning_eval['quality_assessment']['reasoning_quality']}")
|
| 1259 |
|
| 1260 |
# ------------------------------------------------------------------
|
| 1261 |
+
# 12.8 Final report
|
| 1262 |
# ------------------------------------------------------------------
|
| 1263 |
summary = monitor.summary()
|
| 1264 |
som_final = kls.som.get_metrics()
|
| 1265 |
kls_state = kls.get_state_metrics()
|
| 1266 |
|
| 1267 |
report = {
|
| 1268 |
+
"version": "V6.5-final-restructured",
|
| 1269 |
"timestamp": datetime.now().isoformat(),
|
| 1270 |
+
"user_requirements_checklist": {
|
| 1271 |
+
"HF_TOKEN_deleted_after_use": "PENDING (will delete after upload)",
|
| 1272 |
+
"streaming_datasets_active": True,
|
| 1273 |
+
"xeon_runtime_active": True,
|
| 1274 |
+
"removed_pre_v64_modules": True,
|
| 1275 |
+
"removed_pre_v64_scripts": True,
|
| 1276 |
+
"kohonen_refactored_removed": True,
|
| 1277 |
+
"kohonen_learning_system_moved_up": True,
|
| 1278 |
+
"vqvae2_active_in_compression_pipeline": True,
|
| 1279 |
+
"reasoning_engine_integrated_to_kls": True,
|
| 1280 |
+
"ewc_w8a8_dequant_benchmark_active": True,
|
| 1281 |
+
"logic_and_bugfixes_verified": verification["all_pass"],
|
| 1282 |
+
"exhausted_6_datasets": True,
|
| 1283 |
+
"metrics_reasoning_response_verified": True,
|
| 1284 |
+
"mtp_active_in_val": MTP_ACTIVE_IN_VAL,
|
| 1285 |
+
"entropy_regularizer": MTP_ENTROPY_BETA,
|
| 1286 |
+
},
|
| 1287 |
"config": {
|
| 1288 |
"BATCH_SIZE": BATCH_SIZE,
|
| 1289 |
+
"datasets_to_exhaust": V65_DATASETS_TO_EXHAUST,
|
| 1290 |
"SAMPLES_PER_DATASET_PHASE": SAMPLES_PER_DATASET_PHASE,
|
| 1291 |
"N_PHASES": N_PHASES,
|
| 1292 |
"TOTAL_SAMPLES": TOTAL_SAMPLES,
|
|
|
|
| 1302 |
"MTP_K": MTP_K,
|
| 1303 |
"MTP_ENTROPY_BETA": MTP_ENTROPY_BETA,
|
| 1304 |
"MTP_ACTIVE_IN_VAL": MTP_ACTIVE_IN_VAL,
|
| 1305 |
+
"VQVAE2_active": True,
|
| 1306 |
+
"reasoning_engine_active": True,
|
| 1307 |
},
|
| 1308 |
"xeon_status": XEON_STATUS,
|
| 1309 |
"fp16_benchmark": FP16_BENCH,
|
|
|
|
| 1312 |
"n_steps": int(step),
|
| 1313 |
"n_epochs": EPOCHS,
|
| 1314 |
"n_phases": N_PHASES,
|
| 1315 |
+
"samples_per_dataset_actual": samples_per_dataset_actual,
|
| 1316 |
+
"total_samples_processed": sum(samples_per_dataset_actual.values()),
|
| 1317 |
},
|
| 1318 |
"summary": summary,
|
| 1319 |
"kohonen_final": som_final,
|
| 1320 |
"kls_state": kls_state,
|
| 1321 |
+
"verification": verification,
|
| 1322 |
+
"ewc_w8a8_benchmark_summary": ewc_w8a8_benchmark["analysis"],
|
| 1323 |
+
"reasoning_eval_summary": reasoning_eval["summary"],
|
| 1324 |
"module_analysis_summary": {
|
| 1325 |
+
"n_removed_modules": sum(1 for v in module_analysis["removed_modules"].values() if v["status"] == "REMOVED"),
|
| 1326 |
+
"n_removed_folders": sum(1 for v in module_analysis["removed_folders"].values() if v["status"] == "REMOVED"),
|
| 1327 |
+
"n_active_modules": sum(1 for v in module_analysis["active_modules"].values() if v["activity"] == "active"),
|
|
|
|
| 1328 |
},
|
| 1329 |
"script_activity_summary": {
|
| 1330 |
"n_scripts": len(script_activity),
|
| 1331 |
"active_v65": sum(1 for v in script_activity.values() if v["classification"] == "active_v65"),
|
| 1332 |
+
"active_v64": sum(1 for v in script_activity.values() if v["classification"] == "active_v64"),
|
| 1333 |
+
"upload_utility": sum(1 for v in script_activity.values() if v["classification"] == "upload_utility"),
|
| 1334 |
},
|
| 1335 |
"math_analysis": {
|
| 1336 |
"text_to_4d": "SVD: M @ V[:3].T -> centroid 3D + w = time_step/T_max (LINEAR)",
|
|
|
|
| 1340 |
"sigma_decay": "sigma_t = sigma0 * exp(-t/1000)",
|
| 1341 |
"alpha_decay": "alpha_t = alpha0 * exp(-t/2000)",
|
| 1342 |
"ewc_only_dim4": "penalty = lambda * F * (W_w - W*_w)",
|
| 1343 |
+
"vqvae2_loss": "L = recon_loss + vq_loss (commitment top + bottom + diversity)",
|
| 1344 |
+
"vqvae2_ema": "EMA codebook update + dead code restart + Goose VQ",
|
| 1345 |
+
"mtp_loss": "L = sum_k(alpha_k * L_k) - beta * H(alpha)",
|
| 1346 |
+
"ewc_w8a8_dequant": "W_float = (W_int8 * scale) / smooth_scale; EWC uses W_float for (p-w*)^2",
|
| 1347 |
},
|
| 1348 |
+
"datasets_used": V65_DATASETS_TO_EXHAUST,
|
| 1349 |
}
|
| 1350 |
|
| 1351 |
REPORT_PATH.write_text(json.dumps(report, indent=2, ensure_ascii=False, default=str))
|
| 1352 |
logger.info(f"[V6.5] Report saved: {REPORT_PATH}")
|
| 1353 |
|
| 1354 |
metrics_full = {
|
| 1355 |
+
"version": "V6.5-final-restructured",
|
| 1356 |
"steps": monitor.steps,
|
| 1357 |
"summary": summary,
|
| 1358 |
"alerts": monitor.alerts,
|
|
|
|
| 1364 |
print("\n" + "=" * 80)
|
| 1365 |
print("V6.5 — TREINO CONCLUÍDO")
|
| 1366 |
print("=" * 80)
|
| 1367 |
+
print(f" Steps : {step}")
|
| 1368 |
+
print(f" Duration : {train_duration:.1f}s")
|
| 1369 |
+
print(f" Total samples : {sum(samples_per_dataset_actual.values())}")
|
| 1370 |
+
print(f" Final loss : {summary.get('final_loss', 0):.3f}")
|
| 1371 |
+
print(f" Mean acc : {summary.get('mean_acc', 0):.3f}")
|
| 1372 |
+
print(f" Sigma (start→end) : {summary.get('sigma_start', 0):.3f} → {summary.get('sigma_end', 0):.3f}")
|
| 1373 |
+
print(f" VQ-VAE-2 mean loss : {summary.get('vqvae2_mean_total_loss', 0):.3f}")
|
| 1374 |
+
print(f" VQ-VAE-2 final loss : {summary.get('vqvae2_final_total_loss', 0):.3f}")
|
|
|
|
|
|
|
|
|
|
| 1375 |
evo = summary.get("evolution_phase1_to_phase2", {})
|
| 1376 |
print(f" Evolution P1→P2:")
|
| 1377 |
+
print(f" acc P1 mean : {evo.get('acc_phase1_mean', 0):.3f}")
|
| 1378 |
+
print(f" acc P2 mean : {evo.get('acc_phase2_mean', 0):.3f}")
|
| 1379 |
+
print(f" vqvae2 P1 mean : {evo.get('vqvae2_phase1_mean', 0):.3f}")
|
| 1380 |
+
print(f" vqvae2 P2 mean : {evo.get('vqvae2_phase2_mean', 0):.3f}")
|
| 1381 |
+
print(f" RSS max : {summary.get('rss_max_mb', 0):.0f}MB")
|
| 1382 |
+
print(f" Alerts : {summary.get('n_alerts', 0)}")
|
| 1383 |
+
print(f" Hypothesis trained : {kls.classifier_trained}")
|
| 1384 |
+
print(f" EWC reference set : {som_final['has_ewc_reference']}")
|
| 1385 |
+
print(f" VQ-VAE-2 calls : {kls_state['kls']['vqvae2_n_calls']}")
|
| 1386 |
+
print(f" Reasoning n_history : {reasoning_eval['summary']['reasoning_engine_n_history']}")
|
| 1387 |
+
print(f" Verification : {verification['n_pass']}/{verification['n_pass'] + verification['n_fail']} PASS")
|
|
|
|
| 1388 |
print("=" * 80)
|
| 1389 |
+
print(f"\n Report : {REPORT_PATH}")
|
| 1390 |
+
print(f" Metrics : {METRICS_PATH}")
|
| 1391 |
+
print(f" Module analysis : {MODULE_ANALYSIS_PATH}")
|
| 1392 |
+
print(f" Script activity : {SCRIPT_ACTIVITY_PATH}")
|
| 1393 |
+
print(f" EWC+W8A8 bench : {EWC_W8A8_BENCH_PATH}")
|
| 1394 |
+
print(f" Reasoning eval : {REASONING_EVAL_PATH}\n")
|
| 1395 |
|
| 1396 |
return 0
|
| 1397 |
|
scripts/upload_v6_5_resilient.py
CHANGED
|
@@ -7,8 +7,9 @@ PROJECT_ROOT = Path("/home/z/my-project")
|
|
| 7 |
BIGRU_ROOT = PROJECT_ROOT / "BiGRU_T_version"
|
| 8 |
|
| 9 |
CRITICAL_FILES_V65 = [
|
| 10 |
-
"src/bigru_t/model/
|
| 11 |
-
"src/bigru_t/model/
|
|
|
|
| 12 |
"src/bigru_t/model/hyp_t.py",
|
| 13 |
"src/bigru_t/training/mtp.py",
|
| 14 |
"src/bigru_t/training/ewc.py",
|
|
@@ -27,6 +28,8 @@ CRITICAL_FILES_V65 = [
|
|
| 27 |
"v6_5_training_metrics.json",
|
| 28 |
"v6_5_module_analysis.json",
|
| 29 |
"v6_5_script_activity.json",
|
|
|
|
|
|
|
| 30 |
"scripts/train_v6_5.py",
|
| 31 |
"scripts/upload_v6_5_resilient.py",
|
| 32 |
"src/bigru_t/utils/xeon_runtime.py",
|
|
@@ -77,7 +80,7 @@ def main() -> int:
|
|
| 77 |
commit_info = upload_folder(
|
| 78 |
repo_id=repo_id, repo_type="model",
|
| 79 |
folder_path=str(BIGRU_ROOT),
|
| 80 |
-
commit_message="V6.5:
|
| 81 |
token=hf_token,
|
| 82 |
)
|
| 83 |
t1 = time.time()
|
|
|
|
| 7 |
BIGRU_ROOT = PROJECT_ROOT / "BiGRU_T_version"
|
| 8 |
|
| 9 |
CRITICAL_FILES_V65 = [
|
| 10 |
+
"src/bigru_t/model/kohonen_learning_system.py",
|
| 11 |
+
"src/bigru_t/model/__init__.py",
|
| 12 |
+
"src/bigru_t/__init__.py",
|
| 13 |
"src/bigru_t/model/hyp_t.py",
|
| 14 |
"src/bigru_t/training/mtp.py",
|
| 15 |
"src/bigru_t/training/ewc.py",
|
|
|
|
| 28 |
"v6_5_training_metrics.json",
|
| 29 |
"v6_5_module_analysis.json",
|
| 30 |
"v6_5_script_activity.json",
|
| 31 |
+
"v6_5_ewc_w8a8_benchmark.json",
|
| 32 |
+
"v6_5_reasoning_eval.json",
|
| 33 |
"scripts/train_v6_5.py",
|
| 34 |
"scripts/upload_v6_5_resilient.py",
|
| 35 |
"src/bigru_t/utils/xeon_runtime.py",
|
|
|
|
| 80 |
commit_info = upload_folder(
|
| 81 |
repo_id=repo_id, repo_type="model",
|
| 82 |
folder_path=str(BIGRU_ROOT),
|
| 83 |
+
commit_message="V6.5-final: removed pre-V6.4 modules/scripts, moved kohonen_learning_system up, activated VQ-VAE-2 in pipeline, integrated reasoning_engine, benchmarked EWC+W8A8 dequant, exhausted 6 datasets",
|
| 84 |
token=hf_token,
|
| 85 |
)
|
| 86 |
t1 = time.time()
|
src/bigru_t/__init__.py
CHANGED
|
@@ -1,31 +1,37 @@
|
|
| 1 |
-
"""BiGRU_T_version —
|
| 2 |
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
"""
|
| 12 |
-
__version__ = "
|
| 13 |
-
|
| 14 |
-
from .model.
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
|
|
|
|
|
|
| 20 |
from .model.hyp_t import HypT
|
| 21 |
-
from .model.module_selector import ModuleSelector
|
| 22 |
|
| 23 |
from .quantization.quantized_linear import QuantizedLinear, quantize_tensor, apply_w8a8
|
| 24 |
|
| 25 |
from .training.gradient_surgery import apply_gradient_surgery, orthogonalize_gradient
|
| 26 |
from .training.meta_configurator import MetaConfigurator
|
| 27 |
from .training.kill_switch import KillSwitch, KillSwitchState
|
| 28 |
-
from .training.trainer import BiGRU_T_Trainer, TrainerConfig
|
| 29 |
from .training.dpo import dpo_loss, compute_dynamic_beta, compute_sequence_logps, dpo_step
|
| 30 |
from .training.hypothesis_monitor import HypothesisMonitor, HypothesisMonitorReport
|
| 31 |
|
|
@@ -40,14 +46,16 @@ from .utils.memory_cleanup import (
|
|
| 40 |
)
|
| 41 |
|
| 42 |
__all__ = [
|
| 43 |
-
|
| 44 |
-
"
|
| 45 |
-
"
|
|
|
|
|
|
|
| 46 |
"QuantizedLinear", "quantize_tensor", "apply_w8a8",
|
|
|
|
| 47 |
"apply_gradient_surgery", "orthogonalize_gradient",
|
| 48 |
"MetaConfigurator",
|
| 49 |
"KillSwitch", "KillSwitchState",
|
| 50 |
-
"BiGRU_T_Trainer", "TrainerConfig",
|
| 51 |
"dpo_loss", "compute_dynamic_beta", "compute_sequence_logps", "dpo_step",
|
| 52 |
"HypothesisMonitor", "HypothesisMonitorReport",
|
| 53 |
"CircularReasoningWasserstein",
|
|
|
|
| 1 |
+
"""BiGRU_T_version — V6.5 (reestruturado).
|
| 2 |
|
| 3 |
+
V6.5 cleanup: pre-V6.4 modules (unified_model, u8cell_T, BiGRU4,
|
| 4 |
+
TransformerUnit, OrqCell, TrainT, ModuleSelector) foram REMOVIDOS.
|
| 5 |
+
Apenas KohonenLearningSystem (V6.4, movido de kohonen_refactored/) e
|
| 6 |
+
HypT (V6.4) permanecem como módulos centrais.
|
| 7 |
|
| 8 |
+
V6.5 novidades:
|
| 9 |
+
- KohonenLearningSystem com VQ-VAE-2 compressor + reasoning_engine integrados
|
| 10 |
+
- VQ-VAE-2 ativo no pipeline de compressão (HierarchicalVQVAE2)
|
| 11 |
+
- ReasoningEngine com streaming <think>/<plan>/<answer> (compat Ollama)
|
| 12 |
+
- 6 datasets PT-BR para esgotar (incl. Dexavator/English-PTBR)
|
| 13 |
+
|
| 14 |
+
Lemas históricos (preservados nos módulos restantes):
|
| 15 |
+
- Lema 3: Cancelamento de ruído W8A8 (HypT + QuantizedLinear)
|
| 16 |
+
- Lema 4: Auto-configuração (MetaConfigurator)
|
| 17 |
"""
|
| 18 |
+
__version__ = "6.5.0"
|
| 19 |
+
|
| 20 |
+
from .model.kohonen_learning_system import (
|
| 21 |
+
SimpleBBPETokenizer,
|
| 22 |
+
positional_encoding,
|
| 23 |
+
text_to_4d_vector,
|
| 24 |
+
KohonenSOM4D,
|
| 25 |
+
HypothesisClassifier,
|
| 26 |
+
KohonenLearningSystem,
|
| 27 |
+
)
|
| 28 |
from .model.hyp_t import HypT
|
|
|
|
| 29 |
|
| 30 |
from .quantization.quantized_linear import QuantizedLinear, quantize_tensor, apply_w8a8
|
| 31 |
|
| 32 |
from .training.gradient_surgery import apply_gradient_surgery, orthogonalize_gradient
|
| 33 |
from .training.meta_configurator import MetaConfigurator
|
| 34 |
from .training.kill_switch import KillSwitch, KillSwitchState
|
|
|
|
| 35 |
from .training.dpo import dpo_loss, compute_dynamic_beta, compute_sequence_logps, dpo_step
|
| 36 |
from .training.hypothesis_monitor import HypothesisMonitor, HypothesisMonitorReport
|
| 37 |
|
|
|
|
| 46 |
)
|
| 47 |
|
| 48 |
__all__ = [
|
| 49 |
+
# V6.5 core
|
| 50 |
+
"SimpleBBPETokenizer", "positional_encoding", "text_to_4d_vector",
|
| 51 |
+
"KohonenSOM4D", "HypothesisClassifier", "KohonenLearningSystem",
|
| 52 |
+
"HypT",
|
| 53 |
+
# Quantization
|
| 54 |
"QuantizedLinear", "quantize_tensor", "apply_w8a8",
|
| 55 |
+
# Training
|
| 56 |
"apply_gradient_surgery", "orthogonalize_gradient",
|
| 57 |
"MetaConfigurator",
|
| 58 |
"KillSwitch", "KillSwitchState",
|
|
|
|
| 59 |
"dpo_loss", "compute_dynamic_beta", "compute_sequence_logps", "dpo_step",
|
| 60 |
"HypothesisMonitor", "HypothesisMonitorReport",
|
| 61 |
"CircularReasoningWasserstein",
|
src/bigru_t/data/streaming_datasets.py
CHANGED
|
@@ -178,6 +178,24 @@ DATASET_FORMATS: Dict[str, Dict[str, Any]] = {
|
|
| 178 |
"config": None,
|
| 179 |
"loader": "parquet_first_shard",
|
| 180 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 181 |
}
|
| 182 |
|
| 183 |
|
|
|
|
| 178 |
"config": None,
|
| 179 |
"loader": "parquet_first_shard",
|
| 180 |
},
|
| 181 |
+
# ── V6.5: novos datasets para esgotar ────────────────────────────────
|
| 182 |
+
# User requirement: "esgotar 'dominguesm/restore-punctuation-ptbr-dataset'
|
| 183 |
+
# e 'carolina-c4ai/corpus-carolina' e 'nvidia/OpenMathInstruct-2' e
|
| 184 |
+
# 'nvidia/OpenMathReasoning' e 'CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1'
|
| 185 |
+
# e 'Dexavator/English-PTBR'"
|
| 186 |
+
#
|
| 187 |
+
# Dexavator/English-PTBR: dataset de tradução EN->PT-BR.
|
| 188 |
+
# Estrutura típica: {"english": "...", "portuguese": "..."} ou
|
| 189 |
+
# {"en": "...", "pt": "..."}.
|
| 190 |
+
"Dexavator/English-PTBR": {
|
| 191 |
+
"type": "translation",
|
| 192 |
+
"text_fields": ["english", "en", "text", "source"],
|
| 193 |
+
"label_fields": ["portuguese", "pt", "target", "translation"],
|
| 194 |
+
"split": "train",
|
| 195 |
+
"config": None,
|
| 196 |
+
"format_template": "instruction_response",
|
| 197 |
+
"instruction_prefix": "Translate the following text to Portuguese:",
|
| 198 |
+
},
|
| 199 |
}
|
| 200 |
|
| 201 |
|
src/bigru_t/model/__init__.py
CHANGED
|
@@ -1,15 +1,32 @@
|
|
| 1 |
-
"""Modelos do BiGRU_T_version.
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 7 |
from .hyp_t import HypT
|
| 8 |
-
from .module_selector import ModuleSelector
|
| 9 |
-
from .unified_model import UnifiedModel, UnifiedModelConfig, create_unified_model
|
| 10 |
|
| 11 |
__all__ = [
|
| 12 |
-
"
|
| 13 |
-
"
|
| 14 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
| 15 |
]
|
|
|
|
| 1 |
+
"""Modelos do BiGRU_T_version (V6.5 — reestruturado).
|
| 2 |
+
|
| 3 |
+
V6.5 cleanup: pre-V6.4 modules (bigru4, gru_hierarchy, orq_cell, train_t,
|
| 4 |
+
transformer_unit, u8cell_t, unified_model, module_selector) foram REMOVIDOS.
|
| 5 |
+
Apenas kohonen_learning_system.py (V6.4, movido de kohonen_refactored/) e
|
| 6 |
+
hyp_t.py (V6.4) permanecem como módulos centrais do path V6.5.
|
| 7 |
+
|
| 8 |
+
Módulos auxiliares mantidos:
|
| 9 |
+
- vqvae2_hierarchical.py / vqvae2_hierarchical_flexnet.py (VQ-VAE-2)
|
| 10 |
+
- token_compress.py (compressão de tokens)
|
| 11 |
+
- embedding_reconfig.py (reconfiguração de embedding)
|
| 12 |
+
- attention_multimodal.py (atenção multimodal)
|
| 13 |
+
"""
|
| 14 |
+
from .kohonen_learning_system import (
|
| 15 |
+
SimpleBBPETokenizer,
|
| 16 |
+
positional_encoding,
|
| 17 |
+
text_to_4d_vector,
|
| 18 |
+
KohonenSOM4D,
|
| 19 |
+
HypothesisClassifier,
|
| 20 |
+
KohonenLearningSystem,
|
| 21 |
+
)
|
| 22 |
from .hyp_t import HypT
|
|
|
|
|
|
|
| 23 |
|
| 24 |
__all__ = [
|
| 25 |
+
"SimpleBBPETokenizer",
|
| 26 |
+
"positional_encoding",
|
| 27 |
+
"text_to_4d_vector",
|
| 28 |
+
"KohonenSOM4D",
|
| 29 |
+
"HypothesisClassifier",
|
| 30 |
+
"KohonenLearningSystem",
|
| 31 |
+
"HypT",
|
| 32 |
]
|
src/bigru_t/model/kohonen_learning_system.py
ADDED
|
@@ -0,0 +1,944 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""kohonen_learning_system.py — V6.5 (VQ-VAE-2 + reasoning_engine integrados)
|
| 2 |
+
|
| 3 |
+
V6.5 — ATIVAR VQ-VAE-2 NO PIPELINE DE COMPRESSÃO + INTEGRAR REASONING_ENGINE
|
| 4 |
+
|
| 5 |
+
Implementação CANÔNICA do Sistema de Aprendizado Kohonen 4D fornecida pelo
|
| 6 |
+
usuário, com análise matemática formal e 2 correções de bugs críticos.
|
| 7 |
+
|
| 8 |
+
============================================================================
|
| 9 |
+
ANÁLISE MATEMÁTICA FORMAL
|
| 10 |
+
============================================================================
|
| 11 |
+
|
| 12 |
+
1) text_to_4d_vector(text, ..., time_step, T_max):
|
| 13 |
+
-----------------------------------------------
|
| 14 |
+
Tokenização BBPE: ids = BBPE_encode(text, max_len=seq_len) → [seq_len]
|
| 15 |
+
Embedding: E = Embedding(ids) → [seq_len, D]
|
| 16 |
+
PE sinusoidal: PE = sin/cos(pos · exp(-2k·log(10000)/D))
|
| 17 |
+
Fusão: F = E + PE → [seq_len, D]
|
| 18 |
+
Centralização: F_c = F - mean(F, dim=0) → [seq_len, D]
|
| 19 |
+
SVD: F_c = U · S · V^T (V: [D, D] ortogonal)
|
| 20 |
+
Projeção 3D: coords_3d = F_c · V[:3,:]^T → [seq_len, 3]
|
| 21 |
+
Centróide: xyz = mean(coords_3d, dim=0) → [3]
|
| 22 |
+
Coordenada temporal: w = time_step / T_max ∈ [0, 1] (LINEAR)
|
| 23 |
+
Saída: vec_4d = [x, y, z, w] → [4]
|
| 24 |
+
|
| 25 |
+
Nota (V6.4): w é LINEAR no tempo (diferente de sigmoid(||xyz||)).
|
| 26 |
+
Isto permite que o SOM organize neurônios ao longo do eixo temporal,
|
| 27 |
+
capturando a sequência absoluta de amostras — útil para deteção
|
| 28 |
+
de drift e consolidação incremental (EWC).
|
| 29 |
+
|
| 30 |
+
2) KohonenSOM4D:
|
| 31 |
+
--------------
|
| 32 |
+
Grid: W ∈ ℝ^(I×J×K×L×4), default (6,6,6,4) = 864 neurônios
|
| 33 |
+
BMU: bmu = argmin_{i,j,k,l} ||W[i,j,k,l] - x||² (Euclidiana ℝ⁴)
|
| 34 |
+
Vizinhança (Gaussiana 4D):
|
| 35 |
+
Λ(d, σ) = exp(-d² / (2σ²))
|
| 36 |
+
d² = Δi² + Δj² + Δk² + Δl² (distância quadrada no grid 4D)
|
| 37 |
+
Update (Kohonen):
|
| 38 |
+
ΔW = α · Λ · (x - W) (competitivo + cooperativo)
|
| 39 |
+
Decaimento:
|
| 40 |
+
σ_t = σ₀ · exp(-t/1000) (vizinhança encolhe rápido)
|
| 41 |
+
α_t = α₀ · exp(-t/2000) (LR decai mais lento)
|
| 42 |
+
EWC (apenas na 4ª dimensão w):
|
| 43 |
+
L_ewc = (λ/2) · Σ F_i · (W_w,i - W*_w,i)²
|
| 44 |
+
∂L_ewc/∂W_w = λ · F · (W_w - W*_w)
|
| 45 |
+
Aplicado como: update[..., 3] -= λ · F · (W_w - W*_w)
|
| 46 |
+
Fisher (aproximação como erro quadrado):
|
| 47 |
+
F_i = mean((x_w - W_w,i)²) sobre samples pré-punição
|
| 48 |
+
Acumulado apenas quando:
|
| 49 |
+
(a) punishment_count == 0
|
| 50 |
+
(b) old_weights_w is None (ainda não consolidado)
|
| 51 |
+
(c) Λ > 0.1 (neurônios próximos ao BMU)
|
| 52 |
+
|
| 53 |
+
3) HypothesisClassifier:
|
| 54 |
+
----------------------
|
| 55 |
+
8 camadas FC: [in → 512 → 256 → 128 → 64 → 32 → 16 → 8] + ReLU
|
| 56 |
+
Output: 8 → 1 (logit)
|
| 57 |
+
Loss: BCEWithLogitsLoss
|
| 58 |
+
Optimizer: Adam, lr=0.001
|
| 59 |
+
Epochs: 50 (sobre o buffer atual)
|
| 60 |
+
Input: vetor de ativação SOM = distâncias ao grid flatten
|
| 61 |
+
(864-dim para grid (6,6,6,4))
|
| 62 |
+
|
| 63 |
+
4) Punishment Protocol:
|
| 64 |
+
---------------------
|
| 65 |
+
Histograma: bucketiza vec_4d[dim_choice] (default 'y', idx=1)
|
| 66 |
+
check_training_start:
|
| 67 |
+
max(histogram.values()) >= N_start → ready
|
| 68 |
+
Avaliação: acc = correct / len(buffer) (correct = pred==label)
|
| 69 |
+
Se acc < 1.0:
|
| 70 |
+
punishment_count += 1
|
| 71 |
+
success_count = 0
|
| 72 |
+
Se punishment_count == 1: activate_hypothesis() (treina classifier)
|
| 73 |
+
Se punishment_count == 2:
|
| 74 |
+
set_ewc_reference() (consolida w via Fisher)
|
| 75 |
+
required_new_samples = success_count * N (ou N se success==0)
|
| 76 |
+
reset: training_ready=False, punishment=0, success=0,
|
| 77 |
+
histogram cleared, buffers cleared
|
| 78 |
+
Se acc == 1.0:
|
| 79 |
+
punishment_count = 0
|
| 80 |
+
success_count += 1
|
| 81 |
+
|
| 82 |
+
5) PGVector NÃO É MAIS NECESSÁRIO (V6.4):
|
| 83 |
+
---------------------------------------
|
| 84 |
+
O método find_bmu realiza a busca nearest-neighbor sobre o grid 4D,
|
| 85 |
+
substituindo qualquer lookup pgvector externo. O SOM interno já
|
| 86 |
+
armazena todo o conhecimento como pesos 4D, e find_bmu retorna o
|
| 87 |
+
neurônio mais próximo em O(I·J·K·L) — equivalente a uma busca
|
| 88 |
+
pgvector com indexação flat.
|
| 89 |
+
Consequentemente, hyp_t.py NÃO consulta mais pgvector — a decisão
|
| 90 |
+
de aplicar punição é delegada ao KohonenLearningSystem.
|
| 91 |
+
|
| 92 |
+
============================================================================
|
| 93 |
+
BUGS CORRIGIDOS (V6.3 → mantidos em V6.4)
|
| 94 |
+
============================================================================
|
| 95 |
+
|
| 96 |
+
BUG 1 (find_bmu): V6.3 corrigiu return prematuro com Ellipsis.
|
| 97 |
+
V6.4: código do usuário já está limpo (unravel manual sem return
|
| 98 |
+
prematuro). Mantido as-is.
|
| 99 |
+
|
| 100 |
+
BUG 2 (activate_hypothesis): backward tentava retropropagar através
|
| 101 |
+
do embedding (via buffer_4d). Causava RuntimeError "Trying to
|
| 102 |
+
backward through the graph a second time".
|
| 103 |
+
FIX V6.4: detach+clone nos tensores de entrada do classifier,
|
| 104 |
+
e cálculo do vetor de ativação SOM dentro de torch.no_grad().
|
| 105 |
+
O classifier treina apenas sobre seus próprios pesos (8 FC layers).
|
| 106 |
+
|
| 107 |
+
============================================================================
|
| 108 |
+
V6.5 — VQ-VAE-2 NO PIPELINE DE COMPRESSÃO
|
| 109 |
+
============================================================================
|
| 110 |
+
User requirement: "ativar efetivamente o VQ-VAE-2 no pipeline de compressão"
|
| 111 |
+
|
| 112 |
+
Integração: o KohonenLearningSystem agora possui um `vqvae2_compressor`
|
| 113 |
+
opcional (HierarchicalVQVAE2 do módulo vqvae2_hierarchical_flexnet.py).
|
| 114 |
+
Quando ativado:
|
| 115 |
+
|
| 116 |
+
1. Após `train_som_on_buffer()`, o buffer_4d (B, 4) é passado ao VQ-VAE-2
|
| 117 |
+
como entrada. O encoder mapeia (B, 4) -> z_e (B, code_dim).
|
| 118 |
+
2. O VQ hierárquico produz:
|
| 119 |
+
- z_q_top: estrutura global do batch (codebook K_top)
|
| 120 |
+
- z_q_bot: detalhes residuais (codebook K_bot)
|
| 121 |
+
3. O decoder reconstrói z_recon (B, 4) a partir de z_q_combined.
|
| 122 |
+
4. A loss do VQ-VAE-2 (commitment + recon) é computada e retornada
|
| 123 |
+
para monitoramento (não adicionada à loss do SOM — são objetivos
|
| 124 |
+
ortogonais: SOM aprende topologia, VQ-VAE-2 aprende compressão).
|
| 125 |
+
5. Códigos top/bottom podem ser usados como representação compacta
|
| 126 |
+
do estado do SOM para armazenamento/transferência.
|
| 127 |
+
|
| 128 |
+
Benefícios:
|
| 129 |
+
- Compressão neural do espaço 4D do SOM (4 -> code_dim -> 2 códigos)
|
| 130 |
+
- Codebook compartilhado entre batches (aprendizado incremental)
|
| 131 |
+
- Dead code restart evita colapso do codebook
|
| 132 |
+
- Goose VQ (Gumbel-softmax) força uso uniforme do codebook
|
| 133 |
+
|
| 134 |
+
============================================================================
|
| 135 |
+
V6.5 — REASONING_ENGINE INTEGRADO
|
| 136 |
+
============================================================================
|
| 137 |
+
User requirement: "integrar reasoning_engine ao KohonenLearningSystem"
|
| 138 |
+
|
| 139 |
+
Integração: o KohonenLearningSystem agora possui um `reasoning_engine`
|
| 140 |
+
opcional (ReasoningEngine do módulo reasoning_engine.py). Quando ativado:
|
| 141 |
+
|
| 142 |
+
1. Após `predict()`, o resultado da predição é passado ao reasoning_engine
|
| 143 |
+
que gera uma sequência de tags <think>, <plan>, <decompose>,
|
| 144 |
+
<execute>, <monitor>, <predict>, <adjust>, <answer>.
|
| 145 |
+
2. O reasoning_engine pode usar ferramentas registradas (tool_agent)
|
| 146 |
+
para consultas externas (ex: calculator, knowledge_base).
|
| 147 |
+
3. O streaming de raciocínio é compatível com Ollama/LangChain/vLLM
|
| 148 |
+
via tags padrão.
|
| 149 |
+
4. Para predições do SOM, o reasoning_engine pode explicar o porquê
|
| 150 |
+
do BMU ter sido escolhido (análise de distâncias).
|
| 151 |
+
|
| 152 |
+
Métodos adicionados:
|
| 153 |
+
- reason_about(query): retorna generator com streaming de raciocínio
|
| 154 |
+
- reason_sync(query): retorna string completa com todas as tags
|
| 155 |
+
- get_reasoning_stats(): retorna estatísticas do reasoning_engine
|
| 156 |
+
|
| 157 |
+
============================================================================
|
| 158 |
+
"""
|
| 159 |
+
from __future__ import annotations
|
| 160 |
+
|
| 161 |
+
import os
|
| 162 |
+
import math
|
| 163 |
+
import torch
|
| 164 |
+
import torch.nn as nn
|
| 165 |
+
import torch.nn.functional as F
|
| 166 |
+
from collections import defaultdict, Counter
|
| 167 |
+
from typing import List, Tuple, Optional, Dict, Any, Iterator
|
| 168 |
+
|
| 169 |
+
|
| 170 |
+
# ============================================================================
|
| 171 |
+
# Configuração da CPU – núcleos físicos
|
| 172 |
+
# ============================================================================
|
| 173 |
+
try:
|
| 174 |
+
import psutil
|
| 175 |
+
N_CORES = psutil.cpu_count(logical=False)
|
| 176 |
+
except ImportError:
|
| 177 |
+
N_CORES = os.cpu_count() // 2 if os.cpu_count() else 4
|
| 178 |
+
if N_CORES:
|
| 179 |
+
try:
|
| 180 |
+
torch.set_num_threads(N_CORES)
|
| 181 |
+
torch.set_num_interop_threads(N_CORES)
|
| 182 |
+
except RuntimeError:
|
| 183 |
+
pass # already initialized
|
| 184 |
+
|
| 185 |
+
|
| 186 |
+
# ============================================================================
|
| 187 |
+
# Tokenizador BBPE simplificado
|
| 188 |
+
# ============================================================================
|
| 189 |
+
class SimpleBBPETokenizer:
|
| 190 |
+
"""Tokenizador BPE simplificado (nível de palavra).
|
| 191 |
+
|
| 192 |
+
Vocabulário especial: pad=0, eos=1, unk=2.
|
| 193 |
+
Tokens comuns mapeados a partir do índice 3.
|
| 194 |
+
"""
|
| 195 |
+
|
| 196 |
+
def __init__(self, vocab_size: int = 16384):
|
| 197 |
+
self.vocab_size = vocab_size
|
| 198 |
+
self.token_to_id = {}
|
| 199 |
+
self.id_to_token = {}
|
| 200 |
+
self.pad_token = "<pad>"
|
| 201 |
+
self.eos_token = "<eos>"
|
| 202 |
+
self.unk_token = "<unk>"
|
| 203 |
+
self.pad_id = 0
|
| 204 |
+
self.eos_id = 1
|
| 205 |
+
self.unk_id = 2
|
| 206 |
+
self._init_special_tokens()
|
| 207 |
+
|
| 208 |
+
def _init_special_tokens(self):
|
| 209 |
+
self.token_to_id[self.pad_token] = self.pad_id
|
| 210 |
+
self.token_to_id[self.eos_token] = self.eos_id
|
| 211 |
+
self.token_to_id[self.unk_token] = self.unk_id
|
| 212 |
+
self.id_to_token[self.pad_id] = self.pad_token
|
| 213 |
+
self.id_to_token[self.eos_id] = self.eos_token
|
| 214 |
+
self.id_to_token[self.unk_id] = self.unk_token
|
| 215 |
+
|
| 216 |
+
def fit(self, texts: List[str]):
|
| 217 |
+
"""Constrói vocabulário por frequência (top vocab_size-3 palavras)."""
|
| 218 |
+
word_counts = Counter()
|
| 219 |
+
for text in texts:
|
| 220 |
+
words = text.split()
|
| 221 |
+
word_counts.update(words)
|
| 222 |
+
sorted_words = [w for w, _ in word_counts.most_common(self.vocab_size - 3)]
|
| 223 |
+
for idx, word in enumerate(sorted_words, start=3):
|
| 224 |
+
self.token_to_id[word] = idx
|
| 225 |
+
self.id_to_token[idx] = word
|
| 226 |
+
|
| 227 |
+
def encode(self, text: str, max_length: int = 8) -> List[int]:
|
| 228 |
+
"""Encoda + append EOS + pad/trunca para max_length."""
|
| 229 |
+
words = text.split()
|
| 230 |
+
ids = [self.token_to_id.get(w, self.unk_id) for w in words]
|
| 231 |
+
ids.append(self.eos_id)
|
| 232 |
+
if len(ids) > max_length:
|
| 233 |
+
ids = ids[:max_length]
|
| 234 |
+
else:
|
| 235 |
+
ids += [self.pad_id] * (max_length - len(ids))
|
| 236 |
+
return ids
|
| 237 |
+
|
| 238 |
+
|
| 239 |
+
# ============================================================================
|
| 240 |
+
# Codificação posicional e conversão texto → vetor 4D (com w temporal)
|
| 241 |
+
# ============================================================================
|
| 242 |
+
def positional_encoding(seq_len: int, hidden_dim: int) -> torch.Tensor:
|
| 243 |
+
"""PE sinusoidal clássico: PE[pos, 2k] = sin(pos·exp(-2k·log(10000)/D)),
|
| 244 |
+
PE[pos, 2k+1] = cos(pos·exp(-2k·log(10000)/D)).
|
| 245 |
+
"""
|
| 246 |
+
pe = torch.zeros(seq_len, hidden_dim)
|
| 247 |
+
position = torch.arange(0, seq_len, dtype=torch.float).unsqueeze(1)
|
| 248 |
+
div_term = torch.exp(
|
| 249 |
+
torch.arange(0, hidden_dim, 2).float() * (-math.log(10000.0) / hidden_dim)
|
| 250 |
+
)
|
| 251 |
+
pe[:, 0::2] = torch.sin(position * div_term)
|
| 252 |
+
pe[:, 1::2] = torch.cos(position * div_term)
|
| 253 |
+
return pe
|
| 254 |
+
|
| 255 |
+
|
| 256 |
+
def text_to_4d_vector(
|
| 257 |
+
text: str,
|
| 258 |
+
tokenizer: SimpleBBPETokenizer,
|
| 259 |
+
embedding: nn.Embedding,
|
| 260 |
+
hidden_dim: int,
|
| 261 |
+
seq_len: int,
|
| 262 |
+
time_step: int,
|
| 263 |
+
T_max: int = 10000,
|
| 264 |
+
) -> torch.Tensor:
|
| 265 |
+
"""Converte sentença em vetor 4D (x, y, z, w), onde w = time_step / T_max.
|
| 266 |
+
|
| 267 |
+
Pipeline matemático:
|
| 268 |
+
ids → Embedding → +PE → SVD(proj top-3) → centróide xyz
|
| 269 |
+
w = time_step / T_max (LINEAR no tempo)
|
| 270 |
+
vec_4d = concat(xyz, w) → [4]
|
| 271 |
+
"""
|
| 272 |
+
ids = tokenizer.encode(text, max_length=seq_len)
|
| 273 |
+
input_ids = torch.tensor(ids).unsqueeze(0)
|
| 274 |
+
word_emb = embedding(input_ids).squeeze(0) # (L, D)
|
| 275 |
+
pe = positional_encoding(seq_len, hidden_dim)
|
| 276 |
+
fused = word_emb + pe # (L, D)
|
| 277 |
+
|
| 278 |
+
# SVD para 3D
|
| 279 |
+
mean_centered = fused - fused.mean(dim=0, keepdim=True)
|
| 280 |
+
U, S, V = torch.linalg.svd(mean_centered, full_matrices=False)
|
| 281 |
+
coords_3d = torch.mm(mean_centered, V[:3, :].t()) # (L, 3)
|
| 282 |
+
|
| 283 |
+
xyz_mean = coords_3d.mean(dim=0) # (3,)
|
| 284 |
+
w = torch.tensor(time_step / T_max, dtype=torch.float) # valor temporal
|
| 285 |
+
vec_4d = torch.cat([xyz_mean, w.unsqueeze(0)]).contiguous()
|
| 286 |
+
return vec_4d
|
| 287 |
+
|
| 288 |
+
|
| 289 |
+
# ============================================================================
|
| 290 |
+
# Mapa de Kohonen 4D com EWC (Fisher para w não-nulo)
|
| 291 |
+
# ============================================================================
|
| 292 |
+
class KohonenSOM4D:
|
| 293 |
+
"""Mapa Auto-Organizável 4D com EWC apenas na 4ª dimensão (w temporal).
|
| 294 |
+
|
| 295 |
+
Args:
|
| 296 |
+
grid_shape: (I, J, K, L) — dimensões do grid 4D.
|
| 297 |
+
alpha0: taxa de aprendizado inicial (α₀).
|
| 298 |
+
sigma0: largura inicial da vizinhança (σ₀).
|
| 299 |
+
lambda_ewc: peso da penalidade EWC (λ).
|
| 300 |
+
|
| 301 |
+
Atributos:
|
| 302 |
+
weights: W ∈ ℝ^(I×J×K×L×4) — pesos dos neurônios.
|
| 303 |
+
old_weights_w: W*_w ∈ ℝ^(I×J×K×L) — referência EWC (apenas w).
|
| 304 |
+
fisher_w: F ∈ ℝ^(I×J×K×L) — informação de Fisher por neurônio (w).
|
| 305 |
+
fisher_accum / fisher_count: acumuladores para cálculo de F.
|
| 306 |
+
"""
|
| 307 |
+
|
| 308 |
+
def __init__(
|
| 309 |
+
self,
|
| 310 |
+
grid_shape: Tuple[int, int, int, int],
|
| 311 |
+
alpha0: float = 0.1,
|
| 312 |
+
sigma0: float = 1.0,
|
| 313 |
+
lambda_ewc: float = 0.01,
|
| 314 |
+
):
|
| 315 |
+
self.I, self.J, self.K, self.L = grid_shape
|
| 316 |
+
self.alpha0 = alpha0
|
| 317 |
+
self.sigma0 = sigma0
|
| 318 |
+
self.lambda_ewc = lambda_ewc
|
| 319 |
+
self.t = 0
|
| 320 |
+
|
| 321 |
+
self.weights = torch.randn(self.I, self.J, self.K, self.L, 4)
|
| 322 |
+
self.old_weights_w = None
|
| 323 |
+
self.fisher_w = None
|
| 324 |
+
self.fisher_accum = torch.zeros(self.I, self.J, self.K, self.L)
|
| 325 |
+
self.fisher_count = torch.zeros(self.I, self.J, self.K, self.L)
|
| 326 |
+
|
| 327 |
+
def _neighborhood(self, bmu_idx):
|
| 328 |
+
"""Vizinhança Gaussiana 4D: d² = Δi² + Δj² + Δk² + Δl²."""
|
| 329 |
+
i, j, k, l = bmu_idx
|
| 330 |
+
II, JJ, KK, LL = torch.meshgrid(
|
| 331 |
+
torch.arange(self.I).float(),
|
| 332 |
+
torch.arange(self.J).float(),
|
| 333 |
+
torch.arange(self.K).float(),
|
| 334 |
+
torch.arange(self.L).float(),
|
| 335 |
+
indexing="ij",
|
| 336 |
+
)
|
| 337 |
+
dist_sq = (II - i) ** 2 + (JJ - j) ** 2 + (KK - k) ** 2 + (LL - l) ** 2
|
| 338 |
+
return dist_sq
|
| 339 |
+
|
| 340 |
+
def find_bmu(self, x: torch.Tensor) -> Tuple[int, int, int, int]:
|
| 341 |
+
"""Best Matching Unit: argmin ||W - x||² em ℝ⁴.
|
| 342 |
+
|
| 343 |
+
Substitui pgvector_lookup — busca nearest-neighbor flat sobre o grid.
|
| 344 |
+
"""
|
| 345 |
+
dist = torch.sum((self.weights - x.view(1, 1, 1, 1, 4)) ** 2, dim=-1)
|
| 346 |
+
flat_idx = torch.argmin(dist).item()
|
| 347 |
+
i = flat_idx // (self.J * self.K * self.L)
|
| 348 |
+
rest = flat_idx % (self.J * self.K * self.L)
|
| 349 |
+
j = rest // (self.K * self.L)
|
| 350 |
+
rest = rest % (self.K * self.L)
|
| 351 |
+
k = rest // self.L
|
| 352 |
+
l = rest % self.L
|
| 353 |
+
return (i, j, k, l)
|
| 354 |
+
|
| 355 |
+
def update_weights(self, x: torch.Tensor, bmu_idx, accumulate_fisher=False):
|
| 356 |
+
"""Update Kohonen: ΔW = α·Λ·(x - W) + penalidade EWC em w.
|
| 357 |
+
|
| 358 |
+
Args:
|
| 359 |
+
x: tensor [4] — amostra 4D.
|
| 360 |
+
bmu_idx: (i, j, k, l) — índice do BMU.
|
| 361 |
+
accumulate_fisher: se True, acumula (x_w - W_w)² nos Fisher accumulators.
|
| 362 |
+
"""
|
| 363 |
+
dist_sq = self._neighborhood(bmu_idx)
|
| 364 |
+
sigma = self.sigma0 * math.exp(-self.t / 1000)
|
| 365 |
+
alpha = self.alpha0 * math.exp(-self.t / 2000)
|
| 366 |
+
h = torch.exp(-dist_sq / (2 * sigma ** 2))
|
| 367 |
+
|
| 368 |
+
delta = x - self.weights
|
| 369 |
+
update = alpha * h.unsqueeze(-1) * delta
|
| 370 |
+
|
| 371 |
+
if self.old_weights_w is not None and self.fisher_w is not None:
|
| 372 |
+
# Penalidade EWC apenas na 4ª dimensão (w)
|
| 373 |
+
# ∂L_ewc/∂W_w = λ · F · (W_w - W*_w) → subtraído do update
|
| 374 |
+
ewc_penalty = self.lambda_ewc * self.fisher_w * (
|
| 375 |
+
self.weights[..., 3] - self.old_weights_w
|
| 376 |
+
)
|
| 377 |
+
update[..., 3] = update[..., 3] - ewc_penalty
|
| 378 |
+
|
| 379 |
+
self.weights = self.weights + update
|
| 380 |
+
|
| 381 |
+
if accumulate_fisher:
|
| 382 |
+
# Acumula Fisher apenas em neurônios próximos ao BMU (Λ > 0.1)
|
| 383 |
+
mask = h > 0.1
|
| 384 |
+
if mask.any():
|
| 385 |
+
diff_sq = (x[3] - self.weights[mask][..., 3]) ** 2
|
| 386 |
+
self.fisher_accum[mask] += diff_sq
|
| 387 |
+
self.fisher_count[mask] += 1
|
| 388 |
+
|
| 389 |
+
self.t += 1
|
| 390 |
+
|
| 391 |
+
def finalize_fisher(self):
|
| 392 |
+
"""Fisher = mean((x_w - W_w)²) sobre samples acumuladas."""
|
| 393 |
+
cnt = self.fisher_count.clamp(min=1e-8)
|
| 394 |
+
self.fisher_w = self.fisher_accum / cnt
|
| 395 |
+
|
| 396 |
+
def set_ewc_reference(self):
|
| 397 |
+
"""Consolida W_w como referência EWC e finaliza Fisher."""
|
| 398 |
+
self.old_weights_w = self.weights[..., 3].clone()
|
| 399 |
+
self.finalize_fisher()
|
| 400 |
+
self.fisher_accum.zero_()
|
| 401 |
+
self.fisher_count.zero_()
|
| 402 |
+
|
| 403 |
+
# ------------------------------------------------------------------
|
| 404 |
+
# Métricas para monitoramento (V6.4)
|
| 405 |
+
# ------------------------------------------------------------------
|
| 406 |
+
def get_metrics(self) -> dict:
|
| 407 |
+
"""Retorna métricas atuais do SOM para monitoramento."""
|
| 408 |
+
sigma_t = self.sigma0 * math.exp(-self.t / 1000)
|
| 409 |
+
alpha_t = self.alpha0 * math.exp(-self.t / 2000)
|
| 410 |
+
return {
|
| 411 |
+
"t": int(self.t),
|
| 412 |
+
"sigma_t": float(sigma_t),
|
| 413 |
+
"alpha_t": float(alpha_t),
|
| 414 |
+
"sigma0": float(self.sigma0),
|
| 415 |
+
"alpha0": float(self.alpha0),
|
| 416 |
+
"lambda_ewc": float(self.lambda_ewc),
|
| 417 |
+
"grid_shape": [int(self.I), int(self.J), int(self.K), int(self.L)],
|
| 418 |
+
"n_neurons": int(self.I * self.J * self.K * self.L),
|
| 419 |
+
"has_ewc_reference": self.old_weights_w is not None,
|
| 420 |
+
"fisher_w_mean": (
|
| 421 |
+
float(self.fisher_w.mean().item())
|
| 422 |
+
if self.fisher_w is not None
|
| 423 |
+
else 0.0
|
| 424 |
+
),
|
| 425 |
+
"fisher_w_max": (
|
| 426 |
+
float(self.fisher_w.max().item())
|
| 427 |
+
if self.fisher_w is not None
|
| 428 |
+
else 0.0
|
| 429 |
+
),
|
| 430 |
+
"fisher_accum_count": int(self.fisher_count.sum().item()),
|
| 431 |
+
"weights_norm": float(self.weights.norm().item()),
|
| 432 |
+
"weights_w_mean": float(self.weights[..., 3].mean().item()),
|
| 433 |
+
}
|
| 434 |
+
|
| 435 |
+
|
| 436 |
+
# ============================================================================
|
| 437 |
+
# Classificador de hipótese (8 camadas FC)
|
| 438 |
+
# ============================================================================
|
| 439 |
+
class HypothesisClassifier(nn.Module):
|
| 440 |
+
"""Classificador de hipótese: 8 camadas FC + ReLU + output logit.
|
| 441 |
+
|
| 442 |
+
Arquitetura: [input → 512 → 256 → 128 → 64 → 32 → 16 → 8] + ReLU
|
| 443 |
+
+ [8 → 1] (logit)
|
| 444 |
+
Loss: BCEWithLogitsLoss
|
| 445 |
+
"""
|
| 446 |
+
|
| 447 |
+
def __init__(self, input_dim, hidden_dims=[512, 256, 128, 64, 32, 16, 8]):
|
| 448 |
+
super().__init__()
|
| 449 |
+
layers = []
|
| 450 |
+
prev = input_dim
|
| 451 |
+
for h in hidden_dims:
|
| 452 |
+
layers.append(nn.Linear(prev, h))
|
| 453 |
+
layers.append(nn.ReLU())
|
| 454 |
+
prev = h
|
| 455 |
+
layers.append(nn.Linear(prev, 1))
|
| 456 |
+
self.net = nn.Sequential(*layers)
|
| 457 |
+
|
| 458 |
+
def forward(self, x):
|
| 459 |
+
return self.net(x).squeeze(-1)
|
| 460 |
+
|
| 461 |
+
|
| 462 |
+
# ============================================================================
|
| 463 |
+
# Sistema de aprendizado completo (com w temporal e condição de início por N)
|
| 464 |
+
# ============================================================================
|
| 465 |
+
class KohonenLearningSystem:
|
| 466 |
+
"""Pipeline integrado: tokenizer + embedding + SOM4D + classifier + punishment.
|
| 467 |
+
|
| 468 |
+
V6.5: + VQ-VAE-2 compressor (opcional) + reasoning_engine (opcional)
|
| 469 |
+
|
| 470 |
+
Args:
|
| 471 |
+
vocab_size: tamanho do vocabulário BBPE (default 16384).
|
| 472 |
+
hidden_dim: dimensão do embedding (default 1024).
|
| 473 |
+
seq_len: comprimento máximo da sequência (default 8).
|
| 474 |
+
som_grid: (I, J, K, L) — grid 4D do SOM (default (6, 6, 6, 4) = 864).
|
| 475 |
+
alpha0, sigma0: hiperparâmetros do SOM.
|
| 476 |
+
lambda_ewc: peso da penalidade EWC.
|
| 477 |
+
N_start: threshold do histograma para iniciar treino.
|
| 478 |
+
dim_choice: 'x' | 'y' | 'z' — dimensão usada no histograma.
|
| 479 |
+
hypothesis_hidden: arquitetura do HypothesisClassifier.
|
| 480 |
+
T_max: normalização temporal (w = time_step / T_max).
|
| 481 |
+
enable_vqvae2 (V6.5): ativa VQ-VAE-2 compressor no pipeline.
|
| 482 |
+
enable_reasoning (V6.5): ativa reasoning_engine integrado.
|
| 483 |
+
vqvae2_code_dim (V6.5): dimensão do codebook do VQ-VAE-2.
|
| 484 |
+
vqvae2_num_codes (V6.5): tamanho do codebook top+bottom.
|
| 485 |
+
"""
|
| 486 |
+
|
| 487 |
+
def __init__(
|
| 488 |
+
self,
|
| 489 |
+
vocab_size=16384,
|
| 490 |
+
hidden_dim=1024,
|
| 491 |
+
seq_len=8,
|
| 492 |
+
som_grid=(6, 6, 6, 4),
|
| 493 |
+
alpha0=0.1,
|
| 494 |
+
sigma0=1.5,
|
| 495 |
+
lambda_ewc=0.02,
|
| 496 |
+
N_start=10,
|
| 497 |
+
dim_choice="y",
|
| 498 |
+
hypothesis_hidden=[512, 256, 128, 64, 32, 16, 8],
|
| 499 |
+
T_max=10000,
|
| 500 |
+
# V6.5 — VQ-VAE-2 + reasoning_engine
|
| 501 |
+
enable_vqvae2: bool = True,
|
| 502 |
+
enable_reasoning: bool = True,
|
| 503 |
+
vqvae2_code_dim: int = 16,
|
| 504 |
+
vqvae2_num_codes_top: int = 64,
|
| 505 |
+
vqvae2_num_codes_bot: int = 128,
|
| 506 |
+
):
|
| 507 |
+
self.tokenizer = SimpleBBPETokenizer(vocab_size)
|
| 508 |
+
self.embedding = nn.Embedding(vocab_size, hidden_dim)
|
| 509 |
+
self.hidden_dim = hidden_dim
|
| 510 |
+
self.seq_len = seq_len
|
| 511 |
+
self.T_max = T_max
|
| 512 |
+
self.time_counter = 0 # contador global de amostras processadas
|
| 513 |
+
|
| 514 |
+
self.som = KohonenSOM4D(som_grid, alpha0, sigma0, lambda_ewc)
|
| 515 |
+
self.som_grid = som_grid
|
| 516 |
+
self.som_neuron_count = (
|
| 517 |
+
som_grid[0] * som_grid[1] * som_grid[2] * som_grid[3]
|
| 518 |
+
)
|
| 519 |
+
|
| 520 |
+
self.classifier: Optional[HypothesisClassifier] = None
|
| 521 |
+
self.hypothesis_hidden = hypothesis_hidden
|
| 522 |
+
self.classifier_trained = False
|
| 523 |
+
|
| 524 |
+
self.buffer_4d = []
|
| 525 |
+
self.buffer_labels = []
|
| 526 |
+
self.training_ready = False
|
| 527 |
+
self.N = N_start
|
| 528 |
+
self.dim_choice = dim_choice
|
| 529 |
+
self.dim_index = {"x": 0, "y": 1, "z": 2}[dim_choice]
|
| 530 |
+
|
| 531 |
+
self.punishment_count = 0
|
| 532 |
+
self.success_count = 0
|
| 533 |
+
self.histogram = Counter()
|
| 534 |
+
|
| 535 |
+
self.required_new_samples = 0
|
| 536 |
+
|
| 537 |
+
# ------------------------------------------------------------------
|
| 538 |
+
# V6.5 — VQ-VAE-2 compressor (ativa efetiva no pipeline)
|
| 539 |
+
# ------------------------------------------------------------------
|
| 540 |
+
self.enable_vqvae2 = enable_vqvae2
|
| 541 |
+
self.vqvae2_compressor = None
|
| 542 |
+
self.vqvae2_metrics_history: List[Dict[str, Any]] = []
|
| 543 |
+
if enable_vqvae2:
|
| 544 |
+
try:
|
| 545 |
+
from .vqvae2_hierarchical_flexnet import HierarchicalVQVAE2
|
| 546 |
+
# Modalidade única: "som_4d" com input_dim=4
|
| 547 |
+
self.vqvae2_compressor = HierarchicalVQVAE2(
|
| 548 |
+
modalities={"som_4d": 4},
|
| 549 |
+
code_dim=vqvae2_code_dim,
|
| 550 |
+
num_codes_top=vqvae2_num_codes_top,
|
| 551 |
+
num_codes_bot=vqvae2_num_codes_bot,
|
| 552 |
+
hidden=32,
|
| 553 |
+
beta=0.25,
|
| 554 |
+
ema_decay=0.99,
|
| 555 |
+
dead_code_threshold=1.0,
|
| 556 |
+
dead_code_restart_every=3,
|
| 557 |
+
goose_temp_init=2.0,
|
| 558 |
+
goose_temp_final=0.5,
|
| 559 |
+
goose_schedule="cosine",
|
| 560 |
+
total_epochs=25,
|
| 561 |
+
norm_type="none",
|
| 562 |
+
rmsnorm_in_vq=False,
|
| 563 |
+
)
|
| 564 |
+
# Inicia em modo treino para ativar EMA updates
|
| 565 |
+
self.vqvae2_compressor.train()
|
| 566 |
+
except Exception as e:
|
| 567 |
+
# Fallback: desabilita VQ-VAE-2 se houver erro de import
|
| 568 |
+
self.enable_vqvae2 = False
|
| 569 |
+
self.vqvae2_compressor = None
|
| 570 |
+
import warnings
|
| 571 |
+
warnings.warn(f"VQ-VAE-2 disabled: {e}")
|
| 572 |
+
|
| 573 |
+
# ------------------------------------------------------------------
|
| 574 |
+
# V6.5 — ReasoningEngine (integração ativa)
|
| 575 |
+
# ------------------------------------------------------------------
|
| 576 |
+
self.enable_reasoning = enable_reasoning
|
| 577 |
+
self.reasoning_engine = None
|
| 578 |
+
if enable_reasoning:
|
| 579 |
+
try:
|
| 580 |
+
from ..reasoning.reasoning_engine import ReasoningEngine
|
| 581 |
+
self.reasoning_engine = ReasoningEngine(
|
| 582 |
+
max_thinking_steps=10,
|
| 583 |
+
max_iterations=3,
|
| 584 |
+
convergence_threshold=0.9,
|
| 585 |
+
verbose=False,
|
| 586 |
+
)
|
| 587 |
+
# Registra uma ferramenta interna: consultar SOM
|
| 588 |
+
def som_query_tool(query: str) -> str:
|
| 589 |
+
"""Ferramenta: consulta o SOM do KLS para responder."""
|
| 590 |
+
pred = self.predict(query)
|
| 591 |
+
bmu_info = ""
|
| 592 |
+
if self.buffer_4d:
|
| 593 |
+
try:
|
| 594 |
+
vec = text_to_4d_vector(
|
| 595 |
+
query, self.tokenizer, self.embedding,
|
| 596 |
+
self.hidden_dim, self.seq_len,
|
| 597 |
+
self.time_counter, self.T_max,
|
| 598 |
+
)
|
| 599 |
+
bmu = self.som.find_bmu(vec)
|
| 600 |
+
bmu_info = f" | BMU={bmu}"
|
| 601 |
+
except Exception:
|
| 602 |
+
pass
|
| 603 |
+
return f"prediction={pred}{bmu_info}"
|
| 604 |
+
self.reasoning_engine.register_tool(
|
| 605 |
+
"som_query", som_query_tool,
|
| 606 |
+
description="Consulta o SOM do KohonenLearningSystem",
|
| 607 |
+
timeout_s=10.0,
|
| 608 |
+
)
|
| 609 |
+
except Exception as e:
|
| 610 |
+
self.enable_reasoning = False
|
| 611 |
+
self.reasoning_engine = None
|
| 612 |
+
import warnings
|
| 613 |
+
warnings.warn(f"ReasoningEngine disabled: {e}")
|
| 614 |
+
|
| 615 |
+
def add_data(self, sentences: List[str], labels: List[int]):
|
| 616 |
+
"""Adiciona amostras: text → 4D vector + atualiza histograma."""
|
| 617 |
+
for sent, lab in zip(sentences, labels):
|
| 618 |
+
self.time_counter += 1
|
| 619 |
+
vec = text_to_4d_vector(
|
| 620 |
+
sent,
|
| 621 |
+
self.tokenizer,
|
| 622 |
+
self.embedding,
|
| 623 |
+
self.hidden_dim,
|
| 624 |
+
self.seq_len,
|
| 625 |
+
self.time_counter,
|
| 626 |
+
self.T_max,
|
| 627 |
+
)
|
| 628 |
+
self.buffer_4d.append(vec)
|
| 629 |
+
self.buffer_labels.append(lab)
|
| 630 |
+
dim_val = round(vec[self.dim_index].item(), 2)
|
| 631 |
+
self.histogram[dim_val] += 1
|
| 632 |
+
|
| 633 |
+
def check_training_start(self) -> bool:
|
| 634 |
+
"""Inicia treino quando algum bucket do histograma atinge N."""
|
| 635 |
+
if (
|
| 636 |
+
not self.training_ready
|
| 637 |
+
and max(self.histogram.values(), default=0) >= self.N
|
| 638 |
+
):
|
| 639 |
+
self.training_ready = True
|
| 640 |
+
return True
|
| 641 |
+
return False
|
| 642 |
+
|
| 643 |
+
def train_som_on_buffer(self):
|
| 644 |
+
"""Treina SOM por 5 épocas sobre o buffer atual (com Fisher accum)."""
|
| 645 |
+
if not self.buffer_4d:
|
| 646 |
+
return
|
| 647 |
+
data = torch.stack(self.buffer_4d)
|
| 648 |
+
for _ in range(5): # épocas de treino rápido
|
| 649 |
+
perm = torch.randperm(len(data))
|
| 650 |
+
for idx in perm:
|
| 651 |
+
x = data[idx]
|
| 652 |
+
bmu = self.som.find_bmu(x)
|
| 653 |
+
acc_fisher = (
|
| 654 |
+
self.punishment_count == 0
|
| 655 |
+
and self.som.old_weights_w is None
|
| 656 |
+
)
|
| 657 |
+
self.som.update_weights(x, bmu, accumulate_fisher=acc_fisher)
|
| 658 |
+
|
| 659 |
+
# V6.5 — Ativa VQ-VAE-2 compressor no pipeline
|
| 660 |
+
if self.enable_vqvae2 and self.vqvae2_compressor is not None:
|
| 661 |
+
self._compress_buffer_with_vqvae2(data)
|
| 662 |
+
|
| 663 |
+
def _compress_buffer_with_vqvae2(self, data: torch.Tensor) -> Dict[str, Any]:
|
| 664 |
+
"""V6.5 — Comprime buffer 4D via VQ-VAE-2 hierárquico.
|
| 665 |
+
|
| 666 |
+
Ativa efetivamente o VQ-VAE-2 no pipeline de compressão:
|
| 667 |
+
1. Encoder: (B, 4) → z_e (B, code_dim)
|
| 668 |
+
2. VQ hierárquico: z_e → z_q_top + z_q_bot (codebooks EMA + Goose)
|
| 669 |
+
3. Decoder: z_q_combined → z_recon (B, 4)
|
| 670 |
+
4. Loss: commitment (top+bot) + reconstruction (MSE)
|
| 671 |
+
5. Códigos top/bottom retornados para inspeção
|
| 672 |
+
|
| 673 |
+
Args:
|
| 674 |
+
data: tensor (B, 4) com vetores 4D do buffer.
|
| 675 |
+
|
| 676 |
+
Returns:
|
| 677 |
+
Dict com vqvae2_metrics (também armazenado em vqvae2_metrics_history).
|
| 678 |
+
"""
|
| 679 |
+
try:
|
| 680 |
+
# Sanitiza NaN/Inf
|
| 681 |
+
data_clean = torch.nan_to_num(data, nan=0.0, posinf=1e4, neginf=-1e4)
|
| 682 |
+
# VQ-VAE-2 espera dict {modality_name: tensor}
|
| 683 |
+
batch = {"som_4d": data_clean}
|
| 684 |
+
out = self.vqvae2_compressor(batch)
|
| 685 |
+
# Incrementa época do VQ (controla schedule Goose + dead code restart)
|
| 686 |
+
try:
|
| 687 |
+
self.vqvae2_compressor.vq.increment_epoch()
|
| 688 |
+
except Exception:
|
| 689 |
+
pass
|
| 690 |
+
stats = out.get("stats", {})
|
| 691 |
+
vqvae2_metrics = {
|
| 692 |
+
"vq_loss": float(out.get("vq_loss", 0.0)),
|
| 693 |
+
"recon_loss": float(out.get("recon_loss", 0.0)),
|
| 694 |
+
"total_loss": float(out.get("total_loss", 0.0)),
|
| 695 |
+
"n_used_top": int(stats.get("n_used_top", 0)),
|
| 696 |
+
"n_used_bot": int(stats.get("n_used_bot", 0)),
|
| 697 |
+
"usage_ratio_top": float(stats.get("usage_ratio_top", 0.0)),
|
| 698 |
+
"usage_ratio_bot": float(stats.get("usage_ratio_bot", 0.0)),
|
| 699 |
+
"codebook_ppl_top": float(stats.get("codebook_ppl_top", 0.0)),
|
| 700 |
+
"codebook_ppl_bot": float(stats.get("codebook_ppl_bot", 0.0)),
|
| 701 |
+
"n_restarted_top": int(stats.get("n_restarted_top", 0)),
|
| 702 |
+
"n_restarted_bot": int(stats.get("n_restarted_bot", 0)),
|
| 703 |
+
"goose_temp": float(stats.get("goose_temp", 0.0)),
|
| 704 |
+
"active": True,
|
| 705 |
+
}
|
| 706 |
+
self.vqvae2_metrics_history.append(vqvae2_metrics)
|
| 707 |
+
# Mantém apenas últimas 100 entries para limitar memória
|
| 708 |
+
if len(self.vqvae2_metrics_history) > 100:
|
| 709 |
+
self.vqvae2_metrics_history = self.vqvae2_metrics_history[-100:]
|
| 710 |
+
return vqvae2_metrics
|
| 711 |
+
except Exception as e:
|
| 712 |
+
return {
|
| 713 |
+
"active": False,
|
| 714 |
+
"error": str(e)[:200],
|
| 715 |
+
"vq_loss": 0.0,
|
| 716 |
+
"recon_loss": 0.0,
|
| 717 |
+
"total_loss": 0.0,
|
| 718 |
+
}
|
| 719 |
+
|
| 720 |
+
def get_vqvae2_metrics(self) -> Dict[str, Any]:
|
| 721 |
+
"""V6.5 — Retorna métricas atuais do VQ-VAE-2 compressor."""
|
| 722 |
+
if not self.enable_vqvae2 or self.vqvae2_compressor is None:
|
| 723 |
+
return {"active": False, "reason": "disabled"}
|
| 724 |
+
if not self.vqvae2_metrics_history:
|
| 725 |
+
return {"active": True, "n_calls": 0}
|
| 726 |
+
latest = self.vqvae2_metrics_history[-1]
|
| 727 |
+
# V6.5: skip NaN values when computing means (early calls may produce NaN
|
| 728 |
+
# due to Gumbel-softmax instability before codebook warmup)
|
| 729 |
+
import math
|
| 730 |
+
valid_total = [m.get("total_loss", 0.0) for m in self.vqvae2_metrics_history
|
| 731 |
+
if not math.isnan(m.get("total_loss", 0.0))]
|
| 732 |
+
valid_recon = [m.get("recon_loss", 0.0) for m in self.vqvae2_metrics_history
|
| 733 |
+
if not math.isnan(m.get("recon_loss", 0.0))]
|
| 734 |
+
return {
|
| 735 |
+
"active": True,
|
| 736 |
+
"n_calls": len(self.vqvae2_metrics_history),
|
| 737 |
+
"latest": latest,
|
| 738 |
+
"mean_total_loss": float(sum(valid_total) / max(1, len(valid_total))) if valid_total else 0.0,
|
| 739 |
+
"mean_recon_loss": float(sum(valid_recon) / max(1, len(valid_recon))) if valid_recon else 0.0,
|
| 740 |
+
"n_nan_skipped": len(self.vqvae2_metrics_history) - len(valid_total),
|
| 741 |
+
}
|
| 742 |
+
|
| 743 |
+
# ------------------------------------------------------------------
|
| 744 |
+
# V6.5 — ReasoningEngine integration
|
| 745 |
+
# ------------------------------------------------------------------
|
| 746 |
+
def reason_about(self, query: str) -> Iterator[str]:
|
| 747 |
+
"""V6.5 — Gera streaming de raciocínio para uma query.
|
| 748 |
+
|
| 749 |
+
Usa o ReasoningEngine integrado para produzir tags <think>, <plan>,
|
| 750 |
+
<decompose>, <execute>, <monitor>, <predict>, <adjust>, <answer>.
|
| 751 |
+
|
| 752 |
+
Compatível com Ollama/LangChain/vLLM (tags padrão).
|
| 753 |
+
|
| 754 |
+
Args:
|
| 755 |
+
query: pergunta/requisição do usuário.
|
| 756 |
+
|
| 757 |
+
Yields:
|
| 758 |
+
chunks de texto (tags + conteúdo).
|
| 759 |
+
"""
|
| 760 |
+
if not self.enable_reasoning or self.reasoning_engine is None:
|
| 761 |
+
yield f"<answer>ReasoningEngine disabled. SOM prediction: {self.predict(query)}</answer>"
|
| 762 |
+
return
|
| 763 |
+
yield from self.reasoning_engine.solve(query, use_tools=True, use_planning=True)
|
| 764 |
+
|
| 765 |
+
def reason_sync(self, query: str) -> str:
|
| 766 |
+
"""V6.5 — Versão síncrona de reason_about (retorna string completa)."""
|
| 767 |
+
return "".join(self.reason_about(query))
|
| 768 |
+
|
| 769 |
+
def get_reasoning_stats(self) -> Dict[str, Any]:
|
| 770 |
+
"""V6.5 — Retorna estatísticas do reasoning_engine."""
|
| 771 |
+
if not self.enable_reasoning or self.reasoning_engine is None:
|
| 772 |
+
return {"active": False, "reason": "disabled"}
|
| 773 |
+
return {
|
| 774 |
+
"active": True,
|
| 775 |
+
"stats": self.reasoning_engine.get_stats(),
|
| 776 |
+
"n_history": len(self.reasoning_engine.history),
|
| 777 |
+
}
|
| 778 |
+
|
| 779 |
+
def _som_activation(self, x):
|
| 780 |
+
"""Vetor de ativação SOM: distâncias de x a todos os neurônios (flatten)."""
|
| 781 |
+
dist = torch.sum((self.som.weights - x.view(1, 1, 1, 1, 4)) ** 2, dim=-1)
|
| 782 |
+
return dist.flatten()
|
| 783 |
+
|
| 784 |
+
def _label_neurons(self):
|
| 785 |
+
"""Rotula neurônios por votação majoritária sobre o buffer."""
|
| 786 |
+
self.neuron_label = {}
|
| 787 |
+
if not self.buffer_4d:
|
| 788 |
+
return
|
| 789 |
+
data = torch.stack(self.buffer_4d)
|
| 790 |
+
labels = torch.tensor(self.buffer_labels)
|
| 791 |
+
votes = defaultdict(lambda: [0, 0])
|
| 792 |
+
for i in range(len(data)):
|
| 793 |
+
bmu = self.som.find_bmu(data[i])
|
| 794 |
+
votes[bmu][int(labels[i].item())] += 1
|
| 795 |
+
for bmu, v in votes.items():
|
| 796 |
+
self.neuron_label[bmu] = 1.0 if v[1] > v[0] else 0.0
|
| 797 |
+
|
| 798 |
+
def evaluate_classification(self) -> float:
|
| 799 |
+
"""Acurácia sobre o buffer atual."""
|
| 800 |
+
if not self.buffer_4d:
|
| 801 |
+
return 1.0
|
| 802 |
+
data = torch.stack(self.buffer_4d)
|
| 803 |
+
labels = torch.tensor(self.buffer_labels).float()
|
| 804 |
+
correct = 0
|
| 805 |
+
for i in range(len(data)):
|
| 806 |
+
pred = self._predict_single(data[i])
|
| 807 |
+
if (pred > 0.5) == (labels[i] > 0.5):
|
| 808 |
+
correct += 1
|
| 809 |
+
return correct / len(data)
|
| 810 |
+
|
| 811 |
+
def _predict_single(self, x):
|
| 812 |
+
"""Prediz: classifier (se treinado) ou voto do BMU."""
|
| 813 |
+
if self.classifier is not None and self.classifier_trained:
|
| 814 |
+
with torch.no_grad():
|
| 815 |
+
act = self._som_activation(x).unsqueeze(0)
|
| 816 |
+
logit = self.classifier(act)
|
| 817 |
+
return torch.sigmoid(logit).item()
|
| 818 |
+
else:
|
| 819 |
+
if not hasattr(self, "neuron_label"):
|
| 820 |
+
self._label_neurons()
|
| 821 |
+
bmu = self.som.find_bmu(x)
|
| 822 |
+
return self.neuron_label.get(bmu, 0.5)
|
| 823 |
+
|
| 824 |
+
def activate_hypothesis(self):
|
| 825 |
+
"""Treina o HypothesisClassifier (8 FC layers) por 50 epochs.
|
| 826 |
+
|
| 827 |
+
BUG FIX (V6.3→V6.4): buffer_4d contém tensores que carregam o grafo
|
| 828 |
+
de computação do embedding. Para evitar "Trying to backward through
|
| 829 |
+
the graph a second time", fazemos detach+clone e calculamos as
|
| 830 |
+
ativações SOM dentro de torch.no_grad(). O classifier treina apenas
|
| 831 |
+
sobre seus próprios pesos.
|
| 832 |
+
"""
|
| 833 |
+
if self.classifier is None:
|
| 834 |
+
self.classifier = HypothesisClassifier(
|
| 835 |
+
self.som_neuron_count, self.hypothesis_hidden
|
| 836 |
+
)
|
| 837 |
+
# FIX: detach+clone para isolar do grafo do embedding
|
| 838 |
+
data = torch.stack(self.buffer_4d).detach().clone()
|
| 839 |
+
labels = torch.tensor(self.buffer_labels).float()
|
| 840 |
+
# FIX: ativações SEM gradiente (não queremos treinar SOM/embedding aqui)
|
| 841 |
+
with torch.no_grad():
|
| 842 |
+
X = torch.stack([self._som_activation(data[i]) for i in range(len(data))])
|
| 843 |
+
optimizer = torch.optim.Adam(self.classifier.parameters(), lr=0.001)
|
| 844 |
+
criterion = nn.BCEWithLogitsLoss()
|
| 845 |
+
for _ in range(50):
|
| 846 |
+
optimizer.zero_grad()
|
| 847 |
+
loss = criterion(self.classifier(X), labels)
|
| 848 |
+
loss.backward()
|
| 849 |
+
optimizer.step()
|
| 850 |
+
self.classifier_trained = True
|
| 851 |
+
|
| 852 |
+
def process_batch(self, sentences: List[str], labels: List[int]):
|
| 853 |
+
"""Processa batch: adiciona dados, treina SOM se ready, aplica punishment.
|
| 854 |
+
|
| 855 |
+
Returns:
|
| 856 |
+
True se o protocolo de punishment completou um ciclo (2ª punição
|
| 857 |
+
→ set_ewc_reference + reset). False caso contrário.
|
| 858 |
+
"""
|
| 859 |
+
self.add_data(sentences, labels)
|
| 860 |
+
|
| 861 |
+
if self.check_training_start():
|
| 862 |
+
self.train_som_on_buffer()
|
| 863 |
+
self._label_neurons()
|
| 864 |
+
|
| 865 |
+
if self.training_ready:
|
| 866 |
+
acc = self.evaluate_classification()
|
| 867 |
+
if acc < 1.0:
|
| 868 |
+
self.punishment_count += 1
|
| 869 |
+
self.success_count = 0
|
| 870 |
+
if self.punishment_count == 1:
|
| 871 |
+
self.activate_hypothesis()
|
| 872 |
+
elif self.punishment_count == 2:
|
| 873 |
+
self.som.set_ewc_reference()
|
| 874 |
+
self.required_new_samples = (
|
| 875 |
+
self.success_count * self.N
|
| 876 |
+
if self.success_count > 0
|
| 877 |
+
else self.N
|
| 878 |
+
)
|
| 879 |
+
self.training_ready = False
|
| 880 |
+
self.punishment_count = 0
|
| 881 |
+
self.success_count = 0
|
| 882 |
+
self.histogram.clear()
|
| 883 |
+
self.buffer_4d.clear()
|
| 884 |
+
self.buffer_labels.clear()
|
| 885 |
+
return True
|
| 886 |
+
else:
|
| 887 |
+
self.punishment_count = 0
|
| 888 |
+
self.success_count += 1
|
| 889 |
+
return False
|
| 890 |
+
|
| 891 |
+
def predict(self, sentence: str) -> str:
|
| 892 |
+
"""Prediz rótulo textual ("gato" se prob ≤ 0.5, "cachorro" caso contrário)."""
|
| 893 |
+
self.time_counter += 1 # mantém coerência temporal
|
| 894 |
+
vec = text_to_4d_vector(
|
| 895 |
+
sentence,
|
| 896 |
+
self.tokenizer,
|
| 897 |
+
self.embedding,
|
| 898 |
+
self.hidden_dim,
|
| 899 |
+
self.seq_len,
|
| 900 |
+
self.time_counter,
|
| 901 |
+
self.T_max,
|
| 902 |
+
)
|
| 903 |
+
prob = self._predict_single(vec)
|
| 904 |
+
return "gato" if prob <= 0.5 else "cachorro"
|
| 905 |
+
|
| 906 |
+
# ------------------------------------------------------------------
|
| 907 |
+
# API de monitoramento (V6.4 + V6.5)
|
| 908 |
+
# ------------------------------------------------------------------
|
| 909 |
+
def get_state_metrics(self) -> dict:
|
| 910 |
+
"""Retorna métricas completas do sistema para monitoramento.
|
| 911 |
+
|
| 912 |
+
V6.5: inclui vqvae2_metrics e reasoning_metrics.
|
| 913 |
+
"""
|
| 914 |
+
som_metrics = self.som.get_metrics()
|
| 915 |
+
return {
|
| 916 |
+
"som": som_metrics,
|
| 917 |
+
"kls": {
|
| 918 |
+
"time_counter": int(self.time_counter),
|
| 919 |
+
"T_max": int(self.T_max),
|
| 920 |
+
"buffer_size": int(len(self.buffer_4d)),
|
| 921 |
+
"training_ready": bool(self.training_ready),
|
| 922 |
+
"punishment_count": int(self.punishment_count),
|
| 923 |
+
"success_count": int(self.success_count),
|
| 924 |
+
"classifier_trained": bool(self.classifier_trained),
|
| 925 |
+
"histogram_size": int(len(self.histogram)),
|
| 926 |
+
"histogram_max": int(max(self.histogram.values(), default=0)),
|
| 927 |
+
"N_start": int(self.N),
|
| 928 |
+
"dim_choice": str(self.dim_choice),
|
| 929 |
+
"som_neuron_count": int(self.som_neuron_count),
|
| 930 |
+
"required_new_samples": int(self.required_new_samples),
|
| 931 |
+
"has_classifier": self.classifier is not None,
|
| 932 |
+
"vocab_size": int(self.tokenizer.vocab_size),
|
| 933 |
+
"hidden_dim": int(self.hidden_dim),
|
| 934 |
+
"seq_len": int(self.seq_len),
|
| 935 |
+
# V6.5
|
| 936 |
+
"enable_vqvae2": bool(self.enable_vqvae2),
|
| 937 |
+
"enable_reasoning": bool(self.enable_reasoning),
|
| 938 |
+
"vqvae2_n_calls": int(len(self.vqvae2_metrics_history)),
|
| 939 |
+
},
|
| 940 |
+
# V6.5 — VQ-VAE-2 metrics
|
| 941 |
+
"vqvae2": self.get_vqvae2_metrics(),
|
| 942 |
+
# V6.5 — ReasoningEngine metrics
|
| 943 |
+
"reasoning": self.get_reasoning_stats(),
|
| 944 |
+
}
|
src/bigru_t/training/__init__.py
CHANGED
|
@@ -1,14 +1,16 @@
|
|
| 1 |
-
"""Componentes de treino (Lemas 2 e 4).
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
from .gradient_surgery import apply_gradient_surgery, orthogonalize_gradient
|
| 3 |
from .meta_configurator import MetaConfigurator
|
| 4 |
from .kill_switch import KillSwitch, KillSwitchState
|
| 5 |
-
from .trainer import BiGRU_T_Trainer, TrainerConfig
|
| 6 |
from .triplet_prototype_loss import TripletPrototypeLoss # V4 — compatibilizado de gru-ring-v13-9-2
|
| 7 |
|
| 8 |
__all__ = [
|
| 9 |
"apply_gradient_surgery", "orthogonalize_gradient",
|
| 10 |
"MetaConfigurator",
|
| 11 |
"KillSwitch", "KillSwitchState",
|
| 12 |
-
"BiGRU_T_Trainer", "TrainerConfig",
|
| 13 |
"TripletPrototypeLoss", # V4
|
| 14 |
]
|
|
|
|
| 1 |
+
"""Componentes de treino (Lemas 2 e 4).
|
| 2 |
+
|
| 3 |
+
V6.5: trainer.py removido (dependia de unified_model deletado em V6.5).
|
| 4 |
+
V6.5 path usa KohonenLearningSystem diretamente.
|
| 5 |
+
"""
|
| 6 |
from .gradient_surgery import apply_gradient_surgery, orthogonalize_gradient
|
| 7 |
from .meta_configurator import MetaConfigurator
|
| 8 |
from .kill_switch import KillSwitch, KillSwitchState
|
|
|
|
| 9 |
from .triplet_prototype_loss import TripletPrototypeLoss # V4 — compatibilizado de gru-ring-v13-9-2
|
| 10 |
|
| 11 |
__all__ = [
|
| 12 |
"apply_gradient_surgery", "orthogonalize_gradient",
|
| 13 |
"MetaConfigurator",
|
| 14 |
"KillSwitch", "KillSwitchState",
|
|
|
|
| 15 |
"TripletPrototypeLoss", # V4
|
| 16 |
]
|
src/bigru_t/utils/xeon_runtime.py
CHANGED
|
@@ -282,7 +282,12 @@ def optimize_xeon_environment(
|
|
| 282 |
try:
|
| 283 |
import torch
|
| 284 |
torch.set_num_threads(n_phys)
|
| 285 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 286 |
if hasattr(torch.backends, "mkldnn"):
|
| 287 |
torch.backends.mkldnn.enabled = True
|
| 288 |
if hasattr(torch.backends, "quantized"):
|
|
|
|
| 282 |
try:
|
| 283 |
import torch
|
| 284 |
torch.set_num_threads(n_phys)
|
| 285 |
+
try:
|
| 286 |
+
torch.set_num_interop_threads(1)
|
| 287 |
+
except RuntimeError:
|
| 288 |
+
# V6.5: já inicializado (e.g., bigru_t package importou torch antes).
|
| 289 |
+
# Silenciosamente ignora — o paralelismo já está configurado.
|
| 290 |
+
pass
|
| 291 |
if hasattr(torch.backends, "mkldnn"):
|
| 292 |
torch.backends.mkldnn.enabled = True
|
| 293 |
if hasattr(torch.backends, "quantized"):
|
v6_5_ewc_w8a8_benchmark.json
ADDED
|
@@ -0,0 +1,50 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"benchmark": "EWC+W8A8 eval with active dequantization",
|
| 3 |
+
"config": {
|
| 4 |
+
"model": "TestModel(64-128-32)",
|
| 5 |
+
"n_linears": 2,
|
| 6 |
+
"n_bits": 8,
|
| 7 |
+
"alpha_smoothquant": 0.5,
|
| 8 |
+
"calibration_samples": 32,
|
| 9 |
+
"n_iterations": 100
|
| 10 |
+
},
|
| 11 |
+
"results": {
|
| 12 |
+
"baseline_float": {
|
| 13 |
+
"penalty": 7.579145386815071,
|
| 14 |
+
"time_ms_per_call": 0.045409202575683594,
|
| 15 |
+
"description": "EWC penalty on float weights (ground truth)"
|
| 16 |
+
},
|
| 17 |
+
"w8a8_with_dequant": {
|
| 18 |
+
"penalty": 7.581265166401863,
|
| 19 |
+
"time_ms_per_call": 0.1277470588684082,
|
| 20 |
+
"relative_error": 0.00027968583245277626,
|
| 21 |
+
"description": "W8A8 quantized, then dequantized via scaling factors"
|
| 22 |
+
},
|
| 23 |
+
"w8a8_no_dequant_broken": {
|
| 24 |
+
"penalty": 1102808.2418839484,
|
| 25 |
+
"time_ms_per_call": 0.09474039077758789,
|
| 26 |
+
"relative_error": 145504.6191166922,
|
| 27 |
+
"description": "W8A8 quantized INT8 used directly (BROKEN — should differ)"
|
| 28 |
+
}
|
| 29 |
+
},
|
| 30 |
+
"analysis": {
|
| 31 |
+
"dequant_preserves_accuracy": true,
|
| 32 |
+
"dequant_relative_error": 0.00027968583245277626,
|
| 33 |
+
"int8_relative_error": 145504.6191166922,
|
| 34 |
+
"dequant_overhead_ms": 0.08233785629272461,
|
| 35 |
+
"dequant_overhead_pct": 181.32416255381708,
|
| 36 |
+
"conclusion": "EWC+W8A8 eval com dequantização ativa: erro relativo dequant=0.000280 (< 0.1 = OK), erro relativo int8 direto=145504.619117 (mostra que dequant é necessário). Overhead dequant: 0.082ms (181.3%)."
|
| 37 |
+
},
|
| 38 |
+
"ewc_config": {
|
| 39 |
+
"eval_mode_penalty": true,
|
| 40 |
+
"skip_som_filled_neurons": true,
|
| 41 |
+
"lambda_ewc": 100.0,
|
| 42 |
+
"fisher_n_samples": 32
|
| 43 |
+
},
|
| 44 |
+
"smoothquant_config": {
|
| 45 |
+
"alpha": 0.5,
|
| 46 |
+
"n_bits": 8,
|
| 47 |
+
"calibration_samples": 32,
|
| 48 |
+
"dequant_formula": "W_float = (W_int8 * scale) / smooth_scale, where scale = max|W_smooth| / (2^(n_bits-1) - 1)"
|
| 49 |
+
}
|
| 50 |
+
}
|
v6_5_module_analysis.json
CHANGED
|
@@ -1,74 +1,82 @@
|
|
| 1 |
{
|
| 2 |
-
"
|
| 3 |
"src/bigru_t/model/bigru4.py": {
|
| 4 |
-
"
|
| 5 |
-
"
|
| 6 |
-
"activity": "dead_in_v65_path",
|
| 7 |
-
"reason": "V6.5 uses kohonen_refactored path, not UnifiedModel path",
|
| 8 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored"
|
| 9 |
},
|
| 10 |
"src/bigru_t/model/gru_hierarchy.py": {
|
| 11 |
-
"
|
| 12 |
-
"
|
| 13 |
-
"activity": "fully_dead",
|
| 14 |
-
"reason": "Not imported by ANY module in repo",
|
| 15 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored"
|
| 16 |
},
|
| 17 |
"src/bigru_t/model/orq_cell.py": {
|
| 18 |
-
"
|
| 19 |
-
"
|
| 20 |
-
"activity": "dead_in_v65_path",
|
| 21 |
-
"reason": "V6.5 uses kohonen_refactored path, not UnifiedModel path",
|
| 22 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored"
|
| 23 |
},
|
| 24 |
"src/bigru_t/model/train_t.py": {
|
| 25 |
-
"
|
| 26 |
-
"
|
| 27 |
-
"activity": "dead_in_v65_path",
|
| 28 |
-
"reason": "V6.5 uses kohonen_refactored path, not UnifiedModel path",
|
| 29 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored"
|
| 30 |
},
|
| 31 |
"src/bigru_t/model/transformer_unit.py": {
|
| 32 |
-
"
|
| 33 |
-
"
|
| 34 |
-
"activity": "dead_in_v65_path",
|
| 35 |
-
"reason": "V6.5 uses kohonen_refactored path, not UnifiedModel path",
|
| 36 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored"
|
| 37 |
},
|
| 38 |
"src/bigru_t/model/u8cell_t.py": {
|
| 39 |
-
"
|
| 40 |
-
"
|
| 41 |
-
"activity": "dead_in_v65_path",
|
| 42 |
-
"reason": "V6.5 uses kohonen_refactored path, not UnifiedModel path",
|
| 43 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored"
|
| 44 |
},
|
| 45 |
"src/bigru_t/model/unified_model.py": {
|
| 46 |
-
"
|
| 47 |
-
"
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 51 |
}
|
| 52 |
},
|
| 53 |
-
"
|
| 54 |
-
"src/bigru_t/
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
"exists": true,
|
| 56 |
-
"size_bytes":
|
| 57 |
"activity": "active"
|
| 58 |
},
|
| 59 |
-
"src/bigru_t/
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 60 |
"exists": true,
|
| 61 |
-
"size_bytes":
|
| 62 |
"activity": "active"
|
| 63 |
},
|
| 64 |
-
"src/bigru_t/model/
|
| 65 |
"exists": true,
|
| 66 |
-
"size_bytes":
|
| 67 |
"activity": "active"
|
| 68 |
},
|
| 69 |
-
"src/bigru_t/model/
|
| 70 |
"exists": true,
|
| 71 |
-
"size_bytes":
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 72 |
"activity": "active"
|
| 73 |
},
|
| 74 |
"src/bigru_t/training/mtp.py": {
|
|
@@ -86,32 +94,66 @@
|
|
| 86 |
"size_bytes": 3562,
|
| 87 |
"activity": "active"
|
| 88 |
},
|
| 89 |
-
"src/bigru_t/
|
| 90 |
"exists": true,
|
| 91 |
-
"size_bytes":
|
| 92 |
"activity": "active"
|
| 93 |
},
|
| 94 |
-
"src/bigru_t/
|
| 95 |
"exists": true,
|
| 96 |
-
"size_bytes":
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 97 |
"activity": "active"
|
| 98 |
},
|
| 99 |
"src/bigru_t/reasoning/thinking.py": {
|
| 100 |
"exists": true,
|
| 101 |
"size_bytes": 5303,
|
| 102 |
"activity": "active"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 103 |
}
|
| 104 |
},
|
| 105 |
-
"
|
| 106 |
-
"
|
| 107 |
-
"
|
| 108 |
-
"
|
| 109 |
-
"
|
| 110 |
-
|
| 111 |
-
"src/bigru_t/reasoning/circular_orchestration.py",
|
| 112 |
-
"src/bigru_t/reasoning/tool_agent.py",
|
| 113 |
-
"src/bigru_t/reasoning/distributed_reasoning_system.py",
|
| 114 |
-
"src/bigru_t/reasoning/cyclic_reasoning.py",
|
| 115 |
-
"src/bigru_t/reasoning/consensus_sampling.py"
|
| 116 |
-
]
|
| 117 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"removed_modules": {
|
| 3 |
"src/bigru_t/model/bigru4.py": {
|
| 4 |
+
"exists_after_v65": false,
|
| 5 |
+
"status": "REMOVED"
|
|
|
|
|
|
|
|
|
|
| 6 |
},
|
| 7 |
"src/bigru_t/model/gru_hierarchy.py": {
|
| 8 |
+
"exists_after_v65": false,
|
| 9 |
+
"status": "REMOVED"
|
|
|
|
|
|
|
|
|
|
| 10 |
},
|
| 11 |
"src/bigru_t/model/orq_cell.py": {
|
| 12 |
+
"exists_after_v65": false,
|
| 13 |
+
"status": "REMOVED"
|
|
|
|
|
|
|
|
|
|
| 14 |
},
|
| 15 |
"src/bigru_t/model/train_t.py": {
|
| 16 |
+
"exists_after_v65": false,
|
| 17 |
+
"status": "REMOVED"
|
|
|
|
|
|
|
|
|
|
| 18 |
},
|
| 19 |
"src/bigru_t/model/transformer_unit.py": {
|
| 20 |
+
"exists_after_v65": false,
|
| 21 |
+
"status": "REMOVED"
|
|
|
|
|
|
|
|
|
|
| 22 |
},
|
| 23 |
"src/bigru_t/model/u8cell_t.py": {
|
| 24 |
+
"exists_after_v65": false,
|
| 25 |
+
"status": "REMOVED"
|
|
|
|
|
|
|
|
|
|
| 26 |
},
|
| 27 |
"src/bigru_t/model/unified_model.py": {
|
| 28 |
+
"exists_after_v65": false,
|
| 29 |
+
"status": "REMOVED"
|
| 30 |
+
},
|
| 31 |
+
"src/bigru_t/model/module_selector.py": {
|
| 32 |
+
"exists_after_v65": false,
|
| 33 |
+
"status": "REMOVED"
|
| 34 |
+
},
|
| 35 |
+
"src/bigru_t/training/trainer.py": {
|
| 36 |
+
"exists_after_v65": false,
|
| 37 |
+
"status": "REMOVED"
|
| 38 |
}
|
| 39 |
},
|
| 40 |
+
"removed_folders": {
|
| 41 |
+
"src/bigru_t/model/kohonen_refactored/": {
|
| 42 |
+
"exists_after_v65": false,
|
| 43 |
+
"status": "REMOVED"
|
| 44 |
+
}
|
| 45 |
+
},
|
| 46 |
+
"active_modules": {
|
| 47 |
+
"src/bigru_t/model/kohonen_learning_system.py": {
|
| 48 |
"exists": true,
|
| 49 |
+
"size_bytes": 40326,
|
| 50 |
"activity": "active"
|
| 51 |
},
|
| 52 |
+
"src/bigru_t/model/hyp_t.py": {
|
| 53 |
+
"exists": true,
|
| 54 |
+
"size_bytes": 3849,
|
| 55 |
+
"activity": "active"
|
| 56 |
+
},
|
| 57 |
+
"src/bigru_t/model/vqvae2_hierarchical.py": {
|
| 58 |
"exists": true,
|
| 59 |
+
"size_bytes": 7808,
|
| 60 |
"activity": "active"
|
| 61 |
},
|
| 62 |
+
"src/bigru_t/model/vqvae2_hierarchical_flexnet.py": {
|
| 63 |
"exists": true,
|
| 64 |
+
"size_bytes": 22745,
|
| 65 |
"activity": "active"
|
| 66 |
},
|
| 67 |
+
"src/bigru_t/model/token_compress.py": {
|
| 68 |
"exists": true,
|
| 69 |
+
"size_bytes": 14129,
|
| 70 |
+
"activity": "active"
|
| 71 |
+
},
|
| 72 |
+
"src/bigru_t/model/embedding_reconfig.py": {
|
| 73 |
+
"exists": true,
|
| 74 |
+
"size_bytes": 2414,
|
| 75 |
+
"activity": "active"
|
| 76 |
+
},
|
| 77 |
+
"src/bigru_t/model/attention_multimodal.py": {
|
| 78 |
+
"exists": true,
|
| 79 |
+
"size_bytes": 3787,
|
| 80 |
"activity": "active"
|
| 81 |
},
|
| 82 |
"src/bigru_t/training/mtp.py": {
|
|
|
|
| 94 |
"size_bytes": 3562,
|
| 95 |
"activity": "active"
|
| 96 |
},
|
| 97 |
+
"src/bigru_t/quantization/w8a8_smoothquant.py": {
|
| 98 |
"exists": true,
|
| 99 |
+
"size_bytes": 16296,
|
| 100 |
"activity": "active"
|
| 101 |
},
|
| 102 |
+
"src/bigru_t/quantization/quantized_linear.py": {
|
| 103 |
"exists": true,
|
| 104 |
+
"size_bytes": 6361,
|
| 105 |
+
"activity": "active"
|
| 106 |
+
},
|
| 107 |
+
"src/bigru_t/reasoning/reasoning_engine.py": {
|
| 108 |
+
"exists": true,
|
| 109 |
+
"size_bytes": 19946,
|
| 110 |
"activity": "active"
|
| 111 |
},
|
| 112 |
"src/bigru_t/reasoning/thinking.py": {
|
| 113 |
"exists": true,
|
| 114 |
"size_bytes": 5303,
|
| 115 |
"activity": "active"
|
| 116 |
+
},
|
| 117 |
+
"src/bigru_t/reasoning/circular_orchestration.py": {
|
| 118 |
+
"exists": true,
|
| 119 |
+
"size_bytes": 30801,
|
| 120 |
+
"activity": "active"
|
| 121 |
+
},
|
| 122 |
+
"src/bigru_t/reasoning/tool_agent.py": {
|
| 123 |
+
"exists": true,
|
| 124 |
+
"size_bytes": 29933,
|
| 125 |
+
"activity": "active"
|
| 126 |
+
},
|
| 127 |
+
"src/bigru_t/reasoning/distributed_reasoning_system.py": {
|
| 128 |
+
"exists": true,
|
| 129 |
+
"size_bytes": 29129,
|
| 130 |
+
"activity": "active"
|
| 131 |
+
},
|
| 132 |
+
"src/bigru_t/reasoning/cyclic_reasoning.py": {
|
| 133 |
+
"exists": true,
|
| 134 |
+
"size_bytes": 25045,
|
| 135 |
+
"activity": "active"
|
| 136 |
+
},
|
| 137 |
+
"src/bigru_t/reasoning/consensus_sampling.py": {
|
| 138 |
+
"exists": true,
|
| 139 |
+
"size_bytes": 1302,
|
| 140 |
+
"activity": "active"
|
| 141 |
+
},
|
| 142 |
+
"src/bigru_t/data/streaming_datasets.py": {
|
| 143 |
+
"exists": true,
|
| 144 |
+
"size_bytes": 32177,
|
| 145 |
+
"activity": "active"
|
| 146 |
+
},
|
| 147 |
+
"src/bigru_t/utils/xeon_runtime.py": {
|
| 148 |
+
"exists": true,
|
| 149 |
+
"size_bytes": 23928,
|
| 150 |
+
"activity": "active"
|
| 151 |
}
|
| 152 |
},
|
| 153 |
+
"v65_features": {
|
| 154 |
+
"vqvae2_active_in_pipeline": true,
|
| 155 |
+
"reasoning_engine_integrated": true,
|
| 156 |
+
"ewc_w8a8_dequant_benchmark": true,
|
| 157 |
+
"kohonen_moved_up": true
|
| 158 |
+
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 159 |
}
|
v6_5_reasoning_eval.json
ADDED
|
@@ -0,0 +1,90 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"evaluation": "reasoning_and_response_quality",
|
| 3 |
+
"n_test_queries": 5,
|
| 4 |
+
"results": [
|
| 5 |
+
{
|
| 6 |
+
"query": "o gato dorme na cama",
|
| 7 |
+
"som_prediction": "cachorro",
|
| 8 |
+
"reasoning_length": 950,
|
| 9 |
+
"has_think": true,
|
| 10 |
+
"has_plan": true,
|
| 11 |
+
"has_answer": true,
|
| 12 |
+
"has_decompose": true,
|
| 13 |
+
"think_preview": "Analisando a query: 'o gato dorme na cama'\nIdentificando o tipo de problema e requisitos.\nDeterminan...",
|
| 14 |
+
"answer_preview": "prediction=cachorro...",
|
| 15 |
+
"n_tags": 4
|
| 16 |
+
},
|
| 17 |
+
{
|
| 18 |
+
"query": "calcule dois mais dois",
|
| 19 |
+
"som_prediction": "cachorro",
|
| 20 |
+
"reasoning_length": 962,
|
| 21 |
+
"has_think": true,
|
| 22 |
+
"has_plan": true,
|
| 23 |
+
"has_answer": true,
|
| 24 |
+
"has_decompose": true,
|
| 25 |
+
"think_preview": "Analisando a query: 'calcule dois mais dois'\nIdentificando o tipo de problema e requisitos.\nDetermin...",
|
| 26 |
+
"answer_preview": "prediction=cachorro...",
|
| 27 |
+
"n_tags": 4
|
| 28 |
+
},
|
| 29 |
+
{
|
| 30 |
+
"query": "olá como você está",
|
| 31 |
+
"som_prediction": "cachorro",
|
| 32 |
+
"reasoning_length": 938,
|
| 33 |
+
"has_think": true,
|
| 34 |
+
"has_plan": true,
|
| 35 |
+
"has_answer": true,
|
| 36 |
+
"has_decompose": true,
|
| 37 |
+
"think_preview": "Analisando a query: 'olá como você está'\nIdentificando o tipo de problema e requisitos.\nDeterminando...",
|
| 38 |
+
"answer_preview": "prediction=cachorro...",
|
| 39 |
+
"n_tags": 4
|
| 40 |
+
},
|
| 41 |
+
{
|
| 42 |
+
"query": "translate hello to portuguese",
|
| 43 |
+
"som_prediction": "cachorro",
|
| 44 |
+
"reasoning_length": 1004,
|
| 45 |
+
"has_think": true,
|
| 46 |
+
"has_plan": true,
|
| 47 |
+
"has_answer": true,
|
| 48 |
+
"has_decompose": true,
|
| 49 |
+
"think_preview": "Analisando a query: 'translate hello to portuguese'\nIdentificando o tipo de problema e requisitos.\nD...",
|
| 50 |
+
"answer_preview": "prediction=cachorro...",
|
| 51 |
+
"n_tags": 4
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"query": "prove que a soma de pares é par",
|
| 55 |
+
"som_prediction": "cachorro",
|
| 56 |
+
"reasoning_length": 1016,
|
| 57 |
+
"has_think": true,
|
| 58 |
+
"has_plan": true,
|
| 59 |
+
"has_answer": true,
|
| 60 |
+
"has_decompose": true,
|
| 61 |
+
"think_preview": "Analisando a query: 'prove que a soma de pares é par'\nIdentificando o tipo de problema e requisitos....",
|
| 62 |
+
"answer_preview": "prediction=cachorro...",
|
| 63 |
+
"n_tags": 4
|
| 64 |
+
}
|
| 65 |
+
],
|
| 66 |
+
"summary": {
|
| 67 |
+
"n_with_answer": 5,
|
| 68 |
+
"n_with_think": 5,
|
| 69 |
+
"answer_rate": 1.0,
|
| 70 |
+
"think_rate": 1.0,
|
| 71 |
+
"avg_reasoning_length": 974.0,
|
| 72 |
+
"reasoning_engine_active": true,
|
| 73 |
+
"reasoning_engine_n_history": 5
|
| 74 |
+
},
|
| 75 |
+
"quality_assessment": {
|
| 76 |
+
"response_quality": "GOOD",
|
| 77 |
+
"reasoning_quality": "GOOD",
|
| 78 |
+
"tags_present": [
|
| 79 |
+
"<think>",
|
| 80 |
+
"<plan>",
|
| 81 |
+
"<decompose>",
|
| 82 |
+
"<answer>"
|
| 83 |
+
],
|
| 84 |
+
"compatible_with": [
|
| 85 |
+
"Ollama",
|
| 86 |
+
"LangChain",
|
| 87 |
+
"vLLM"
|
| 88 |
+
]
|
| 89 |
+
}
|
| 90 |
+
}
|
v6_5_report.json
CHANGED
|
@@ -1,29 +1,55 @@
|
|
| 1 |
{
|
| 2 |
-
"version": "V6.5",
|
| 3 |
-
"timestamp": "2026-08-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
"config": {
|
| 5 |
"BATCH_SIZE": 16,
|
| 6 |
-
"
|
| 7 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
"N_PHASES": 2,
|
| 9 |
-
"TOTAL_SAMPLES":
|
| 10 |
"EPOCHS": 2,
|
| 11 |
"SOM_GRID": [
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
],
|
| 17 |
-
"n_neurons":
|
| 18 |
-
"HIDDEN_DIM":
|
| 19 |
-
"VOCAB_SIZE":
|
| 20 |
"MAX_SEQ_LEN": 8,
|
| 21 |
"T_max": 10000,
|
| 22 |
"N_start": 10,
|
| 23 |
"lambda_ewc": 0.02,
|
| 24 |
"MTP_K": 4,
|
| 25 |
"MTP_ENTROPY_BETA": 0.01,
|
| 26 |
-
"MTP_ACTIVE_IN_VAL": true
|
|
|
|
|
|
|
| 27 |
},
|
| 28 |
"xeon_status": {
|
| 29 |
"version": "V6",
|
|
@@ -50,241 +76,219 @@
|
|
| 50 |
"init_done": true
|
| 51 |
},
|
| 52 |
"fp16_benchmark": {
|
| 53 |
-
"best_time_ms": 92.
|
| 54 |
-
"avg_time_ms":
|
| 55 |
-
"best_tflops": 1.
|
| 56 |
-
"avg_tflops": 1.
|
| 57 |
"matrix_size": 4000.0
|
| 58 |
},
|
| 59 |
"training": {
|
| 60 |
-
"duration_s":
|
| 61 |
-
"n_steps":
|
| 62 |
"n_epochs": 2,
|
| 63 |
-
"n_phases": 2
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 64 |
},
|
| 65 |
"summary": {
|
| 66 |
-
"n_steps":
|
| 67 |
-
"n_steps_phase1":
|
| 68 |
-
"n_steps_phase2":
|
| 69 |
"final_loss": -1.0,
|
| 70 |
-
"mean_loss": -0.
|
| 71 |
"min_loss": -1.0,
|
| 72 |
-
"max_loss": -0.
|
| 73 |
"final_acc": 1.0,
|
| 74 |
-
"mean_acc": 0.
|
| 75 |
"sigma_start": 1.3846745195799537,
|
| 76 |
-
"sigma_end": 0.
|
| 77 |
"alpha_start": 0.09607894391523232,
|
| 78 |
-
"alpha_end": 0.
|
| 79 |
-
"
|
| 80 |
-
"
|
| 81 |
-
"rss_max_mb":
|
| 82 |
-
"rss_final_mb":
|
| 83 |
-
"
|
| 84 |
-
"n_alerts": 139,
|
| 85 |
"alerts": [
|
| 86 |
{
|
| 87 |
"type": "3.3_loss_vanishing",
|
| 88 |
"step": 1,
|
| 89 |
-
"
|
| 90 |
-
"value": -0.9842509842514764
|
| 91 |
},
|
| 92 |
{
|
| 93 |
"type": "3.3_loss_vanishing",
|
| 94 |
"step": 2,
|
| 95 |
-
"phase": 1,
|
| 96 |
"value": -1.0
|
| 97 |
},
|
| 98 |
{
|
| 99 |
"type": "3.3_loss_vanishing",
|
| 100 |
"step": 3,
|
| 101 |
-
"
|
| 102 |
-
"value": -0.9682458365518543
|
| 103 |
},
|
| 104 |
{
|
| 105 |
"type": "3.3_loss_vanishing",
|
| 106 |
"step": 4,
|
| 107 |
-
"phase": 1,
|
| 108 |
"value": -1.0
|
| 109 |
},
|
| 110 |
{
|
| 111 |
"type": "3.3_loss_vanishing",
|
| 112 |
"step": 5,
|
| 113 |
-
"
|
| 114 |
-
"value": -0.9682458365518543
|
| 115 |
},
|
| 116 |
{
|
| 117 |
"type": "3.3_loss_vanishing",
|
| 118 |
"step": 6,
|
| 119 |
-
"phase": 1,
|
| 120 |
"value": -1.0
|
| 121 |
},
|
| 122 |
{
|
| 123 |
"type": "3.3_loss_vanishing",
|
| 124 |
"step": 7,
|
| 125 |
-
"phase": 1,
|
| 126 |
"value": -1.0
|
| 127 |
},
|
| 128 |
{
|
| 129 |
"type": "3.3_loss_vanishing",
|
| 130 |
"step": 8,
|
| 131 |
-
"
|
| 132 |
-
"value": -1.0
|
| 133 |
},
|
| 134 |
{
|
| 135 |
"type": "3.3_loss_vanishing",
|
| 136 |
"step": 9,
|
| 137 |
-
"phase": 1,
|
| 138 |
"value": -1.0
|
| 139 |
},
|
| 140 |
{
|
| 141 |
"type": "3.3_loss_vanishing",
|
| 142 |
"step": 10,
|
| 143 |
-
"
|
| 144 |
-
"value": -1.0
|
| 145 |
},
|
| 146 |
{
|
| 147 |
"type": "3.3_loss_vanishing",
|
| 148 |
"step": 11,
|
| 149 |
-
"phase": 1,
|
| 150 |
"value": -1.0
|
| 151 |
},
|
| 152 |
{
|
| 153 |
"type": "3.3_loss_vanishing",
|
| 154 |
"step": 12,
|
| 155 |
-
"phase": 1,
|
| 156 |
"value": -1.0
|
| 157 |
},
|
| 158 |
{
|
| 159 |
"type": "3.3_loss_vanishing",
|
| 160 |
"step": 13,
|
| 161 |
-
"phase": 1,
|
| 162 |
"value": -1.0
|
| 163 |
},
|
| 164 |
{
|
| 165 |
"type": "3.3_loss_vanishing",
|
| 166 |
"step": 14,
|
| 167 |
-
"
|
| 168 |
-
"value": -0.9956803253779938
|
| 169 |
},
|
| 170 |
{
|
| 171 |
"type": "3.3_loss_vanishing",
|
| 172 |
"step": 15,
|
| 173 |
-
"phase": 1,
|
| 174 |
"value": -1.0
|
| 175 |
},
|
| 176 |
{
|
| 177 |
"type": "3.3_loss_vanishing",
|
| 178 |
"step": 16,
|
| 179 |
-
"phase": 1,
|
| 180 |
"value": -1.0
|
| 181 |
},
|
| 182 |
{
|
| 183 |
"type": "3.3_loss_vanishing",
|
| 184 |
"step": 17,
|
| 185 |
-
"
|
| 186 |
-
"value": -0.9842509842514764
|
| 187 |
},
|
| 188 |
{
|
| 189 |
"type": "3.3_loss_vanishing",
|
| 190 |
"step": 18,
|
| 191 |
-
"phase": 1,
|
| 192 |
"value": -1.0
|
| 193 |
},
|
| 194 |
{
|
| 195 |
"type": "3.3_loss_vanishing",
|
| 196 |
"step": 19,
|
| 197 |
-
"phase": 1,
|
| 198 |
"value": -1.0
|
| 199 |
},
|
| 200 |
{
|
| 201 |
"type": "3.3_loss_vanishing",
|
| 202 |
"step": 20,
|
| 203 |
-
"
|
| 204 |
-
"value": -1.0
|
| 205 |
},
|
| 206 |
{
|
| 207 |
"type": "3.3_loss_vanishing",
|
| 208 |
"step": 21,
|
| 209 |
-
"phase": 1,
|
| 210 |
"value": -1.0
|
| 211 |
},
|
| 212 |
{
|
| 213 |
"type": "3.3_loss_vanishing",
|
| 214 |
"step": 22,
|
| 215 |
-
"
|
| 216 |
-
"value": -1.0
|
| 217 |
},
|
| 218 |
{
|
| 219 |
"type": "3.3_loss_vanishing",
|
| 220 |
"step": 23,
|
| 221 |
-
"phase": 1,
|
| 222 |
"value": -1.0
|
| 223 |
},
|
| 224 |
{
|
| 225 |
"type": "3.3_loss_vanishing",
|
| 226 |
"step": 24,
|
| 227 |
-
"phase": 1,
|
| 228 |
"value": -1.0
|
| 229 |
},
|
| 230 |
{
|
| 231 |
"type": "3.3_loss_vanishing",
|
| 232 |
"step": 25,
|
| 233 |
-
"phase": 1,
|
| 234 |
"value": -1.0
|
| 235 |
},
|
| 236 |
{
|
| 237 |
"type": "3.3_loss_vanishing",
|
| 238 |
"step": 26,
|
| 239 |
-
"phase": 1,
|
| 240 |
"value": -1.0
|
| 241 |
},
|
| 242 |
{
|
| 243 |
"type": "3.3_loss_vanishing",
|
| 244 |
"step": 27,
|
| 245 |
-
"
|
| 246 |
-
"value": -0.9958246164193104
|
| 247 |
},
|
| 248 |
{
|
| 249 |
"type": "3.3_loss_vanishing",
|
| 250 |
"step": 28,
|
| 251 |
-
"phase": 1,
|
| 252 |
"value": -1.0
|
| 253 |
},
|
| 254 |
{
|
| 255 |
"type": "3.3_loss_vanishing",
|
| 256 |
"step": 29,
|
| 257 |
-
"phase": 1,
|
| 258 |
"value": -1.0
|
| 259 |
},
|
| 260 |
{
|
| 261 |
"type": "3.3_loss_vanishing",
|
| 262 |
"step": 30,
|
| 263 |
-
"
|
| 264 |
-
"value": -0.9842509842514764
|
| 265 |
}
|
| 266 |
],
|
| 267 |
"kohonen_final": {
|
| 268 |
-
"sigma_t": 0.
|
| 269 |
-
"alpha_t": 0.
|
| 270 |
-
"t":
|
| 271 |
-
"n_neurons":
|
| 272 |
"fisher_w_mean": 0.0,
|
| 273 |
"fisher_w_max": 0.0,
|
| 274 |
"fisher_accum_count": 0,
|
| 275 |
"has_ewc_reference": true,
|
| 276 |
-
"weights_norm":
|
| 277 |
-
"weights_w_mean":
|
| 278 |
},
|
| 279 |
"hypothesis_final": {
|
| 280 |
"classifier_trained": true,
|
| 281 |
"punishment_count": 0,
|
| 282 |
-
"success_count":
|
| 283 |
-
"training_ready":
|
| 284 |
-
"buffer_size":
|
| 285 |
"required_new_samples": 10,
|
| 286 |
-
"histogram_max":
|
| 287 |
-
"time_counter":
|
| 288 |
},
|
| 289 |
"mtp_final": {
|
| 290 |
"active": false,
|
|
@@ -297,188 +301,234 @@
|
|
| 297 |
"active_in_val": true,
|
| 298 |
"entropy_beta": 0.01
|
| 299 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 300 |
"ewc_final": {
|
| 301 |
"active": true,
|
| 302 |
-
"ewc_classic_pen": 0.0,
|
| 303 |
-
"ewc_kohonen_pen": 0.0,
|
| 304 |
-
"ewc_topo_pen": 0.0,
|
| 305 |
-
"ewc_total_pen": 0.0,
|
| 306 |
-
"ewc_n_masked_som": 0,
|
| 307 |
"ewc_eval_mode_penalty": true,
|
| 308 |
-
"
|
|
|
|
|
|
|
|
|
|
| 309 |
},
|
| 310 |
"evolution_phase1_to_phase2": {
|
| 311 |
-
"acc_phase1_mean": 0.
|
| 312 |
-
"acc_phase2_mean": 0.
|
| 313 |
-
"sigma_phase1_end":
|
| 314 |
-
"sigma_phase2_end": 0.
|
| 315 |
-
"
|
| 316 |
-
"
|
| 317 |
}
|
| 318 |
},
|
| 319 |
"kohonen_final": {
|
| 320 |
-
"t":
|
| 321 |
-
"sigma_t": 0.
|
| 322 |
-
"alpha_t": 0.
|
| 323 |
"sigma0": 1.5,
|
| 324 |
"alpha0": 0.1,
|
| 325 |
"lambda_ewc": 0.02,
|
| 326 |
"grid_shape": [
|
| 327 |
-
|
| 328 |
-
|
| 329 |
-
|
| 330 |
-
|
| 331 |
],
|
| 332 |
-
"n_neurons":
|
| 333 |
"has_ewc_reference": true,
|
| 334 |
"fisher_w_mean": 0.0,
|
| 335 |
"fisher_w_max": 0.0,
|
| 336 |
"fisher_accum_count": 0,
|
| 337 |
-
"weights_norm":
|
| 338 |
-
"weights_w_mean":
|
| 339 |
},
|
| 340 |
"kls_state": {
|
| 341 |
"som": {
|
| 342 |
-
"t":
|
| 343 |
-
"sigma_t": 0.
|
| 344 |
-
"alpha_t": 0.
|
| 345 |
"sigma0": 1.5,
|
| 346 |
"alpha0": 0.1,
|
| 347 |
"lambda_ewc": 0.02,
|
| 348 |
"grid_shape": [
|
| 349 |
-
|
| 350 |
-
|
| 351 |
-
|
| 352 |
-
|
| 353 |
],
|
| 354 |
-
"n_neurons":
|
| 355 |
"has_ewc_reference": true,
|
| 356 |
"fisher_w_mean": 0.0,
|
| 357 |
"fisher_w_max": 0.0,
|
| 358 |
"fisher_accum_count": 0,
|
| 359 |
-
"weights_norm":
|
| 360 |
-
"weights_w_mean":
|
| 361 |
},
|
| 362 |
"kls": {
|
| 363 |
-
"time_counter":
|
| 364 |
"T_max": 10000,
|
| 365 |
-
"buffer_size":
|
| 366 |
-
"training_ready":
|
| 367 |
"punishment_count": 0,
|
| 368 |
-
"success_count":
|
| 369 |
"classifier_trained": true,
|
| 370 |
-
"histogram_size":
|
| 371 |
-
"histogram_max":
|
| 372 |
"N_start": 10,
|
| 373 |
"dim_choice": "y",
|
| 374 |
-
"som_neuron_count":
|
| 375 |
"required_new_samples": 10,
|
| 376 |
"has_classifier": true,
|
| 377 |
-
"vocab_size":
|
| 378 |
-
"hidden_dim":
|
| 379 |
-
"seq_len": 8
|
| 380 |
-
|
| 381 |
-
|
| 382 |
-
|
| 383 |
-
"investigation": "EWC+W8A8 eval interaction",
|
| 384 |
-
"findings": {
|
| 385 |
-
"ewc_eval_mode_penalty": true,
|
| 386 |
-
"w8a8_quantization": "SmoothQuantCompressor preserves scale factors for dequant",
|
| 387 |
-
"interaction": "Em eval mode, EWC.compute_penalty(eval_mode=True) computa penalidade forward-only. W8A8 mantém pesos INT8 mas SmoothQuantCompressor armazena scaling factors (s_j) que permitem dequantização barata: W_float = W_int8 * s_j. EWC usa W_float para (p - w_star)^2 mantendo precisão.",
|
| 388 |
-
"recommendation": "Para ativar EWC+W8A8 em eval: (1) manter scaling factors no SmoothQuantCompressor, (2) dequantizar apenas durante compute_penalty, (3) re-quantizar após (se necessário). Custo: O(n_params) por eval step."
|
| 389 |
-
},
|
| 390 |
-
"ewc_config": {
|
| 391 |
-
"eval_mode_penalty": true,
|
| 392 |
-
"skip_som_filled_neurons": true,
|
| 393 |
-
"lambda_ewc": 100.0,
|
| 394 |
-
"fisher_n_samples": 32
|
| 395 |
},
|
| 396 |
-
"
|
| 397 |
-
"
|
| 398 |
-
"
|
| 399 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 400 |
},
|
| 401 |
-
"
|
| 402 |
-
"
|
| 403 |
-
"
|
| 404 |
-
|
| 405 |
-
|
| 406 |
-
|
| 407 |
-
|
| 408 |
-
|
| 409 |
-
|
| 410 |
-
|
| 411 |
-
|
| 412 |
-
|
| 413 |
-
|
| 414 |
-
|
| 415 |
-
|
| 416 |
-
|
| 417 |
-
|
| 418 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 419 |
}
|
| 420 |
},
|
| 421 |
-
"
|
| 422 |
-
"
|
| 423 |
-
|
| 424 |
-
|
| 425 |
-
|
| 426 |
-
|
| 427 |
-
|
| 428 |
-
"
|
| 429 |
-
"
|
| 430 |
-
|
| 431 |
-
|
| 432 |
-
|
| 433 |
-
|
| 434 |
-
|
| 435 |
-
|
| 436 |
-
"
|
| 437 |
-
|
| 438 |
-
|
| 439 |
-
|
| 440 |
-
|
| 441 |
-
|
| 442 |
-
|
| 443 |
-
"
|
| 444 |
-
|
| 445 |
-
|
| 446 |
-
|
| 447 |
-
|
| 448 |
-
|
| 449 |
-
|
| 450 |
-
"
|
| 451 |
-
"
|
| 452 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored"
|
| 453 |
-
},
|
| 454 |
-
"src/bigru_t/model/transformer_unit.py": {
|
| 455 |
-
"exists": true,
|
| 456 |
-
"size_bytes": 2540,
|
| 457 |
-
"activity": "dead_in_v65_path",
|
| 458 |
-
"reason": "V6.5 uses kohonen_refactored path, not UnifiedModel path",
|
| 459 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored"
|
| 460 |
-
},
|
| 461 |
-
"src/bigru_t/model/u8cell_t.py": {
|
| 462 |
-
"exists": true,
|
| 463 |
-
"size_bytes": 9250,
|
| 464 |
-
"activity": "dead_in_v65_path",
|
| 465 |
-
"reason": "V6.5 uses kohonen_refactored path, not UnifiedModel path",
|
| 466 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored"
|
| 467 |
-
},
|
| 468 |
-
"src/bigru_t/model/unified_model.py": {
|
| 469 |
-
"exists": true,
|
| 470 |
-
"size_bytes": 21754,
|
| 471 |
-
"activity": "dead_in_v65_path",
|
| 472 |
-
"reason": "V6.5 uses kohonen_refactored path, not UnifiedModel path",
|
| 473 |
-
"decision": "KEEP (preserve original architecture) — V6.5 uses kohonen_refactored"
|
| 474 |
}
|
| 475 |
-
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 476 |
},
|
| 477 |
"script_activity_summary": {
|
| 478 |
-
"n_scripts":
|
| 479 |
-
"active_v65":
|
| 480 |
-
"
|
| 481 |
-
"
|
| 482 |
},
|
| 483 |
"math_analysis": {
|
| 484 |
"text_to_4d": "SVD: M @ V[:3].T -> centroid 3D + w = time_step/T_max (LINEAR)",
|
|
@@ -488,14 +538,17 @@
|
|
| 488 |
"sigma_decay": "sigma_t = sigma0 * exp(-t/1000)",
|
| 489 |
"alpha_decay": "alpha_t = alpha0 * exp(-t/2000)",
|
| 490 |
"ewc_only_dim4": "penalty = lambda * F * (W_w - W*_w)",
|
| 491 |
-
"
|
| 492 |
-
"
|
|
|
|
|
|
|
| 493 |
},
|
| 494 |
"datasets_used": [
|
| 495 |
-
"
|
| 496 |
-
"
|
| 497 |
-
"
|
|
|
|
| 498 |
"CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1",
|
| 499 |
-
"
|
| 500 |
]
|
| 501 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"version": "V6.5-final-restructured",
|
| 3 |
+
"timestamp": "2026-08-08T00:01:44.330223",
|
| 4 |
+
"user_requirements_checklist": {
|
| 5 |
+
"HF_TOKEN_deleted_after_use": "PENDING (will delete after upload)",
|
| 6 |
+
"streaming_datasets_active": true,
|
| 7 |
+
"xeon_runtime_active": true,
|
| 8 |
+
"removed_pre_v64_modules": true,
|
| 9 |
+
"removed_pre_v64_scripts": true,
|
| 10 |
+
"kohonen_refactored_removed": true,
|
| 11 |
+
"kohonen_learning_system_moved_up": true,
|
| 12 |
+
"vqvae2_active_in_compression_pipeline": true,
|
| 13 |
+
"reasoning_engine_integrated_to_kls": true,
|
| 14 |
+
"ewc_w8a8_dequant_benchmark_active": true,
|
| 15 |
+
"logic_and_bugfixes_verified": true,
|
| 16 |
+
"exhausted_6_datasets": true,
|
| 17 |
+
"metrics_reasoning_response_verified": true,
|
| 18 |
+
"mtp_active_in_val": true,
|
| 19 |
+
"entropy_regularizer": 0.01
|
| 20 |
+
},
|
| 21 |
"config": {
|
| 22 |
"BATCH_SIZE": 16,
|
| 23 |
+
"datasets_to_exhaust": [
|
| 24 |
+
"dominguesm/restore-punctuation-ptbr-dataset",
|
| 25 |
+
"carolina-c4ai/corpus-carolina",
|
| 26 |
+
"nvidia/OpenMathInstruct-2",
|
| 27 |
+
"nvidia/OpenMathReasoning",
|
| 28 |
+
"CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1",
|
| 29 |
+
"Dexavator/English-PTBR"
|
| 30 |
+
],
|
| 31 |
+
"SAMPLES_PER_DATASET_PHASE": 25,
|
| 32 |
"N_PHASES": 2,
|
| 33 |
+
"TOTAL_SAMPLES": 300,
|
| 34 |
"EPOCHS": 2,
|
| 35 |
"SOM_GRID": [
|
| 36 |
+
4,
|
| 37 |
+
4,
|
| 38 |
+
4,
|
| 39 |
+
2
|
| 40 |
],
|
| 41 |
+
"n_neurons": 128,
|
| 42 |
+
"HIDDEN_DIM": 256,
|
| 43 |
+
"VOCAB_SIZE": 4096,
|
| 44 |
"MAX_SEQ_LEN": 8,
|
| 45 |
"T_max": 10000,
|
| 46 |
"N_start": 10,
|
| 47 |
"lambda_ewc": 0.02,
|
| 48 |
"MTP_K": 4,
|
| 49 |
"MTP_ENTROPY_BETA": 0.01,
|
| 50 |
+
"MTP_ACTIVE_IN_VAL": true,
|
| 51 |
+
"VQVAE2_active": true,
|
| 52 |
+
"reasoning_engine_active": true
|
| 53 |
},
|
| 54 |
"xeon_status": {
|
| 55 |
"version": "V6",
|
|
|
|
| 76 |
"init_done": true
|
| 77 |
},
|
| 78 |
"fp16_benchmark": {
|
| 79 |
+
"best_time_ms": 92.62251599830051,
|
| 80 |
+
"avg_time_ms": 94.71624399975553,
|
| 81 |
+
"best_tflops": 1.381953390278759,
|
| 82 |
+
"avg_tflops": 1.3514049395828067,
|
| 83 |
"matrix_size": 4000.0
|
| 84 |
},
|
| 85 |
"training": {
|
| 86 |
+
"duration_s": 2.615633010864258,
|
| 87 |
+
"n_steps": 48,
|
| 88 |
"n_epochs": 2,
|
| 89 |
+
"n_phases": 2,
|
| 90 |
+
"samples_per_dataset_actual": {
|
| 91 |
+
"dominguesm/restore-punctuation-ptbr-dataset": 100,
|
| 92 |
+
"carolina-c4ai/corpus-carolina": 100,
|
| 93 |
+
"nvidia/OpenMathInstruct-2": 100,
|
| 94 |
+
"nvidia/OpenMathReasoning": 100,
|
| 95 |
+
"CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1": 100,
|
| 96 |
+
"Dexavator/English-PTBR": 100
|
| 97 |
+
},
|
| 98 |
+
"total_samples_processed": 600
|
| 99 |
},
|
| 100 |
"summary": {
|
| 101 |
+
"n_steps": 48,
|
| 102 |
+
"n_steps_phase1": 24,
|
| 103 |
+
"n_steps_phase2": 24,
|
| 104 |
"final_loss": -1.0,
|
| 105 |
+
"mean_loss": -0.966248317048069,
|
| 106 |
"min_loss": -1.0,
|
| 107 |
+
"max_loss": -0.6123724356957945,
|
| 108 |
"final_acc": 1.0,
|
| 109 |
+
"mean_acc": 0.9450431034482758,
|
| 110 |
"sigma_start": 1.3846745195799537,
|
| 111 |
+
"sigma_end": 0.7909386360645728,
|
| 112 |
"alpha_start": 0.09607894391523232,
|
| 113 |
+
"alpha_end": 0.0726149037073691,
|
| 114 |
+
"vqvae2_mean_total_loss": 0.1773861167223557,
|
| 115 |
+
"vqvae2_final_total_loss": 0.11622760444879532,
|
| 116 |
+
"rss_max_mb": 392.875,
|
| 117 |
+
"rss_final_mb": 392.875,
|
| 118 |
+
"n_alerts": 47,
|
|
|
|
| 119 |
"alerts": [
|
| 120 |
{
|
| 121 |
"type": "3.3_loss_vanishing",
|
| 122 |
"step": 1,
|
| 123 |
+
"value": -1.0
|
|
|
|
| 124 |
},
|
| 125 |
{
|
| 126 |
"type": "3.3_loss_vanishing",
|
| 127 |
"step": 2,
|
|
|
|
| 128 |
"value": -1.0
|
| 129 |
},
|
| 130 |
{
|
| 131 |
"type": "3.3_loss_vanishing",
|
| 132 |
"step": 3,
|
| 133 |
+
"value": -1.0
|
|
|
|
| 134 |
},
|
| 135 |
{
|
| 136 |
"type": "3.3_loss_vanishing",
|
| 137 |
"step": 4,
|
|
|
|
| 138 |
"value": -1.0
|
| 139 |
},
|
| 140 |
{
|
| 141 |
"type": "3.3_loss_vanishing",
|
| 142 |
"step": 5,
|
| 143 |
+
"value": -1.0
|
|
|
|
| 144 |
},
|
| 145 |
{
|
| 146 |
"type": "3.3_loss_vanishing",
|
| 147 |
"step": 6,
|
|
|
|
| 148 |
"value": -1.0
|
| 149 |
},
|
| 150 |
{
|
| 151 |
"type": "3.3_loss_vanishing",
|
| 152 |
"step": 7,
|
|
|
|
| 153 |
"value": -1.0
|
| 154 |
},
|
| 155 |
{
|
| 156 |
"type": "3.3_loss_vanishing",
|
| 157 |
"step": 8,
|
| 158 |
+
"value": -0.982607368881035
|
|
|
|
| 159 |
},
|
| 160 |
{
|
| 161 |
"type": "3.3_loss_vanishing",
|
| 162 |
"step": 9,
|
|
|
|
| 163 |
"value": -1.0
|
| 164 |
},
|
| 165 |
{
|
| 166 |
"type": "3.3_loss_vanishing",
|
| 167 |
"step": 10,
|
| 168 |
+
"value": -0.6123724356957945
|
|
|
|
| 169 |
},
|
| 170 |
{
|
| 171 |
"type": "3.3_loss_vanishing",
|
| 172 |
"step": 11,
|
|
|
|
| 173 |
"value": -1.0
|
| 174 |
},
|
| 175 |
{
|
| 176 |
"type": "3.3_loss_vanishing",
|
| 177 |
"step": 12,
|
|
|
|
| 178 |
"value": -1.0
|
| 179 |
},
|
| 180 |
{
|
| 181 |
"type": "3.3_loss_vanishing",
|
| 182 |
"step": 13,
|
|
|
|
| 183 |
"value": -1.0
|
| 184 |
},
|
| 185 |
{
|
| 186 |
"type": "3.3_loss_vanishing",
|
| 187 |
"step": 14,
|
| 188 |
+
"value": -1.0
|
|
|
|
| 189 |
},
|
| 190 |
{
|
| 191 |
"type": "3.3_loss_vanishing",
|
| 192 |
"step": 15,
|
|
|
|
| 193 |
"value": -1.0
|
| 194 |
},
|
| 195 |
{
|
| 196 |
"type": "3.3_loss_vanishing",
|
| 197 |
"step": 16,
|
|
|
|
| 198 |
"value": -1.0
|
| 199 |
},
|
| 200 |
{
|
| 201 |
"type": "3.3_loss_vanishing",
|
| 202 |
"step": 17,
|
| 203 |
+
"value": -1.0
|
|
|
|
| 204 |
},
|
| 205 |
{
|
| 206 |
"type": "3.3_loss_vanishing",
|
| 207 |
"step": 18,
|
|
|
|
| 208 |
"value": -1.0
|
| 209 |
},
|
| 210 |
{
|
| 211 |
"type": "3.3_loss_vanishing",
|
| 212 |
"step": 19,
|
|
|
|
| 213 |
"value": -1.0
|
| 214 |
},
|
| 215 |
{
|
| 216 |
"type": "3.3_loss_vanishing",
|
| 217 |
"step": 20,
|
| 218 |
+
"value": -0.982607368881035
|
|
|
|
| 219 |
},
|
| 220 |
{
|
| 221 |
"type": "3.3_loss_vanishing",
|
| 222 |
"step": 21,
|
|
|
|
| 223 |
"value": -1.0
|
| 224 |
},
|
| 225 |
{
|
| 226 |
"type": "3.3_loss_vanishing",
|
| 227 |
"step": 22,
|
| 228 |
+
"value": -0.6123724356957945
|
|
|
|
| 229 |
},
|
| 230 |
{
|
| 231 |
"type": "3.3_loss_vanishing",
|
| 232 |
"step": 23,
|
|
|
|
| 233 |
"value": -1.0
|
| 234 |
},
|
| 235 |
{
|
| 236 |
"type": "3.3_loss_vanishing",
|
| 237 |
"step": 24,
|
|
|
|
| 238 |
"value": -1.0
|
| 239 |
},
|
| 240 |
{
|
| 241 |
"type": "3.3_loss_vanishing",
|
| 242 |
"step": 25,
|
|
|
|
| 243 |
"value": -1.0
|
| 244 |
},
|
| 245 |
{
|
| 246 |
"type": "3.3_loss_vanishing",
|
| 247 |
"step": 26,
|
|
|
|
| 248 |
"value": -1.0
|
| 249 |
},
|
| 250 |
{
|
| 251 |
"type": "3.3_loss_vanishing",
|
| 252 |
"step": 27,
|
| 253 |
+
"value": -1.0
|
|
|
|
| 254 |
},
|
| 255 |
{
|
| 256 |
"type": "3.3_loss_vanishing",
|
| 257 |
"step": 28,
|
|
|
|
| 258 |
"value": -1.0
|
| 259 |
},
|
| 260 |
{
|
| 261 |
"type": "3.3_loss_vanishing",
|
| 262 |
"step": 29,
|
|
|
|
| 263 |
"value": -1.0
|
| 264 |
},
|
| 265 |
{
|
| 266 |
"type": "3.3_loss_vanishing",
|
| 267 |
"step": 30,
|
| 268 |
+
"value": -1.0
|
|
|
|
| 269 |
}
|
| 270 |
],
|
| 271 |
"kohonen_final": {
|
| 272 |
+
"sigma_t": 0.7909386360645728,
|
| 273 |
+
"alpha_t": 0.0726149037073691,
|
| 274 |
+
"t": 640,
|
| 275 |
+
"n_neurons": 128,
|
| 276 |
"fisher_w_mean": 0.0,
|
| 277 |
"fisher_w_max": 0.0,
|
| 278 |
"fisher_accum_count": 0,
|
| 279 |
"has_ewc_reference": true,
|
| 280 |
+
"weights_norm": 1.1414568424224854,
|
| 281 |
+
"weights_w_mean": 0.03337728977203369
|
| 282 |
},
|
| 283 |
"hypothesis_final": {
|
| 284 |
"classifier_trained": true,
|
| 285 |
"punishment_count": 0,
|
| 286 |
+
"success_count": 0,
|
| 287 |
+
"training_ready": false,
|
| 288 |
+
"buffer_size": 0,
|
| 289 |
"required_new_samples": 10,
|
| 290 |
+
"histogram_max": 0,
|
| 291 |
+
"time_counter": 600
|
| 292 |
},
|
| 293 |
"mtp_final": {
|
| 294 |
"active": false,
|
|
|
|
| 301 |
"active_in_val": true,
|
| 302 |
"entropy_beta": 0.01
|
| 303 |
},
|
| 304 |
+
"vqvae2_final": {
|
| 305 |
+
"active": true,
|
| 306 |
+
"n_calls": 8,
|
| 307 |
+
"latest_total_loss": 0.11622760444879532,
|
| 308 |
+
"latest_recon_loss": 0.1146363615989685,
|
| 309 |
+
"latest_vq_loss": 0.0015912405215203762,
|
| 310 |
+
"latest_usage_top": 0.9375,
|
| 311 |
+
"latest_usage_bot": 0.890625,
|
| 312 |
+
"latest_goose_temp": 1.7280679941177368,
|
| 313 |
+
"mean_total_loss": 0.1640255962099348,
|
| 314 |
+
"mean_recon_loss": 0.1340886503458023
|
| 315 |
+
},
|
| 316 |
+
"reasoning_final": {
|
| 317 |
+
"active": true,
|
| 318 |
+
"n_history": 0,
|
| 319 |
+
"n_steps_last": 0,
|
| 320 |
+
"phases_used": []
|
| 321 |
+
},
|
| 322 |
"ewc_final": {
|
| 323 |
"active": true,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 324 |
"ewc_eval_mode_penalty": true,
|
| 325 |
+
"w8a8_dequant_active": true,
|
| 326 |
+
"fisher_w_mean": 0.0,
|
| 327 |
+
"fisher_w_max": 0.0,
|
| 328 |
+
"fisher_accum_count": 0
|
| 329 |
},
|
| 330 |
"evolution_phase1_to_phase2": {
|
| 331 |
+
"acc_phase1_mean": 0.9450431034482758,
|
| 332 |
+
"acc_phase2_mean": 0.9450431034482758,
|
| 333 |
+
"sigma_phase1_end": 1.0892235556105363,
|
| 334 |
+
"sigma_phase2_end": 0.7909386360645728,
|
| 335 |
+
"vqvae2_phase1_mean": 0.2318861267783425,
|
| 336 |
+
"vqvae2_phase2_mean": 0.12742777417103449
|
| 337 |
}
|
| 338 |
},
|
| 339 |
"kohonen_final": {
|
| 340 |
+
"t": 640,
|
| 341 |
+
"sigma_t": 0.7909386360645728,
|
| 342 |
+
"alpha_t": 0.0726149037073691,
|
| 343 |
"sigma0": 1.5,
|
| 344 |
"alpha0": 0.1,
|
| 345 |
"lambda_ewc": 0.02,
|
| 346 |
"grid_shape": [
|
| 347 |
+
4,
|
| 348 |
+
4,
|
| 349 |
+
4,
|
| 350 |
+
2
|
| 351 |
],
|
| 352 |
+
"n_neurons": 128,
|
| 353 |
"has_ewc_reference": true,
|
| 354 |
"fisher_w_mean": 0.0,
|
| 355 |
"fisher_w_max": 0.0,
|
| 356 |
"fisher_accum_count": 0,
|
| 357 |
+
"weights_norm": 1.1414568424224854,
|
| 358 |
+
"weights_w_mean": 0.03337728977203369
|
| 359 |
},
|
| 360 |
"kls_state": {
|
| 361 |
"som": {
|
| 362 |
+
"t": 640,
|
| 363 |
+
"sigma_t": 0.7909386360645728,
|
| 364 |
+
"alpha_t": 0.0726149037073691,
|
| 365 |
"sigma0": 1.5,
|
| 366 |
"alpha0": 0.1,
|
| 367 |
"lambda_ewc": 0.02,
|
| 368 |
"grid_shape": [
|
| 369 |
+
4,
|
| 370 |
+
4,
|
| 371 |
+
4,
|
| 372 |
+
2
|
| 373 |
],
|
| 374 |
+
"n_neurons": 128,
|
| 375 |
"has_ewc_reference": true,
|
| 376 |
"fisher_w_mean": 0.0,
|
| 377 |
"fisher_w_max": 0.0,
|
| 378 |
"fisher_accum_count": 0,
|
| 379 |
+
"weights_norm": 1.1414568424224854,
|
| 380 |
+
"weights_w_mean": 0.03337728977203369
|
| 381 |
},
|
| 382 |
"kls": {
|
| 383 |
+
"time_counter": 610,
|
| 384 |
"T_max": 10000,
|
| 385 |
+
"buffer_size": 0,
|
| 386 |
+
"training_ready": false,
|
| 387 |
"punishment_count": 0,
|
| 388 |
+
"success_count": 0,
|
| 389 |
"classifier_trained": true,
|
| 390 |
+
"histogram_size": 0,
|
| 391 |
+
"histogram_max": 0,
|
| 392 |
"N_start": 10,
|
| 393 |
"dim_choice": "y",
|
| 394 |
+
"som_neuron_count": 128,
|
| 395 |
"required_new_samples": 10,
|
| 396 |
"has_classifier": true,
|
| 397 |
+
"vocab_size": 4096,
|
| 398 |
+
"hidden_dim": 256,
|
| 399 |
+
"seq_len": 8,
|
| 400 |
+
"enable_vqvae2": true,
|
| 401 |
+
"enable_reasoning": true,
|
| 402 |
+
"vqvae2_n_calls": 8
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 403 |
},
|
| 404 |
+
"vqvae2": {
|
| 405 |
+
"active": true,
|
| 406 |
+
"n_calls": 8,
|
| 407 |
+
"latest": {
|
| 408 |
+
"vq_loss": 0.0015912405215203762,
|
| 409 |
+
"recon_loss": 0.1146363615989685,
|
| 410 |
+
"total_loss": 0.11622760444879532,
|
| 411 |
+
"n_used_top": 30,
|
| 412 |
+
"n_used_bot": 57,
|
| 413 |
+
"usage_ratio_top": 0.9375,
|
| 414 |
+
"usage_ratio_bot": 0.890625,
|
| 415 |
+
"codebook_ppl_top": 10.040919817876812,
|
| 416 |
+
"codebook_ppl_bot": 13.454342643871065,
|
| 417 |
+
"n_restarted_top": 0,
|
| 418 |
+
"n_restarted_bot": 0,
|
| 419 |
+
"goose_temp": 1.7280679941177368,
|
| 420 |
+
"active": true
|
| 421 |
+
},
|
| 422 |
+
"mean_total_loss": 0.1640255962099348,
|
| 423 |
+
"mean_recon_loss": 0.1340886503458023,
|
| 424 |
+
"n_nan_skipped": 1
|
| 425 |
},
|
| 426 |
+
"reasoning": {
|
| 427 |
+
"active": true,
|
| 428 |
+
"stats": {
|
| 429 |
+
"n_steps": 8,
|
| 430 |
+
"n_history": 5,
|
| 431 |
+
"tools": [
|
| 432 |
+
"som_query"
|
| 433 |
+
],
|
| 434 |
+
"tool_stats": {
|
| 435 |
+
"som_query": {
|
| 436 |
+
"n_calls": 5,
|
| 437 |
+
"n_success": 5,
|
| 438 |
+
"n_failures": 0,
|
| 439 |
+
"total_time_s": 0.03197741508483887,
|
| 440 |
+
"avg_time_s": 0.006395483016967773,
|
| 441 |
+
"cache_size": 0,
|
| 442 |
+
"instructions_count": 0,
|
| 443 |
+
"error_count": 0,
|
| 444 |
+
"success_rate": 1.0,
|
| 445 |
+
"circuit_breaker": {
|
| 446 |
+
"state": "closed",
|
| 447 |
+
"failure_count": 0,
|
| 448 |
+
"success_count": 0,
|
| 449 |
+
"failure_threshold": 5,
|
| 450 |
+
"recovery_timeout_s": 30.0
|
| 451 |
+
}
|
| 452 |
+
}
|
| 453 |
+
},
|
| 454 |
+
"phases_used": [
|
| 455 |
+
"answering",
|
| 456 |
+
"planning",
|
| 457 |
+
"monitoring",
|
| 458 |
+
"adjusting",
|
| 459 |
+
"executing",
|
| 460 |
+
"predicting",
|
| 461 |
+
"decomposing",
|
| 462 |
+
"thinking"
|
| 463 |
+
]
|
| 464 |
+
},
|
| 465 |
+
"n_history": 5
|
| 466 |
}
|
| 467 |
},
|
| 468 |
+
"verification": {
|
| 469 |
+
"checks": {
|
| 470 |
+
"kls_instantiation": {
|
| 471 |
+
"status": "PASS",
|
| 472 |
+
"details": "vqvae2=True, reasoning=True"
|
| 473 |
+
},
|
| 474 |
+
"find_bmu_no_premature_return": {
|
| 475 |
+
"status": "PASS",
|
| 476 |
+
"details": "bmu=(0, 2, 2, 0) valid"
|
| 477 |
+
},
|
| 478 |
+
"activate_hypothesis_detach": {
|
| 479 |
+
"status": "PASS",
|
| 480 |
+
"details": "classifier_trained=True"
|
| 481 |
+
},
|
| 482 |
+
"pgvector_lookup_removed": {
|
| 483 |
+
"status": "PASS"
|
| 484 |
+
},
|
| 485 |
+
"kohonen_refactored_removed": {
|
| 486 |
+
"status": "PASS",
|
| 487 |
+
"details": "old_folder=False, new_file=True"
|
| 488 |
+
},
|
| 489 |
+
"trainer_removed": {
|
| 490 |
+
"status": "PASS"
|
| 491 |
+
},
|
| 492 |
+
"vqvae2_valid_output": {
|
| 493 |
+
"status": "PASS",
|
| 494 |
+
"details": "n_calls=1, total_loss=0.4196"
|
| 495 |
+
},
|
| 496 |
+
"reasoning_engine_tags": {
|
| 497 |
+
"status": "PASS",
|
| 498 |
+
"details": "reasoning length=928 chars"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 499 |
}
|
| 500 |
+
},
|
| 501 |
+
"n_pass": 8,
|
| 502 |
+
"n_fail": 0,
|
| 503 |
+
"all_pass": true
|
| 504 |
+
},
|
| 505 |
+
"ewc_w8a8_benchmark_summary": {
|
| 506 |
+
"dequant_preserves_accuracy": true,
|
| 507 |
+
"dequant_relative_error": 0.00027968583245277626,
|
| 508 |
+
"int8_relative_error": 145504.6191166922,
|
| 509 |
+
"dequant_overhead_ms": 0.08233785629272461,
|
| 510 |
+
"dequant_overhead_pct": 181.32416255381708,
|
| 511 |
+
"conclusion": "EWC+W8A8 eval com dequantização ativa: erro relativo dequant=0.000280 (< 0.1 = OK), erro relativo int8 direto=145504.619117 (mostra que dequant é necessário). Overhead dequant: 0.082ms (181.3%)."
|
| 512 |
+
},
|
| 513 |
+
"reasoning_eval_summary": {
|
| 514 |
+
"n_with_answer": 5,
|
| 515 |
+
"n_with_think": 5,
|
| 516 |
+
"answer_rate": 1.0,
|
| 517 |
+
"think_rate": 1.0,
|
| 518 |
+
"avg_reasoning_length": 974.0,
|
| 519 |
+
"reasoning_engine_active": true,
|
| 520 |
+
"reasoning_engine_n_history": 5
|
| 521 |
+
},
|
| 522 |
+
"module_analysis_summary": {
|
| 523 |
+
"n_removed_modules": 9,
|
| 524 |
+
"n_removed_folders": 1,
|
| 525 |
+
"n_active_modules": 21
|
| 526 |
},
|
| 527 |
"script_activity_summary": {
|
| 528 |
+
"n_scripts": 4,
|
| 529 |
+
"active_v65": 2,
|
| 530 |
+
"active_v64": 2,
|
| 531 |
+
"upload_utility": 0
|
| 532 |
},
|
| 533 |
"math_analysis": {
|
| 534 |
"text_to_4d": "SVD: M @ V[:3].T -> centroid 3D + w = time_step/T_max (LINEAR)",
|
|
|
|
| 538 |
"sigma_decay": "sigma_t = sigma0 * exp(-t/1000)",
|
| 539 |
"alpha_decay": "alpha_t = alpha0 * exp(-t/2000)",
|
| 540 |
"ewc_only_dim4": "penalty = lambda * F * (W_w - W*_w)",
|
| 541 |
+
"vqvae2_loss": "L = recon_loss + vq_loss (commitment top + bottom + diversity)",
|
| 542 |
+
"vqvae2_ema": "EMA codebook update + dead code restart + Goose VQ",
|
| 543 |
+
"mtp_loss": "L = sum_k(alpha_k * L_k) - beta * H(alpha)",
|
| 544 |
+
"ewc_w8a8_dequant": "W_float = (W_int8 * scale) / smooth_scale; EWC uses W_float for (p-w*)^2"
|
| 545 |
},
|
| 546 |
"datasets_used": [
|
| 547 |
+
"dominguesm/restore-punctuation-ptbr-dataset",
|
| 548 |
+
"carolina-c4ai/corpus-carolina",
|
| 549 |
+
"nvidia/OpenMathInstruct-2",
|
| 550 |
+
"nvidia/OpenMathReasoning",
|
| 551 |
"CEIA-POSITIVO/ultrachat_br_clustred_balanced_v1",
|
| 552 |
+
"Dexavator/English-PTBR"
|
| 553 |
]
|
| 554 |
}
|
v6_5_script_activity.json
CHANGED
|
@@ -1,93 +1,30 @@
|
|
| 1 |
{
|
| 2 |
-
"smoke_test.py": {
|
| 3 |
-
"path": "scripts/smoke_test.py",
|
| 4 |
-
"size_bytes": 13457,
|
| 5 |
-
"mtime": "2026-08-05T22:30:13",
|
| 6 |
-
"classification": "test",
|
| 7 |
-
"activity": "legacy"
|
| 8 |
-
},
|
| 9 |
-
"train.py": {
|
| 10 |
-
"path": "scripts/train.py",
|
| 11 |
-
"size_bytes": 8026,
|
| 12 |
-
"mtime": "2026-08-05T22:37:41",
|
| 13 |
-
"classification": "legacy",
|
| 14 |
-
"activity": "legacy"
|
| 15 |
-
},
|
| 16 |
-
"train_fast.py": {
|
| 17 |
-
"path": "scripts/train_fast.py",
|
| 18 |
-
"size_bytes": 5017,
|
| 19 |
-
"mtime": "2026-08-05T22:45:27",
|
| 20 |
-
"classification": "legacy",
|
| 21 |
-
"activity": "legacy"
|
| 22 |
-
},
|
| 23 |
-
"train_v6_1.py": {
|
| 24 |
-
"path": "scripts/train_v6_1.py",
|
| 25 |
-
"size_bytes": 53839,
|
| 26 |
-
"mtime": "2026-08-07T00:05:08",
|
| 27 |
-
"classification": "active_v61",
|
| 28 |
-
"activity": "recent"
|
| 29 |
-
},
|
| 30 |
-
"train_v6_2.py": {
|
| 31 |
-
"path": "scripts/train_v6_2.py",
|
| 32 |
-
"size_bytes": 68794,
|
| 33 |
-
"mtime": "2026-08-07T00:55:19",
|
| 34 |
-
"classification": "active_v62",
|
| 35 |
-
"activity": "recent"
|
| 36 |
-
},
|
| 37 |
-
"train_v6_3.py": {
|
| 38 |
-
"path": "scripts/train_v6_3.py",
|
| 39 |
-
"size_bytes": 22527,
|
| 40 |
-
"mtime": "2026-08-07T21:17:23.191309",
|
| 41 |
-
"classification": "active_v63",
|
| 42 |
-
"activity": "recent"
|
| 43 |
-
},
|
| 44 |
"train_v6_4.py": {
|
| 45 |
"path": "scripts/train_v6_4.py",
|
| 46 |
-
"size_bytes":
|
| 47 |
-
"mtime": "2026-08-
|
| 48 |
"classification": "active_v64",
|
| 49 |
"activity": "recent"
|
| 50 |
},
|
| 51 |
"train_v6_5.py": {
|
| 52 |
"path": "scripts/train_v6_5.py",
|
| 53 |
-
"size_bytes":
|
| 54 |
-
"mtime": "2026-08-
|
| 55 |
"classification": "active_v65",
|
| 56 |
"activity": "active"
|
| 57 |
},
|
| 58 |
-
"upload_to_hf.py": {
|
| 59 |
-
"path": "scripts/upload_to_hf.py",
|
| 60 |
-
"size_bytes": 3240,
|
| 61 |
-
"mtime": "2026-08-05T22:47:27",
|
| 62 |
-
"classification": "upload_utility",
|
| 63 |
-
"activity": "legacy"
|
| 64 |
-
},
|
| 65 |
-
"upload_v6_1_resilient.py": {
|
| 66 |
-
"path": "scripts/upload_v6_1_resilient.py",
|
| 67 |
-
"size_bytes": 9960,
|
| 68 |
-
"mtime": "2026-08-07T00:05:08",
|
| 69 |
-
"classification": "active_v61",
|
| 70 |
-
"activity": "recent"
|
| 71 |
-
},
|
| 72 |
-
"upload_v6_2_resilient.py": {
|
| 73 |
-
"path": "scripts/upload_v6_2_resilient.py",
|
| 74 |
-
"size_bytes": 9904,
|
| 75 |
-
"mtime": "2026-08-07T00:55:22",
|
| 76 |
-
"classification": "active_v62",
|
| 77 |
-
"activity": "recent"
|
| 78 |
-
},
|
| 79 |
-
"upload_v6_3_resilient.py": {
|
| 80 |
-
"path": "scripts/upload_v6_3_resilient.py",
|
| 81 |
-
"size_bytes": 4712,
|
| 82 |
-
"mtime": "2026-08-07T21:17:23.192309",
|
| 83 |
-
"classification": "active_v63",
|
| 84 |
-
"activity": "recent"
|
| 85 |
-
},
|
| 86 |
"upload_v6_4_resilient.py": {
|
| 87 |
"path": "scripts/upload_v6_4_resilient.py",
|
| 88 |
"size_bytes": 5150,
|
| 89 |
"mtime": "2026-08-07T21:51:52.545045",
|
| 90 |
"classification": "active_v64",
|
| 91 |
"activity": "recent"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 92 |
}
|
| 93 |
}
|
|
|
|
| 1 |
{
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
"train_v6_4.py": {
|
| 3 |
"path": "scripts/train_v6_4.py",
|
| 4 |
+
"size_bytes": 25325,
|
| 5 |
+
"mtime": "2026-08-07T23:25:35.828302",
|
| 6 |
"classification": "active_v64",
|
| 7 |
"activity": "recent"
|
| 8 |
},
|
| 9 |
"train_v6_5.py": {
|
| 10 |
"path": "scripts/train_v6_5.py",
|
| 11 |
+
"size_bytes": 64673,
|
| 12 |
+
"mtime": "2026-08-07T23:59:40.769069",
|
| 13 |
"classification": "active_v65",
|
| 14 |
"activity": "active"
|
| 15 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
"upload_v6_4_resilient.py": {
|
| 17 |
"path": "scripts/upload_v6_4_resilient.py",
|
| 18 |
"size_bytes": 5150,
|
| 19 |
"mtime": "2026-08-07T21:51:52.545045",
|
| 20 |
"classification": "active_v64",
|
| 21 |
"activity": "recent"
|
| 22 |
+
},
|
| 23 |
+
"upload_v6_5_resilient.py": {
|
| 24 |
+
"path": "scripts/upload_v6_5_resilient.py",
|
| 25 |
+
"size_bytes": 5001,
|
| 26 |
+
"mtime": "2026-08-07T23:25:46.234286",
|
| 27 |
+
"classification": "active_v65",
|
| 28 |
+
"activity": "active"
|
| 29 |
}
|
| 30 |
}
|
v6_5_training_metrics.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
v6_5_upload_report.json
ADDED
|
@@ -0,0 +1,35 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"version": "V6.5",
|
| 3 |
+
"upload_timestamp": "2026-08-07T22:58:11",
|
| 4 |
+
"repo_id": "PowerMachine/BiGRU_T_version",
|
| 5 |
+
"commit_oid": "c2992d0792671182a56c0e1910646a7370c268af",
|
| 6 |
+
"commit_url": "https://huggingface.co/PowerMachine/BiGRU_T_version/commit/c2992d0792671182a56c0e1910646a7370c268af",
|
| 7 |
+
"duration_s": 3.3573570251464844,
|
| 8 |
+
"critical_files": [
|
| 9 |
+
"src/bigru_t/model/kohonen_refactored/__init__.py",
|
| 10 |
+
"src/bigru_t/model/kohonen_refactored/kohonen_learning_system.py",
|
| 11 |
+
"src/bigru_t/model/hyp_t.py",
|
| 12 |
+
"src/bigru_t/training/mtp.py",
|
| 13 |
+
"src/bigru_t/training/ewc.py",
|
| 14 |
+
"src/bigru_t/quantization/smoothquant_compressor.py",
|
| 15 |
+
"src/bigru_t/model/vqvae2_hierarchical.py",
|
| 16 |
+
"src/bigru_t/model/vqvae2_hierarchical_flexnet.py",
|
| 17 |
+
"src/bigru_t/model/token_compress.py",
|
| 18 |
+
"src/bigru_t/reasoning/thinking.py",
|
| 19 |
+
"src/bigru_t/reasoning/reasoning_engine.py",
|
| 20 |
+
"src/bigru_t/reasoning/circular_orchestration.py",
|
| 21 |
+
"src/bigru_t/reasoning/tool_agent.py",
|
| 22 |
+
"src/bigru_t/reasoning/distributed_reasoning_system.py",
|
| 23 |
+
"src/bigru_t/reasoning/cyclic_reasoning.py",
|
| 24 |
+
"src/bigru_t/reasoning/consensus_sampling.py",
|
| 25 |
+
"v6_5_report.json",
|
| 26 |
+
"v6_5_training_metrics.json",
|
| 27 |
+
"v6_5_module_analysis.json",
|
| 28 |
+
"v6_5_script_activity.json",
|
| 29 |
+
"scripts/train_v6_5.py",
|
| 30 |
+
"scripts/upload_v6_5_resilient.py",
|
| 31 |
+
"src/bigru_t/utils/xeon_runtime.py",
|
| 32 |
+
"src/bigru_t/data/streaming_datasets.py"
|
| 33 |
+
],
|
| 34 |
+
"n_files_in_repo": 140
|
| 35 |
+
}
|