@Hoglet-33 did you have attention layers dominating the gradient? im guessing an all ratio sweep at like 100% mamba 7:1 3:1 and 1:1 are already planned for a wide curve. seeing where the attention layers sit too
i64 Systems PRO
i64systems
Β·
AI & ML interests
deterministic, auditable local AI Β· integer/low-bit ML Β· memory systems
Recent Activity
View all activity
Organizations
None yet