An NLA trained with random-length truncation of the verbalizer's output, so the most important information comes first. Checkpoints, data, baseline.
Smitty PRO
syvb
·
AI & ML interests
None yet
Recent Activity
updated a Space 1 day ago
syvb/nla-qwen36-27b-explorer-std updated a Space 1 day ago
syvb/nla-qwen36-27b-explorer published a model 4 days ago
syvb/nla-qwen2.5-7b-L20-rl-overlappenOrganizations
None yet
Matryoshka NLA (Qwen2.5-7B L20)
An NLA trained with random-length truncation of the verbalizer's output, so the most important information comes first. Checkpoints, data, baseline.
NLA length penalty
nanoNLA multi-input affine experiment (Qwen3-8B L24)
16 repeated injection markers w/ per-slot learned affine vs 1-slot control. Warm-started from nanonla-qwen3-8b-L24-av.
Matryoshka NLA (earlier tests)
v2 item-truncation matryoshka NLA (Qwen2.5-7B L20): warm-start AV verbalizer + AR critic checkpoints, warm-start data, base NLA.