YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

skiplens-ar-L62-allidx β€” truncation-resistant activation reconstructor (AR), Qwen3.6-27B

LoRA (r64 / Ξ±16, rsLoRA, lora_scope=all) + Linear(5120,5120) value head on a 63-layer-truncated Qwen3.6-27B. Reconstructs the layer-62 residual-stream activation from a short text span (≀12 tokens), trained with the dense --ar-all-idx objective (reconstruct at every causal prefix β†’ read-position robust). Used as the frozen reward in futurelens-NLA GRPO.

  • Base: Qwen/Qwen3.6-27B (truncated to 63 layers, final RMSNorm stripped)
  • Training: 776k on-policy ≀12-token spans, 1 epoch, lr 1e-4, batch 32
  • Held-out FVE (norm-constrained): 24.9% averaged over 4–12-tok spans (~30% at full length); baseline (predict-the-mean) β‰ˆ 0
  • Read template: Summary of the following text: <text>{span}</text> <summary> (reads at the last token; but all-idx β†’ robust at any read position)
  • Files: ar_lora_value_head.safetensors (LoRA + value head), ar_meta.json, nla_meta.yaml
  • ⚠ Loading gotcha: ar_meta.json["target_modules"] was regenerated to the true set, but if you hit a stale copy, derive target_modules from the .lora_ keys in the safetensors (full set = q/k/v/o_proj, in_proj_a/b, in_proj_qkv/z, out_proj, gate/up/down_proj).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support