--- library_name: peft base_model: Qwen/Qwen3-8B tags: - lora - transformers pipeline_tag: text-generation --- # qwen3-8b-navigation-lora-persistent Anonymous supplementary release for a double-blind workshop submission. This is one of four LoRA adapters (rule_diagnosis / navigation task family x persistent/stateless training regime), fine-tuned on the navigation agentic task (graph exploration with a per-turn tool-call budget). It is the second-family generalization arm alongside the primary Opaque Knapsack result (see the sibling Qwen3-8B knapsack release). - **Base model:** [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B) - **Training regime:** persistent (trained with a persistent Python interpreter runtime (state carries over across agent turns)) - **Seed:** 3407 ## Training configuration Fine-tuned with [Axolotl](https://github.com/axolotl-ai-cloud/axolotl), LoRA adapter, 4-bit NF4 quantized base: | Hyperparameter | Value | |---|---| | lora_r | 64 | | lora_alpha | 128 | | lora_dropout | 0.05 | | lora_target_modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj | | learning_rate | 1e-4 | | lr_scheduler | cosine | | optimizer | adamw_torch | | epochs | 3.0 | | micro_batch_size | 1 | | gradient_accumulation_steps | 16 | | sequence_len | 16384 | | sample_packing | false | | seed | 3407 | | training data | paired traces for the "persistent" regime on navigation, see paper Appendix for pairing/filtering procedure | ## Provenance Released anonymously alongside a NeurIPS workshop submission for reproducibility review. Non-anonymous release (paper citation, full code, full training traces) will follow after the review process concludes.