--- library_name: peft base_model: meta-llama/Llama-3.1-8B tags: - lora - transformers pipeline_tag: text-generation --- # llama31-8b-knapsack-lora-persistent Anonymous supplementary release for a double-blind workshop submission. This is one of four LoRA adapters (Mistral-7B-v0.3 / Llama-3.1-8B base model x persistent/stateless training regime) fine-tuned on the Opaque Knapsack agentic task, extending a prior single-base-model result (see the sibling Qwen3-8B release) to a second base model family for the same reproducibility review. - **Base model:** [meta-llama/Llama-3.1-8B](https://huggingface.co/meta-llama/Llama-3.1-8B) - **Training regime:** persistent (trained with a persistent Python interpreter runtime (state carries over across agent turns)) - **Seed:** 3407 ## Training configuration Fine-tuned with [Axolotl](https://github.com/axolotl-ai-cloud/axolotl), LoRA adapter, 4-bit NF4 quantized base: | Hyperparameter | Value | |---|---| | lora_r | 64 | | lora_alpha | 128 | | lora_dropout | 0.05 | | lora_target_modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj | | learning_rate | 1e-4 | | lr_scheduler | cosine | | optimizer | adamw_torch | | epochs | 3.0 | | micro_batch_size | 1 | | gradient_accumulation_steps | 16 | | sequence_len | 16384 | | sample_packing | false | | seed | 3407 | | training data | paired traces for the "persistent" regime (see paper Appendix for pairing/filtering procedure) | Base checkpoint's own chat-format tokens (e.g. `<|eot_id|>`) are untrained on this base (non-instruct) checkpoint -- trained and served with a hand-written minimal template using only real trained tokens (BOS/EOS + plain-text role prefixes), not Llama's native instruct template. ## Provenance Released anonymously alongside a NeurIPS workshop submission for reproducibility review. Non-anonymous release (paper citation, full code, full training traces) will follow after the review process concludes.