llama31-8b-knapsack-lora-persistent

Anonymous supplementary release for a double-blind workshop submission. This is one of four LoRA adapters (Mistral-7B-v0.3 / Llama-3.1-8B base model x persistent/stateless training regime) fine-tuned on the Opaque Knapsack agentic task, extending a prior single-base-model result (see the sibling Qwen3-8B release) to a second base model family for the same reproducibility review.

  • Base model: meta-llama/Llama-3.1-8B
  • Training regime: persistent (trained with a persistent Python interpreter runtime (state carries over across agent turns))
  • Seed: 3407

Training configuration

Fine-tuned with Axolotl, LoRA adapter, 4-bit NF4 quantized base:

Hyperparameter Value
lora_r 64
lora_alpha 128
lora_dropout 0.05
lora_target_modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
learning_rate 1e-4
lr_scheduler cosine
optimizer adamw_torch
epochs 3.0
micro_batch_size 1
gradient_accumulation_steps 16
sequence_len 16384
sample_packing false
seed 3407
training data paired traces for the "persistent" regime (see paper Appendix for pairing/filtering procedure)

Base checkpoint's own chat-format tokens (e.g. <|eot_id|>) are untrained on this base (non-instruct) checkpoint -- trained and served with a hand-written minimal template using only real trained tokens (BOS/EOS + plain-text role prefixes), not Llama's native instruct template.

Provenance

Released anonymously alongside a NeurIPS workshop submission for reproducibility review. Non-anonymous release (paper citation, full code, full training traces) will follow after the review process concludes.

Downloads last month
39
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hfunknown/llama31-8b-knapsack-lora-persistent

Adapter
(1188)
this model

Collection including hfunknown/llama31-8b-knapsack-lora-persistent