--- license: other base_model: qwen/Qwen3-4B tags: - text-generation - 3d - mesh --- # output_40k_cont4ep Qwen3-4B fine-tuned for **3D mesh understanding & generation** (SPAC pipeline, Spatial-PE enabled), continued from the `output_40k_cont2ep` weights for 4 more epochs. ## Training data (`bpe_40k_mix`, 40k samples) - 20k mesh -> text: given a mesh-token sequence, produce a text description (understanding). - 20k text -> mesh: given a caption + "Reconstruct this 3D shape in mesh token format:", output the mesh token sequence (generation). Both directions are paired on the same shapes. ## Training details - Full fine-tune (all params), DeepSpeed ZeRO-3, bf16, FlashAttention-2, packing (8192, neat). - +4 epochs continuation from `output_40k_cont2ep` (fresh optimizer, fresh LR schedule). - LR 5e-5 cosine decay to 5e-6, warmup 3%, 4 GPUs, batch 1 each. - Train loss: epoch means 0.844 -> 0.660 -> 0.254 -> 0.052 (whole-run avg 0.4524). ## Inference notes - Chat template: `qwen3_nothink` (no thinking). Mesh shapes are tokenized as `` tokens (see `added_tokens.json`). - Requires the SPAC Spatial-PE inference wrapper (`spatial_pe_scale 0.08`, merge table from `Jack04810/spac-qwen3-4b-spatial-pe-8ep`) to run with the spatial positional encoding used in training.