yunfeixie commited on Apr 22

Commit

5c6d204

verified ·

1 Parent(s): 3d6bcd8

Upload folder using huggingface_hub

Browse files

Files changed (22) hide show

INFERENCE.md +162 -0
RESULTS.md +96 -0
configs/qwen2.5-vl-3b-instruct/LICENSE +54 -0
configs/qwen2.5-vl-3b-instruct/README.md +525 -0
configs/qwen2.5-vl-3b-instruct/chat_template.json +3 -0
configs/qwen2.5-vl-3b-instruct/config.json +61 -0
configs/qwen2.5-vl-3b-instruct/generation_config.json +12 -0
configs/qwen2.5-vl-3b-instruct/merges.txt +0 -0
configs/qwen2.5-vl-3b-instruct/model-00001-of-00002.safetensors +3 -0
configs/qwen2.5-vl-3b-instruct/model-00002-of-00002.safetensors +3 -0
configs/qwen2.5-vl-3b-instruct/model.safetensors.index.json +831 -0
configs/qwen2.5-vl-3b-instruct/preprocessor_config.json +19 -0
configs/qwen2.5-vl-3b-instruct/tokenizer.json +0 -0
configs/qwen2.5-vl-3b-instruct/tokenizer_config.json +207 -0
configs/qwen2.5-vl-3b-instruct/vocab.json +0 -0
dataset_statistics_bridge.json +116 -0
patches/convert_ckpt_standalone.py +18 -0
patches/eval_calvin_model_wrapper.py.patch +39 -0
patches/eval_simpler_main_inference.py.patch +38 -0
project.json +166 -0
requirements-eval.txt +212 -0
stepstep=0030000.fp32.pt +3 -0

INFERENCE.md ADDED Viewed

	@@ -0,0 +1,162 @@

+# Run (b) – best reproducible checkpoint for SimplerBridge
+This bundle contains everything needed to load and evaluate our run-(b) Qwen2.5-VL-3B freeze-vision VLA on SimplerBridge and reproduce (and exceed) the paper's reported 23.95% SR.
+- Checkpoint: `stepstep=0030000.fp32.pt` (15 GB, FP32 single-file state dict)
+- Target task suite: SimplerBridge visual-matching (4 tasks × 24 episodes)
+- Best-of-sweep SR: **30.21%** (at `execute_step=2`) – paper target: 23.95%
+- Vision encoder: **frozen** during training (Qwen2.5-VL ViT, 668 M params)
+- Trainable params at eval time: 3.09 B (language tower + word-embed + FC action head)
+## What's in this bundle
+```
+run_b_best/
+├── stepstep=0030000.fp32.pt          # checkpoint weights (15 GB, FP32)
+├── project.json                       # training-time config (Lightning snapshot)
+├── dataset_statistics_bridge.json     # BridgeV2 action stats (used for un-normalization)
+├── configs/
+│   └── qwen2.5-vl-3b-instruct/        # exact HF base-model files used at training time
+├── patches/
+│   ├── eval_calvin_model_wrapper.py.patch   # fix: `get_text_function` 4-arg -> 2-arg
+│   ├── eval_simpler_main_inference.py.patch # fix: set args.policy_model from configs["model"]
+│   └── convert_ckpt_standalone.py           # DS stage-2 -> FP32 converter (thin wrapper)
+├── requirements-eval.txt              # `pip freeze` of the eval conda env
+├── INFERENCE.md                       # this file
+└── RESULTS.md                         # full 30-cell sweep matrix + best-of
+```
+## Environment (known-good)
+- Linux + CUDA 12.4
+- 1x A100-80GB (for a single-ckpt eval); 8x A100-80GB for a full 30-cell sweep
+- `conda` (miniforge3 works), Python 3.10
+- Disk: ~25 GB for this bundle + ~1-2 GB of per-task eval artifacts
+### Install (fresh machine)
+```bash
+# 1. Clone the fork with the eval fixes already applied
+git clone https://github.com/yunfeixie233/VLM4VLA.git
+cd VLM4VLA
+git checkout b4ddb40       # or the latest main with patches applied
+# 2. Training / inference deps
+conda create -y -n vlm4vla_eval python=3.10
+conda activate vlm4vla_eval
+pip install -e .
+pip install -e ./openvla
+git clone https://github.com/moojink/dlimp_openvla.git /tmp/dlimp_openvla
+pip install -e /tmp/dlimp_openvla
+pip install -U hydra-core
+pip install bitsandbytes pretty_errors deepspeed qwen-vl-utils decord accelerate
+pip install 'huggingface_hub<1.0,>=0.34.0'
+pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl
+# 3. SimplerEnv stack (conflicts in numpy, do it last)
+git clone https://github.com/simpler-env/SimplerEnv --recurse-submodules /tmp/SimplerEnv
+cd /tmp/SimplerEnv/ManiSkill2_real2sim && pip install -e .
+cd /tmp/SimplerEnv && pip install -e .
+pip install mediapy 'numpy==1.24.4' 'opencv-python<4.11' 'setuptools<81'
+# 4. Asset symlink expected by scripts/bridge.bash
+ln -sfn /tmp/SimplerEnv/ManiSkill2_real2sim/data/real_inpainting ${VLM4VLA_REPO}/real_inpainting
+```
+If you *cannot* use the fork, apply the two patches in `patches/` to the upstream repo:
+```bash
+git apply patches/eval_calvin_model_wrapper.py.patch
+git apply patches/eval_simpler_main_inference.py.patch
+```
+(These are `git show` dumps; use `-p1` through `git apply` as usual.)
+`requirements-eval.txt` is the exact `pip freeze` of the conda env that produced the sweep results below; use it if you need to pin precisely.
+## Fix up paths in `project.json`
+The training config hard-codes paths that were valid on the training box (`/workspace/models/...`, `/workspace/data/...`). Before inference, point them at this bundle:
+```bash
+python - <<'PY'
+import json, os
+BUNDLE = os.path.abspath(".")
+p = os.path.join(BUNDLE, "project.json")
+c = json.load(open(p))
+qwen = os.path.join(BUNDLE, "configs/qwen2.5-vl-3b-instruct")
+c["model_path"] = qwen
+c["model_config"] = os.path.join(qwen, "config.json")
+c["tokenizer"]["pretrained_model_name_or_path"] = qwen
+c["vlm"]["pretrained_model_name_or_path"] = qwen
+# data_root_dir is unused at inference; leave as-is (or point to a Bridge copy if you resume training)
+json.dump(c, open(p, "w"), indent=4)
+print("patched", p)
+PY
+```
+## Also copy the BridgeV2 action stats into the repo
+The `BaseModelInference` class reads `configs/data/oxe_dataset_stats/dataset_statistics_bridge.json` relative to `$CWD`. Inside the repo:
+```bash
+cp /path/to/run_b_best/dataset_statistics_bridge.json \
+   VLM4VLA/configs/data/oxe_dataset_stats/dataset_statistics_bridge.json
+```
+## One-shot eval (best cell)
+Our headline result is `execute_step=2` at step 30000, avg SR = 30.21%. Run it with a single call of the canonical bridge script:
+```bash
+cd VLM4VLA
+conda activate vlm4vla_eval
+export TF_CPP_MIN_LOG_LEVEL=2
+BUNDLE=/abs/path/to/run_b_best
+bash scripts/bridge.bash \
+    ${BUNDLE}/stepstep=0030000.fp32.pt \
+    ${BUNDLE}/project.json \
+    2 \
+    0            # GPU index
+```
+Wall time: ~17-25 min on one A100 (4 tasks, 24 episodes each, up to 120 env steps/episode).
+Per-task `Average success` lines get printed into stdout; find them with:
+```bash
+grep -n 'Average success' eval.log
+```
+Expected output (identical up to sim-stochasticity):
+```
+PutCarrotOnPlate:           16.7 %
+StackGreenCubeOnYellow:      0.0 %
+PutSpoonOnTableCloth:       12.5 %
+PutEggplantInBasket:        91.7 %
+-> avg = 30.21 %
+```
+## Full sweep (reproduces paper's `eval_ckpts_bridge.py` protocol)
+The paper's protocol iterates over `execute_step in [4, 2, 1]` and picks the best cell, which on 8 GPUs finishes in ~100 min. Use the parallel launcher from the fork:
+```bash
+cd VLM4VLA
+conda activate vlm4vla_eval
+python eval/simpler/sweep_parallel_bridge.py \
+    --base-path /abs/path/to/run_b_best \
+    --ngpu 8 --exec-steps 4,2,1
+```
+See `RESULTS.md` for the full 30-cell SR matrix we obtained.
+## Known quirks
+- `StackGreenCubeOnYellow` has **0 % SR** across all 10 checkpoints × 3 execute_steps. This is a known hard task in SimplerBridge; the paper's run (b) also scores near-zero on it (the 23.95 % headline is dominated by eggplant).
+- `PutEggplantInBasket` is the dominant SR driver (up to 100 %). Evaluate it at both exec=4 and exec=2 for the most stable number.
+- `execute_step=1` is consistently the slowest (~30 min per cell) and rarely the best. `execute_step=4` is fastest (~21 min).
+- Vision freeze verification after load: at startup the code prints `Trainable Model Parameters: 3092.51M` and `-- trainable backbone Parameters: 3085.94M`. If you see 3761 M the freeze did not take effect (the ViT is trainable).
+## Caveats on reproduction
+- Training was done on this fork with `learning_rate=5e-5` (in the LOCAL config) vs the paper's non-LOCAL config setting of `2e-5`. Despite the larger LR, our converged SR *exceeds* the paper's, so the LR difference is not limiting reproduction.
+- We trained `max_steps=50000`. The paper's `max_steps` value is the same; however the peak performance in our sweep is at **step 30000**, not step 50000. If you are retraining, save often; don't assume the last step is best.
+- SimplerEnv upstream may drift. We pinned the specific SimplerEnv + ManiSkill2_real2sim HEAD as of Tue Apr 21 2026; future versions may perturb SR by a few pp due to sim changes.

RESULTS.md ADDED Viewed

	@@ -0,0 +1,96 @@

+# Run (b) – full SimplerBridge sweep (30 cells, 10 ckpts × 3 exec_steps)
+Training: Qwen2.5-VL-3B-Instruct + FCDecoder (action head), **vision encoder frozen**, BridgeData V2 RLDS, Lightning + DeepSpeed stage-2, bf16, 50 000 opt steps, global batch 512 (8 per GPU × 8 accumulate × 8x A100-80GB), LR 5e-5. See `project.json` for the complete config snapshot.
+Eval: `scripts/bridge.bash` – 4 visual-matching tasks, 24 episodes per task, `max_episode_steps=60` (120 for `PutEggplantInBasket`), `control_freq=5`, `sim_freq=500`, 8x A100 running 30 `(ckpt, execute_step)` cells in parallel (~107 min wall).
+Numbers below are per-task success rates in **percent** (24 episodes per task), followed by the 4-task mean (`avg`). The paper target for run (b) is **23.95 %**.
+## Per-cell SR matrix
+| step   | exec | carrot | stack | spoon | eggplant | avg SR |
+|-------:|-----:|-------:|------:|------:|---------:|-------:|
+|  5000  |   1  |   0.0  |  0.0  |   0.0 |    12.5  |   3.12 |
+|  5000  |   2  |   0.0  |  0.0  |   0.0 |     0.0  |   0.00 |
+|  5000  |   4  |   0.0  |  0.0  |   0.0 |    33.3  |   8.33 |
+| 10000  |   1  |   0.0  |  0.0  |   0.0 |     0.0  |   0.00 |
+| 10000  |   2  |   0.0  |  0.0  |   0.0 |     0.0  |   0.00 |
+| 10000  |   4  |   0.0  |  0.0  |   0.0 |     0.0  |   0.00 |
+| 15000  |   1  |   0.0  |  0.0  |   0.0 |    75.0  |  18.75 |
+| 15000  |   2  |   0.0  |  0.0  |   0.0 |    91.7  |  22.92 |
+| 15000  |   4  |   0.0  |  0.0  |   0.0 |    79.2  |  19.79 |
+| 20000  |   1  |   0.0  |  0.0  |   0.0 |    70.8  |  17.71 |
+| 20000  |   2  |   0.0  |  0.0  |   0.0 |    62.5  |  15.62 |
+| 20000  |   4  |   0.0  |  0.0  |   0.0 |    75.0  |  18.75 |
+| 25000  |   1  |   0.0  |  0.0  |   4.2 |    58.3  |  15.62 |
+| 25000  |   2  |   4.2  |  0.0  |   4.2 |    37.5  |  11.46 |
+| 25000  |   4  |   0.0  |  0.0  |   4.2 |    20.8  |   6.25 |
+| **30000** | **2** | **16.7** | **0.0** | **12.5** | **91.7** | **30.21** |
+| 30000  |   4  |  12.5  |  0.0  |   4.2 |   100.0  |  29.17 |
+| 30000  |   1  |  12.5  |  0.0  |   4.2 |    87.5  |  26.04 |
+| 35000  |   1  |   0.0  |  0.0  |   4.2 |    70.8  |  18.75 |
+| 35000  |   2  |   4.2  |  0.0  |   0.0 |    91.7  |  23.96 |
+| 35000  |   4  |   0.0  |  0.0  |   0.0 |    95.8  |  23.96 |
+| 40000  |   1  |   8.3  |  0.0  |   0.0 |    45.8  |  13.54 |
+| 40000  |   2  |   8.3  |  0.0  |   0.0 |    50.0  |  14.58 |
+| 40000  |   4  |   0.0  |  0.0  |   0.0 |    16.7  |   4.17 |
+| 45000  |   1  |   0.0  |  0.0  |   0.0 |    83.3  |  20.83 |
+| 45000  |   2  |   0.0  |  0.0  |   0.0 |    91.7  |  22.92 |
+| 45000  |   4  |   0.0  |  0.0  |   0.0 |    95.8  |  23.96 |
+| 50000  |   1  |   4.2  |  0.0  |   0.0 |    58.3  |  15.62 |
+| 50000  |   2  |   4.2  |  0.0  |   0.0 |    79.2  |  20.83 |
+| 50000  |   4  |   8.3  |  0.0  |   4.2 |    45.8  |  14.58 |
+## Aggregates
+### Best cell per `execute_step`
+| exec | best step | avg SR |
+|-----:|----------:|-------:|
+|   4  | 30000     | 29.17  |
+|   2  | 30000     | **30.21** |
+|   1  | 30000     | 26.04  |
+### Best cell per step (across `execute_step`)
+| step   | best exec | avg SR |
+|-------:|----------:|-------:|
+|  5000  |    4      |   8.33 |
+| 10000  |    1      |   0.00 |
+| 15000  |    2      |  22.92 |
+| 20000  |    4      |  18.75 |
+| 25000  |    1      |  15.62 |
+| 30000  |    2      | **30.21** |
+| 35000  |  4 or 2   |  23.96 |
+| 40000  |    2      |  14.58 |
+| 45000  |    4      |  23.96 |
+| 50000  |    2      |  20.83 |
+### Overall best
+- **step=30000, exec_step=2, avg SR = 30.21 %** (paper: 23.95 %, **+6.26 pp**)
+- Per-task: carrot 16.7 · stack 0.0 · spoon 12.5 · eggplant 91.7
+## Key observations
+1. **The peak is at step 30 000, not step 50 000.** Evaluating only the final checkpoint (our earlier single-cell run) gave 14.58 % at exec=4 – 16 pp below step-30k. A single-step eval would not have reproduced the paper.
+2. **Both execute_step 4 and 2 are competitive.** Exec=1 is consistently weakest and ~50 % slower (because every env step requires a fresh forward pass). Exec=4 finishes fastest; exec=2 generally gives the best SR in the converged regime (step 30k-45k).
+3. **StackGreenCubeOnYellow is 0.0 % everywhere.** That is consistent with the paper's own per-task breakdown; this task is dominated by other policies' zero-shot priors and is near-impossible without large-scale stacking data in the finetune mix.
+4. **Eggplant-in-basket carries the SR.** Peaking at 100 % (step 30k, exec=4) and mostly 80-95 % in the converged regime.
+5. **Learning-rate comparison.** Our LOCAL config uses `lr=5e-5`; the paper's non-LOCAL config uses `lr=2e-5`. Despite the 2.5× larger LR, our converged SR exceeds the paper's headline by 6 pp. The larger LR may have pushed the peak earlier (step 30k) relative to where the paper's run peaks.
+6. **Training-time freeze verification held end-to-end.** The converter-side log at `convert_ckpt_standalone.py` reports `Reconstructed Frozen fp32 state dict with 390 params 668 684 288 elements` at every save step – exactly the Qwen2.5-VL ViT parameter count. The loaded state dict matches `BaseTrainer` with `missing=0, unexpected=0`.
+7. **Inference-time shape check.** The policy's `inference_step` returns an action chunk of shape `(B=1, seq=1, chunk=4, act_dim=7)` – i.e. 4 distinct 7-D actions per inference. This matches the `execute_step` semantics in `eval/calvin/model_wrapper.py`.
+## Reproduction command
+```bash
+cd VLM4VLA
+conda activate vlm4vla_eval
+python eval/simpler/sweep_parallel_bridge.py \
+    --base-path /abs/path/to/run_b_best \
+    --ngpu 8 --exec-steps 4,2,1
+```
+Wall time on 8× A100-80GB: **~107 min** (first wave starts at t=0, last cell finishes at t≈107m). RAM peak ~20 GB per process (transient during ckpt load); steady VRAM ~9 GB per process.
+See `INFERENCE.md` for the single-cell command that reproduces just the best cell.

configs/qwen2.5-vl-3b-instruct/LICENSE ADDED Viewed

	@@ -0,0 +1,54 @@

+Qwen RESEARCH LICENSE AGREEMENT
+Qwen RESEARCH LICENSE AGREEMENT Release Date: September 19, 2024
+By clicking to agree or by using or distributing any portion or element of the Qwen Materials, you will be deemed to have recognized and accepted the content of this Agreement, which is effective immediately.
+1. Definitions
+    a. This Qwen RESEARCH LICENSE AGREEMENT (this "Agreement") shall mean the terms and conditions for use, reproduction, distribution and modification of the Materials as defined by this Agreement.
+    b. "We" (or "Us") shall mean Alibaba Cloud.
+    c. "You" (or "Your") shall mean a natural person or legal entity exercising the rights granted by this Agreement and/or using the Materials for any purpose and in any field of use.
+    d. "Third Parties" shall mean individuals or legal entities that are not under common control with us or you.
+    e. "Qwen" shall mean the large language models, and software and algorithms, consisting of trained model weights, parameters (including optimizer states), machine-learning model code, inference-enabling code, training-enabling code, fine-tuning enabling code and other elements of the foregoing distributed by us.
+    f. "Materials" shall mean, collectively, Alibaba Cloud's proprietary Qwen and Documentation (and any portion thereof) made available under this Agreement.
+    g. "Source" form shall mean the preferred form for making modifications, including but not limited to model source code, documentation source, and configuration files.
+    h. "Object" form shall mean any form resulting from mechanical transformation or translation of a Source form, including but not limited to compiled object code, generated documentation, and conversions to other media types.
+    i. "Non-Commercial" shall mean for research or evaluation purposes only.
+2. Grant of Rights
+    a. You are granted a non-exclusive, worldwide, non-transferable and royalty-free limited license under Alibaba Cloud's intellectual property or other rights owned by us embodied in the Materials to use, reproduce, distribute, copy, create derivative works of, and make modifications to the Materials FOR NON-COMMERCIAL PURPOSES ONLY.
+    b. If you are commercially using the Materials, you shall request a license from us.
+3. Redistribution
+You may distribute copies or make the Materials, or derivative works thereof, available as part of a product or service that contains any of them, with or without modifications, and in Source or Object form, provided that you meet the following conditions:
+    a. You shall give any other recipients of the Materials or derivative works a copy of this Agreement;
+    b. You shall cause any modified files to carry prominent notices stating that you changed the files;
+    c. You shall retain in all copies of the Materials that you distribute the following attribution notices within a "Notice" text file distributed as a part of such copies: "Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) Alibaba Cloud. All Rights Reserved."; and
+    d. You may add your own copyright statement to your modifications and may provide additional or different license terms and conditions for use, reproduction, or distribution of your modifications, or for any such derivative works as a whole, provided your use, reproduction, and distribution of the work otherwise complies with the terms and conditions of this Agreement.
+4. Rules of use
+    a. The Materials may be subject to export controls or restrictions in China, the United States or other countries or regions. You shall comply with applicable laws and regulations in your use of the Materials.
+    b. If you use the Materials or any outputs or results therefrom to create, train, fine-tune, or improve an AI model that is distributed or made available, you shall prominently display “Built with Qwen” or “Improved using Qwen” in the related product documentation.
+5. Intellectual Property
+    a. We retain ownership of all intellectual property rights in and to the Materials and derivatives made by or for us. Conditioned upon compliance with the terms and conditions of this Agreement, with respect to any derivative works and modifications of the Materials that are made by you, you are and will be the owner of such derivative works and modifications.
+    b. No trademark license is granted to use the trade names, trademarks, service marks, or product names of us, except as required to fulfill notice requirements under this Agreement or as required for reasonable and customary use in describing and redistributing the Materials.
+    c. If you commence a lawsuit or other proceedings (including a cross-claim or counterclaim in a lawsuit) against us or any entity alleging that the Materials or any output therefrom, or any part of the foregoing, infringe any intellectual property or other right owned or licensable by you, then all licenses granted to you under this Agreement shall terminate as of the date such lawsuit or other proceeding is commenced or brought.
+6. Disclaimer of Warranty and Limitation of Liability
+    a. We are not obligated to support, update, provide training for, or develop any further version of the Qwen Materials or to grant any license thereto.
+    b. THE MATERIALS ARE PROVIDED "AS IS" WITHOUT ANY EXPRESS OR IMPLIED WARRANTY OF ANY KIND INCLUDING WARRANTIES OF MERCHANTABILITY, NONINFRINGEMENT, OR FITNESS FOR A PARTICULAR PURPOSE. WE MAKE NO WARRANTY AND ASSUME NO RESPONSIBILITY FOR THE SAFETY OR STABILITY OF THE MATERIALS AND ANY OUTPUT THEREFROM.
+    c. IN NO EVENT SHALL WE BE LIABLE TO YOU FOR ANY DAMAGES, INCLUDING, BUT NOT LIMITED TO ANY DIRECT, OR INDIRECT, SPECIAL OR CONSEQUENTIAL DAMAGES ARISING FROM YOUR USE OR INABILITY TO USE THE MATERIALS OR ANY OUTPUT OF IT, NO MATTER HOW IT’S CAUSED.
+    d. You will defend, indemnify and hold harmless us from and against any claim by any third party arising out of or related to your use or distribution of the Materials.
+7. Survival and Termination.
+    a. The term of this Agreement shall commence upon your acceptance of this Agreement or access to the Materials and will continue in full force and effect until terminated in accordance with the terms and conditions herein.
+    b. We may terminate this Agreement if you breach any of the terms or conditions of this Agreement. Upon termination of this Agreement, you must delete and cease use of the Materials. Sections 6 and 8 shall survive the termination of this Agreement.
+8. Governing Law and Jurisdiction.
+    a. This Agreement and any dispute arising out of or relating to it will be governed by the laws of China, without regard to conflict of law principles, and the UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement.
+    b. The People's Courts in Hangzhou City shall have exclusive jurisdiction over any dispute arising out of this Agreement.
+9. Other Terms and Conditions.
+    a. Any arrangements, understandings, or agreements regarding the Material not stated herein are separate from and independent of the terms and conditions of this Agreement. You shall request a separate license from us, if you use the Materials in ways not expressly agreed to in this Agreement.
+    b. We shall not be bound by any additional or different terms or conditions communicated by you unless expressly agreed.

configs/qwen2.5-vl-3b-instruct/README.md ADDED Viewed

	@@ -0,0 +1,525 @@

+---
+license_name: qwen-research
+license_link: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE
+language:
+- en
+pipeline_tag: image-text-to-text
+tags:
+- multimodal
+library_name: transformers
+---
+# Qwen2.5-VL-3B-Instruct
+<a href="https://chat.qwenlm.ai/" target="_blank" style="margin: 2px;">
+    <img alt="Chat" src="https://img.shields.io/badge/%F0%9F%92%9C%EF%B8%8F%20Qwen%20Chat%20-536af5" style="display: inline-block; vertical-align: middle;"/>
+</a>
+## Introduction
+In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL.
+#### Key Enhancements:
+* **Understand things visually**: Qwen2.5-VL is not only proficient in recognizing common objects such as flowers, birds, fish, and insects, but it is highly capable of analyzing texts, charts, icons, graphics, and layouts within images.
+* **Being agentic**: Qwen2.5-VL directly plays as a visual agent that can reason and dynamically direct tools, which is capable of computer use and phone use.
+* **Understanding long videos and capturing events**: Qwen2.5-VL can comprehend videos of over 1 hour, and this time it has a new ability of cpaturing event by pinpointing the relevant video segments.
+* **Capable of visual localization in different formats**: Qwen2.5-VL can accurately localize objects in an image by generating bounding boxes or points, and it can provide stable JSON outputs for coordinates and attributes.
+* **Generating structured outputs**: for data like scans of invoices, forms, tables, etc. Qwen2.5-VL supports structured outputs of their contents, benefiting usages in finance, commerce, etc.
+#### Model Architecture Updates:
+* **Dynamic Resolution and Frame Rate Training for Video Understanding**:
+We extend dynamic resolution to the temporal dimension by adopting dynamic FPS sampling, enabling the model to comprehend videos at various sampling rates. Accordingly, we update mRoPE in the time dimension with IDs and absolute time alignment, enabling the model to learn temporal sequence and speed, and ultimately acquire the ability to pinpoint specific moments.
+<p align="center">
+    <img src="https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen2.5-VL/qwen2.5vl_arc.jpeg" width="80%"/>
+<p>
+* **Streamlined and Efficient Vision Encoder**
+We enhance both training and inference speeds by strategically implementing window attention into the ViT. The ViT architecture is further optimized with SwiGLU and RMSNorm, aligning it with the structure of the Qwen2.5 LLM.
+We have three models with 3, 7 and 72 billion parameters. This repo contains the instruction-tuned 3B Qwen2.5-VL model. For more information, visit our [Blog](https://qwenlm.github.io/blog/qwen2.5-vl/) and [GitHub](https://github.com/QwenLM/Qwen2.5-VL).
+## Evaluation
+### Image benchmark
+| Benchmark | InternVL2.5-4B |Qwen2-VL-7B |Qwen2.5-VL-3B |
+| :--- | :---:  | :---: | :---: |
+| MMMU<sub>val</sub>  | 52.3 | 54.1 | 53.1|
+| MMMU-Pro<sub>val</sub>  | **32.7** | 30.5 | 31.6|
+| AI2D<sub>test</sub> | 81.4 | **83.0** | 81.5 |
+| DocVQA<sub>test</sub>  | 91.6 | 94.5 | **93.9** |
+| InfoVQA<sub>test</sub>  | 72.1 | 76.5 | **77.1** |
+| TextVQA<sub>val</sub>  | 76.8 | **84.3** | 79.3|
+| MMBench-V1.1<sub>test</sub>  | 79.3 | **80.7** | 77.6 |
+| MMStar | 58.3 | **60.7** | 55.9 |
+| MathVista<sub>testmini</sub>  | 60.5 | 58.2 | **62.3** |
+| MathVision<sub>full</sub>  | 20.9 | 16.3  | **21.2** |
+### Video benchmark
+| Benchmark | InternVL2.5-4B | Qwen2-VL-7B | Qwen2.5-VL-3B |
+| :--- | :---:  | :---: | :---: |
+| MVBench | 71.6 | 67.0 | 67.0 |
+| VideoMME | 63.6/62.3 | 69.0/63.3 | 67.6/61.5 |
+| MLVU | 48.3 | - | 68.2 |
+| LVBench | - | - | 43.3 |
+| MMBench-Video | 1.73 | 1.44 | 1.63 |
+| EgoSchema | - | - | 64.8 |
+| PerceptionTest | - | - | 66.9 |
+| TempCompass | - | - | 64.4 |
+| LongVideoBench | 55.2 | 55.6 | 54.2 |
+| CharadesSTA/mIoU | - | - | 38.8 |
+### Agent benchmark
+| Benchmarks              | Qwen2.5-VL-3B |
+|-------------------------|---------------|
+| ScreenSpot              |     55.5    |
+| ScreenSpot Pro          |     23.9    |
+| AITZ_EM                 |  	76.9    |
+| Android Control High_EM |    	63.7    |
+| Android Control Low_EM  |  	22.2    |
+| AndroidWorld_SR         | 	90.8  	|
+| MobileMiniWob++_SR      | 	67.9    |
+## Requirements
+The code of Qwen2.5-VL has been in the latest Hugging face transformers and we advise you to build from source with command:
+```
+pip install git+https://github.com/huggingface/transformers accelerate
+```
+or you might encounter the following error:
+```
+KeyError: 'qwen2_5_vl'
+```
+## Quickstart
+Below, we provide simple examples to show how to use Qwen2.5-VL with 🤖 ModelScope and 🤗 Transformers.
+The code of Qwen2.5-VL has been in the latest Hugging face transformers and we advise you to build from source with command:
+```
+pip install git+https://github.com/huggingface/transformers accelerate
+```
+or you might encounter the following error:
+```
+KeyError: 'qwen2_5_vl'
+```
+We offer a toolkit to help you handle various types of visual input more conveniently, as if you were using an API. This includes base64, URLs, and interleaved images and videos. You can install it using the following command:
+```bash
+# It's highly recommanded to use `[decord]` feature for faster video loading.
+pip install qwen-vl-utils[decord]==0.0.8
+```
+If you are not using Linux, you might not be able to install `decord` from PyPI. In that case, you can use `pip install qwen-vl-utils` which will fall back to using torchvision for video processing. However, you can still [install decord from source](https://github.com/dmlc/decord?tab=readme-ov-file#install-from-source) to get decord used when loading video.
+### Using 🤗  Transformers to Chat
+Here we show a code snippet to show you how to use the chat model with `transformers` and `qwen_vl_utils`:
+```python
+from transformers import Qwen2_5_VLForConditionalGeneration, AutoTokenizer, AutoProcessor
+from qwen_vl_utils import process_vision_info
+# default: Load the model on the available device(s)
+model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
+    "Qwen/Qwen2.5-VL-3B-Instruct", torch_dtype="auto", device_map="auto"
+)
+# We recommend enabling flash_attention_2 for better acceleration and memory saving, especially in multi-image and video scenarios.
+# model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
+#     "Qwen/Qwen2.5-VL-3B-Instruct",
+#     torch_dtype=torch.bfloat16,
+#     attn_implementation="flash_attention_2",
+#     device_map="auto",
+# )
+# default processer
+processor = AutoProcessor.from_pretrained("Qwen/Qwen2.5-VL-3B-Instruct")
+# The default range for the number of visual tokens per image in the model is 4-16384.
+# You can set min_pixels and max_pixels according to your needs, such as a token range of 256-1280, to balance performance and cost.
+# min_pixels = 256*28*28
+# max_pixels = 1280*28*28
+# processor = AutoProcessor.from_pretrained("Qwen/Qwen2.5-VL-3B-Instruct", min_pixels=min_pixels, max_pixels=max_pixels)
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {
+                "type": "image",
+                "image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
+            },
+            {"type": "text", "text": "Describe this image."},
+        ],
+    }
+]
+# Preparation for inference
+text = processor.apply_chat_template(
+    messages, tokenize=False, add_generation_prompt=True
+)
+image_inputs, video_inputs = process_vision_info(messages)
+inputs = processor(
+    text=[text],
+    images=image_inputs,
+    videos=video_inputs,
+    padding=True,
+    return_tensors="pt",
+)
+inputs = inputs.to("cuda")
+# Inference: Generation of the output
+generated_ids = model.generate(**inputs, max_new_tokens=128)
+generated_ids_trimmed = [
+    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
+]
+output_text = processor.batch_decode(
+    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
+)
+print(output_text)
+```
+<details>
+<summary>Multi image inference</summary>
+```python
+# Messages containing multiple images and a text query
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {"type": "image", "image": "file:///path/to/image1.jpg"},
+            {"type": "image", "image": "file:///path/to/image2.jpg"},
+            {"type": "text", "text": "Identify the similarities between these images."},
+        ],
+    }
+]
+# Preparation for inference
+text = processor.apply_chat_template(
+    messages, tokenize=False, add_generation_prompt=True
+)
+image_inputs, video_inputs = process_vision_info(messages)
+inputs = processor(
+    text=[text],
+    images=image_inputs,
+    videos=video_inputs,
+    padding=True,
+    return_tensors="pt",
+)
+inputs = inputs.to("cuda")
+# Inference
+generated_ids = model.generate(**inputs, max_new_tokens=128)
+generated_ids_trimmed = [
+    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
+]
+output_text = processor.batch_decode(
+    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
+)
+print(output_text)
+```
+</details>
+<details>
+<summary>Video inference</summary>
+```python
+# Messages containing a images list as a video and a text query
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {
+                "type": "video",
+                "video": [
+                    "file:///path/to/frame1.jpg",
+                    "file:///path/to/frame2.jpg",
+                    "file:///path/to/frame3.jpg",
+                    "file:///path/to/frame4.jpg",
+                ],
+            },
+            {"type": "text", "text": "Describe this video."},
+        ],
+    }
+]
+# Messages containing a local video path and a text query
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {
+                "type": "video",
+                "video": "file:///path/to/video1.mp4",
+                "max_pixels": 360 * 420,
+                "fps": 1.0,
+            },
+            {"type": "text", "text": "Describe this video."},
+        ],
+    }
+]
+# Messages containing a video url and a text query
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {
+                "type": "video",
+                "video": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen2-VL/space_woaudio.mp4",
+            },
+            {"type": "text", "text": "Describe this video."},
+        ],
+    }
+]
+#In Qwen 2.5 VL, frame rate information is also input into the model to align with absolute time.
+# Preparation for inference
+text = processor.apply_chat_template(
+    messages, tokenize=False, add_generation_prompt=True
+)
+image_inputs, video_inputs, video_kwargs = process_vision_info(messages, return_video_kwargs=True)
+inputs = processor(
+    text=[text],
+    images=image_inputs,
+    videos=video_inputs,
+    fps=fps,
+    padding=True,
+    return_tensors="pt",
+    **video_kwargs,
+)
+inputs = inputs.to("cuda")
+# Inference
+generated_ids = model.generate(**inputs, max_new_tokens=128)
+generated_ids_trimmed = [
+    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
+]
+output_text = processor.batch_decode(
+    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
+)
+print(output_text)
+```
+Video URL compatibility largely depends on the third-party library version. The details are in the table below. change the backend by `FORCE_QWENVL_VIDEO_READER=torchvision` or `FORCE_QWENVL_VIDEO_READER=decord` if you prefer not to use the default one.
+| Backend     | HTTP | HTTPS |
+|-------------|------|-------|
+| torchvision >= 0.19.0 | ✅  | ✅   |
+| torchvision < 0.19.0  | ❌  | ❌   |
+| decord      | ✅  | ❌   |
+</details>
+<details>
+<summary>Batch inference</summary>
+```python
+# Sample messages for batch inference
+messages1 = [
+    {
+        "role": "user",
+        "content": [
+            {"type": "image", "image": "file:///path/to/image1.jpg"},
+            {"type": "image", "image": "file:///path/to/image2.jpg"},
+            {"type": "text", "text": "What are the common elements in these pictures?"},
+        ],
+    }
+]
+messages2 = [
+    {"role": "system", "content": "You are a helpful assistant."},
+    {"role": "user", "content": "Who are you?"},
+]
+# Combine messages for batch processing
+messages = [messages1, messages2]
+# Preparation for batch inference
+texts = [
+    processor.apply_chat_template(msg, tokenize=False, add_generation_prompt=True)
+    for msg in messages
+]
+image_inputs, video_inputs = process_vision_info(messages)
+inputs = processor(
+    text=texts,
+    images=image_inputs,
+    videos=video_inputs,
+    padding=True,
+    return_tensors="pt",
+)
+inputs = inputs.to("cuda")
+# Batch Inference
+generated_ids = model.generate(**inputs, max_new_tokens=128)
+generated_ids_trimmed = [
+    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
+]
+output_texts = processor.batch_decode(
+    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
+)
+print(output_texts)
+```
+</details>
+### 🤖 ModelScope
+We strongly advise users especially those in mainland China to use ModelScope. `snapshot_download` can help you solve issues concerning downloading checkpoints.
+### More Usage Tips
+For input images, we support local files, base64, and URLs. For videos, we currently only support local files.
+```python
+# You can directly insert a local file path, a URL, or a base64-encoded image into the position where you want in the text.
+## Local file path
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {"type": "image", "image": "file:///path/to/your/image.jpg"},
+            {"type": "text", "text": "Describe this image."},
+        ],
+    }
+]
+## Image URL
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {"type": "image", "image": "http://path/to/your/image.jpg"},
+            {"type": "text", "text": "Describe this image."},
+        ],
+    }
+]
+## Base64 encoded image
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {"type": "image", "image": "data:image;base64,/9j/..."},
+            {"type": "text", "text": "Describe this image."},
+        ],
+    }
+]
+```
+#### Image Resolution for performance boost
+The model supports a wide range of resolution inputs. By default, it uses the native resolution for input, but higher resolutions can enhance performance at the cost of more computation. Users can set the minimum and maximum number of pixels to achieve an optimal configuration for their needs, such as a token count range of 256-1280, to balance speed and memory usage.
+```python
+min_pixels = 256 * 28 * 28
+max_pixels = 1280 * 28 * 28
+processor = AutoProcessor.from_pretrained(
+    "Qwen/Qwen2.5-VL-3B-Instruct", min_pixels=min_pixels, max_pixels=max_pixels
+)
+```
+Besides, We provide two methods for fine-grained control over the image size input to the model:
+1. Define min_pixels and max_pixels: Images will be resized to maintain their aspect ratio within the range of min_pixels and max_pixels.
+2. Specify exact dimensions: Directly set `resized_height` and `resized_width`. These values will be rounded to the nearest multiple of 28.
+```python
+# min_pixels and max_pixels
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {
+                "type": "image",
+                "image": "file:///path/to/your/image.jpg",
+                "resized_height": 280,
+                "resized_width": 420,
+            },
+            {"type": "text", "text": "Describe this image."},
+        ],
+    }
+]
+# resized_height and resized_width
+messages = [
+    {
+        "role": "user",
+        "content": [
+            {
+                "type": "image",
+                "image": "file:///path/to/your/image.jpg",
+                "min_pixels": 50176,
+                "max_pixels": 50176,
+            },
+            {"type": "text", "text": "Describe this image."},
+        ],
+    }
+]
+```
+### Processing Long Texts
+The current `config.json` is set for context length up to 32,768 tokens.
+To handle extensive inputs exceeding 32,768 tokens, we utilize [YaRN](https://arxiv.org/abs/2309.00071), a technique for enhancing model length extrapolation, ensuring optimal performance on lengthy texts.
+For supported frameworks, you could add the following to `config.json` to enable YaRN:
+```
+{
+	...,
+    "type": "yarn",
+    "mrope_section": [
+        16,
+        24,
+        24
+    ],
+    "factor": 4,
+    "original_max_position_embeddings": 32768
+}
+```
+However, it should be noted that this method has a significant impact on the performance of temporal and spatial localization tasks, and is therefore not recommended for use.
+At the same time, for long video inputs, since MRoPE itself is more economical with ids, the max_position_embeddings can be directly modified to a larger value, such as 64k.
+## Citation
+If you find our work helpful, feel free to give us a cite.
+```
+@misc{qwen2.5-VL,
+    title = {Qwen2.5-VL},
+    url = {https://qwenlm.github.io/blog/qwen2.5-vl/},
+    author = {Qwen Team},
+    month = {January},
+    year = {2025}
+}
+@article{Qwen2VL,
+  title={Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution},
+  author={Wang, Peng and Bai, Shuai and Tan, Sinan and Wang, Shijie and Fan, Zhihao and Bai, Jinze and Chen, Keqin and Liu, Xuejing and Wang, Jialin and Ge, Wenbin and Fan, Yang and Dang, Kai and Du, Mengfei and Ren, Xuancheng and Men, Rui and Liu, Dayiheng and Zhou, Chang and Zhou, Jingren and Lin, Junyang},
+  journal={arXiv preprint arXiv:2409.12191},
+  year={2024}
+}
+@article{Qwen-VL,
+  title={Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond},
+  author={Bai, Jinze and Bai, Shuai and Yang, Shusheng and Wang, Shijie and Tan, Sinan and Wang, Peng and Lin, Junyang and Zhou, Chang and Zhou, Jingren},
+  journal={arXiv preprint arXiv:2308.12966},
+  year={2023}
+}
+```

configs/qwen2.5-vl-3b-instruct/chat_template.json ADDED Viewed

	@@ -0,0 +1,3 @@

+{
+    "chat_template": "{% set image_count = namespace(value=0) %}{% set video_count = namespace(value=0) %}{% for message in messages %}{% if loop.first and message['role'] != 'system' %}<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n{% endif %}<|im_start|>{{ message['role'] }}\n{% if message['content'] is string %}{{ message['content'] }}<|im_end|>\n{% else %}{% for content in message['content'] %}{% if content['type'] == 'image' or 'image' in content or 'image_url' in content %}{% set image_count.value = image_count.value + 1 %}{% if add_vision_id %}Picture {{ image_count.value }}: {% endif %}<|vision_start|><|image_pad|><|vision_end|>{% elif content['type'] == 'video' or 'video' in content %}{% set video_count.value = video_count.value + 1 %}{% if add_vision_id %}Video {{ video_count.value }}: {% endif %}<|vision_start|><|video_pad|><|vision_end|>{% elif 'text' in content %}{{ content['text'] }}{% endif %}{% endfor %}<|im_end|>\n{% endif %}{% endfor %}{% if add_generation_prompt %}<|im_start|>assistant\n{% endif %}"
+}

configs/qwen2.5-vl-3b-instruct/config.json ADDED Viewed

	@@ -0,0 +1,61 @@

+{
+  "architectures": [
+    "Qwen2_5_VLForConditionalGeneration"
+  ],
+  "attention_dropout": 0.0,
+  "bos_token_id": 151643,
+  "eos_token_id": 151645,
+  "vision_start_token_id": 151652,
+  "vision_end_token_id": 151653,
+  "vision_token_id": 151654,
+  "image_token_id": 151655,
+  "video_token_id": 151656,
+  "hidden_act": "silu",
+  "hidden_size": 2048,
+  "initializer_range": 0.02,
+  "intermediate_size": 11008,
+  "max_position_embeddings": 128000,
+  "max_window_layers": 70,
+  "model_type": "qwen2_5_vl",
+  "num_attention_heads": 16,
+  "num_hidden_layers": 36,
+  "num_key_value_heads": 2,
+  "rms_norm_eps": 1e-06,
+  "rope_theta": 1000000.0,
+  "sliding_window": 32768,
+  "tie_word_embeddings": true,
+  "torch_dtype": "bfloat16",
+  "transformers_version": "4.41.2",
+  "use_cache": true,
+  "use_sliding_window": false,
+  "vision_config": {
+    "depth": 32,
+    "hidden_act": "silu",
+    "hidden_size": 1280,
+    "intermediate_size": 3420,
+    "num_heads": 16,
+    "in_chans": 3,
+    "out_hidden_size": 2048,
+    "patch_size": 14,
+    "spatial_merge_size": 2,
+    "spatial_patch_size": 14,
+    "window_size": 112,
+    "fullatt_block_indexes": [
+      7,
+      15,
+      23,
+      31
+    ],
+    "tokens_per_second": 2,
+    "temporal_patch_size": 2
+  },
+  "rope_scaling": {
+    "type": "mrope",
+    "mrope_section": [
+      16,
+      24,
+      24
+    ]
+  },
+  "vocab_size": 151936
+}

configs/qwen2.5-vl-3b-instruct/generation_config.json ADDED Viewed

	@@ -0,0 +1,12 @@

+{
+  "bos_token_id": 151643,
+  "pad_token_id": 151643,
+  "do_sample": true,
+  "eos_token_id": [
+    151645,
+    151643
+  ],
+  "repetition_penalty": 1.05,
+  "temperature": 0.000001,
+  "transformers_version": "4.49.0"
+}

configs/qwen2.5-vl-3b-instruct/merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

configs/qwen2.5-vl-3b-instruct/model-00001-of-00002.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:41a8895c164b4d32bae6b302f4603fcbc1797f32dafa45c7e9bcda23c6755df8
+size 3982649232

configs/qwen2.5-vl-3b-instruct/model-00002-of-00002.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:365531ff8752420e89dee707b79d021fb2d6e25abafe486f080555a4fe6972e4
+size 3526688744

configs/qwen2.5-vl-3b-instruct/model.safetensors.index.json ADDED Viewed

	@@ -0,0 +1,831 @@

+{
+  "metadata": {
+    "total_size": 7509245952
+  },
+  "weight_map": {
+    "model.embed_tokens.weight": "model-00001-of-00002.safetensors",
+    "model.layers.0.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.0.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.0.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.0.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.0.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.0.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.1.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.1.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.1.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.1.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.1.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.1.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.10.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.10.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.10.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.10.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.10.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.10.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.10.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.10.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.10.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.10.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.10.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.10.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.11.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.11.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.11.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.11.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.11.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.11.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.11.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.11.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.11.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.11.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.11.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.11.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.12.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.12.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.12.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.12.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.12.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.12.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.12.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.12.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.12.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.12.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.12.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.12.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.13.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.13.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.13.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.13.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.13.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.13.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.13.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.13.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.13.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.13.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.13.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.13.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.14.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.14.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.14.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.14.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.14.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.14.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.14.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.14.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.14.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.14.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.14.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.14.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.15.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.15.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.15.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.15.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.15.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.15.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.15.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.15.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.15.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.15.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.15.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.15.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.16.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.16.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.16.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.16.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.16.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.16.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.16.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.16.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.16.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.16.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.16.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.16.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.17.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.17.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.17.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.17.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.17.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.17.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.17.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.17.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.17.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.17.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.17.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.17.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.18.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.18.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.18.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.18.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.18.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.18.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.18.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.18.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.18.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.18.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.18.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.18.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.19.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.19.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.19.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.19.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.19.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.19.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.19.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.19.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.19.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.19.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.19.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.19.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.2.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.2.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.2.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.2.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.2.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.2.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.2.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.2.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.2.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.20.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.20.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.20.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.20.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.20.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.20.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.20.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.20.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.20.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.20.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.20.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.20.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.21.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.21.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.21.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.21.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.21.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.21.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.21.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.21.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.21.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.21.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.21.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.21.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.22.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.22.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.22.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.22.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.22.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.22.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.22.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.22.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.22.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.22.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.22.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.22.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.23.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.23.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.23.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.23.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.23.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.23.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.23.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.23.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.23.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.23.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.23.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.23.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.24.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.24.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.24.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.24.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.24.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.24.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.24.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.24.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.24.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.24.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.24.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.24.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.25.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.25.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.25.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.25.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.25.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.25.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.25.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.25.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.25.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.25.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.25.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.25.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.26.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.26.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.26.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.26.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.26.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.26.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.26.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.26.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.26.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.26.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.26.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.26.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.27.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.27.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.27.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.27.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.27.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.27.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.27.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.27.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.27.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.27.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.27.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.27.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.28.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.28.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.28.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.28.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.28.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.28.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.28.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.28.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.28.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.28.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.28.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.28.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.29.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.29.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.29.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.29.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.29.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.29.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.29.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.29.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.29.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.29.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.29.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.29.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.3.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.3.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.3.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.3.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.3.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.3.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.3.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.3.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.3.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.3.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.30.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.30.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.30.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.30.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.30.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.30.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.30.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.30.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.30.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.30.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.30.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.30.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.31.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.31.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.31.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.31.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.31.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.31.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.31.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.31.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.31.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.31.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.31.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.31.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.32.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.32.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.32.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.32.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.32.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.32.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.32.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.32.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.32.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.32.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.32.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.32.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.33.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.33.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.33.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.33.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.33.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.33.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.33.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.33.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.33.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.33.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.33.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.33.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.34.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.34.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.34.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.34.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.34.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.34.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.34.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.34.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.34.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.34.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.34.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.34.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.35.input_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.35.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.35.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.35.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.35.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
+    "model.layers.35.self_attn.k_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.35.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.35.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.35.self_attn.q_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.35.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.35.self_attn.v_proj.bias": "model-00002-of-00002.safetensors",
+    "model.layers.35.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
+    "model.layers.4.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.4.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.4.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.4.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.4.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.4.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.4.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.4.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.4.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.4.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.5.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.5.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.5.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.5.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.5.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.5.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.5.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.5.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.5.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.5.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.5.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.5.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.6.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.6.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.6.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.6.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.6.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.6.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.6.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.6.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.6.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.6.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.6.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.6.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.7.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.7.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.7.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.7.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.7.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.7.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.7.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.7.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.7.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.7.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.7.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.7.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.8.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.8.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.8.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.8.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.8.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.8.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.8.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.8.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.8.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.8.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.8.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.8.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.9.input_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.9.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.9.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.9.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.9.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
+    "model.layers.9.self_attn.k_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.9.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.9.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.9.self_attn.q_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.9.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
+    "model.layers.9.self_attn.v_proj.bias": "model-00001-of-00002.safetensors",
+    "model.layers.9.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
+    "model.norm.weight": "model-00002-of-00002.safetensors",
+    "visual.blocks.0.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.0.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.1.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.10.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.11.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.12.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.13.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.14.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.15.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.16.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.17.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.18.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.19.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.2.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.20.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.21.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.22.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.23.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.24.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.25.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.26.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.27.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.28.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.29.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.3.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.30.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.31.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.4.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.5.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.6.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.7.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.8.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.attn.proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.attn.proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.attn.qkv.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.attn.qkv.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.mlp.down_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.mlp.gate_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.mlp.up_proj.bias": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.norm1.weight": "model-00001-of-00002.safetensors",
+    "visual.blocks.9.norm2.weight": "model-00001-of-00002.safetensors",
+    "visual.merger.ln_q.weight": "model-00001-of-00002.safetensors",
+    "visual.merger.mlp.0.bias": "model-00001-of-00002.safetensors",
+    "visual.merger.mlp.0.weight": "model-00001-of-00002.safetensors",
+    "visual.merger.mlp.2.bias": "model-00001-of-00002.safetensors",
+    "visual.merger.mlp.2.weight": "model-00001-of-00002.safetensors",
+    "visual.patch_embed.proj.weight": "model-00001-of-00002.safetensors"
+  }
+}

configs/qwen2.5-vl-3b-instruct/preprocessor_config.json ADDED Viewed

	@@ -0,0 +1,19 @@

+{
+  "min_pixels": 3136,
+  "max_pixels": 12845056,
+  "patch_size": 14,
+  "temporal_patch_size": 2,
+  "merge_size": 2,
+  "image_mean": [
+    0.48145466,
+    0.4578275,
+    0.40821073
+  ],
+  "image_std": [
+    0.26862954,
+    0.26130258,
+    0.27577711
+  ],
+  "image_processor_type": "Qwen2VLImageProcessor",
+  "processor_class": "Qwen2_5_VLProcessor"
+}

configs/qwen2.5-vl-3b-instruct/tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

configs/qwen2.5-vl-3b-instruct/tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,207 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "151643": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151644": {
+      "content": "<|im_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151645": {
+      "content": "<|im_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151646": {
+      "content": "<|object_ref_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151647": {
+      "content": "<|object_ref_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151648": {
+      "content": "<|box_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151649": {
+      "content": "<|box_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151650": {
+      "content": "<|quad_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151651": {
+      "content": "<|quad_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151652": {
+      "content": "<|vision_start|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151653": {
+      "content": "<|vision_end|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151654": {
+      "content": "<|vision_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151655": {
+      "content": "<|image_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151656": {
+      "content": "<|video_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "151657": {
+      "content": "<tool_call>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151658": {
+      "content": "</tool_call>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151659": {
+      "content": "<|fim_prefix|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151660": {
+      "content": "<|fim_middle|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151661": {
+      "content": "<|fim_suffix|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151662": {
+      "content": "<|fim_pad|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151663": {
+      "content": "<|repo_name|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "151664": {
+      "content": "<|file_sep|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    }
+  },
+  "additional_special_tokens": [
+    "<|im_start|>",
+    "<|im_end|>",
+    "<|object_ref_start|>",
+    "<|object_ref_end|>",
+    "<|box_start|>",
+    "<|box_end|>",
+    "<|quad_start|>",
+    "<|quad_end|>",
+    "<|vision_start|>",
+    "<|vision_end|>",
+    "<|vision_pad|>",
+    "<|image_pad|>",
+    "<|video_pad|>"
+  ],
+  "bos_token": null,
+  "chat_template": "{% set image_count = namespace(value=0) %}{% set video_count = namespace(value=0) %}{% for message in messages %}{% if loop.first and message['role'] != 'system' %}<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n{% endif %}<|im_start|>{{ message['role'] }}\n{% if message['content'] is string %}{{ message['content'] }}<|im_end|>\n{% else %}{% for content in message['content'] %}{% if content['type'] == 'image' or 'image' in content or 'image_url' in content %}{% set image_count.value = image_count.value + 1 %}{% if add_vision_id %}Picture {{ image_count.value }}: {% endif %}<|vision_start|><|image_pad|><|vision_end|>{% elif content['type'] == 'video' or 'video' in content %}{% set video_count.value = video_count.value + 1 %}{% if add_vision_id %}Video {{ video_count.value }}: {% endif %}<|vision_start|><|video_pad|><|vision_end|>{% elif 'text' in content %}{{ content['text'] }}{% endif %}{% endfor %}<|im_end|>\n{% endif %}{% endfor %}{% if add_generation_prompt %}<|im_start|>assistant\n{% endif %}",
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "<|im_end|>",
+  "errors": "replace",
+  "model_max_length": 131072,
+  "pad_token": "<|endoftext|>",
+  "split_special_tokens": false,
+  "tokenizer_class": "Qwen2Tokenizer",
+  "unk_token": null,
+  "add_bos_token": false
+}

configs/qwen2.5-vl-3b-instruct/vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff

dataset_statistics_bridge.json ADDED Viewed

	@@ -0,0 +1,116 @@

+{
+    "action": {
+        "mean": [
+            0.0002334193413844332,
+            0.0001300490548601374,
+            -0.0001276246621273458,
+            -0.00015565502690151334,
+            -0.0004039333143737167,
+            0.0002355769247515127,
+            0.5764579772949219
+        ],
+        "std": [
+            0.009765916503965855,
+            0.013689138926565647,
+            0.012667354196310043,
+            0.02853417582809925,
+            0.0306379534304142,
+            0.07691461592912674,
+            0.49737000465393066
+        ],
+        "max": [
+            0.41691166162490845,
+            0.25864794850349426,
+            0.21218234300613403,
+            3.122201919555664,
+            1.8618112802505493,
+            6.280478477478027,
+            1.0
+        ],
+        "min": [
+            -0.4007510244846344,
+            -0.13874775171279907,
+            -0.22553899884223938,
+            -3.2010786533355713,
+            -1.8618112802505493,
+            -6.279075622558594,
+            0.0
+        ],
+        "q01": [
+            -0.02872725307941437,
+            -0.04170349963009357,
+            -0.026093858778476715,
+            -0.08092105075716972,
+            -0.09288699507713317,
+            -0.20718276381492615,
+            0.0
+        ],
+        "q99": [
+            0.028309678435325586,
+            0.040855254605412394,
+            0.040161586627364146,
+            0.08192047759890528,
+            0.07792850524187081,
+            0.20382574498653397,
+            1.0
+        ]
+    },
+    "proprio": {
+        "mean": [
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0
+        ],
+        "std": [
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0
+        ],
+        "max": [
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0
+        ],
+        "min": [
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0
+        ],
+        "q01": [
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0
+        ],
+        "q99": [
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0,
+            0.0
+        ]
+    },
+    "num_transitions": 2135463,
+    "num_trajectories": 60064
+}

patches/convert_ckpt_standalone.py ADDED Viewed

	@@ -0,0 +1,18 @@

+"""Standalone DeepSpeed stage-2 -> FP32 single-file converter.
+Usage:
+    python scripts/convert_ckpt_standalone.py <ds_ckpt_dir> <fp32_out_path>
+"""
+import os
+import sys
+from vlm4vla.utils.zero_to_fp32 import convert_zero_checkpoint_to_fp32_state_dict
+src = sys.argv[1]
+dst = sys.argv[2]
+assert os.path.isdir(src), f"not a directory: {src}"
+os.makedirs(os.path.dirname(dst), exist_ok=True)
+print(f"converting {src} -> {dst}")
+convert_zero_checkpoint_to_fp32_state_dict(src, dst)
+print(f"done, size: {os.path.getsize(dst) / 1e9:.2f} GB")

patches/eval_calvin_model_wrapper.py.patch ADDED Viewed

	@@ -0,0 +1,39 @@

+commit b4ddb404e6bce2e116b04c598b9495a99bf40fdc
+Author: yunfeixie <x908717327@gmail.com>
+Date:   Wed Apr 22 02:33:50 2026 +0000
+    Eval: SimplerBridge eval harness fixes + parallel sweep launcher
+    - eval/simpler/eval_ckpts_bridge.py: parameterize base_path/exec_steps/device
+      via env vars (STEP_FILTER, EXEC_STEPS, CUDA_DEV); point default at
+      /workspace/ckpts_archive/run_b_all so reproduction on this box is one command.
+    - eval/simpler/main_inference.py: set args.policy_model from configs["model"]
+      so maniskill2_evaluator -> get_robot_control_mode no longer
+      AttributeErrors on Qwen2.5-VL configs.
+    - eval/calvin/model_wrapper.py: call get_text_function with the arity the
+      function actually has (2 args); the 4-arg call was dead since data_utils.py
+      defines get_text_function(tokenizer, tokenizer_type, max_length=256).
+    - eval/simpler/sweep_parallel_bridge.py: new launcher that schedules
+      (ckpt, execute_step) cells across NGPU GPUs concurrently, one cell per GPU.
+      Ports cleanly to other boxes by --base-path/--ngpu flags; full porting
+      guide lives in the module docstring. Expected wall time for 10 ckpts x
+      3 exec_steps on 8 A100s: ~90 min.
+    - eval/simpler/diag_one_episode.sh: single-episode helper used to confirm
+      FCDecoder output shape is (1, 1, 4, 7) at inference time, i.e. the harness
+      correctly replays 4 distinct actions at execute_step=4.
+    - .gitignore: ignore the /real_inpainting symlink that points at
+      SimplerEnv/ManiSkill2_real2sim/data/real_inpainting.
+diff --git a/eval/calvin/model_wrapper.py b/eval/calvin/model_wrapper.py
+index 11fcef1..4a89788 100644
+--- a/eval/calvin/model_wrapper.py
++++ b/eval/calvin/model_wrapper.py
+@@ -141,7 +141,7 @@ class CustomModel:
+         else:
+             robot_prompt = None
+         print('robot_prompt', robot_prompt)
+-        self.text_preprocess = get_text_function(self.model.model.tokenizer, configs["model"], qwen25_seq_id, robot_prompt)
++        self.text_preprocess = get_text_function(self.model.model.tokenizer, configs["model"])
+         self.action_space = self.configs["act_head"].get("action_space", "continuous")

patches/eval_simpler_main_inference.py.patch ADDED Viewed

	@@ -0,0 +1,38 @@

+commit b4ddb404e6bce2e116b04c598b9495a99bf40fdc
+Author: yunfeixie <x908717327@gmail.com>
+Date:   Wed Apr 22 02:33:50 2026 +0000
+    Eval: SimplerBridge eval harness fixes + parallel sweep launcher
+    - eval/simpler/eval_ckpts_bridge.py: parameterize base_path/exec_steps/device
+      via env vars (STEP_FILTER, EXEC_STEPS, CUDA_DEV); point default at
+      /workspace/ckpts_archive/run_b_all so reproduction on this box is one command.
+    - eval/simpler/main_inference.py: set args.policy_model from configs["model"]
+      so maniskill2_evaluator -> get_robot_control_mode no longer
+      AttributeErrors on Qwen2.5-VL configs.
+    - eval/calvin/model_wrapper.py: call get_text_function with the arity the
+      function actually has (2 args); the 4-arg call was dead since data_utils.py
+      defines get_text_function(tokenizer, tokenizer_type, max_length=256).
+    - eval/simpler/sweep_parallel_bridge.py: new launcher that schedules
+      (ckpt, execute_step) cells across NGPU GPUs concurrently, one cell per GPU.
+      Ports cleanly to other boxes by --base-path/--ngpu flags; full porting
+      guide lives in the module docstring. Expected wall time for 10 ckpts x
+      3 exec_steps on 8 A100s: ~90 min.
+    - eval/simpler/diag_one_episode.sh: single-episode helper used to confirm
+      FCDecoder output shape is (1, 1, 4, 7) at inference time, i.e. the harness
+      correctly replays 4 distinct actions at execute_step=4.
+    - .gitignore: ignore the /real_inpainting symlink that points at
+      SimplerEnv/ManiSkill2_real2sim/data/real_inpainting.
+diff --git a/eval/simpler/main_inference.py b/eval/simpler/main_inference.py
+index 59ce4d5..f438f74 100644
+--- a/eval/simpler/main_inference.py
++++ b/eval/simpler/main_inference.py
+@@ -235,6 +235,7 @@ if __name__ == "__main__":
+     args.model_name += f'_{configs["exp_name"]}'
+     if args.double_step:
+         args.model_name += "double"
++    args.policy_model = configs.get("model", "vlm4vla")
+     os.environ["DISPLAY"] = ""
+     # prevent a single jax process from taking up all the GPU memory
+     os.environ["XLA_PYTHON_CLIENT_PREALLOCATE"] = "false"

project.json ADDED Viewed

	@@ -0,0 +1,166 @@

+{
+    "robovlm_name": "RoboQwen25VL",
+    "parent": null,
+    "task_name": "bridge_finetune",
+    "model": "qwen25vl",
+    "model_url": "https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct",
+    "seq_len": 1,
+    "image_size": 224,
+    "image_mean": [
+        0.48145466,
+        0.4578275,
+        0.40821073
+    ],
+    "image_std": [
+        0.26862954,
+        0.26130258,
+        0.27577711
+    ],
+    "window_size": 1,
+    "fwd_pred_next_n": 4,
+    "arm_gripper_loss_ratio": 0.01,
+    "cap_loss_ratio": 0.05,
+    "fwd_loss_ratio": 0,
+    "seed": 123,
+    "batch_size": 8,
+    "num_workers": 4,
+    "data_scale": 1,
+    "optimizer": "adam",
+    "learning_rate": 5e-05,
+    "min_lr_scale": 0.01,
+    "weight_decay": 0,
+    "warmup_epochs": 0.25,
+    "warmup_steps": 0,
+    "warmup_ratio": null,
+    "use_hand_rgb": false,
+    "use_time_causal_attn": false,
+    "use_mim_obs_loss": false,
+    "use_pixel_loss": true,
+    "use_obs_queries": true,
+    "use_vision_resampler": false,
+    "vision_masked_ratio": 0.9,
+    "use_tube_mask": false,
+    "cache_root": "/dev/shm/vlm4vla/cache/qwen25vl",
+    "model_load_path": null,
+    "model_load_source": "torch",
+    "resume": null,
+    "model_path": "/workspace/models/Qwen2.5-VL-3B-Instruct",
+    "model_config": "/workspace/models/Qwen2.5-VL-3B-Instruct/config.json",
+    "train_setup": {
+        "precision": "bf16",
+        "predict_action": true,
+        "predict_forward": false,
+        "predict_forward_hand": false,
+        "predict_caption": false,
+        "train_vision": false,
+        "bits": -1,
+        "freeze_mm_mlp_adapter": false,
+        "freeze_backbone": false,
+        "freeze_resampler": false,
+        "tune_mm_mlp_adapter": false,
+        "mm_use_im_start_end": false,
+        "mm_use_im_patch_token": false,
+        "gradient_checkpointing": false,
+        "lora_enable": false,
+        "mm_projector_lr": 0.0001,
+        "lora_r": 64,
+        "lora_alpha": 16,
+        "lora_dropout": 0.05,
+        "lora_bias": "none",
+        "train_text_embedding": true
+    },
+    "vision_resampler": {
+        "vis_dim": 1024,
+        "depth": 8,
+        "dim_head": 64,
+        "heads": 8,
+        "num_latents": 64
+    },
+    "act_encoder": null,
+    "act_head": {
+        "type": "FCDecoder",
+        "hidden_size": 1024,
+        "action_dim": 7,
+        "down_sample": "none",
+        "latent": 1,
+        "fwd_pred_next_n": 1,
+        "window_size": 1,
+        "action_space": "continuous",
+        "with_history": true,
+        "history_type": "post"
+    },
+    "fwd_head": null,
+    "tokenizer": {
+        "type": "AutoProcessor",
+        "pretrained_model_name_or_path": "/workspace/models/Qwen2.5-VL-3B-Instruct",
+        "tokenizer_type": "qwen25vl",
+        "additional_special_tokens": null
+    },
+    "vlm": {
+        "type": "Qwen2_5_VLForConditionalGeneration",
+        "pretrained_model_name_or_path": "/workspace/models/Qwen2.5-VL-3B-Instruct",
+        "name": "qwen25vl"
+    },
+    "trainer": {
+        "accelerator": "gpu",
+        "strategy": "deepspeed_stage_2",
+        "precision": "bf16",
+        "logger": [
+            "wandb"
+        ],
+        "gradient_clip_val": 1.0,
+        "use_distributed_sampler": false,
+        "log_every_n_steps": 10,
+        "max_epochs": 5,
+        "val_check_interval": 40000,
+        "check_val_every_n_epoch": null,
+        "max_steps": 50000,
+        "accumulate_grad_batches": 8
+    },
+    "train_dataset": {
+        "type": "OpenVLADataset",
+        "data_root_dir": "/workspace/data",
+        "model_name": "qwen25vl",
+        "image_aug": true,
+        "mode": "train",
+        "data_mix": "bridge",
+        "window_sample": "sliding",
+        "organize_type": "interleave",
+        "shuffle_buffer_size": 51200,
+        "train": true
+    },
+    "val_dataset": {
+        "type": "OpenVLADataset",
+        "data_root_dir": "/workspace/data",
+        "model_name": "qwen25vl",
+        "mode": "train",
+        "data_mix": "bridge",
+        "window_sample": "sliding",
+        "organize_type": "interleave",
+        "shuffle_buffer_size": 10000,
+        "train": false
+    },
+    "norm_action": true,
+    "norm_min": -0.65,
+    "norm_max": 0.65,
+    "raw_config_path": "/workspace/VLM4VLA/configs/oxe_training/bridge/finetune_qwen25vl-3b_bridge_LOCAL_freezevis.json",
+    "num_nodes": 1,
+    "config": "/workspace/VLM4VLA/configs/oxe_training/bridge/finetune_qwen25vl-3b_bridge_LOCAL_freezevis.json",
+    "gpus": 8,
+    "log_dir": "/dev/shm/vlm4vla/logs/qwen25vl/bridge_finetune/2026-04-18/bridge_-bs512-lr5e-05-ws1-FCDecoder-latent1-freeze_vision",
+    "output_dir": "/dev/shm/vlm4vla/ckpts/qwen25vl/bridge_finetune/2026-04-18/bridge_-bs512-lr5e-05-ws1-FCDecoder-latent1-freeze_vision",
+    "data_dir": null,
+    "annotation_file": null,
+    "data_subfolder": null,
+    "task_num": null,
+    "exp_name": "21-04",
+    "use_multi_modal_emb": false,
+    "no_video_pretrained_model": false,
+    "finetune": false,
+    "llm": {
+        "type": null,
+        "n_embd": null,
+        "n_layer": null,
+        "n_head": null
+    }
+}

requirements-eval.txt ADDED Viewed

	@@ -0,0 +1,212 @@

+absl-py==2.4.0
+accelerate==1.13.0
+aiohappyeyeballs==2.6.1
+aiohttp==3.13.5
+aiosignal==1.4.0
+annotated-types==0.7.0
+antlr4-python3-runtime==4.9.3
+anyio==4.13.0
+array_record==0.8.1
+asttokens==3.0.1
+astunparse==1.6.3
+async-timeout==5.0.1
+attrs==26.1.0
+av==17.0.1
+beautifulsoup4==4.14.3
+bitsandbytes==0.49.2
+certifi==2026.2.25
+cffi==2.0.0
+charset-normalizer==3.4.7
+click==8.3.2
+cloudpickle==3.1.2
+colorama==0.4.6
+colorlog==6.10.1
+contourpy==1.3.2
+cryptography==46.0.7
+cycler==0.12.1
+datasets==2.12.0
+decorator==5.2.1
+decord==0.6.0
+deepspeed==0.18.9
+diffusers==0.37.1
+dill==0.3.6
+-e git+https://github.com/moojink/dlimp_openvla.git@040105d256bd28866cc6620621a3d5f7b6b91b46#egg=dlimp
+dm-tree==0.1.10
+draccus==0.8.0
+einops==0.8.2
+einops-exts==0.0.4
+etils==1.13.0
+exceptiongroup==1.3.1
+executing==2.2.1
+Farama-Notifications==0.0.4
+filelock==3.29.0
+flamingo-pytorch==0.1.2
+flash_attn @ https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp310-cp310-linux_x86_64.whl#sha256=ffe17686fa1a0f288de9eae7c32af209d32a27b037ef28614f042b377af5b15a
+flatbuffers==25.12.19
+fonttools==4.62.1
+frozenlist==1.8.0
+fsspec==2026.3.0
+ftfy==6.3.1
+gast==0.7.0
+gdown==6.0.0
+gitdb==4.0.12
+GitPython==3.1.46
+google-auth==2.49.2
+google-auth-oauthlib==1.3.1
+google-pasta==0.2.0
+grpcio==1.80.0
+gymnasium==0.29.1
+h11==0.16.0
+h5py==3.16.0
+hf-xet==1.4.3
+hjson==3.1.0
+httpcore==1.0.9
+httpx==0.28.1
+huggingface_hub==0.36.2
+hydra-colorlog==1.2.0
+hydra-core==1.3.2
+idna==3.12
+ImageIO==2.37.3
+imageio-ffmpeg==0.6.0
+importlib_metadata==9.0.0
+importlib_resources==7.1.0
+ipython==8.39.0
+jedi==0.19.2
+Jinja2==3.1.6
+joblib==1.5.3
+json-numpy==2.1.1
+jsonlines==4.0.0
+keras==2.15.0
+kiwisolver==1.5.0
+libclang==18.1.1
+lightning==2.6.1
+lightning-lite==1.8.6
+lightning-utilities==0.15.3
+-e git+https://github.com/simpler-env/ManiSkill2_real2sim@ef7a4d4fdf4b69f2c2154db5b15b9ac8dfe10682#egg=mani_skill2_real2sim&subdirectory=../../ManiSkill2_real2sim
+Markdown==3.10.2
+markdown-it-py==4.0.0
+MarkupSafe==3.0.3
+matplotlib==3.10.8
+matplotlib-inline==0.2.1
+mdurl==0.1.2
+mediapy==1.2.6
+mergedeep==1.3.4
+ml-dtypes==0.2.0
+mpmath==1.3.0
+msgpack==1.1.2
+multidict==6.7.1
+multiprocess==0.70.14
+mypy_extensions==1.1.0
+networkx==3.4.2
+ninja==1.13.0
+nltk==3.9.4
+numpy==1.24.4
+nvidia-cublas-cu12==12.4.5.8
+nvidia-cuda-cupti-cu12==12.4.127
+nvidia-cuda-nvrtc-cu12==12.4.127
+nvidia-cuda-runtime-cu12==12.4.127
+nvidia-cudnn-cu12==9.1.0.70
+nvidia-cufft-cu12==11.2.1.3
+nvidia-curand-cu12==10.3.5.147
+nvidia-cusolver-cu12==11.6.1.9
+nvidia-cusparse-cu12==12.3.1.170
+nvidia-cusparselt-cu12==0.6.2
+nvidia-nccl-cu12==2.21.5
+nvidia-nvjitlink-cu12==12.4.127
+nvidia-nvtx-cu12==12.4.127
+oauthlib==3.3.1
+omegaconf==2.3.0
+open-clip-torch==2.20.0
+opencv-python==4.10.0.84
+OpenEXR==3.4.10
+-e git+https://github.com/yunfeixie233/VLM4VLA.git@b4ddb404e6bce2e116b04c598b9495a99bf40fdc#egg=openvla&subdirectory=openvla
+opt_einsum==3.4.0
+packaging @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_packaging_1776209387/work
+pandas==2.3.3
+parso==0.8.6
+pexpect==4.9.0
+pillow==12.2.0
+platformdirs==4.9.6
+pretty-errors==1.2.25
+promise==2.3
+prompt_toolkit==3.0.52
+propcache==0.4.1
+protobuf==4.25.9
+psutil==7.2.2
+ptyprocess==0.7.0
+pure_eval==0.2.3
+py-cpuinfo==9.0.0
+pyarrow==24.0.0
+pyasn1==0.6.3
+pyasn1_modules==0.4.2
+pycparser==3.0
+pydantic==2.13.3
+pydantic_core==2.46.3
+Pygments==2.20.0
+pyparsing==3.3.2
+PySocks==1.7.1
+python-dateutil==2.9.0.post0
+pytorch-lightning==2.6.1
+pytz==2026.1.post1
+PyYAML==6.0.3
+pyyaml-include==1.4.1
+qwen-vl-utils==0.0.14
+regex==2026.4.4
+requests==2.33.1
+requests-oauthlib==2.0.0
+responses==0.18.0
+rich==15.0.0
+rtree==1.4.1
+ruckig==0.17.3
+safetensors==0.7.0
+sapien==2.2.2
+scikit-learn==1.7.2
+scipy==1.15.3
+sentence-transformers==2.2.2
+sentencepiece==0.1.99
+sentry-sdk==2.58.0
+-e git+https://github.com/simpler-env/SimplerEnv@06accaca93535902d408da4855f21cece12bceb7#egg=simpler_env
+six==1.17.0
+smmap==5.0.3
+soupsieve==2.8.3
+stack-data==0.6.3
+sympy==1.13.1
+tabulate==0.10.0
+tensorboard==2.15.2
+tensorboard-data-server==0.7.2
+tensorboardX==2.6.5
+tensorflow==2.15.0
+tensorflow-addons==0.23.0
+tensorflow-datasets==4.9.3
+tensorflow-estimator==2.15.0
+tensorflow-graphics==2021.12.3
+tensorflow-io-gcs-filesystem==0.37.1
+tensorflow-metadata==1.17.3
+termcolor==3.3.0
+threadpoolctl==3.6.0
+timm==1.0.26
+tokenizers==0.22.2
+toml==0.10.2
+torch==2.6.0
+torchmetrics==1.9.0
+torchvision==0.21.0
+tqdm==4.67.3
+traitlets==5.14.3
+transformers==4.57.0
+transforms3d==0.4.2
+trimesh==4.11.5
+triton==3.2.0
+typeguard==2.13.3
+typing-inspect==0.9.0
+typing-inspection==0.4.2
+typing_extensions==4.15.0
+tzdata==2026.1
+urllib3==2.6.3
+-e git+https://github.com/yunfeixie233/VLM4VLA.git@b4ddb404e6bce2e116b04c598b9495a99bf40fdc#egg=vlm4vla
+wandb==0.25.0
+wcwidth==0.6.0
+Werkzeug==3.1.8
+wrapt==1.14.2
+xxhash==3.6.0
+yarl==1.23.0
+zipp==3.23.1

stepstep=0030000.fp32.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:594f4440cc6d83069eb5a0a0f65711e1671cce064e00ddbc11fde9fc670697b0
+size 15045187610