--- tags: - visual-language-navigation - robotics - multimodal - fp8 - orcalab base_model: a8cheng/navila-llama3-8b-8f library_name: pytorch --- # Orca_VLN NaVILA FP8 E4M3 FP8 inference artifact for the [Orca_VLN](https://github.com/openverse-orca/Orca_VLN/tree/fp8) visual-language navigation baseline. The artifact quantizes 225 Linear layers in the NaVILA language model. The vision tower and multimodal projector remain FP16. It is approximately 9.0 GiB, compared with approximately 15.8 GiB for the source checkpoint. ## Download and run Clone the `fp8` branch and run its installer: ```bash git clone --branch fp8 https://github.com/openverse-orca/Orca_VLN.git cd Orca_VLN ./NaVILA-Orca/scripts/setup_all.sh ``` Verify the exact FP8 CUDA kernel and start the service: ```bash ./NaVILA-Orca/scripts/test_navila_fp8_gpu.py ./NaVILA-Orca/scripts/start_navvlm_server.sh -fp8 ``` The runtime requires the project-owned `FP8Linear` implementation and is not a drop-in Transformers checkpoint. Use the loader shipped in the `fp8` branch. ## Hardware - NVIDIA compute capability 8.9 or newer - PyTorch 2.7 with CUDA 12.8 - RTX 4090 (Ada) and RTX 5070 Ti (Blackwell) are supported by the runtime - A 16 GiB RTX 5070 Ti should be dedicated to NaVILA inference; running OrcaLab on the same GPU may exceed available VRAM ## Artifact | File | Purpose | | --- | --- | | `model-fp8.pt` | Serialized NaVILA model with E4M3 LLM Linear weights | | `fp8_manifest.json` | Format, quantized module list, size, and checksum | | `config.json` | NaVILA model configuration | `model-fp8.pt` SHA-256: ```text 78ef32bf7a586321b105157a359385afd2998ce09e91525b9bf783c51cdd6f0a ``` The artifact was derived from [`a8cheng/navila-llama3-8b-8f`](https://huggingface.co/a8cheng/navila-llama3-8b-8f) at revision `b2294e96581454468d6b94f38201f4f965ef48b7`. Usage remains subject to the upstream NaVILA model and underlying model licenses.