Instructions to use CoderViking/birefnet-lite-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- BiRefNet
How to use CoderViking/birefnet-lite-onnx with BiRefNet:
# Option 1: use with transformers from transformers import AutoModelForImageSegmentation birefnet = AutoModelForImageSegmentation.from_pretrained("CoderViking/birefnet-lite-onnx", trust_remote_code=True)# Option 2: use with BiRefNet # Install from https://github.com/ZhengPeng7/BiRefNet from models.birefnet import BiRefNet model = BiRefNet.from_pretrained("CoderViking/birefnet-lite-onnx") - Notebooks
- Google Colab
- Kaggle
BiRefNet_lite — browser-tuned ONNX export (static 1024², opset 17)
Custom ONNX export of BiRefNet_lite (bilateral reference network for dichotomous image segmentation, Swin-v1-tiny backbone) from the official ZhengPeng7/BiRefNet_lite weights, re-hosted for AllPrivate — where every model runs in the visitor's browser and nothing is uploaded.
Why a custom export
The community export (onnx-community/BiRefNet_lite-ONNX)
throws OrtRun std::bad_alloc in ONNX Runtime Web on every EP (fp32/fp16 ×
wasm/webgpu, tested 2026-07 on an M-series MacBook): it is a dynamic-shape
trace whose deformable convolutions decompose into GatherND/ScatterND/Clip
chains that fall back to CPU on the WebGPU EP and materialize im2col tensors of
hundreds of MB inside the 32-bit wasm heap.
This export replaces DeformableConv2d.forward with a numerically identical
per-kernel-tap GridSample decomposition (one bilinear GridSample + 1×1 Conv
per tap, accumulated — peak extra memory is one [1,C,H,W] tensor per tap) and
traces with a static input shape, so all shape dynamism constant-folds
away. It runs to completion on both the wasm and WebGPU execution providers of
ONNX Runtime Web (verified 1.26-dev).
Provenance
- Source weights:
model.safetensorsat pinned revision7838f1c- sha256:
4417d89795250e698c3cb0ae8df15743810065f646f48a694fdfa7ca052d0815
- sha256:
- Exported with
torch.onnx.export(PyTorch 2.8.0, TorchScript exporter), opset 17, fp32, constant folding on, post-processed with onnxslim 0.1.94 (constant-folds the Swin attention-mask construction; Gemm fusion disabled) plus a content-hash initializer dedupe (the backbone is traced twice for the multi-scale 'cat' input) — seeexport_birefnet_lite.pyfor the exact reproducible script - I/O:
input_image[1, 3, 1024, 1024](NCHW, RGB, float32, ImageNet normalization:(x/255 - mean) / std, mean[0.485, 0.456, 0.406], std[0.229, 0.224, 0.225]) →output_image[1, 1, 1024, 1024]logits; apply sigmoid for the [0,1] matte - Graph ops:
Add, BatchNormalization, Concat, Conv, Div, Erf, Gather, GlobalAveragePool, GridSample, LayerNormalization, MatMul, Mul, Pad, Relu, Reshape, Resize, Shape, Sigmoid, Slice, Softmax, Transpose, Unsqueeze— no GatherND, no ScatterND, no Clip - Validated against the unpatched PyTorch reference (torchvision
deform_conv2d): GridSample decomposition max abs dev3.1e-5vs torchvision; full-model ONNX (onnxruntime CPU) max abs logits dev1.0e-4, max post-sigmoid dev1.3e-9 - File sha256:
50a57872cc739192446da2a934159f957c81af8b5a161dfda8e3daa51660ca67
License
MIT, inherited from BiRefNet (© Zheng Peng et al.). This repo only re-packages the officially released weights in ONNX form.
Citation
@article{BiRefNet,
title = {Bilateral Reference for High-Resolution Dichotomous Image Segmentation},
author = {Zheng, Peng and Gao, Dehong and Fan, Deng-Ping and Liu, Li and Laaksonen, Jorma and Ouyang, Wanli and Sebe, Nicu},
journal = {CAAI Artificial Intelligence Research},
year = {2024}
}