TripoSG, compressed ONNX for the browser
TripoSG (SIGGRAPH 2025, VAST-AI-Research) is a 1.5B rectified-flow diffusion transformer over an SDF VAE: one image in, high fidelity geometry out. This is that pipeline compressed so it runs entirely inside a web browser on WebGPU.
7.8 GB down to 3.5 GB, with the graphs checked one at a time against the originals.
Try it without downloading anything
Files
| File | Precision | Size | Runs |
|---|---|---|---|
triposg_image_encoder_q8.onnx + .data |
int8, block 32 | 345 MB | once |
triposg_vae_latents_q8.onnx + .data |
int8, block 32 | 227 MB | once |
triposg_dit_step_fp16.onnx + .data |
float16 | 2881 MB | 100 times |
triposg_vae_decoder.onnx |
float32 | 51 MB | per chunk of points |
Why the graphs are not all treated the same
Run-once graphs quantize cleanly; iterated ones do not, and this pipeline runs its transformer a hundred times per generation with classifier-free guidance. So the two single-shot graphs are int8 and the DiT is float16. Measured on the same input, against the float32 originals:
| Graph | Cosine |
|---|---|
| image encoder, int8 | 0.999899 |
| vae latents, int8 | 0.999554 |
| DiT step, float16 | 0.999993, no NaN |
The field decoder stays float32: it is small, it is evaluated at every sample point, and its sign is what decides which side of the surface a point is on.
Inference contract
image -> cut out, composited over WHITE, cropped to the subject with 10%
padding, shortest edge to 256, centre crop to 224, divided by 255
(ImageNet mean and std are inside the graph)
-> image_embeds [1,257,1024]
latents = randn(1, 2048, 64)
for i in 0..49:
sigma = 1 - i/50
sigma_next = 1 - (i+1)/50
v_uncond = dit(latents, 1000*sigma, zeros[1,257,1024])
v_cond = dit(latents, 1000*sigma, image_embeds)
latents += (sigma - sigma_next) * (v_uncond + 7.0*(v_cond - v_uncond))
kv_cache = vae_latents(latents) once
sdf = vae_decoder(kv_cache, points) chunks of <= 8192
mesh = marching_cubes(-sdf, level=0) over (-1.005 .. 1.005)
Four ways to get this wrong
- The subject must be cut out. Handed a photograph with a room behind it this model reconstructs the room: the field comes back half inside and half outside, which is a long way of saying noise. This is the failure that looks most like a broken port and is not one.
- The background is white, not the grey TripoSR wants.
- The update adds
(sigma_i - sigma_next) * v. Stock diffusers'FlowMatchEulerDiscreteScheduleruses the opposite sign, because its models predictnoise - x0where this one predictsx0 - noise. - The unconditional branch is a zero embedding, not an encoded black image. DINOv2 turns a black square into a confident, very non-zero embedding, and guidance of 7.0 applied against it produces noise.
On the field's polarity
The upstream export notes say the decoder ends in * -1 and is inside
positive; the standalone model card says it is outside positive and must be
negated. Measured on a cut out chair, 95.2% of the box comes back positive, and
an object filling 95% of its bounding box is not an object. This export is
outside positive: negate it, or take the object to be where the field is
negative.
Verification
scripts/triposg-verify.py compares each compressed graph with its original on
identical input, and separately runs the whole pipeline to a mesh, because the
two fail differently: a compression problem shows up as a cosine, and a contract
problem shows up as a field that never crosses zero or crosses it everywhere. On
a cut out chair the finished field is negative over 4.8% of its box.
Note on the upstream ONNX export
Derived from the export by
fernandotonon for
QtMeshEditor, whose
docs/TRIPOSG_EXPORT_NOTES.md is the most useful description of this model's
inference contract anywhere, and was read from the upstream sources rather than
guessed at.
Licence
MIT, from the upstream code and weights. The DINOv2-large encoder is Apache-2.0 (credit Meta AI). Both permit commercial use.
Note for anyone reproducing this: TripoSG's own reference pipeline mattes with BriaRMBG, which is non-commercial. It is not used here and not needed; any permissively licensed matting model works.
Model tree for cgb/triposg-onnx-webgpu
Base model
VAST-AI/TripoSG