GLINER-INT8 / USAGE.md
mircq's picture
GLINER quantized INT8
29cd9ec verified
|
Raw
History Blame Contribute Delete
941 Bytes

gliner2-multi-v1 — INT8 ONNX (CPU)

Dynamically-quantized INT8 ONNX export of fastino/gliner2-multi-v1. This bundle contains only the INT8 graphs — load it with precision="int8".

from gliner2_onnx import GLiNER2ONNXRuntime

rt = GLiNER2ONNXRuntime(
    "gliner2-multi-v1-int8",          # this folder
    precision="int8",                  # required: no fp32 graphs are included
    providers=["CPUExecutionProvider"],
)

rt.extract_entities("pagamento polizza tfr A4983AS", ["amount", "reference_number"])
rt.classify("acquisto gasolio automezzi", ["carburanti", "polizze"], multi_label=True)

For best CPU throughput set intra-op threads to your physical core count via ORT SessionOptions.

Contents

  • gliner2_config.json — model config, references the INT8 graphs only
  • config.json, tokenizer.json, tokenizer_config.json
  • onnx/*_int8.onnx — encoder / classifier / span_rep / count_embed