zlm-v1-iab-classify-edge

A 24M-parameter MiniLM model that tags English text with IAB content and audience categories, small enough to run on the device: a 25 MB int8 ONNX file that classifies a 128-token text in 4–10 ms on a laptop CPU.

Give it an article, a page summary, or any short English text and it returns IAB Tech Lab content categories (704 labels, IAB Content Taxonomy 3.0) and audience segments (1,567 labels, IAB Audience Taxonomy 1.1) with a score for each. In a blind, 10,000-text comparison judged by GPT-5.5, its labels were preferred over GPT-5.4-nano's in 66% of decided cases, and on independent human labels it scores 0.406 F1 vs 0.375 for nano. It never returns a label outside the taxonomy.

This is the exact bundle ZeroGPU runs on edge devices (browsers, Node.js / Docker workers, Android) in production. It is laid out for transformers.js and also runs directly in onnxruntime.

Typical uses: contextual ad targeting, content labelling, and audience enrichment where a page's text is available and an LLM call per page is too slow or too expensive.

Model description

  • Backbone: sentence-transformers/all-MiniLM-L6-v2 (6 layers, hidden size 384), fine-tuned end to end. Mean pooling and L2 normalisation over the encoder output feed two heads.
  • Heads: two multi-label MLP heads, Linear(384→512) → GELU → Linear(512→N): IAB Content (704 labels) and IAB Audience (1,567 labels).
  • Parameters: 24.1M (encoder 22.6M, heads 1.6M).
  • Output: one logits tensor of shape [batch, 2271], the 704 content logits followed by the 1,567 audience logits. Labels in config.json carry a head prefix: content|Soccer, audience|Interest | Sports | Soccer |.
  • Decoding (as in production): sigmoid per label (multi-label, not softmax), keep scores ≥ 0.5, at most 6 labels per head, and display the last |-separated segment of each label (audience|Interest | Sports | Soccer | → Soccer).
  • Input window: 128 tokens. The model was trained at 512 but is served at 128; longer texts are classified from their first 128 word-pieces. Production sets no minimum device memory for it.
  • Quantization: int8 dynamic quantization (QUInt8, per-channel) of the fp32 model. It reaches the same accuracy as fp32 on the independent benchmark below (F1 0.4062 vs 0.4060) and picks the same top content label on 95.3% of its 2,784 pages.

Label space

The 704 content labels follow IAB Content Taxonomy 3.0: 701 of them are exact 3.0 category names, and they cover all 36 tier-1 categories and their tier-2 to tier-4 children. The audience labels follow IAB Audience Taxonomy 1.1 and include its Interest, Purchase Intent and Demographic branches.

Files

File Purpose
onnx/model_quantized.onnx int8 graph, 24.8 MB. Inputs input_ids, attention_mask, token_type_ids (int64); output logits
config.json id2label / label2id over the 2,271 prefixed labels, problem_type: multi_label_classification
tokenizer.json, tokenizer_config.json BERT WordPiece tokenizer (uncased, 30,522 tokens)

tokenizer.json carries a built-in truncation and fixed padding to 128 tokens. Disable the padding when you classify one text at a time: with dynamic int8 quantization, padded positions shift the activation ranges and nudge the scores (for example 0.908 unpadded vs 0.920 padded for the top label of one sports text). Production runs one text at a time without padding.

config.json describes the packaging (BertForSequenceClassification, 2,271 labels) so that runtimes can read the label map. The repository contains no PyTorch weights.

Usage

Python (onnxruntime)

import json
import numpy as np
import onnxruntime as ort
from huggingface_hub import snapshot_download
from tokenizers import Tokenizer

root = snapshot_download("ZeroGPU/zlm-v1-iab-classify-edge")
sess = ort.InferenceSession(f"{root}/onnx/model_quantized.onnx", providers=["CPUExecutionProvider"])
tok = Tokenizer.from_file(f"{root}/tokenizer.json")
tok.enable_truncation(max_length=128)  # the window the model is served at
tok.no_padding()                       # one text at a time: no padding (see Files)
id2label = json.load(open(f"{root}/config.json", encoding="utf-8"))["id2label"]

def classify(text: str, min_score: float = 0.5, max_results: int = 6) -> dict:
    enc = tok.encode(text)
    input_ids = np.array([enc.ids], dtype=np.int64)
    logits = sess.run(["logits"], {
        "input_ids": input_ids,
        "attention_mask": np.array([enc.attention_mask], dtype=np.int64),
        "token_type_ids": np.zeros_like(input_ids),
    })[0][0]                                    # (2271,) = 704 content + 1567 audience logits
    probs = 1.0 / (1.0 + np.exp(-logits))       # multi-label: independent sigmoid per label
    out = {"content": [], "audience": []}
    for i in np.argsort(-probs):
        if probs[i] < min_score:
            break
        head, name = id2label[str(i)].split("|", 1)
        name = [s.strip() for s in name.split("|") if s.strip()][-1]  # display name = last segment
        if len(out[head]) < max_results:
            out[head].append((name, round(float(probs[i]), 3)))
    return out

print(classify("Mortgage rates fell to a six-month low this week, giving first-time buyers some relief."))

Output:

{'content': [('Home Financing', 0.912), ('Housing Market', 0.746), ('Interest Rates', 0.667),
             ('Real Estate Buying and Selling', 0.661), ('Real Estate', 0.651), ('Personal Finance', 0.587)],
 'audience': [('Home Owners', 0.755), ('First Time Homeowner', 0.63), ('Personal Finance', 0.587),
              ('Real Estate Buying and Selling', 0.575), ('Household Income (USD)', 0.532), ('Real Estate', 0.522)]}

JavaScript (transformers.js — browser, Node.js, workers)

import { AutoModelForSequenceClassification, AutoTokenizer, Tensor } from '@huggingface/transformers';

const repo = 'ZeroGPU/zlm-v1-iab-classify-edge';
const tokenizer = await AutoTokenizer.from_pretrained(repo);
const model = await AutoModelForSequenceClassification.from_pretrained(repo, { dtype: 'q8' });
const { id2label } = model.config;

// Cut to the 128-token production window and keep the closing [SEP]
// (transformers.js' own `truncation` option drops that last token, which shifts the scores).
function encode(text, maxLength = 128) {
  let ids = Array.from(tokenizer(text).input_ids.data, Number);
  if (ids.length > maxLength) ids = [...ids.slice(0, maxLength - 1), ids.at(-1)];
  const tensor = (data) => new Tensor('int64', BigInt64Array.from(data, BigInt), [1, data.length]);
  return { input_ids: tensor(ids), attention_mask: tensor(ids.map(() => 1)), token_type_ids: tensor(ids.map(() => 0)) };
}

async function classify(text, minScore = 0.5, maxResults = 6) {
  const { logits } = await model(encode(text));
  const probs = Array.from(logits.data, (z) => 1 / (1 + Math.exp(-z)));   // multi-label: sigmoid per label
  const out = { content: [], audience: [] };
  for (const i of [...probs.keys()].sort((a, b) => probs[b] - probs[a])) {
    if (probs[i] < minScore) break;
    const [head, ...path] = id2label[i].split('|');
    const name = path.map((s) => s.trim()).filter(Boolean).at(-1);          // display name = last segment
    if (out[head].length < maxResults) out[head].push([name, Number(probs[i].toFixed(3))]);
  }
  return out;
}

console.log(await classify('Mortgage rates fell to a six-month low this week, giving first-time buyers some relief.'));

The same files load in transformers.js v2 (@xenova/transformers) with { quantized: true }.

Evaluation

Two benchmarks against GPT-5.4-nano prompted for the same task with the full label lists in its prompt. Neither benchmark's texts were used in training (excluded by id and by text).

Blind head-to-head, 10,000 English texts (GPT-5.5 judge)

7,216 real-world article and summary texts plus 2,784 web pages from the Figure Eight contextual dataset. For each text, GPT-5.5 saw both label sets in random A/B order and picked the better one.

Outcome Texts Share of decided
This model preferred 6,329 66.0%
GPT-5.4-nano preferred 3,263 34.0%
Tie 393
Both wrong 15

By source: article text 71.0% (4,922 vs 2,011); web pages 52.9% (1,407 vs 1,252). Across the 53 topic buckets with at least 40 texts, the model was preferred in 48 and tied in one; it trailed in Health, Games, News & Media, and Career & Education.

Output validity: the model returns 0 labels outside the taxonomy (closed label set); GPT-5.4-nano returned invalid audience labels on 1,796 of the 10,000 texts.

Independent human labels (Figure Eight, 2,784 web pages)

The Figure Eight pages carry human IAB 1.0 labels; model and nano outputs were mapped to IAB 1.0 codes and scored against them at the production decoding settings (score ≥ 0.5, at most 6 labels). Tier-2 counts a match on the exact tier-2 code, over the pages that have a tier-2 gold label; tier-1 counts a match on the tier-1 category.

System Tier-2 precision Tier-2 recall Tier-2 F1 Tier-1 F1
This model, int8 (this repository) 0.2951 0.6515 0.4062 0.3986
This model, fp32 0.2954 0.6489 0.4060 0.3975
GPT-5.4-nano 0.2672 0.6255 0.3745 0.3820

The head-to-head run was scored on the fp32 model; the table above shows the int8 file matches it.

Latency

int8 model, onnxruntime 1.30, one text, CPU (Intel Core Ultra 9 275HX), median of 50 runs:

Threads 22 tokens 128 tokens 512 tokens
1 1.7 ms 9.9 ms 62.7 ms
4 1.2 ms 4.0 ms 18.1 ms

For comparison, the GPT-5.4-nano API calls in the benchmark took a median of 1,911 ms (network round trip included).

Training

  • Fine-tuned from all-MiniLM-L6-v2 in two stages: an initial IAB fine-tune of the same architecture, then continued training (warm start) on about 170,000 English texts: ~164,000 real-world article and summary texts and 5,383 web pages from the Figure Eight contextual dataset (CC BY 4.0), sampled across the IAB tier-1 categories with extra weight on weak ones.
  • Labels: every text was labelled twice by a teacher LLM (GPT-5.4-mini) whose output was constrained to the model's 704 + 1,567 labels, and only the labels both passes agreed on were kept. The Figure Eight pages also keep their human labels. The validation split was labelled by GPT-5.5.
  • Recipe: Asymmetric Loss, 512-token sequences, AdamW (encoder LR 1e-5, heads LR 1e-4); checkpoint chosen by per-tier-1 content F1 on the validation split.
  • Exported to ONNX and quantized to int8 (dynamic, per-channel QUInt8).

The training data is not published.

Limitations and intended use

  • Absolute F1 is modest (~0.41 on human labels): IAB categories overlap and a text often fits several of them. Treat the output as ranked enrichment signals.
  • The web-page margin over GPT-5.4-nano is small (52.9% of decided cases); the larger margin on article text comes from the same kind of text the model was trained on.
  • The judge (GPT-5.5) and the teacher (GPT-5.4-mini) come from the same model family, which may favour the model's labelling style.
  • English only, uncased. Translate other languages first.
  • Only the first 128 tokens are read.
  • Near-miss sibling labels are common (for example Tennis or Sports Video Games alongside Soccer).
  • The label set is a fixed snapshot of the IAB taxonomies.
  • The audience head includes Demographic segments (age, income, household, education, occupation). Text rarely supports them; treat them as contextual hints about a page's likely audience, and do not use the model to infer attributes of individual people. The content taxonomy also contains sensitive categories; review predictions in those before acting on them.

License and attribution

Citation

@misc{zerogpu2026iabclassifyedge,
  title  = {zlm-v1-iab-classify-edge: on-device IAB content and audience classification with a fine-tuned MiniLM},
  author = {ZeroGPU},
  year   = {2026},
  url    = {https://huggingface.co/ZeroGPU/zlm-v1-iab-classify-edge}
}
Downloads last month
73
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZeroGPU/zlm-v1-iab-classify-edge

Evaluation results

  • Tier-2 F1 (int8 on-device model) on Figure Eight contextual dataset, 2,784 web pages (human IAB labels)
    self-reported
    0.406
  • Tier-2 precision on Figure Eight contextual dataset, 2,784 web pages (human IAB labels)
    self-reported
    0.295
  • Tier-2 recall on Figure Eight contextual dataset, 2,784 web pages (human IAB labels)
    self-reported
    0.651