--- language: eu library_name: transformers license: apache-2.0 base_model: ixa-ehu/berteus-base-cased tags: - onnx - bert - basque - euskara - feature-extraction - int8 - int4 - quantized inference: false --- # BERTeus ONNX (int8 + int4 quantized) > ⚠️ **This is an unofficial community conversion.** This is not the original BERTeus repository. > The original model was created by **Rodrigo Agerri, Iñaki San Vicente, Jon Ander Campos, Ander Barrena, Xabier Saralegi, Aitor Soroa, and Eneko Agirre** at the [IXA NLP Group](https://ixa.eus/) (University of the Basque Country, UPV/EHU). > > 🔗 **Original model:** [`ixa-ehu/berteus-base-cased`](https://huggingface.co/ixa-ehu/berteus-base-cased) > > 📄 **Paper:** [Give your Text Representation Models some Love: the Case for Basque](https://arxiv.org/abs/2004.00033) (LREC 2020) --- ## What is this? This repository provides **ONNX-quantized** versions of [BERTeus](https://huggingface.co/ixa-ehu/berteus-base-cased), a monolingual BERT model pre-trained on Basque (Euskara). Two quantization levels are available: - **int8** (`model_quantized.onnx`, 119 MB) — dynamic int8 quantization - **int4** (`model_q4.onnx`, 85 MB) — weight-only int4 (MatMul) + int8 (embeddings) Both enable efficient inference in: - **JavaScript / TypeScript** via [`@huggingface/transformers`](https://github.com/huggingface/transformers.js) (Transformers.js) — runs in the browser via WebAssembly - **Python** via [`optimum`](https://github.com/huggingface/optimum) with ONNX Runtime - Any runtime that supports the ONNX format ### Why use this instead of the original? | | Original (`ixa-ehu/berteus-base-cased`) | This repo (`itzune/berteus-onnx`) | |---|---|---| | Format | PyTorch / Safetensors | **ONNX (int8 or int4 quantized)** | | Size | ~470 MB | **85 MB** (int4) or 119 MB (int8) + 74 MB (embeddings) | | Python (PyTorch) | ✅ | — | | Python (ONNX Runtime) | — | ✅ | | Browser (WASM) | ❌ | ✅ via Transformers.js | | Speed (CPU) | Baseline | **~2–4× faster** (quantized + ONNX Runtime) | The original PyTorch model remains the best choice for training, fine-tuning, or GPU inference. Use this ONNX version when you need **browser deployment, small download size, or fast CPU inference**. ### int4 vs int8 — which should I use? | | int4 (`model_q4.onnx`) | int8 (`model_quantized.onnx`) | |---|---|---| | File size | **85 MB** | 119 MB | | Transformers.js `dtype` | `'q4'` | `'q8'` | | Quality (benchmark) | **+110 net** (85.5% accuracy) | +105 net (85.0% accuracy) | | Ranking agreement vs fp32 | 96.7% | 96.7% | | Mean score diff vs fp32 | 0.011 | 0.013 | | Browser load time | ~2.8s | ~3.1s | **Recommendation:** Use **int4** (`dtype: 'q4'`) for browser deployment — it is smaller, faster to load, and produces equivalent or better quality on the re-ranking benchmark. Use int8 if you need maximum backward compatibility with older Transformers.js versions. --- ## Model details | Property | Value | |---|---| | Architecture | BERT (encoder only, `BertModel`) | | Parameters | ~110M (encoder) + ~39M (embeddings) | | Hidden size | 768 | | Layers | 12 | | Attention heads | 12 | | Vocab size | 50,099 | | Max sequence length | 512 | | Training corpus | 224.6M tokens (Basque news + Wikipedia) | | Quantization | int4 (weight-only, MatMul) + int8 (embeddings), or int8 (dynamic) | **Note:** The original BERTeus checkpoint was saved as `BertModel` (encoder only) — it does **not** include the masked LM (MLM) prediction head. This means `BertForMaskedLM` will produce random MLM outputs. For masked LM tasks, use the static word embeddings approach described below, or fine-tune an MLM head on the original model. ### Quantization details **int4** (`model_q4.onnx`): Two-step quantization using ONNX Runtime's `MatMulNBitsQuantizer`: 1. Encoder MatMul ops → int4 (weight-only, block_size=128, asymmetric) via `MatMulNBits` operator 2. Embedding Gather ops → int8 (QDQ: `DequantizeLinear`) via dynamic quantization This approach keeps the embedding table in int8 (not int4) to avoid `GatherBlockQuantized`, which is not implemented on the WASM backend. The `MatMulNBits` operator IS supported on WASM in Transformers.js v3+. **int8** (`model_quantized.onnx`): Standard dynamic int8 quantization using `MatMulInteger` + `DynamicQuantizeLinear` operators. --- ## Files | File | Size | Description | |---|---|---| | `onnx/model_q4.onnx` | 85 MB | **int4 quantized** ONNX model (recommended for browser) | | `onnx/model_quantized.onnx` | 119 MB | int8 quantized ONNX model | | `word_embeddings_f16.bin` | 74 MB | Word embedding matrix (50,099 × 768) in float16 | | `config.json` | 578 B | Model configuration | | `tokenizer.json` | 1.2 MB | Fast tokenizer (Transformers.js compatible) | | `tokenizer_config.json` | 1.3 KB | Tokenizer configuration | | `special_tokens_map.json` | 125 B | Special tokens mapping | | `vocab.txt` | 413 KB | WordPiece vocabulary | | `embedding_meta.json` | 187 B | Embedding matrix metadata (vocab size, token IDs) | ### Why a separate embedding file? The ONNX model outputs contextual hidden states, but the **static word embedding matrix** (the input embedding layer) is useful for tasks like lexical similarity and candidate scoring. We extract it as a separate float16 binary file (74 MB) to avoid loading the full model when only embeddings are needed. --- ## Usage ### Python (ONNX Runtime via Optimum) ```python from optimum.onnxruntime import ORTModelForFeatureExtraction from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("itzune/berteus-onnx") model = ORTModelForFeatureExtraction.from_pretrained("itzune/berteus-onnx") inputs = tokenizer("Euskara munduko hizkuntzarik ederrena da.", return_tensors="pt") outputs = model(**inputs) # Contextual embeddings: [1, seq_len, 768] hidden_states = outputs.last_hidden_state ``` ### JavaScript / Browser (Transformers.js) **int4 (recommended — smaller, faster):** ```javascript import { AutoModel, AutoTokenizer, env } from '@huggingface/transformers'; const model = await AutoModel.from_pretrained('itzune/berteus-onnx', { dtype: 'q4', // int4 quantization — 85 MB device: 'wasm', }); const tokenizer = await AutoTokenizer.from_pretrained('itzune/berteus-onnx'); const inputs = tokenizer('Euskara mundiko hizkuntzarik ederrena da.', { truncation: true, max_length: 512, padding: false, }); const outputs = await model(inputs); // outputs.last_hidden_state.data → Float32Array of shape [1, seq_len, 768] ``` **int8 (alternative — slightly larger):** ```javascript const model = await AutoModel.from_pretrained('itzune/berteus-onnx', { dtype: 'q8', // int8 quantization — 119 MB device: 'wasm', }); ``` ### Using the embedding matrix (lexical similarity) ```python import numpy as np import struct # Load float16 embeddings with open("word_embeddings_f16.bin", "rb") as f: raw = np.frombuffer(f.read(), dtype=np.float16) embeddings = raw.reshape(50099, 768).astype(np.float32) # Get embedding for a token ID token_id = tokenizer.convert_tokens_to_ids("euskara") word_embedding = embeddings[token_id] # 768-dim vector ``` ```javascript // Browser: load and convert float16 → float32 const response = await fetch('word_embeddings_f16.bin'); const buffer = await response.arrayBuffer(); const uint16 = new Uint16Array(buffer); const float32 = new Float32Array(uint16.length); // Convert each float16 → float32 (bit manipulation, see source) ``` --- ## Use cases ### ✅ What this model is good for | Use case | How | |---|---| | **Contextual embeddings** | Run the encoder, extract `last_hidden_state` for any token | | **Sentence embeddings** | Mean-pool the `last_hidden_state` across tokens | | **Lexical similarity** | Compare static word embeddings (from the `.bin` file) | | **Candidate re-ranking** | Mask a position, compare `[MASK]` hidden state against candidate embeddings (cosine similarity) | | **Feature extraction** | Feed hidden states to a downstream classifier (NER, POS, sentiment) | | **Text classification** | Use `[CLS]` token output as a sentence representation | ### ❌ What this model cannot do | Limitation | Why | |---|---| | **Masked LM fill-in-the-blank** | The checkpoint has no MLM head (`cls` weights missing). Use the embedding similarity approach instead. | | **Text generation** | BERT is encoder-only, not a generative model. | | **Fine-tuning** | ONNX is an inference format. Fine-tune the [original PyTorch model](https://huggingface.co/ixa-ehu/berteus-base-cased) instead. | ### Masked embedding similarity (no MLM head needed) Since the checkpoint lacks an MLM head, the recommended approach for "which word fits here?" tasks is: 1. Replace the target position with `[MASK]` (token ID 4) 2. Run the BERT encoder on the full sentence (bidirectional context) 3. Extract the `[MASK]` position's hidden state (768-dim) 4. For each candidate word, compute its static embedding (mean of subword piece embeddings) 5. Score by cosine similarity between the `[MASK]` hidden state and the candidate embedding This is based on the lexical substitution approach from Paetzold & Specia (2017) and works well for spell-checker candidate re-ranking. --- ## Performance ### int4 vs int8 vs fp32 (933-case spell-correction re-ranking benchmark) | Model | Size | Best weight | Accuracy | Net improvement | Mean score diff vs fp32 | |---|---|---|---|---|---| | PyTorch fp32 | 496 MB | w=12 | 84.2% | +99 | — | | ONNX int8 | 119 MB | w=15 | 85.0% | +105 | 0.013 | | **ONNX int4** | **85 MB** | **w=18** | **85.5%** | **+110** | **0.011** | Both quantized variants slightly outperform fp32 PyTorch on this benchmark — quantization noise occasionally helps break ties in the re-ranking task. The int4 model is the best overall: smallest, fastest, and highest accuracy. ### Browser validation (30-case subset, Transformers.js WASM) | Variant | Model load | T2 accuracy | Match Python ranking | Mean score diff | |---|---|---|---|---| | int4 (`q4`) | 2.8s | 80.0% | 96.7% | 0.011 | | int8 (`q8`) | 3.1s | 73.3% | 96.7% | 0.013 | The float16 embedding matrix has negligible precision loss compared to float32. --- ## Citation If you use this model, please cite the **original** BERTeus paper: ```bibtex @inproceedings{agerri2020give, title = {Give your Text Representation Models some Love: the Case for Basque}, author = {Rodrigo Agerri and Iñaki San Vicente and Jon Ander Campos and Ander Barrena and Xabier Saralegi and Aitor Soroa and Eneko Agirre}, booktitle = {Proceedings of the 12th International Conference on Language Resources and Evaluation (LREC)}, year = {2020} } ``` --- ## License The original BERTeus model is released under the Apache 2.0 license. This ONNX conversion is provided under the same license. ## Acknowledgements All credit for the BERTeus model goes to the [IXA NLP Group](https://ixa.eus/) at the University of the Basque Country (UPV/EHU). This repository is merely a format conversion for deployment convenience.