Instructions to use masahiroid/bert-base-ner-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use masahiroid/bert-base-ner-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir bert-base-ner-mlx masahiroid/bert-base-ner-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
bert-base-ner-mlx
Model Summary
This is an unofficial MLX conversion of dslim/bert-base-NER (a standard BERT-base with a named-entity-recognition head on top, trained on CoNLL-2003: PER/ORG/LOC/MISC). All credit for the original model goes to its author.
This cannot be loaded with mlx-embeddings
mlx-embeddings targets pooled embedding/reranker outputs and doesn't
support a per-token classification head. This model was reimplemented from
scratch for MLX and requires the bundled bert_ner_mlx.py. The
architecture itself is simple: a standard post-norm BERT encoder plus a
linear classification head.
Usage
import mlx.core as mx
from mlx.utils import tree_unflatten
from transformers import AutoTokenizer
from bert_ner_mlx import BertNerMLX
tokenizer = AutoTokenizer.from_pretrained("dslim/bert-base-NER")
model = BertNerMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())
text = "My name is Wolfgang and I live in Berlin, working at Hugging Face."
inputs = tokenizer(text, return_tensors="np")
input_ids = mx.array(inputs["input_ids"])
attention_mask = mx.array(inputs["attention_mask"])
logits = model(input_ids, attention_mask=attention_mask) # (1, seq_len, 9)
labels = logits.argmax(-1)[0].tolist()
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
for tok, label_id in zip(tokens, labels):
print(tok, model.config.id2label[label_id] if hasattr(model, "config") else label_id)
Labels are the same as the original model: {0: "O", 1: "B-MISC", 2: "I-MISC", 3: "B-PER", 4: "I-PER", 5: "B-ORG", 6: "I-ORG", 7: "B-LOC", 8: "I-LOC"}.
Accuracy
Compared against the PyTorch fp32 reference on a real sentence containing person/location/organization entities:
| Precision | Logits cosine sim. | Label agreement |
|---|---|---|
| MLX fp32 | 1.0 | 100% |
| MLX fp16 (this release) | 0.99999994 | 100% |
During fp16 conversion, using an attention-mask value like -1e9 (which
doesn't fit in fp16) caused 0 * (-inf) = NaN at unmasked (mask=1)
positions. Fixed by switching to -1e4, which fits fp16's range (full
details are in the conversion-toolkit repo, not SECURITY.md).
Specs
| Item | Value |
|---|---|
| Base model | dslim/bert-base-NER (BERT-base, 108M params) |
| Precision | float16 |
| Framework | MLX (from-scratch bert_ner_mlx.py) |
Notes
- This is a community conversion, not an official release from the original author.
- Security audit uses model-audit-lite
(see
SECURITY.mdfor details).
Security
Audited against its upstream with model-audit-lite: weight format, bundled code, and a machine-readable lineage (ML-BOM). Details, checksums and how to reproduce: SECURITY.md.
モデルの概要
dslim/bert-base-NER(標準的なBERT-baseに 固有表現抽出ヘッドを乗せたモデル。CoNLL-2003、PER/ORG/LOC/MISCの4種)の MLX版です。元モデルの著作権はその作者に帰属します。
mlx-embeddingsでは読み込めません
mlx-embeddingsはembedding/reranker用のpooling出力をターゲットにしており、トークン単位の
分類ヘッド(各トークンに対して独立にラベルを出す構成)には対応していないため、MLXでの
実装をゼロから書き起こして変換しています。同梱のbert_ner_mlx.pyが必要です。構造自体は
標準的なpost-norm BERTエンコーダー + 線形分類ヘッドとシンプルです。
使い方
import mlx.core as mx
from mlx.utils import tree_unflatten
from transformers import AutoTokenizer
from bert_ner_mlx import BertNerMLX
tokenizer = AutoTokenizer.from_pretrained("dslim/bert-base-NER")
model = BertNerMLX()
weights = mx.load("model.safetensors")
model.update(tree_unflatten(list(weights.items())))
mx.eval(model.parameters())
text = "My name is Wolfgang and I live in Berlin, working at Hugging Face."
inputs = tokenizer(text, return_tensors="np")
input_ids = mx.array(inputs["input_ids"])
attention_mask = mx.array(inputs["attention_mask"])
logits = model(input_ids, attention_mask=attention_mask) # (1, seq_len, 9)
labels = logits.argmax(-1)[0].tolist()
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
for tok, label_id in zip(tokens, labels):
print(tok, model.config.id2label[label_id] if hasattr(model, "config") else label_id)
ラベルは元モデルと同じ: {0: "O", 1: "B-MISC", 2: "I-MISC", 3: "B-PER", 4: "I-PER", 5: "B-ORG", 6: "I-ORG", 7: "B-LOC", 8: "I-LOC"}。
精度検証
PyTorch fp32リファレンスと、実際の文章1件(人名・地名・組織名を含む)で比較:
| 精度 | Logitsコサイン類似度 | ラベル一致率 |
|---|---|---|
| MLX fp32 | 1.0 | 100% |
| MLX fp16(本リリース) | 0.99999994 | 100% |
fp16変換の際、attention maskのマスク値に-1e9のようなfp16で表現できない大きな負数を使うと、
マスクされていない(mask=1)位置で0 * (-inf)が発生してNaNになる罠があった。
fp16の範囲内に収まる-1e4に変更して解消している(詳細はSECURITY.mdではなく、変換ノウハウ
リポジトリに記載)。
Specs
| Item | Value |
|---|---|
| ベースモデル | dslim/bert-base-NER(BERT-base、108M params) |
| 精度 | float16 |
| フレームワーク | MLX(ゼロから実装したbert_ner_mlx.py) |
備考
- 本変換は非公式のコミュニティ版です。
- セキュリティー監査にはmodel-audit-liteを
使用しています(詳細は
SECURITY.md)。
セキュリティー
model-audit-lite で変換元と突き合わせて監査済みです(重みの形式、同梱コード、機械可読な系譜=ML-BOM)。詳細・チェックサム・再現方法は SECURITY.md をご覧ください。
Quantized
Model tree for masahiroid/bert-base-ner-mlx
Base model
dslim/bert-base-NER