codemap / README.md
bogdanraduta's picture
Add OpenNER model, card, NOTICE (Apache-2.0, FlowX.AI)
4454d65 verified
|
Raw
History Blame Contribute Delete
1.71 kB
metadata
license: apache-2.0
library_name: transformers
pipeline_tag: token-classification
base_model: answerdotai/ModernBERT-base
tags:
  - ner
  - on-device
  - privacy
  - flowx
  - openner
  - insurance
  - de-identification
  - token-classification
metrics:
  - f1

CodeMap

CodeMap is a small, on-device insurance NER model from the FlowX OpenNER family. Developed by FlowX.AI. Runs 100% on-premise / air-gapped, so no data leaves your boundary.

What it does

  • Task: token-classification
  • Base model: answerdotai/ModernBERT-base
  • Entity types (5): CAT, CPT, ICD, NAIC, NCCI
  • Held-out F1: 1.0000
  • Runtime: CPU, Apple Silicon, one GPU, or browser/edge via ONNX (INT8). ~100-160 ms/doc on CPU.

Why a small model

Fine-tuned encoders match or beat frontier LLMs on structured, convention-bound extraction, at a fraction of the latency and cost, with zero data egress. Identifiers are validated by checksum (IBAN mod-97, card Luhn, ISIN/LEI, container ISO-6346, VIN, national IDs), a correctness guarantee general LLMs lack. See the FlowX OpenNER benchmark for measured results.

Usage

from transformers import AutoTokenizer, AutoModelForTokenClassification
tok = AutoTokenizer.from_pretrained("flowxai/codemap")
model = AutoModelForTokenClassification.from_pretrained("flowxai/codemap")

License & attribution

Licensed under the Apache License 2.0. Copyright 2026 FlowX.AI (https://flowx.ai). See the NOTICE file. Trained on synthetic, checksum-validated data.

Part of the FlowX OpenNER model family. Synthetic-data F1 reflects an in-distribution synthetic distribution; validate on real documents before production use.