Instructions to use AlfredJames/jobbert-zh-1m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AlfredJames/jobbert-zh-1m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="AlfredJames/jobbert-zh-1m")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("AlfredJames/jobbert-zh-1m") model = AutoModelForMaskedLM.from_pretrained("AlfredJames/jobbert-zh-1m", device_map="auto") - Notebooks
- Google Colab
- Kaggle
JobBERT-zh 1M (DAPT contrast)
This repository provides the 1M domain-adaptive pretraining contrast of JobBERT-zh. It preserves the smaller-corpus run used in the historical comparison with the 3M initialization.
| Checkpoint | Hub | Test | Typed exact F1 | Typed relaxed F1 |
|---|---|---|---|---|
| 3M + V4 CRF (historical comparison) | AlfredJames/jobbert-zh |
V4 hybrid 2601 + jieba | 0.433118 | 0.587322 |
| 1M + V4 CRF (this repo) | AlfredJames/jobbert-zh-1m |
V4 hybrid 2601 + jieba | 0.427162 | 0.595170 |
The values from tables/hybrid_cws_simhuman980_all_models.csv compare the 1M and 3M runs under the same historical hybrid-reference/jieba protocol: 0.4272 and 0.4331 typed exact F1, respectively. The later human-reference continuation result (0.5536±0.0054) and the archived Gold v2 ChatGPT result (0.6365) use separate evaluation conditions and should be interpreted within their own protocols.
- 3M initialization: https://huggingface.co/AlfredJames/jobbert-zh
- Silver fine-tuned model: https://huggingface.co/AlfredJames/jobbert-zh-v6a
- Code and data: https://github.com/AlfredJamesLi/chinese-skillspan-benchmark
- Archive (
v0.1.3): https://doi.org/10.5281/zenodo.22698504
This Chinese JobBERT implementation uses a domain-adapted encoder and custom CRF. Extraction requires the project's BertCRF class and local inference code. Other models sharing the JobBERT name include English jjzha/jobbert-base-cased and TechWolf JobBERT-v3, each with its own architecture and protocol.
What this repository contains
The repository contains an MLM encoder (config.json, model.safetensors, tokenizer files) and the historical V4 silver CRF checkpoint (crf/best.pt). The 1M and 3M encoders come from separate domain-adaptive pretraining runs.
The historical 3M result of 0.4331 corresponds to AlfredJames/jobbert-zh; this repository supplies the 1M run shown in the table above.
Intended uses
- Ablation / contrast against the 3M JobBERT-zh encoder on Chinese-SkillSpan
- Research on domain-adaptive pre-training scale for Chinese job-ad spans
Evaluation scope
The historical V4 hybrid reference combines annotation sources. The current manuscript's human-reference comparisons use a separate character-span protocol, documented for the Silver fine-tuned model linked above. Hosted Inference Providers and the default token-classification widget are unavailable for this custom CRF package. Applicant screening, hiring automation and ESCO concept-ID prediction remain outside the intended use.
Loading
Use the same BertCRF loading pattern as the 3M initialization (scripts/train_cn_roberta_crf.py). To reproduce the historical hybrid-reference results above, apply jieba alignment after tag decoding and run scorer/score_lskt.py --align-mode official. The evaluation entry guide distinguishes this historical protocol from current human-reference scoring.
import torch
from torch import nn
from torchcrf import CRF
from huggingface_hub import hf_hub_download
from transformers import AutoModel, AutoTokenizer
REPO = "AlfredJames/jobbert-zh-1m"
class BertCRF(nn.Module):
def __init__(self, model_dir: str, n_labels: int = 9, dropout: float = 0.1):
super().__init__()
self.encoder = AutoModel.from_pretrained(model_dir)
self.dropout = nn.Dropout(dropout)
self.emissions = nn.Linear(self.encoder.config.hidden_size, n_labels)
self.crf = CRF(n_labels, batch_first=True)
tok = AutoTokenizer.from_pretrained(REPO)
model = BertCRF(REPO)
crf_path = hf_hub_download(REPO, "crf/best.pt")
model.load_state_dict(torch.load(crf_path, map_location="cpu"))
Licence
This checkpoint retains license: other while job-advertisement text rights remain unresolved. The backbone hfl/chinese-roberta-wwm-ext is listed as Apache-2.0 on Hugging Face; reuse of this adapted checkpoint is governed by its own accompanying notices.
Funding: National Social Science Fund of China, Grant No. 21BGL142. Authors: same as JobBERT-zh (corresponding author Xiangyu Zhao, xianzhao@cityu.edu.hk).
- Downloads last month
- 49
Model tree for AlfredJames/jobbert-zh-1m
Base model
hfl/chinese-roberta-wwm-ext