JobBERT-zh 1M (DAPT contrast)

This repository provides the 1M domain-adaptive pretraining contrast of JobBERT-zh. It preserves the smaller-corpus run used in the historical comparison with the 3M initialization.

Checkpoint Hub Test Typed exact F1 Typed relaxed F1
3M + V4 CRF (historical comparison) AlfredJames/jobbert-zh V4 hybrid 2601 + jieba 0.433118 0.587322
1M + V4 CRF (this repo) AlfredJames/jobbert-zh-1m V4 hybrid 2601 + jieba 0.427162 0.595170

The values from tables/hybrid_cws_simhuman980_all_models.csv compare the 1M and 3M runs under the same historical hybrid-reference/jieba protocol: 0.4272 and 0.4331 typed exact F1, respectively. The later human-reference continuation result (0.5536±0.0054) and the archived Gold v2 ChatGPT result (0.6365) use separate evaluation conditions and should be interpreted within their own protocols.

This Chinese JobBERT implementation uses a domain-adapted encoder and custom CRF. Extraction requires the project's BertCRF class and local inference code. Other models sharing the JobBERT name include English jjzha/jobbert-base-cased and TechWolf JobBERT-v3, each with its own architecture and protocol.


What this repository contains

The repository contains an MLM encoder (config.json, model.safetensors, tokenizer files) and the historical V4 silver CRF checkpoint (crf/best.pt). The 1M and 3M encoders come from separate domain-adaptive pretraining runs.

The historical 3M result of 0.4331 corresponds to AlfredJames/jobbert-zh; this repository supplies the 1M run shown in the table above.


Intended uses

  • Ablation / contrast against the 3M JobBERT-zh encoder on Chinese-SkillSpan
  • Research on domain-adaptive pre-training scale for Chinese job-ad spans

Evaluation scope

The historical V4 hybrid reference combines annotation sources. The current manuscript's human-reference comparisons use a separate character-span protocol, documented for the Silver fine-tuned model linked above. Hosted Inference Providers and the default token-classification widget are unavailable for this custom CRF package. Applicant screening, hiring automation and ESCO concept-ID prediction remain outside the intended use.


Loading

Use the same BertCRF loading pattern as the 3M initialization (scripts/train_cn_roberta_crf.py). To reproduce the historical hybrid-reference results above, apply jieba alignment after tag decoding and run scorer/score_lskt.py --align-mode official. The evaluation entry guide distinguishes this historical protocol from current human-reference scoring.

import torch
from torch import nn
from torchcrf import CRF
from huggingface_hub import hf_hub_download
from transformers import AutoModel, AutoTokenizer

REPO = "AlfredJames/jobbert-zh-1m"

class BertCRF(nn.Module):
    def __init__(self, model_dir: str, n_labels: int = 9, dropout: float = 0.1):
        super().__init__()
        self.encoder = AutoModel.from_pretrained(model_dir)
        self.dropout = nn.Dropout(dropout)
        self.emissions = nn.Linear(self.encoder.config.hidden_size, n_labels)
        self.crf = CRF(n_labels, batch_first=True)

tok = AutoTokenizer.from_pretrained(REPO)
model = BertCRF(REPO)
crf_path = hf_hub_download(REPO, "crf/best.pt")
model.load_state_dict(torch.load(crf_path, map_location="cpu"))

Licence

This checkpoint retains license: other while job-advertisement text rights remain unresolved. The backbone hfl/chinese-roberta-wwm-ext is listed as Apache-2.0 on Hugging Face; reuse of this adapted checkpoint is governed by its own accompanying notices.

Funding: National Social Science Fund of China, Grant No. 21BGL142. Authors: same as JobBERT-zh (corresponding author Xiangyu Zhao, xianzhao@cityu.edu.hk).

Downloads last month
49
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AlfredJames/jobbert-zh-1m

Finetuned
(72)
this model