JobBERT-zh Silver fine-tuned

JobBERT model fine-tuned on the Silver training/development sets (JobBERT) and evaluated on the 150-sentence human reference set.

Project · Data · Reproduction guide

Use this model

This package combines a Transformers encoder with a custom CRF. It is not a standard AutoModelForTokenClassification export. Loading the encoder alone does not reproduce span extraction. See loading and technical details for the custom model class and checkpoint paths.

The model predicts flat character spans in Chinese job advertisements using language (L), knowledge (K), occupational-skill (S), and transversal-competence (T) labels. It does not assign ESCO concept IDs or measure individual applicants' abilities.

Released checkpoints

Checkpoint Human reference set exact F1
Default / seed 42 0.5544
Seed 43 0.5479
Seed 44 0.5587
Three-seed mean ± sample SD 0.5536 ± 0.0054

The default crf/best.pt is seed 42. Other checkpoints are under crf/seed43/ and crf/seed44/. The mean is a summary of three runs, not a downloadable checkpoint. The human reference set (archive name: Gold150) and the Silver training/development sets (Qwen and encoders) are included in Zenodo v0.1.3.

The Silver training/development sets (JobBERT) contain 2,156/169 records; the revised-label condition is archived as B2 (v6a). The Silver training/development sets (Qwen and encoders), archived as v6a_nocross, contain 2,150/169 records. These lists are not interchangeable. The executed configuration does not record an encoder-freeze flag. See the model guide.

Access and scope

The model licence remains other; source-text permissions differ from the software licence. Use for research on recruitment text, with appropriate data permissions and validation. Applicant profiling and automated hiring decisions are outside the intended use.

Recruitment texts in Chinese-SkillSpan were obtained through three routes: purchased collections from MacroData (马克数据网), operated by 重庆马禾锐信息科技有限公司; research-team collection of public recruitment pages across regions of China; and the 招聘数据集 / Recruitment Dataset, ID 163746 on Alibaba Cloud Tianchi. For the public-institution subset, team members manually identified pages, registered selected URLs in a crawling framework that revisited those sources on weekly or monthly schedules, and then screened and cleaned the records. This card does not establish the source composition of the 3M pretraining corpus. See data sources and collection. Source attribution does not add redistribution rights or a promise to supply additional data by email.

Loading instructions, checkpoint information, and evaluation details are in TECHNICAL_DETAILS.md. The resource naming guide maps current names to preserved archive labels, filenames and model IDs. B1 and B2 denote archived label conditions, not separate public datasets.

Downloads last month
60
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AlfredJames/jobbert-zh-v6a

Finetuned
(1)
this model