Instructions to use AlfredJames/jobbert-zh-v6a with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AlfredJames/jobbert-zh-v6a with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="AlfredJames/jobbert-zh-v6a")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("AlfredJames/jobbert-zh-v6a") model = AutoModel.from_pretrained("AlfredJames/jobbert-zh-v6a", device_map="auto") - Notebooks
- Google Colab
- Kaggle
JobBERT-zh Silver fine-tuned
JobBERT model fine-tuned on the Silver training/development sets (JobBERT) and evaluated on the 150-sentence human reference set.
Project · Data · Reproduction guide
Use this model
This package combines a Transformers encoder with a custom CRF. It is not a standard AutoModelForTokenClassification export. Loading the encoder alone does not reproduce span extraction. See loading and technical details for the custom model class and checkpoint paths.
The model predicts flat character spans in Chinese job advertisements using language (L), knowledge (K), occupational-skill (S), and transversal-competence (T) labels. It does not assign ESCO concept IDs or measure individual applicants' abilities.
Released checkpoints
| Checkpoint | Human reference set exact F1 |
|---|---|
| Default / seed 42 | 0.5544 |
| Seed 43 | 0.5479 |
| Seed 44 | 0.5587 |
| Three-seed mean ± sample SD | 0.5536 ± 0.0054 |
The default crf/best.pt is seed 42. Other checkpoints are under crf/seed43/ and crf/seed44/. The mean is a summary of three runs, not a downloadable checkpoint. The human reference set (archive name: Gold150) and the Silver training/development sets (Qwen and encoders) are included in Zenodo v0.1.3.
The Silver training/development sets (JobBERT) contain 2,156/169 records; the revised-label condition is archived as B2 (v6a). The Silver training/development sets (Qwen and encoders), archived as v6a_nocross, contain 2,150/169 records. These lists are not interchangeable. The executed configuration does not record an encoder-freeze flag. See the model guide.
Access and scope
The model licence remains other; source-text permissions differ from the software licence. Use for research on recruitment text, with appropriate data permissions and validation. Applicant profiling and automated hiring decisions are outside the intended use.
Recruitment texts in Chinese-SkillSpan were obtained through three routes: purchased collections from MacroData (马克数据网), operated by 重庆马禾锐信息科技有限公司; research-team collection of public recruitment pages across regions of China; and the 招聘数据集 / Recruitment Dataset, ID 163746 on Alibaba Cloud Tianchi. For the public-institution subset, team members manually identified pages, registered selected URLs in a crawling framework that revisited those sources on weekly or monthly schedules, and then screened and cleaned the records. This card does not establish the source composition of the 3M pretraining corpus. See data sources and collection. Source attribution does not add redistribution rights or a promise to supply additional data by email.
Loading instructions, checkpoint information, and evaluation details are in TECHNICAL_DETAILS.md. The resource naming guide maps current names to preserved archive labels, filenames and model IDs. B1 and B2 denote archived label conditions, not separate public datasets.
- Downloads last month
- 60