Instructions to use karma689/mmbert-base-layer-detection-v1.2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use karma689/mmbert-base-layer-detection-v1.2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="karma689/mmbert-base-layer-detection-v1.2")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("karma689/mmbert-base-layer-detection-v1.2") model = AutoModelForTokenClassification.from_pretrained("karma689/mmbert-base-layer-detection-v1.2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| language: | |
| - bo | |
| library_name: transformers | |
| pipeline_tag: token-classification | |
| tags: | |
| - tibetan | |
| - token-classification | |
| - layer-detection | |
| base_model: jhu-clsp/mmBERT-base | |
| datasets: | |
| - karma689/layer_detection | |
| # mmbert_base_v1.2 | |
| Tibetan annotation-layer token classifier fine-tuned from | |
| [`jhu-clsp/mmBERT-base`](https://huggingface.co/jhu-clsp/mmBERT-base). | |
| ## Validation metrics | |
| The values below are from the **best** evaluation record. | |
| The complete evaluation history and original Trainer state are included in | |
| `training/metrics.json` and `training/trainer_state.json`. | |
| | Metric | Value | | |
| |---|---:| | |
| | `epoch` | `8.0` | | |
| | `eval_f1` | `0.078113428181641` | | |
| | `eval_f1_author` | `0.6329113924050632` | | |
| | `eval_f1_chapter` | `0.5358490566037736` | | |
| | `eval_f1_quote` | `0.022925944252020453` | | |
| | `eval_f1_sabche` | `0.19551934826883913` | | |
| | `eval_f1_tsawa` | `0.03828125` | | |
| | `eval_f1_yigchung` | `0.02306079664570231` | | |
| | `eval_loss` | `0.15565519034862518` | | |
| | `eval_precision` | `0.052703555649063746` | | |
| | `eval_recall` | `0.15083546590395905` | | |
| | `eval_runtime` | `165.0116` | | |
| | `eval_samples_per_second` | `8.581` | | |
| | `eval_soft_f1_tol0` | `0.078113428181641` | | |
| | `eval_soft_f1_tol1` | `0.09284739816799845` | | |
| | `eval_soft_precision_tol0` | `0.052703555649063746` | | |
| | `eval_soft_precision_tol1` | `0.06264464548706081` | | |
| | `eval_soft_recall_tol0` | `0.15083546590395905` | | |
| | `eval_soft_recall_tol1` | `0.17928646695769984` | | |
| | `eval_steps_per_second` | `4.291` | | |
| | `step` | `10832` | | |
| ## Training parameters | |
| Every value from the experiment's `config.yaml` is listed below. The original | |
| file is also included as `training/config.yaml`. | |
| | Parameter | Value | | |
| |---|---| | |
| | `experiment_name` | `mmbert_base_v1.2` | | |
| | `dataset_repo` | `karma689/layer_detection` | | |
| | `dataset_revision` | `null` | | |
| | `scope_policy` | `confirmed` | | |
| | `local_fallback.repos_dir` | `../../Data/repos` | | |
| | `local_fallback.split_csv` | `../../Data/audit/split.csv` | | |
| | `checkpoint` | `jhu-clsp/mmBERT-base` | | |
| | `fallback_checkpoint` | `KoichiYasuoka/bert-base-tibetan` | | |
| | `attn_implementation` | `flash_attention_2` | | |
| | `attn_fallback_implementation` | `sdpa` | | |
| | `max_length` | `8192` | | |
| | `stride` | `5120` | | |
| | `pad_to_multiple_of` | `8` | | |
| | `window_cache_dir` | `window_cache` | | |
| | `tokenize_num_workers` | `8` | | |
| | `masked_layers` | `["BookTitle"]` | | |
| | `learning_rate` | `1e-05` | | |
| | `train_batch_size` | `2` | | |
| | `eval_batch_size` | `2` | | |
| | `gradient_accumulation_steps` | `4` | | |
| | `num_train_epochs` | `10` | | |
| | `weight_decay` | `0.01` | | |
| | `warmup_ratio` | `0.2` | | |
| | `max_grad_norm` | `0.3` | | |
| | `class_weighting` | `True` | | |
| | `class_weight_cap` | `5.0` | | |
| | `o_class_weight` | `0.07` | | |
| | `fp16` | `False` | | |
| | `bf16` | `True` | | |
| | `gradient_checkpointing` | `False` | | |
| | `dataloader_num_workers` | `4` | | |
| | `dataloader_pin_memory` | `True` | | |
| | `tf32` | `True` | | |
| | `seed` | `42` | | |
| | `logging_steps` | `20` | | |
| | `save_total_limit` | `3` | | |
| | `resume_from_checkpoint` | `null` | | |
| | `eval_strategy` | `epoch` | | |
| | `save_strategy` | `epoch` | | |
| | `load_best_model_at_end` | `True` | | |
| | `metric_for_best_model` | `soft_f1_tol1` | | |
| | `greater_is_better` | `True` | | |
| | `early_stopping_patience` | `3` | | |
| | `report_to` | `["wandb"]` | | |
| | `wandb.project` | `layer-detection` | | |
| | `wandb.mode` | `online` | | |
| | `output_dir` | `outputs` | | |
| ## Labels | |
| | ID | BIO label | | |
| |---:|---| | |
| | 0 | `O` | | |
| | 1 | `B-QUOTE` | | |
| | 2 | `I-QUOTE` | | |
| | 3 | `B-SABCHE` | | |
| | 4 | `I-SABCHE` | | |
| | 5 | `B-TSAWA` | | |
| | 6 | `I-TSAWA` | | |
| | 7 | `B-YIGCHUNG` | | |
| | 8 | `I-YIGCHUNG` | | |
| | 9 | `B-CHAPTER` | | |
| | 10 | `I-CHAPTER` | | |
| | 11 | `B-AUTHOR` | | |
| | 12 | `I-AUTHOR` | | |
| | 13 | `B-BOOKTITLE` | | |
| | 14 | `I-BOOKTITLE` | | |
| ## Usage | |
| ```python | |
| from transformers import AutoModelForTokenClassification, AutoTokenizer | |
| model_id = "karma689/mmbert-base-layer-detection-v1.2" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForTokenClassification.from_pretrained(model_id) | |
| ``` | |
| Long texts should use the same window settings shown in the training parameters. | |