Instructions to use swetlanas/physbert-tag-recommender with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use swetlanas/physbert-tag-recommender with PEFT:
from peft import PeftModel from transformers import AutoModelForSequenceClassification base_model = AutoModelForSequenceClassification.from_pretrained("thellert/physbert_uncased") model = PeftModel.from_pretrained(base_model, "swetlanas/physbert-tag-recommender") - Notebooks
- Google Colab
- Kaggle
Model Card for Model ID
PhysBERT-tag-recommender is a parameter-efficient fine-tuned (LoRA) adapter built on top of thellert/physbert_uncased. It is tailored for multi-label PhySH tag recommendation for condensed matter physics research articles, specifically relevant to APS Physical Review B (PRB).
Model Details
Model Description
- Developed by: swetlanas
- Model type: Transformer-based sequence classification with LoRA adapter (PEFT)
- Language(s) (NLP): English
- License: apache-2.0
- Finetuned from model : thellert/physbert_uncased
Model Sources
- Repository: PhySH-Tank
- Base model : thellert/physbert_uncased
- Demo : [Coming soon]
Uses
Direct Use
This model is designed to recommend Physics subject heading (PhySH) tags for APS Physical Review B articles based on their titles and abstracts.
Out-of-Scope Use
- General domain text classification (non-physics domains).
- Processing full-length papers beyond standard sequence length limits without truncation.
- Recommending tags outside of the PhySH scheme.
Bias, Risks, and Limitations
- Domain Specificity: The model is heavily biased toward condensed matter, materials science, and solid-state physics terminology. Performance will degrade on astrophysics, high-energy physics, or non-physics fields.
- Taxonomy Dependency: Label suggestions rely on the label set present during training and may omit newly introduced APS/PhySH terms past July 2026.
How to Get Started with the Model
Use the code snippet below to run inference on a title and abstract pair:
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from peft import PeftModel
MODEL_ID = "swetlanas/physbert-tag-recommender"
#Load Tokenizer & Model
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_ID)
model.eval()
#Title + Abstract Input
input_title = "Three-dimensional zigzag correlations in the van der Waals Kitaev magnet RuBr3."
input_abstract = "Ruthenium trihalides RuX3 (X=Cl,Br,I) provide a tunable platform for Kitaev magnetism..."
input_text = input_title + " " + input_abstract
inputs = tokenizer(
input_text,
padding=True,
truncation=True,
max_length=512,
return_tensors="pt"
)
#Compute Inference
threshold = 0.3 # Probability threshold for multi-label tag acceptance
with torch.no_grad():
outputs = model(**inputs)
probabilities = torch.sigmoid(outputs.logits)[0]
#Extract Predicted Tags
predicted_tags = [
model.config.id2label[i]
for i, prob in enumerate(probabilities)
if prob > threshold
]
print("Recommended Tags:", predicted_tags)
Training Details
Training Data
- Dataset: Abstract and title metadata from condensed matter physics publications (targeting APS Physical Review B categories / PhySH taxonomy). More info
- Target Format: Multi-label binary indicator vectors.
Training Procedure
The model was fine-tuned using Hugging Face's Trainer and peft library with Low-Rank Adaptation (LoRA) on top of the base architecture thellert/physbert_uncased.
Preprocessing
- Input texts were constructed by concatenating the paper Title and Abstract.
- Tokenized with a maximum sequence length of 512 tokens.
Parameter-Efficient Fine-Tuning (PEFT/LoRA) Configuration
- Task Type: Sequence Classification (
SEQ_CLS) - LoRA Rank ($r$): 8
- LoRA Alpha ($\alpha$): 16
- LoRA Dropout: 0.1
- Target Modules:
["query", "value"] - Modules to Save (Fully Fine-Tuned):
["classifier"]
trainable params: 2,156,661 || all params: 113,500,650 || trainable%: 1.9001
Training Hyperparameters
- Objective: Multi-label classification with Binary Cross-Entropy Loss (BCEWithLogitsLoss).
- Adapter Technique: Low-Rank Adaptation (LoRA / PEFT).
- Base Architecture: PhysBERT Uncased.
- Learning Rate:
3e-4 - Learning Rate Scheduler:
linear - Warmup Ratio:
0.1(10% of total steps) - Weight Decay:
0.01 - Gradient Accumulation Steps:
2 - Gradient Checkpointing: Enabled (
True) - Precision:
bf16=True(bfloat16 mixed precision) withtf32=True(TensorFloat-32) - Evaluation & Save Strategy: Evaluated and checkpointed at the end of every
epoch - Best Model Selection Metric:
precision_at_5(maximizing Precision@5) - Early Stopping: Triggered after
2epochs of non-improvingprecision_at_5(EarlyStoppingCallback(patience=2)) - Load Best Model at End: Enabled (
True)
Evaluation
Testing Data, Factors & Metrics
Testing Data
Held-out test split from the ~50,000 APS Physical Review B abstracts (2016–2026), using Hugging Face Dataset dataset.train_test_split(test_size=0.2, seed=0) for consistency across all model tracks.
Factors
Evaluation was run across the full label space (~3,000 PhySH tags) with no per-subfield breakdown; performance is expected to be stronger on condensed matter / materials science / solid-state topics, which dominate the training distribution.
Metrics
Precision@5 (P@5) was used as the primary metric and the early-stopping / best-checkpoint criterion, since the task is to recommend the top 5 most relevant tags per paper.
Recall@5, NDCG@5, and Hit@5 were also tracked via a custom compute_metrics function in the HuggingFace Trainer.
Results
| Model | Precision@5 |
|---|---|
| TF-IDF + MultinomialNB (baseline) | 0.332 |
| napkinXC (PLT) + frozen PhysBERT embeddings | 0.371 |
| LoRA fine-tuned PhysBERT (this model) | 0.408 |
Summary
LoRA fine-tuning of PhysBERT outperforms both a classical TF-IDF baseline and a tree-based extreme multi-label classifier (napkinXC) using frozen PhysBERT embeddings, while updating only 2.16M parameters (1.9% of the base model's total).
Technical Specifications
Model Architecture and Objective
Base architecture: PhysBERT thellert/physbert_uncased, a BERT-style transformer encoder pretrained on physics (arXiv) text, with a sequence classification head for multi-label PhySH tag prediction.
A LoRA (Low-Rank Adaptation) adapter (r=8, α=16, dropout=0.1) was applied to the query and value attention projection matrices, with the classification head fully fine-tuned. This updates 2.16M parameters (1.9% of the base model's total).
The training objective is multi-label binary classification: each of the ~3,000 PhySH tags is treated as an independent binary target, optimized with BCEWithLogitsLoss (Binary Cross-Entropy with logits) over sigmoid outputs, rather than a single softmax over mutually exclusive classes.
Compute Infrastructure
Hardware
RTX 3060 12GB GPU
Software
PyTorch, Hugging Face Transformers, PEFT
Citation
BibTeX:
@misc{physh-tank,
author = {Swetlana S},
title = {PhySH-Tank: PhySH Tag Recommendation for PRB Abstracts},
year = {2026},
howpublished = {\url{https://github.com/swetlanas/PhySH-Tank}}
}
Model Card Authors
- Downloads last month
- -
Model tree for swetlanas/physbert-tag-recommender
Base model
thellert/physbert_uncased