GeistBERT
/

GeistBERT_base

Model card Files Files and versions

rjschmitt commited on Dec 5, 2025

Commit

af3f629

·

verified ·

1 Parent(s): 7db859f

Update README.md

Files changed (1) hide show

README.md +12 -8

README.md CHANGED Viewed

@@ -88,13 +88,17 @@ Get the fairseq checkpoint [here](https://drive.proton.me/urls/P83GCPNM40#2f0f87
 If you use GeistBERT in your research, please cite the following paper:
 ```
-@misc{scheibleschmitt2025geistbertbreathinglifegerman,
-      title={GeistBERT: Breathing Life into German NLP},
-      author={Raphael Scheible-Schmitt and Johann Frei},
-      year={2025},
-      eprint={2506.11903},
-      archivePrefix={arXiv},
-      primaryClass={cs.CL},
-      url={https://arxiv.org/abs/2506.11903},
 }
 ```

 If you use GeistBERT in your research, please cite the following paper:
 ```
+@book{scheible-schmitt-frei-2025-geistbert,
+  author    = {\textbf{Scheible-Schmitt}, \textbf{Raphael}  and  Frei, Johann},
+  title     = {GeistBERT: Breathing Life into German NLP},
+  booktitle      = {Proceedings of the Workshop on Beyond English: Natural Language Processing for all Languages in an Era of Large Language Models},
+  month          = {September},
+  year           = {2025},
+  address        = {Varna, Bulgaria},
+  publisher      = {INCOMA Ltd., Shoumen, BULGARIA},
+  pages     = {42--50},
+  abstract  = {Advances in transformer-based language models have highlighted the benefits of language-specific pre-training on high-quality corpora. In this context, German NLP stands to gain from updated architectures and modern datasets tailored to the linguistic characteristics of the German language. GeistBERT seeks to improve German language processing by incrementally training on a diverse corpus and optimizing model performance across various NLP tasks. We pre-trained GeistBERT using fairseq, following the RoBERTa base configuration with Whole Word Masking (WWM), and initialized from GottBERT weights. The model was trained on a 1.3 TB German corpus with dynamic masking and a fixed sequence length of 512 tokens. For evaluation, we fine-tuned the model on standard downstream tasks, including NER (CoNLL 2003, GermEval 2014), text classification (GermEval 2018 coarse/fine, 10kGNAD), and NLI (German XNLI), using $F_1$ score and accuracy as evaluation metrics. GeistBERT achieved strong results across all tasks, leading among base models and setting a new state-of-the-art (SOTA) in GermEval 2018 fine text classification. It also outperformed several larger models, particularly in classification benchmarks. To support research in German NLP, we release GeistBERT under the MIT license.},
+  url       = {https://aclanthology.org/2025.globalnlp-1.6},
+  doi       = {https://doi.org/10.26615/978-954-452-105-9-006}
 }
 ```