xlm_roberta_large / README.md
IMvision12's picture
Super-squash branch 'main' using huggingface_hub
d6ceddd
|
Raw
History Blame Contribute Delete
4.26 kB
metadata
pipeline_tag: fill-mask
license: mit
base_model: FacebookAI/xlm-roberta-large
library_name: kerasformers
tags:
  - keras
  - kerasformers
  - xlm-roberta
  - fill-mask
  - multilingual
  - text-encoder
  - arxiv:1911.02116
  - pytorch
  - jax
  - tf

See our collection for all versions of XLM-RoBERTa.

Run XLM-RoBERTa with Keras 3: JAX, PyTorch, or TensorFlow

GitHub Docs Collection

kerasformers/xlm_roberta_large

Paper: Unsupervised Cross-lingual Representation Learning at Scale (arXiv:1911.02116) · HF Papers

XLM-RoBERTa is the multilingual RoBERTa: same encoder architecture, pretrained on 2.5TB CommonCrawl across 100 languages, with a 250k SentencePiece vocabulary (mask token <mask>).

For more details on the model, please go to the upstream model card.

Pure-Keras 3 conversion of FacebookAI/xlm-roberta-large for kerasformers. One implementation runs unmodified on TensorFlow / Torch / JAX.

This is a fill-mask / encoder checkpoint (XLMRobertaMaskedLM, large). Task heads load via hf: fine-tunes.

✨ Quick start (multilingual fill-mask)

import os
os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

from kerasformers.models.xlm_roberta import (
    XLMRobertaMaskedLM,
    XLMRobertaTokenizer,
)

mlm = XLMRobertaMaskedLM.from_weights("kerasformers/xlm_roberta_large")
tokenizer = XLMRobertaTokenizer.from_weights("kerasformers/xlm_roberta_large")

# Multilingual: same <mask> API as RoBERTa, 100-language SentencePiece vocab.
inputs = tokenizer("La capitale de la France est <mask>.")
logits = mlm(inputs)  # (1, L, vocab_size)
mask = int((inputs["input_ids"][0] == tokenizer.mask_token_id).argmax())
print(tokenizer.decode([int(logits[0, mask].argmax())]))

Load any XLM-RoBERTa variant the same way with from_weights("kerasformers/<variant>"):

Variant Hub
xlm_roberta_base kerasformers/xlm_roberta_base
xlm_roberta_large kerasformers/xlm_roberta_large

Available classes

Load any of these from this repo with from_weights("kerasformers/xlm_roberta_large") (or on the fly via the hf: prefix). The pretrained backbone is shared; task heads not stored in this checkpoint start randomly initialized, ready for fine-tuning (or load a hf: fine-tune).

Class Task
XLMRobertaModel Encoder backbone
XLMRobertaMaskedLM Masked language modeling (fill-mask)
XLMRobertaSequenceClassify Sequence classification
XLMRobertaTokenClassify Token classification (NER / POS)
XLMRobertaQnA Extractive question answering
XLMRobertaMultipleChoice Multiple choice
from kerasformers.models.xlm_roberta import XLMRobertaSequenceClassify
model = XLMRobertaSequenceClassify.from_weights("kerasformers/xlm_roberta_large")

Tips

  • Set KERAS_BACKEND before importing Keras / kerasformers.
  • Prefer XLMRobertaTokenizer.from_weights(...) so the SentencePiece vocab matches.
  • Use <mask> (not [MASK]).
  • See XLM-RoBERTa docs and Loading Weights.
  • Community / upstream safetensors still work via the hf: prefix, e.g. XLMRobertaMaskedLM.from_weights("hf:FacebookAI/xlm-roberta-large").

Special Thanks

A huge thank you to the Facebook AI XLM-RoBERTa authors for creating and releasing these models.

License: MIT.