AnushS's picture
Update README.md
ca93b7b
|
Raw
History Blame
2.9 kB
metadata
license: apache-2.0
tags:
  - generated_from_trainer
model-index:
  - name: Hieroglyph-Translator-Using-Gardiner-Codes
    results: []

Hieroglyph Gardiner Code Translator Model

This model was created to translate hieroglyphs into english.

Egyptian Hieroglyphs have been grouped into different classes and given a referencing method called Gardiner Codes using Gardiner Classification.

Using the Gardiner Codes we can assign meanings to different combinations of hieroglyphs.

To Translate any sequence of hieroglyphs using this model, provide the following input :-

"Translate hieroglyph unicode sequence to English: {Gardiner Codes of the Hieroglyphs}"

Examples :

"Translate hieroglyph unicode sequence to English: A1 B6 F8"

"Translate hieroglyph unicode sequence to English: G4 H9 P3"

Model description

This model is a fine-tuned version of t5-small on a custom dataset derived from the Dictionary of Middle Egyptian.

The Inference Api on the hugging face model page doesn't work well, load the model in jupyter notebook using the following code snippet:

  text = "" # add your hieroglyph gardiner code combination in the string
  
  from transformers import AutoTokenizer
  
  tokenizer = AutoTokenizer.from_pretrained("AnushS/hieroglyph_unicode_translator_t5_small")
  inputs = tokenizer(text, return_tensors="pt").input_ids
  
  from transformers import AutoModelForSeq2SeqLM
  
  
  model = AutoModelForSeq2SeqLM.from_pretrained("AnushS/hieroglyph_unicode_translator_t5_small")
  outputs = model.generate(inputs, max_new_tokens=40, do_sample=True, top_k=30, top_p=0.95)
  translated_keywords = str(tokenizer.decode(outputs[0], skip_special_tokens=True))

Intended uses & limitations

The Model is intended to be used to translate hieroglyphs. The model does not provide full sentences, it only outputs bits and keywords.

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 16
  • eval_batch_size: 16
  • seed: 42
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • num_epochs: 2

Training results

Training Loss Epoch Step Validation Loss Gen Len
5.0665 1.0 688 4.2034 6.946
4.4621 2.0 1376 4.1388 6.946

Framework versions

  • Transformers 4.27.4
  • Pytorch 2.2.0.dev20231113
  • Datasets 2.12.0
  • Tokenizers 0.13.3