Technical Documentation
Data
The projects requires additional runtime data to be able to performs its operations. These are the dictionaries in csv format:
data/lang/es: Spanish dictionaries.
data/lang/gl: Gallego dictionaries.
data/lang/gl/abr.txt: Contains a series of abbreviations, along with their expanded form.
data/lang/gl/acron.txt: Contains a series of acronyms, along with their expanded form.
data/lang/gl/adxectivos.txt: Contains a series of adjectives, along with their number and gender.
data/lang/gl/bigramas.dat: Binary data of a 2-gram neural network.
data/lang/gl/bigramas_real.dat: Binary data of a 2-gram neural network.
data/lang/gl/clases_ambiguedad.dat: Binary data of a neural network that collapses the grammar category of a word.
data/lang/gl/derivativos.txt: Contain rules for generating derivative words from certain rules.
data/lang/gl/desinencias.txt: Contains a series of verb endings.
data/lang/gl/locucions.txt: Contains a series of verb locucions, along with their grammar category.
data/lang/gl/nomes.txt: Contains a series nouns, with their number and gender.
data/lang/gl/palabrasFunction.txt: Dictionary with all the function words.
data/lang/gl/pentagramas.dat: Binary data of a 5-gram neural network.
data/lang/gl/pentagramas_real.dat: Binary data of a 5-gram neural network.
data/lang/gl/perifrase.txt: Dictionary with all the periphrases.
data/lang/gl/tetragramas.dat: Binary data of a 4-gram neural network.
data/lang/gl/tetragramas_real.dat: Binary data of a 4-gram neural network.
data/lang/gl/trigramas.dat: Binary data of a 3-gram neural network.
data/lang/gl/trigramas.dat: Binary data of a 3-gram neural network.
data/lang/gl/unigramas.dat: Binary data of a 1-gram neural network.
data/lang/gl/utf8chars.csv: Contains a series of custom transformation from UTF-8 words to latin1
data/lang/gl/verbos.txt: Contains all the verbs, along with their roots.
Classes
- Cotovia: The primary entry point of the application. It handles the loading and unloading of dictionaries and processes sentences.
- PreprocessorUTF8::Processor: Handles the transformation of strings in UTF-8 into latin1.
- Tokenizar: Manages the tokenization of an input string
- Preproceso: Applies a series of transformations to a tokenized input string to ensure it is in a canonical form.
- Analisis_morfologico: Applies a morphologic analysis in which it deduces the classification of each token without contextual information.
- Analisis_morfosintactico: Uses a neural network and an n-gram model to compute the classification of each word.
- Sintagma: Performs syntagma analysis on the input tokens, clustering the tokens into groups.
- Transcripcion: Takes a tokenized sentence and transforms it into multiple phonemes.
Tests
In the "tests" folder, there is a python script that performs integration tests on the cotovia projects. The following aspects are tested:
- Allomorphs of articles in affirmative sentences with a period
- Allomorphs of articles in affirmative sentences
- Allomorphs of articles in exclamatory sentences
- Allomorphs of articles in interrogative sentences
- Galician recognition
- Galician voice
- Feminine names
- Masculine names
- Number recognition
- Number pronunciation
- Ordinal numbers
- Roman numerals
- Punctuation marks
- Punctuation marks (variant N)
- Phonemes