# Technical Documentation ## Data The projects requires additional runtime data to be able to performs its operations. These are the dictionaries in csv format: - data/lang/es: Spanish dictionaries. - data/lang/gl: Gallego dictionaries. - data/lang/gl/abr.txt: Contains a series of abbreviations, along with their expanded form. - data/lang/gl/acron.txt: Contains a series of acronyms, along with their expanded form. - data/lang/gl/adxectivos.txt: Contains a series of adjectives, along with their number and gender. - data/lang/gl/bigramas.dat: Binary data of a 2-gram neural network. - data/lang/gl/bigramas_real.dat: Binary data of a 2-gram neural network. - data/lang/gl/clases_ambiguedad.dat: Binary data of a neural network that collapses the grammar category of a word. - data/lang/gl/derivativos.txt: Contain rules for generating derivative words from certain rules. - data/lang/gl/desinencias.txt: Contains a series of verb endings. - data/lang/gl/locucions.txt: Contains a series of verb locucions, along with their grammar category. - data/lang/gl/nomes.txt: Contains a series nouns, with their number and gender. - data/lang/gl/palabrasFunction.txt: Dictionary with all the function words. - data/lang/gl/pentagramas.dat: Binary data of a 5-gram neural network. - data/lang/gl/pentagramas_real.dat: Binary data of a 5-gram neural network. - data/lang/gl/perifrase.txt: Dictionary with all the periphrases. - data/lang/gl/tetragramas.dat: Binary data of a 4-gram neural network. - data/lang/gl/tetragramas_real.dat: Binary data of a 4-gram neural network. - data/lang/gl/trigramas.dat: Binary data of a 3-gram neural network. - data/lang/gl/trigramas.dat: Binary data of a 3-gram neural network. - data/lang/gl/unigramas.dat: Binary data of a 1-gram neural network. - data/lang/gl/utf8chars.csv: Contains a series of custom transformation from UTF-8 words to latin1 - data/lang/gl/verbos.txt: Contains all the verbs, along with their roots. ## Classes - Cotovia: The primary entry point of the application. It handles the loading and unloading of dictionaries and processes sentences. - PreprocessorUTF8::Processor: Handles the transformation of strings in UTF-8 into latin1. - Tokenizar: Manages the tokenization of an input string - Preproceso: Applies a series of transformations to a tokenized input string to ensure it is in a canonical form. - Analisis_morfologico: Applies a morphologic analysis in which it deduces the classification of each token without contextual information. - Analisis_morfosintactico: Uses a neural network and an n-gram model to compute the classification of each word. - Sintagma: Performs syntagma analysis on the input tokens, clustering the tokens into groups. - Transcripcion: Takes a tokenized sentence and transforms it into multiple phonemes. ## Tests In the "tests" folder, there is a python script that performs integration tests on the cotovia projects. The following aspects are tested: - Allomorphs of articles in affirmative sentences with a period - Allomorphs of articles in affirmative sentences - Allomorphs of articles in exclamatory sentences - Allomorphs of articles in interrogative sentences - Galician recognition - Galician voice - Feminine names - Masculine names - Number recognition - Number pronunciation - Ordinal numbers - Roman numerals - Punctuation marks - Punctuation marks (variant N) - Phonemes