Comparative analysis of preprocessing tasks over social media texts in Spanish

2019 
One of the key aspects of the texts coming from social media is that they tend to be very noisy. This is mainly because of the usage of informal language and non-standard grammatical structures. Therefore, in order to use these contents as input for a text analysis process, it is highly recommended to previously clean and reduce the noise of the data. This work focuses on measuring the effectiveness that diverse cleaning and repairing tasks have on the data. The results obtained, indicate that the tasks of "tokens with no letters removal", and "stressed words processing" are the most effective. In addition, some tasks like hashtags or usernames processing, which behave very well in other datasets, are not that relevant in this one. This research is part of a more general one that pursues to build an automatic emotion classifier that makes use of the preprocessed comments as input.
    • Correction
    • Source
    • Cite
    • Save
    • Machine Reading By IdeaReader
    24
    References
    1
    Citations
    NaN
    KQI
    []