Kuncahyo Setyo Nugroho, Fitra Abdurrachman Bachtiar, Wayan Firdaus Mahmudy, Matthew Martianus Henry, Mahmud Isnan, Gusti Pangestu, Bens Pardamean
The EmoTweetID dataset is a large-scale open resource of Indonesian tweets curated for emotion classification and word embedding, addressing the scarcity of publicly available datasets for this low-resource language. Tweets were collected from the X platform (formerly Twitter) using the snscrape library, yielding >4.5 million raw tweets based on basic and derived emotion keywords from Ekman's six basic emotions. After cleaning to remove duplicates, irrelevant tweets, and sensitive elements, a corpus of 3,126,987 clean unlabeled tweets was compiled from the first scraping stage. The second stage gathered 2,243 clean tweets annotated through lexicon-based and manual methods, into six emotion classes based on Ekman's basic emotions: anger, disgust, fear, joy, sadness, and surprise. The manual annotation by three psychology students used majority voting and achieved substantial inter-annotator agreement, with a Fleiss' Kappa score of 0.7323. Two pretrained embeddings, Word2Vec and fastText, were trained on the corpus using a 300-dimensional skip-gram architecture to enrich semantic representation. Baseline evaluations using BiLSTM with fastText achieved a weighted F1-score of 0.8285 on the human-annotated set, demonstrating the practical utility of the dataset and embeddings for the downstream emotion classification task. Publicly available on Mendeley Data, the EmoTweetID dataset provides a valuable foundation for advancing Indonesian natural language processing by enabling unsupervised pre-training and supervised multi-class emotion classification.