A Robustly Optimized BERT using Random Oversampling for Analyzing Imbalanced Stock News Sentiment Data

Salsabila Mazya Permataning Tyas*, Riyanarto Sarno, Agus Tri Haryono, Kelly Rossa Sungkono

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

2 Citations (Scopus)

Abstract

Stock news is one of the information sources that can used to monitor stock prices. The information from stock news usually contains positive and negative sentiments that can affect stock prices. Therefore, sentiment analysis is needed to process the sentiment of stock news. The stock news dataset is taken from Kaggle. From these data, there is an imbalanced class between positive and negative sentiment. This research proposed a method to solve the imbalance dataset with random oversampling which worked by randomly replicating several minority classes. This research presents several scenarios of pre-processing text with different stages, intending to get high accuracy. The classification method used in this paper is a robustly optimized Bidirectional Transformer Encoder Representation (RoBERTa). Besides that, this paper also compared with baseline of Machine Learning (ML) such as Multinomial Naïve Bayes, Bernoulli Naïve Bayes, Support Vector Machine, Random Forest Classifier, Logistic Regression and used two different text representation such as TF-IDF and Word2Vec. The best result in this research is obtained using RoBERTa method with the fourth scenario of pre-processing text, in which the stage of pre-processing in this scenario only removing hashtag, without removing punctuation, removing the number, converting number, stop word removal, and lemmatization. The performance result is 0.85 precision, 0,84 recall, 0,84 F1-score, and 86% for accuracy result.

Original languageEnglish
Title of host publicationICCoSITE 2023 - International Conference on Computer Science, Information Technology and Engineering
Subtitle of host publicationDigital Transformation Strategy in Facing the VUCA and TUNA Era
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages897-902
Number of pages6
ISBN (Electronic)9798350320954
DOIs
Publication statusPublished - 2023
Event2023 International Conference on Computer Science, Information Technology and Engineering, ICCoSITE 2023 - Jakarta, Indonesia
Duration: 16 Feb 2023 → …

Publication series

NameICCoSITE 2023 - International Conference on Computer Science, Information Technology and Engineering: Digital Transformation Strategy in Facing the VUCA and TUNA Era

Conference

Conference2023 International Conference on Computer Science, Information Technology and Engineering, ICCoSITE 2023
Country/TerritoryIndonesia
CityJakarta
Period16/02/23 → …

Keywords

  • pre-processing text
  • random oversampling
  • robustly optimized BERT
  • sentiment analysis
  • stock news

Fingerprint

Dive into the research topics of 'A Robustly Optimized BERT using Random Oversampling for Analyzing Imbalanced Stock News Sentiment Data'. Together they form a unique fingerprint.

Cite this