•  
  •  
 

Middle East Journal of Communication Studies

DOI

10.71220/2790-5616.1099

Abstract

Objectives: This study develops and evaluates an Arabic scientific misinformation detection system by fine-tuning AraBERT-base-v2. It examines the effects of early stopping and input sequence length on model performance, interprets selected linguistic characteristics associated with misleading content, and discusses the limitations of using machine-translated data.

Methodology: The study adopted a mixed-methods design, employing a systematic integration of quantitative and qualitative approaches, supported by an interpretive qualitative reading. The initial database consisted of 23,546 records, including 123 Arabic articles collected from the Akeed, Sheek, and Taqeen platforms, and 23,423 foreign-language records drawn from the GossipCop and PolitiFact collections within FakeNewsNet. After data cleaning, 73 Arabic articles were retained for descriptive and interpretive analysis, while the foreign-language data were machine-translated into Arabic using the MarianMT model (Helsinki-NLP/opus-mt-en-ar), and subsequently cleaned and linguistically normalized. The final dataset used in the model experiments comprised 21,595 records, divided into 17,276 records for training and 4,319 records for validation. The binary classification into genuine and misleading content was based on the sources' original labels. Fine-tuning of the AraBERT-base-v2 model was then carried out through three experiments: the baseline model, early stopping, and an increased sequence length of up to 512 tokens.

Results: The findings revealed that the baseline model achieved the highest classification accuracy (80.9%) in detecting scientific news; however, it exhibited signs of overfitting. In contrast, applying the Early Stopping strategy improved the balance between performance and stability, achieving an accuracy of 75.7% while enhancing the model's generalization capability. Increasing the maximum sequence length to 512 tokens resulted in a lower accuracy (73.1%), likely due to the inclusion of irrelevant information. Furthermore, the model demonstrated superior performance in identifying fake scientific news compared with genuine scientific news.

Conclusion: The findings indicate that AraBERT is an effective tool for classifying Arabic scientific misinformation. However, achieving optimal performance requires a careful balance between data quality and the training strategies employed. The results also emphasize the importance of regularization techniques, such as Early Stopping, in mitigating overfitting and improving model generalization, whereas increasing the input sequence length does not necessarily lead to better performance. Accordingly, the study recommends prioritizing improvements in dataset quality and diversity, together with careful hyper parameters tuning, to support the development of more efficient and reliable misinformation detection systems for real-world applications.

Share

COinS
 
 

To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.