NLP-Based Approach to Multilingual Fake News Detection through Social Media in Low-Resource Languages: A Review

Authors

  • Sibgha Munir* Department of Artificial Intelligence, University of Engineering and Technology (UET), Pakistan
  • Haris Munir Department of Computer Science, COMSATS University Islamabad, Sahiwal Campus, Pakistan

DOI:

https://doi.org/10.18178/JAAI.2026.4.2.76-93

Keywords:

fake news detection, low-resource languages, multilingual Natural Language Processing (NLP), social media misinformation

Abstract

Especially in multilingual and impoverished language situations, the quick proliferation of false news on social media jeopardizes social stability. Emphasizing difficulties specific to low-resource environments, this paper offers a thorough examination of current Natural Language Processing (NLP) techniques for fake news detection across several languages. It examines prominent techniques, including named entity recognition, sentiment analysis, and text categorization, highlighting their uses, advantages, and drawbacks. Particularly focused on advanced methods like transfer learning, multilingual embeddings, and cross-lingual models all of which attempt to get around the scarcity of labeled data and the complexity of linguistic variety. The article also draws attention to deficiencies in present techniques and stresses the need for flexible models able to handle developing disinformation problems. The research provides ideas to help in the creation of strong, inclusive, and efficient tools for reducing the world-wide spread of false information by combining present progress with gaps.

References

[1] Dhiman, B. (2023). The rise and impact of misinformation and fake news on digital youth: A critical review. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.4438362

[2] De, A., Bandyopadhyay, D., Gain, B., & Ekbal, A. (2022). A transformer-based approach to multilingual fake news detection in low-resource languages. ACM Transactions on Asian and Low-Resource Language Information Processing, 21(1), 1. https://doi.org/10.1145/3472619

[3] Jand, H. (2018). Natural language processing with speech recognition. International Journal of Advanced Research Trends in Engineering and Technology (IJARTET).

[4] Bhardwaj, A., Khanna, P., Kumar, S., & Pragya. (2020). Generative model for NLP applications based on component extraction. In Procedia Computer Science. https://doi.org/10.1016/j.procs.2020.03.391

[5] Tsai, C. M. (2023). Stylometric fake news detection based on natural language processing using named entity recognition: In-domain and cross-domain analysis. Electronics, 12(17), 3676. https://doi.org/10.3390/electronics12173676

[6] Mohawesh, R., Liu, X., Arini, H. M., Wu, Y., & Yin, H. (2023). Semantic graph based topic modelling framework for multilingual fake news detection. AI Open, 4. https://doi.org/10.1016/j.aiopen.2023.08.004

[7] Zhang, Y., et al. (2023). Stance-level sarcasm detection with BERT and stance-centered graph attention networks. ACM Transactions on Internet Technology, 23(2), 20. https://doi.org/10.1145/3533430

[8] Sharifani, K., Amini, M., Akbari, Y., & Godarzi, J. A. (2022). Operating machine learning across natural language processing techniques for improvement of fabricated news model. International Journal of Science and Information System Research, 12(9).

[9] Alnabhan, M. Q., Branco, P., & Alnabhan, M. (2023). Fake news detection using deep learning: A systematic literature review. IEEE Access. https://doi.org/10.1109/ACCESS.2023

[10] Alghamdi, J., Lin, Y., & Luo, S. (2024). Fake news detection in low-resource languages: A novel hybrid summarization approach. Knowledge-Based Systems, 296, 111884. https://doi.org/10.1016/j.knosys.2024.111884

[11] Fagundes, M. J. G., Roman, N. T., & Digiampietri, L. A. (2024). The use of syntactic information in fake news detection: A systematic review. SBC Reviews on Computer Science, 4(1), 1–10. https://doi.org/10.5753/reviews.2024.2718

[12] Krasadakis, P., Sakkopoulos, E., & Verykios, V. S. (2024). A survey on challenges and advances in natural language processing with a focus on legal informatics and low-resource languages. Electronics, 13(3), 648. https://doi.org/10.3390/electronics13030648

[13] Meesad, P. (2021). Thai fake news detection based on information retrieval, natural language processing and machine learning. SN Computer Science, 2(6), 434. https://doi.org/10.1007/s42979-021-00775-6

[14] Agras, K., & Atay, B. (2024). A novel transformer-based deep learning pipeline for multilingual fake news detection. International Journal of Applied Mathematics Electronics and Computers, 6(4), 12. https://doi.org/10.36838/v6i4.12

[15] Cekinel, R. F., Karagoz, P., & Coltekin, C. (2024). Cross-Lingual Learning vs. Low-Resource Fine-Tuning: A Case Study with Fact-Checking in Turkish. arXiv preprint, arXiv:2403.00411

[16] Han, S. (2022). Cross-Lingual Transfer Learning for Fake News Detector in a Low-Resource Language. arXiv preprint, arXiv:2208.12482

[17] Zovikoğlu, M., & Çetin, U. (2024). Detecting misinformation on social networks with natural language processing. DergiPark. Retrieved from https://dergipark.org.tr/

[18] Ziyaden, A., Yelenov, A., Hajiyev, F., Rustamov, S., & Pak, A. (2024). Text data augmentation and pre-trained language model for enhancing text classification of low-resource languages. PeerJ Computer Science, 10, e1974. https://doi.org/10.7717/peerj-cs.1974

[19] PMookdarsanit, P., & Mookdarsanit, L. (2021). The COVID-19 fake news detection in Thai social texts. Bulletin of Electrical Engineering and Informatics, 10(2), 981–988. https://doi.org/10.11591/eei.v10i2.2745

[20] Farhangian, F., Cruz, R. M. O., & Cavalcanti, G. D. C. (2024). Fake news detection: Taxonomy and comparative study. Information Fusion, 103, 102140. https://doi.org/10.1016/j.inffus.2023.102140

[21] Hu, B., Mao, Z., & Zhang, Y. (2024). An overview of fake news detection: From a new perspective. Fundamental Research. https://doi.org/10.1016/j.fmre.2024.01.017

[22] Nasir, J. A., & Din, Z. U. (2021). Syntactic structured framework for resolving reflexive anaphora in Urdu discourse using multilingual NLP. KSII Transactions on Internet and Information Systems, 15(4), 1460–1479. https://doi.org/10.3837/tiis.2021.04.012

[23] Ranathunga, S., Lee, E. S. A., Skenduli, M. P., Shekhar, R., Alam, M., & Kaur, R. (2023). Neural machine translation for low-resource languages: A survey. ACM Computing Surveys, 55(11), 229. https://doi.org/10.1145/3567592

[24] Reddy, R., B., Tejaswi, T., Naveen, P., Giresh, J., & Vamsi, B. (2023). Optimising the detection of fake news in multilingual exposition using machine learning techniques. Proceedings of the 7th International Conference on Intelligent Computing and Control Systems (ICICCS 2023). IEEE. https://doi.org/10.1109/ICICCS56967.2023.10142901

[25] Wijayanti, R., Khodra, M. L., Surendro, K., & Widyantoro, D. H. (2023). Learning bilingual word embedding for automatic text summarization in low resource language. Journal of King Saud University —Computer and Information Sciences, 35(4), 101–112. https://doi.org/10.1016/j.jksuci.2023.03.015

[26] Mahmud, T., Ptaszynski, M., Eronen, J., & Masui, F. (2023). Cyberbullying detection for low-resource languages and dialects: Review of the state of the art. Information Processing & Management, 60(5), 103454. https://doi.org/10.1016/j.ipm.2023.103454

[27] Na, S., Na, P., Na, T., & Nasim, Z. (2022). Verify: Breakthrough accuracy in the Urdu fake news detection using text classification. Proceedings of the 36th Pacific Asia Conference on Language, Information and Computation.

[28] Azhar, S. A. F., Hidayat, F., Azfarezat, M. H., Nabiilah, G. Z., & Rojali. (2023). Efficiency of fake news detection with text classification using natural language processing. Journal of Theoretical and Applied Information Technology, 101(22).

[29] Yang, Y., et al. (2020). Multilingual universal sentence encoder for semantic retrieval. Proceedings of the Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-demos.12

[30] Cuenca, W., González-Fernández, C., Fernández-Isabel, A., Martín de Diego, I., & Martín, A. G. (2022). Combining conceptual graphs and sentiment analysis for fake news detection. In Studies in Computational Intelligence (Vol. 955). Springer. https://doi.org/10.1007/978-3-030-88817-6_15

[31] Mohamed, B., Haytam, H., & Abdelhadi, F. (2022). Applying fuzzy logic and neural network in sentiment analysis for fake news detection: Case of Covid-19. In Studies in Computational Intelligence (Vol. 1001). Springer. https://doi.org/10.1007/978-3-030-90087-8_19

[32] Taherdoost, H., & Madanchian, M. (2023). Artificial intelligence and sentiment analysis: A review in competitive research. Computers, 12(2), 37. https://doi.org/10.3390/computers12020037

[33] Ligthart, A., Catal, C., & Tekinerdogan, B. (2021). Systematic reviews in sentiment analysis: A tertiary study. Artificial Intelligence Review, 54(7), 4997–5053. https://doi.org/10.1007/s10462-021-09973-3

[34] Li, J., Sun, A., Han, J., & Li, C. (2022). A survey on deep learning for named entity recognition. IEEE Transactions on Knowledge and Data Engineering, 34(1), 50–70. https://doi.org/10.1109/TKDE.2020.2981314

[35] Ehrmann, M., Hamdi, A., Pontes, E. L., Romanello, M., & Doucet, A. (2023). Named entity recognition and classification in historical documents: A survey. ACM Computing Surveys, 56(2), 44. https://doi.org/10.1145/3604931

[36]Budi, I., & Suryono, R. R. (2023). Application of named entity recognition method for Indonesian datasets: A review. Bulletin of Electrical Engineering and Informatics, 12(2), 1089–1098. https://doi.org/10.11591/eei.v12i2.4529

[37] Sun, Z., & Li, X. (2023). Named entity recognition model based on feature fusion. Information, 14(2), 133. https://doi.org/10.3390/info14020133

[38] Wang, H., Zhou, L., Duan, J., & He, L. (2023). Cross-lingual named entity recognition based on attention and adversarial training. Applied Sciences, 13(4), 2548. https://doi.org/10.3390/app13042548

[39] Hernández, M. S. (2023). Beliefs and attitudes of canarians towards the Chilean linguistic variety. Lenguas Modernas, 62, 183–209. https://doi.org/10.13039/501100011033

[40] Dhyani, B. (2021). Transfer learning in natural language processing: A survey. Mathematical Statistician and Engineering Applications, 70(1). https://doi.org/10.17762/msea.v70i1.2312

[41] Palani, B., & Elango, S. (2023). CTrL-FND: Content-based transfer learning approach for fake news detection on social media. International Journal of System Assurance Engineering and Management, 14(3), 1013–1025. https://doi.org/10.1007/s13198-023-01891-7

[42] Ghayoomi, M., & Mousavian, M. (2022). Deep transfer learning for COVID-19 fake news detection in Persian. Expert Systems, 39(8), e13008. https://doi.org/10.1111/exsy.13008

[43] Schwarz, S., Theophilo, A., & Rocha, A. (2020). Emet: Embeddings from multilingual-encoder transformer for fake news detection. Proceedings of ICASSP 2020—IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE. https://doi.org/10.1109/ICASSP40776.2020.9054673

[44] Kasim, S. (2023). One true pairing: Evaluating effective language pairings for fake news detection employing zero-shot cross-lingual transfer. In Communications in Computer and Information Science (Vol. 1780). Springer. https://doi.org/10.1007/978-3-031-27609-5_2

[45] Han, X., Zhao, W., Ding, N., Liu, Z., & Sun, M. (2022). PTR: Prompt tuning with rules for text classification. AI Open, 3, 27–37. https://doi.org/10.1016/j.aiopen.2022.11.003

[46] Yang, J., Hu, X., Xiao, G., & Shen, Y. (2024). A survey of knowledge enhanced pre-trained language models. ACM Transactions on Asian and Low-Resource Language Information Processing.

[47] Howard, J., & Ruder, S. (2018). Universal language model fine-tuning for text classification. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers). https://doi.org/10.18653/v1/P18-1031

[48] Gururaja, S., Dutt, R., Liao, T., & Rosé, C. (2023). Linguistic representations for fewer-shot relation extraction across domains. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.414

[49] Wang, Z., & Hershcovich, D. (2023). On evaluating multilingual compositional generalization with translated datasets. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.93

[50] Farokhian, M., Rafe, V., & Veisi, H. (2024). Fake news detection using dual BERT deep neural networks. Multimedia Tools and Applications, 83(15), 44245–44268. https://doi.org/10.1007/s11042-023-17115-w

[51] lghamdi, J., Lin, Y., & Luo, S. (2023). Towards COVID-19 fake news detection using transformer-based models. Knowledge-Based Systems, 274, 110642. https://doi.org/10.1016/j.knosys.2023.110642

[52] G Kuling, G., Curpen, B., & Martel, A. L. (2022). BI-RADS BERT and using section segmentation to understand radiology reports. Journal of Imaging, 8(5), 131. https://doi.org/10.3390/jimaging8050131

[53] Wang, Y., et al. (2018). A comparison of word embeddings for the biomedical natural language processing. Journal of Biomedical Informatics, 87, 12–20. https://doi.org/10.1016/j.jbi.2018.09.008

[54] Shi, X., Hu, M., Deng, J., Ren, F., Shi, P., & Yang, J. (2023). Integration of multi-branch GCNs enhancing aspect sentiment triplet extraction. Applied Sciences, 13(7), 4345. https://doi.org/10.3390/app13074345

[55] Hao, Y., Xue, T., & Liu, G. (2023). Research on the robustness of neural machine translation systems in word order perturbation. Chinese Journal of Network and Information Security, 9(5), 120–130. https://doi.org/10.11959/j.issn.2096-109x.2023078

[56] Mohammed, I., & Prasad, R. (2023). Building lexicon-based sentiment analysis model for low-resource languages. MethodsX, 11, 102460. https://doi.org/10.1016/j.mex.2023.102460

[57] Haq, I., Qiu, W., Guo, J., & Tang, P. (2023). NLPashto: NLP toolkit for low-resource Pashto language. International Journal of Advanced Computer Science and Applications, 14(6). https://doi.org/10.14569/IJACSA.2023.01406142

[58] Zeng, Z., & Bhat, S. (2021). Idiomatic expression identification using semantic compatibility. Transactions of the Association for Computational Linguistics, 9, 154–170. https://doi.org/10.1162/tacl_a_00442

[59] Singh, S., & Singh, P. P. (2023). Transitive and intransitive verb analysis for idiomatic expression understanding: An NLP-based framework. International Journal of Advanced Research in Science, Communication and Technology, 123–130. https://doi.org/10.48175/ijarsct-11682

[60] Raffel, C., et al. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21, 1–67.

[61] Careem, R., Johar, G., & Khatibi, A. (2024). Deep neural networks optimization for resource-constrained environments: Techniques and models. Indonesian Journal of Electrical Engineering and Computer Science, 33(3), 1843–1854. https://doi.org/10.11591/ijeecs.v33.i3.pp1843-1854

[62] Kreutzer, J., et al. (2022). Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10, 50–72. https://doi.org/10.1162/tacl_a_00447

[63] Fuchs, C. (2020). Cultural and contextual affordances in language MOOCs: Student perspectives. International Journal of Online Pedagogy and Course Design, 10(4). https://doi.org/10.4018/IJOPCD.2020040104

[64] Wang, L., Feng, M., Zhou, B., Xiang, B., & Mahadevan, S. (2015). Efficient hyper-parameter optimization for NLP applications. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP). https://doi.org/10.18653/v1/D15-1253

[65] Dementieva, D., & Panchenko, A. (2020). Fake news detection using multilingual evidence. Proceedings of 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (DSAA). IEEE. https://doi.org/10.1109/DSAA49011.2020.00111

[66] Tan, Z., et al. (2024). Large Language Models for Data Annotation and Synthesis: A Survey. arXiv preprint, arXiv:2402.13446.

[67] Alzubaidi, L., et al. (2023). A survey on deep learning tools dealing with data scarcity: Definitions, challenges, solutions, tips, and applications. Journal of Big Data, 10(1), 46. https://doi.org/10.1186/s40537-023-00727-2

[68] Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.560

[69] Edwards, J. (2012). Multilingualism: Understanding Linguistic Diversity. Continuum. https://doi.org/10.5860/choice.50-1302

[70] Abdedaiem, A., Dahou, A. H., & Cheragui, M. A. (2023). Fake news detection in low resource languages using SetFit framework. Inteligencia Artificial, 26(72), 178–201. https://doi.org/10.4114/intartif.vol26iss72pp178-201

[71] Raja, E., Soni, B., & Borgohain, S. K. (2023). Fake news detection in Dravidian languages using transfer learning with adaptive finetuning. Engineering Applications of Artificial Intelligence, 126, 106877. https://doi.org/10.1016/j.engappai.2023.106877

[72] Gereme, F., Zhu, W., Ayall, T., & Alemu, D. (2021). Combating fake news in ‘low-resource’ languages: Amharic fake news detection accompanied by resource crafting. Information, 12(1), 20. https://doi.org/10.3390/info12010020

[73] Zoph, B., Yuret, D., May, J., & Knight, K. (2016). Transfer learning for low-resource neural machine translation. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP). https://doi.org/10.18653/v1/D16-1163

[74] Chen, Y., et al. (2022). Cross-modal ambiguity learning for multimodal fake news detection. Proceedings of the ACM Web Conference 2022. https://doi.org/10.1145/3485447.3511968

[75] Shu, K., Wang, S., & Liu, H. (2019). Beyond news contents: The role of social context for fake news detection. Proceedings of the 12th ACM International Conference on Web Search and Data Mining. https://doi.org/10.1145/3289600.3290994

[76] Dementieva, D., Kuimov, M., & Panchenko, A. (2023). Multiverse: Multilingual evidence for fake news detection. Journal of Imaging, 9(4), 77. https://doi.org/10.3390/jimaging9040077

[77] Park, M., & Chai, S. (2023). Constructing a user-centered fake news detection model by using classification algorithms in machine learning techniques. IEEE Access, 11, 78945–78958. https://doi.org/10.1109/ACCESS.2023.3294613

[78] Dou, Y., Shu, K., Xia, C., Yu, P. S., & Sun, L. (2021). User preference-aware fake news detection. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. https://doi.org/10.1145/3404835.3462990

Downloads

Published

2026-04-28

Issue

Section

Article