Optimization Strategies for Large-scale Network Traffic Analysis Using Spark

Authors

  • Trust Museta Department of Computer Science, The University of West Florida, Pensacola, Florida, USA
  • Sikha S. Bagui* Department of Computer Science, The University of West Florida, Pensacola, Florida, USA
  • Dustin Mink Department of Cybersecurity, The University of West Florida, Pensacola, Florida, USA
  • Subhash C. Bagui Department of Mathematics and Statistics, The University of West Florida, Pensacola, Florida, USA
  • Aaron Webb Department of Computer Science, The University of West Florida, Pensacola, Florida, USA
  • Myles Brown Department of Computer Science, The University of West Florida, Pensacola, Florida, USA
  • Venkata SriKrishna Garikapati Department of Computer Science, The University of West Florida, Pensacola, Florida, USA

DOI:

https://doi.org/10.18178/JAAI.2026.4.3.164-189

Keywords:

big data analytics, cybersecurity, Apache spark, principal component analysis, network traffic analysis, machine learning, random forest, support vector machines, naï ve bayes

Abstract

This paper presents a comprehensive big data analytics approach for cybersecurity threat detection using a network flow dataset, UWF-ZeekData22. A scalable solution using Apache Spark for distributed processing is presented. Principal Component Analysis (PCA) was used for dimensionality reduction and three different classifiers were used for threat detection. 2,044,734 network traffic records with 23 features were processed, achieving 99.999% variance retention through PCA dimensionality reduction and 100% classification accuracy. The study demonstrates significant performance optimization through systematic parameter tuning, reducing execution time from 9.024 s to 4.567 s (62.7% improvement) while maintaining perfect classification accuracy. These findings contribute to the field of big data cybersecurity analytics by providing evidence-based optimization strategies for large-scale network traffic analysis.

References

[1] Buczak, A. L., & Guven, E. (2016). A survey of data mining and ML methods for cyber security intrusion detection. IEEE Communications Surveys & Tutorials, 18(2), 1153–1176.

doi: 10.1109/COMST.2015.2494502

[2] Armbrust, M., Xin, R. S., Lian, C., Huai, Y., Liu, D., Bradley, J. K., Meng, X., Kaftan, T., Franklin, M. J., Ghodsi, A., & Zaharia, M. (May 31–June 4, 2015). Spark SQL: Relational data processing in Spark. Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data (pp. 1383–1394). Melbourne, Victoria, Australia. doi: 10.1145/2723372.2742797

[3] Bagui, S., & Spratlin, S. (2018). A review of data mining algorithms on Hadoop’s MapReduce. International Journal of Data Science, 3(2), 146–169. doi: 10.1504/ijds.2018.092285

[4] University of West Florida. (2023). UWF-Zeekdata22 dataset (.csv files). Retrieved from https://datasets.uwf.edu/

[5] Trellix. (2024). What is the MITRE ATT&CK framework? Retrieved from https://www.trellix.com/

[6] Bagui, S. S., Mink, D., Bagui, S. C., Madhyala, P., Uppal, N., McElroy, T., Plenkers, R., Elam, M., & Prayaga, S. (2023). Introducing the UWF-ZeekDataFall22 dataset to classify attack tactics from Zeek Conn logs using Spark’s ML in a big data framework. Electronics, 12(24), 5039. doi: 10.3390/electronics12245039

[7] Bagui, S. S., Eller, C., Armour, R., Singh, S., Bagui, S. C., Mink, D. (2025). Analyzing performance of data preprocessing techniques on CPUs vs. GPUs with and without the MapReduce environment. Electronics, 14, 3597. doi: 10.3390/electronics14183597

[8] Breiman, L. (1996). Bagging predictors. Machine Learning, 24(2), 123–140. doi: 10.1007/BF00058655

[9] Vapnik, V. (1999). The Nature of Statistical Learning Theory. New York, NY: Springer.

[10] Domingos, P., & Pazzani, M. (1997). On the optimality of the simple Bayesian classifier under zero-one loss. Machine Learning, 29(2–3), 103–130. doi: 10.1023/A:1007413511361

[11] Apache Spark. (2024). PySpark MLlib PCA. Retrieved from https://spark.apache.org/docs/3.5.1/api/python/index.html

[12] Shaikh, E., Mohiuddin, I. A., Alufaisan, Y., & Nahvi, I. (November 19–21, 2019). Apache Spark: A big data processing engine. Proceedings of the 2019 2nd IEEE Middle East and North

Africa Communications Conference (MENACOMM) (pp. 1–6). Manama, Bahrain. IEEE.

doi: 10.1109/MENACOMM46666.2019.8988541

[13] Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., & Stoica, I. (April 25–27, 2012). Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI) (pp. 15–28). San Jose, California. USENIX Association.

[14] Zaharia, M., Chowdhury, M., Das, T., Dave, A., Ma, J., McCauley, M., Franklin, M. J., Shenker, S., & Stoica, I. (June 22–25, 2010). Apache Spark: Cluster computing with working sets. Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing (HotCloud). Boston, Massachusetts. USENIX Association.

[15] Zaharia, M., Xin, R. S., Wendell, P., Das, T., Armbrust, M., Dave, A., Meng, X., Rosen, J., Venkataraman, S., Franklin, M. J., Ghodsi, A., Gonzalez, J., Shenker, S., & Stoica, I. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65. doi: 10.1145/2934664

[16] Meng, X., Bradley, J., Yavuz, B., Sparks, E., Venkataraman, S., Liu, D., Freeman, J., Tsai, D., Amde, M., Owen, S., Xin, D., Xin, R., Franklin, M. J., Zadeh, R., Zaharia, M., & Talwalkar, A. (2016). MLlib: Machine Learning in Apache Spark. Journal of Machine Learning Research, 17, 1–7.

[17] Jolliffe, I. T. (1986). Principal Component Analysis. New York, NY: Springer. doi: 10.1007/978-1-4757-1904-8

[18] Farnaaz, N., & Jabbar, M. A. (2016). Random Forest modeling for intrusion detection system. Procedia Computer Science, 89, 213–217. doi: 10.1016/j.procs.2016.06.047

[19] Ahn, G., Kim, J., Choi, C., Kim, K., & Kim, S. (2022). Malicious file detection using machine learning and the MITRE ATT&CK framework. Applied Sciences, 12(19), 9841. doi: 10.3390/app12199841.

[20] Disha, R. A., & Waheed, S. (2022). Performance analysis of Machine Learning models for intrusion detection system using Gini Impurity-Based Weighted RF (GIWRF) feature selection technique. Cybersecurity, 5(1), 1. doi: 10.1186/s42400-021-00103-8

[21] Ding, Q., & Kolaczyk, E. (2012). Compressed PCA subspace method. arXiv preprint, arXiv:1109.4408.

[22] Guller, M. (2015). Big Data Analytics with Spark: A Practitioner’s Guide to Using Spark for Large Scale Data Analysis. Berkeley, CA: Apress. doi: 10.1007/978-1-4842-0964-6

[23] Quinto, B. (2020). Next-Generation Machine Learning with Spark: Generative AI, Large Language Models, and Pyspark. New York, NY: Apress. doi: 10.1007/978-1-4842-5666-4

[24] Mukkamala, S., Janoski, G., & Sung, A. H. (May 12–17, 2002). Intrusion detection using neural networks and support vector machines. Proceedings of the International Joint Conference on Neural Networks (IJCNN) (Vol. 2, pp. 1702–1707). Honolulu, HI, USA, IEEE. doi: 10.1109/IJCNN.2002.1007774

[25] Ahmad, Z., Shahid Khan, A., Wai Shiang, C., Abdullah, J., & Ahmad, F. (2021). Network intrusion detection system: A systematic study of ML and deep learning approaches. Transactions on Emerging Telecommunications Technologies, 32(1), e4150. doi: 10.1002/ett.4150

[26] Miller, E., Mink, D., Spellings, P., Bagui, S. S., & Bagui, S. C. (2025). Classifying cyber ranges: A case-based analysis using the UWF cyber range. Encyclopedia, 5(4), 162. doi: 10.3390/encyclopedia5040162

Downloads

Published

2026-08-28

Issue

Section

Article