Machine Learning-Based Predictive Analytics: Investigating the Influence of Feature Selection and Model Optimization on Prediction Accuracy

Authors

  • Muhammad Ahsan Tariq MS Project Management, Bahria University; BSc Electrical Engineering, University of South Asia. Currently working at Nokia, Riyadh, Saudi Arabia Author
  • Fasie Haider Department of Computational Engineering, LUT University Author
  • Taaha Shahzad COMSATS University Islamabad, Vehari Campus Author
  • Muhammad Umar Farooq COMSATS University Islamabad, Vehari Campus Author

DOI:

https://doi.org/10.71317/jgst.2.9(s).2026.614

Keywords:

Machine Learning, Classification, Feature Selection, Model Optimization, Random Forest, Predictive Analytics, Classification Performance

Abstract

This study comparatively evaluated the predictive performance of machine learning classification models and examined whether feature selection and model optimization improved classification accuracy and overall predictive effectiveness. A quantitative computational research design was employed using a structured dataset containing the predictor variables required for classification, which was preprocessed and divided into training and testing subsets before model development. Four machine learning algorithms, namely Logistic Regression, Decision Tree, Random Forest, and Support Vector Machine (SVM), were trained under baseline, feature-selection, model-optimization, and combined feature-selection-plus-optimization configurations. Feature selection was performed using SelectKBest and Recursive Feature Elimination (RFE), while model optimization was applied to improve the predictive configurations of the algorithms. Model performance was evaluated using accuracy, precision, recall, and F1-score, allowing comparative assessment of classification effectiveness. The results showed that feature selection consistently improved predictive accuracy, with RFE producing stronger improvements than SelectKBest. Model optimization also enhanced performance, with accuracy improvements of 5.50% for Logistic Regression, 6.20% for Decision Tree, 6.10% for Random Forest, and 5.40% for SVM compared with their respective baseline models. When feature selection and optimization were combined, mean accuracy increased from 83.25% in the baseline configuration to 93.00%, demonstrating a substantial improvement in overall predictive performance. Under the combined approach, Logistic Regression achieved 92.30% accuracy, Decision Tree achieved 90.40%, SVM achieved 94.10%, and Random Forest achieved the highest accuracy of 95.20%, with precision, recall, and F1-score all reaching 0.95. Overall, the findings demonstrate that integrating feature selection with model optimization substantially enhances classification performance, with the optimized Random Forest emerging as the most accurate and balanced model for the evaluated dataset.

Downloads

Published

2026-09-15

How to Cite

Tariq, M. A., Haider, F., Shahzad, T., & Farooq, M. U. (2026). Machine Learning-Based Predictive Analytics: Investigating the Influence of Feature Selection and Model Optimization on Prediction Accuracy. Journal of Global Social Transformation, 2(9.1), 313-326. https://doi.org/10.71317/jgst.2.9(s).2026.614