Machine Learning-Based Predictive Analytics: Investigating the Influence of Feature Selection and Model Optimization on Prediction Accuracy
DOI:
https://doi.org/10.71317/jgst.2.9(s).2026.614Keywords:
Machine Learning, Classification, Feature Selection, Model Optimization, Random Forest, Predictive Analytics, Classification PerformanceAbstract
This study comparatively evaluated the predictive performance of machine learning classification models and examined whether feature selection and model optimization improved classification accuracy and overall predictive effectiveness. A quantitative computational research design was employed using a structured dataset containing the predictor variables required for classification, which was preprocessed and divided into training and testing subsets before model development. Four machine learning algorithms, namely Logistic Regression, Decision Tree, Random Forest, and Support Vector Machine (SVM), were trained under baseline, feature-selection, model-optimization, and combined feature-selection-plus-optimization configurations. Feature selection was performed using SelectKBest and Recursive Feature Elimination (RFE), while model optimization was applied to improve the predictive configurations of the algorithms. Model performance was evaluated using accuracy, precision, recall, and F1-score, allowing comparative assessment of classification effectiveness. The results showed that feature selection consistently improved predictive accuracy, with RFE producing stronger improvements than SelectKBest. Model optimization also enhanced performance, with accuracy improvements of 5.50% for Logistic Regression, 6.20% for Decision Tree, 6.10% for Random Forest, and 5.40% for SVM compared with their respective baseline models. When feature selection and optimization were combined, mean accuracy increased from 83.25% in the baseline configuration to 93.00%, demonstrating a substantial improvement in overall predictive performance. Under the combined approach, Logistic Regression achieved 92.30% accuracy, Decision Tree achieved 90.40%, SVM achieved 94.10%, and Random Forest achieved the highest accuracy of 95.20%, with precision, recall, and F1-score all reaching 0.95. Overall, the findings demonstrate that integrating feature selection with model optimization substantially enhances classification performance, with the optimized Random Forest emerging as the most accurate and balanced model for the evaluated dataset.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Muhammad Ahsan Tariq, Fasie Haider, Taaha Shahzad, Muhammad Umar Farooq (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.



