Enhanced Diabetes Prediction Accuracy through Cutting-Edge Deep Learning and Ensemble Machine Learning Techniques
Main Article Content
Abstract
Diabetes mellitus is a long-term metabolic dis-order, which is relevant to over 537 million individuals across the globe, and the prevalence is expected to rise to 46 percent by 2045. Predication of risks at an early and precise stage is very necessary to minimize compli-cations, mainly in healthcare systems that are short of re-sources. This research will offer a hybrid predictive model that integrates ensemble machine learning and deep learn-ing models in identifying diabetes at its early stages based on clinical and demographic data. The Pima Indian Di-abetes Dataset (PIDD) comprising of 768 female patients was experimented on. The pre-processing of the data con-sisted of imputation of missing values, elimination of out-liers, balancing the classes by using SMOTE and augmen-tation of the dataset to 10,000 training balanced samples using multivariate Gaussian sampling used exclusively on the training set. Four ensemble models, which are Bagging, Gradient Boosting, Random Forest and Stacking, were op-timized with gridsearchCV. An early stopping and dropout Deep Neural Network (DNN) as well as a Long Short-Term Memory (LSTM) model were trained. The evaluation of performance was based on the accuracy, precision, recall, F1-score, and ROC-AUC on a strictly held-out test set. On the augmented dataset, the DNN data obtained a 99.9% ac-curacy and ROC-AUC of 1.000, whereas ensemble models obtained 97%– 99.1% accuracy. The ensemble accuracies observed on the original dataset were between 73.4 per-cent and 77.3 percent as expected in literature. The key predictors were found to be glucose, BMI and age. There was cross-dataset generalization that was verified by out-side validation. The suggested framework has a high pre-dictive power, a high generalization, and results that are clinically interpretable, which evidences its possible use as a healthcare decision-support system.
Keywords- Diabetes Prediction; Ensemble Learning; Deep Learning; Bagging and Boosting; Random Forest; Pima In-dian Diabetes Dataset (PIDD).
Article Details
The author transfers all copyright ownership of the manuscript entitled (title of article) to the Technical Journal in the event the work is published.