Development of Machine Learning Based System for Prediction of Diabetes
Keywords:
Diabetes Prediction, Machine Learning, Ensemble Learning, Voting Classifier, Random Forest, AdaBoost, XGBoost, Recursive Feature Elimination, Streamlit, Early Disease DetectionAbstract
The early prediction of diabetes mellitus is of great importance in controlling the progression of the disease and in managing its complications. Since traditional diagnostic methods are unable to predict the disease in its early stages, the present study suggests a diabetes prediction system based on machine learning in order to improve the rate at which diabetes can be predicted. In order to achieve this aim, the PIMA Indians Diabetes Dataset was used to create the prediction model. The data obtained from the dataset were preprocessed and put in order for analysis. Recursive feature elimination was carried out in order to select the most relevant features for the prediction model. A hybrid model based on the Voting Classifier was designed by using Random Forest, AdaBoost, and XGBoost classifiers. The performance of the proposed model was assessed by means of the 5-fold cross-validation technique. Statistical parameters such as accuracy, precision, recall, and AUC-ROC curves were calculated in order to determine the performance metrics of the model. The model suggested had an accuracy of 88 per cent and a recall of 87 per cent. Moreover, an optimal threshold value of 0.52 was found for the model in order to decrease prediction errors, such as false positives and false negatives. Lastly, a user-friendly web application was created using the Streamlit library in order to predict whether or not a person has diabetes. The present study has developed an efficient machine learning-based model which is able to accurately predict the onset of diabetes. The model proposed can be altered and put into use in a clinical setting to help healthcare professionals in the effective management of the disease.