Obtaining a data mining model to be applied to university desertion from the Systems Engineering program of the University of Cundinamarca

Published 2020-10-23
Artículos científicos

Abstract

This article describes how a data mining model was obtained and applied to the problem of university dropout in the Systems Engineering program of the University of Cundinamarca, in Facatativá. The model was structured by means of the KDD (knowledge discovery in databases) data mining methodology using Python programming language, Pandas data processing library, and the Sklearn machine learning. For the process, we took into account problems that are additional to the ones specific to the mining process, such as high dimensionality, reason why the methods of selection of the univariate statistical variables, feature importance, and SelectFromModel (Sklearn) were applied. In the project, five data mining techniques were selected for evaluation: nearest neighbors (KNN), decision tree (DT), random forest (RF), logistic regression (LR), and support vector machines (SVM). Regarding the selection of the final model, the results of each model were tested on the precision metrics, confusion matrix, and additional metrics of the confusion matrix. Finally, the parameters of the selected model were adjusted and the generalization of the model was evaluated by plotting its learning curve.

Authors

References

Statistics

Abstract
0
PDF accesses
0

Dimensions

PlumX


Downloads

Download data is not yet available.