Publicación: Sistema predictivo basado en técnicas de machine learning para la detección temprana del riesgo académico y la deserción estudiantil en una universidad peruana
Portada
Citas bibliográficas
Código QR
Autor corporativo
Recolector de datos
Otros/Desconocido
Director audiovisual
Editor
Tipo de Material
Fecha
Citación
Título de serie/ reporte/ volumen/ colección
Es Parte de
Resumen en español
En Perú, la tasa de deserción universitaria suele estar entre el 10 y el 13% según el III Informe Bienal de SUNEDU. Esto representa carreras interrumpidas y recursos institucionales que han quedado fuera de alcance y no pueden ser reembolsados en el futuro. La pregunta que motiva este proyecto sería: ¿es posible detectar a los estudiantes en riesgo antes de que abandonen, no después de que ya hayan abandonado? Para responder a esta pregunta, desarrollamos y verificamos un sistema predictivo de aprendizaje automático para encontrar a los estudiantes al final del primer ciclo universitario que tendrían la mayor probabilidad de no continuar en la universidad. El modelo combina datos académicos del SIAD y del Backoffice de UPCH con calificaciones de educación secundaria obtenidas mediante un proceso de reconocimiento óptico de caracteres (OCR) en certificados oficiales del Minedu. Este último paso fue el más difícil ya que los diferentes tipos de documentos utilizados para representar los datos en los documentos escolares tuvieron que ser limpiados varias veces antes de que los datos pudieran ser recuperados. Se compararon Random Forest y XGBoost con un modelo de regresión logística; dado que existe un desequilibrio significativo entre la clase mayoritaria (continuación) y la minoritaria (no continuación), el recall y las calibraciones probabilísticas fueron más importantes en la evaluación que la precisión general. El prototipo resultante tiene dos componentes: un panel en Dash/Plotly que muestra alertas basadas en cohorte, programa y estudiante y un servicio de inferencia en FastAPI con control de acceso basado en roles. Ninguno está conectado a sistemas de producción institucionales y no genera alertas sobre estudiantes reales automáticamente. Toda la canalización está estructurada utilizando CRISP-DM en Python con bibliotecas de código abierto como Pandas, scikit-learn y XGBoost con trazabilidad comprobada en cada paso. Tomamos la cohorte 2024-1 como un conjunto de validación externo independiente ylogramos un AUC-ROC de 85.7% y un Brier Score de 0.142, ambos por encima del umbral en la hipótesis. El sistema no está destinado a reemplazar el juicio del tutor, sino a proporcionar información más oportuna y estructurada que el informe promedio actualmente utilizado.
Resumen en inglés
In Peru, the university dropout rate is usually between 10% and 13% according to SUNEDU's III Biennial Report. This represents interrupted careers and institutional resources that are lost and cannot be recouped in the future. The driving question behind this project is: is it possible to detect at-risk students before they drop out, rather than after they have already left? To answer this question, we developed and verified a predictive machine learning system to identify students at the end of their first university term who have the highest probability of not continuing at the university. The model combines academic data from UPCH's SIAD and Backoffice with high school grades obtained through an optical character recognition (OCR) process on official Minedu certificates. This last step was the most difficult, as the different types of documents used to represent the data in school records had to be cleaned multiple times before the data could be retrieved. Random Forest and XGBoost were compared with a logistic regression model; given the significant imbalance between the majority class (continuation) and the minority class (non-continuation), recall and probabilistic calibrations were more important in the evaluation than overall accuracy. The resulting prototype has two components: a Dash/Plotly dashboard that displays alerts based on cohort, program, and student, and an inference service in FastAPI with role-based access control. Neither is connected to institutional production systems, and they do not automatically generate alerts about real students. The entire pipeline is structured using CRISP-DM in Python with open-source libraries such as Pandas, scikit-learn, and XGBoost, featuring proven traceability at every step. We used the 2024-1 cohort as an independent external validation set and achieved an AUC-ROC of 85.7% and a Brier Score of 0.142, both above the hypothesis threshold. The system is not intended to replace the tutor's judgment, but rather to provide more timely and structured information than the average report currently used.

PDF
FLIP 
