A Machine Learning-Enhanced Data Envelopment Analysis
Framework for Predicting Institutional Efficiency in Higher
Education
Hiteshkumar Solanki1,* Paresh Virparia2 Devika Madalli3
1 Scientist-D (CS), Information and Library Network (INFLIBNET) Centre, Gandhinagar, India
2 Professor, P G Department of Computer Science & Technology, Sardar Patel University, Vallabh Vidyanagar,
India
3 Director, Information and Library Network (INFLIBNET) Centre, Gandhinagar, India
Emails: hitesh@inflibnet.ac.in · pvvirparia@yahoo.com · dmadalli@gmail.com
Received: November 07, 2025 Revised: December 17 2025 Accepted: January 17, 2026 Corresponding author
ABSTRACT
The paper presents a novel approach for predicting the efficiency of engineering higher educational institutions in
India. The prediction model uses data envelopment analysis to solve linear programming problems and support
vector regression as a supervised machine learning algorithm. DEA is a nonparametric tool for computing the
relative efficiency of HEIs across multiple inputs and outputs. A total of 25 featured variables is considered, out
of which 15 are input-oriented, whereas 10 are output-oriented. Input-oriented features are normalised using the
standard z-score method, whereas output-oriented features are normalised using the Min–Max normalised method. A
sample of 7432 engineering institutions is considered for the featured dataset. The featured dataset is split at a 90:10
ratio for training and testing. The hybrid model, combining traditional DEA with a supervised machine learning
SVR, is trained on a training dataset. The prediction model is tested for estimating 743 engineering institutions. The
model provides higher accuracy, along with a precision score, in the confusion matrix comparing the actual and
predicted HEI performance categories. The best-fitting hybrid model also yields a predicted efficiency index for
engineering HEIs, along with MAE, MSE, and RMSE, which are 0.0703, 0.0080, and 0.0894, respectively. The
R-squared score of the prediction model is 0.7642. The developed predictive model can be generalised across sectors
such as agriculture, banking, industry, and education.
Keywords: Supervised Machine Learning Data Envelopment Analysis Relative Efficiency Support Vector Regression
Prediction Model Higher Education
1. INTRODUCTION
The Indian Higher Education System is the largest diversified
education sector in the world. A total of 1,168 university-level
institutions, 45,473 colleges and 12,002 stand-alone institutions
are registered in the All-India Survey of Higher Educational
Institutions as per the AISHE Report 2021–22 [1].
It offers a wide range of programs, including undergraduate,
postgraduate, and doctoral studies, across disciplines such
as science, engineering, humanities, management, and vocational
education, as per the All-India Survey of Higher
Education [1]. Higher educational institutions are classified
into several categories, including central universities, state
universities, deemed-to-be-universities and private universities,
institutes of national importance, affiliated colleges,
and stand-alone institutions. Higher education institutions
are compromising on infrastructure, manpower, experimental