A Machine Learning-Enhanced Data Envelopment Analysis

Framework for Predicting Institutional Efficiency in Higher

Education

Hiteshkumar Solanki1,* Paresh Virparia2 Devika Madalli3

1 Scientist-D (CS), Information and Library Network (INFLIBNET) Centre, Gandhinagar, India

2 Professor, P G Department of Computer Science & Technology, Sardar Patel University, Vallabh Vidyanagar,

India

3 Director, Information and Library Network (INFLIBNET) Centre, Gandhinagar, India

Emails: hitesh@inflibnet.ac.in · pvvirparia@yahoo.com · dmadalli@gmail.com

Received: November 07, 2025 Revised: December 17 2025 Accepted: January 17, 2026 Corresponding author

ABSTRACT

The paper presents a novel approach for predicting the efficiency of engineering higher educational institutions in

India. The prediction model uses data envelopment analysis to solve linear programming problems and support

vector regression as a supervised machine learning algorithm. DEA is a nonparametric tool for computing the

relative efficiency of HEIs across multiple inputs and outputs. A total of 25 featured variables is considered, out

of which 15 are input-oriented, whereas 10 are output-oriented. Input-oriented features are normalised using the

standard z-score method, whereas output-oriented features are normalised using the Min–Max normalised method. A

sample of 7432 engineering institutions is considered for the featured dataset. The featured dataset is split at a 90:10

ratio for training and testing. The hybrid model, combining traditional DEA with a supervised machine learning

SVR, is trained on a training dataset. The prediction model is tested for estimating 743 engineering institutions. The

model provides higher accuracy, along with a precision score, in the confusion matrix comparing the actual and

predicted HEI performance categories. The best-fitting hybrid model also yields a predicted efficiency index for

engineering HEIs, along with MAE, MSE, and RMSE, which are 0.0703, 0.0080, and 0.0894, respectively. The

R-squared score of the prediction model is 0.7642. The developed predictive model can be generalised across sectors

such as agriculture, banking, industry, and education.

Keywords: Supervised Machine Learning Data Envelopment Analysis Relative Efficiency Support Vector Regression

Prediction Model Higher Education

1. INTRODUCTION

The Indian Higher Education System is the largest diversified

education sector in the world. A total of 1,168 university-level

institutions, 45,473 colleges and 12,002 stand-alone institutions

are registered in the All-India Survey of Higher Educational

Institutions as per the AISHE Report 2021–22 [1].

It offers a wide range of programs, including undergraduate,

postgraduate, and doctoral studies, across disciplines such

as science, engineering, humanities, management, and vocational

education, as per the All-India Survey of Higher

Education [1]. Higher educational institutions are classified

into several categories, including central universities, state

universities, deemed-to-be-universities and private universities,

institutes of national importance, affiliated colleges,

and stand-alone institutions. Higher education institutions

are compromising on infrastructure, manpower, experimental