Discrimination or Leniency? A Statistical and Machine-Learning Analysis of Internal versus External Assessment across Institution Types in a District of Nepal
DOI:
https://doi.org/10.65091/icicset.v3i1.117Abstract
The basic level examination for grade 8 has implemented an internal assessment system in Nepal's
basic school curriculum. Schools combine internal marks with board-external marks in many
high-stakes examination systems, but the question remains whether internal marks serve as a valid,
consistent substitute for external attainment. Whether the validity depends on the type of institution
awarding the marks remains largely unexamined in Nepal, where prior scholarship on the topic
is predominantly qualitative. Using district-level records for 2788 students drawn from three types
of institutions (Institutional/private, Community/public, and Alternative), this study characterizes
the internal-external relationship by analyzing the agreement using statistics, ceiling-aware
regression, and machine learning.
Concordance is weak, with Lin's concordance correlation coefficient not exceeding 0.15
in any type, and the form of the discrepancy differs systematically by the type of institution.
Institutional marks are compressed near the scale of the ceiling and fail to discriminate attainment
(area under the ROC curve, AUC = 0.62; a censoring-corrected Tobit slope of 0.20; a 90thpercentile with quantile slope of 0.00), whereas Community marks discriminate attainment
considerably better (AUC = 0.78) but sit systematically above external performance. The study
found that invalid internal grading has two distinct aspects, such as a discrimination failure and
a level (leniency) failure, that implicate different types of institutions and that the discrimination
finding is harmful to the quality of students. The results imply that internal marks are not
interchangeable across institution types and that any high-stakes collection of internal and
external marks should apply type moderation.