The Scaling Law of Evaluation Failure: Why Simple Averaging Collapses Under Data Sparsity and Item Difficulty Gaps, and How Item Response Theory Recovers Ground Truth Across Domains
यह शोध पत्र प्रदर्शित करता है कि कई डोमेनों में विषम आइटम कठिनाइयों वाले विरल मूल्यांकन मैट्रिसेस (sparse evaluation matrices) में सरल औसत विफल रहता है, जबकि आइटम रिस्पॉन्स थ्योरी (IRT) मॉडल इन परिस्थितियों में भी सुदृढ़ता से वास्तविक रैंकिंग को पुनः प्राप्त कर लेते हैं।