LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning
LinAlg-Bench यह प्रकट करता है कि लार्ज लैंग्वेज मॉडल्स 4x4 मैट्रिक्स आयामों पर एक संरचनात्मक व्यवहार संबंधी दहलीज (structural behavioral threshold) प्रदर्शित करते हैं, जहाँ वे छोटे कार्यों में निष्पादन त्रुटियों (execution errors) से बड़े कार्यों में व्यवस्थित कम्प्यूटेशनल परित्याग (systematic computational abandonment) और संरचित मतिभ्रम (structured hallucination) की ओर स्थानांतरित हो जाते हैं, जो ज्ञान की कमी के बजाय एक मौलिक वर्किंग मेमोरी सीमा का संकेत देता है।