Memory or Generalization? Temporal Probing of Data Contamination in Bengali LLM Benchmarks
This paper introduces "temporal probing," a direct validity test using difficulty-matched, post-cutoff items to measure data contamination in Bengali LLM benchmarks, finding that current models on the BoolQ-bn dataset show no significant inflation from memorization despite the prevalence of static, translated test sets.