💻 computer science

What Makes a Programming Problem Hard for a Language Model? An Empirical Study of Item Difficulty Across Code LLMs on Two Benchmarks

本文通过一项实证研究表明,代码生成基准测试中的问题难度是一个稳定的、可迁移的指标,其由 HumanEval 中的规范特征(如示例和提示词长度)以及 MBPP 中的解法复杂度所驱动,这为在模型聚合评分趋于饱和时改进基准测试、自动评分和教育工具设计提供了关键见解。

TANZIM ISLAM KHAN2026-07-13
💻 computer science

Detecting ovarian endometriomas from ultrasound using vision transformers and cross-modality transfer learning

本文提出了一种开创性的机器学习模型,该模型利用视觉 Transformer 和跨模态迁移学习从超声图像中检测卵巢内膜异位囊肿,在数据稀缺和不平衡的情况下实现了高性能,同时证明了低级特征在医学成像模态间的可迁移性。

Matthew Watson, Miliani Fraser-Fletcher, Tom Willshare, Molly Jowsey, Noura Al Moubayed2026-07-13
💻 computer science

Human–AI collaboration in deductive coding of classroom dialogue: prompt engineering and collaboration patterns

本研究采用基于设计的研究方法,旨在证明当定制化 GPT 编程助手与结构化提示词及以人为本的协作工作流相结合时,能有效支持课堂对话的演绎编码,同时强调成功的机人协作取决于提示词架构、工作流排序以及研究者对人工智能态度的动态相互作用。

Luwei Bai, Dongkeun Han, Sara Hennessy2026-07-13
💻 computer science

When Does Context Help Machine Vision? Decomposing the Risk and Reliability of Contextual Conditioning in Vision-Language Models

本文系统地分析了不同架构和数据集下视觉语言模型中的上下文条件化现象,揭示了虽然领域级提示(domain-level prompting)始终具有益处,但全量单图上下文条件化(full per-image contextual conditioning)存在高方差和过拟合风险,且其效用主要受模型规模和基准准确率的驱动。

Alaa Alahmadi2026-07-13
💻 computer science

Mining AI-Assisted Course Design Workflows at Production Scale

本文通过从 CourseFactory 挖掘一个保护隐私的数据集以提取四个工作流界面,展示了首次在生产规模上进行的 AI 辅助课程设计研究,证明了结构质量分拣和题目到作业的路由模型显著提高了评审效率与准确性,同时证实了被隐藏的提示词文本并未增加任何可衡量的信号。

Aleksandr Volkov, Taras Pustovoy, Dmitriy Istomin, Viacheslav Istomin, Tatiana Otbetkina2026-07-13
💻 computer science

H. Dilpriya's Momentum (HDM) : A Multi-Strategy Gradient-Aligned Optimizer with Adaptive Per-Parameter Corrections, Cosine-Annealed Scheduling, and Convergence Guarantees for Deep Neural Networks

本文介绍了 H. Dilpriya's Momentum (HDM),这是一种结合了自适应逐参数修正与余弦退火调度的新型多策略优化器,旨在实现严格的收敛保证和最先进的梯度对齐,并在病态合成问题以及 MNIST 和 CIFAR-10 等深度学习基准测试上展示了卓越的性能。

Janaka Ishan Senarathna2026-07-13