💻 computer science

Every-successful-replay route admission changes first-token latency rankings in clinical question answering

This study demonstrates that enforcing explicit evidence-admission requirements in clinical question answering significantly alters route rankings and latency metrics while reducing material evidence-validity errors, highlighting the need for standardized admission rules in clinical latency benchmarks.

Rui Li, Jason Zhao, Shuang Cao, Alexandre Duprey, Ruihua Liu2026-09-15
💻 computer science

Condence, Calibration, and Abstention in Small Open-Weight Language Models on Psychiatric Knowledge Questions: A Cross-Domain Replication

This pilot study demonstrates that small, open-weight language models applied to psychiatric knowledge questions exhibit poor confidence calibration, domain-dependent accuracy rankings, and that prompting them to abstain from answering often degrades rather than improves performance when evaluated with rigorous paired statistics.

Kunal Dhanda2026-09-15
💻 computer science

Collective Epistemic Network Framework with Adaptive Field Control: Stability and Robustness in Adversarial Multi-Agent Systems

This study introduces the Collective Epistemic Network Framework (CENF), a computational model integrating adaptive field control and trust plasticity that demonstrates significant improvements in consensus accuracy and stability against adversarial attacks in synthetic multi-agent networks, while establishing formal convergence guarantees under specific design conditions.

Ali Moslemi Tabrizi2026-09-15
💻 computer science

Accuracy and Deployability of Deep Neural Networks for Human Action Recognition: A Six-Axis Survey and Controlled Benchmark

This paper presents a six-axis survey and controlled benchmark demonstrating that while deep neural networks like R3D-18, R(2+1)D-18, and MC3-18 achieve high accuracy in human action recognition, practical deployment requires balancing recognition performance with computational efficiency, temporal robustness, and resource constraints rather than relying on accuracy alone.

Nousheen Taj2026-09-15
💻 computer science

Token Probability Beats Verbalized Condence for Error Detection in Small Open-Weight Language Models: A Medical and Psychiatric Benchmark Study

This pilot study demonstrates that for small open-weight language models (1.2B–9.2B parameters) applied to medical and psychiatric tasks, extracting internal token probabilities is a significantly more reliable method for error detection than relying on the models' verbally expressed confidence, particularly for the smallest models where verbalized confidence performs at chance levels.

Kunal Dhanda2026-09-15
💻 computer science

Does a Safety-Priming Instruction Reduce Risk-Fact Omission in Psychiatric Note Summarization? A Synthetic-Vignette Study of Small Open-Weight Language Models

This synthetic-vignette study demonstrates that while adding a safety-priming instruction to small open-weight language models can significantly reduce the omission of critical risk factors in psychiatric note summarization for some models, it often creates a trade-off by increasing the omission of routine clinical details, suggesting that such safety interventions must be validated on a per-model basis rather than assumed to be universally effective.

Kunal Dhanda2026-09-15