💻 computer science
Condence, Calibration, and Abstention in Small Open-Weight Language Models on Psychiatric Knowledge Questions: A Cross-Domain Replication
This pilot study demonstrates that small, open-weight language models applied to psychiatric knowledge questions exhibit poor confidence calibration, domain-dependent accuracy rankings, and that prompting them to abstain from answering often degrades rather than improves performance when evaluated with rigorous paired statistics.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.