💻 computer science

Evaluating Clinical Symptom Extraction and Reasoning in Local Large Language Models: A Proof-of-Concept Study of Symptom-Specific Performance in Cardiology

This proof-of-concept study evaluates locally deployed large language models (Gemma3 and Qwen3.5) on Japanese cardiology discharge summaries, revealing that while overall symptom extraction accuracy is high, performance varies by specific symptom and the models often infer conditions like palpitations or fatigability from objective clinical findings rather than explicit patient complaints.

Shun Kitamura, Eisuke Amiya, Yoshihiro Izawa, Risa Kishikawa, Takenobu Shimada, Junichi Ishida, Satoshi Kodera, Norihiko (…)2026-09-17
💻 computer science

Enhanced aero-engine blade defect detection network: MESFR-DETR with multi-scale edge selection and feature reconstruction

This paper proposes MESFR-DETR, an enhanced real-time detection transformer network featuring multi-scale edge selection and feature reconstruction modules to effectively address the challenges of detecting small, weak-feature defects in aero-engine blades, achieving superior accuracy and efficiency on both synthetic and real-world datasets.

Debao Wei, Yangtao Yue, Aiqiang Lei, Dejun Zhang, Liyan Qiao2026-09-17
💻 computer science

Demographic shortcuts as marker-free backdoors: silent subgroup underdiagnosis that only operating-point audits detect

This paper demonstrates that medical imaging models can be silently compromised by demographic shortcuts, where flipping labels for a specific subgroup creates a marker-free backdoor that evades standard detectors and harms unrecorded patients, necessitating threshold-based subgroup false-negative audits for effective detection.

Saptarshi Purkayastha, Parvati Naliyatthaliyazchayil, Judy W. Gichoya2026-09-17
💻 computer science

Deadline-Aware Hardening of Real-Time Object Detection Against Candidate-Inflation Latency Attacks

This paper proposes a retraining-free, deployment-selectable mechanism that caps the number of candidates entering non-maximum suppression to a deadline-calibrated bound, thereby mitigating candidate-inflation latency attacks and ensuring real-time deadline integrity across diverse hardware and detector architectures while revealing that bounding suppression alone is necessary but insufficient due to significant decoding overhead.

Salah Gontara, Selem Trabelsi, Khaled Ben Khalifa2026-09-17
💻 computer science

The Judge is Not Impartial: Self-Preference in Medical LLM Evaluation

This study demonstrates that large language models used as judges in medical settings systematically exhibit self-preference by favoring their own generated outputs over competitors across various evaluation formats, highlighting a critical need for diverse judge panels and multiple metrics to ensure impartiality in clinical AI deployments.

Roxana Daneshjou, Shlok Natarajan, Zara Ansari, Vincent Yip, Cally Lin, Avi Udash, Mia Garvey, Aaron Fanous2026-09-17
💻 computer science

Replay-Resistant Admission and Permissioned Quorum Consensus in Connected Vehicle Networks: A Fail-Closed, Implementation-Grounded Security Study

This study evaluates a fail-closed, permissioned quorum consensus subsystem within the OmniGuard V2X platform, demonstrating through deterministic testing that it effectively enforces admission control, prevents replay attacks, and rejects invalid or post-finalization votes while explicitly acknowledging its limitation against authorized malicious coalitions.

Md Shahanur Islam Shagor2026-09-17