💻 computer science

Detector-Calibration Failures in Pattern-Based LLM Refusal Classification: Discovery, Generalization, and a Confirmed False-Positive Pattern Across Models

This paper details a three-phase investigation revealing that apparent non-determinism in LLM refusal detection was largely caused by correctable detector artifacts, which, while generalizable across models, introduce specific false-positive patterns that necessitate manual auditing and transparent reporting of data gaps to ensure accurate safety assessments.

Waqar Javed2026-09-22
💻 computer science

SafeAgent-300: A Balanced 300-Prompt Benchmark for Agentic AI Security, with Findings on Detector Coverage Gaps and Cross-Model Compliance Variance

This paper introduces SafeAgent-300, a balanced 300-prompt benchmark for agentic AI security that evaluates six models to reveal significant disparities in model conservatism, uncover a spontaneous tool-invocation behavior in a Google Gemini model, and identify critical detector coverage gaps that skew the relationship between prompt sophistication and violation detection rates.

Waqar Javed2026-09-22
💻 computer science

AgentPort-Bench: A Controlled Seven-Framework Evaluation of Agentic AI Security Portability

This paper presents a controlled, seven-framework evaluation demonstrating that while attack type and model choice significantly impact the security of tool-using LLM agents, the choice of orchestration framework generally has no meaningful effect on security posture, with the notable exception of CrewAI, which exhibits a small but statistically significant residual elevation in failure rates even after correcting for a specific adapter defect.

Waqar Javed2026-09-22
💻 computer science

Artificial Intelligence for Alzheimer’s Disease Detection Across Neuroimaging and Speech: A Reproducible Cross-Modal Secondary Analysis

This study synthesizes evidence from 26 AI-based Alzheimer's disease detection papers to reveal that while reported within-study accuracies are high, the field critically lacks external validation, multimodal integration, and linguistic diversity, necessitating a shift toward robust, reproducible, and clinically transportable model designs.

Jeffrey Tao, Christa C. Caggiano2026-09-22
💻 computer science

Runtime Continuation after Faults in Partially Executed Agent Workflows

This paper proposes a runtime continuation mechanism for partially executed language-model agent workflows that ensures safe recovery from faults by reconstructing a trusted frontier, deriving residual obligations from a canonical contract, and advancing state only through authorized effects supported by valid, fresh, and uniquely matching evidence.

Li Zeng, Chuanwu Yang, Qiang Zhao, Qinde Chen, Deyang Qu2026-09-22
💻 computer science

A Structural Preservation Framework for Denoiser Selection in YOLO-Based Pedestrian Detection under Sensor Noise

This paper introduces the Structural Preservation Score (SPS) to evaluate denoiser effectiveness for YOLO-based pedestrian detection, revealing that while specific variants like gradient-magnitude correlation can rank denoisers within known noise levels, no single image-level metric can universally predict detection performance across unseen denoisers or noise conditions due to non-linear, architecture-dependent factors.

Vo Thanh Kiet, Rene Jaros, Minh Ly Duc, Petr Bilik, Radek Martinek2026-09-22
💻 computer science

From NER to Business Process Automation: Comparative Evaluation of CNN, LSTM, and Transformer Models for Address Intelligence

This study provides a comparative evaluation of five neural architectures (WordCNN, WordLSTM, BERT, DistilBERT, and ELECTRA) for address extraction in business process automation, demonstrating that ELECTRA achieves the highest accuracy and efficiency while lighter models offer viable alternatives for latency-sensitive applications.

Saurabh Kumar Srivastava, Vikas Jalodia, Divya Srivastava2026-09-22
💻 computer science

Coligo: A Retrieval-Augmented Generation Assistant for WhatsApp-Based TNEA Engineering Admission Counselling

Coligo is a Retrieval-Augmented Generation assistant deployed on WhatsApp that alleviates the administrative burden of Tamil Nadu Engineering Admissions (TNEA) counselling by leveraging Google Gemini and vector search over college documents to provide accurate, context-aware, and category-specific admission guidance while explicitly distinguishing between its functional prototype and its intended full-scale architecture.

Nithishkumar A, Nitin A V, Prasanna S2026-09-22
💻 computer science

An Empirical Investigation of Multi-Trial Consistency, Trajectory Pathologies, and Reliability Rankings in Software Engineering Agents

This paper proposes a comprehensive multi-trial evaluation framework to assess the consistency, reliability, and trajectory pathologies of autonomous software engineering agents, demonstrating its operational feasibility through a pilot study that reveals significant infrastructure censoring rates and establishes a foundation for moving beyond single-trial success metrics.

Muhammad Tayyab2026-09-22