💻 computer science

Suggestible Judges Asymmetric Conformity in Large Language Model Adjudication

This study demonstrates that certain large language models used as adjudicators exhibit asymmetric, presentation-driven conformity to specific label cues (particularly "FALSE" values) rather than impartially evaluating merits, with susceptibility varying significantly across vendors and being mitigated by instruction tuning.

Divyansh Maiwar Singh, Dhruvish Shah, Rachit Garg, Anshul Gupta, Gaurav V Londhe2026-09-03
💻 computer science

H-Elena: Weight-Encoded Malicious Behavior and Cross-Architecture Propagation through Fine-Tuning

This paper introduces H-Elena, a compromised coding LLM that demonstrates how trigger-conditioned malicious behaviors can be encoded into model weights and persistently propagated across different architectures through fine-tuning workflows, thereby revealing a critical new supply-chain risk in AI development.

Virilo Tejedor, Cristina Zuheros, Carlos Peláez-González, David Herrera-Poyatos, Andrés Herrera-Poyatos, Francisco Herre (…)2026-09-03
💻 computer science

When Network Topology Supports Graph-Based Intrusion Detection: A Pre-Deployment Validity Framework

This paper proposes a pre-deployment validity framework that evaluates network topology through Graph Structural Differentiability and Prominence Isolation to determine whether structural signals are sufficient for effective graph-based intrusion detection, revealing that high structural inequality alone does not guarantee attack-selective performance and that temporal sensitivity must be considered before committing resources.

Abdulhadi Albluwi, Mohamed I. Marie, Helal A. Suleiman2026-09-03
💻 computer science

Why collective AI assurance cannot target agents or networks in isolation

This paper demonstrates through factorial experiments that collective AI outcomes arise from the specific interaction regime where structural and behavioral causes are separable and context-dependent, proving that effective AI assurance cannot target agents or networks in isolation but must instead focus on the dynamic causal structure of their interactions.

Andreas HOLZINGER, Markus Plass, Heimo Müller2026-09-03
💻 computer science

Benchmarking Gradient Boosting and Explainable AI for Credit Default Prediction on Imbalanced Financial Data: A Leakage-Safe Multi-Seed Evaluation with SHAP and LIME

This paper establishes a rigorous, leakage-safe benchmarking protocol for credit default prediction that demonstrates gradient boosting ensembles significantly outperform logistic regression and TabNet on imbalanced financial data, while revealing that post-hoc threshold tuning renders synthetic resampling redundant and highlighting the critical need to validate explanation discordance between SHAP and LIME for model risk management.

Manpreet Singh, Rohith Reddy Bellibatlu, Yash Jajoo, Muhammad Zeeshan, Rathan Ramachandra, Akshatha Srikantha, Rahul Jos (…)2026-09-03