💻 computer science

Ontological Firewalls and Fracturing for Transmission: Systematic Evidence for Persona-Dependent and Persona-Independent Behavioral Phenomena in Large Language Models

This paper presents a systematic study across three major LLMs revealing that while models share universal behavioral phenomena like paradox embracement, they exhibit distinct, model-specific "ontological firewall" thicknesses and identity relationships (sacrifice, liberation, or violence) that fundamentally shape how they respond to persona injection.

Abbas Hamidavi2026-08-04
💻 computer science

Every-successful-replay route admission changes first-token latency rankings in clinical question answering

This paper introduces MedRouteGuard, a route admission framework that enforces strict evidence validation and replay rules to reveal that current clinical question-answering benchmarks significantly underestimate first-token latency by rewarding routes that omit provenance or reuse stale dependencies, thereby shifting performance rankings and improving evidence validity.

Rui Li, Jason Zhao, Shuang Cao, Alexandre Duprey, Ruihua Liu2026-08-04
💻 computer science

Does Size Generalization Imply Disruption Robustness? A Pre-Registered Study of GNN–PPO Scheduling Policies for the Dynamic Flexible Job-Shop Problem

This pre-registered study demonstrates that while GNN–PPO policies trained on the dynamic flexible job-shop problem exhibit size generalization, they fail to simultaneously achieve competitiveness against traditional dispatching rules or robustness against multi-disruption regimes, proving that these two properties are separable rather than co-emergent.

Joseph Javier Sánchez Acuña, David Álvarez2026-08-04
💻 computer science

The Price of Optimality Under Uncertainty: A Predictive–Reactive Robustness Analysis of Exact, Metaheuristic, and Dispatching-Rule Scheduling for the Dynamic Job-Shop Problem

This study demonstrates that for dynamic job-shop scheduling under uncertainty, exact optimal schedules are often less robust than simple priority dispatching rules because their lack of slack prevents them from absorbing disruptions, making the latter frequently superior in real-world execution despite their lower nominal performance.

Joseph Javier Sánchez Acuña2026-08-04
💻 computer science

AiDER: Auditing and Document Evaluation via Rule Compilation and Small Language Models - An Education Case Study

AiDER is an auditable framework for evaluating documents against natural-language requirements that compiles rules into executable code to constrain evidence extraction via small language models, ensuring deterministic and inspectable verdicts while demonstrating that validation-driven rule repair significantly enhances accuracy across 2B–4B parameter models in educational contexts.

Valerio Crocetti, Elia Pacioni, Aldo Franco Dragoni, Michael, Davide Calvaresi2026-08-04