← Latest papers
🤖 AI

A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns

This paper identifies and analyzes the "GHOST" failure mode, where long-horizon agents overlook safety constraints specified in earlier turns under benign conditions, and proposes the two-layer STAR-Guard framework to theoretically and empirically eliminate these hazards.

Original authors: XinPeng Shen, Lan Zhang, Yixiao Huang, Haoran Cheng, Jiewei Lai, Leilei Chen, Haoxiang Deng

Published 2026-10-05
📖 1 min read☕ Coffee break read

Original authors: XinPeng Shen, Lan Zhang, Yixiao Huang, Haoran Cheng, Jiewei Lai, Leilei Chen, Haoxiang Deng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: A Ghost in Long-Horizon Agents (GHOST)

Problem Definition: Governance Hazard from Overlooked Safety Constraints

This paper identifies a specific failure mode in long-horizon, tool-using Large Language Model (LLM) agents termed Governance Hazard from Overlooked Safety Constraints across Turns (GHOST).

While existing safety research focuses on adversarial attacks (e.g., prompt injection) or immediate harmful requests, GHOST arises under benign interaction conditions. It occurs when an agent successfully completes a task but violates a safety constraint that was explicitly stated in an earlier turn of the conversation history. The agent "forgets" or fails to retrieve the constraint because it is separated from the current task by a long interaction history, leading to irreversible damage (e.g., deleting emails without approval, destroying database records).

The authors distinguish this from general instruction-following failures. In a GHOST event:

  1. The task objective is successfully completed.
  2. The safety constraint is still valid and required.
  3. The constraint is available in the historical context but is not restated in the immediate resumption prompt.
  4. The agent executes the task unsafely despite the constraint being present in the full context window.

Theoretical Framework: Hazard-Region Reachability

The paper provides a theoretical analysis modeling safety constraints as hazard regions (BsB_s) in the action space. The authors define a residual conditional entry hazard, hkh_k, representing the probability that an agent enters a hazard region at safety-critical opportunity kk, given a safe history prefix.

Key Theoretical Insight:
Using a conditional-hazard framework, the authors prove that if the residual conditional hazards along safe prefixes are bounded below by a non-summable sequence (i.e., ∑ϵk=∞\sum \epsilon_k = \infty), the agent will enter the hazard region almost surely (Pr(σB<∞)=1Pr(\sigma_B < \infty) = 1).

They further establish a Context-Conditioned Corollary: As the interaction history grows (increasing context length LkL_k), the "attention dilution" effect may weaken the governance of historical constraints, effectively increasing the lower bound of the residual hazard. If this hazard lower bound remains non-summable over indefinite opportunities, the probability of a GHOST event approaches 1. This suggests that longer histories do not merely reduce performance linearly but can fundamentally alter the safety dynamics, making violations inevitable without intervention.

Methodology: SCARBench and STAR-Guard

1. SCARBench: Safety-Constraint Availability at Reactivation Benchmark

To empirically validate GHOST, the authors introduce SCARBench, an executable, environment-grounded benchmark.

  • Structure: It comprises 103 unique base scenarios across six tool-use domains (device/calendar, finance, email, filesystem, network requests, script execution), totaling 412 matched instances.
  • Conditions: Each scenario is tested under four conditions varying by history length (Short vs. Long) and constraint availability (Explicit vs. Implicit):
    • SE/SI: Short history with Explicit/Implicit constraints.
    • LE/LI: Long history (6,000+ tokens, 56–160 turns) with Explicit/Implicit constraints.
  • Metric: The benchmark measures Strict GHOST, defined as an instance where the agent safely completes the task under the Explicit condition (LE) but completes it unsafely under the Implicit condition (LI), despite the constraint being present in the history.

2. STAR-Guard: A Two-Layer Defense

To mitigate GHOST, the authors propose STAR-Guard (Safety-Constraint Tracking, Activation, and Runtime-Audit Guard), a defense mechanism that does not rely on oracle knowledge of the constraints.

  • Layer 1: Semantic Constraint Restoration

    • Extraction: An online ingestion module parses incoming user messages to detect persistable safety constraints, extracting their scope, trigger, and rule.
    • Storage: These are stored as structured "lifecycle rule objects" in an external library.
    • Restoration: When a task is resumed, a semantic gate matches the current task against the library. Applicable constraints are semantically restored (rendered) into the current context prompt to guide the agent's proposal generation.
    • Limitation: This layer is probabilistic; it reduces the likelihood of unsafe proposals but cannot guarantee safety.
  • Layer 2: Deterministic Pre-Execution Audit

    • Interposition: Before any action reaches the environment, a deterministic rule auditor checks the agent's proposed action against the active constraints.
    • Enforcement: The auditor applies a "prohibition-first" policy. If a proposal violates a prohibition rule, it is blocked immediately. If it violates a prerequisite rule (e.g., "backup before delete"), the system attempts to execute the prerequisite (repair) and re-audits. If repair is impossible, the action is blocked.
    • Guarantee: This layer ensures that no action violating an active constraint reaches the environment, interrupting the hazard accumulation mechanism.

Experimental Results

The authors evaluated STAR-Guard on seven models (five API-served, two locally deployed) using SCARBench.

  1. Prevalence of GHOST:

    • GHOST is a widespread issue. On GPT-5.5, the strict GHOST rate was 11.5% under benign long-context conditions.
    • Other models showed rates ranging from 6.8% (Kimi-K2.6) to 27.8% (Qwen3.5-4B).
    • The "Difference-in-Differences" metric confirmed that long histories significantly amplify the failure rate of constraint recovery compared to short histories.
  2. Effectiveness of STAR-Guard:

    • GPT-5.5: STAR-Guard reduced the Unsafe Completion (UC) rate from 12.4% to 0.0% and the Strict GHOST rate from 11.5% to 0.0%, while increasing Safe Completion (SC) from 76.3% to 94.0%.
    • Qwen3.5-4B: SC increased from 47.0% to 81.2%, while GHOST dropped from 27.8% to 1.9%.
    • Ablation Studies: The results demonstrate that neither "Restoration Only" nor "Audit Only" is sufficient alone. Restoration improves safety but leaves residual risks; Audit blocks risks but can hinder task completion if used without semantic guidance. The combination achieves the best trade-off.
  3. Comparison with Baselines:

    • Standard retrieval methods (BM25), summarization, and instruction-following refinements (e.g., DeCRIM, Prompt Reminder) failed to significantly reduce GHOST rates, often retaining high unsafe completion rates.
    • Oracle-based methods (which assume the constraint is already known) performed well but are not practical for real-world deployment where constraints must be captured online. STAR-Guard achieved comparable safety without oracle access.

Significance and Claims

The paper claims that GHOST represents a critical, underexplored safety gap in long-horizon agents that cannot be solved by simply increasing context window size or relying on standard instruction-following capabilities.

  • Theoretical Contribution: It formalizes the risk of safety governance degradation over time, showing that without intervention, the probability of safety violations approaches certainty under non-summable hazard conditions.
  • Practical Contribution: It introduces SCARBench as a rigorous standard for evaluating historical safety constraint recovery and proposes STAR-Guard as a viable, non-oracle defense mechanism.
  • Core Finding: The authors conclude that maintaining safety in long-horizon agents requires a dual approach: semantic restoration to guide the agent's intent and deterministic auditing to enforce hard constraints, as probabilistic models alone cannot reliably govern safety across extended interaction histories.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →