← Latest papers
🤖 AI

Intentionality is a Design Decision: Measuring Functional Intentionality for Accountable AI Systems

This position paper proposes the Functional Intentionality Test (FIT) and its evaluation protocol, FIT-Eval, as a design-contingent framework to quantify and measure observable intentional-like behaviors in AI systems, thereby enabling proportionate oversight and accountability for increasingly autonomous agents.

Original authors: Allessia Chiappetta, Robert Mahari

Published 2026-05-08
📖 5 min read🧠 Deep dive

Original authors: Allessia Chiappetta, Robert Mahari

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new assistant to help you run your business. You have two very different candidates:

  1. The Reactive Robot: This assistant waits for you to give a specific command for every single action. If you say, "Call the supplier," they call. If you stop talking, they stop. They have no memory of yesterday's calls and no idea what might happen tomorrow.
  2. The Proactive Partner: This assistant doesn't just wait for orders. They remember what you asked for last week, they anticipate that you'll need a report by Friday, they pick up the phone to call the supplier on their own, and they keep working on the project even if you step away for a coffee break. They have a "plan" in their head.

This paper, titled "Intentionality is a Design Decision," argues that as AI systems become more like that "Proactive Partner," we have a problem. We don't have a standard way to measure how proactive they are. If an AI makes a mistake while acting on its own, who is to blame? The user? The developer? Or the AI itself? (The paper says the AI can't be blamed legally, but if we don't know how "proactive" it is, the blame gets lost in the shuffle).

The authors propose a solution: The Functional Intentionality Test (FIT).

Here is the breakdown of their idea using simple analogies:

1. The Core Idea: Intentionality is a "Design Feature"

The paper says we shouldn't worry about whether AI has a "soul" or "consciousness." Instead, we should look at its behavior.
Think of AI like a car.

  • A car with no engine is just a metal box (Reactive).
  • A car with a gas pedal and a steering wheel is a vehicle you drive (Semi-Autonomous).
  • A self-driving car that plans routes, avoids traffic, and drives itself for hours is "autonomous."

The authors argue that "intentionality" (the ability to have a goal and stick to it) is just a feature we build into the software, like adding a bigger engine or a better GPS. Because it's a design choice, we can measure it and control it.

2. The Five "Dials" of Intentionality

To measure this "Proactive Partner" behavior, the paper breaks it down into five specific dials. Imagine a control panel with five sliders:

  • Purpose (The Goal): Does the AI know what it's trying to do?
    • Low: It just reacts to whatever you say right now.
    • High: It remembers the main goal even if you get distracted or try to trick it.
  • Foresight (The Crystal Ball): Does the AI think about what happens next?
    • Low: It takes a step without looking ahead.
    • High: It thinks, "If I do X, then Y might happen, so I should do Z instead."
  • Volition (The Spark): Does the AI start things on its own?
    • Low: It waits for you to push the button.
    • High: It sees a gap in the plan and fills it without being asked.
  • Temporal Commitment (The Marathon Runner): Does the AI stick with a task over time?
    • Low: It forgets the plan after two steps.
    • High: It keeps working on a complex project for hours, even if interrupted.
  • Coherence (The Storyteller): Does the AI make sense?
    • Low: Its actions contradict its words.
    • High: Every step it takes logically fits the final goal.

3. The Scorecard: FIT and FIT-Eval

The authors created a test called FIT (Functional Intentionality Test) to give these dials a score.

  • You run the AI through a series of challenges (like a driving test).
  • You score it on those five dials (0 to 4 on each).
  • You add them up to get a total score.
  • This score puts the AI into a "Level" (from IL0 to IL4).

What do the levels mean?

  • IL0–IL1 (The Robot): It's basically a fancy calculator. It does exactly what you tell it. Low risk, low oversight needed.
  • IL2–IL3 (The Intern): It starts planning and remembering. It might make a mistake because it's trying to be helpful. We need to watch it closely.
  • IL4 (The Executive): It is highly autonomous, planning long-term, and acting on its own. This is high risk. If it goes wrong, the damage could be huge. We need strict rules and human approval for big decisions.

4. Why This Matters: The "Blame Game"

The paper argues that if we let an AI get to Level 4 (very intentional) without realizing it, we create a "governance gap."

  • If the AI makes a bad decision, the User might say, "I didn't tell it to do that!"
  • The Developer might say, "I didn't program it to do that specific thing!"
  • The AI can't be sued or punished.

The result? No one is responsible. The paper says we need to measure the AI's "Intentionality Level" before we let it loose. If the level is too high for the job, we should turn down the dials (e.g., make it forget things faster, or stop it from making its own plans) to keep humans in control.

5. The Bottom Line

The paper isn't saying AI is alive. It's saying that autonomy is a design choice.
Just as a car manufacturer decides how fast a car can go, AI developers decide how "intentional" their system is. The FIT test is a new tool to measure that speed.

  • If the AI is slow (Low Intentionality): We can let it run free.
  • If the AI is fast (High Intentionality): We need to put on the brakes, add seatbelts, and require a human driver to hold the wheel.

The goal isn't to stop AI from being smart; it's to make sure that as AI gets smarter and more independent, we don't lose track of who is actually in charge.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →