Process-Constituted Intelligence: A Shared Criterion for Humans and Machines
This paper argues that true intelligence is defined by the iterative cognitive process rather than the output, proposing a framework of "strong equivalence" based on seven process features to distinguish human cognition from current generative AI, which merely reproduces output traces, while offering design principles and audits to ensure AI tools augment rather than erode human generative capacities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Intelligence is often mistaken for the final result: the solved equation, the finished painting, the correct answer given in a conversation. For decades, scientists have debated whether a machine that produces the same result as a human is truly thinking. This question has moved from philosophy into the daily reality of generative artificial intelligence, tools that can write essays, generate code, and create images that look indistinguishable from human work. The core issue is not just whether the output is good, but how it got there. In the study of the mind, a distinction has long existed between a system that merely matches human behavior and one that actually replicates the internal activity required to produce that behavior. One is like a map that shows the destination; the other is the journey itself, with all its wrong turns, pauses, and moments of doubt. Understanding the difference is no longer an academic exercise, because when humans rely on machines to do the thinking for them, they may be skipping the very steps that build their own ability to think.
A team of researchers from Macquarie University and other institutions has proposed a new way to measure intelligence that focuses entirely on this journey. They argue that true intelligence is constituted by the process—the iterative activity of trying, failing, revising, and learning—rather than the output itself. To test this, they developed a framework that looks at seven specific features of how a problem is solved. These features include the willingness to generate many attempts and discard the ones that fail, the ability to sit with uncertainty without rushing to a confident answer, and the capacity to learn from the resistance offered by the material or the environment. They also look at how a solver engages in dialogue, holds themselves accountable to others, and how the struggle itself shapes the person doing the work. The researchers suggest that for a machine to be truly equivalent to a human mind, it must not just produce the right answer, but must demonstrate these same seven features in its own way.
When the researchers applied this framework to current generative AI, they found a significant gap. These systems are trained on the "traces" of human thought: finished papers, published code, and edited images. They learn to reproduce the final product by sampling from a vast collection of these completed works. While the output often looks like reasoning or creativity, the internal activity that generates it in humans is largely absent. The machine does not struggle with the problem, nor does it experience the slow accrual of judgment that comes from years of practice. It simply predicts the next likely word or pixel based on patterns it has seen before. The researchers describe this as "weak equivalence": the machine matches the input and the output, but the process that connects them is missing or opaque. Even when these systems appear to "think" by showing their work, such as listing steps before an answer, this is often just a cosmetic display. The steps do not necessarily drive the computation; they are just another layer of output designed to look like reasoning.
The study explicitly argues against the idea that simply making these models larger or giving them more data will fix this problem. Adding more text to the training set only provides more finished traces, not more of the messy, iterative process that builds intelligence. The researchers also reject the notion that machines must copy the exact biological architecture of the human brain to be intelligent. They propose a middle ground called "strong equivalence." This does not require a machine to use the same algorithms as a human brain, which is neither possible nor desirable. Instead, it requires that the machine's behavior is organized by the same seven process features, even if the physical machinery is different. For example, a machine could demonstrate "engagement with uncertainty" not by feeling doubt, but by having a mechanism that flags ambiguous problems and refuses to give a confident answer until more information is gathered.
The implications for human users are just as critical as the technical diagnosis. The researchers point out that when a person uses a tool that performs the generative work for them, the capacity that work would have built is never formed. Evidence from studies on students using AI suggests that while they may perform better with the tool present, they often perform worse when it is removed, having skipped the cognitive effort required to learn. The danger is not just a loss of skill, but a failure of formation. The difficult, uncertain, and dialogical activities that shape a person's judgment and wisdom are the very things that AI tools are most tempted to bypass. If a tool resolves an ambiguity for a user, the user never learns to sit with that uncertainty. If a tool supplies the answer, the user never practices the trial and error needed to develop taste and judgment.
To address this, the authors propose a new approach to both designing AI and using it. On the machine side, they suggest building architectures that force the system to engage with the process, such as agents that must revise their approach when challenged by another agent, rather than just voting on an answer. On the human side, they advocate for "process-preserving" assistance. This means tools that supply the next question or a counterexample to extend the user's thinking, rather than supplying the answer. The goal is to keep the generative steps with the person, ensuring that the tool supports the activity of thinking rather than replacing it. The researchers also outline a method for "process audits," where both humans and machines are tested on the same tasks using these seven features as a scoring rubric. This would allow scientists to measure whether a machine is truly engaging in the process or just mimicking the shape of it, and whether human users are retaining their cognitive capacities when assisted by these tools.
Ultimately, the paper suggests that the difference between a system that produces intelligence and one that constitutes it is the difference between a record of an event and the event itself. As these tools become more powerful, the focus must shift from optimizing the output to engineering the process. The researchers do not claim to have solved the mystery of intelligence, nor do they say that current AI is useless. Instead, they offer a clear criterion for what is missing and a path forward for building systems that respect the process of thinking, both in silicon and in the human mind. By focusing on the activity rather than the result, we can ensure that as machines become more capable, they do not inadvertently erode the very human capacities they are meant to support.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.