← Latest papers
💻 computer science

Using profiles of cognitive capability to assess AI suitability for workplace tasks

This paper introduces a framework that profiles both AI systems and workplace tasks using a shared set of core cognitive capabilities to overcome the limitations of aggregate benchmarks and provide a dynamic, comparative tool for determining optimal human-AI task allocation.

Original authors: Jonathan Prunty, Marko Tešić, Patrick Quinn, José Hernández-Orallo, Lucy Cheke

Published 2026-08-27
📖 6 min read🧠 Deep dive

Original authors: Jonathan Prunty, Marko Tešić, Patrick Quinn, José Hernández-Orallo, Lucy Cheke

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern workplace, a quiet but urgent question is reshaping how organizations think about their future: which jobs can be safely handed over to artificial intelligence, and which must remain in human hands? For years, the answer has been elusive. Companies often rely on broad, aggregate test scores to judge a machine's intelligence, much like a school principal might judge a student's entire year based on a single final exam grade. These scores offer a snapshot of general performance, but they rarely explain why a system succeeds or fails, nor do they predict how it will handle the messy, unpredictable complexity of real-world work. When a machine fails in a controlled test, it is often a mystery; when it fails in a real office or factory, the consequences can be costly and dangerous. The core challenge is that artificial intelligence does not think like a human. Its strengths and weaknesses are jagged and uneven; it might possess a vast library of facts but struggle with the simple logic of planning a sequence of events. To deploy these systems safely, leaders need a way to map the specific cognitive demands of a job against the specific cognitive abilities of a machine, moving beyond guesswork to a clear, evidence-based assessment.

A new technical report from researchers at the University of Cambridge and other institutions offers a solution to this problem. Instead of asking whether a machine can perform a specific task, the team developed a pipeline to profile the underlying mental abilities required for that task and compare them directly to the machine's capabilities. They began by breaking down the complex world of artificial intelligence testing into eighteen core cognitive skills, such as memory, reasoning, planning, and social understanding. Using a massive collection of existing test questions, they annotated each one to determine exactly which of these mental skills it required and how difficult it was. This created a detailed "demand map" of the tests. They then used this map to infer the true cognitive profile of six different AI systems, not just by looking at their final scores, but by analyzing how they performed across the different types of mental challenges. The result was a profile for each machine that showed exactly where its abilities were strong and where they were fragile, revealing that while the systems varied in overall power, they shared a remarkably similar shape of strengths and weaknesses.

To make these profiles useful for real-world decisions, the researchers needed to know what human workers actually need to do their jobs. They surveyed 410 employees across six different occupational fields, ranging from warehouse logistics and manufacturing to customer service and data analysis. Rather than asking these workers to predict how well a machine would do, the researchers asked them to identify the most critical mental skills required for their daily tasks. The employees ranked the importance of skills like problem-solving, planning, and communication for their specific roles. This process revealed a "cognitive core" common to almost all jobs: a heavy reliance on memory, language, and the ability to plan ahead. While different jobs emphasized different secondary skills—such as spatial reasoning for warehouse workers or social understanding for sales staff—the fundamental mental machinery required for work was surprisingly consistent across the board.

When the researchers overlaid the AI profiles onto the job requirements, a clear picture of suitability emerged. The study found that the most advanced AI systems were best suited for tasks that relied heavily on language, factual knowledge, and social interaction, areas where current machines excel. However, they consistently struggled with tasks requiring complex planning, reasoning about physical objects, or understanding cause-and-effect relationships in dynamic environments. Crucially, the researchers discovered that the differences between the various AI models were far smaller than the differences between the different types of cognitive skills. In other words, one AI system was not radically different from another in its overall "personality"; rather, all of them were strong in some areas and weak in others. The most capable systems showed only modest improvements in planning and control compared to their peers, suggesting that while machines are getting better at organizing information, they are still far from mastering the kind of flexible, goal-directed behavior that humans take for granted.

The framework also provided a way to calculate a "suitability score" for any specific job, allowing organizations to see which tasks are ready for automation and which are not. By combining the machine's capability profile with the importance of the task to the business, the system could highlight high-priority opportunities where AI would be both effective and valuable. For instance, tasks like data analysis and administrative research emerged as strong candidates for automation, while roles requiring deep interpersonal rapport or complex physical reasoning remained firmly in the human domain. The researchers tested this approach on a single company, showing how the same method could be tailored to a specific organization's unique needs, moving from broad industry trends to the specific duties of individual employees.

The study explicitly argues against the idea that a single, high test score is enough to guarantee a machine's reliability in the workplace. It demonstrates that a system with a high average score can still fail catastrophically if it lacks a specific, hidden skill required for a particular job. The researchers also caution that their findings are based on the current generation of AI models and the specific tests available today. As new machines are built and new tests are developed, the profiles will need to be updated. However, the method itself remains robust. By separating the measurement of what a machine can do from the measurement of what a job requires, the framework provides a stable, scientific language for discussing the future of work. It suggests that the path forward is not about finding a machine that can do everything, but about carefully matching the specific, uneven capabilities of our current tools to the specific, uneven demands of human labor, ensuring that automation is deployed where it will truly succeed.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →