← Latest papers
🤖 AI

Lomekwi: Resource-Bounded Tool Discovery in LLM Agents

This paper proposes a cognitive science-inspired framework that decomposes tool discovery into curiosity, recognition, and efficiency, revealing that recognition capabilities inversely scale with model size through analysis of combinatorial games and real-world task environments.

Original authors: Roshan Klein-Seetharaman, Daniel Wang, Andrew Xu

Published 2026-07-21
📖 4 min read☕ Coffee break read

Original authors: Roshan Klein-Seetharaman, Daniel Wang, Andrew Xu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a robot learn to survive in a new world. For a long time, scientists have been testing these robots by handing them a toolbox and asking, "Can you use this hammer to fix the wall?" or "Can you use this wrench to tighten that bolt?" If the robot fixes the wall, we say it's smart. But this is like testing a chef by only asking them to use a knife they were just handed, never seeing if they would ever think to grab one themselves. In the real world, and in the wild, being smart isn't just about using tools you already know; it's about looking at a pile of random junk and realizing, "Hey, if I stick these two things together, I can make a tool!" This paper dives into that missing piece: the moment of discovery. It asks a simple but tricky question: As robots (specifically, AI language models) get bigger and smarter, do they actually get better at inventing their own tools, or do they just get better at using the ones they are told about?

The researchers behind this study, working with a team from Yale and Sea12 Technologies, decided to stop just measuring if the robot wins the game. Instead, they broke the game down into three distinct steps, like a recipe for invention. First, there's Curiosity: Does the robot even look for the parts it needs? Second, there's Recognition: Once it has the parts, does it realize, "Oh, I can combine these to make a machine"? And third, there's Efficiency: Once it builds the machine, can it actually use it to win? They built a special digital playground called "Lomekwi" (named after a real place where ancient humans first used stone tools) to test this. In this game, the robot has to open locked doors. It can either try to brute-force every single door until it finds the keys (a slow, boring grind), or it can find specific shards, combine them into a machine, and use that machine to open the doors instantly.

Here is where things get weird and counter-intuitive. You might expect that the bigger, more powerful AI models would be the best at everything: finding the parts, realizing how to build the machine, and using it. But the paper suggests something surprising happens. While the biggest models are great at the "grind" (solving problems by hand) and very efficient once they have a tool, they seem to lose their Recognition. It's as if the smartest models get so confident in their own ability to solve problems by brute force that they stop looking for shortcuts. They ignore the shiny new machine they could build because they think, "I can just do this myself."

In fact, the researchers found that as the models got larger, they actually became worse at spotting the opportunity to build a tool. In one experiment, a smaller model was much more likely to discover the tool recipe than a massive, super-smart model. The big models were so busy trying to solve the puzzle with their own "muscle" that they missed the "brain" solution. However, the paper also suggests this might be part of a bigger pattern. When the researchers gave the big models a little hint that a tool might exist, the big models suddenly woke up and started looking again, showing that they weren't incapable, just overconfident.

So, what does this all mean? It turns out that being a "tool user" and being a "tool inventor" are two very different skills. A model can be a genius at using a tool once it's handed to it, but a total flop at realizing it needs one in the first place. The paper suggests that as AI gets bigger, we might be seeing a dip in this creative spark of discovery, not because the AI is getting dumber, but because it's getting too sure of its own ability to grind through problems without help. It's a reminder that sometimes, the smartest thing you can do is stop trying to do it all yourself and look for a tool to help you out.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →