A Taxonomy of Construction Task Activities for Robot Workers
This paper introduces TARCAT, a human-interpretable taxonomy of 41 construction action primitives derived from occupational data and instructional videos, designed to standardize the analysis of human work and enable the development of general-purpose robot workers through reusable, composable skills.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
To understand the challenge of teaching a robot to build a house, one must first understand the difference between a machine that follows a single command and a worker who understands a job. Modern robots are often like highly skilled specialists who can perform one perfect motion, such as welding a seam or painting a single wall, but they struggle when the environment changes or when a task requires a sequence of different actions. In recent years, a new type of artificial intelligence has emerged that can see an image, understand a sentence, and decide on a physical action, offering a path toward machines that can adapt to new situations. However, for construction, where materials, tools, and sites vary wildly, simply having a smart brain is not enough. The robot needs a clear, shared language to describe exactly what it must do, breaking down complex human jobs into small, learnable pieces that can be recombined. Without this vocabulary, a robot cannot know what it is capable of learning or what it is missing.
This is the problem a team of researchers at the University of California, Irvine, set out to solve with a new framework called TARCAT. Rather than guessing what a construction robot should do, the team started with the people who actually do the work. They analyzed the official job descriptions for seven of the most common construction trades in the United States, including carpenters, electricians, roofers, and laborers. These descriptions, which come from a government database used to track the economy, list hundreds of specific tasks that humans perform daily. The researchers then watched thirty instructional videos of these workers in action, carefully mapping every movement and decision the humans made to the official job descriptions. From this deep dive into real-world labor, they extracted a precise inventory of forty-one basic actions, which they call "primitives."
These primitives are the building blocks of construction work. They are not tied to a specific trade or a specific tool, but rather to the fundamental physical or mental capability required to get something done. For example, the action of "measuring" is a primitive that applies whether a worker is using a tape measure to check a wall or a pressure gauge to check a pipe. Similarly, "communicating" is a primitive that covers both talking to a coworker and writing down a report. The researchers organized these forty-one actions into three main categories: information processing, which involves understanding plans and checking conditions; communication, which covers talking and signaling; and tool use, which encompasses everything from the fine movements of fingers to the heavy lifting of the whole body. By defining these actions this way, the team created a system where a robot can learn a single skill, like "tightening a screw," and understand that it is made of smaller, reusable parts like "grabbing the tool," "positioning the hand," and "rotating the wrist."
The power of this system lies in how these small actions are combined. The researchers showed that by stringing these primitives together in a specific order, a robot can form a "skill" that accomplishes a sub-goal, such as disconnecting a wire or attaching a tile. These skills can then be mixed and matched to create complex procedures for entire construction tasks. To prove that this abstract list of actions could actually work in the real world, the team tested four of these tool-use primitives on a physical robot arm equipped with a dexterous hand. They successfully programmed the machine to position a power screwdriver, place a hammer in a box, hammer a nail, and brush objects aside. These demonstrations confirmed that the taxonomy is not just a theoretical list but a set of instructions that a machine can physically execute.
The ultimate goal of TARCAT is to provide a common language that allows robots to learn from human demonstrations and to help software agents figure out how to build new skills. If a robot needs to learn how to install a roof, it does not need to be taught from scratch; instead, it can retrieve the known skills for measuring, cutting, and fastening, and compose them into a new sequence. The researchers acknowledge that their current work covers only seven occupations and a limited number of video examples, so the system is not yet complete. However, by grounding their work in the actual tasks of human workers, they have created a stable foundation for developing general-purpose construction robots that can eventually handle the messy, varied, and complex reality of building sites.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.