A Primer on Computational Semantics for Artificial Intelligence Systems
This paper provides a primer on computational semantics by exploring formal, grounded, and distributional theories of meaning to help readers understand how transformer-based language models learn and represent language compared to humans.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language is the most human thing we possess, a tool we use to share thoughts, feelings, and plans with one another. Yet, despite using it every day, most of us cannot explain how it actually works or what it truly means when we say a word like "red." We know the color when we see it, but we cannot easily define it without pointing to the world around us. This gap between using language and understanding its mechanics is the central puzzle facing scientists who study how machines learn to speak. For decades, researchers have tried to teach computers to understand meaning by feeding them logic, rules, or vast amounts of text. But a new perspective suggests that these machines are missing something fundamental: the physical experience of being alive. To understand why a computer might sound like a person but not think like one, we must look at how humans learn language as children and compare it to how artificial intelligence systems are built today.
This paper, written by Casey Kennington of Boise State University, serves as a guide through the different ways scientists have tried to define and compute meaning. The author argues that while modern artificial intelligence, specifically the powerful models behind tools like ChatGPT, can mimic human conversation with startling accuracy, they do not understand language in the way humans do. The paper explores three main ways researchers have tried to teach machines about meaning: using strict logic, connecting words to physical experiences, and analyzing how words appear next to each other in text. Kennington finds that while the third method—learning from text patterns—is currently the most successful at producing fluent speech, it leaves a critical hole. Machines trained only on text have never seen a red apple, felt the weight of a chair, or experienced the relief of drinking water. Without these physical and emotional connections, the words they process remain ungrounded symbols, lacking the deep, lived meaning that humans possess.
To understand the problem, one must first look at how humans learn. When a child learns the word "dog," they do not start with a dictionary definition. They learn through interaction. A parent points to a furry animal, the child looks at it, and the word is linked to the sight, the sound, and the feeling of the animal. This process relies on shared attention and physical presence. The child learns that words refer to real things in the world, and that these things have specific uses, such as sitting on a chair or kicking a ball. This connection between a word and the physical reality it describes is called grounding. For humans, meaning is not just a list of definitions; it is tied to our bodies, our senses, and our emotions. We know what "thirst" means because we have felt it, and we know what "chair" means because we have sat in one.
For a long time, computer scientists tried to teach machines to understand language using formal logic, similar to the math used in engineering. They treated words like variables in an equation, where "big" and "gray" and "elephant" could be combined to create a logical statement about a specific animal. This approach works well for computers because it is precise and unambiguous. However, it fails when it comes to the messy, real world. A computer following these rules can know that a "gray elephant" is an elephant that is gray, but it does not know what gray looks like or what an elephant feels like. It has the structure of the meaning but not the substance. The paper points out that this approach hits a wall because the machine has no way to know what the words actually refer to in the physical world.
Another approach, known as grounded semantics, attempts to fix this by connecting words to sensory data. In this view, a computer should learn what "red" means by looking at pictures of red objects, or what "kick" means by seeing a video of a foot hitting a ball. The idea is that if a machine can see and touch the world, it can build a dictionary of meanings that matches human experience. While this is a promising step, the paper notes that it is difficult to implement. Most current systems still rely heavily on text, and even when they use images, they often miss the full range of human experience, such as smell, touch, or the feeling of muscle memory. A machine might recognize a picture of a chair, but it does not know the relief of sitting down after a long walk.
The most dominant method in modern artificial intelligence is called distributional semantics. This approach is based on a simple idea: you can understand a word by looking at the company it keeps. If the word "apple" often appears near words like "red," "fruit," and "pie," the computer learns that "apple" is related to those concepts. By analyzing billions of sentences, these systems build a map where words with similar meanings are located close to each other. This method has led to the creation of large language models that can write stories, answer questions, and translate languages with incredible fluency. These models are trained on vast amounts of text from the internet, learning to predict the next word in a sentence based on the words that came before.
However, the paper argues that this success comes with a hidden cost. Because these models learn only from text, they never experience the world directly. They have never seen a sunset, tasted a strawberry, or felt the pain of a stubbed toe. They know that "apple" is related to "red" because they have seen those words together in books and articles, but they do not know what red looks like or what an apple tastes like. The author uses the example of asking a language model if it has ever seen an apple. The model will truthfully say no, because it has never seen anything. It has only processed the symbols. This creates a situation where the machine can talk about the world perfectly, but it does not understand the world. It is like a person who has read every book about swimming but has never touched water.
The paper also highlights the role of emotion and intention in human language. When we speak, we are not just exchanging data; we are expressing feelings, desires, and goals. A child learns to speak because they have a need to communicate with others, to get help, or to share joy. This drive is tied to our bodies and our emotions. Current language models do not have feelings or intentions. They do not get hungry, tired, or curious. They do not speak because they want to; they speak because they are programmed to predict the next word. The author suggests that without these internal drives and emotional connections, machines will always struggle to grasp the deeper, more nuanced meanings of language that humans take for granted.
The conclusion of the paper is that while artificial intelligence has made remarkable progress in processing language, it has not yet solved the problem of understanding meaning. The current models are powerful tools that can mimic human conversation, but they lack the embodied experience that gives words their true weight. To move forward, the author suggests that we need to combine the strengths of different approaches. We need systems that can not only analyze text but also interact with the physical world, feel emotions, and have intentions. Until machines can experience the world in the same way humans do, their understanding of language will remain incomplete. They will continue to be excellent at using words, but they will never truly know what those words mean.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.