← Latest papers
💻 computer science

Spreadsheet Modeling Experiments Using GPTs on Small Problem Statements and the Wall Task

This paper evaluates the capabilities of GPT-based tools, specifically Excel AI, in generating reusable spreadsheet models, finding that while they can produce well-structured drafts, their inconsistency, lack of reproducibility, and the resulting need for skilled user verification currently limit their reliability for professional use.

Original authors: Thomas A. Grossman, Yuan Chen, Sopiko Datuashvili

Published 2026-04-29
📖 6 min read🧠 Deep dive

Original authors: Thomas A. Grossman, Yuan Chen, Sopiko Datuashvili

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a very smart, but sometimes scatterbrained, robot assistant how to build a spreadsheet. You give it a simple description of a math problem, and you hope it hands you back a perfectly organized Excel file that you can use forever.

This paper is a report card on how well that robot assistant actually did. The researchers, Thomas Grossman and his team, put five different "AI helpers" to the test to see which one could build a reusable spreadsheet model best. They picked the winner, a tool called Excel AI, and then put it through a series of stress tests.

Here is the breakdown of their findings, using some everyday analogies:

1. The Goal: Building a "Lego Set," Not a "Sandcastle"

The researchers wanted the AI to build a spreadsheet that is reusable.

  • The Bad Way (Sandcastle): Imagine the AI writes the numbers directly into the cells (like "10" and "12"). If you want to change the room size later, you have to erase and rewrite everything. This is like building a sandcastle; it looks good for a moment, but you can't easily rebuild it.
  • The Good Way (Lego Set): The researchers wanted the AI to put the numbers in special "input" boxes and use formulas (like =A1*B1) for the rest. This is like a Lego set. You can swap out the bricks (the input numbers), and the whole structure automatically updates. This is what they call ERFR (Essential Requirements for Reusability).

2. The Screening: Finding the Best Assistant

The team tried out five different AI tools. It was like hiring five different contractors to build a shed.

  • Some contractors handed back broken blueprints.
  • Some gave them a picture of a shed instead of the shed itself.
  • One contractor, Excel AI, was the most consistent. It usually handed over a clean, downloadable file with the right structure. So, the researchers decided to focus all their future testing on this one.

3. The Experiments: How the Robot Handled Different Tasks

The researchers gave Excel AI eight different simple math problems to solve, like "Calculate the area of a room" or "Figure out the cost of apples over a month." They tested three specific things:

  • The "Recipe" Test (Data vs. Variables):

    • Scenario A: "Here are the numbers: 10 and 12."
    • Scenario B: "Here are the names of the numbers: Length and Width."
    • Result: The robot handled both perfectly. It knew to leave blank boxes for the user to fill in later if no numbers were given.
  • The "Calendar" Test (30 Days vs. A Month):

    • Scenario A: "Do this for 30 days."
    • Scenario B: "Do this for a month."
    • Result: This is where the robot stumbled. When told "30 days," it built a perfect list of 30 rows. But when told "a month," it got confused. Sometimes it built a list for the current month (which has 31 days or fewer), and sometimes it put actual numbers in the cells instead of formulas. It was like a chef who can follow a recipe for "30 minutes" perfectly, but gets confused when you say "an hour" and starts guessing.
  • The "Nonsense Word" Test:

    • They used a made-up word (like "Zorp") instead of a real object (like "Apple").
    • Result: The robot didn't care at all. It treated "Zorp" exactly like "Apple." This was a success.

4. The Big Problem: The "Magic 8-Ball" Effect

The biggest issue the researchers found is reproducibility.

  • If you ask the robot the exact same question twice, you might get two different answers.
  • Analogy: Imagine asking a human assistant, "Make me a sandwich." The first time, they give you a perfect turkey sandwich. The second time, you ask the exact same thing, and they give you a peanut butter sandwich with the crusts cut off.
  • In the "Wall Task" (a slightly more complex problem about building a wall), the robot sometimes built a perfect, professional-grade model. Other times, it built a broken model with missing formulas that crashed Excel. You never knew which version you were going to get.

5. The "Confidence" and "Workflow" Problems

The paper concludes with two major headaches for anyone trying to use this technology:

  • The Problem of Confidence:
    If a robot hands you a spreadsheet, how do you know it's right?

    • Analogy: If a student hands you a math test, you can check the answers. But if the student is a robot that might be hallucinating, you have to be an expert mathematician yourself to verify its work. You can't just trust it. The paper says you need a skilled human to double-check the robot's work, which defeats the purpose of using the robot to save time.
  • The Problem of Workflow:
    Some people argue, "Maybe the robot should just give us a 'rough draft' that we can fix."

    • Analogy: It's like asking a robot to write a first draft of a novel. If the draft is full of plot holes and typos, is it faster to fix it or to write the book from scratch? The researchers found that for complex spreadsheets, fixing the robot's messy draft often takes just as much time (or more) than doing it yourself.

The Final Verdict

The paper's conclusion is cautious.

  • The Good: The AI can sometimes build beautiful, professional-looking spreadsheets from simple instructions.
  • The Bad: It is inconsistent, unreliable, and often unpredictable.
  • The Reality: Right now, these tools are not ready for prime time. They are not reliable enough for professional use where mistakes can cost money. They are best viewed as a "drafting assistant" that requires a skilled human to supervise, verify, and fix the work.

In short: The robot is a talented apprentice who sometimes builds a masterpiece and sometimes builds a house of cards. Until it learns to be consistent, you still need a master builder (a human) to hold the clipboard.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →