← Latest papers
🤖 AI

Assessing the Business Process Modeling Competences of Large Language Models

This paper introduces the BEF4LLM framework to systematically evaluate large language models on BPMN generation across four quality dimensions, revealing that while LLMs excel in syntactic and pragmatic aspects, they still lag slightly behind human experts in semantic quality and validity, yet demonstrate competitive potential for future business process modeling applications.

Original authors: Chantale Lauer, Peter Pfeiffer, Alexander Rombach, Nijat Mehdiyev

Published 2026-01-30
📖 5 min read🧠 Deep dive

Original authors: Chantale Lauer, Peter Pfeiffer, Alexander Rombach, Nijat Mehdiyev

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to give a recipe to a very smart, but slightly literal, robot chef. You want the robot to draw a map of how a business works (like a flowchart for a factory or a bank) based on your written instructions. This map is called a BPMN model.

For years, drawing these maps has been a job for expensive experts who know both the business and the strict rules of the drawing language. But now, we have Large Language Models (LLMs)—the same kind of AI that writes emails and stories. The big question is: Can these AI chefs draw these complex business maps just as well as human experts?

This paper is a massive "cooking competition" to find out. Here is what they did and what they found, explained simply.

The Contest: The BEF4LLM Framework

The researchers built a new scoring system called BEF4LLM. Think of it as a judge's scorecard with four specific categories:

  1. Validity (The "Is it real?" check): Did the AI actually produce a file that the computer can open and read? If the AI writes gibberish or breaks the file format, it gets a zero here. This is the most basic hurdle.
  2. Syntactic Quality (The "Grammar" check): Did the AI follow the strict rules of the drawing language? (e.g., "Every loop must have a start and an end," "Arrows must point the right way").
  3. Pragmatic Quality (The "Readability" check): Is the map easy for a human to look at and understand? Is it too cluttered? Are the lines crossing everywhere? A perfect grammar map that looks like a spiderweb is useless.
  4. Semantic Quality (The "Meaning" check): Does the map actually tell the right story? If you said "The customer pays, then the item ships," does the map show that order, or did the AI get confused and swap them?

The Participants

They tested 17 different open-source AI models of various sizes.

  • Small models: Like a smart calculator.
  • Medium models: Like a well-read librarian.
  • Large models: Like a super-computer with a massive brain.

They gave these AIs 105 different business descriptions (like "A customer orders a pizza, pays, and waits for delivery") and asked them to draw the maps.

The Big Surprises (The Results)

1. Bigger isn't always better.
You might think the biggest, most expensive AI would win every time. But the paper found that bigger models didn't necessarily make better maps.

  • The Analogy: Imagine a giant, over-enthusiastic chef (the big AI) who adds so many extra ingredients and steps to the recipe that the dish becomes a mess, even if the flavors are technically correct. Sometimes, a smaller, more focused chef (a medium AI) made a cleaner, easier-to-read map.
  • The Catch: The biggest models were better at following the strict rules (Grammar) and getting the story right (Meaning), but they often made the maps too complicated to read (Readability).

2. The "Valid File" Problem.
This was the biggest hurdle. Many AIs, especially the smaller ones, failed to produce a file that the computer could actually open.

  • The Analogy: It's like the AI wrote a beautiful letter, but it forgot to put it in an envelope, or the envelope was sealed with glue that the mailman couldn't open. Even if the letter inside was perfect, it's useless if it can't be delivered. Only a few top-tier models could consistently produce "openable" files.

3. AI vs. Human Experts.
The researchers also had human experts draw maps for the same descriptions.

  • The Result: The AI and the humans were surprisingly close in overall skill.
    • AI Strengths: The AI was actually better at following the strict grammar rules and making the maps look neat and organized.
    • Human Strengths: Humans were slightly better at capturing the deep meaning and logic of the story.
    • The Verdict: The AI is no longer a clumsy beginner; it has reached a level where it can compete with a human professional, though it still struggles with the deepest nuances of the story.

4. The "Refinement" Loop.
The researchers gave the AI a second chance. If the AI made a mistake, they told it, "Hey, your file is broken, fix it," and let it try again. This helped a lot, but it wasn't a magic cure-all.

The Bottom Line

This paper tells us that AI is ready to help with business process modeling, but we have to be careful about which AI we pick.

  • Don't just pick the biggest model: Sometimes a medium-sized model is the "Goldilocks" choice—just right for the job.
  • Validity is key: If the AI can't even generate a file that opens, it doesn't matter how smart the map inside is.
  • The Future: We need to teach these AIs to be more consistent and to understand the "story" behind the business process better.

In short, the AI is a very talented apprentice who knows the rules perfectly and draws neat lines, but it still needs a human supervisor to make sure the story it's telling actually makes sense.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →