← Latest papers
💻 computer science

A Comparative Study of Competency Question Elicitation Methods from Ontology Requirements

This paper presents an empirical comparative evaluation of manual, pattern-based, and LLM-driven Competency Question (CQ) elicitation methods, revealing that while LLMs can effectively generate initial CQs, the resulting outputs vary significantly in quality and typically require human refinement before being used for ontology modeling.

Original authors: Reham Alharbi, Valentina Tamma, Terry R. Payne, Jacopo de Berardinis

Published 2026-04-03
📖 5 min read🧠 Deep dive

Original authors: Reham Alharbi, Valentina Tamma, Terry R. Payne, Jacopo de Berardinis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are building a massive, digital library for a museum's collection of music memorabilia. Before you can organize the books, records, and photos, you need a blueprint. In the world of computer science, this blueprint is called an Ontology. It's a map that tells the computer how everything is connected (e.g., "This song was written by that artist").

But how do you know if your blueprint is good? You need a checklist of questions that the computer must be able to answer. These are called Competency Questions (CQs).

  • Example: "Who wrote this song?" or "Which albums were released in 1990?"

If the questions are vague, the blueprint is useless. If they are too complicated, the computer gets confused.

The Big Question: Who is the best "Question Writer"?

The authors of this paper wanted to find out: Who is better at writing these checklist questions?
They compared three different "writers":

  1. The Human Expert: A seasoned computer scientist who knows the rules of the game.
  2. The Template Filler: A method that uses pre-made sentence structures (like a "Mad Libs" game) where you just plug in the words.
  3. The AI Robot: A super-smart Large Language Model (like GPT-4 or Gemini) that writes questions automatically.

They gave all three writers the exact same story about the museum's music collection and asked them to write their lists of questions. Then, they graded the results.

The Results: The Race for the Best Questions

Here is what they found, using some fun analogies:

1. The Human Expert (The Master Chef) 🧑‍🍳

  • Performance: The humans wrote the best questions.
  • Why? Their questions were clear, easy to read, and hit the nail on the head. They didn't just ask what was written in the story; they asked what was implied.
  • Analogy: Think of the Human Expert as a Master Chef. They know exactly what ingredients are needed to make a perfect dish. They don't just follow a recipe; they know why you need salt and how to balance the flavors. Their questions were short, punchy, and exactly what the computer needed to understand the museum.

2. The Template Filler (The Assembly Line Worker) 🏭

  • Performance: This method was okay, but a bit stiff and repetitive.
  • Why? It used pre-made patterns (like "Which [Item] has [Feature]?"). It worked, but it felt robotic and sometimes missed the nuance of the story.
  • Analogy: This is like an Assembly Line Worker stamping out parts. They are fast and consistent, but they can't improvise. If the story needed a creative twist, the template just didn't fit, leaving some gaps in the blueprint.

3. The AI Robot (The Over-Enthusiastic Intern) 🤖

  • Performance: The AI wrote a lot of questions, but they were often too long, too complicated, and sometimes missed the point.
  • Why? The AI tried too hard. It used big words, complex sentence structures, and asked questions that were technically correct but hard for a human to understand. It also missed the "hidden" requirements that the human experts caught.
  • Analogy: The AI is like an Over-Enthusiastic Intern who wants to impress the boss. They write a 10-page essay when a one-sentence answer would do. They use fancy vocabulary to sound smart, but they sometimes miss the simple, practical needs of the job. They also tend to go off on tangents, asking questions that aren't really in the story.

The "Secret Sauce" Discovery

The study found something really interesting about Relevance:

  • Humans asked questions about things that weren't explicitly written in the story but were obviously necessary (e.g., "What is the file format?" even if the story didn't say "file format" explicitly). They used their brain to fill in the blanks.
  • AI stuck strictly to what was written. If the story didn't say it, the AI didn't ask about it. It lacked the "common sense" or "domain knowledge" that humans have.

The Verdict

Can we just let the AI do all the work?
Not yet.

The study concludes that while AI is a great tool for brainstorming (getting the ball rolling), it cannot replace the human expert.

  • AI is like a rough draft. It gives you a pile of ideas, but they need editing.
  • Humans are the editors. They need to take the AI's messy, complex draft, simplify it, fix the grammar, and add the "common sense" questions that the AI missed.

The Takeaway

If you are building a digital map of knowledge, don't just ask the robot to do it alone. Use the robot to get ideas, but let the human expert refine them. The best results come from a hybrid team: the speed of the AI combined with the wisdom of the human.

In short: The AI is a fast typist with a big vocabulary, but the Human is the one who actually knows what the story is really about.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →