← Latest papers
💬 NLP

Building a Multimodal Dataset of Academic Paper for Keyword Extraction

This study addresses the scarcity of multimodal resources for keyword extraction by constructing a new dataset of 1,000 academic papers containing text, images, and audio, demonstrating that fusing information from these diverse modalities significantly improves extraction performance compared to using text alone.

Original authors: Jingyu Zhang, Xinyi Yan, Yi Xiang, Yingyi Zhang, Chengzhi Zhang

Published 2026-07-01
📖 4 min read☕ Coffee break read

Original authors: Jingyu Zhang, Xinyi Yan, Yi Xiang, Yingyi Zhang, Chengzhi Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to find the most important ingredients in a giant, complex recipe book. Usually, when people try to pick out the key ingredients (which researchers call "keywords"), they only read the written text of the recipe. They ignore the pictures of the finished dish or the audio recording of the chef explaining the steps.

This paper argues that by ignoring the pictures and the audio, we are missing out on a lot of flavor. The authors built a new "kitchen" (a dataset) to test if looking at the whole meal—text, images, and audio together—helps us identify the key ingredients better.

Here is a simple breakdown of what they did and what they found:

1. The Problem: A One-Sided Conversation

For a long time, computers have been very good at reading text to find keywords. But in the real world, information isn't just text. It's a mix of:

  • Text: The actual words on the page.
  • Images: Slides, charts, and photos.
  • Audio: The sound of a speaker talking.

The authors say that by only listening to the "text" part of the conversation, we miss the clues hidden in the pictures and the voice. Also, there wasn't really a "recipe book" available that had all three of these things lined up together for researchers to study.

2. The Solution: Building a New "Recipe Book"

To fix this, the team created a brand-new dataset called a Multimodal Dataset.

  • What's inside? They gathered 1,000 samples from academic conferences.
  • The ingredients: For every single sample, they collected:
    1. The written paper (the text).
    2. The presentation slides (the images).
    3. The recording of the speaker (the audio).
    4. The list of keywords the authors originally chose (the "answer key").

Think of this like creating a study guide where every page has the essay, the diagrams, and a recording of the teacher explaining it, all perfectly synced up.

3. The Experiment: The Taste Test

Once they had their dataset, they ran a "taste test" using different computer programs (models) to see which method could find the keywords best. They tested three scenarios:

  1. The Text-Only Diet: Feeding the computer only the written paper.
  2. The Audio & Image Diet: Feeding the computer only the text transcribed from the audio or the text recognized from the images (using OCR, which is like a robot reading a sign).
  3. The Buffet: Feeding the computer a mix of all three text sources combined.

4. The Results: What Worked Best?

Here is what the "taste test" revealed:

  • The Written Paper is the Star: The text from the actual academic papers was the most reliable source. It had the richest information, so the computers did the best job finding keywords here.
  • Audio is a Good Side Dish: The text transcribed from the speaker's voice was helpful, though not quite as good as the written paper.
  • Images are the Tricky Part: The text pulled from the slides (images) performed the worst. The authors explain this is like trying to read a menu written on a blurry, crumpled napkin; the computer often got confused by the background noise or bad lighting, making the text hard to read.
  • The Power of the Buffet (Fusion): The biggest discovery was that combining all three sources worked better than using just one. Even though the image text was messy, when the computer looked at the paper, the audio, and the image text all at once, it could fill in the gaps. If a keyword was missing from the paper but mentioned in the audio, the "Buffet" method caught it.

5. The Takeaway

The main lesson from this paper is that while the written text is the most important ingredient, ignoring the other parts of the presentation (the voice and the slides) leaves information on the table.

By building this new dataset and proving that mixing these different types of information together improves the computer's ability to find keywords, the authors have given researchers a new tool. It's like upgrading from reading a recipe in the dark to reading it with a flashlight, a picture of the dish, and a recording of the chef all at once.

In short: They built a new library of mixed-media academic papers and proved that when computers read the text, listen to the audio, and look at the slides all together, they get a much clearer picture of what the paper is actually about.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →