← Latest papers
⚡ electrical engineering

GameScope: A Multi-Attribute, Multi-Codec Benchmark Dataset for Gaming Video Quality Assessment

This paper introduces GameScope, the largest and most diverse gaming video quality benchmark dataset to date, which features 4,048 samples across H.264, H.265, and AV1 codecs with both user and professional content, detailed quality attributes, and extensive human annotations to advance multi-attribute, multi-codec video quality assessment.

Original authors: Rajesh Sureddi, Shreshth Saini, Avinab Saha, Alan C. Bovik

Published 2026-05-05
📖 4 min read☕ Coffee break read

Original authors: Rajesh Sureddi, Shreshth Saini, Avinab Saha, Alan C. Bovik

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a film critic, but instead of reviewing movies, you are reviewing video games being streamed live on the internet. The problem is that every streaming platform (like Twitch or YouTube) uses a different "compression tool" (a codec) to shrink the video so it can travel fast over the internet. Sometimes these tools work great, and sometimes they turn a crisp, beautiful game into a blurry, blocky mess.

Until now, scientists trying to build computers that can "see" and rate this video quality have been working with very small, limited libraries of examples. It's like trying to teach a dog to recognize all animals in the world, but you only show it pictures of three different types of cats.

Enter "GameScope": The Ultimate Video Game Taste Test

The authors of this paper have built the biggest, most diverse library of gaming video samples ever created. Think of it as a massive, organized "tasting menu" for video quality.

Here is what makes this "menu" special:

  • A Mix of Home Cooking and Fine Dining: They didn't just use professional studio recordings (PGC); they also included clips made by regular gamers (UGC). This covers everything from a pro player's perfect stream to a YouTuber's video with their face in the corner and custom logos.
  • The Three Big Compression Tools: They tested the videos using the three most popular "shrinking tools" used today: H.264, H.265, and AV1. It's like testing how a cake tastes when baked in three different ovens.
  • Thousands of Samples: They created over 4,000 video clips. To get the "true" taste, they didn't just ask one person; they asked an average of 37 different people to rate each video.
  • More Than Just a Score: Instead of just asking, "Is this good or bad?", they asked specific questions:
    • Clarity: Is it blurry?
    • Artifacts: Are there weird blocky squares?
    • Immersion: Does the visual quality ruin the feeling of being inside the game?

How They Did It

They set up a giant online "taste test" using a platform called Amazon Mechanical Turk. Imagine a digital room where thousands of volunteers watched these clips on their computers. To make sure the results were fair:

  • They blocked people from using phones or tablets (so everyone watched on a similar screen).
  • They gave the volunteers a quiz to make sure they were paying attention.
  • They used a special math trick (called SUREAL) to filter out people who were just guessing or being inconsistent, ensuring the final "score" was as accurate as possible.

What They Learned

Once they had all these human ratings, they used them as a "gold standard" to test the smartest computer programs (AI) currently available for judging video quality.

  • The Old Guard: Older computer programs struggled. They were like a student who studied hard but only learned from a tiny textbook; they couldn't handle the variety of the new dataset.
  • The New Contenders: The authors tested some very new, powerful AI models that can understand both images and language (Vision-Language Models).
  • The Winner: One model, named Qwen3-VL-4B, performed the best. It was like a super-critic that could not only give a score but also explain why the video looked bad (e.g., "too much pixelation") in a single glance. It outperformed all the previous methods.

The Bottom Line

The paper introduces a massive new dataset called GameScope. It is a public resource that allows researchers to finally train and test their computer vision systems on a wide variety of real-world gaming scenarios. The authors found that the newest generation of AI models is much better at understanding gaming video quality than the older tools, especially when they can look at the video and "read" the specific problems with it.

They have made this entire dataset available for anyone to use, hoping it will help build better systems for judging how good our favorite games look when we stream them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →