GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection
This paper introduces GBU-Palm, a large-scale multimodal video dataset and benchmark featuring 21,326 synchronized RGB-NIR samples across diverse environments, to systematically evaluate and reveal the limitations of current palm presentation attack detection methods under cross-environment conditions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to unlock your phone with a picture of your hand instead of your actual hand. This is the world of biometrics, where computers recognize us by our unique body parts, like our fingerprints or the veins in our palms. It's supposed to be super secure, but just like a thief can pick a lock with a fake key, hackers can trick these systems with "presentation attacks." They might hold up a high-resolution photo of your palm (a Print attack) or play a video of your hand on a screen (a Replay attack). To stop this, scientists build "Presentation Attack Detection" (PAD) systems—digital bouncers that check if the hand in front of the camera is real or a fake.
For a long time, these digital bouncers were trained mostly on still pictures, like looking at a single snapshot of a suspect. But in the real world, hands move, light changes, and screens flicker. To build a truly smart bouncer, we need to see the whole movie, not just a photo. We also need to know if using two different "eyes" for the camera—one seeing normal colors (RGB) and one seeing invisible infrared light (NIR)—actually helps the bouncer spot the fake, or if it just adds confusion. The big question is: Can we build a system that stays sharp even when the lighting changes from a sunny park to a dark room, and does having two types of vision always make it smarter?
Enter GBU-Palm, a massive new playground for testing these digital bouncers. Think of this dataset as a giant, carefully organized "attack simulator" created by researchers. Instead of just a few photos, they recorded 21,326 videos from 105 different people (covering 210 palms). They didn't just film in one spot; they set up cameras in six different environments, ranging from bright, sunny outdoors to dim, tricky indoor lighting. They filmed real hands, hands printed on paper (in color, black-and-white, and different textures), and hands being replayed on various electronic screens. Crucially, they synchronized two cameras for thousands of these clips: one seeing the world in normal color and the other seeing it in near-infrared (NIR), which reveals details invisible to the naked eye.
The researchers used this massive library to test four different types of "brain" architectures (the computer models that do the thinking). They asked: "If we train a model in a sunny room, will it still work in a dark room?" and "Does giving the model both color and infrared vision actually make it better?"
The results were surprising and taught them a lot about how these systems fail. First, they found that not all computer brains are built the same. Some models handled the change from a sunny day to a dark room like a champ, while others completely lost their cool, with their accuracy dropping significantly. This proved that "environment shift" isn't just a general difficulty; it hits different models in very specific ways.
Second, the idea that "more vision is always better" turned out to be a myth. The researchers discovered that adding the infrared (NIR) camera didn't automatically make every model smarter. For some models, the extra infrared data was like a superpower, helping them spot fakes much better. But for others, the extra data actually made them worse at their job, or didn't help at all. It turns out, just having the data isn't enough; the model has to know how to use it.
Finally, they looked at how the models made mistakes. They broke down errors into four categories: accepting a real hand, rejecting a fake hand, accepting a fake hand (a security disaster), and rejecting a real hand (an annoying user error). They found that when the environment changed, some models started letting fakes in (security risk), while others started locking out real people (usability risk). They also tested if the models were actually watching the movement in the video or just the still frames. Some models relied heavily on the order of the frames (the motion), while others seemed to ignore the timing entirely, treating the video like a stack of photos.
In short, GBU-Palm isn't just a new dataset; it's a reality check. It shows that building a perfect palm scanner isn't about just throwing more data or more cameras at the problem. It requires carefully matching the right type of computer brain to the right kind of data, because what works in a sunny office might fail miserably in a shadowy hallway. The researchers are releasing this dataset to the public so that everyone can build better, more reliable security systems that don't get fooled by a clever photo or a tricky screen.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.