A Comparative Study of Faster R-CNN, Mask R-CNN, and YOLOv8 for Object Detection and Instance Segmentation, with a Custom-Class Transfer-Learning Case Study
This paper provides a theoretical and qualitative comparative analysis of Faster R-CNN, Mask R-CNN, and YOLOv8 for object detection and instance segmentation, demonstrating through a custom military vehicle case study that lightweight single-stage detectors can effectively adapt to novel classes via transfer learning while outperforming two-stage models trained solely on general datasets, all within the context of evolving real-time detection architectures.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of computer vision, machines are learning to see the world not just as a collection of colors and shapes, but as a scene filled with distinct objects. This field, known as object detection, is the technology that allows a self-driving car to recognize a pedestrian, a security camera to spot an intruder, or a medical scanner to identify a tumor. For a machine to do this, it must perform two difficult tasks simultaneously: it must find where an object is located within a picture, and it must decide exactly what that object is. For years, researchers have debated the best way to teach a computer to do this. Some approaches act like a careful inspector, scanning an image slowly and methodically to ensure every detail is correct, while others act like a quick glance, processing the entire scene in a single pass to prioritize speed. The central question for engineers is how to balance this trade-off between being perfectly accurate and being fast enough to work in real time.
A recent study by Muhammad Usman Aslam at the University of West Florida explores this balance by testing three different computer systems designed to see the world. The research compares two established methods that work in stages, one of which can also draw precise outlines around objects, against a newer, faster method that processes images all at once. The study does more than just compare these systems on standard test images; it tackles a practical problem that standard tests often miss: what happens when a machine encounters an object it has never been taught to recognize? To answer this, the researcher trained a fast, modern system to identify military armored vehicles, a category of objects that does not exist in the standard library of images used to train most computer vision models.
The investigation began by examining how these three systems, known as Faster R-CNN, Mask R-CNN, and YOLOv8, are built. The first two systems operate like a two-step process. First, they scan an image to find regions that might contain an object, and then they analyze those specific regions to decide what the object is and where its edges are. Mask R-CNN adds a third step, allowing it to draw a pixel-perfect outline around each object, separating it from the background and from other overlapping objects. These methods are known for high accuracy but require more computing power and time. The third system, YOLOv8, takes a different approach. It looks at the entire image in a single pass, predicting the location and type of every object at once. This design is built for speed, making it ideal for real-time applications like video surveillance, though it has historically been considered slightly less precise on very small or crowded objects.
To test these systems in the real world, the researcher used a standard set of images containing common items like people, cars, and animals to see how well the pre-trained systems performed. When the Faster R-CNN system was shown images of beetles, it correctly identified them but also produced extra, lower-confidence guesses for "insects" in the same spots, showing that the system sometimes struggles when different categories of objects overlap or look very similar. When shown a picture of a bird, the system hesitated between calling it a "bird," an "animal," or a specific type of bird, revealing that its vocabulary is limited to the 80 specific categories it was originally taught. The Mask R-CNN system performed impressively on standard images, successfully separating overlapping zebras and people in a crowd with precise outlines, but it required significantly more computing resources to do so.
The most significant part of the study, however, was the attempt to teach these systems to see something they had never seen before: tanks. Standard computer vision models are "closed-vocabulary" systems, meaning they can only recognize the specific list of objects they were trained on. When the researcher showed the standard Faster R-CNN and Mask R-CNN models pictures of military armored vehicles, the machines failed to identify them as tanks. Instead, they forced the images into the closest categories they knew, labeling the tanks as "trucks" or "cars." This was not a failure of the machine's ability to see the shape of the vehicle, but a limitation of its vocabulary; it simply did not have the word "tank" in its dictionary.
To solve this, the researcher turned to the YOLOv8 system and used a technique called transfer learning. Instead of starting from scratch, the system was given a small, custom dataset of just twenty images of tanks, which had been manually labeled by a human. The system was then fine-tuned on this small set of images. The result was striking. The YOLOv8 system, which had been adapted to this new class, successfully identified tanks in museum displays and outdoor photographs with high confidence. It even recognized a moving armored vehicle in a field, despite the blur and distance, proving that a lightweight, fast system could be quickly adapted to a new, specific job with very little data.
The study concludes that while the older, two-stage systems remain excellent for general-purpose tasks where high precision is needed for known objects, they are rigid when it comes to new categories. They cannot recognize what they have not been explicitly taught. In contrast, the single-stage YOLOv8 system demonstrated a flexibility that is highly valuable for practical applications. It showed that with a small amount of human effort to label a few new images, a fast detector can be extended to recognize entirely new types of objects. This suggests that for many real-world problems, where the goal is to detect specific, custom items rather than just general categories, the speed and adaptability of the newer single-stage approach may be more useful than the raw accuracy of the older, slower methods. The research also notes that the field is moving quickly, with newer models emerging that aim to combine the speed of the single-stage systems with the accuracy of the older ones, but the core lesson remains clear: the best tool depends on whether you need a machine that knows everything about the world, or one that can quickly learn to see something new.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.