Novel Vision-Based Camera Distance Estimation for Indoor Positioning
This paper proposes a novel vision-based framework that combines human pose detection, monocular depth features, homography mapping, and a learned residual correction model to accurately estimate the distance between a fixed CCTV camera and a person in GPS-denied indoor environments, achieving an average absolute error of 0.159 meters.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet, complex world of indoor navigation, the familiar tools we rely on outdoors often fail. The Global Positioning System, which guides us through cities and across fields, struggles to penetrate the thick walls and ceilings of buildings, leaving us without a reliable signal. To solve this, scientists have turned to the cameras already watching over our offices, hospitals, and universities. These closed-circuit television systems, or CCTV, offer a unique advantage: they do not need radio signals to see. Instead, they capture the visual geometry of a space, providing a rich map of where people are relative to the camera. However, turning a flat, two-dimensional image into an accurate measurement of real-world distance is a tricky puzzle. A single photo cannot inherently tell how far away an object is; a person standing close to the lens looks large, while the same person far away looks small, creating a confusing ambiguity that has long challenged researchers.
A new study from Koya University in Iraq addresses this challenge by teaching a computer how to measure the distance between a fixed security camera and a person walking through a hallway. The researchers did not attempt to guess the person's exact location on a global map, a task that often leads to errors. Instead, they focused on a more manageable and reliable goal: determining the precise physical distance from the camera to the person's feet. To do this, they built a system that first identifies the person in the video feed and locates their ankles, the specific point where their body meets the floor. If the ankles are hidden or hard to see, the system uses a backup method based on the bottom of the person's outline. The researchers then applied a series of digital corrections to account for the way camera lenses naturally distort images and used a geometric mapping technique to translate the pixel location on the screen into a real-world position on the floor.
To refine this initial calculation, the team introduced a learning step that acts like a fine-tuning mechanism. The system analyzes the image for clues about depth and the size of the person, comparing these visual features against known distances to correct any remaining small errors. When tested on footage from a university corridor, this combined approach proved remarkably effective. The system estimated the distance to a person with an average error of just 0.159 meters, or about 15.9 centimeters. In statistical terms, the root-mean-square error, which accounts for larger mistakes, was 0.218 meters. These results suggest that the method is stable and accurate enough for practical use in indoor positioning.
The study also compared this new method against a more common, off-the-shelf approach that relies solely on detecting a person's general shape without pinpointing their feet or applying geometric corrections. The standard method, which simply draws a box around a person, produced an average error of 0.386 meters, more than double the error of the new system. This difference highlights that simply spotting a person is not enough; understanding exactly where they touch the ground and correcting for the camera's specific angle and lens quirks are essential for precision. The researchers noted that their success relied on a specific set of conditions, using data collected in a single, well-lit corridor with fixed cameras. They acknowledge that real-world environments with moving crowds, changing light, or heavy shadows might present new difficulties.
Despite these limitations, the work demonstrates that existing security infrastructure can be repurposed to provide accurate spatial awareness without the need for expensive new sensors or specialized hardware. By combining the ability to detect human posture with mathematical mapping and a learned correction step, the researchers showed that a standard camera can reliably measure how far away a person is. This capability could eventually help guide people through large buildings, assist emergency responders in locating individuals, or help robots navigate complex indoor spaces, all by simply looking at the world through a lens that already exists.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.