← Latest papers
💻 computer science

Visual Timelines of Police Encounters in Body-Worn Camera Footage: Operational Context and Activity Cataloging for Training and Analysis in OpenBWC

This paper presents a privacy-conscious system that converts body-worn camera footage into visual timelines by classifying 10-second windows for operational context and motion intensity, thereby accelerating incident review and enhancing officer training workflows.

Original authors: Angela Srbinovska, Christopher Homan, Adrian Martin, Ernest Fokoué

Published 2026-05-19
📖 4 min read☕ Coffee break read

Original authors: Angela Srbinovska, Christopher Homan, Adrian Martin, Ernest Fokoué

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a police officer's body camera as a continuous, unedited movie that runs for hours. For a trainer or a reviewer trying to find a specific moment—like when an officer steps out of their car or when a chase begins—watching every single minute of this movie is like trying to find a needle in a haystack by staring at the whole haystack for days. It's slow, tedious, and inefficient.

This paper presents a new way to organize these videos, turning a long, blurry movie into a visual timeline that acts like a "table of contents" for the action.

Here is how they did it, broken down into simple concepts:

1. The "10-Second Snapshot" Method

Instead of watching the whole video, the researchers chopped every police encounter into small, non-overlapping 10-second chunks (like slicing a loaf of bread into even slices).

  • The Goal: They wanted to label each slice with two simple pieces of information:
    1. Where are we? (The "Context"): Is the officer in a patrol car, outside on the street, inside a house, or is it too dark to see?
    2. What's the energy level? (The "Activity"): Is everything calm and routine, is there a foot chase, or is there chaotic, high-intensity movement?

2. The "Privacy-First" Labeling

To protect people's privacy, the system never tries to recognize faces or identify specific people. It's like looking at a video through a foggy window; you can see the shapes, the lighting, and how fast things are moving, but you can't see who is who.

  • If the video is too dark, blurry, or blocked by a hand, the system doesn't guess. It simply labels that slice as "Low Evidence" (or "I can't see enough to tell"). This prevents the computer from making up stories about what it thinks it sees.

3. Teaching the Computer to "See"

The researchers taught a computer model to look at these 10-second slices and guess the labels. They used two different "eyes" for the computer:

  • The "Photo Eye" (CLIP): This looks at the still pictures inside the video to understand the setting (e.g., "That looks like a dashboard" or "That looks like a hallway").
  • The "Motion Eye" (Optical Flow): This tracks how pixels move from one frame to the next to measure intensity (e.g., "Everything is blurring because the camera is shaking fast" vs. "The scene is very still").

They found that combining these two "eyes" gave the best results.

4. The Results: A Clearer Picture

When they tested this system, here is what happened:

  • Location (Context): The computer was quite good at guessing where the officer was (about 79% accurate). It could easily tell the difference between being in a car, outside, or inside.
  • Action (Activity): It was also good at spotting routine movement (about 88% accurate). However, spotting the most intense, chaotic moments was harder. The computer sometimes confused "unclear video" with "high activity," likely because when things are blurry, it's hard to tell if it's just bad lighting or a real fight.
  • The "Low Confidence" Safety Net: When the computer wasn't sure of its answer, it flagged the video. Interestingly, these "unsure" moments were usually the same ones that human reviewers also found difficult to label. This means the computer is honest about its own confusion.

5. Why This Matters

The paper argues that this method creates a reproducible, auditable index. Think of it like a music player that doesn't just show the song title, but gives you a visual bar showing where the quiet parts are and where the loud, intense parts are.

  • For Reviewers: Instead of watching 45 minutes of footage, they can scan the timeline to jump straight to the "High Activity" or "Foot Pursuit" sections.
  • For Training: Instructors can quickly pull up examples of specific scenarios (like "indoor encounters") without digging through hours of irrelevant video.

Summary

The paper doesn't claim this system can solve crimes or replace human judgment. Instead, it offers a practical tool to organize massive amounts of video data. It turns a chaotic, hours-long stream of footage into a structured, easy-to-read map that highlights where things happened and how intense they were, all while keeping the identities of the people involved private.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →