← Latest papers
💻 computer science

Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage

This paper proposes a novel interdisciplinary framework that leverages advanced AI, including multimodal analysis and large language models, to automatically detect, classify, and summarize behavioral dynamics such as escalation and respect in police body-worn camera footage to enhance law enforcement accountability and training.

Original authors: Anita Srbinovska, Angela Srbinovska, Vivek Senthil, Jonathan Bateman, Adrian Martin, John McCluskey, Ernest Fokoué

Published 2026-09-11
📖 5 min read🧠 Deep dive

Original authors: Anita Srbinovska, Angela Srbinovska, Vivek Senthil, Jonathan Bateman, Adrian Martin, John McCluskey, Ernest Fokoué

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Every day, police officers in cities across the United States wear cameras that record their interactions with the public. These devices, known as body-worn cameras, capture a vast and growing library of video and audio, offering a raw, unfiltered look at the moments when law enforcement meets the community. For decades, researchers and police leaders have hoped to sift through this footage to understand what happens during a traffic stop or a neighborhood call, looking for patterns of respect, tension, or de-escalation. However, the sheer volume of data has made this task nearly impossible for human eyes and ears alone. A single officer might generate hours of footage in a week, and a whole department accumulates thousands of hours every year. To make sense of this deluge, scientists are turning to artificial intelligence, not to replace human judgment, but to help organize and highlight the most important parts of these complex encounters. The challenge lies in teaching machines to hear over the noise of sirens and shouting, to distinguish between different voices in a chaotic scene, and to understand the meaning behind the words spoken in high-pressure situations.

A team of researchers from the Rochester Institute of Technology, working alongside the Rochester Police Department, has built a new system designed to tackle this problem. They call their framework OpenBWC, an open-source tool that uses a combination of audio processing, speech recognition, and language analysis to turn raw video files into structured, searchable information. The goal is not simply to transcribe what was said, but to identify the dynamics of the interaction: Was the officer respectful? Did the situation escalate or calm down? To do this, the researchers fed the system 1,225 videos obtained through public records requests. These recordings, which had been edited to protect the privacy of the people involved by blurring faces, ranged from brief eleven-second clips to nearly twelve hours of continuous footage, totaling over 1,877 hours of video. The team broke these long files into manageable thirty-second segments to help the computer process the audio without losing the flow of conversation.

The core of their method involves a multi-step process that mimics how a human might listen to a difficult recording. First, the system separates the mixed audio into individual voices, a technique known as speaker separation. This is crucial because body-worn camera footage often contains overlapping speech, background noise, and multiple people talking at once. The researchers used a specialized model to isolate the officer's voice from the civilian's, allowing the computer to focus on one speaker at a time. Once the voices were separated, the system transcribed the audio into text using advanced speech-to-text technology. However, the researchers found that not all transcription tools worked equally well in these messy, real-world environments. They tested two versions of a popular speech recognition system and discovered that the smaller, faster version frequently missed words, repeated phrases, or failed to capture the full conversation. In contrast, the larger, more detailed version produced much more accurate transcripts, though it still required human review to catch subtle errors.

After generating the text, the researchers applied a large language model to read through the transcripts and summarize the interactions. This step allowed the system to identify key themes, such as whether an officer used de-escalation tactics or if the tone of the conversation shifted from calm to aggressive. The results were then stored in a database that could be searched by topic, speaker, or specific behavior, making it possible to find patterns across thousands of incidents. The team found that while the system worked well for routine interactions like traffic stops, where the audio is clearer and the speech patterns are predictable, it struggled significantly in chaotic scenarios. In high-stress situations involving shouting, sirens, or multiple people speaking over one another, the accuracy of the transcription dropped, and the system sometimes confused who was speaking or missed critical details.

The researchers were careful to note that their system is not a perfect solution and should not be used as the sole basis for disciplinary decisions or legal judgments. They explicitly ruled out the idea that fully automated transcription could currently replace human review in critical incidents. The study suggests that the most effective approach is a hybrid one, where the AI handles the heavy lifting of organizing and summarizing the data, but human experts verify the results, especially in complex or high-stakes cases. The team also highlighted a technical limitation: the officer's voice is often captured much more clearly than the civilian's because the camera is worn on the officer's body, pointing toward them. This imbalance means that the system is naturally better at understanding the officer's side of the conversation, which could skew the analysis if not accounted for.

Despite these challenges, the project demonstrates a practical path forward for using artificial intelligence in policing. By combining open-source tools with rigorous testing, the researchers have created a framework that can help police departments review their own practices, train officers on better communication, and hold themselves accountable to the public. The work does not claim to have solved the problem of analyzing police footage, but it offers a reliable method for starting the conversation. The findings suggest that while machines can help us see patterns in the noise, the most important insights still come from combining computational power with human understanding. As the technology improves and the datasets grow, this kind of interdisciplinary approach could transform how law enforcement agencies learn from their daily work, turning vast archives of video into actionable knowledge that improves safety and trust for everyone involved.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →