← Latest papers
💻 computer science

Real-Time Acoustic Keylogging: A Convolutional Neural Network Based Approach

This paper presents a novel real-time acoustic keylogging system that utilizes a convolutional neural network trained on log-mel spectrograms to achieve high-accuracy keystroke inference with low latency (0.35 seconds) and robust performance even under limited training data conditions.

Original authors: Joshua Edgley, Aikaterini Kanta, Athanasios Paraskelidis

Published 2026-09-18
📖 4 min read☕ Coffee break read

Original authors: Joshua Edgley, Aikaterini Kanta, Athanasios Paraskelidis

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Every time a person types on a standard computer keyboard, they create a tiny, distinct sound. To the human ear, these clicks and clacks blend together into a uniform rhythm, indistinguishable from one another. However, the physics of the machine tells a different story. Each key, depending on its location and the specific mechanism beneath it, produces a unique acoustic signature. Just as a fingerprint identifies a person, these subtle variations in sound can identify a specific key. This phenomenon forms the basis of a security threat known as acoustic keylogging. Unlike traditional hacking methods that require installing malicious software on a victim's computer, this attack relies on capturing sound from the outside. With the microphones in modern smartphones and laptops becoming increasingly sensitive, the potential for an attacker to record these sounds and reconstruct a password or confidential message has moved from a theoretical possibility to a practical concern.

Researchers at the University of Portsmouth set out to determine how feasible this attack is when it happens in real time, rather than after the fact. While previous studies had shown that machines could learn to distinguish these sounds, they often relied on analyzing full recordings of typing sessions after they were completed. This approach missed a critical question: could a system listen to a person typing, identify each key as it was pressed, and do so fast enough to be useful to an attacker while the typing was still occurring? To answer this, the team built a prototype system designed to mimic a live listening scenario. They used a standard smartphone to record the sounds of a mechanical keyboard, then fed that audio into a specialized computer program. This program, built on a type of artificial intelligence known as a convolutional neural network, was trained to recognize the visual patterns of sound waves, converting the audio into images that the computer could analyze.

The team's experiments revealed that such a system is indeed capable of real-time operation. In their tests, the system processed the audio and identified the typed keys with an average delay of just 0.35 seconds after each key was pressed. This speed is fast enough to be considered instantaneous for most practical purposes. When the researchers tested the system with common passwords, it successfully reconstructed the typed sequences with high accuracy, often recovering the entire string correctly. The system did make occasional mistakes, but these errors were rarely random; when the machine guessed wrong, it almost always chose a key located physically next to the correct one on the keyboard. This suggests that the system was struggling to tell the difference between two very similar sounds, rather than failing completely.

However, the study also highlighted significant weaknesses that could protect users. The system's success depended heavily on the environment in which it was trained. When the researchers introduced background noise, such as the chatter and hum of a busy office, the system's ability to identify keys dropped dramatically. Similarly, if the recording device was moved even a short distance away from the keyboard, or if the typist changed their speed or pressure, the accuracy suffered. The most striking finding was that the system could not easily adapt to a different keyboard. A model trained on one specific mechanical keyboard failed almost entirely when asked to listen to a different brand, suggesting that the attack relies on very specific hardware characteristics.

Perhaps most concerning for security experts was the discovery that the system did not need vast amounts of data to learn. The researchers found that the artificial intelligence could be trained effectively with as few as four recordings of each key being pressed. This low requirement means that an attacker does not need to spend weeks collecting data to build a working tool; a brief, targeted recording session could be sufficient. Despite these capabilities, the study concluded that the threat is not invincible. The researchers proposed several practical defenses, such as typing in noisy environments, varying one's typing speed, or using software that fills in passwords automatically to avoid physical key presses altogether. While the technology to listen to keystrokes in real time is undeniably present, the study suggests that simple changes in behavior and environment can effectively disrupt the attack, turning a potential vulnerability into a manageable risk.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →