← Latest papers
📄 medicine

Validation of Software-Based Hearing Tests Against Clinical Audiometry

This cross-sectional validation study found that while the Apple Health hearing test application demonstrated high accuracy and convenience comparable to clinical audiometry at mid-to-high frequencies, the MiMi application showed significant deviations, highlighting the variable reliability of consumer software-based hearing tests.

Original authors: Pouya Ariamanesh, Tomasz Przewoźny, Izabela Gajda

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Pouya Ariamanesh, Tomasz Przewoźny, Izabela Gajda

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Hearing is a silent conversation between the world and the brain, a process that relies on delicate structures in the ear to translate sound waves into meaning. When this system falters, the result is hearing loss, a condition affecting more than 1.5 billion people globally. For decades, the only way to measure the extent of this loss was through a clinical test called audiometry. This gold-standard procedure requires a person to sit in a quiet room, wear specialized headphones, and press a button whenever they hear a series of beeps at different pitches. While accurate, this method demands expensive equipment, a trained technician, and a controlled environment, making it difficult for many people to access for regular check-ups.

In recent years, the rise of smartphones has promised a simpler solution. Apps now claim to test hearing using the device's speakers or connected wireless earbuds, offering a quick, free, and private way to check one's ears. But a critical question remains: can a test performed in a noisy living room with consumer headphones be trusted as much as one done in a doctor's office? A team of researchers at the Medical University of Gdańsk in Poland set out to answer this by putting two popular hearing apps to the test against the clinical standard. They wanted to know if these digital tools could reliably spot hearing loss or if they were merely guessing.

The study involved forty-six adults who volunteered to undergo three different hearing assessments in a single session. First, they took the traditional test using a calibrated medical audiometer, the reference point against which all other methods are judged. Next, they used the Apple Hearing Test, an app built into iPhones that works with Apple's AirPods Pro II. Finally, they tried the MiMi Hearing Test, a separate application available on the App Store that is highly rated by users. To ensure fairness, the researchers controlled the environment as much as possible, monitoring background noise and ensuring the participants had healthy ears free of blockages or infection. The order of the tests was randomized so that the sequence would not influence the results.

The findings revealed a clear split in performance between the two apps. The Apple Hearing Test showed a remarkable ability to match the clinical machine, but only for specific sounds. When testing higher pitches, from 2,000 Hz up to 8,000 Hz, the app's results were nearly identical to the medical device, with differences so small they were considered negligible. However, at lower pitches, the app struggled. It consistently reported that participants could hear sounds that were actually quieter than they truly were, underestimating the hearing threshold by as much as 6.1 decibels at the lowest frequencies tested. This suggests that while the app is excellent for detecting high-frequency hearing loss, it may give a false sense of security regarding low-frequency hearing.

In contrast, the MiMi Hearing Test behaved differently. Instead of underestimating hearing ability, it consistently overestimated the difficulty of hearing. Across almost every frequency tested, the app told participants they needed louder sounds to hear than the medical machine indicated. The difference was significant, with the app suggesting hearing thresholds were up to 7.9 decibels higher than reality in the middle range of frequencies. This systematic error means the app is more likely to flag a person as having hearing loss when they might actually hear perfectly well, potentially leading to unnecessary worry or follow-up visits.

Beyond the numbers, the researchers asked the participants how they felt about each experience. The results here were telling. When asked which test they trusted the most, the majority pointed to the clinical audiometer, followed by the Apple app, with the MiMi app receiving the lowest marks for trustworthiness. However, when asked about convenience, the Apple app took the lead. Participants found it the easiest and most pleasant to use, likely because it is integrated directly into a device they already carry and use every day. The MiMi app was seen as moderately convenient, while the clinical machine, though trusted, was viewed as the least convenient option.

The study concludes that while these software-based tests are not ready to replace the professional audiometer for making medical diagnoses, they have a valuable role to play. The Apple Hearing Test, in particular, appears to be a strong tool for preliminary screening, especially for detecting high-frequency hearing loss, which is often the first sign of age-related decline. Its accuracy in the higher ranges and its ease of use make it a practical option for people to monitor their hearing health at home. The MiMi app, while popular, showed larger errors that make it less reliable for precise measurement. Ultimately, these apps serve best as a first step, a way to raise awareness and prompt those who might be losing their hearing to seek a professional evaluation, rather than as a final verdict on their auditory health.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →