← Latest papers
📄 health informatics

Collecting health narratives at scale: A multi-study evaluation of speech-to-text in health surveys

This multi-study evaluation demonstrates that while speech-to-text technology can significantly enrich the length and content of patient health narratives at scale, its adoption varies widely across populations and depends heavily on contextual factors such as privacy, interface design, and individual preferences.

Original authors: Baumer, A. M., Naoumis, P., Fabian, J., Strukova, S., Wolf, M., Ballouz, T., Farnham, A., Hofmann, V. X. C., Spitale, G., Dellwo, V., Glaessel, A., Puhan, M. A., von Wyl, V.

Published 2026-09-25
📖 3 min read☕ Coffee break read

Original authors: Baumer, A. M., Naoumis, P., Fabian, J., Strukova, S., Wolf, M., Ballouz, T., Farnham, A., Hofmann, V. X. C., Spitale, G., Dellwo, V., Glaessel, A., Puhan, M. A., von Wyl, V.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Human beings often carry stories about their health that numbers and checklists simply cannot capture. While doctors and researchers rely on standardized questions to track symptoms and conditions, these rigid formats can miss the texture of a person's lived experience. The nuance of how a patient describes their pain, the unexpected details they share about their daily struggles, and the unique way they connect their symptoms to their life all exist in open-ended stories. Yet, gathering these narratives from large groups of people has historically been a slow, expensive process, usually requiring a researcher to sit down and interview each person individually. This limitation leaves a gap in our understanding of population health, as the most detailed accounts often remain uncollected because they are too difficult to gather at scale.

A team of researchers set out to see if modern technology could bridge this gap by using speech-to-text tools in online surveys. They wanted to know if people would be willing to speak their answers into a microphone instead of typing them, and if doing so would actually result in richer, more detailed stories. To test this, they invited participants from four very different groups: people living with post-COVID-19 condition, healthy adults, individuals who engage in sex work, and older adults. In each of these surveys, the researchers gave people a choice. They could answer open-ended questions by typing on a keyboard, or they could use a speech-to-text feature to speak their answers aloud, which the computer would then transcribe into written words.

The results showed that speaking changed the nature of the responses in measurable ways. When people chose to speak, their answers were significantly longer. In two of the studies, the spoken responses contained hundreds more characters than the typed ones, with one group averaging an increase of nearly nine hundred characters per answer. The spoken words also included more unique content words—the specific nouns and verbs that carry the core meaning of a story—suggesting that speaking allowed people to share more distinct ideas. However, the spoken answers were not perfectly polished; they contained a higher proportion of filler words and showed slightly less variety in the types of words used compared to the typed answers. This indicates that while speech unlocked more volume and specific details, it also introduced the natural, unedited flow of conversation.

Adoption of the technology, however, was not uniform across all groups. The rate at which people chose to speak rather than type varied dramatically, ranging from less than one percent in some groups to eighty percent in others. This wide gap suggests that the decision to speak is deeply personal and heavily influenced by the specific situation. Participants who did use the tool generally praised it for its convenience and the spontaneity it offered, noting that it felt easier to just talk than to type. Yet, many others hesitated due to concerns about privacy, the awkwardness of speaking in public spaces, or simply a preference for the control they felt when typing.

The study concludes that speech-to-text is a viable tool for collecting richer health narratives from large populations, but it is not a one-size-fits-all solution. The technology works best when the implementation is carefully tailored to the specific group being studied and the environment in which they are answering. For some, speaking opens a door to sharing more of their story; for others, the barriers of context and comfort remain too high. The path forward lies in understanding these differences and designing surveys that respect the unique needs and preferences of the people sharing their health experiences.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →