← Latest papers
📄 medicine

A head-to-head comparison of the experimental version of EQ-TIPS-3L and EQ-TIPS-5L (EV3.0) in Chinese preterm infants and toddlers

This study comparing the EQ-TIPS-3L and EQ-TIPS-5L (EV3.0) in Chinese preterm infants and toddlers found that the newer 5L version did not consistently demonstrate superiority over the 3L version in psychometric properties, though the findings require cautious interpretation due to a fixed administration order that may have introduced bias.

Original authors: Lin Xu, Meiying Gao, Chai Ji, Lejing Guan, Mingyan Li, Guannan Bai

Published 2026-08-19
📖 6 min read🧠 Deep dive

Original authors: Lin Xu, Meiying Gao, Chai Ji, Lejing Guan, Mingyan Li, Guannan Bai

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Measuring how a child feels about their health is a delicate task, especially when that child is too young to speak for themselves. For infants and toddlers, doctors and researchers rely on parents or guardians to describe the child's well-being, a process known as proxy reporting. This is particularly crucial for babies born prematurely, who face higher risks of developmental challenges and need careful monitoring as they grow. To track their progress, scientists use questionnaires that ask caregivers to rate specific areas of a child's life, such as their ability to move, play, eat, sleep, and interact with others. These tools are designed to capture a broader picture of health than just survival or physical growth; they aim to understand the quality of life itself. However, creating a questionnaire that is both simple enough for a parent to answer quickly and detailed enough to catch subtle changes in a very young child's condition is a significant scientific challenge.

In the world of health measurement, there is a constant effort to refine these tools. Recently, a widely used system called EQ-TIPS was updated. The original version asked parents to choose from three options for each question: no problems, some problems, or a lot of problems. A newer experimental version expanded this to five options, adding two intermediate choices like "a little bit of a problem" and "more than some problems." The logic behind this change was intuitive: more choices should allow for finer distinctions, capturing mild issues that a three-option scale might miss. It was expected that this newer, more detailed version would be better at telling the difference between children with varying health statuses and would provide more reliable information over time. But until now, no one had directly tested whether this assumption held true for the youngest patients, specifically preterm infants and toddlers.

A team of researchers at the Children's Hospital of Zhejiang University School of Medicine in Hangzhou decided to put these two versions to the test. They gathered a group of 262 caregivers of preterm children, ranging from newborns to just under four years old, who were visiting the hospital for routine developmental check-ups. The study was straightforward in its design but rigorous in its execution. Each caregiver was asked to fill out both the old three-level questionnaire and the new five-level questionnaire, along with a simple visual scale where they rated their child's overall health from zero to one hundred. To check for consistency, a subset of these parents returned to complete the same questionnaires again one to two weeks later. The researchers then compared the results to see which version of the tool performed better at identifying differences between children, matching up with other standard developmental tests, and giving the same answers when asked twice.

The results of this head-to-head comparison were surprising. Contrary to the expectation that the newer, more detailed five-level version would be superior, it did not outperform the original three-level version in any consistent way. In fact, the older three-level tool often worked better. When the researchers looked at how well the questionnaires could distinguish between groups of children—for example, separating those with good growth and development from those with fair or poor development—the three-level version succeeded in identifying five distinct groups, while the five-level version only managed to separate two. The three-level version also showed a stronger ability to align with results from other standard developmental screening tools, suggesting it was more accurately reflecting the child's actual condition.

Perhaps the most telling difference appeared when the researchers checked for reliability. They asked parents to answer the same questions twice, a few days apart, to see if the answers remained stable. The three-level version proved to be much more consistent. When parents answered the same questions again, their responses on the three-level scale stayed remarkably similar. The five-level version, however, showed more variation; parents were less likely to give the exact same answer on the second try. This suggests that the extra choices in the five-level version might have introduced confusion or made it harder for parents to decide on a specific level of problem, leading to less stable results. Interestingly, the single question asking for an overall health rating on a scale of zero to one hundred actually performed better at distinguishing between different groups of children than either of the multi-question versions.

The researchers also noticed a peculiar pattern in how the questionnaires were filled out. Because every parent had to complete the three-level version first, followed immediately by the five-level version, the order of the questions may have influenced the answers. Many parents who reported "no problems" on the first survey also chose "no problems" on the second, even though the second survey offered more nuanced options. It is possible that by the time parents reached the second questionnaire, they were simply repeating their initial judgment or felt that the extra effort to distinguish between "a little bit of a problem" and "some problems" was unnecessary, especially if they had already decided their child was doing well. This "ceiling effect," where most people report the best possible health, was actually higher in the five-level version, with nearly two-thirds of parents selecting the best health state, compared to just over half for the three-level version.

The study concludes that for preterm infants and toddlers, the newer five-level version of the EQ-TIPS does not offer a clear advantage over the original three-level version. The older tool was more reliable, better at spotting differences between children with different health statuses, and more consistent with other developmental measures. However, the researchers are careful to note that the fixed order of the surveys might have skewed these results, making the first tool look better simply because it was answered first. They suggest that future studies need to randomize the order—asking some parents to do the five-level version first and others to do the three-level version first—to confirm these findings. Until then, the evidence suggests that for the very youngest patients, a simpler, three-choice approach may be just as effective, and perhaps even more dependable, than a more complex system with more choices.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →