← Latest papers
📄 chemistry

A Comparative Evaluation of Sample Selection Algorithms for Multivariate Calibration in Near-Infrared Spectroscopic Analysis of Pharmaceutical Formulations

This study demonstrates that the Kennard–Stone algorithm, evaluated via Gaussian process regression on paracetamol tablet NIR spectra, outperforms other sample selection methods in generating robust multivariate calibration models, as confirmed by superior predictive statistics and rigorous non-parametric testing.

Original authors: Amanita Sow, Harouna Sangaré, Fadaba Danioko, Tidiane Diallo

Published 2026-09-09
📖 4 min read☕ Coffee break read

Original authors: Amanita Sow, Harouna Sangaré, Fadaba Danioko, Tidiane Diallo

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of pharmaceutical manufacturing, ensuring that every pill contains the exact right amount of medicine is a matter of safety and trust. For decades, chemists have relied on a technique called near-infrared spectroscopy to check this quality. Imagine shining a special kind of light onto a pile of crushed medicine powder; the way the powder bounces that light back creates a unique fingerprint that reveals its chemical makeup. This method is fast, does not destroy the sample, and requires very little preparation. However, turning these light patterns into a reliable tool for measuring drug strength is not as simple as pointing a camera and taking a picture. The computer needs to learn how to read these fingerprints, a process called building a model. To teach the computer, scientists must show it a collection of samples with known drug amounts. The critical challenge lies in deciding which specific samples to show the computer for learning and which to save for a final test. If the learning group is chosen poorly, the computer might memorize the training examples perfectly but fail completely when it encounters a new, unseen pill.

A team of researchers in Bamako, Mali, set out to solve this specific puzzle of sample selection. They worked with fifty-eight commercial paracetamol tablets, representing different batches from pharmacies across the city. Their goal was to test four different computer strategies for splitting these fifty-eight tablets into a learning group and a testing group. One strategy, known as the Kennard–Stone algorithm, works by picking samples that are as different from each other as possible, ensuring the computer sees the full range of variations. Another, called the Honigs method, picks samples that add the most new information at each step. A third approach, Duplex, tries to build the learning and testing groups at the same time to keep them balanced, while the fourth, Naes, groups similar tablets together and picks a representative from each group. To see which strategy worked best, the researchers used a sophisticated mathematical approach called Gaussian process regression to build their models, a method they had previously found to be highly effective for this type of data.

The results were clear and decisive. When the researchers compared the performance of these four strategies, the Kennard–Stone algorithm emerged as the most reliable, followed closely by the Honigs method. The models built using these two strategies predicted the drug content with near-perfect accuracy, achieving a score of 0.99999 on a scale where one is perfect, and making errors so small they were measured in millionths. In contrast, the Duplex and Naes strategies, while showing excellent results during the training phase, performed significantly worse when tested on new data. This indicated that those methods had likely led the computer to memorize the training set rather than learn the underlying patterns, a problem known as overfitting. The researchers also tested a simple random selection method, where tablets were split into groups by chance, and found it produced the poorest results of all, confirming that a thoughtful, systematic approach is essential.

To ensure these findings were not just a lucky fluke, the team subjected the data to rigorous statistical testing. The analysis confirmed that the superior performance of the Kennard–Stone and Honigs methods was statistically significant and not due to random chance. Interestingly, the statistical tests showed that the Kennard–Stone and Honigs methods were effectively equivalent in their ability to produce accurate models, giving scientists flexibility in choosing between them. The researchers also investigated how the size of the learning group affected the outcome. They found a steady, predictable relationship: as the learning group grew larger, taking up more of the total samples, the model's accuracy improved. The best results were seen when the learning group comprised between 80 and 90 percent of the total samples, suggesting that for this specific type of medicine, a large training set yields the most robust predictions.

The study also examined the mathematical tools used to measure how different the samples were from one another. For the top-performing Kennard–Stone algorithm, using a specific type of distance measurement that accounts for how the different parts of the light spectrum relate to each other proved superior to a simpler measurement. This nuance mattered, as it allowed the algorithm to better understand the complex structure of the data. The researchers concluded that for anyone developing these rapid testing methods for pharmaceuticals, using a systematic selection strategy like Kennard–Stone or Honigs is far better than guessing or picking samples at random. Their work provides a clear, evidence-based guide for ensuring that the computer models used to check our medicine are built on the strongest possible foundation, ensuring that the pills we take are exactly what they are supposed to be.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →