← Latest papers
💻 bioinformatics

A strategy for ionic current analysis of nanopore protein readouts

This paper presents an integrated computational strategy for analyzing nanopore protein ionic current signals that combines denoising, segmentation, and machine learning techniques to achieve high accuracy in charge-based binary classification and to model peptide translocation, despite challenges in distinguishing all 20 amino acids.

Original authors: Stein, A. J., Ghohabi Esfahani, N., Akeson, S., Kakhaki, P. D., Kontogiorgos-Heintz, D., Nivala, J., Jain, M.

Published 2026-10-08
📖 5 min read🧠 Deep dive

Original authors: Stein, A. J., Ghohabi Esfahani, N., Akeson, S., Kakhaki, P. D., Kontogiorgos-Heintz, D., Nivala, J., Jain, M.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Proteins are the workhorses of life, performing the vast majority of tasks that keep our bodies functioning. While our genetic code provides the blueprint for building these molecules, the final product is often more complex than the instructions suggest. Proteins can be chemically altered after they are made, folded into intricate three-dimensional shapes, or combined in different ways to create a dizzying array of distinct forms. To truly understand health and disease, scientists need to read these individual protein molecules directly, rather than just inferring their existence from genetic data. For decades, a technology called nanopore sequencing has allowed researchers to read DNA and RNA with high precision by watching them thread through a microscopic hole. However, applying this same technique to proteins has proven far more difficult. Unlike the four-letter alphabet of DNA, proteins are built from twenty different building blocks called amino acids, each with unique physical properties. Furthermore, proteins are often folded and carry uneven electrical charges, creating a tangled signal that is hard to untangle.

A team of researchers has now developed a new computational strategy to make sense of the electrical signals generated when proteins pass through a nanopore. Their work focuses on a specific method where a molecular motor, acting like a tiny engine, pulls a protein strand through the pore one step at a time. As the protein moves, it disrupts the flow of electricity in a way that depends on which amino acids are currently inside the hole. The challenge lies in decoding this noisy, fluctuating current to identify the sequence of amino acids. The researchers created a pipeline to clean up the raw data, break it into meaningful chunks, and then use computer models to identify the specific amino acids responsible for the signal. They found that while telling every single amino acid apart remains a difficult task, the method is highly effective at distinguishing between groups of amino acids based on their electrical charge.

The study began with a set of synthetic protein strands designed specifically to test this technology. These strands were engineered to contain repeating sections, each holding a single, unique amino acid in the center, separated by markers that create a distinct dip in the electrical signal. This design allowed the researchers to isolate the signal produced by individual amino acids. First, they had to clean the raw electrical data, which was full of random noise that could hide the true biological signal. They used a multi-step process to remove these artifacts while preserving the sharp transitions that mark the movement of the protein. Once the data was clean, they needed to figure out where one amino acid's signal ended and the next began. They tested two different mathematical approaches to find these boundaries. One method automatically detected changes in the signal, while the other forced the data into a fixed number of segments to ensure consistency. Both methods produced reliable results, dividing the protein's journey into roughly thirty-five distinct steps, a number that aligns with the known speed of the molecular motor pulling the protein through.

With the signals segmented, the researchers turned to the task of identification. They used two different types of computer learning models to analyze the electrical patterns. The first approach involved training a deep learning network to recognize statistical features within each segment, such as the average current, the range of fluctuations, and the shape of the signal curve. When asked to identify all twenty amino acids at once, the model struggled, achieving an accuracy of only about twenty-one percent. This result suggests that the electrical signatures of many amino acids are too similar to distinguish when all twenty are competing against each other. However, the model performed much better when asked to choose between just two amino acids at a time. In these head-to-head comparisons, the average accuracy rose to roughly seventy-five percent. Certain amino acids, particularly those with a negative electrical charge, stood out clearly, while others with similar physical properties, like size or lack of charge, were frequently confused with one another.

The second approach used a different kind of model that treats the protein's movement as a sequence of steps, much like a hidden map that the computer tries to follow. This model was particularly good at capturing the overall pattern of the protein's journey, including moments where the motor might slip backward or pause. Like the deep learning model, this approach excelled at distinguishing between positively and negatively charged amino acids, achieving an accuracy of over ninety-three percent in binary tests. It also performed well at separating amino acids by their general size, though it struggled with the middle-sized groups. The researchers found that the most reliable way to read these proteins was not to try to name every single letter in the sequence, but to group them by their most obvious physical traits. The electrical charge of an amino acid proved to be the strongest signal, creating a distinct fingerprint that the computer could recognize with high confidence.

The study concludes that while reading a complete protein sequence from scratch is not yet a solved problem, the technology is ready for more targeted applications. The ability to reliably distinguish between charged and uncharged regions, or to identify specific mutations that change the charge of a protein, offers a practical path forward. The researchers suggest that future improvements will come from using larger datasets and designing protein strands that expose amino acids in more varied contexts. For now, this work provides a solid foundation for understanding how proteins interact with nanopores. It demonstrates that even with complex, noisy data, it is possible to extract meaningful biological information by combining careful signal processing with smart computational models. The path to full protein sequencing is long, but this strategy offers a clear, step-by-step method for navigating the electrical landscape of the proteome.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →