← Latest papers
💻 computer science

From NER to Business Process Automation: Comparative Evaluation of CNN, LSTM, and Transformer Models for Address Intelligence

This study provides a comparative evaluation of five neural architectures (WordCNN, WordLSTM, BERT, DistilBERT, and ELECTRA) for address extraction in business process automation, demonstrating that ELECTRA achieves the highest accuracy and efficiency while lighter models offer viable alternatives for latency-sensitive applications.

Original authors: Saurabh Kumar Srivastava, Vikas Jalodia, Divya Srivastava

Published 2026-09-22
📖 5 min read🧠 Deep dive

Original authors: Saurabh Kumar Srivastava, Vikas Jalodia, Divya Srivastava

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Every day, businesses move mountains of paper and digital records, from shipping packages to managing student enrollments. A critical part of this work involves reading addresses. While humans can glance at a messy handwritten note or a typed line of text and instantly know the city, state, and zip code, computers often struggle. To a machine, a string of words is just a sequence of characters without inherent meaning. For businesses to automate their workflows—letting software handle tasks without human help—they need a way to turn these unstructured lines of text into neat, organized data. This is where a field of artificial intelligence called Named Entity Recognition comes in. Think of it as a digital highlighter that scans a sentence and tags specific words as "city," "state," or "country." The challenge for engineers is finding the right kind of digital brain to do this highlighting quickly and accurately, especially when the addresses come from different sources and look very different from one another.

A team of researchers at Bennett University in India set out to solve this specific puzzle. They wanted to know which of the five most popular types of neural networks—computer systems designed to mimic the way human brains learn from patterns—works best for breaking down addresses. They tested three distinct families of models: those that look for small, local patterns; those that read text like a sentence, keeping track of the order of words; and the newest, most powerful type that understands the full context of a sentence at once. The researchers did not just look at which model got the most answers right. They also measured how long each model took to make a decision, because in the real world of business automation, speed is just as important as accuracy. They tested these systems on two very different sets of real-world data: records from aerospace suppliers and a massive list of public schools across the United States.

The study began by feeding these models thousands of address lines. The researchers were careful to ensure a fair test. They split the data so that the models were trained on some addresses and tested on completely different ones they had never seen before, preventing the computers from simply memorizing the answers. They also tested the models under different conditions, sometimes giving them only half the data to learn from, and other times giving them the full set. This helped them understand how well each system could learn when information was scarce versus when it was abundant. The goal was to find a system that could handle the messy reality of business data without slowing down the entire operation.

The results painted a clear picture of the trade-offs involved. The models that relied on reading text word-by-word in sequence, known as LSTM networks, performed well and were better at understanding how words relate to each other than the simpler models that just looked for local patterns. However, the most powerful results came from the transformer-based models, which are designed to understand the entire context of a sentence all at once. Among these, one model named ELECTRA stood out as the most efficient. When the researchers gave the models the full amount of data to work with, ELECTRA correctly identified address components 99.4% of the time for the aerospace data and 99.6% of the time for the school data.

What made ELECTRA particularly valuable was not just its high score, but its speed. While the most famous and powerful model, BERT, achieved similar accuracy, it took significantly longer to process each address. In the tests, ELECTRA was roughly six times faster than BERT, completing its work in about 56 milliseconds compared to over 300 milliseconds for the other. This difference is crucial for businesses that need to process thousands of records in real-time. The researchers found that while the simpler, faster models were good enough for very standardized addresses, they could not match the precision of the more advanced systems when the data was complex. Conversely, the most powerful systems were often too slow for high-volume tasks unless a model like ELECTRA was chosen.

The study also revealed that having more data helped all the models improve, but the advanced transformer models benefited the most. When the researchers gave the models 100% of the available data instead of just 50%, the accuracy of the top models climbed even higher, while the simpler models saw more modest gains. This suggests that for the most difficult address parsing tasks, the extra computational power of the advanced models is worth the investment, provided a fast version like ELECTRA is used. The researchers noted that these findings are specific to the structured data they tested, which mostly followed North American address formats. They cautioned that the results might look different if the models were asked to read messy, handwritten notes or addresses from other countries with different rules.

Ultimately, this research provides a practical guide for businesses trying to automate their data processing. It shows that there is no single "best" model for every situation. If a company needs to process millions of standardized addresses quickly, a simpler, lightweight model might be the right choice. If the addresses are complex and require deep understanding, a transformer model is necessary. However, for those who need the best of both worlds—high accuracy and fast speed—ELECTRA emerged as the clear winner in this comparison. The study confirms that while artificial intelligence can now handle these tasks with near-perfect accuracy, the key to successful automation lies in choosing the right tool for the specific job, balancing the need for precision with the need for speed.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →