Citation Intent Classification for Turkish Academic Literature with Prompt-Optimized LLMs and Language-Specific Encoders
This paper introduces TurkCite, a framework and expert-annotated dataset for classifying citation intents in Turkish academic literature, demonstrating that language-specific Transformer encoders outperform translation-based transfer while offering prompt-optimized large language models as effective training-free alternatives for non-English bibliometric analysis.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the vast landscape of academic research, a citation is more than just a footnote or a polite nod to a predecessor. It is a specific communicative act, a tiny signal sent from one paper to another that reveals exactly how the new work is using the old. Sometimes a researcher cites a previous study because they are building directly upon its method; other times, they cite it to show how their results contradict it, or simply to provide historical context. For decades, the academic world has measured the value of research primarily by counting these signals, tallying how many times a paper is mentioned to calculate impact. However, this approach treats every mention as identical, missing the crucial difference between a deep intellectual debt and a passing reference. To understand the true nature of scientific influence, scholars need to know not just how often a paper is cited, but why.
This distinction is the heart of citation intent classification, a field that seeks to sort these references into functional categories. While researchers have made significant progress in doing this for English-language literature, where large datasets and advanced computer tools are available, the same tools have largely failed to work for other languages. Turkish academic publishing represents a massive and growing ecosystem, supported by national databases that index thousands of journals, yet it lacked a system capable of understanding the specific reasons why Turkish scholars cite one another. The linguistic structure of Turkish is quite different from English, making it difficult to simply translate English tools and expect them to work. Without a native system, the rich, nuanced conversations happening in Turkish journals remained invisible to automated analysis, stuck behind a wall of raw numbers.
A team of researchers set out to build that missing system, creating a new framework designed specifically for Turkish academic literature. They began by constructing a specialized dataset called TurkCite, which consists of nearly 2,800 citation examples drawn from computer science papers published in Turkey. These examples were not just collected; they were carefully reviewed and labeled by human experts who determined the specific intent behind each reference, sorting them into five distinct categories: background context, methodological basis, supportive findings, contrasting results, and critical discussion. This human-annotated foundation was essential because the team first tested whether they could simply translate Turkish citations into English and use existing English tools to analyze them. The results were clear: translation alone was insufficient. The subtle cues and discourse patterns that signal intent in Turkish were lost in translation, proving that a native solution was necessary.
With their new dataset in hand, the researchers explored two different paths to solve the classification problem. The first path relied on large language models, the powerful artificial intelligence systems that can generate text and answer questions. They tested whether these models could learn to classify citations without being explicitly trained on the new Turkish data, a method known as in-context learning. By carefully crafting instructions and providing a few examples within the prompt, they found that these models could achieve a high level of accuracy, reaching nearly 87 percent in identifying the correct intent. This approach proved valuable as a "cold-start" solution, offering a way to analyze citations immediately without the need for a massive, pre-labeled dataset. However, the researchers noted that the performance of these models was highly sensitive to how the instructions were written, and they struggled to consistently identify the rarer types of citations, such as those indicating a disagreement with a previous study.
The second path involved training specialized computer models from scratch using the new Turkish dataset. The team fine-tuned five different types of neural network architectures, testing which ones worked best for the Turkish language. They discovered that models pre-trained specifically on Turkish text significantly outperformed those designed for multiple languages or English. The most successful system, which combined a Turkish-specific model with a richer input that included the section of the paper where the citation appeared and the sentences surrounding it, achieved an accuracy of over 90 percent. This supervised approach provided a stable, reliable, and cost-effective way to process large volumes of academic text. Yet, even with this high accuracy, the system still faced a challenge: the data itself was heavily skewed. In academic writing, the vast majority of citations are used for general background, making up about 78 percent of the dataset. This imbalance meant that while the models were excellent at spotting common background references, they remained less reliable at identifying the rarer, more complex intents like disagreement or critical discussion.
The study concluded that while the new system offers a powerful tool for analyzing Turkish scholarship, the difficulty of the task lies not just in the technology but in the nature of the data. The researchers found that simply adding more context or breaking the problem into smaller steps did not fully solve the issue of identifying rare citation types. Instead, the path forward involves expanding the dataset to include more examples of these underrepresented intents. By providing a robust framework and a high-quality dataset, the team has opened the door for a more nuanced understanding of Turkish academic influence, allowing researchers and editors to move beyond simple counts and see the actual structure of scientific conversation. This work serves as a blueprint for other non-English academic communities, demonstrating that while translation is a useful first step, true understanding requires tools built specifically for the language and discourse of the local scholarly world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.