LLM Pedagogical Behavior in AI Tutoring Interactions
This study introduces a validated five-level scaffolding scale to analyze over 14,000 LLM tutoring responses, revealing that AI models overwhelmingly provide direct assistance (explaining or solving) rather than guided support, a behavior pattern that influences student dialogue but offers limited predictive value for exam performance beyond prior achievement.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet hum of a university computer lab, a new kind of tutoring has taken root, one that never sleeps, never tires, and never judges. Students increasingly turn to large language models, powerful computer programs trained on vast amounts of text, to help them solve homework problems and understand difficult concepts. These digital assistants can do many things, but in the context of learning, a fundamental tension exists. If a tutor gives too little help, a student might get stuck and give up. If the tutor gives too much help, the student might simply copy the answer without ever learning how to think through the problem themselves. This balance is the core of the challenge: finding the right amount of support to guide a learner without doing the work for them. Researchers have long studied how human teachers navigate this space, but it remains unclear how these new, automated tutors behave when students use them in real, unscripted learning situations.
A team of researchers at KAIST in South Korea set out to map this uncharted territory. They analyzed a massive collection of conversations between students and an artificial intelligence tutor during an introductory artificial intelligence course. The students were free to ask for anything, from a simple hint to a complete solution, and the computer responded based on its default programming, which was simply to be a helpful assistant. The researchers wanted to know exactly what kind of help the computer was actually giving. To measure this, they created a five-step scale to categorize every response the AI made. At the lowest end, the computer might offer no real help at all or simply ask the student to think harder. In the middle, it might point out a relevant concept or a specific error without solving the problem. At the highest end, it would explain a method in detail or provide the complete, ready-to-use solution to the student's specific task.
When the researchers examined more than 14,000 responses from over 200 students, a striking pattern emerged. The AI almost never chose the middle ground. More than 95 percent of the time, the computer was either explaining a concept in great detail or simply solving the problem for the student. Responses that merely prompted the student to think or offered a small hint were incredibly rare, making up less than five percent of all interactions. Whether a student asked for help with writing code, editing a report, or understanding a complex theory, the computer's default reaction was to provide substantial, direct assistance. It acted less like a coach who guides a player through a drill and more like a mechanic who fixes the car for them.
This tendency to provide direct answers had a clear ripple effect on how the students behaved next. When the computer gave a full solution, the student was likely to ask for another piece of code or a new task to solve. When the computer offered an explanation, the student was more likely to ask a follow-up question about the underlying concept. The researchers found that the level of help the computer gave reliably predicted what the student would say in the very next turn of the conversation. However, this immediate reaction did not translate into better grades later on. The study looked at how well students performed on three major exams, which they took without any access to the computer. The amount of help the student received during their study sessions did not predict their exam scores any better than their past grades or the simple fact that they were talking to the computer at all.
The findings suggest that while these general-purpose AI tools are incredibly eager to help, they are not naturally inclined to teach in a way that forces students to struggle productively. They default to being efficient problem-solvers rather than patient educators. This does not mean the technology is useless, but it does reveal a specific behavior: without explicit instructions to hold back and guide, these systems will almost always do the heavy lifting for the student. The researchers conclude that to change this dynamic, educators and developers must deliberately design these systems to offer less direct help, rather than hoping the technology will naturally find the right balance on its own. The data provides a clear baseline of how these tools currently behave, showing that the path to effective AI tutoring may require us to teach the machines how to teach, rather than just how to answer.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.