The comparative evaluation of explanations for algorithmic programming problems provided by a Human vs. ChatGPT: an exploratory online user study
This exploratory study of 37 upper secondary students reveals that while participants could not significantly distinguish between human and ChatGPT-generated programming solutions, their acceptance of a 3D avatar tutor was unaffected by prior programming experience or avatar familiarity, highlighting both optimism for and skepticism about integrating generative AI into education.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the classroom, the relationship between a student and a teacher has long been defined by the exchange of knowledge, where a human guide breaks down complex ideas into understandable steps. Today, a new kind of guide has entered the room: artificial intelligence. These systems can generate text, solve math problems, and write computer code with a speed that rivals human experts. But a critical question remains for educators and parents alike: when a student is trying to learn, does it matter who is doing the explaining? Is there a fundamental difference between a solution crafted by a person and one generated by a machine, and can a learner tell the difference? This inquiry sits at the heart of a recent study exploring how students perceive explanations for difficult programming tasks, specifically when those explanations are delivered by a digital character rather than a living person.
The researchers, a team of scientists from universities in Kazakhstan, set out to test the limits of this emerging technology with a group of thirty-seven high school students. These were not beginners; the participants were aged fifteen to seventeen and came from a specialized school known for training students in mathematics and science competitions. They had already spent years learning to code, with many having started as early as primary school. The goal was to see if these experienced young programmers could distinguish between solutions written by a human instructor and those produced by a large language model, a type of artificial intelligence capable of generating text and code. To make the test fair and remove any bias based on a person's voice or appearance, the researchers used a 3D animated character, or avatar, to deliver the explanations. This digital figure spoke with a computer-generated voice, ensuring that the students could not rely on human cues like tone or emotion to guess the source of the answer.
The experiment was structured like a series of challenges. The students were shown five different programming problems, each accompanied by a solution and an explanation. These problems were taken from a popular online platform used by competitive programmers. For each problem, the explanation was delivered by the same 3D avatar, but the source of the code varied. In some instances, the code and explanation came from a human expert who had solved the problem on the website. In others, the code was generated by an artificial intelligence system. The students watched the avatar explain the solution, and then they were asked a series of questions. They had to rate how well they understood the material, guess whether the code was written by a human or a machine, and evaluate the avatar itself. The researchers also checked if the students' prior experience with programming or their previous interactions with digital avatars influenced their ability to spot the difference.
The results of the study offered a surprising insight into how students interact with this technology. When asked to identify the source of the code, the data revealed a nuanced pattern: participants selected "provided by ChatGPT" more often and resolutely for tasks that were actually provided by ChatGPT. However, for tasks provided by humans, they were more hesitant and selected the source more equably between human and AI. This suggests that while students could identify AI-generated content with a degree of confidence, distinguishing human solutions was more challenging, leading to a mix of correct and incorrect guesses rather than a clear, consistent ability to tell them apart. The students rated their understanding of the solutions similarly, regardless of whether the code came from a person or a machine. Furthermore, the study found that a student's background did not change the outcome. Whether a student had started coding in primary school or high school, or whether they had interacted with digital assistants before, did not significantly affect their ability to distinguish the source or their level of engagement.
While the students could not easily tell the difference between the two sources, their feedback on the delivery method was more critical. The 3D avatar, which served as the narrator for both human and machine solutions, received mixed reviews. Many students noted that the character's voice was monotonous and lacked the natural rhythm of human speech. They pointed out that the avatar's movements were sometimes repetitive and did not always match the words being spoken, which made it harder to focus on the complex logic of the code. Some students felt that the robotic nature of the character made the learning experience feel less personal and more difficult to follow. Despite these criticisms, the overall sentiment was optimistic. Most students believed that such technology could be useful in the future, particularly as a tool to explain difficult concepts when a teacher is not available. They acknowledged that while the current version of the avatar had flaws, improvements in voice and animation could make it a valuable addition to education.
The study concludes that while artificial intelligence is rapidly becoming capable of producing explanations that are indistinguishable from human ones in terms of clarity and accuracy, the medium of delivery still matters. The students were willing to accept the technology as a learning aid, but they were quick to identify the limitations of the current digital avatars. The research suggests that as these tools evolve, the focus must shift not just on the quality of the code they generate, but on the quality of the interaction they provide. For educators, the findings imply that students may soon be able to learn effectively from machine-generated content, but the human touch of a teacher remains unique in its ability to connect, adapt, and engage in a way that a static, monotonous avatar cannot yet replicate. The future of education may well involve a partnership between human teachers and intelligent machines, but the path to that partnership requires refining the digital tools to be as engaging as the people they are meant to assist.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.