← Latest papers
📄 social_science

Influence of Generative AI on Graduate Students’ Writing Style and Critical Reasoning in Written Paper Critiques

This mixed-methods study of 120 graduate engineering critiques reveals that while generative AI significantly improves surface-level writing scores, it risks distorting assessment validity by masking reduced methodological specificity and fostering cognitive offloading, thereby challenging the accurate measurement of disciplinary reasoning.

Original authors: Alejandro H. Espera

Published 2026-06-29
📖 5 min read🧠 Deep dive

Original authors: Alejandro H. Espera

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a graduate engineering class where students are asked to write "book reports" on complex scientific papers. Instead of just summarizing the plot, they have to act like critics: pointing out flaws in the experiments, questioning the data, and explaining how the findings fit into the bigger picture of energy technology.

This study, conducted by Alejandro H. Espera, is like a detective story about what happens when these students are allowed to use a super-smart AI writing assistant versus when they have to write entirely on their own.

Here is the breakdown of what happened, using simple analogies:

1. The Setup: The "Polished Car" vs. The "Rough Engine"

The researcher looked at 120 of these "book reports" (critiques) from a sustainable energy course.

  • Phase 1 (The AI-Permitted Zone): Students could use AI to help them write.
  • Phase 2 (The AI-Free Zone): Students were strictly forbidden from using AI.

The Result: When AI was allowed, the papers looked like brand-new, shiny luxury cars. They were perfectly organized, flowed smoothly, and had no grammar bumps. The grades were very high (an average of 82 out of 100).

When AI was banned, the papers looked more like cars that had been driven off-road. They were a bit rougher, the sentences were clunkier, and the grades dropped significantly (an average of 60 out of 100).

2. The Twist: The "Mannequin" Problem

Here is the catch. The study argues that the "shiny cars" (AI papers) might be hiding a problem.

Think of a mannequin in a store window. It is dressed in a perfect suit, standing perfectly straight, and looks flawless. But a mannequin has no brain, no thoughts, and no ability to think for itself.

The study found that while the AI-permitted papers looked perfect on the outside (great "Communication" scores), the inside was sometimes hollow.

  • The AI Effect: The AI acted like a "cognitive offloader." It took the heavy lifting of organizing thoughts and finding words, which made the writing look great. But in doing so, it sometimes smoothed over the student's actual thinking.
  • The "Generic" Trap: Some of the high-scoring AI papers were described as "generic." They sounded like a smooth, generic summary of the topic rather than a sharp, specific critique of the specific paper the student was assigned. It was like a student giving a speech that sounded perfect but didn't actually answer the specific question asked.

3. The "Rough" Truth

In the AI-free zone, the papers were messier. But the teacher's feedback suggested these "messy" papers often contained something valuable: authentic struggle.

When a student writes without AI, you can see them wrestling with the ideas. They might stumble over a sentence, but they are actually doing the hard work of figuring out if a scientific method was flawed or if a conclusion was weak. The "roughness" was actually proof that their brain was doing the work the class was supposed to test.

4. The "Fake" Alarm

The study also looked at a tool (Turnitin) that tries to detect AI. In the AI-free zone, the tool flagged many papers as "100% AI-generated."

  • The Reality: Sometimes these were students who actually used AI (violating the rules).
  • The Warning: But the study warns that these detectors aren't perfect. They can be like a smoke alarm that goes off when you're just burning toast. You can't rely on the alarm alone to prove someone cheated; you have to look at the toast (the actual content) to see if it's burnt or just a little crispy.

5. The Big Lesson: What Are We Actually Testing?

The main point of this paper is about validity.

Imagine a driving test.

  • Scenario A: A student drives a car with a self-driving mode turned on. The car arrives at the destination perfectly, no swerving, perfect braking. The instructor gives them an "A" for the drive.
  • Scenario B: A student drives manually. They swerve a bit, brake late, and look nervous. They get a "C."

If the goal of the test is to see if the student can drive, Scenario A is a failure. The car did the work, not the student.

The paper argues that in graduate school, the goal of a critique is to see if the student can think critically, not just write pretty sentences.

  • When AI is allowed, the "Communication" score goes up, but the evidence that the student actually understood the complex science might go down.
  • The "shiny" paper might be a valid piece of writing, but it is an invalid piece of evidence for proving the student's critical thinking skills.

Summary

The study concludes that AI is a powerful tool that can make writing look beautiful, like putting a fresh coat of paint on a house. But if the foundation (the student's own reasoning) is weak, the paint doesn't fix the house.

For teachers, the takeaway is: Don't just grade the paint job. If you want to know if a student can think like an engineer, you need to look past the polished surface to see if the reasoning underneath is real, specific, and their own. The "messy" work without AI might actually be the better proof of learning.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →