← Latest papers
🤖 AI

Framing, Judging, Steering: An Assessable Competency Model for Teach-ing Students to Reason With Generative AI

This paper introduces CoRe-3, a competency model that assesses students' ability to effectively use generative AI by separating the distinct skills of Framing tasks, Judging outputs, and Steering models, and validates this framework through an open platform demonstrating the dissociation and convergent validity of these three skills.

Original authors: Alexander Apartsin, Yehudit Aperstein

Published 2026-06-05
📖 5 min read🧠 Deep dive

Original authors: Alexander Apartsin, Yehudit Aperstein

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Magic Answer" Trap

Imagine you have a magic genie (Generative AI) that can instantly write your homework, solve your math problems, or draft your emails. The paper argues that schools are currently testing students on how well they can do these tasks without the genie. But in the real world, the genie is always there.

The danger isn't using the genie; it's using it uncritically. If you just ask the genie for an answer and copy it, you aren't learning. You're just "offloading" your brain work to the machine. You get a correct answer, but your mind stays the same.

The authors say the real skill we need to teach and test isn't "how to get the answer," but "how to work with the genie."

The Solution: The "FJS" Model (CoRe-3)

The authors propose a new way to measure how well a student works with AI. They call it CoRe-3 (Co-Reasoning). They break the process down into three distinct skills, which they abbreviate as FJS:

  1. Framing (The Architect)

    • What it is: Before you even ask the genie for help, you have to figure out exactly what problem you are solving.
    • The Analogy: Imagine you hire a contractor to build a house. If you just say, "Build me a house," you'll get a mess. You need to say, "Build a 3-bedroom house with a solar roof, on this specific plot of land, under $200k."
    • The Skill: Taking a vague, messy problem and turning it into a clear, specific instruction before the AI starts working.
  2. Judging (The Inspector)

    • What it is: The genie gives you an answer. Your job is to look at it and say, "Is this actually right?"
    • The Analogy: The contractor hands you the blueprints. You don't just nod and say "Great!" You check: "Wait, you put the bathroom on the roof? That's a mistake. And you used wood for the foundation? That's wrong."
    • The Skill: Critically spotting errors, missing details, or hidden assumptions in the AI's output.
  3. Steering (The Pilot)

    • What it is: Once you find the mistakes, you have to tell the genie how to fix them.
    • The Analogy: You don't just say, "Fix it." You say, "Move the bathroom to the second floor and switch the foundation to concrete. Also, make sure the kitchen faces south."
    • The Skill: Giving specific, corrective feedback to guide the AI toward a better result.

The Big Idea: Why Separate Them?

Previous ways of testing AI skills just gave students one score for "Prompting." The authors say this is like grading a chef on "Cooking" without checking if they bought the right ingredients (Framing), tasted the soup (Judging), or adjusted the salt (Steering).

  • Framing happens before the AI speaks.
  • Judging happens after the AI speaks.
  • Steering happens after you judge.

The paper's main claim is that these are three different skills. You can be great at Framing (asking good questions) but terrible at Judging (missing the AI's mistakes). Or you can be great at Steering (fixing things) but have started with a bad plan. The authors proved that you can measure these separately.

How They Tested It (The "Robot Student" Experiment)

To prove these skills are separate, the authors built a digital lab called CoReasoningLab.

  • The Setup: They created 80 "simulated students" (computer programs acting like humans). They programmed some to be good at Framing, some at Judging, and some at Steering, and mixed and matched these skills.
  • The Test: They gave these robot students a series of challenges.
  • The Result: The grading system worked perfectly.
    • When a robot was programmed to be "bad" at Framing, its Framing score dropped, but its Judging score stayed the same.
    • When a robot was "bad" at Steering, its Steering score dropped, but its Framing score stayed high.
    • The Conclusion: The system successfully proved that these three skills are distinct. You can be good at one and bad at another. They are not just one big "AI skill."

Why This Matters

The paper argues that as AI gets smarter, the human's job changes. We don't need to be the ones doing the heavy lifting (the AI does that). Instead, we need to be the Architect, Inspector, and Pilot.

  • Old Education: "Can you solve this math problem alone?"
  • New Education (CoRe-3): "Can you define the problem, spot the AI's errors, and guide it to the right solution?"

The authors have released their tools and the "lab" so teachers can start testing and teaching these three specific skills right now. They aren't saying AI is bad; they are saying that to use AI well, we need to stop treating it like a magic answer machine and start treating it like a partner that needs clear direction and constant supervision.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →