The ICSE 2026 Shadow PC: Training the Next Generation of Reviewers Through Deliberate Practice
This paper presents the ICSE 2026 Shadow PC, a scalable training program utilizing deliberate practice and structured feedback that successfully equipped 102 participants to review 117 papers while achieving high satisfaction rates and demonstrating a viable pathway for developing future software engineering research leaders.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a giant, high-stakes game of "peer review" where scientists submit their best work to a panel of judges. In the world of software engineering, this is how new ideas get tested and approved. But here's the catch: most of these judges are learning on the job. They don't have a training manual; they just hope they get good advice from their mentors or learn by reading other people's reviews. It's like trying to become a professional chef by only watching other people cook, without ever being told why a dish tastes good or bad. The big question is: Can we build a real, structured training program that turns total beginners into expert judges, and can we do it for a huge number of people at once?
This paper tells the story of the ICSE 2026 Shadow PC, a massive experiment in teaching people how to review scientific papers. The organizers didn't just hand out a checklist; they treated reviewing like a sport that requires deliberate practice. Think of it like a sports team that doesn't just play a game once a week, but spends weeks doing drills, watching game tapes, getting instant feedback from coaches, and even critiquing each other's moves. The team used a "shadow" system, meaning these trainees reviewed the exact same papers as the real judges, but their opinions didn't count toward the final decision. This created a safe "practice field" where mistakes were okay and learning was the only goal.
The results were surprisingly successful. Out of 183 people who started the program, 102 stuck with it all the way through, reviewing 117 real papers. The trainees loved it: 97% said they would recommend the experience to others. Even the authors of the papers being reviewed felt the feedback was helpful 67% of the time. The study suggests that by using a structured, multi-step approach—where you try, fail, get corrected, and try again—you can train a large group of people to become competent reviewers. However, the paper also notes that keeping everyone engaged over a long period was tricky, and they suggest that future programs might need "team captains" (called Area Chairs) to help manage the flow and keep the energy up.
The Problem: Learning to Judge by Accident
In the world of software research, peer review is the gatekeeper. It's the process where experts read a new paper and decide if it's good enough to be published. But how do these experts learn to do it? Usually, they don't. They learn "implicitly," which is a fancy way of saying they figure it out by osmosis. A PhD student might get a harsh review from a senior professor and think, "Okay, I guess I need to be more critical," or they might get a vague one and be confused. Some lucky students get direct coaching, but for many, it's a guessing game.
As more people submit papers, the system is drowning in requests for more judges. We need more people who know how to spot a great idea from a flawed one, but there is no clear path to get there. It's like a city trying to hire 1,000 new firefighters but only teaching them by throwing them into a burning building and hoping they learn to hold the hose.
The Solution: The "Shadow" Gym
The authors of this paper decided to fix this by creating a Shadow PC (Program Committee). Think of the main PC as the official Olympic judges. The Shadow PC is the "training camp" right next door. The trainees (mostly early-career researchers like PhD students and new professors) reviewed the exact same papers as the real judges, but they did it in a separate, safe space. Their reviews didn't decide who got in; they were just for practice.
The key innovation here was Deliberate Practice. Instead of just saying, "Here's a paper, write a review," the program was designed like a video game with levels. It was built on the idea that you learn best when you try something, realize you're bad at it, get specific feedback, and try again.
How the Training Camp Worked
The program was broken down into several phases, each designed to teach a specific skill:
The "Try First" Phase (Generation before Instruction):
Usually, teachers give you the rules before you start. This program did the opposite. Participants were asked to write a review before they got any checklists or guides. Why? Because it's hard to learn if you don't know what you're missing. By trying first, they realized, "Oh, I didn't even know I needed to check for this!" It turned their "unconscious incompetence" (not knowing what they didn't know) into "conscious incompetence" (knowing they needed to learn).The "Contrast" Phase (Calibration):
To teach them what a "good" paper looks like, they were given two special papers. One was a paper that had already been accepted to a top conference, and the other was a rejected paper that had been secretly edited to have obvious flaws. The trainees had to review both. Then, they saw how the experts reviewed them. This helped them see the difference between a "flawed but good" paper and a "broken" paper, sharpening their judgment.The "Feedback Loop" (Peer Review of Reviews):
Since there were too many people for one teacher to grade everyone, the trainees graded each other's reviews. This is based on the idea that "teaching is the best way to learn." By critiquing someone else's review, they had to think deeply about what makes a review good, which helped them improve their own.The "Real Deal" Phase:
After all the drills, they were assigned three real papers to review on their own. Then, they had to discuss these papers with a group and write a "meta-review" (a summary of the group's opinion), just like the real judges do.The "Debrief" (The Reveal):
Finally, they got to see what the real judges thought about the same papers. They could compare their reviews with the official ones and see where they matched up and where they missed the mark.
The Results: Did It Work?
The experiment was a hit, but with some growing pains.
- Scale: They managed to get 183 people to sign up. 102 of them finished the whole program. That's a lot of people learning at once!
- The Reviews: These trainees reviewed 117 papers.
- The Verdict: 97% of the participants said they would recommend this experience to others. They felt they learned a lot and gained a new perspective.
- Author Feedback: The people who wrote the papers being reviewed were asked if the trainees' reviews were helpful. 67% said yes. Even more interestingly, 63% of the authors noticed that the trainees' reviews matched the official judges' opinions, meaning the training actually worked to align their thinking with the experts.
- The "Conservative" Bias: The trainees were a bit more cautious than the real judges. They rejected papers more often and asked for "major revisions" more frequently. The authors suggest this is because the trainees didn't have to deal with the hassle of re-reading a paper later if it was fixed, so they were stricter.
The Hiccups and Future Plans
It wasn't perfect. The biggest challenge was keeping people engaged. The program took about three months, and by the end, some people stopped participating, especially during the discussion phase. The organizers had to step in and write some of the final summaries because the groups got stuck.
The authors suggest a few fixes for next time:
- More "Team Captains": They propose creating a new role called "Shadow PC Area Chairs." These would be experienced trainees who manage a small group of reviewers, keeping the discussion moving and helping out. This would also create a career ladder: Trainee Team Captain Real Judge.
- Better Sync: The live meetings were great but too short. They suggest making them longer and recording them for different time zones.
- Community: They want to turn this into a permanent club where past participants come back to help train the next group, creating a self-sustaining community.
The Takeaway
The ICSE 2026 Shadow PC proves that you can teach people how to be good scientific judges, even if they start as beginners. By using a structured, practice-heavy approach—where you try, fail, get feedback, and try again—you can train a large group of people to a high standard. It's not just about making more reviewers; it's about building a pipeline of leaders who know how to judge quality, ensuring that the future of software research stays strong and fair. The paper suggests that if we keep refining this "training camp" model, we can solve the shortage of good reviewers without sacrificing quality.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.