← Latest papers
💻 computer science

Impacts of Generative AI on Agile Teams' Productivity: A Multi-Case Longitudinal Study

This 13-month longitudinal multi-case study demonstrates that Generative AI significantly enhances agile teams' performance and well-being by increasing the value density of development work rather than its volume, a nuanced impact that is only visible through multi-dimensional frameworks like SPACE rather than traditional activity metrics.

Original authors: Rafael Tomaz, Paloma Guenes, Allysson Allex Araújo, Maria Teresa Baldassarre, Marcos Kalinowski

Published 2026-02-17
📖 5 min read🧠 Deep dive

Original authors: Rafael Tomaz, Paloma Guenes, Allysson Allex Araújo, Maria Teresa Baldassarre, Marcos Kalinowski

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a group of professional chefs working in a busy restaurant kitchen. For years, they've been chopping vegetables, measuring spices, and stirring pots by hand. They are fast, but they are tired, and some tasks are just boring repetition.

Then, the restaurant introduces a Magic Assistant Robot. This robot doesn't cook the whole meal for them, but it can chop onions in a split second, instantly recall the perfect recipe for a sauce, and even suggest a garnish.

This paper is a 13-month diary of three different kitchen teams (Agile Teams) in a massive global restaurant chain (an IT consulting firm) as they started using this robot (Generative AI tools like GitHub Copilot). The researchers wanted to know: Does this robot actually make the kitchen run better, or does it just make the chefs feel like they are working harder?

Here is the story of what they found, broken down simply:

1. The Old Way of Measuring Success (The "Activity" Trap)

Before this study, if you wanted to know if a chef was productive, you might count how many onions they chopped or how many pots they stirred. In software, this is called counting "lines of code" or "commits."

The researchers expected that with the Magic Robot, the chefs would chop way more onions. They thought the "Activity" meter would go through the roof.

The Surprise: The "Activity" meter didn't move at all. The chefs didn't chop more onions. They didn't write more code. In fact, the amount of raw work they did stayed exactly the same.

2. The Real Magic: "Value Density"

If they didn't do more work, why did the restaurant owner (the business) think the teams were super productive?

Because the quality and value of what they produced skyrocketed.

Think of it like this:

  • Before: A chef spends 2 hours chopping onions, 1 hour looking for a recipe, and 30 minutes fixing a burnt sauce. They make one mediocre dish.
  • After: The robot chops the onions in seconds and gives the perfect recipe. The chef spends that saved time designing a complex, beautiful sauce that makes the dish a masterpiece.

The chefs made the same number of dishes, but the dishes were much better and more valuable. The researchers call this "Value Density." The work became "denser" with value, even though the volume of work stayed flat.

3. The "P-A-E" Divergence (The Big Discovery)

The study used a special framework called SPACE to measure productivity. It looked at five things: Satisfaction, Performance, Activity, Communication, and Efficiency.

They found a strange and wonderful split, which they call the P-A-E Divergence:

  • Performance (P) went UP (+59%): The teams delivered way more value (more "story points" or completed tasks).
  • Efficiency (E) went UP: The chefs felt faster and less stressed. They felt like they were in a "flow state."
  • Activity (A) stayed FLAT: The raw number of lines of code written didn't change.

The Analogy: Imagine a runner. Before, they ran 10 miles but only moved 1 mile forward because they were running in circles. With the robot, they still run 10 miles (same effort/activity), but now they run in a straight line and move 5 miles forward (double the performance).

4. The Good, The Bad, and The "Robot Glitch"

The study wasn't all sunshine and rainbows. The Magic Robot had its limits:

  • The Good: It was amazing at boring, repetitive tasks like writing unit tests or fixing simple bugs. The chefs loved it for this. It reduced their stress and made them feel more confident.
  • The Bad: When the task was complex (like integrating a weird, old system or a massive data migration), the robot sometimes got confused. It would give suggestions that looked okay but were actually wrong.
  • The Catch: The chefs still had to do the "heavy lifting" of checking the robot's work. If they trusted the robot blindly, they would serve burnt food. They had to become editors rather than just writers.

5. What This Means for the Future

The biggest lesson from this paper is that we need to stop counting "how much" work people do and start measuring "how good" the work is.

If a manager only looks at the "lines of code" (Activity), they would think the Magic Robot is useless because the numbers didn't go up. But if they look at the value delivered (Performance), they see a massive improvement.

In a nutshell:
Generative AI isn't a machine that makes developers work faster in a frantic, chaotic way. It's a tool that lets them work smarter. It takes the boring, repetitive "toil" out of the day, allowing the human brain to focus on the creative, complex, and high-value problems that actually matter. The team didn't become a machine; they became a more effective team of humans.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →