Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
This mixed-method study at BNY Mellon, involving 2,989 survey responses and 11 interviews, argues that evaluating AI coding assistants requires a holistic, multifaceted approach that incorporates long-term human-centered factors like technical expertise and work ownership, rather than relying solely on traditional short-term productivity metrics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you run a massive construction company. For years, you've measured your builders' productivity by counting how many bricks they lay per hour or how many walls they finish in a day. But recently, you've handed every builder a magical assistant robot that can instantly suggest the next brick, write the blueprints, and even fix mistakes.
Now, you're confused. The robots are popular, but are they actually making the company faster? And if they are, does "faster" mean the same thing it used to?
This paper is like a deep-dive investigation by a team of researchers (from Carnegie Mellon University and BNY Mellon) who went to talk to nearly 3,000 of these "builders" (software developers) to figure out how to measure success in this new age of AI robots.
Here is the story of what they found, broken down into simple parts:
1. The Big Confusion: Happy vs. Fast
The researchers first asked a simple question: "Are you happy with your robot assistant?" and "How much time does it save you?"
- The Result: Most developers said, "Yes, I love the robot! It makes my day easier." (86% were satisfied).
- The Twist: But when asked, "How much time did you save?" most said, "Not much. Maybe 30 minutes a week."
The Analogy: Imagine you have a super-fast car that gets you to work in 5 minutes, but you spend 45 minutes stuck in traffic. You might love the car because it's fun and reliable, but you aren't actually getting to work any faster than before. The study found that developers love the feeling of using the AI, but it doesn't always translate to huge chunks of saved time. This proves you can't just use one number (like "time saved") to judge if the tool is working.
2. The Six New Rules of the Game
Since the old way of counting (bricks per hour) doesn't work anymore, the researchers interviewed 11 developers to find new ways to measure success. They found six distinct factors that matter, which they grouped into three stages of a project:
Stage A: While Building (The "In the Moment" Feel)
- Self-Sufficiency (The "Do-It-Yourself" Superpower):
- Before: If a builder didn't know how to fix a leaky pipe, they had to stop, call a senior expert, or search a giant library of manuals.
- Now: The robot whispers the answer right in their ear. They feel like a superhero who can solve problems without leaving their desk.
- Frustration & Brain Load (The "Mental Tug-of-War"):
- The Catch: The robot isn't perfect. Sometimes it suggests a solution that looks right but is actually wrong. The developer has to stop, think hard, and double-check everything. This can actually make them more tired and frustrated, even if they are typing faster.
Stage B: Handing Over the Keys (The Team Check)
- Task Completion Speed (The "Throughput" Check):
- This is the old-school metric: How fast did we finish the job? The study found that while AI helps, it doesn't always mean the job is done faster. Sometimes it just means the job is done with less effort, but the time saved is small.
- Peer Review (The "Safety Inspection"):
- Before, a senior builder would check a junior's work. Now, if the junior uses the robot, the senior has to ask: "Did you write this, or did the robot?" If the robot wrote it, the senior has to spend extra time understanding the code to make sure it's safe. Sometimes, the robot makes the code look "too perfect" or confusing, making the safety check harder.
Stage C: The Long Haul (The Future of the Builder)
- Technical Expertise (The "Learning Curve"):
- The Risk: If a junior builder relies on the robot to do all the thinking, they might never learn how to fix a leaky pipe themselves. They might become great at pressing buttons but terrible at understanding the plumbing. The study warns that if we aren't careful, we might create a generation of developers who can't work without the robot.
- Ownership (The "Pride of Creation"):
- The Feeling: Developers love to say, "I built this." If the robot wrote 90% of the code, do they still feel proud? Do they feel responsible if it breaks? The study found that developers worry that if they didn't write the code themselves, they won't feel a deep connection to it, and they might be slower to fix it when it breaks later.
3. It Depends on What You Are Doing
The researchers also found that the AI helps differently depending on the task:
- Building something new: The robot is great at giving a head start, but you have to be careful not to just copy-paste blindly.
- Fixing old code: The robot struggles here because it needs a lot of context. It's like trying to fix a 50-year-old house with a robot that only knows how to build new houses.
- Writing manuals or tests: This is where the robot shines. It's like having a robot that can instantly write the instruction manual for the house you just built. This saves the most time.
The Bottom Line
The paper concludes that we need to stop looking for a single "magic number" to measure productivity.
The Analogy: Imagine trying to judge a chef's skill just by counting how many plates they serve. If they use a robot to chop vegetables, they might serve more plates, but if the robot makes the food taste bad or the chef forgets how to cook, the restaurant fails in the long run.
To truly understand if AI coding assistants are helping, we need to look at the whole picture:
- Are developers happy?
- Are they learning, or just copying?
- Do they feel responsible for the code?
- Is the team checking the work effectively?
The authors say we need a "holistic" view—a balanced scorecard that values the human experience and long-term growth, not just the speed of the output.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.