Applying a Requirements-Focused Agile Management Approach for Machine Learning-Enabled Systems
This paper presents the practical application and evaluation of RefineML, a requirements-focused agile approach tailored for Machine Learning-enabled systems, which was shown in an industry-academia collaboration to improve communication, facilitate early feasibility assessments, and enable dual-track governance, despite remaining challenges in operationalizing ML concerns and estimating effort.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to build a custom car, but instead of a standard engine, you are trying to build a "learning" engine that figures out how to drive itself by looking at millions of photos of roads. This is what building Machine Learning (ML) systems feels like. It's messy, unpredictable, and very different from building regular software.
This paper is a story about a team (a university lab in Brazil and a cybersecurity company called EXA) who tried to build a "smart guard" to stop online scams. They needed a new way to manage this project because old rules didn't work. They created a method called RefineML.
Here is how they did it, explained simply:
The Problem: The "Black Box" vs. The "Assembly Line"
Usually, building software is like an assembly line: you know exactly what parts you need, you build them, and you put them together.
But building AI is more like teaching a dog to fetch. You don't know exactly how long it will take to learn, or if it will even work, until you start training it with data.
- The Conflict: The business people wanted a finished product on a strict schedule. The AI experts needed time to experiment, fail, and try again. They were speaking different languages, and the project was getting stuck.
The Solution: RefineML (The "Dual-Track" Construction Site)
The team invented RefineML, a management style that acts like a construction site with two parallel tracks.
1. The Blueprint Phase (Initial Specification)
Before digging, they used a special checklist called PerSpecML. Think of this as a "Master Blueprint" that forces everyone to agree on:
- What are we building? (The Goal)
- What does the user need to see? (The Experience)
- Do we have enough "bricks" (Data)?
- Is the foundation strong enough? (Infrastructure)
- Analogy: Instead of just saying "Build a house," they agreed on "Build a house with 3 bedrooms, a solar roof, and a garage that fits a truck."
2. The "Two-Track" System (Conception & Feasibility)
This is the core of their innovation. They split the work into two separate but connected lines:
- Track A (The Software Team): They build the car's body, the dashboard, and the doors. They need a "fake engine" to test if the dashboard works.
- Track B (The AI Team): They are busy training the "learning engine." This takes time and is unpredictable.
The Magic Trick: The "Demo API" (The Dummy Engine)
To keep Track A moving while Track B is still training, the AI team built a Dummy Engine (called a Demo API). It wasn't the real learning brain yet, but it pretended to be.
- Analogy: Imagine the software team is building the car's interior. They plug in a cardboard box that looks like an engine. They can test the steering wheel and pedals without waiting for the real engine to be finished. This stops the whole project from stalling.
3. The "Two Sprints Ahead" Rule
The AI team is always two steps ahead of the software team.
- Analogy: The AI team is baking a cake. They start baking the batter (training the model) two days before the software team needs to put the cake in the box (integrate it). This gives the bakers time to fix a burnt cake without delaying the delivery of the box.
4. The "Layers of Done" (LoD)
Instead of waiting for a "perfect" AI, they delivered versions in layers:
- Layer 0: The Dummy Engine (Demo API).
- Layer 1: A "Minimum Viable Model" (MVM) – a model that works well enough to be useful, even if it's not perfect.
- Layer 2+: Continuous improvements.
- Analogy: Instead of waiting for a Ferrari, they delivered a working bicycle first. Then a scooter. Then a car. The customer got value immediately, and the team kept upgrading it.
What Happened in the Real World?
They applied this to a cybersecurity project to stop scams.
- The Result: They successfully built tools to detect scam messages, unsafe websites, and even analyze screenshots of scams.
- The Win: The "Minimum Viable Model" they delivered early on was actually better than the company's old solution. They didn't have to wait years for a perfect AI; they delivered value immediately and kept improving it.
What Worked Well?
- Communication: The "Blueprint" (PerSpecML) stopped the business people and the AI experts from talking past each other. Everyone knew what "success" looked like.
- No More Blocking: Because of the "Dummy Engine," the software team never had to wait around doing nothing while the AI team was experimenting.
- Early Reality Checks: They checked if they had enough data before starting. This saved them from wasting time on projects that were impossible from the start.
What Was Still Hard?
Even with this great system, two big problems remained:
- The "Translation" Gap: It was still hard for the team to turn the high-level "Blueprint" into specific daily tasks. It required an experienced guide (a facilitator) to help translate the big ideas into small steps.
- The "Gamble" of Estimation: You still can't perfectly predict how long it will take to train an AI. Sometimes you try for two weeks and the model gets worse. The paper admits that estimating effort for AI is still a mystery that no management tool can fully solve yet.
The Bottom Line
RefineML is a way to manage AI projects by accepting that AI is unpredictable. It uses a "Dual-Track" system to keep the business moving while the AI learns, uses "Dummy" versions to keep everyone connected, and delivers value in small, improving steps rather than waiting for perfection. It didn't solve the mystery of how long AI takes to learn, but it did solve the problem of how to keep a team working together without getting stuck.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.