← Latest papers
💬 NLP

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

This paper introduces ComboShoppingBench, a novel benchmark designed to evaluate LLM agents on complex, budget-constrained basket shopping tasks by combining semantic assessment with deterministic validation to address the challenges of verifying feasible, coupon-optimized, and compatible product combinations.

Original authors: Adrian Li, Kelong Mao, Yudong Guo, Heming Xia, Xinwei Yang, Lirui Luo, Jace Wong, Pu Yao, Sulong Xu, Simiu Gu

Published 2026-08-11
📖 4 min read☕ Coffee break read

Original authors: Adrian Li, Kelong Mao, Yudong Guo, Heming Xia, Xinwei Yang, Lirui Luo, Jace Wong, Pu Yao, Sulong Xu, Simiu Gu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to go grocery shopping. In the early days, you might have asked the robot to find "a specific brand of milk." That's a simple task: look, find, grab. But real life is rarely that simple. Often, you need to build a whole "combo" meal: a main dish, a side that fits perfectly with it, a drink that pairs well, all while staying under a specific budget and using a stack of confusing coupons that might cancel each other out. This is the world of LLM Agents (Large Language Models acting as digital assistants). These are smart computer programs that can read your request, search through millions of products, and try to make a purchase. The big question scientists are asking is: Can these digital shoppers actually handle the messy, complicated reality of building a full shopping basket without getting confused, overspending, or buying things that don't fit together?

Enter ComboShoppingBench, a new "exam" designed by researchers at JD.com to test just how good these AI shoppers really are. Think of this benchmark not as a simple quiz, but as a high-stakes obstacle course. The researchers created a simulated world filled with real products, coupons, and strict rules. They didn't just ask the AI to pick a few items; they gave it complex missions like "Build a gaming PC that fits in a small case" or "Order a group meal for five people where the drinks come from the same store as the food." The catch? The AI has to make sure every item works with every other item (like making sure the computer parts fit together), use the coupons in the smartest way possible to save money, and stay within a tight budget. It's like asking a robot to be a master chef, a budget accountant, and a coupon-clipping wizard all at once.

The results of this exam were a bit of a reality check for the tech world. The researchers tested 11 different AI agents, including some of the most powerful ones available today. Even the "star student," a model called GPT-5.5 with its thinking mode turned on, only passed the full test about 61.2% of the time. That means it failed nearly 4 out of every 10 tasks. The paper suggests that while these AIs are getting better at finding individual items, they still struggle with the "combo" part of the job. They often miss the subtle connections between products (like buying a phone case that doesn't fit the phone model) or get tripped up by the math of stacking coupons. They might pick the right items but fail to calculate the final price correctly, or they might choose a coupon that looks good on paper but actually costs more in the end.

The researchers built this test using a clever "solution-first" method. Instead of guessing if a shopping list is possible, they first had a computer agent find a working basket of items. Then, they worked backward to create the shopping request, the budget, and the coupons based on that working basket. This ensures the test is fair and solvable. They also used a mix of human-like judges (other AI models) and strict rule-checkers to grade the answers. The rule-checkers acted like a strict cashier, verifying that the math added up and the coupons were legal, while the judges checked if the shopping list actually made sense for the user's request.

The study found that the hardest part for the AI wasn't finding a single product; it was juggling multiple constraints at once. When a task required the AI to remember that "Item A must fit Item B" and "Item C must be from the same store as Item D" and "The total must be under $100," the AI often dropped the ball on one of those rules. The "thinking" mode helped some agents plan better, but it didn't fix everything. In fact, for some models, thinking too hard actually made them worse at using their calculator tools, leading to more math errors. The paper concludes that while we are making progress, today's shopping agents are still far from being the reliable, all-knowing personal shoppers we might hope for. They are like enthusiastic interns who know a lot about products but still need a human to double-check the receipt and make sure the puzzle pieces actually fit together.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →