💻 computer science

Multi-Agent LLM Collaborative Reasoning and Task Planning for Complex Task Solving

This paper proposes a multi-agent LLM framework featuring distinct planner, executor, and verifier roles operating on a dependency-aware task graph with calibrated parameters and shared-memory mechanisms, which significantly outperforms single-agent baselines in complex task solving by improving completion rates, accuracy, and factual consistency while reducing tool errors.

Wentao Zhang, Tian Liao, Yihui Feng2026-07-06
💻 computer science

Stratum-eval: A Normative-First Evaluation Framework for Clinical Machine Learning

The paper introduces Stratum-eval, a normative-first evaluation framework for clinical machine learning that mandates a validated Normative Specification Document defining use cases, fairness trade-offs, and stakeholder boundaries before computing any performance metrics, thereby ensuring ethical and regulatory alignment is established prior to model deployment.

Hassan Farooq, Abdullah Jawad, Muhammad Salman Butt2026-07-06
💻 computer science

Secure Peer-to-Peer Entropy Generatorfor Fairness in Blockchain-Based Games

This paper proposes and validates a decentralized, peer-to-peer entropy generator that utilizes a commit-reveal protocol and hybrid aggregation to provide statistically uniform, low-cost, and attack-resistant randomness for blockchain-based games, outperforming centralized oracles like Chainlink VRF in both economic efficiency and security.

Jonathas Tavares Neves, Jessica Mariella de Carvalho Oliveira, Carlos Augusto de Moraes Cruz2026-07-06
💻 computer science

When Generative AI Writes Test Cases: Scenario-Driven Evaluation of Generated Tests

This paper introduces a scenario-driven evaluation methodology to compare GenAI-generated test suites against human-written ones, revealing that while direct prompts recover more expected behaviors, ideate-then-implement strategies yield more efficient suites with fewer duplicates, and ideate-only approaches effectively uncover additional test scenarios.

Baris Ardic, Idil Isil Yildirim, Carolin Brandt, Andy Zaidman2026-07-06
💻 computer science

Conscious infrastructure for fairness aware detection of household water insecurity in Hyderabad informal settlements

This paper proposes a Conscious Infrastructure Neural Systems (CINS) framework that utilizes household-burden data rather than traditional infrastructure metrics to significantly improve the fairness and accuracy of AI-driven detection of water insecurity in Hyderabad's informal settlements, thereby bridging the gap between lived experiences and institutional decision-making.

Anil Kumar Palakodeti2026-07-06