cs.AI papers | Gist.Science

Do Deployment Constraints Make LLMs Hallucinate Citations? An Empirical Study across Four Models and Five Prompting Regimes

This empirical study demonstrates that deployment-motivated prompting constraints significantly exacerbate citation hallucinations across four large language models, with no model achieving a citation existence rate above 47.5% and a substantial portion of unverifiable outputs being fabricated, thereby underscoring the critical need for post-hoc verification in academic and software engineering contexts.

Chen Zhao, Yuan Tang, Yitian Qian2026-03-10💻 cs

MAviS: A Multimodal Conversational Assistant For Avian Species

This paper introduces MAviS, a domain-adaptive multimodal conversational assistant for avian species that leverages the newly created MAviS-Dataset and is evaluated on the MAviS-Bench to achieve state-of-the-art performance in fine-grained bird species understanding and multimodal question answering.

Yevheniia Kryklyvets, Mohammed Irfan Kurpath, Sahal Shaji Mullappilly, Jinxing Zhou, Fahad Shabzan Khan, Rao Anwer, Salman Khan, Hisham Cholakkal2026-03-10💻 cs

A Cortically Inspired Architecture for Modular Perceptual AI

This paper proposes a modular, cortically inspired architecture for perceptual AI that leverages neuroscientific principles like predictive processing and specialized modules to overcome the interpretability and generalization limitations of current monolithic models, thereby enabling more transparent and human-aligned reasoning.

Prerna Luthra2026-03-10💻 cs

Spectral Discovery of Continuous Symmetries via Generalized Fourier Transforms

This paper proposes a novel framework for discovering continuous one-parameter symmetries by leveraging the Generalized Fourier Transform to detect structured sparsity patterns in the spectral domain, offering a principled and interpretable alternative to existing generator-based optimization methods.

Pavan Karjol, Kumar Shubham, Prathosh AP2026-03-10🤖 cs.LG

Data-Driven Hints in Intelligent Tutoring Systems

This chapter reviews the evolution of data-driven hint generation in intelligent tutoring systems, highlighting how historical student data enables the creation of next-step hints, strategic subgoals, and timely interventions, while also exploring future adaptations involving behavioral data and Large Language Models.

Sutapa Dey Tithi, Kimia Fazeli, Dmitri Droujkov, Tahreem Yasir, Xiaoyi Tian, Tiffany Barnes2026-03-10💻 cs

Adversarial Latent-State Training for Robust Policies in Partially Observable Domains

This paper introduces an adversarial latent-initial-state POMDP framework that theoretically establishes a minimax principle and finite-sample guarantees, while empirically demonstrating that targeted adversarial training significantly reduces robustness gaps in partially observable reinforcement learning.

Angad Singh Ahuja2026-03-10🤖 cs.LG

Shutdown Safety Valves for Advanced AI

This paper explores the unorthodox proposal of programming advanced AI systems with a primary goal of being turned off to mitigate the risk of them resisting shutdown, while analyzing the conditions under which such an approach would be effective.

Vincent Conitzer2026-03-10🤖 cs.LG

FinSheet-Bench: From Simple Lookups to Complex Reasoning, Where LLMs Break on Financial Spreadsheets

FinSheet-Bench introduces a synthetic benchmark modeled on real private equity fund structures to evaluate LLMs on financial spreadsheet tasks, revealing that even the best-performing models currently lack the accuracy required for unsupervised professional use, particularly on complex, large-scale documents, and suggesting that reliable extraction will require separating document understanding from deterministic computation.

Jan Ravnik, Matjaž Ličen, Felix Bührmann, Bithiah Yuan, Felix Stinson, Tanvi Singh2026-03-10💻 cs

Norm-Hierarchy Transitions in Representation Learning: When and Why Neural Networks Abandon Shortcuts

This paper introduces the Norm-Hierarchy Transition (NHT) framework, which explains that neural networks delay learning structured representations in favor of spurious shortcuts because weight decay slowly drives the model from high-norm solutions to lower-norm ones, with the transition delay logarithmically scaling to the ratio between these norms.

Truong Xuan Khanh, Truong Quynh Hoa2026-03-10🤖 cs.LG

The Third Ambition: Artificial Intelligence and the Science of Human Behavior

This paper proposes a "third ambition" for artificial intelligence research, advocating for the use of large language models as scientific instruments to study human behavior, culture, and moral reasoning by treating them as computationally accessible condensates of collective discourse while addressing their methodological and epistemic limitations.

W. Russell Neuman, Chad Coleman2026-03-10💬 cs.CL

VisualScratchpad: Inference-time Visual Concepts Analysis in Vision Language Models

This paper introduces VisualScratchpad, an interactive inference-time analysis tool that leverages sparse autoencoders and attention mechanisms to visualize and debug vision language models by linking visual concepts to text tokens, thereby revealing previously underexplored failure modes such as limited cross-modal alignment and misleading visual concepts.

Hyesu Lim, Jinho Choi, Taekyung Kim, Byeongho Heo, Jaegul Choo, Dongyoon Han2026-03-10💻 cs

Agora: Teaching the Skill of Consensus-Finding with AI Personas Grounded in Human Voice

The paper introduces Agora, an AI-powered platform that leverages LLMs to simulate diverse human perspectives on policy issues, enabling users to practice consensus-building and demonstrating through a preliminary study that access to authentic voice explanations significantly enhances problem-solving skills and the quality of collective decisions compared to viewing aggregate data alone.

Suyash Fulay, Prerna Ravi, Emily Kubin, Shrestha Mohanty, Michiel Bakker, Deb Roy2026-03-10💻 cs

Learning Concept Bottleneck Models from Mechanistic Explanations

This paper introduces Mechanistic CBM (M-CBM), a novel pipeline that extracts concepts directly from black-box models using Sparse Autoencoders and Multimodal LLMs to create interpretable Concept Bottleneck Models that outperform prior methods in predictive accuracy and explanation quality while maintaining strict control over information leakage.

Antonio De Santis, Schrasing Tong, Marco Brambilla, Lalana Kagal2026-03-10🤖 cs.LG

AgrI Challenge: A Data-Centric AI Competition for Cross-Team Validation in Agricultural Vision

The AgrI Challenge introduces a data-centric competition framework featuring Cross-Team Validation to demonstrate that while single-source training suffers from significant generalization gaps in agricultural vision, collaborative multi-source training on independently collected, heterogeneous datasets dramatically improves model robustness and real-world performance.

Mohammed Brahimi, Karim Laabassi, Mohamed Seghir Hadj Ameur, Aicha Boutorh, Badia Siab-Farsi, Amin Khouani, Omar Farouk Zouak, Seif Eddine Bouziane, Kheira Lakhdari, Abdelkader Nabil Benghanem2026-03-10🤖 cs.LG

Latent Generative Models with Tunable Complexity for Compressed Sensing and other Inverse Problems

This paper introduces tunable-complexity priors for generative models like diffusion models, normalizing flows, and VAEs by leveraging nested dropout, demonstrating that adaptively adjusting model dimensionality significantly improves reconstruction performance across various inverse problems compared to fixed-complexity baselines.

Sean Gunn, Jorio Cocola, Oliver De Candido, Vaggos Chatziafratis, Paul Hand2026-03-10🤖 cs.LG

The Yerkes-Dodson Curve for AI Agents: Emergent Cooperation Under Environmental Pressure in Multi-Agent LLM Simulations

This paper demonstrates that environmental pressure in multi-agent LLM simulations follows a Yerkes-Dodson inverted-U relationship, where medium stress optimizes emergent cooperative trade while extreme pressure causes behavioral collapse, and suggests that calibrating such pressure serves as an effective curriculum design strategy for agent development.

Ivan Pasichnyk2026-03-10💻 cs

Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes

This paper reveals that in the sub-20M parameter "tiny" regime, models follow steeper but non-uniform scaling laws where increasing size not only reduces overall error but fundamentally alters the structure of mistakes, shifts capacity from easy to hard classes, and paradoxically degrades calibration, necessitating validation at the specific target model size for edge AI deployment.

Mohammed Alnemari, Rizwan Qureshi, Nader Begrazadah2026-03-10🤖 cs.LG

Position: LLMs Must Use Functor-Based and RAG-Driven Bias Mitigation for Fairness

This position paper proposes a dual-pronged framework for mitigating biases in large language models by integrating category-theoretic functor-based transformations to structurally map semantic domains to unbiased forms and retrieval-augmented generation to dynamically inject diverse external knowledge during inference.

Ravi Ranjan, Utkarsh Grover, Agorista Polyzou2026-03-10💬 cs.CL

ConfHit: Conformal Generative Design with Oracle Free Guarantees

ConfHit is a distribution-free framework that enables reliable, oracle-free generative design in drug discovery by providing statistical guarantees that generated molecular batches contain at least one valid hit while allowing for the refinement of these batches into compact sets.

Siddhartha Laghuvarapu, Ying Jin, Jimeng Sun2026-03-10🤖 cs.LG

Domain-Specific Quality Estimation for Machine Translation in Low-Resource Scenarios

This paper addresses the challenge of domain-specific machine translation quality estimation in low-resource scenarios by demonstrating that while prompt-only methods are fragile for open-weight models, adapting intermediate Transformer layers via Low-Rank Adaptation (ALOPE) and Low-Rank Multiplicative Adaptation (LoRMA) significantly improves robustness and performance across English-to-Indic language pairs.

Namrata Patil Gurav, Akashdeep Ranu, Archchana Sindhujan, Diptesh Kanojia2026-03-10🤖 cs.LG

← Previous Next →