cs.LG papers | Gist.Science

LongAudio-RAG: Event-Grounded Question Answering over Multi-Hour Long Audio

LongAudio-RAG is a hybrid edge-cloud framework that enables precise, low-hallucination question answering over multi-hour audio streams by converting recordings into timestamped event records for SQL-based retrieval, which then grounds Large Language Model responses in structured evidence rather than raw audio.

Naveen Vakada, Kartik Hegde, Arvind Krishna Sridhar, Yinyi Guo, Erik Visser2026-03-10🤖 cs.LG

Accelerated Predictive Coding Networks via Direct Kolen-Pollack Feedback Alignment

This paper introduces Direct Kolen-Pollack Predictive Coding (DKP-PC), a novel algorithm that enhances the efficiency and scalability of biologically inspired predictive coding by establishing direct learnable feedback connections from the output to all hidden layers, thereby reducing error propagation time complexity from O(L) to O(1) while mitigating vanishing updates and maintaining local learning.

Davide Casnici, Martin Lefebvre, Justin Dauwels, Charlotte Frenkel2026-03-10🤖 cs.LG

On the Power of Source Screening for Learning Shared Feature Extractors

This paper demonstrates that strategically screening and training on a carefully selected subset of high-quality, relevant data sources is sufficient to achieve statistically optimal shared feature extraction, even when discarding a substantial portion of available data.

Leo Muxing Wang, Connor Mclaughlin, Lili Su2026-03-10🤖 cs.LG

Emotion Collider: Dual Hyperbolic Mirror Manifolds for Sentiment Recovery via Anti Emotion Reflection

The paper introduces Emotion Collider (EC-Net), a hyperbolic hypergraph framework that leverages Poincaré-ball embeddings, bidirectional message passing, and contrastive learning to achieve robust and noise-resilient multimodal sentiment analysis by preserving high-order semantic relations and enhancing class separation.

Rong Fu, Ziming Wang, Shuo Yin, Haiyun Wei, Kun Liu, Xianda Li, Zeli Su, Simon Fong2026-03-10🤖 cs.LG

ModalImmune: Immunity Driven Unlearning via Self Destructive Training

ModalImmune is a training framework that enhances the robustness of multimodal systems against input channel loss by intentionally collapsing selected modality information during training through a combination of adaptive regularization, targeted intervention, and certified meta-parameter adaptation.

Rong Fu, Jia Yee Tan, Zijian Zhang, Ziming Wang, Zhaolu Kang, Muge Qi, Shuning Zhang, Simon Fong2026-03-10🤖 cs.LG

Whole-Brain Connectomic Graph Model Enables Whole-Body Locomotion Control in Fruit Fly

This paper introduces FlyGM, a whole-brain connectomic graph model that leverages the exact static neural architecture of an adult fruit fly to achieve stable, sample-efficient whole-body locomotion control in embodied reinforcement learning without task-specific tuning.

Zehao Jin, Yaoye Zhu, Chen Zhang, Yanan Sui2026-03-10🤖 cs.LG

Conformal Tradeoffs: Guarantees Beyond Coverage

This paper introduces a framework for operational certification of split conformal predictors that moves beyond marginal coverage by providing finite-sample guarantees for critical deployment metrics like commitment frequency and error exposure through Small-Sample Beta Correction, an independent audit-based auditing protocol, and a geometric analysis of Pareto trade-offs.

Petrus H. Zwart2026-03-10🤖 cs.LG

Latent Equivariant Operators for Robust Object Recognition: Promise and Challenges

This paper demonstrates that neural networks learning equivariant operators in a latent space can effectively generalize to out-of-distribution symmetric transformations on simple datasets like rotated MNIST, while also highlighting the significant challenges involved in scaling this approach to more complex data.

Minh Dinh, Stéphane Deny2026-03-10🤖 cs.LG

Characterizing MARL for Energy Control: A Multi-KPI Benchmark on the CityLearn Environment

This paper establishes a comprehensive multi-KPI benchmark for Multi-Agent Reinforcement Learning in urban energy management using the CityLearn environment, demonstrating that Decentralized Training with Decentralized Execution (DTDE) consistently outperforms Centralized Training with Decentralized Execution (CTDE) in both average and worst-case performance while offering greater resilience and sustainability.

Aymen Khouja, Imen Jendoubi, Oumayma Mahjoub, Oussama Mahfoudhi, Ruan De Kock, Siddarth Singh, Claude Formanek2026-03-10🤖 cs.LG

RAmmStein: Regime Adaptation in Mean-reverting Markets with Stein Thresholds -- Optimal Impulse Control in Concentrated AMMs

This paper introduces RAmmStein, a deep reinforcement learning framework that optimizes liquidity provision in concentrated Automated Market Makers by solving an impulse control problem via a Hamilton-Jacobi-Bellman quasi-variational inequality, thereby significantly reducing rebalancing frequency and gas costs while maximizing net returns through regime-aware, mean-reversion-informed decision-making.

Pranay Anchuri2026-03-10🤖 cs.LG

Benchmarking GNN Models on Molecular Regression Tasks with CKA-Based Representation Analysis

This paper benchmarks four GNN architectures on molecular regression tasks, demonstrating that a hierarchical fusion framework combining GNNs with molecular fingerprints outperforms standalone models by over 7% in RMSE, while CKA analysis reveals that GNN and fingerprint embeddings occupy highly independent latent spaces despite high convergence among isotopic GNN architectures.

Rajan, Ishaan Gupta2026-03-10🤖 cs.LG

MrBERT: Modern Multilingual Encoders via Vocabulary, Domain, and Dimensional Adaptation

The paper introduces MrBERT, a family of efficient, open-source multilingual encoders built on the ModernBERT architecture that achieves state-of-the-art performance in specific languages and specialized domains while leveraging Matryoshka Representation Learning to reduce inference and storage costs.

Daniel Tamayo, Iñaki Lacunza, Paula Rivera-Hidalgo, Severino Da Dalt, Javier Aula-Blasco, Aitor Gonzalez-Agirre, Marta Villegas2026-03-10🤖 cs.LG

Autoregressive Visual Decoding from EEG Signals

The paper introduces AVDE, a lightweight and efficient autoregressive framework that leverages contrastive learning and multi-scale token prediction to decode EEG signals into coherent images, outperforming state-of-the-art methods with significantly fewer parameters while mimicking the hierarchical nature of human visual perception.

Sicheng Dai, Hongwang Xiao, Shan Yu, Qiwei Ye2026-03-10🤖 cs.LG

CeRA: Breaking the Linear Ceiling of Low-Rank Adaptation via Manifold Expansion

CeRA overcomes the linear performance ceiling of Low-Rank Adaptation (LoRA) in complex reasoning tasks by introducing a weight-level parallel adapter with SiLU gating and structural dropout to induce manifold expansion, thereby achieving superior spectral efficiency and preventing rank collapse.

Hung-Hsuan Chen2026-03-10🤖 cs.LG

Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments

This paper addresses the scarcity of expert textual relevance labels in large-scale app store search by leveraging a specialized, fine-tuned LLM to generate millions of high-quality labels, which, when used to augment the production ranker, significantly improves both offline metrics and real-world conversion rates, particularly for tail queries lacking reliable behavioral data.

Evangelia Christakopoulou, Vivekkumar Patel, Hemanth Velaga, Sandip Gaikwad, Sean Suchter, Venkat Sundaranatha2026-03-10🤖 cs.LG

End-to-end Differentiable Calibration and Reconstruction for Optical Particle Detectors

This paper introduces the first end-to-end differentiable optical particle detector simulator that unifies simulation, calibration, and reconstruction into a single gradient-based framework, demonstrating improved accuracy, speed, and flexibility for analyzing large-scale neutrino detectors compared to traditional methods.

Omar Alterkait, César Jesús-Valls, Ryo Matsumoto, Patrick de Perio, Kazuhiro Terao2026-03-10🤖 cs.LG

Attn-QAT: 4-Bit Attention With Quantization-Aware Training

This paper introduces Attn-QAT, the first systematic 4-bit quantization-aware training framework for attention mechanisms that ensures stable FP4 training and inference by matching low-precision recomputation in the backward pass and correcting implicit precision assumptions, thereby eliminating quality drops and delivering up to 1.5x speedup on FP4-capable GPUs without relying on outlier-mitigation heuristics.

Peiyuan Zhang, Matthew Noto, Wenxuan Tan, Chengquan Jiang, Will Lin, Wei Zhou, Hao Zhang2026-03-10🤖 cs.LG

The Partition Principle Revisited: Non-Equal Volume Designs Achieve Minimal Expected Star Discrepancy

This paper introduces a new class of non-equal volume partitions that achieve a lower expected star discrepancy and improved upper bounds compared to classical jittered sampling, thereby providing a theoretical foundation for their use in high-dimensional numerical integration.

Xiaoda Xu2026-03-10🤖 cs.LG

How Well Do Multimodal Models Reason on ECG Signals?

This paper introduces a reproducible, scalable framework for evaluating multimodal models on ECG signals by decomposing reasoning into "Perception" (verified via code generation) and "Deduction" (verified via retrieval against clinical criteria) to address the limitations of existing manual or superficial evaluation methods.

Maxwell A. Xu, Harish Haresamudram, Catherine W. Liu, Patrick Langer, Jathurshan Pradeepkumar, Wanting Mao, Sunita J. Ferns, Aradhana Verma, Jimeng Sun, Paul Schmiedmayer, Xin Liu, Daniel McDuff, Emily B. Fox, James M. Rehg2026-03-10🤖 cs.LG

Opponent State Inference Under Partial Observability: An HMM-POMDP Framework for 2026 Formula 1 Energy Strategy

This paper proposes a tractable two-layer framework combining a Hidden Markov Model for inferring rival energy states and a Deep Q-Network for decision-making to optimize 2026 Formula 1 energy strategies under partial observability, specifically addressing the "counter-harvest trap" where opponents deliberately mask their deployment signals.

Kalliopi Kleisarchaki2026-03-10🤖 cs.LG

← Previous Next →