cs.CL papers | Gist.Science

Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?

This paper demonstrates that commonly used bias metrics fail to reliably capture allocational harms from large language models because they overlook the critical discrepancy between model predictions and the actual decisions made when allocating limited resources.

Hannah Cyberey, Yangfeng Ji, David Evans2026-03-09💬 cs.CL

Goldfish: Monolingual Language Models for 350 Languages

The paper introduces Goldfish, a suite of over 1,000 small monolingual language models for 350 languages that outperform large multilingual models in grammaticality and perplexity, particularly for low-resource languages where existing models often struggle.

Tyler A. Chang, Catherine Arnett, Zhuowen Tu, Benjamin K. Bergen2026-03-09💬 cs.CL

Towards Resource Efficient and Interpretable Bias Mitigation in Large Language Models

This paper proposes a resource-efficient and interpretable framework that mitigates biases in large language models at decoding time by leveraging small, fine-tuned expert models to generate debiasing signals, achieving significant reductions in gender, race, and religion biases while preserving overall model performance.

Schrasing Tong, Eliott Zemour, Jessica Lu, Rawisara Lohanimit, Lalana Kagal2026-03-09💬 cs.CL

SpecFuse: Ensembling Large Language Models via Next-Segment Prediction

The paper introduces SpecFuse (referred to as SpecEM in the abstract), a training-free ensemble framework that enhances large language model performance by enabling segment-level semantic collaboration through speculative decoding and dynamically adjusting model weights via an online feedback mechanism to prioritize stronger contributors.

Bo Lv, Nayu Liu, Chen Tang, Xin Liu, Yue Yu, Ping Luo2026-03-09🤖 cs.AI

Rethinking the Mixture of Vision Encoders Paradigm for Enhanced Visual Understanding in Multimodal LLMs

This paper introduces LEO, a streamlined multimodal large language model architecture that employs a lightweight fusion strategy of post-adaptation projectors, tile-level sequence interleaving, and dynamic tiling to significantly enhance visual understanding across diverse benchmarks and specialized domains like autonomous driving.

Mozhgan Nasr Azadani, James Riddell, Sean Sedwards, Krzysztof Czarnecki2026-03-09💬 cs.CL

Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation

This survey provides a comprehensive overview of the emerging ecosystem of large language models and tools that support researchers across the scientific lifecycle, covering key tasks from literature search and idea generation to content creation, experimentation, and evaluation, while addressing associated datasets, methods, limitations, and ethical concerns.

Steffen Eger, Yong Cao, Jennifer D'Souza, Andreas Geiger, Christian Greisinger, Stephanie Gross, Yufang Hou, Brigitte Krenn, Anne Lauscher, Yizhi Li, Chenghua Lin, Nafise Sadat Moosavi, Wei Zhao, Tristan Miller2026-03-09🤖 cs.AI

Conditioning LLMs to Generate Code-Switched Text

This paper proposes a methodology to fine-tune Large Language Models for generating fluent English-Spanish code-switched text by leveraging back-translated parallel corpora, demonstrating that while traditional metrics fail to correlate with human preferences, LLM-based evaluation aligns well with human judgment and the approach significantly advances CS text generation capabilities.

Maite Heredia, Gorka Labaka, Jeremy Barnes, Aitor Soroa2026-03-09🤖 cs.AI

CAReDiO: Cultural Alignment via Representativeness and Distinctiveness Guided Data Optimization

The paper introduces CAReDiO, a novel data optimization framework that enhances cultural alignment in Large Language Models by iteratively optimizing data for representativeness and distinctiveness, enabling efficient and superior alignment across 15 diverse cultures with as few as 200 training samples.

Jing Yao, Xiaoyuan Yi, Jindong Wang, Zhicheng Dou, Xing Xie2026-03-09💬 cs.CL

RM-R1: Reward Modeling as Reasoning

The paper introduces Reasoning Reward Models (ReasRMs), specifically the RM-R1 family, which reformulate reward modeling as a reasoning task using a chain-of-rubrics mechanism and a two-stage training pipeline to achieve superior interpretability and performance compared to existing large-scale models.

Xiusi Chen, Gaotang Li, Ziqi Wang, Bowen Jin, Cheng Qian, Yu Wang, Hongru Wang, Yu Zhang, Denghui Zhang, Tong Zhang, Hanghang Tong, Heng Ji2026-03-09🤖 cs.AI

Maximizing Asynchronicity in Event-based Neural Networks

This paper introduces EVA, a novel event-by-event asynchronous-to-synchronous (A2S) framework inspired by language modeling that generates highly expressive features, outperforming prior methods in recognition tasks and achieving state-of-the-art results in detection for event-based vision.

Haiqing Hao, Nikola Zubic, Weihua He, Zhipeng Sui, Davide Scaramuzza, Wenhui Wang2026-03-09🤖 cs.AI

Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering

This paper proposes K-CAST, a novel fine-grained conditional activation steering method that dynamically mitigates content biases in large language models, significantly improving formal reasoning accuracy by up to 15% while maintaining robustness across prompts and languages.

Marco Valentino, Geonhee Kim, Dhairya Dalal, Zhixue Zhao, André Freitas2026-03-09🤖 cs.AI

AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference

This paper introduces AdAEM, a novel self-extensible evaluation framework that automatically generates adaptive test questions by probing the internal value boundaries of diverse LLMs to overcome the limitations of static benchmarks and provide more informative, distinguishable insights into models' value differences and alignment dynamics.

Jing Yao, Shitong Duan, Xiaoyuan Yi, Dongkuan Xu, Peng Zhang, Tun Lu, Ning Gu, Zhicheng Dou, Xing Xie2026-03-09🤖 cs.AI

From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise

This paper introduces a deterministic, automated pipeline that transforms raw domain corpora into completion-style benchmarks to provide scalable, unbiased, and LLM-independent evaluations of domain expertise in both base and instruction-tuned models, effectively addressing issues of benchmark contamination and multiple-choice bias.

Nitin Sharma, Thomas Wolfers, Ça\u{g}atay Yıldız2026-03-09💬 cs.CL

Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts

The paper introduces Sysformer, a novel approach that safeguards frozen large language models by learning to adapt system prompts in the embedding space to significantly improve safety robustness against harmful inputs and jailbreaking attacks without requiring costly model fine-tuning.

Kartik Sharma, Yiqiao Jin, Vineeth Rakesh, Yingtong Dou, Menghai Pan, Mahashweta Das, Srijan Kumar2026-03-09🤖 cs.AI

VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models

This paper introduces VLMQ, a post-training quantization framework tailored for vision-language models that leverages a gradient-driven importance factor to address visual over-representation and modality gaps, thereby achieving state-of-the-art performance across various model sizes and low-bit settings.

Yufei Xue, Yushi Huang, Jiawei Shao, Lunjie Zhu, Chi Zhang, Xuelong Li, Jun Zhang2026-03-09🤖 cs.AI

Agri-Query: A Case Study on RAG vs. Long-Context LLMs for Cross-Lingual Technical Question Answering

This paper presents a case study demonstrating that Hybrid Retrieval-Augmented Generation (RAG) consistently outperforms direct long-context prompting across English, French, and German technical agricultural manuals, achieving over 85% accuracy with models like Gemini 2.5 Flash and Qwen 2.5 7B in cross-lingual question answering.

Julius Gun, Timo Oksanen2026-03-09💬 cs.CL

CMRAG: Co-modality-based visual document retrieval and question answering

The paper proposes CMRAG, a novel framework that enhances visual document retrieval and question answering by unifying text and image modalities through a shared embedding space and a normalized co-modality retrieval method, thereby outperforming existing single-modality approaches.

Wang Chen, Wenhan Yu, Guanqiang Qi, Weikang Li, Yang Li, Lei Sha, Deguo Xia, Jizhou Huang2026-03-09💬 cs.CL

MERLIN: Multi-Stage Curriculum Alignment for Multilingual Encoder-LLM Integration in Cross-Lingual Reasoning

MERLIN is a two-stage, curriculum-based framework that integrates multilingual encoders with LLMs using efficient DoRA fine-tuning to significantly enhance cross-lingual reasoning performance, particularly in low-resource languages where existing methods and even GPT-4o-mini fall short.

Kosei Uemura, David Guzmán, Quang Phuoc Nguyen, Jesujoba Oluwadara Alabi, En-shiun Annie Lee, David Ifeoluwa Adelani2026-03-09💬 cs.CL

Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation

This paper addresses the inconsistency and structural biases in existing latency metrics for simultaneous speech-to-text translation by introducing a comprehensive meta-evaluation, proposing new metrics (YAAL and LongYAAL) and a resegmentation tool (SoftSegmenter), and implementing these solutions within the OmniSTEval toolkit to enable more reliable system assessments.

Peter Polák, Sara Papi, Luisa Bentivogli, Ondřej Bojar2026-03-09🤖 cs.AI

Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs

This paper demonstrates that while standard decoder-only models underperform compared to encoder-only architectures in cross-modal adaptation for partial differential equations, introducing novel bidirectionality-mimicking techniques like Parallel Flipping and Sequence Doubling effectively closes this performance gap.

Paloma García-de-Herreros, Philipp Slusallek, Dietrich Klakow, Vagrant Gautam2026-03-09🤖 cs.LG

← Previous Next →