← Latest papers
💻 computer science

Evaluating Multimodal Input Combinations for Zero-Shot Summarization Across Indian Languages

This paper presents a benchmarking study evaluating five pre-trained seq2seq transformer models on the COSMMIC dataset across nine Indian languages to demonstrate that incorporating reader comments and image URLs alongside headlines and content significantly enhances zero-shot multimodal summarization performance.

Original authors: Saptadipa Mazumder

Published 2026-09-24
📖 1 min read☕ Coffee break read

Original authors: Saptadipa Mazumder

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: Evaluating Multimodal Input Combinations for Zero-Shot Summarization Across Indian Languages

Problem Statement
The rapid proliferation of digital news in regional Indian languages has created a demand for automated summarization systems capable of handling heterogeneous content. Traditional summarization pipelines typically rely solely on article text (headlines and body), failing to leverage the rich contextual information available in multimodal news articles, specifically accompanying images and reader comments. While multimodal summarization has advanced for English and high-resource languages, Indian language Natural Language Processing (NLP) remains underexplored due to a lack of multimodal, comment-sensitive datasets and reproducible open-source frameworks. Existing evaluations often rely on proprietary Large Language Models (LLMs) with limited transparency regarding input modality combinations. This study addresses the gap in understanding how lightweight, open-source sequence-to-sequence models perform in a zero-shot setting across nine Indian languages when exposed to various combinations of headlines, content, image URLs, and reader comments.

Methodology
The research proposes a fully automated, end-to-end Python pipeline implemented in a Google Colab environment to benchmark five pre-trained sequence-to-sequence transformer models: mT5, T5, BART, mBART, and PEGASUS.

  • Dataset: The study utilizes the COSMMIC (Comment sensitive multimodal multilingual Indian corpus) dataset, which contains 4,959 article-image pairs and 24,484 reader comments across nine languages (Bengali, Hindi, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, and Odia). The data includes human-written reference summaries.
  • Experimental Design: The evaluation is conducted in a zero-shot manner, meaning none of the models are fine-tuned on the COSMMIC dataset. This isolates the models' inherent pre-training capabilities.
  • Input Modalities: The system systematically tests 15 distinct input combinations formed from four modalities: Headline (H), Content (C), Image URL (I), and Comments (Co). These range from single-modality inputs to the full quadrimodal combination (H+C+I+Co). Notably, Image URLs are processed as raw text strings rather than visual features.
  • Evaluation Metrics: Generated summaries are evaluated against reference summaries using seven metrics: ROUGE-1, ROUGE-2, ROUGE-L, METEOR, BERTScore Precision, BERTScore Recall, BERTScore F1, and CIDEr.
  • Quality Classification: A confusion matrix-based classifier labels generated summaries as "GOOD" (ROUGE-L ≥\ge 0.50) or "BAD" (< 0.50) to provide a binary quality assessment beyond aggregate statistics.
  • Case Study: A detailed analysis is performed on the Hindi subset, the largest in the corpus, to investigate the impact of specific comment types ("Good" vs. "Bad") and the marginal utility of Image URLs.

Key Contributions

  1. Reproducible Open-Source Pipeline: The author introduces the first fully automated, open-source pipeline for multimodal summarization on the COSMMIC dataset. It automates data acquisition, splitting, model loading, generation, and multi-metric evaluation without requiring manual annotation or proprietary API access.
  2. Comprehensive Zero-Shot Benchmark: The study provides the first systematic benchmark of five distinct sequence-to-sequence models (mT5, PEGASUS, BART, mBART, T5) across all 15 possible input combinations for nine Indian languages.
  3. Modality Contribution Analysis: By testing all subsets of input modalities, the work quantifies the specific contribution of headlines, comments, and image URLs to summarization quality, offering guidance for system architects on resource allocation.
  4. Quality Classification Framework: The integration of a confusion matrix-based quality classifier allows for the visualization of "GOOD" vs. "BAD" summary distributions, offering a more intuitive measure of model reliability than raw scores alone.

Results and Analysis

  • Model Performance: Among the evaluated models, mBART demonstrated the most robust performance in the zero-shot setting, particularly for the Hindi language. This is attributed to its explicit pre-training on Hindi data within the CC25 corpus. Conversely, T5 and mT5 showed lower performance, likely due to their English-centric or less specific pre-training for these languages.
  • Input Modality Impact:
    • Content (C) alone provided the strongest baseline performance.
    • Adding Reader Comments generally improved metrics, but the quality of comments mattered; "Good" comments enhanced summary quality, while unfiltered or "Bad" comments introduced noise that degraded performance.
    • Image URLs, when included as text strings, provided marginal but consistent improvements in ROUGE-L scores across most languages, suggesting that URL metadata can serve as weak supervisory signals.
    • The full combination of all four modalities (H+C+I+Co) did not always outperform the Content-only baseline, indicating that unfiltered multimodal inputs can introduce noise in a zero-shot context.
  • Cross-Language Variance: Performance varied significantly across languages. Hindi yielded the best results, while languages with complex scripts or agglutinative morphology (e.g., Malayalam, Tamil) showed lower scores, highlighting the "data size effect" and script complexity challenges.
  • Zero-Shot Limitations: The study observed extremely low ROUGE-2 scores (near 0.00) across most models and languages. This underscores that zero-shot abstractive summarization for Indian languages remains a significant challenge and that practical deployment likely requires at least minimal task-specific fine-tuning.
  • Artifact Identification: The analysis of mBART outputs revealed the generation of special tokens (e.g., <extra_id_0>) due to the model's span-masking pre-training objective, necessitating post-processing or prompt engineering (e.g., adding "summarize") for optimal results.

Significance
The paper establishes a foundational, reproducible framework for evaluating multimodal summarization in low-resource Indian languages. By demonstrating that mBART is currently the most suitable architecture for zero-shot tasks in this domain and that comment quality is a critical variable, the study provides actionable insights for researchers and developers. The work emphasizes that while multimodal inputs hold promise, their integration requires careful filtering to avoid noise. Ultimately, the study argues that while zero-shot baselines are valuable for establishing limits, the path to practical, high-quality summarization systems for Indian languages necessitates moving beyond zero-shot inference toward fine-tuned, language-specific adaptations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →