💻 computer science

BMTransUNet: Boundary-Aware Gated Multimodal Transformer for Remote Sensing Semantic Segmentation

The paper proposes BMTransUNet, a boundary-aware gated multimodal Transformer U-Net that integrates RGB and DSM data through adaptive fusion, Vision Transformer-based context modeling, and edge-enhanced skip connections to achieve superior semantic segmentation accuracy in remote sensing images.

Xiaona Peng, Chengyun Liu, Yaping Zhao, Zhenyan Wang, Zhong Chen, Zhenxue Chen2026-08-10
💻 computer science

Scene-aware Transformer encoder with Multi Instance Learning-guided Attention Mechanism for Anomaly Detection in Video Surveillance

This paper proposes the STAM framework, which integrates contrastive learning for scene-aware embeddings with a transformer encoder guided by Multi-Instance Learning attention, to effectively address contextual and temporal limitations in video anomaly detection and achieve superior accuracy on benchmark datasets.

Rifa Nizam Khan, Mohd. Amjad2026-08-10
💻 computer science

Evaluating Large language models on Understanding Korean indirect Speech acts

This study evaluates the ability of various large language models to understand Korean indirect speech acts by constructing a specialized dataset and employing both automated and human assessments, revealing that while proprietary models like Claude3-Opus outperform open-source alternatives, none yet match human-level proficiency in interpreting context-dependent intentions.

Youngeun Koo, Jiwoo Lee, Dojun Park, Seohyun Park, Sungeun Lee2026-08-10