💻 computer science

Automated Fracture Image Captioning Using Multimodal Vision-Language Models: A Comprehensive Comparative Study on a Clinically Curated Dataset

This paper presents a comprehensive benchmark of nine encoder-decoder architectures for automated fracture radiograph captioning on a clinically curated dataset, demonstrating that the vision-language pre-trained BLIP-Base model outperforms other combinations of visual encoders and GPT-2 decoders while providing critical insights into failure modes and architectural selection for low-resource medical imaging tasks.

Nikosi2026-07-14
💻 computer science

Foresight Is Not Enough: Sentence-Level Future Signals, Self-Loop Hard Negatives, and the Calibration Gap in Small Language Models

This paper demonstrates that while the ForesightLM-v2 model successfully learns a stable sentence-level future prediction signal, this representation fails to directly control generation behavior, revealing that its apparent benefits stem primarily from semantic reranking rather than the learned future term itself and highlighting a critical gap between learned representations and calibrated behavioral outcomes in small language models.

Ahmet Rıfat Öztürk, Yağız Ekrem Dalar, Ömer Faruk Aksoy, Nedim Mutlu Sezer, Feyzi Arda Salihoğlu2026-07-14
💻 computer science

From Telemetry to Techniques: Behavior-Centric MITRE ATT&CK Technique Classification Through Sysmon Event Correlation

This paper proposes a behavior-centric framework that reconstructs Sysmon events into GUID-correlated sequences to improve MITRE ATT&CK technique classification, demonstrating that preserving process-level context outperforms temporal grouping and that a baseline SecureBERT model achieves superior results without explicit class imbalance mitigation.

Amir Hossein Hemmati, Babak Sadeghiyan, Mahsa Saeidi2026-07-14
💻 computer science

GeoDualNet: Geometry-Aware Dual-Latent Neural Codec with Epipolar Cross-Attention for Joint Disparity Vector and Depth Prediction in Multi-View Video Coding

This paper proposes GeoDualNet, the first end-to-end neural codec that jointly learns disparity vector prediction and depth map coding through a unified geometric latent space featuring five novel components, achieving significant bitrate savings and improved synthesized view quality over traditional 3D-HEVC and existing learned multiview codecs.

Reka Sandaruwan Gallena Watthage, Anil Fernando2026-07-14