TY - GEN
T1 - Focus on What Matters
T2 - 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
AU - Chang, Aofei
AU - Huang, Le
AU - Boyd, Alex James
AU - Bhatia, Parminder
AU - Kass-Hout, Taha
AU - Xiao, Cao
AU - Ma, Fenglong
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Medical Large Vision-Language Models (Med-LVLMs) often exhibit suboptimal attention distribution on visual inputs, leading to hallucinated or inaccurate outputs. Existing mitigation methods primarily rely on inference-time interventions, which are limited in attention adaptation or require additional supervision. To address this, we propose A3TUNE, a novel fine-tuning framework for Automatic Attention Alignment Tuning. A3TUNE leverages zero-shot weak labels from SAM, refines them into prompt-aware labels using BiomedCLIP, and then selectively modifies visually-critical attention heads to improve alignment while minimizing interference. Additionally, we introduce a A3MOE module, enabling adaptive parameter selection for attention tuning across diverse prompts and images. Extensive experiments on medical VQA and report generation benchmarks show that A3TUNE outperforms state-of-the-art baselines, achieving enhanced attention distributions and performance.
AB - Medical Large Vision-Language Models (Med-LVLMs) often exhibit suboptimal attention distribution on visual inputs, leading to hallucinated or inaccurate outputs. Existing mitigation methods primarily rely on inference-time interventions, which are limited in attention adaptation or require additional supervision. To address this, we propose A3TUNE, a novel fine-tuning framework for Automatic Attention Alignment Tuning. A3TUNE leverages zero-shot weak labels from SAM, refines them into prompt-aware labels using BiomedCLIP, and then selectively modifies visually-critical attention heads to improve alignment while minimizing interference. Additionally, we introduce a A3MOE module, enabling adaptive parameter selection for attention tuning across diverse prompts and images. Extensive experiments on medical VQA and report generation benchmarks show that A3TUNE outperforms state-of-the-art baselines, achieving enhanced attention distributions and performance.
UR - https://www.scopus.com/pages/publications/105021053985
UR - https://www.scopus.com/pages/publications/105021053985#tab=citedBy
U2 - 10.18653/v1/2025.acl-long.460
DO - 10.18653/v1/2025.acl-long.460
M3 - Conference contribution
AN - SCOPUS:105021053985
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 9357
EP - 9372
BT - Long Papers
A2 - Che, Wanxiang
A2 - Nabende, Joyce
A2 - Shutova, Ekaterina
A2 - Pilehvar, Mohammad Taher
PB - Association for Computational Linguistics (ACL)
Y2 - 27 July 2025 through 1 August 2025
ER -