Structured Cross-Modal Representation Alignment via Dual Mixing Blocks in Vision–Language Architectures
學年 114
學期 2
發表日期 2026-07-01
作品名稱 Structured Cross-Modal Representation Alignment via Dual Mixing Blocks in Vision–Language Architectures
作品名稱(其他語言)
著者 Zong-Yan Yang; Jing-Ming Guo; Yi-Chong Zeng; Ze-Wen Chen; Wei-Hsiang Huang
作品所屬單位
出版者
會議名稱 IEEE Int'l Conf. on Consumer Electronics-Taiwan 2026 (ICCE-TW)
會議地點 Taoyuan, Taiwan
摘要 Vision-language models (VLMs) often use shallow multilayer perceptron (MLP) connectors to project visual features into the language embedding space, limiting cross-modal reasoning. We propose a structured connector, Dual Mixing Block, that models spatial and channel interactions via lightweight MLP transformations with residual connections. The design improves cross-modal expressiveness while being attention-free. Built on SigLIP and LLaMA3.1-8B with a comparable training scale (1.2 M samples), the model achieves an average gain of +1.8 points over LLaVA-1.5 (Vicuna-13B) on four VQA benchmarks, highlighting the importance of connector design in multimodal reasoning.
關鍵字
語言 en_US
收錄於
會議性質 國際
校內研討會地點 無
研討會時間 20260701~20260703
通訊作者 Yi-Chong Zeng
國別 TWN
公開徵稿
出版型式
出處 IEEE
相關連結

機構典藏連結 ( http://tkuir.lib.tku.edu.tw:8080/dspace/handle/987654321/129918 )