| Structured Cross-Modal Representation Alignment via Dual Mixing Blocks in Vision–Language Architectures | |
|---|---|
| 學年 | 114 |
| 學期 | 2 |
| 發表日期 | 2026-07-01 |
| 作品名稱 | Structured Cross-Modal Representation Alignment via Dual Mixing Blocks in Vision–Language Architectures |
| 作品名稱(其他語言) | |
| 著者 | Zong-Yan Yang; Jing-Ming Guo; Yi-Chong Zeng; Ze-Wen Chen; Wei-Hsiang Huang |
| 作品所屬單位 | |
| 出版者 | |
| 會議名稱 | IEEE Int'l Conf. on Consumer Electronics-Taiwan 2026 (ICCE-TW) |
| 會議地點 | Taoyuan, Taiwan |
| 摘要 | Vision-language models (VLMs) often use shallow multilayer perceptron (MLP) connectors to project visual features into the language embedding space, limiting cross-modal reasoning. We propose a structured connector, Dual Mixing Block, that models spatial and channel interactions via lightweight MLP transformations with residual connections. The design improves cross-modal expressiveness while being attention-free. Built on SigLIP and LLaMA3.1-8B with a comparable training scale (1.2 M samples), the model achieves an average gain of +1.8 points over LLaVA-1.5 (Vicuna-13B) on four VQA benchmarks, highlighting the importance of connector design in multimodal reasoning. |
| 關鍵字 | |
| 語言 | en_US |
| 收錄於 | |
| 會議性質 | 國際 |
| 校內研討會地點 | 無 |
| 研討會時間 | 20260701~20260703 |
| 通訊作者 | Yi-Chong Zeng |
| 國別 | TWN |
| 公開徵稿 | |
| 出版型式 | |
| 出處 | IEEE |
| 相關連結 |
機構典藏連結 ( http://tkuir.lib.tku.edu.tw:8080/dspace/handle/987654321/129918 ) |