| SSDOC-DET: efficient document layout detection via selective state-space modeling | |
|---|---|
| 學年 | 115 |
| 學期 | 1 |
| 發表日期 | 2026-09-13 |
| 作品名稱 | SSDOC-DET: efficient document layout detection via selective state-space modeling |
| 作品名稱(其他語言) | |
| 著者 | Jing-Ming Guo; Chun-Wei Huang; Yi-Chong Zeng; Cheng-Yen Hsiao |
| 作品所屬單位 | |
| 出版者 | |
| 會議名稱 | 2026 IEEE International Conference in Image Processing (ICIP) |
| 會議地點 | Tampere, Finland |
| 摘要 | This paper explores the efficiency–accuracy trade-off in high-resolution document layout detection, introducing the Selective State-space Document Detector (SSDoc-Det), a compact detection framework designed for low-latency, resource-efficient deployment. Our method incorporates selective two-dimensional state-space modeling into both the feature extraction backbone and a hybrid encoder, enabling long-range dependency modeling while avoiding the quadratic overhead of self-attention. To facilitate effective multi-scale feature integration, a bottleneck-style convolutional alignment module allows the decoding stage to be streamlined with fewer layers. For bounding-box localization, SSDoc-Det replaces conventional IoU variants with MPDIoU and retains distribution-aware refinement and self-distillation mechanisms to enhance regression stability and precision. Extensive experiments on standard document layout benchmarks demonstrate that the proposed framework achieves strong detection accuracy while substantially reducing computational complexity and inference latency. These results show that combining selective state-space modeling with lightweight feature fusion and refined localization yields an effective solution for accurate, efficient document layout detection. |
| 關鍵字 | |
| 語言 | en_US |
| 收錄於 | |
| 會議性質 | 國際 |
| 校內研討會地點 | 無 |
| 研討會時間 | 20260913~20260917 |
| 通訊作者 | Yi-Chong Zeng |
| 國別 | FIN |
| 公開徵稿 | |
| 出版型式 | |
| 出處 | IEEE |
| 相關連結 |
機構典藏連結 ( http://tkuir.lib.tku.edu.tw:8080/dspace/handle/987654321/129920 ) |