Temporal Context Modeling for Multi-Class Voice Pathology Detection via BiLSTM and MFCCs
學年 114
學期 2
發表日期 2026-05-27
作品名稱 Temporal Context Modeling for Multi-Class Voice Pathology Detection via BiLSTM and MFCCs
作品名稱(其他語言)
著者 Vainawin Piyamitmongkol;Chii-Jen Chen
作品所屬單位
出版者
會議名稱 The International Conference on Recent Advancements in Computing in AI, IoT and Computer Engineering Technology (CICET 2026)
會議地點 New Taipei, Taiwan
摘要 Voice pathology detection is pivotal for non-invasive medical diagnosis. While traditional approaches rely on binary classification, clinical practice demands differentiation among multiple specific disorders. This paper proposes a Bidirectional Long Short-Term Memory (BiLSTM) framework combined with Mel-Frequency Cepstral Coefficients (MFCCs) for multi-class voice pathology classification. Unlike standard Recurrent Neural Networks (RNNs) or purely spatial models like Residual Networks (ResNet), the proposed BiLSTM leverages bidirectional temporal dependencies to capture both past and future acoustic contexts. Experimental evaluations conducted on a dataset of 19,368 audio samples demonstrate that the BiLSTM model achieves superior performance with an accuracy of 98%, significantly outperforming RNN and ResNet baselines. The results validate the efficacy of bidirectional temporal modeling in identifying complex pathological voice patterns.
關鍵字 Voice Pathology, Deep Learning, BiLSTM, MFCC, Multi-Class Classification
語言 en_US
收錄於
會議性質 國際
校內研討會地點 淡水校園
研討會時間 20260527~20260529
通訊作者 Chii-Jen Chen
國別 TWN
公開徵稿
出版型式
出處