VLDB 2026 Research / reviewers in the wild / expert
Yun Xue 0002
dblp:39/4919-2
· DBLP profile ↗
35ranked-venue papers
0as first author
30since 2021 · last 2026
0000-0002-4048-5298ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 19 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DG-MCTS: Dual-Guidance Monte Carlo Tree Search for Adaptive Emotional Support Dialogue Planning
Benshuo Lin, Yuelei Li, Yun Xue 0002, Yiping Song |
WWW | 5 |
| 2026 | Multi-view dynamic perception framework for Chinese harmful meme detection
Jiapei Hu, Kuntao Li, Yun Xue 0002, Jinghua Liang |
Inf. Process. Manag. | 5 |
| 2025 | Advancing Collaborative Debates with Role Differentiation through Multi-Agent Reinforcement LearningabstractMulti-agent collaborative tasks exhibit exceptional capabilities in natural language applications and generation. By prompting agents to assign clear roles, it is possible to facilitate cooperation and achieve complementary capabilities among LLMs. A common strategy involves adopting a relatively general role assignment mechanism, such as introducing a “judge” or a “summarizer”. However, these approaches lack task-specific role customization based on task characteristics. Another strategy involves decomposing the task based on domain knowledge and task characteristics, followed by assigning appropriate roles according to LLMs’ respective strengths, such as programmers and testers. However, in some given tasks, obtaining domain knowledge related to task characteristics and getting the strengths of different LLMs is hard. To solve these problems, we propose a Multi-LLM Cooperation (MLC) framework with automatic role assignment capabilities. The core idea of the MLC is to initialize role assignments randomly and then allow the role embeddings to be learned jointly with the downstream task. To capture the state transitions of multiple LLMs during turn-based speaking, the role embedding is sequence-aware. At the same time, to avoid role convergence, the role differentiation module in MLC encourages behavioral differentiation between LLMs while ensuring the LLM team consistency, guiding different LLMs to develop complementary strengths from the optimization level. Our experiments on seven datasets demonstrate that MLC significantly enhances collaboration and expertise, which collaboratively addresses multi-agent tasks. Ziyi Su, Yun Xue 0002, Zhiliang Tian, Yiping Song, Minlie Huang |
ACL (1) | 3 |
| 2025 | MSG-LLM: A Multi-scale Interactive Framework for Graph-enhanced Large Language ModelsabstractGraph-enhanced large language models (LLMs) leverage LLMs’ remarkable ability to model language and use graph structures to capture topological relationships. Existing graph-enhanced LLMs typically retrieve similar subgraphs to augment LLMs, where the subgraphs carry the entities related to our target and relations among the entities. However, the retrieving methods mainly focus solely on accurately matching subgraphs between our target subgraph and the candidate subgraphs at the same scale, neglecting that the subgraphs with different scales may also share similar semantics or structures. To tackle this challenge, we introduce a graph-enhanced LLM with multi-scale retrieval (MSG-LLM). It captures similar graph structures and semantics across graphs at different scales and bridges the graph alignment across multiple scales. The larger scales maintain the graph’s global information, while the smaller scales preserve the details of fine-grained sub-structures. Specifically, we construct a multi-scale variation to dynamically shrink the scale of graphs. Further, we employ a graph kernel search to discover subgraphs from the entire graph, which essentially achieves multi-scale graph retrieval in Hilbert space. Additionally, we propose to conduct multi-scale interactions (message passing) over graphs at various scales to integrate key information. The interaction also bridges the graph and LLMs, helping with graph retrieval and LLM generation. Finally, we employ a Chain-of-Thought-based LLM prediction to perform the downstream tasks. We evaluate our approach on two graph-based downstream tasks and the experimental results show that our method achieves state-of-the-art performance. Zhangkai Zheng, Benshuo Lin, Yun Xue 0002, Yiping Song |
COLING | 4 |
| 2025 | Ambiguity-aware Multi-level Incongruity Fusion Network for Multi-Modal Sarcasm DetectionabstractMulti-modal sarcasm detection aims to identify whether a given image-text pair is sarcastic. The pivotal factor of the task lies in accurately capturing incongruities from different modalities. Although existing studies have achieved impressive success, they primarily committed to fusing the textual and visual information to establish cross-modal correlations, overlooking the significance of original unimodal incongruity information at the text-level and image-level. Furthermore, the utilized fusion strategies of cross-modal information neglected the effect of inherent ambiguity within text and image modalities on multimodal fusion. To overcome these limitations, we propose a novel Ambiguity-aware Multi-level Incongruity Fusion Network (AMIF) for multi-modal sarcasm detection. Our method involves a multi-level incongruity learning module to capture the incongruity information simultaneously at the text-level, image-level and cross-modal-level. Additionally, an ambiguity-based fusion module is developed to dynamically learn reasonable weights and interpretably aggregate incongruity features from different levels. Comprehensive experiments conducted on a publicly available dataset demonstrate the superiority of our proposed model over state-of-the-art methods. Kuntao Li, Qiaofeng Wu, Weixing Mai, Fenghuan Li, Yun Xue 0002 |
COLING | 6 |
| 2025 | Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question AnsweringabstractThe knowledge-based visual question answering (KB-VQA) task involves using external knowledge about the image to assist reasoning. Building on the impressive performance of multimodal large language model (MLLM), recent methods have commenced leveraging MLLM as an implicit knowledge base for reasoning. However, the direct employment of MLLM with raw external knowledge might result in reasoning errors due to misdirected knowledge information. Additionally, MLLM may lack fine-grained perception of visual features, which can result in hallucinations during reasoning. To address these challenges, we propose Notes-guided MLLM Reasoning (NoteMR), a novel framework that guides MLLM in better reasoning by utilizing knowledge notes and visual notes. Specifically, we initially obtain explicit knowledge from an external knowledge base. Then, this explicit knowledge, combined with images, is used to assist the MLLM in generating knowledge notes. These notes are designed to filter explicit knowledge and identify relevant internal implicit knowledge within the MLLM. We then identify highly correlated regions between the images and knowledge notes, retaining them as image notes to enhance the model’s fine-grained perception, thereby mitigating MLLM induced hallucinations. Finally, both notes are fed into the MLLM, enabling a more comprehensive understanding of the image-question pair and enhancing the model’s reasoning capabilities. Our method achieves state-of-the-art performance on the OK-VQA and A-OKVQA datasets, demonstrating its robustness and effectiveness across diverse VQA scenarios. Wenlong Fang, Qiaofeng Wu, Yun Xue 0002 |
CVPR | 4 |
| 2025 | SACR: Self-training with Saliency-Augmented Consistency Regularization for Few-Shot LearnersabstractPre-trained language models have made significant strides in natural language processing tasks, enabling flexible fine-tuning for downstream applications. However, in few-shot learning scenarios, pre-trained models face challenges related to overfitting due to limited training samples, which hinders their ability to capture data diversity and robustly handle input variations. To overcome these limitations, we propose the Saliency-Augmented Consistency Regularization (SACR) framework, a novel self-training strategy designed to improve few-shot learning performance and robustness. SACR consists of three key components: (1) Saliency-guided Data Perturbation, which uses saliency analysis to identify and perturb words that significantly influence model predictions, generating semantically consistent pseudo-samples; (2) Semantic-equivalent Sample Mining, which employs K-means clustering to select semantically similar samples and prevent semantic shift; and (3) Consistency Training, where regularization ensures consistency between the semantics and prediction distributions of the original and perturbed samples. We conducted extensive experiments on 15 public datasets to evaluate our approach. SACR significantly outperforms strong baseline models and demonstrates superior generalization capabilities compared to state-of-the-art methods, achieving an average 1.0% improvement in classification accuracy. Yanyan Feng, Yue Zhou 0012, Yun Xue 0002, Fenghuan Li, Zehong Lin |
ICASSP | 3 |
| 2025 | ReHyGen: Relational hypergraph enhanced generative aspect sentiment triplet extractionabstractAspect Sentiment Triplet Extraction (ASTE) has emerged as a pivotal task in sentiment analysis , focusing on extracting the aspect terms along with the corresponding opinion terms and the expressed sentiments. Recently, generative models have achieved significant success in ASTE task. However, existing generative approaches fail to further model the specific relations within the context for ASTE at the encoding phase, making it difficult to establish the nuanced connections between aspect and opinion terms. Additionally, these approaches rely on simple structured templates at the decoding phase to pair aspect terms with opinion terms, which fails to provide effective relation information for the decoding process. To address the aforementioned issues, we propose ReHyGen, a novel relational hypergraph enhanced framework designed to enhance the relational modeling capabilities of generative ASTE models during both the encoding and decoding phases. Specifically, ReHyGen comprises two core components: the Relational Hypergraph Enhanced Module (RHEM) and the Relational Prompt Module (RPM). RHEM leverages the hypergraph attention network and auxiliary relation classification to capture high-order word interactions and boundary-sensitive word pair relations. RPM incorporates relational information into the decoding phase by providing relation-aware prompts, guiding the generation of more accurate target sequences. Extensive experiments on benchmark datasets demonstrate that our proposed framework significantly improve the performance of generative ASTE models. Zehong Lin, Weibo Chen, Yun Xue 0002, Fenghuan Li |
Neurocomputing | 3 |
| 2025 | Dual-level adaptive incongruity-enhanced model for multimodal sarcasm detection
Qiaofeng Wu, Wenlong Fang, Weiyu Zhong, Fenghuan Li, Yun Xue 0002, Bo Chen 0004 |
Neurocomputing | 5 |
| 2025 | Knowledge based attribute completion for heterogeneous graph node classification
Zhangkai Zheng, Yun Xue 0002, Yiping Song, Zhuoming Liang |
Neurocomputing | 3 |
| 2025 | Semantic enhanced bi-syntactic graph convolutional network for aspect-based sentiment analysis
Junyang Xiao, Yun Xue 0002, Fenghuan Li |
Inf. Sci. | 2 |
| 2025 | VS-MRC: A visual semantics-guided machine reading comprehension framework for multimodal named entity recognition with multiple images
Jiapei Hu, Yun Xue 0002 |
Knowl. Based Syst. | 3 |
| 2024 | Semantics-Aware Dual Graph Convolutional Networks for Argument Pair ExtractionabstractArgument pair extraction (APE) is a task that aims to extract interactive argument pairs from two argument passages. Generally, existing works focus on either simple argument interaction or task form conversion, instead of thorough deep-level feature exploitation of argument pairs. To address this issue, a Semantics-Aware Dual Graph Convolutional Networks (SADGCN) is proposed for APE. Specifically, the co-occurring word graph is designed to tackle the lexical and semantic relevance of arguments with a pre-trained Rouge-guided Transformer (ROT). Considering the topic relevance in argument pairs, a topic graph is constructed by the neural topic model to leverage the topic information of argument passages. The two graphs are fused via a gating mechanism, which contributes to the extraction of argument pairs. Experimental results indicate that our approach achieves the state-of-the-art performance. The performance on F1 score is significantly improved by 6.56% against the existing best alternative. Minzhao Guan, Zhixun Qiu, Fenghuan Li, Yun Xue 0002 |
LREC/COLING | 4 |
| 2024 | Representation and Granularity Joint Alignment Framework for Multimodal Sarcasm Detection on Social Media
Jiapei Hu, Yun Xue 0002, Fenghuan Li, Qianhua Cai |
DASFAA (7) | 3 |
| 2024 | D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment DetectionabstractMultimodal sentiment detection aims to classify the sentiment polarity of a given imagetext pair.Existing approaches apply the same fixed framework to all input samples, lacking the flexibility to adapt to different image-text pairs.Furthermore, the interaction patterns of these methods are overly homogenized, limiting the model's capacity to extract multimodal sentiment information effectively.In this paper, we develop a Dual-Branch Dynamic Routing Network (D 2 R), which is the first multimodal dynamic interaction model towards multimodal sentiment detection.Specifically, we design six independent units to simulate inter-and intramodal information interactions without depending on any existing fixed frameworks.Additionally, we configure a soft router in each unit to guide path generation and introduce the path regularization term to optimize these inference paths.Comprehensive experiments on three publicly available datasets demonstrate the superiority of our proposed model over state-ofthe-art methods. Kuntao Li, Weixing Mai, Qiaofeng Wu, Yun Xue 0002, Fenghuan Li |
EMNLP | 5 |
| 2024 | Prompt-enhanced Network for Hateful Meme Classification
Junxi Liu, Yanyan Feng, Jiehai Chen, Yun Xue 0002, Fenghuan Li |
IJCAI | 4 |
| 2024 | GeDa: Improving training data with large language models for Aspect Sentiment Triplet Extraction
Weixing Mai, Zhengxuan Zhang, Kuntao Li, Yun Xue 0002 |
Knowl. Based Syst. | 5 |
| 2024 | Enhanced Syntactic and Semantic Graph Convolutional Network With Contrastive Learning for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) aims to predict the sentiment polarity of a given specific aspect in the sentence. Recent studies focus on leveraging graph convolutional neural networks to encode both syntactic and semantic information. However, current syntactic parsers, which are not specifically for ABSA, introduce noise to the syntactic information. Besides, ongoing studies ignore the distinctiveness of semantics and syntax. To address these issues, we proposed an enhanced syntactic and semantic graph convolutional network (GCN) with contrastive learning in this article. An aspect-oriented syntactic graph is constructed with aspect-specific perturbed masking for reducing the syntactic noise, and a semantic graph is established with self-attention weights from bidirectional encoder representation from transformers (BERT). The semantic and syntactic representations are further enhanced by both sentiment polarity-based supervised contrastive learning and syntactic reliability-based unsupervised contrastive learning. Furthermore, label embeddings of syntactic reliability are learned to determine the weights of syntactic and semantic information. Extensive experiments on four publicly available datasets demonstrate that our model is more competitive than the state-of-the-arts. Minzhao Guan, Fenghuan Li, Yun Xue 0002 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | End-to-End Visual Grounding Framework for Multimodal NER in Social Media PostsabstractMultimodal named entity recognition (MNER) for social media aims to detect named entities in user-generated posts with the aid of visual information from attached images. Existing methods use pretrained visual models or visual grounding (VG) toolkits to learn visual information. However, they still suffer from the mismatch issue, where the visual features extracted from visual encoder are inconsistent with actual requirements for cross-modal interaction. In an ideal scenario, the visual encoder should actively extract visual information guided by the text, which inherently provides the blueprint of desired visual features. In this article, we present an end-to-end VG framework for MNER task (VG-MNER), which adaptively learns the text-related visual features. Specifically, we introduce a backbone network with a feature fusion module to learn and aggregate multisize visual representations. We then develop a text-related visual attention to refine the visual features. Notably, entity-image contrast loss is designed to guide the training of visual encoder. The proposed model outperforms several state-of-the-art methods, achieving F1 scores of 75.62% and 88.11% on two benchmark datasets. Experimental results reveal the effectiveness of leveraging text-related visual information in the MNER task. Jiapei Hu, Yun Xue 0002, Qianhua Cai |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Dynamic Graph Construction Framework for Multimodal Named Entity Recognition in Social MediaabstractMultimodal named entity recognition (MNER) aims to detect named entities and identify the entity types based on texts and attached images, which also generates inputs for other comprehensive tasks, such as multimodal machine translation, visual dialog, and multimodal sentiment analysis. Existing studies have limitations in text-image matching and multimodal semantic disparity reduction. For one thing, current methods fail to resolve both overall and local text-image matching issues in a self-guided way. For another, the static graphs constructed in MNER models are challenging in bridging the semantic gap between different modalities. In this work, a dynamic graph construction framework (DGCF) is proposed to solve the above-mentioned limitations. A similarity vector-based text-image matching inferring strategy is designed to obtain the overall and local matching relation between text and image while the overall matching determines the retained proportion of visual information. Then, a multimodal dynamic graph interaction module is developed. Within each layer of the module, the local matching relations and part of speech (POS)-based multihead attention are integrated to construct a dynamic cross-modal graph and a semantic graph. Lastly, a CRF layer is used to predict entity label. Extensive experiments are performed on two benchmark datasets. The experimental results reveal that our model is a competitive alternative and achieves state-of-the-art performance. Weixing Mai, Zhengxuan Zhang, Kuntao Li, Yun Xue 0002, Fenghuan Li |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Modeling Inter-Aspect Relations With Clause and Contrastive Learning for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) is a fine-grained sentiment analysis task that aims to identify the sentiment polarity of the given aspect. Recent studies fail to establish the relation among multiple aspects in one sentence. To address this issue, a clause-level relational graph attention network with contrastive learning (CLRCL) model is proposed. Specifically, the given sentence is segmented into clauses to obtain the relation between two aspects based on clause-level interaction. Then, to integrate multiple-aspect information, a clause-level relational graph which contains all aspects and inter-aspect relations is developed. Notably, to precisely learn the inter-aspect relations, the supervised contrastive learning strategy is used. Experimental results reveal that the proposed model is a competitive alternative compared with the state-of-the-art methods. Zhixun Qiu, Kehai Chen, Yun Xue 0002, Zhengxuan Zhang |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | A Token-wise Graph-based Framework for Multimodal Named Entity RecognitionabstractMultimodal Named Entity Recognition (MNER) on social media posts is a leading but challenging task. However, most existing MNER methods fail to effectively exploit the visual information from the image. Besides, the multimodal interaction and alignment remains unsettled. In this paper, we propose a novel token-wise graph-based framework to deal with the MNER task. Specifically, a token-wise image processing manner is established. A muti-modal graph is constructed based on the textual token derived from BERT and the visual token derived from SwinT. Then, the muti-modal graph is fed into a multi-layer Transformer-based module for intra- and inter-modal information fusion. In addition, multiple contrastive learning is devised to perform the global and local alignment between textual and visual nodes. Experimental results on two benchmark multimodal datasets indicate that our model achieves state-of-the-art performance in MNER tasks. Zhengxuan Zhang, Weixing Mai, Haoliang Xiong, Chuhan Wu, Yun Xue 0002 |
ICME | 5 |
| 2023 | Span-Based Pair-Wise Aspect and Opinion Term Joint Extraction with Contrastive Learning
Jinjie Yang, Feipeng Dai, Fenghuan Li, Yun Xue 0002 |
NLPCC (2) | 4 |
| 2022 | Exploring fine-grained syntactic information for aspect-based sentiment classification with dual graph neural networks
Luwei Xiao, Yun Xue 0002, Hua Wang 0002, Donghong Gu, Yongsheng Zhu |
Neurocomputing | 2 |
| 2022 | Aspect-Level Sentiment Analysis with Local Semantic and Global Syntactic Features IntegrationabstractAspect-level sentiment analysis aims to predict the sentiment polarity toward a specific aspect in a sentence. Most current approaches are based on deep learning and the attention mechanism. However, these models cannot simultaneously include the context semantic information carried by local words and the global syntactic information possibly carried by remote words. In this paper, we propose a local semantic and global syntactic integration scheme, which employs a local focus mechanism over local context words and exploits improved graph convolutional networks over dependency tree to encode global syntactic information. Moreover, multi-head attention is used to capture both the semantic information and also the interactive information between semantics and syntactic features. Experimental results on five datasets show the effectiveness of our model over a series of latest models. Luwei Xiao, Yue-Cai Huang, Yun Xue 0002, Haoliang Zhao |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2022 | Multi-head self-attention based gated graph convolutional networks for aspect-based sentiment classification
Luwei Xiao, Yun Xue 0002, Bingliang Chen, Donghong Gu, Bixia Tang |
Multim. Tools Appl. | 4 |
| 2021 | Aspect-Based Sentiment Analysis Using Graph Convolutional Networks and Co-attention Mechanism
Zhaowei Chen, Yun Xue 0002, Luwei Xiao, Hao Lan Zhang 0001 |
ICONIP (6) | 2 |
| 2021 | Bilateral-brain-like Semantic and Syntactic Cognitive Network for Aspect-level Sentiment AnalysisabstractAspect-level sentiment analysis (ALSA) is a fine-grained task for classifying the sentiment polarity of a specific aspect in a sentence. In spite of the progress in deep-learning algorithms within natural language processing domain, the methods in line with human cognition are absent. Inspired by the processing principle of our brain, we propose a Bilateral-brain-like Semantic and Syntactic Cognitive Network (BSSCN) for aspect-level sentiment analysis. There are four major modules established in BSSCN, which are left hemisphere semantic activation (LH-SA), semantic selection (SS) module, semantic integration (SI) module and right hemisphere semantic activation (RH-SA). Experimental results on a variety of datasets show that our model stably outperforms the widely-used methods, which establishes a strong evidence of the effectiveness in ALSA tasks. Zuohua Huang, Yun Xue 0002 |
IJCNN | 3 |
| 2021 | A Novel Bi-Branch Graph Convolutional Neural Network for Aspect Level Sentiment ClassificationabstractAspect-level sentiment classification is a fine-grained task in sentiment analysis whose main purpose is to identify the sentiment polarity of a specific aspect. Current Graph Convolutional Network (GCN) has its distinctive superiority in tackling sentiment classification both semantically and syntactically. However, GCN still has deficiencies in introducing the noise during processing and dealing with sentences of complex structure. To address these issues, we propose a novel Bi-branch GCN (Bi-B GCN). In our model, an attention weight graph, by employing the attention mechanism, is constructed to substitute the basic syntax dependency tree and thus to remove the irrelevant information. Furthermore, a semantic dependency graph is devised to supplement the semantic information to the syntax dependency tree, based on which the connection between different words can be captured. In addition, on the task of sentiment classification, the integration of semantic information and the syntactic information is conducted by using a combinational gated mechanism. Substantial experiments to validate the working performance of Bi-B GCN are performed on a variety of datasets. The encouraging results establish a strong evidence of the high accuracy of the proposed model. Bingliang Chen, Guojun Lu, Yun Xue 0002, Qianhua Cai |
IJCNN | 3 |
| 2021 | SIntactical Distance Attention Guided Graph Convolutional Network for aspect-based sentiment analIsisabstractAspect-based sentiment analIsis (ABSA) aims to detect the sentiment polaritI of a specific aspect in an opinionated sentence. Current work focuses on exploiting the sIntactic tree to shorten the distance between the aspect term and context words. However, the “hard-pruning” strategI on the sIntactic tree maI lead to the reduction of importa nt sIntactic information. In this paper, we propose a novel sInt actical distance attention guided graph convolutional network (SDGCN) for ABSA. Our model is capable of fullI exploiting the sIntactic knowledge with a “soft pruning” strategI and learning crucial fine-grain sIntactic distance info rmation. AdditionallI, an effective denselI connected graph convolutional laIer is applied to avoid the over-sm oothing problem of standard GCN. Experiments conducted on three benchmark datasets show that our model achieves promising results comparing to the baseline models. Luwei Xiao, Donghong Gu, Yun Xue 0002, Yongsheng Zhu |
IJCNN | 3 |
| 2019 | Co-attention Networks for Aspect-Level Sentiment Analysis
Haihui Li, Yun Xue 0002, Hongya Zhao, Sancheng Peng |
NLPCC (2) | 2 |
| 2019 | Shared-Private LSTM for Multi-domain Text Classification
Haiming Wu, Yue Zhang 0004, Yun Xue 0002, Ziwen Wang 0004 |
NLPCC (2) | 4 |
| 2019 | A novel feature extraction methodology for sentiment analysis of product reviews
Xin Chen 0046, Yun Xue 0002, Hongya Zhao |
Neural Comput. Appl. | 2 |
| 2016 | Classification of Epileptic EEG Signals with Stacked Sparse Autoencoder Based on Deep Learning
Qin Lin 0002, Shuqun Ye, Xiu-mei Huang, Si-you Li, Meizhen Zhang, Yun Xue 0002 |
ICIC (3) | 6 |
| 2011 | Spoken arabic digits recognition based on wavelet neural networksabstractThe paper describes a novel method for discrete speech recognition based on spoken Arabic digit recognition by means of wavelet neural network in which Morlet wavelet is introduced to the hidden layer. The speech signal is extracted by means of Mel Frequency Cepstral Coefficients (MFCCs) and followed by vector quantization (VQ). The experimental results obtained on a spoken Arabic digit dataset proved that it could achieve better accuracy and need less learning time than the proposed method. Lvjun Zhan, Yun Xue 0002, Weixing Zhou, Liangjun Zhang |
SMC | 3 |