Xinyuan Zhou

dblp:210/1336 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Small object detection using multi-scale detail enhancement and decoupled detection head
Yixin Qiao, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Yao Li 0019, Guonan Deng
Neurocomputing2
2026 STGFMamba: Spatio-temporal graph Fourier-enhanced Mamba for traffic prediction
Xinyuan Zhou, Ruiyi Lu, Zhiang Hou, Yao Ren, Wenwu Wang 0001, Shiyong Lan
Inf. Sci.2
2026 Evolutionary multitasking combined with migration estimation distribution and its application on power electronics system
Xinyuan Zhou, Xuemin Ma, Yue Pan 0016
Inf. Sci.2
2026 DSTFGCN: A dynamic spatial-temporal fusion graph convolution network for traffic flow forecasting
abstract
Traffic flow prediction is one of the core technology of Intelligent Transportation System. Its fundamental challenge is to effectively model the complex spatial-temporal dependencies. Although extensive research has been conducted in this field, the limitations of current methods restrict their effectiveness in accurate predictions. For temporal dependence, existing methods based on recurrent neural networks only focus on local dependencies and ignore global dependencies. For spatial dependencies, existing methods use predefined or adaptive adjacency matrices that cannot accurately reflect the relationships between real traffic flow. To overcome these limitations, we propose a dynamic spatial-temporal fusion graph convolution network (DSTFGCN). In the temporal aspect, we introduce gated dilated causal convolution to capture the local dependencies and node-independent temporal graph convolution to capture the global dependencies specific to each node. In the spatial aspect, we propose a dynamic graph convolution block. It can construct dynamic graphs based on the characteristics of the input data and aggregate both local and global spatial dependencies. Experiments on six real-world datasets have shown that DSTFGCN outperforms current mainstream methods. The codes are available at https://github.com/SYLan2019/DSTFGCN.
Tianyi Pan, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Hongyu Yang 0002, Zhiang Hou, Yao Ren
Neural Networks2
2025 Scalable Data Synthesis through Human-like Cognitive Imitation and Data Recombination
abstract
Large language models (LLMs) rely on massive amounts of training data, however, the quantity of empirically observed data is limited.To alleviate this issue, lots of LLMs leverage synthetic data to enhance the quantity of training data.Despite significant advancements in LLMs, the efficiency and scalability characteristics of data synthesis during pre-training phases remain insufficiently explored.In this work, we propose a novel data synthesis framework, Cognitive Combination Synthesis (CCS), designed to achieve highly efficient and scalable data synthesis.Specifically, our methodology mimics human cognitive behaviors by recombining and interconnecting heterogeneous data from diverse sources thereby enhancing advanced reasoning capabilities in LLMs.Extensive experiments demonstrate that: (1) effective data organization is essential, and our mappingbased combination learning approach significantly improves data utilization efficiency; (2) by enhancing data diversity, accuracy, and complexity, our synthetic data scales beyond 100B tokens, revealing CCS's strong scalability.Our findings highlight the impact of data organization methods on LLM learning efficiency and the significant potential of scalable synthetic data to enhance model reasoning capabilities.
Zhongyi Ye, Weitai Zhang, Xinyuan Zhou, Ninghui Rao, Enhong Chen
EMNLP3
2025 SAR Ship Detector Using Cross-stage Feature Fusion and Decoupled Head with Mutual Guidance
abstract
Deep learning-based SAR ship detection methods enhance resilience to noise, distortion, and interference in ocean environments, establishing them as the foremost approach for ship detection nowadays. Nonetheless, substantial difficulties persist in separating ships from the complex backgrounds found in SAR images: 1) Many ships are highly similar to the sea surface clutter noise, making them susceptible to false alarms; 2) Ship targets exhibit a wide range of variations in size and shape. In this paper, we propose a novel network to address above problems. Firstly, the deformable convolution is incorporated into the backbone network to adapt to the wide-range of ship shapes. Secondly, the cross-stage feature fusion module (CSFFM) is introduced to realize local self-supervised interaction between two adjacent layers, thereby reducing the impact of receptive field differences between different feature layers and mitigating the influence of complex background noise. Finally, the mutually guided decoupled-head (MGDH) is designed to achieve mutual guidance between classification and regression, thus further enhancing the significant regions of the feature maps. Through extensive experiments, it has been verified that our proposed method has achieved the most promising performance compared to well-known baselines. The codes will be available at https://github.com/SYLan2019/CSFF-MGDH.
Yixin Qiao, Xiaoxiao Yin, Xinyuan Zhou, Shiyong Lan, Guonan Deng
ICASSP3
2025 Bridging Modality Gap with Large Speech and Language Models for End-to-End Speech-to-Text Translation
abstract
End-to-end speech-to-text translation (E2E ST) has increasingly aroused interest and attention recently, attempting to address the problem of data scarcity and modeling burden. Several attempts exploring the combination of Large Speech and Language Models into a unified model to improve E2E ST are carried out. However, the inherent differences between speech and text modalities often impede effective cross-modal and cross-lingual transfer. In this study, we introduce LaSaLM-ST, a novel model architecture built upon Pre-trained Large Speech and Language Models for improving E2E ST. Our speech encoder begins with processing the source speech sequence. An adaptor and speech decoder then project speech features into the compatible feature space for the decoder-only Large Language Model (LLM), which then aligns the representation spaces of speech and text modalities with attentive interactions. Besides, we also develop a multi-step fine-tuning method to preserve the pre-trained multilingual knowledge and keep ST fine-tuning stably. Experiments conducted on the IWSLT2023 offline ST task from English to German, Chinese and Japanese demonstrate that our methodology not only achieves state-of-the-art BLEU scores but also outperforms the highly competitive cascaded ST systems in an unrestricted setting.
Weitai Zhang, Simran Naagar, Zhongyi Ye, Peiwang Tang, Xinyuan Zhou, Li-Rong Dai 0001
ICASSP5
2025 MRGNN: Mamba-Register-Based Graph Neural Network for Unsupervised Anomaly Detection in Multivariate Time Series
Danling Meng, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Weihong Yuan, Ruiyi Lu
ICIC (12)2
2025 Diff-DTF: Dynamic Temporal Feature Extraction and Refinement with Diffusion Model in Time Series Anomaly Detection
Ruiyi Lu, Weihong Yuan, Xinyuan Zhou, Shiyong Lan
PRICAI (5)3
2025 Spatiotemporal Trend Fusion Feature Graph Convolution Network for Spatial Interpolation in Traffic Scenes⋆
abstract
Sensors are always sparsely distributed in traffic networks due to high deployment costs, sensor damage, etc. Insufficient data may affect our perception of traffic scenarios, resulting in the inability of intelligent transportation systems (ITS) to efficiently perform traffic monitoring and scenario decisions. Spatial interpolation methods are used to infer the status of locations where no sensors are deployed. However, existing methods still have the following limitations: (1) Mainly interpolated unobserved nodes by extracting spatiotemporal dependencies between nodes, but ignored complex interaction patterns in feature dimensions. (2) Recent methods generally used standard TCN to extract temporal correlations, which is affected by abnormal data. (3) When extracting spatial correlations, the use of deep GCN layer can lead to over-smoothing problem. To mitigate these limitations, we propose a novel spatial interpolation model, namely, Spatiotemporal Trend Fusion Feature Graph Convolution Network (STFGCN). Specifically, a novel feature graph convolution network is used to capture complex interaction patterns. Secondly, the trend capture branch is used to alleviate the impact of abnormal data on TCN. Finally, a dual-stage spatial module is used to solve the degradation of detailed feature representations in deep GCN layer. Experimental results on six real traffic datasets demonstrate that our method outperforms state-of-the-art baseline models.
Zhiang Hou, Shiyong Lan, Xinyuan Zhou, Wujiang Zhu, Yao Ren
SMC3
2025 FE-HGAT: Frequency-Enhanced Hybrid Graph Attention Network For Traffic Prediction
abstract
Traffic flow prediction is crucial for urban traffic management and planning. However, although existing research has achieved promising results, most methods primarily focus on time-domain processing, with insufficient exploration of frequency-domain signal characteristics. Moreover, existing approaches often fail to effectively distinguish and simultaneously model the spatial dependencies between nearby and distant nodes. To address these issues, this paper proposes a Frequency-Enhanced Hybrid Graph Attention Network (FE-HGAT) for traffic flow prediction. Our approach employs a dynamic filter in the temporal dimension, which utilizes Fast Fourier Transform (FFT) to enhance key frequency-domain features, thereby better characterizing the temporal dependencies in traffic data. Besides, to capture spatial dependencies between nodes at varying distances, we design a dynamic threshold module to distinguish between nearby and distant nodes, employing external attention (EA) and a mixture-of-experts-enhanced graph attention network (MOE-GAT) to model local dependencies and long-distance semantic similarities, respectively. Experiments demonstrate that FE-HGAT outperforms existing baseline models on several public transportation datasets, validating its effectiveness in traffic forecasting. The code is available at https://github.com/ry123scuer/FE-HGAT.
Yao Ren, Wujiang Zhu, Shiyong Lan, Xinyuan Zhou, Hongyu Yang 0002, Zhiang Hou
SMC4
2025 Spatial Interpolation Based on Causal Spatiotemporal Modeling
abstract
The problem of spatial interpolation is a common challenge in fields such as traffic flow analysis. However, most existing methods directly utilize information from neighboring nodes to infer status of the unobserved location, without considering whether this information contains confounding factors or whether there is a true causal relationship, despite the fact that some unknown confounding factors are inevitably included in the data collection process. To address this, this paper proposes a Causal Attention Spatial Temporal Interpolation (CASI), which leverages causal relationships between nodes for spatial interpolation. The proposed CASI employs dilated convolutions and gating mechanisms to capture temporal dependencies, and introduces the Causal Spatiotemporal Attention (CSTA) mechanism, to uncover spatial causal dependencies between nodes. Subsequently, a novel Graph-Guided Feature Enhance Module (GFEM) is designed, which leverages causal probabilities from CSTA’s Gumbel-Softmax to weight the adjacency matrix in traditional GCN, composing GS-GCN, then adopts self-attention on temporal and spatial dimension respectively to further enhance the features from the improved GCN. We evaluate CASI on three real-world datasets, where it outperforms the optimal baselines across MAE, RMSE and MAPE by up to 3.7%, 1.7%, and 5.4% on PEMS04, 8.5%, 6.0%, and 5.9% on PEMS08, respectively. Ablation studies further demonstrated the effectiveness of the proposed modules in causal spatiotemporal modeling.
Shiyong Lan, Yao Ren, Weihong Yuan, Xinyuan Zhou, Zhiang Hou
SMC5
2025 LGRDet: A Light Object Detection Network for Gesture Recognition
abstract
ABSTRACT Gesture recognition plays a crucial role in Human‐Machine Interaction (HMI) by enabling interaction with systems without physical contact. Nevertheless, current gesture recognition methods encounter various challenges, including suboptimal lighting conditions, low detection rates, slow processing speeds, and occlusion from protective gloves, which can impede sensor capture of hand movements and consequently degrade recognition accuracy. To overcome the high computational cost and limited robustness observed in existing gesture recognition algorithms, this paper introduces LGRDet, a novel gesture recognition model. LGRDet enhances NanoDet‐Plus by integrating Coordinate Attention (CA) and Squeeze‐and‐Excitation (SE) attention mechanisms into its backbone network, thereby strengthening its capacity to capture long‐range spatial dependencies and effectively detect small target gestures. This enhancement is crucial for capturing the fine‐grained features of gestures, such as finger bends and palm shapes. Furthermore, the Filtration‐Fusion (FF) attention mechanism is incorporated into the original Path Aggregation Network (PAN) to optimize feature fusion across diverse scales. Our proposed LGRDet algorithm represents a notable improvement in gesture recognition accuracy, achieved while upholding a remarkably small model footprint. This characteristic makes LGRDet ideally suited for practical, real‐time gesture detection and recognition. Specifically, LGRDet achieved an accuracy of 92.4 for recognizing 9 gestures on our custom IHGD dataset (involving protective gloves), and a robust 92.9 accuracy on the publicly available HAGRID dataset. Crucially, these high‐accuracy results are coupled with an outstandingly low inference latency of merely 8.32 ms. These compelling experimental findings underscore the efficacy and real‐time capability of the LGRDet algorithm. With a compact model size of just 1.23 MB and its inherently streamlined nature, LGRDet demonstrates immense potential for integration into resource‐constrained real‐world environments.
YaPing Wan, Xinyuan Zhou, ZiJun Guo
Concurr. Comput. Pract. Exp.4
2025 MADFlow: Multimodal difference compensation flow for multimodal anomaly detection
Yao Li 0019, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Yixin Qiao
Neurocomputing2
2025 RemoteDPL: A Semi-Supervised Object Detector With Dense Pseudo-Labels for Remote Sensing
abstract
Deep learning-based object detection has seen substantial advancements, however, its practical deployment is often constrained by the need for large-scale labeled datasets. This limitation becomes even more critical in remote sensing imagery, where objects are densely distributed and exhibit significant scale variations. To address these challenges, we introduce RemoteDPL, a novel semi-supervised object detection (SSOD) framework that leverages dense pseudo-labeling (DPL) and multi-scale learning. RemoteDPL offers three key contributions. First, a fusion module is designed to dynamically integrate spatial and channel features across scales, improving detection across varied object sizes. Second, an instance density prediction branch is introduced to support pseudo-label mining, enhancing detection performance in densely populated regions. Lastly, we propose a two-stage pseudo-label filtering strategy that first selects "pending" class predictions and then refines them using a joint confidence score based on both classification and density information. Extensive experiments on the DOTA-v1.0 and NWPU datasets confirm the effectiveness of RemoteDPL, demonstrating its clear advantage over existing state-of-the-art (SOTA) semi-supervised object detection methods. On the NWPU dataset, RemoteDPL outperforms the SOTA baseline by +3.44%, +1.10%, and +1.62% under the settings of data labelled with 30%, 40%, and 50%, respectively, highlighting its strong capability in low-label remote sensing scenarios.
Yongjie Ma, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Zicheng Sun, Yixin Qiao
IEEE Trans. Geosci. Remote. Sens.2
2024 A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker Extraction
abstract
Target speaker extraction (TSE) based on direction of arrival (DOA) has a wide range of applications in e.g., remote conferencing, hearing aids, in-car speech interaction. Due to the inherent phase uncertainty, existing TSE methods usually suffer from speaker confusion within specific frequency bands. Imprecise DOA measurements caused by e.g., the calibration of the microphone array and ambient noises, can also deteriorate the TSE performance. In order to improve the robustness of TSE, in this work we propose several new multichannel spatiotemporal features to represent the discriminability of the target speaker. The narrow-band Conformer model is applied in combination with the proposed features to facilitate the extraction of the target speaker. In addition, we consider knowledge distillation for improving the model robustness, particularly in the presence of DOA mis-match. Experimental results on a public dataset verify the efficacy of the proposed method.
Yichi Wang 0001, Jie Zhang 0042, Shihao Chen, Weitai Zhang, Zhongyi Ye, Xinyuan Zhou, Li-Rong Dai 0001
ICASSP6
2024 Pre-Trained Acoustic-and-Textual Modeling for End-To-End Speech-To-Text Translation
abstract
End-to-end paradigm has aroused more and more interests and attention for improving speech-to-text translation (ST) recently. Existing end-to-end models mainly attributes and attempts to address the problem of modeling burden and data scarcity, while always fail to maintain both cross-modal and cross-lingual mapping well at the same time. In this work, we investigate methods for improving endto-end ST with pre-trained acoustic-and-textual models. Our acoustic encoder and decoder begins with processing the source speech sequence as usual. A textual encoder and an adaptor module then obtain source acoustic and textual information respectively, alleviating the representation inconsistency with attentive interactions in the textual decoder. Also, we utilize pre-trained models, and develop an adaptation fine-tuning method to preserve the pre-training knowledge. Experimental results on the IWSLT2023 offline ST task from English to German, Japanese and Chinese show that our method achieves state-of-the-art BLEU scores and surpasses the strong cascaded ST counterparts in unrestricted setting.
Weitai Zhang, Hanyi Zhang, Chenxuan Liu, Zhongyi Ye, Xinyuan Zhou, Li-Rong Dai 0001
ICASSP5
2023 DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech Translation
abstract
While Diffusion Generative Models have achieved great success on image generation tasks, how to efficiently and effectively incorporate them into speech generation especially translation tasks remains a non-trivial problem.Specifically, due to the low information density of speech data, the transformed discrete speech unit sequence is much longer than the corresponding text transcription, posing significant challenges to existing auto-regressive models.Furthermore, it is not optimal to brutally apply discrete diffusion on the speech unit sequence while disregarding the continuous space structure, which will degrade the generation performance significantly.In this paper, we propose a novel diffusion model by applying the diffusion forward process in the continuous speech representation space, while employing the diffusion backward process in the discrete speech unit space.In this way, we preserve the semantic structure of the continuous speech representation space in the diffusion process and integrate the continuous and discrete diffusion models.We conduct extensive experiments on the textless direct speech-to-speech translation task, where the proposed method achieves comparable results to the computationally intensive auto-regressive baselines (500 steps on average) with significantly fewer decoding steps (50 steps).
Yongxin Zhu 0003, Zhujin Gao, Xinyuan Zhou, Zhongyi Ye, Linli Xu 0002
EMNLP3
2023 Visual-Haptic-Kinesthetic Object Recognition with Multimodal Transformer
Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Hongyu Yang 0002
ICANN (7)1
2023 CGF: A Category Guidance Based PM$_{2.5}$ Sequence Forecasting Training Framework
abstract
PM$_{2.5}$concentration forecasting is important yet challenging. First, complicated local fluctuations in PM$_{2.5}$concentrations disturb modeling global trends. Second, forecasting errors are often accumulated through an autoregressive process. To contend with the two challenges, we propose aCategoryGuidance based PM${_{2.5}}$sequenceForecasting training framework (CGF) to enhance the performance of existing PM${_{2.5}}$concentration forecasting models. CGF contains a Category based Representation Learning (CRL) module and a Category based Self-paced Learning (CSL) module, both of which utilize PM${_{2.5}}$category information that is easily obtained and publicly available. First, CRL employs category information to guide forecasting models to produce more robust hidden representations that are insensitive to local fluctuations, thus alleviating the negative impact of local fluctuations. Second, CSL adaptively selects real PM${_{2.5}}$concentration values versus autoregressive PM${_{2.5}}$forecast values when training forecasting models, helping alleviate error accumulations. The CGF framework is applied to existing PM${_{2.5}}$forecasting models, and the experimental results on two real-world datasets demonstrate that CGF is able to consistently improve the accuracy of existing forecasting models. Furthermore, to validate the generality of CGF, we conduct extensional experiments in two other time-series prediction tasks, including exchange rate forecasting and electricity forecasting. The experimental results also verify the effectiveness of CGF.
Haomin Yu, Jilin Hu, Xinyuan Zhou, Chenjuan Guo, Bin Yang 0002, Qingyong Li
IEEE Trans. Knowl. Data Eng.3
2022 Decoupled Hyperbolic Graph Attention Network for Modeling Substitutable and Complementary Item Relationships
abstract
Modeling substitutable and complementary item relationships is a fundamental and important topic for recommendation in e-commerce online scenarios. In the real world, item relationships are usually coupled, heterogeneous and they also have abundant side information and hierarchical data structures. Recently, to take full advantage of both sides information and topological structure, graph neural networks are widely explored in relationship modeling. However, the existing methods are crude in decoupling heterogeneous relationships. Their model designs lack deep insight of relationships' coupling mode, i.e. neglects the prior knowledge of how relationships affect each other. In addition, many existing graph methods, regardless of how they handle coupled relationships, are deployed in Euclidean spaces, which distorts hierarchical data structure and limits the expressive power due to the non power law characteristic of Euclidean topology. In this paper, we propose a novel Decoupled Hyperbolic Graph Attention Network (DHGAN). The innovations of our DHGAN can be highlighted as two aspects. Firstly, we design metapaths in an adequate way following an algebraic perspective of relationships coupling mode, which helps achieving better model interpretability. Secondly, DHGAN maps heterogeneous relationships into separate hyperbolic spaces, which can better capture the hierarchical information of graph nodes and helps improving model's representational capacity. We conduct extensive experiments on three public real-world datasets, demonstrating DHGAN is superior to the state-of-the-art graph baselines. We release the codes at https://github.com/wt-tju/DHGAN.
Linfang Hou, Xinyuan Zhou, Mian Ma, Zhuoye Ding
CIKM4
2022 LightNet+: A dual-source lightning forecasting network with bi-direction spatiotemporal transformation
Xinyuan Zhou, Haomin Yu, Qingyong Li, Liangtao Xu, Yijun Zhang 0002
Appl. Intell.1
2022 GREAP: a comprehensive enrichment analysis software for human genomic regions
abstract
The rapid development of genomic high-throughput sequencing has identified a large number of DNA regulatory elements with abundant epigenetics markers, which promotes the rapid accumulation of functional genomic region data. The comprehensively understanding and research of human functional genomic regions is still a relatively urgent work at present. However, the existing analysis tools lack extensive annotation and enrichment analytical abilities for these regions. Here, we designed a novel software, Genomic Region sets Enrichment Analysis Platform (GREAP), which provides comprehensive region annotation and enrichment analysis capabilities. Currently, GREAP supports 85 370 genomic region reference sets, which cover 634 681 107 regions across 11 different data types, including super enhancers, transcription factors, accessible chromatins, etc. GREAP provides widespread annotation and enrichment analysis of genomic regions. To reflect the significance of enrichment analysis, we used the hypergeometric test and also provided a Locus Overlap Analysis. In summary, GREAP is a powerful platform that provides many types of genomic region sets for users and supports genomic region annotations and enrichment analyses. In addition, we developed a customizable genome browser containing >400 000 000 customizable tracks for visualization. The platform is freely available at http://www.liclab.net/Greap/view/index.
Yongsan Yang, Fengcui Qian, Xuecang Li, Yanyu Li, Liwei Zhou, Qiuyu Wang, Xinyuan Zhou, Jian Zhang 0084, Zhengmin Yu, Ting Cui, Chenchen Feng, Desi Shang, Mengfei Sun, Yuexin Zhang, Huifang Tang, Chunquan Li 0002
Briefings Bioinform.7
2021 Multi-Channel Target Speech Extraction with Channel Decorrelation and Target Speaker Adaptation
abstract
The end-to-end approaches for single-channel target speech extraction have attracted widespread attention. However, the studies for end-to-end multi-channel target speech extraction are still relatively limited. In this work, we propose two methods for exploiting the multi-channel spatial information to extract the target speech. The first one is using a target speech adaptation layer in a parallel encoder architecture. The second one is designing a channel decorrelation mechanism to extract the inter-channel differential information to enhance the multi-channel encoder representation. We compare the proposed methods with two strong state-of-the-art baselines. Experimental results on the multi-channel reverberant WSJ0 2-mix dataset demonstrate that our proposed methods achieve up to 11.2% and 11.5% relative improvements in SDR and SiSDR respectively, which are the best reported results on this task to the best of our knowledge.
Jiangyu Han, Xinyuan Zhou, Yanhua Long, Yijie Li 0001
ICASSP2
2021 ComPAT: A Comprehensive Pathway Analysis Tools
Xiaojie Su, Chenchen Feng, Ziyu Ning, Qiuyu Wang, Yuexin Zhang, Ling Wei, Xinyuan Zhou, Chunquan Li 0002
ICIC (3)10
2020 Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-Based LVCSR
abstract
Transformer has shown impressive performance in automatic speech recognition.It uses an encoder-decoder structure with self-attention to learn the relationship between high-level representation of source inputs and embedding of target outputs.In this paper, we propose a novel decoder structure that features a self-and-mixed attention decoder (SMAD) with a deep acoustic structure (DAS) to improve the acoustic representation of Transformer-based LVCSR.Specifically, we introduce a self-attention mechanism to learn a multi-layer deep acoustic structure for multiple levels of acoustic abstraction.We also design a mixed attention mechanism that learns the alignment between different levels of acoustic abstraction and its corresponding linguistic information simultaneously in a shared embedding space.The ASR experiments on Aishell-1 show that the proposed structure achieves CERs of 4.8% on the dev set and 5.1% on the test set, which are the best reported results on this task to the best of our knowledge.
Xinyuan Zhou, Grandee Lee, Emre Yilmaz 0001, Yanhua Long, Jiaen Liang, Haizhou Li 0001
INTERSPEECH1
2020 Multi-Encoder-Decoder Transformer for Code-Switching Speech Recognition
abstract
Code-switching (CS) occurs when a speaker alternates words of two or more languages within a single sentence or across sentences.Automatic speech recognition (ASR) of CS speech has to deal with two or more languages at the same time.In this study, we propose a Transformer-based architecture with two symmetric language-specific encoders to capture the individual language attributes, that improve the acoustic representation of each language.These representations are combined using a language-specific multi-head attention mechanism in the decoder module.Each encoder and its corresponding attention module in the decoder are pre-trained using a large monolingual corpus aiming to alleviate the impact of limited CS training data.We call such a network a multi-encoder-decoder (MED) architecture.Experiments on the SEAME corpus show that the proposed MED architecture achieves 10.2% and 10.8% relative error rate reduction on the CS evaluation sets with Mandarin and English as the matrix language respectively.
Xinyuan Zhou, Emre Yilmaz 0001, Yanhua Long, Yijie Li 0001, Haizhou Li 0001
INTERSPEECH1