EDBT 2026 Demo / reviewers in the wild / expert
Mingjiang Wang
dblp:165/7729
· DBLP profile ↗
33ranked-venue papers
1as first author
26since 2021 · last 2026
0000-0002-4706-009XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 20 since 2021Artificial intelligence and machine learning · 12 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MS-TTBA: Mid-side tokenized text-to-binaural audio with large language models
Changjun He, Lianyu Zhou, Shiyun Xu, Weiping Chen, Mingjiang Wang |
Neurocomputing | 6 |
| 2026 | OACodec: Audio attribute disentanglement via orthogonal disentanglement and mutual information minimization
Yukun Qian, Lianyu Zhou, Xuyi Zhuang, Mingjiang Wang |
Neural Networks | 6 |
| 2026 | MS-VBRVQ: Multi-scale variable bitrate speech residual vector quantization
Yukun Qian, Shiyun Xu, Xuyi Zhuang, Mingjiang Wang |
Speech Commun. | 5 |
| 2026 | ADS-BiMamba: Attentive Dynamic-Split Bidirectional Mamba for Multi-Channel Speech Enhancement
Shiyun Xu, Yinghan Cao, Changjun He, Mingjiang Wang |
IEEE Signal Process. Lett. | 5 |
| 2026 | CARVE: Content-Adaptive Rate-Variable Encoding for Neural Speech CodecsabstractNeural speech codecs that convert continuous wave forms into discrete tokens are an essential component for building compact speech representations in downstream applications. However, most existing codecs require both fixed and high frame rates to maintain reconstruction quality, which reduces the efficiency of downstream applications. To address this issue, we propose CARVE, a Content-Adaptive Rate-Variable Encoding strategy that performs compression at an externally specified frame rate within the latent space of a neural speech codec. CARVE comprises three modules: (i) Unsupervised Density Aware Scoring Module that combines inter-frame similarity with redundancy-aware cues to quantify the relative importance of each frame; (ii) Sparse Frame Segmentation Module that allocates finer segmentation to high-density information regions and coarser segmentation to low-density information regions based on these scores; and (iii) Context-Aware Frame Reconstruction Module that maps segment-level context back to frame-level latent representations. Integrating CARVE into the neural speech codec, extensive experiments show that it can perform compression at an externally specified frame rate while maintaining high reconstruction quality, thereby validating its effectiveness and practicality in low-frame-rate neural speech coding scenarios. Audio samples are available at: https://ethuil.github.io/CARVEdemo/ Yukun Qian, Yinghan Cao, Changjun He, Shiyun Xu, Mingjiang Wang |
IEEE Signal Process. Lett. | 6 |
| 2025 | Dual Position Attention Time-Frequency Network for Binaural Audio SynthesisabstractIn applications such as virtual reality and augmented reality, binaural audio provides listeners with a more immersive experience. To synthesize binaural audio with enhanced spatial localization, especially in scenarios involving moving sound sources, accurate phase estimation is crucial. However, existing deep learning methods have yet to achieve satisfactory results in this area. To address this issue, this paper introduces a Dual Position Attention Time-Frequency Network (DPATFNet). Specifically, our approach targets the interaural differences and Doppler effects induced by sound source movement, guiding the synthesis process from monaural to binaural audio through strong positional conditions in the time-frequency domain. The network employs a Dual Position Attention Block (DPAB) to effectively focus on sound source movement and improve phase estimation performance. The proposed DPATFNet demonstrates a strong capability for synthesizing accurate binaural audio, with experimental results on the Binaural Speech dataset showing that DPATFNet achieves state-of-the-art performance in phase metrics (Phase-L2: 0.717, IPD-L2: 1.020, Wave-L2: 0.148, Amplitude-L2: 0.037). Changjun He, Weiping Chen, Mingjiang Wang |
ICASSP | 3 |
| 2025 | CSMT: Combining Snoring and Metadata-based Text for Sleep Apnea Severity ClassificationabstractSleep apnea is a common sleep disorder that, if untreated, can lead to serious health issues. Snoring is a typical symptom of sleep apnea and can be utilized to develop a noncontact automatic detection method for sleep apnea severity classification (SASC). However, due to patient heterogeneity, the acoustic characteristics of snoring vary significantly among individuals. To address this issue, we introduced a text-audio multimodal model that leverages patient’s metadata to provide valuable supplementary information for SASC task. Specifically, we utilized text descriptions derived from metadata and snoring sounds to fine-tune a pretrained text-audio multimodal model. The metadata includes patient’s physical indicators such as gender, age, BMI, neck circumference, and blood pressure. We constructed a snoring dataset that included four sleep apnea severity levels. On this dataset, our method achieved a classification F-score of 74.34%. We conducted a series of ablation experiments to validate the effectiveness of improving SASC performance by leveraging both metadata-based text and snoring sounds. Additionally, we discussed the model’s performance in scenarios where parts of the metadata are unavailable, a situation that may occur in real-world applications. Heng Li 0013, Yukun Qian, Yun Lu 0004, Mingjiang Wang |
ICASSP | 4 |
| 2025 | PriorSinger: Singing Voice Synthesis Model with Prior Condition Cross AttentionabstractThe singing voice synthesis system is designed to generate realistic and expressive singing based on a given musical score. Generative Adversarial Networks (GANs) or diffusion models generate acoustic features, such as Mel-spectrograms, which are subsequently reconstructed into waveforms by a vocoder. In this work, the musical score is encoded as a prior condition to guide the diffusion denoiser through a novel prior cross-attention Transformer during the denoising process. Moreover, we introduce attention mechanisms in both the time and frequency domains within the diffusion denoiser to enhance the resolution of the generated acoustic features. Additionally, incorporating rotary positional encoding allows the model to better handle temporal and frequency positional information. Our model is capable of synthesizing singing with both higher quality and more vivid expressiveness. In subjective evaluations of the Opencpop dataset, our model outperforms state-of-the-art methods. Bosong Yan, Yinghan Cao, Mingjiang Wang |
ICASSP | 4 |
| 2025 | Joint Training Framework for Accent and Speech Recognition Based on Conformer Low-Rank AdaptationabstractIn real-world scenarios, accent variations often reduce Automatic Speech Recognition (ASR) accuracy. Addressing this typically involves a multi-task ASR and Accent Recognition (ASR-AR) framework, but there is limited research on optimizing task-specific feature extraction and enhancing ASR with AR information. This study introduces the Conformer Low-rank Adaptation for Joint Accent and Speech Recognition (CLAnSR), employing LoRA to augment both ASR and AR capabilities using a shared pre-trained base encoder. This approach significantly reduces the model’s parameter and training resource demands while facilitating the extraction of task-specific features. Additionally, we have incorporated accent-aware multi-channel embedding layers, which through spatially independent embeddings, enhance the model’s capacity to accurately represent tokens across diverse dialectical contexts. Tested on the KeSpeech dataset, CLAnSR reaches state-of-the-art AR accuracy 80.41% and competitive ASR CER 8.39%, outperforming non-LLM systems and matching those with LLMs. It reduces parameters by 34.02% and enhances both ASR and AR performance, effectively handling speech dialect variations and advancing the field. Xuyi Zhuang, Yukun Qian, Shiyun Xu, Mingjiang Wang |
ICASSP | 4 |
| 2025 | SF-AN: A lightweight shuffle Fourier attention network for multi-channel speech enhancement
Shiyun Xu, Yinghan Cao, Yukun Qian, Changjun He, Mingjiang Wang |
Speech Commun. | 6 |
| 2025 | Two-stage UNet with channel and temporal-frequency attention for multi-channel speech enhancement
Shiyun Xu, Yinghan Cao, Mingjiang Wang |
Speech Commun. | 4 |
| 2025 | FSTF-AN: Fused Sparse Temporal-Frequency Attentive Network for Multi-Channel Speech EnhancementabstractThe Transformer has achieved impressive performance in the multi-channel speech enhancement field; however, it struggles to capture local features, which leads to the loss of speech details. To enhance the extraction of local features in the network, we propose a fused sparse temporal-frequency attentive network (FSTF-AN), which aims to fully capture features across the temporal-frequency, frequency, and temporal dimensions. We propose top-$k$fused sparse self-attention, which employs a fusion strategy to adaptively retain the most crucial attention scores when computing self-attention maps, thereby eliminating irrelevant information interference and better aggregating features. Furthermore, we propose a multi-scale fused feed-forward network, which effectively captures multi-scale features, further enhancing the network's ability to capture local features. The experimental results demonstrate that FSTF-AN exhibits significant advantages over other SOTA models, effectively enhancing speech quality and intelligibility. Shiyun Xu, Yinghan Cao, Mingjiang Wang |
IEEE Signal Process. Lett. | 5 |
| 2025 | Hypformer: A Fast Hypothesis-Driven Rescoring Speech Recognition FrameworkabstractRecently, the performance of non-autoregressive ASR models has made significant progress but still lags behind hybrid CTC/attention systems. This paper introduces Hypformer, a fast hypothesis-driven rescoring speech recognition framework. Multiple hypothetical prefixes are realized by fast prefix generation algorithm. With two different rescoring methods, nar-ar rescoring and nar$^{2}$rescoring, Hypformer can flexibly switch between autoregressive and non-autoregressive decoding modes to perform rescoring of hypothesis prefixes. Experiments on the standard Mandarin datasets AISHELL-1 and AISHELL-2 demonstrate that Hypformer outperforms the state-of-the-art Hybrid CTC/Attention systems in ASR performance while achieving a speedup of over six times. Experiments on the Mandarin sub-dialect dataset KeSpeech indicate that Hypformer achieves more accurate recognition by leveraging richer contextual information. Xuyi Zhuang, Yukun Qian, Mingjiang Wang |
IEEE Signal Process. Lett. | 3 |
| 2024 | Lightweight Multi-Axial Transformer with Frequency Prompt for Single Channel Speech EnhancementabstractTime-frequency analysis in single-channel speech enhancement has received considerable attention. While Transformer-based architectures are gaining traction, their computational burden can be substantial, especially when dealing with longer speech samples. To address this, our research introduces the lightweight multi-axial Transformer (LMA-Transformer) optimized for low computational overhead while efficiently extracting features along both temporal and frequency axes. Major to our approach is the Temporal/Frequency MultiDConv head self-attention module (T/F-MDHSA), which not only reduces computational costs but also improves the Transformer’s capability to utilize local features effectively. Moreover, we introduce the frequency prompt block, designed to dynamically guide the recovery of frequency features in speech signals that have experienced varying levels of degradation. Compared to state-of-the-art models, our model has competitive performance with 3.40 PESQ, 95.8% STOI, and 10.15 SSNR on the VoiceBank + Demand dataset. Xingwei Liang, Mingjiang Wang, Ruifeng Xu 0001 |
ICASSP | 3 |
| 2024 | Hybrid Attention Time-Frequency Analysis Network for Single-Channel Speech EnhancementabstractThe time-frequency domain remains central to the speech signal analysis. Enhancing the efficacy of neural network-based speech models demands a detailed multi-scale analysis of time-frequency features. This study presents the Hybrid Attention Time-Frequency Analysis Network (HATFANet), an innovative model that uses a dual-branch structure to concurrently estimate the ideal ratio mask and the enhanced complex spectrum. Each branch incorporates Hybrid Attention Blocks (HABs) to capture local, global, and inter-window attention for more effective deep feature extraction by employing reshaping techniques and gated multi-layer perceptrons to focus on different attention scales. The addition of residual channel attention and window multi-head self-attention mechanism accentuate channel attention features and intra-window attention. Our experiments verify the pivotal role of these HABs across varied attentional scales. HATFANet achieves state-of-the-art results on the Voice Bank + DEMAND dataset, recording 3.37 PESQ, 95.8% STOI, and 10.15 SSNR. Xingwei Liang, Ruifeng Xu 0001, Mingjiang Wang |
ICASSP | 4 |
| 2024 | Lightweight Dynamic Sparse Transformer for Monaural Speech Enhancement
Xuyi Zhuang, Yukun Qian, Mingjiang Wang |
INTERSPEECH | 4 |
| 2024 | Multi-modal Feature Fistillation Emotion Recognition Method For Social MediaabstractWith the rise of social media as a primary channel for information exchange, the spread of online public sentiment has become a key factor in escalating social conflicts and triggering public concern, exerting profound influences on social stability and public values. On some instant interaction platforms, social media texts are often short and noisy, rendering traditional sentiment recognition methods unsuitable for scenarios with scarce short-text content and semantic and emotional uncertainty. To address these challenges, this study proposes a TVE-MGF model (Textual-Visual Context Enhancement and Multi-Granularity Semantic Fusion Model), which utilizes ViLBERT for indepth multimodal feature extraction from both visual context and semantic content. In addition, by enhancing the expression of textual and visual semantics, the model achieves multi-granularity fusion of social media text and associated images, enhancing the precision and comprehensive understanding in sentiment analysis. Furthermore, we adopted feature distillation techniques to optimize the TVE-MGF model, aiming to more effectively extract and utilize implicit knowledge from both visual and textual data to construct higher-level semantic feature representations. This step has bolstered the model’s capacity for generalization and extraction of critical knowledge, significantly improving performance when handling complex multimodal sentiment data. Finally, the methodology was experimentally validated on two multimodal datasets, MVSA-single and MVSA-multiple, with the TVE-MGF model achieving F1 scores of $78.10 \%$ and $79.03 \%$ respectively, thereby demonstrating its effectiveness in enhancing the efficiency of sentiment recognition in social media, particularly for texts with high semantic and emotional uncertainty. Mingjiang Wang |
QRS | 2 |
| 2023 | Half-Temporal and Half-Frequency Attention U2Net for Speech Signal ImprovementabstractDuring communication, volume changes, noise, and reverberation can disturb speech signals, significantly affecting the quality and intelligibility of speech. In the context of the ICASSP 2023 Signal Processing Grand Challenge, the first Speech Signal Improvement Grand Challenge (SIG) is organized to improve the quality of speech signals during communication. This paper proposes half-temporal and half-frequency attention U2Net for improving full-band speech signal. Channel-spectrum attention is proposed for the skip connection between the encoder and decoder. The proposed model achieves 0.353, 1.289, 0.604, 0.625, and 0.924 improvements in signal, noise, overall, reverberation, and loudness, respectively, in the SIG subjective test. The proposed model achieved fourth place in the SIG real-time track, showing excellent denoising and de-reverberation performance. Shiyun Xu, Xuyi Zhuang, Yukun Qian, Lianyu Zhou, Mingjiang Wang |
ICASSP | 6 |
| 2023 | Two-Stage UNet with Multi-Axis Gated Multilayer Perceptron for Monaural Noisy-Reverberant Speech EnhancementabstractIn denoising and de-reverberation tasks, the dominant methods are complex spectral masking and complex spectral mapping. To combine advantages and improve speech enhancement performance, we propose a two-stage UNet (TSUNet) to estimate complex spectral masking and complex spectral mapping. We use a multi-axis gated multilayer perceptron to build global and local attention modules of linear complexity for extracting speech features. Furthermore, we use the residual channel attention block to further filter out important speech features. On the blind test dataset of the Deep Noise Suppression Challenge, our proposed TSUNet has a massive advantage over other state-of-the-art models. TSUNet performs significantly better than the most recent models at noisy-reverberant speech enhancement. Shiyun Xu, Xuyi Zhuang, Lianyu Zhou, Heng Li 0013, Mingjiang Wang |
ICASSP | 6 |
| 2023 | Automatic Speech Recognition Transformer with Global Contextual Information Decoder
Yukun Qian, Xuyi Zhuang, Mingjiang Wang |
INTERSPEECH | 3 |
| 2023 | Toward Explainable Dialogue System Using Two-stage Response GenerationabstractIn recent years, neural networks have achieved impressive performance on dialogue response generation. However, most of these models still suffer from some shortcomings, such as yielding uninformative responses and lacking explainable ability. This article proposes a Two-stage Dialogue Response Generation model (TSRG), which specifies a method to generate diverse and informative responses based on an interpretable procedure between stages. TSRG involves a two-stage framework that generates a candidate response first and then instantiates it as the final response. The positional information and a resident token are injected into the candidate response to stabilize the multi-stage framework, alleviating the shortcomings in the multi-stage framework. Additionally, TSRG allows adjusting and interpreting the interaction pattern between the two generation stages, making the generation response somewhat explainable and controllable. We evaluate the proposed model on three dialogue datasets that contain millions of single-turn message-response pairs between web users. The results show that, compared with the previous multi-stage dialogue generation models, TSRG can produce more diverse and informative responses and maintain fluency and relevance. Shaobo Li 0004, Chengjie Sun, Zhen Xu 0003, Prayag Tiwari, Bingquan Liu, Deepak Gupta 0002, K. Shankar 0002, Zhenzhou Ji, Mingjiang Wang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 9 |
| 2022 | Design of Real-Time System Based on Machine Learning for Snoring and OSA DetectionabstractObstructive sleep apnea (OSA) is a common sleep disorder. The diagnosis of OSA based on snoring is low-cost, convenient and non-invasive. In this study, we place a microphone under the patient’s bed and combined with full-night polysomnography to record audio signals. Five machine learning models and two OSA diagnostic schemes are used to classify night audio as non-snoring, snoring, or OSA-related snoring. Our experiment has achieved good results, and the highest diagnosis rate of OSA can reach 97%. Based on the trained classification model, we design a system that can diagnose OSA in real-time. Tests on the system show that it can diagnose apnea by detecting OSA-related snoring. We hope that this approach can develop into a new tool to help a large number of potential OSA patients understand their sleep health. Huaiwen Luo, Lianyu Zhou, Zehuai Zhang, Mingjiang Wang |
ICASSP | 6 |
| 2022 | FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional NetworkabstractIn recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time enhancement of full-band (48 kHz) speech signals is still very challenging. Because of the low energy of spectral information in the high-frequency part, it is more difficult to directly model and enhance the full-band spectrum using neural networks. To solve this problem, this paper proposes a two-stage real-time speech enhancement model with extraction-interpolation mechanism for a full-band signal. The 48 kHz full-band time-domain signal is divided into three sub-channels by extracting, and a two-stage processing scheme of ‘masking + compensation’ is proposed to enhance the signal in the complex domain. After the two-stage enhancement, the enhanced full-band speech signal is restored by interval interpolation. In the subjective listening and word accuracy test, our proposed model achieves superior performance and outperforms the baseline model overall by 0.59 MOS and 4.0% WAcc for the non-personalized speech denoising task. Lu Zhang 0055, Xuyi Zhuang, Yukun Qian, Heng Li 0013, Mingjiang Wang |
ICASSP | 6 |
| 2022 | Coarse-Grained Attention Fusion With Joint Training Framework for Complex Speech Enhancement and End-to-End Speech Recognition
Xuyi Zhuang, Yukun Qian, Mingjiang Wang |
INTERSPEECH | 5 |
| 2022 | High-Speed VLSI Implementation of an Improved Parallel Delayed LMS Algorithm
Mingxiang Guan, Zhou Wu 0002, Chongwu Sun, Mingjiang Wang |
Mob. Networks Appl. | 6 |
| 2021 | PhaseDCN: A Phase-Enhanced Dual-Path Dilated Convolutional Network for Single-Channel Speech EnhancementabstractRecent deep neural network (DNN) based single-channel speech enhancement methods have achieved remarkable results in the time-frequency (TF) magnitude domain. To further improve the quality and intelligibility of enhanced speech, the attention to phase enhancement is also increasing. In this paper, we propose a novel dilated convolutional network (DCN) model to simultaneously enhance the magnitude and phase of noisy speech. Unlike the direct complex spectral mapping methods, we take the complex spectrum of the signal as the main target and the ideal ratio mask (IRM) as the auxiliary target in a multi-target learning framework to achieve their complementary advantages. Firstly, a feature extraction module is introduced to achieve the fusion of local and long-term features. Two different targets are learned separately, but share the common feature extraction module, which is helpful to extract more general and suitable features. During the joint learning, the intermediate estimation of the IRM target in the auxiliary path, contributing as the attention gating factors, helps to distinguish the speech or non-speech components of the complex-valued signals in the main path. To leverage more fine-grained long-term contextual information, we introduce a multi-scale dilated convolution approach for feature encoding. Moreover, the proposed model is a causal system, which can fully meet the low latency requirements of real-time speech products. Experimental results show that, compared with other advanced systems, the proposed model not only has better speech denoising performance and phase estimation accuracy, but also generalizes better in the speaker, noise, and channel mismatch cases. Lu Zhang 0055, Mingjiang Wang, Qiquan Zhang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Rotate3D: Representing Relations as Rotations in Three-Dimensional Space for Knowledge Graph EmbeddingabstractKnowledge graph embedding, which aims to learn low-dimensional embeddings of entities and relations, plays a vital role in a wide range of applications. It is crucial for knowledge graph embedding models to model and infer various relation patterns, such as symmetry/antisymmetry, inversion, and composition. However, most existing methods fail to model the non-commutative composition pattern, which is essential, especially for multi-hop reasoning. To address this issue, we propose a new model called Rotate3D, which maps entities to the three-dimensional space and defines relations as rotations from head entities to tail entities. By using the non-commutative composition property of rotations in the three-dimensional space, Rotate3D can naturally preserve the order of the composition of relations. Experiments show that Rotate3D outperforms existing state-of-the-art models for link prediction and path query answering. Further case studies demonstrate that Rotate3D can effectively capture various relation patterns with a marked improvement in modeling the composition pattern. Chengjie Sun, Lili Shan, Lei Lin 0001, Mingjiang Wang |
CIKM | 5 |
| 2020 | Multi-Scale TCN: Exploring Better Temporal DNN Model for Causal Speech Enhancement
Lu Zhang 0055, Mingjiang Wang |
INTERSPEECH | 2 |
| 2020 | Probability decision-driven speech enhancement algorithm based on human acoustic perceptionabstractIn this study, a novel human acoustic perception motivated Wiener filter speech enhancement system is presented to cope with real‐world interfering background noises. Guiding by the speech presence probability, two alternative methods are proposed to reduce the noise by adopting the audible sound pressure level (SPL) and the masking characteristic of the human auditory system to achieve better listening comfort level. More specifically, when the probability of speech presence in the noisy signal is less than the decision threshold, a new SPL compressed method effectively reduces the noise. When the speech presence probability is more than the decision threshold, an improved acoustical mask threshold constrained Wiener filter approach enhances the noisy speech. Moreover, in order to evaluate the performance of the new system, the proposed algorithm is compared with the classic prior signal‐to‐noise ratio‐based Wiener filter and three acoustic perception related algorithms. The experimental results show that the proposed algorithm significantly outperforms the four comparing algorithms in terms of speech quality and intelligibility either in stationary or moderate non‐stationary noisy environments. Thus, the intended approach can be employed as the front‐end module for various speech‐related applications. Lu Zhang 0055, Mingjiang Wang, Qiquan Zhang |
IET Signal Process. | 2 |
| 2020 | DeepMMSE: A Deep Learning Approach to MMSE-Based Noise Power Spectral Density EstimationabstractAn accurate noise power spectral density (PSD) tracker is an indispensable component of a single-channel speech enhancement system. Bayesian-motivated minimum mean-square error (MMSE)-based noise PSD estimators have been the most prominent in recent time. However, they lack the ability to track highly non-stationary noise sources due to current methods of a priori signal-to-noise (SNR) estimation. This is caused by the underlying assumption that the noise signal changes at a slower rate than the speech signal. As a result, MMSE-based noise PSD trackers exhibit a large tracking delay and produce noise PSD estimates that require bias compensation. Motivated by this, we propose an MMSE-based noise PSD tracker that employs a temporal convolutional network (TCN) a priori SNR estimator. The proposed noise PSD tracker, called DeepMMSE makes no assumptions about the characteristics of the noise or the speech, exhibits no tracking delay, and produces an accurate estimate that requires no bias correction. Our extensive experimental investigation shows that the proposed DeepMMSE method outperforms state-of-the-art noise PSD trackers and demonstrates the ability to track abrupt changes in the noise level. Furthermore, when employed in a speech enhancement framework, the proposed DeepMMSE method is able to outperform state-of-the-art noise PSD trackers, as well as multiple deep learning approaches to speech enhancement. Availability: DeepMMSE is available at: https://github.com/anicolson/DeepXi. Qiquan Zhang, Aaron Nicolson, Mingjiang Wang, Kuldip K. Paliwal, Chenxu Wang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Cancer Detection in Breast Histopathology with Convolution Neural Network Based ApproachabstractBreast cancer is one of most common causes of mortality in women. However, few limitations, e.g., similar structure statistics in inter-class and textural variations in intra-class images make the breast histology analysis a challenging process. In this paper, the multi-class breast cancer classification is carried out with deep convolution neural network (CNN) based transfer learning approach. To explore the feasibility of transfer learning in breast histology, pre-trained deep CNN model is inherited and simultaneously a multi-scale feature concatenation strategy is used. Moreover, incorporating with stain normalization and channel color modification strategies the designed model can be effectively trained. The experiments on publicly available multi-class ICIAR 2018 breast dataset corroborated the efficiency of ou method. The designed approach outperforms the existing methods by achieving 94.3% and 97.5% accuracy on 4-class and 2-class histology image recognition respectively. Tasleem Kausar, Mingjiang Wang, M. S. S. Malik |
AICCSA | 2 |
| 2019 | Dynamic Working Memory for Context-Aware Response GenerationabstractIn human-to-human conversations, the context generally provides several backgrounds and strategic points for the following response. Therefore, many response generation approaches have explored the methodologies to incorporate the context into the encoder-decoder architecture, to generate context-aware responses that are remarkably relevant and cohesive to the given context. However, most approaches pay less attention to semantic interactions implicitly existing within contextual utterances, which are of great importance to capture semantic clues of the given dialog context, indeed. This paper proposes a dynamic working memory mechanism to model long-term semantic hints in the conversation context, by performing semantic interactions between utterances and updating context representation dynamically. Then, the outputs of the dynamic working memory are employed to provide helpful clues for the encoder-decoder architecture to generate responses to the given dialog. We have evaluated the proposed approach on Twitter Customer Service Corpus and OpenSubtitles Corpus, with several automatic evaluation metrics and the human evaluation, and the empirical results show the effectiveness of the proposed method. Zhen Xu 0003, Chengjie Sun, Yinong Long, Bingquan Liu, Baoxun Wang, Mingjiang Wang, Min Zhang 0005, Xiaolong Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2015 | Improved Azimuth Multichannel SAR Imaging for Configurations With Redundant MeasurementsabstractThe traditional spectrum reconstruction algorithm for nonuniform sampling of a displaced phase center antenna (DPCA) synthetic aperture radar (SAR) system can provide an unambiguous recovery of an original signal through the inverse computation of a joint filter matrix. However, when the sampling positions of different receivers coincide spatially, the joint filter matrix will be non-invertible, thus causing the ineffectiveness of the algorithm. This letter proposes a regularized spectrum reconstruction scheme when the pulse repetition frequency of a multichannel system operates on these critical frequencies, as well as in their neighborhoods. First, a low-order regularization of the joint filter matrix is suggested so that the original spectrum of the multichannel system can be recovered more accurately with a decreased number of nonredundant channels' measurements. Furthermore, in order to improve the output signal-to-noise ratio (SNR) of the reconstruction network, a combinational low-order spectrum recovery scheme is additionally presented, which takes all redundant channels' samples into account. The optimizations bring significant benefits to both the suppression of aliased ambiguities and the advance of the output SNR of the reconstruction network. Mingjiang Wang, Robert Wang 0001, Lei Guo 0030, Xiulian Luo |
IEEE Geosci. Remote. Sens. Lett. | 1 |