VLDB 2026 Research / reviewers in the wild / expert
Meng Ge
dblp:23/7081
· DBLP profile ↗
70ranked-venue papers
7as first author
60since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 4 first-author · 42 since 2021Artificial intelligence and machine learning · 37 · 2 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating the Expressive Appropriateness of Speech in Rich ContextsabstractTianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianrui Wang, Ziyang Ma 0001, Yizhou Peng, Zhikang Niu, Zikang Huang, Yi-Wen Chao, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Yifan Yang 0005, Tianchi Liu 0004, Nana Hou, Meng Ge, Fuming You, Zhongqian Sun, Haifeng Hu 0009, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
ACL (1) | 20 |
| 2026 | Sliding-controllable on-off states in MAX3-based (M = Mn, Ni; A = Si, Ge; X = S, Se) van der Waals tunnel junctions
Jianing Tan, Meng Ge, Huamin Hu, Gang Ouyang |
Sci. China Inf. Sci. | 2 |
| 2026 | Improved ship-radiated noise recognition using a pre-trained noise reduction module combined with feature optimization network
Yunye Feng, Haonan Xing, Meng Ge, Tianrui Wang, Lin Gan 0003, Gaoyan Zhang |
Knowl. Based Syst. | 4 |
| 2026 | Unveiling Fine-Grained Deceptive Patterns in Multimodal Fake News: An Explainable Neuro-Symbolic Framework With LVLMsabstractThe widespread proliferation of fake news on the Internet, especially in multi-modal formats, poses a substantial threat to society. Most deep learning-based approaches for fake news detection yield accurate predictions but lack explainability. Existing models focusing on explainability visualize key components from results or generate surface causes via Large Language Models. However, they can hardly provide the deep rationale behind the fabrication of fake news, which is indispensable for misinformation mitigation. Thus, we approach explainability from a different perspective, focusing on explaining how fake news is fabricated, which we term deceptive patterns, at its very source. First, four types of deceptive patterns are pre-established, namely Image Manipulation, Cross-modal Inconsistency, Image Repurposing and Others. Based on this, we propose GE-NSLM, a General Explainable Neuro-Symbolic Latent Model that integrates the power of Large Vision Language Models, which not only provides accurate judgments but also offers insights on deceptive patterns. Specifically, each deceptive pattern is represented as a binary learnable latent variable, obtained through amortized variational inference and weak supervision guided by logical rules. Experiments show GE-NSLM achieves competitive performance. More importantly, it provides interpretable insights into the underlying reasons why specific news items are fake. Dongxiao He, Yiqi Dong, Xiaobao Wang, Meng Ge, Carl Yang 0001, Di Jin 0001, Witold Pedrycz |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Global Entity Relationship Enhancement Network for Multimodal Sarcasm DetectionabstractSarcasm functions as a distinct mode of communication, intended to convey a meaning contrary to its literal interpretation. With the rapid proliferation of social networks, various manifestations of multimodal sarcasm have become widespread. Consequently, there is a growing emphasis on discerning sarcasm conveyed through multimodal data. Existing research often provides only a superficial interpretation of images, lacking a thorough exploration of the contextual nuances embedded within, particularly in understanding the scene depicted in the image. In this paper, we aim to delve deeper into the information embedded within images. Specifically, we begin by extracting entity relationships from images to capture the contextual information they convey. Additionally, we utilize image text recognition to extract textual information from the images. After conducting a comprehensive analysis of the image information, we establish consistency modeling between the image content and text using external knowledge. Finally, we employ a graph neural network to process the constructed cross-modal graph and make predictions regarding sarcasm. Extensive experiments validate the state-of-the-art performance of our model on publicly available multimodal Twitter datasets. Xiaobao Wang, Meng Ge, Lingshan Li, Di Jin 0001, Kai He 0001, Erik Cambria |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging ModuleabstractThe information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an effective post-processing method to improve ASR performance by balancing noisy and enhanced speech. Determining the OA coefficient is crucial. However, the currently supervised OA coefficient module, called the bridging module, only utilizes simulated noisy speech for training, which has a severe mismatch with real noisy speech. In this paper, we propose training strategies to train the bridging module with real noisy speech. First, DNSMOS is selected to evaluate the perceptual quality of real noisy speech with no need for the corresponding clean label to train the bridging module. Additional constraints during training are introduced to enhance the robustness of the bridging module further. Each utterance is evaluated by the ASR back-end using various OA coefficients to obtain the word error rates (WERs). The WERs are used to construct a multidimensional vector. This vector is introduced into the bridging module with multi-task learning and is used to determine the optimal OA coefficients. The experimental results on the CHiME-4 dataset show that the proposed methods all had significant improvement compared with the simulated data trained bridging module, especially under real evaluation sets. Zhongjian Cui, Chenrui Cui, Tianrui Wang, Mengnan He, Meng Ge, Caixia Gong, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 6 |
| 2025 | Augmenting Short Enrollment Speech via Synthesis for Target Speaker ExtractionabstractA high-quality enrollment speech is crucial to target speaker extraction (TSE), since it provides essential cues for identifying the target speaker in the mixture. However, real applications usually only permit a short enrollment speech, e.g. a wakeup word for a mobile device, that provides limited cues. To address this issue, we propose an enrollment augmentation strategy that allows us to enrich the limited enrollment speech with massive text data through speech synthesis. By doing so, the extended enrollment speech contains enhanced speaker timbre and phonetic content which leads to better extraction quality. Furthermore, we propose a training data augmentation strategy to improve the model’s robustness and generalization in short enrollment speech scenarios. Experiments on Libri2Mix demonstrate that our proposed strategies bring a significant improvement in extreme scenarios where only 0.5s and 1-word enrollment speech is provided. We also release our code at https://github.com/HuangZikang-TJU/Aug4TSE. Zikang Huang, Jingru Lin, Meng Ge, Xiaobao Wang, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 3 |
| 2025 | A Chinese Expressive Long-dialogue Speech Dataset with ScriptsabstractWith the advancement of large-scale models, the demand for emotionally rich, long-context, and highly natural communication in human-computer interaction increases. However, the exploration of long-context or script-level speech conversation tasks remains limited due to the lack of specific supervised data. To address this, we introduce a three-stage data processing pipeline for creating a Chinese expressive long-dialogue speech dataset with scripts (CELSDS). We collect videos from TV series, manually annotate speaker information for each character, apply Optical Character Recognition (OCR) to extract speech content, annotate episode summaries, and use a large language model (LLM) to generate sentence-level scenario descriptions. To our knowledge, this is the first Chinese long-context dialogue dataset that incorporates speaker and content annotations, script-level episode summaries, and sentence-level scenario details. Using this dataset, we develop a baseline model for both speech-to-script and script-to-speech generation tasks. The annotations and data production code are open-sourced at: https://github.com/lijin0120/CELSDS. Tianrui Wang, Meng Ge, Chenrui Cui, Jianrong Wang, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 3 |
| 2025 | Mamba-SEUNet: Mamba UNet for Monaural Speech EnhancementabstractIn recent speech enhancement (SE) research, transformer and its variants have emerged as the predominant methodologies. However, the quadratic complexity of the self-attention mechanism imposes certain limitations on practical deployment. Mamba, as a novel state-space model (SSM), has gained widespread application in natural language processing and computer vision due to its strong capabilities in modeling long sequences and relatively low computational complexity. In this work, we introduce Mamba-SEUNet, an innovative architecture that integrates Mamba with U-Net for SE tasks. By leveraging bidirectional Mamba to model forward and backward dependencies of speech signals at different resolutions, and incorporating skip connections to capture multi-scale information, our approach achieves state-of-the-art (SOTA) performance. Experimental results on the VCTK+DEMAND dataset indicate that Mamba-SEUNet attains a PESQ score of 3.59, while maintaining low computational complexity. When combined with the Perceptual Contrast Stretching technique, Mamba-SEUNet further improves the PESQ score to 3.73. Zizhen Lin, Tianrui Wang, Meng Ge, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 4 |
| 2025 | Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech EnhancementabstractTime-frequency (T-F) domain methods for monaural speech enhancement have benefited from the success of deep learning. Recently, focus has been put on designing two-stream network models to predict amplitude mask and phase separately, or, coupling the amplitude and phase into Cartesian coordinates and constructing real and imaginary pairs. However, most methods suffer from the alignment modeling of amplitude and phase (real and imaginary pairs) in a two-stream network framework, which inevitably incurs performance restrictions. In this paper, we introduce a graph Fourier transform defined with the singular value decomposition (GFT-SVD), resulting in real-valued time-graph representation for neural speech enhancement. This real-valued representation-based GFT-SVD provides an ability to align the modeling of amplitude and phase, leading to avoiding recovering the target speech phase information. Our findings demonstrate the effects of real-valued time-graph representation based on GFT-SVD for neutral speech enhancement. The extensive speech enhancement experiments establish that the combination of GFT-SVD and DNN outperforms the combination of GFT with the eigenvector decomposition (GFT-EVD) and magnitude estimation UNet, and outperforms the short-time Fourier transform (STFT) and DNN, regarding objective intelligibility and perceptual quality. We release our source code at: https://github.com/Wangfighting0015/GFTproject. Tianrui Wang, Meng Ge, Qiquan Zhang, Zirui Ge, Zhen Yang 0001 |
ICASSP | 3 |
| 2025 | Multi-range Random Walk Based Graph Neural Network
Meng Ge, Xiaobao Wang, Di Jin 0001 |
ICIC (10) | 2 |
| 2025 | A Progressive Generation Framework with Speech Pre-trained Model for Expressive Voice ConversionabstractExpressive voice conversion (EVC) aims to modify the speaker identity and emotional style of speech while preserving its content. Existing approaches often focus on disentangling speaker, emotion, and content information but overlook the progressive generation mechanisms in human speech production. To address this, we propose a three-stage framework that includes a speech disentanglement module, a progressive generator, and an acoustic refiner. This framework enables speech pre-trained models to parse linguistic content, emotional style, and speaker identity, which are then progressively integrated into the speech reconstruction branch to generate high-quality speech with replaceable emotional style and speaker identity. Experiments with six different pre-trained models show that our framework activates their disentanglement capabilities, surpassing baseline performance in EVC, and supports speaker and emotion control from different target samples. This framework also provides a valuable reference for evaluating the disentanglement capabilities of speech pre-training models. Tianrui Wang, Meng Ge, Zhikang Niu, Chunyu Qiang, Zikang Huang, Ziyang Ma 0001, Xiaobao Wang, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
ICME | 2 |
| 2025 | A Three-Stage Beamforming with Harmonic Guidance for Multi-Channel Speech Enhancement
Nurali Alip, Tianrui Wang, Meng Ge, Jingru Lin, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 4 |
| 2025 | ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
Tianrui Wang, Meng Ge, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2025 | Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech SynthesisabstractWhile emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modeling multi-emotion transitions and the scarcity of annotated datasets that capture intra-sentence emotional and prosodic variation. In this paper, we propose WeSCon, the first self-training framework that enables word-level control of both emotion and speaking rate in a pretrained zero-shot TTS model, without relying on datasets containing intra-sentence emotion or speed transitions.
Our method introduces a transition-smoothing strategy and a dynamic speed control mechanism to guide the pretrained TTS model in performing word-level expressive synthesis through a multi-round inference process. To further simplify the inference, we incorporate a dynamic emotional attention bias mechanism and fine-tune the model via self-training, thereby activating its ability for word-level expressive control in an end-to-end manner. Experimental results show that WeSCon effectively overcomes data scarcity, achieving state-of-the-art performance in word-level emotional expression control while preserving the strong zero-shot synthesis capabilities of the original TTS model. Tianrui Wang, Meng Ge, Chunyu Qiang, Ziyang Ma 0001, Zikang Huang, Guanrou Yang, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
NeurIPS | 3 |
| 2025 | Dipole-moment modulation of photovoltaic and photoresponse properties in 2H-MoSSe/MoS2 van der Waals heterojunctions under electric field
Yinkang Wang, Meng Ge, Jianing Tan, Gang Ouyang |
Sci. China Inf. Sci. | 2 |
| 2025 | HC-APNet: Harmonic Compensation Auditory Perception Network for low-complexity speech enhancement
Meng Ge, Longbiao Wang, Yang-Hao Zhou, Jianwu Dang 0001 |
Speech Commun. | 2 |
| 2025 | LORT: Locally refined convolution and Taylor transformer for monaural speech enhancement
Zizhen Lin, Tianrui Wang, Meng Ge, Longbiao Wang, Jianwu Dang 0001 |
Speech Commun. | 4 |
| 2025 | Enhancing Semi-Supervised Instance Segmentation Through SAM-Driven Pseudo-Label Generation in Autonomous Driving EnvironmentabstractIn recent years, relying on large amounts of labeled data in specific scenarios, instance segmentation has made great strides in autonomous driving. However, labeling sufficient images for each target domain scene is time-consuming and impractical. Based on this, for effectively utilizing a large amount of unlabeled data in the target domain to improve the model performance, this paper proposes a Segment Anything Model (SAM)-driven Pseudo-Labels gEneration (SAMPLE) framework for enhancing semi-supervised instance segmentation (SSIS) in autonomous driving environments. This framework is the first to apply the visual foundation model SAM to an SSIS framework, using precise prompts and the segmentation capabilities of SAM to generate pseudo-labels. Additionally, to mitigate pseudo-label noise, a Multi-Point Enhancement (MPE) prompting method is proposed, which enables SAM to generate high-quality pseudo-labels and thus improve the performance of the segmentation model. Experiments were conducted on the public datasets Cityscapes and COCO. The results show that using only 5% of the labeled data from Cityscapes, the proposed method achieves a 1.8% improvement over the current state-of-the-art GD method. On COCO, SAMPLE consistently surpasses other methods across various data percentages. Furthermore, compared to the Mixture-of-Experts (MoE) method, our method boasts higher accuracy, reduces model size by 8@, decreases GPU memory usage to 1/3 of the original, and increases inference speed by 110@. These improvements lay the foundation for effectively deploying segmentation models in the autonomous driving environment. Licong Guan, Meng Ge |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Unveiling Implicit Deceptive Patterns in Multi-Modal Fake News via Neuro-Symbolic ReasoningabstractIn the current Internet landscape, the rampant spread of fake news, particularly in the form of multi-modal content, poses a great social threat. While automatic multi-modal fake news detection methods have shown promising results, the lack of explainability remains a significant challenge. Existing approaches provide superficial explainability by displaying learned important components or views from well-trained networks, but they often fail to uncover the implicit deceptive patterns that reveal how fake news is fabricated. To address this limitation, we begin by predefining three typical deceptive patterns, namely image manipulation, cross-modal inconsistency, and image repurposing, which shed light on the mechanisms underlying fake news fabrication. Then, we propose a novel Neuro-Symbolic Latent Model called NSLM, that not only derives accurate judgments on the veracity of news but also uncovers the implicit deceptive patterns as explanations. Specifically, the existence of each deceptive pattern is expressed as a two-valued learnable latent variable, which is acquired through amortized variational inference and weak supervision based on symbolic logic rules. Additionally, we devise pseudo-siamese networks to capture distinct deceptive patterns effectively. Experimental results on two real-world datasets demonstrate that our NSLM achieves the best performance in fake news detection while providing insightful explanations of deceptive patterns. Yiqi Dong, Dongxiao He, Xiaobao Wang, Youzhu Jin, Meng Ge, Carl Yang 0001, Di Jin 0001 |
AAAI | 5 |
| 2024 | Improving Distinguishability of Class for Graph Neural NetworksabstractGraph Neural Networks (GNNs) have received widespread attention and applications due to their excellent performance in graph representation learning. Most existing GNNs can only aggregate 1-hop neighbors in a GNN layer, so they usually stack multiple GNN layers to obtain more information from larger neighborhoods. However, many studies have shown that model performance experiences a significant degradation with the increase of GNN layers. In this paper, we first introduce the concept of distinguishability of class to indirectly evaluate the learned node representations, and verify the positive correlation between distinguishability of class and model performance. Then, we propose a Graph Neural Network guided by Distinguishability of class (Disc-GNN) to monitor the representation learning, so as to learn better node representations and improve model performance. Specifically, we first perform inter-layer filtering and initial compensation based on Local Distinguishability of Class (LDC) in each layer, so that the learned node representations have the ability to distinguish different classes. Furthermore, we add a regularization term based on Global Distinguishability of Class (GDC) to achieve global optimization of model performance. Extensive experiments on six real-world datasets have shown that the competitive performance of Disc-GNN to the state-of-the-art methods on node classification and node clustering tasks. Dongxiao He, Shuwei Liu, Meng Ge, Zhizhi Yu, Guangquan Xu, Zhiyong Feng 0002 |
AAAI | 3 |
| 2024 | Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-Talker SpeechabstractTarget speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the interfering speech. However, this scenario only accounts for a small percentage of real-world conversations. In this paper, we aim at the sparsely overlapped scenarios in which the auxiliary reference needs to perform two tasks simultaneously: detect the activity of the target speaker and disentangle the active speech from any interfering speech. We propose an audio-visual speaker extraction model named ActiveExtract, which leverages speaking activity from audio-visual active speaker detection (ASD). The ASD directly provides the frame-level activity of the target speaker, while its intermediate feature representation is trained to discriminate speech-lip synchronization that could be used for speaker disentanglement. Experimental results show our model outperforms baselines across various overlapping ratios, achieving an average improvement of more than 4 dB in terms of SI-SNR. Ruijie Tao, Zexu Pan, Meng Ge, Shuai Wang 0016, Haizhou Li 0001 |
ICASSP | 4 |
| 2024 | Gradient Weighting for Speaker Verification in Extremely Low Signal-to-Noise RatioabstractSpeaker verification is hampered by background noise, particularly at extremely low Signal-to-Noise Ratio (SNR) under 0 dB. It is difficult to suppress noise without introducing unwanted artifacts, which adversely affects speaker verification. We proposed the mechanism called Gradient Weighting (Grad-W), which dynamically identifies and reduces artifact noise during prediction. The mechanism is based on the property that the gradient indicates which parts of the input the model is paying attention to. Specifically, when the speaker network focuses on a region in the denoised utterance but not on the clean counterpart, we consider it artifact noise and assign higher weights for this region during optimization of enhancement. We validate it by training an enhancement model and testing the enhanced utterance on speaker verification. The experimental results show that our approach effectively reduces artifact noise, improving speaker verification across various SNR levels. Kong-Aik Lee, Ville Hautamäki, Meng Ge, Haizhou Li 0001 |
ICASSP | 4 |
| 2024 | SVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural NetworksabstractSpeech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neural Networks (SNNs) are known to be biologically plausible and power-efficient. However, SNN-based VADs have yet to achieve noise robustness and often require large models for high performance. This paper introduces a novel SNN-based VAD model, referred to as sVAD, which features an auditory encoder with an SNN-based attention mechanism. Particularly, it provides effective auditory feature representation through SincNet and 1D convolution, and improves noise robustness with attention mechanisms. The classifier utilizes Spiking Recurrent Neural Networks (sRNN) to exploit temporal speech information. Experimental results demonstrate that our sVAD achieves remarkable noise robustness and meanwhile maintains low power consumption and a small footprint, making it a promising solution for real-world VAD applications. Qu Yang, Qianhui Liu, Meng Ge, Zeyang Song, Haizhou Li 0001 |
ICASSP | 4 |
| 2024 | An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech EnhancementabstractTransformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encoding exactly impacts speech enhancement based on Transformer architectures. In this paper, we perform a comprehensive empirical study evaluating five positional encoding methods, i.e., Sinusoidal and learned absolute position embedding (APE), T5-RPE, KERPLE, as well as the Transformer without positional encoding (No-Pos), across both causal and noncausal configurations. We conduct extensive speech enhancement experiments, involving spectral mapping and masking methods. Our findings establish that positional encoding is not quite helpful for the models in a causal configuration, which indicates that causal attention may implicitly incorporate position information. In a noncausal configuration, the models significantly benefit from the use of positional encoding. In addition, we find that among the four position embeddings, relative position embeddings outperform APEs. Qiquan Zhang, Meng Ge, Hongxu Zhu, Eliathamby Ambikairajah, Zhaoheng Ni, Haizhou Li 0001 |
ICASSP | 2 |
| 2024 | Multi-Modal Sarcasm Detection Based on Dual Generative Processes
Huiying Ma, Dongxiao He, Xiaobao Wang, Di Jin 0001, Meng Ge, Longbiao Wang |
IJCAI | 5 |
| 2024 | VoiCor: A Residual Iterative Voice Correction Framework for Monaural Speech Enhancement
Tianrui Wang, Meng Ge, Andong Li, Longbiao Wang, Jianwu Dang 0001, Yungang Jia |
INTERSPEECH | 3 |
| 2024 | SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture SpeechabstractIt was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks.However, most such models are trained on singlespeaker speech data, limiting their effectiveness in mixture speech.This motivates us to explore pre-training on mixture speech.This work presents SA-WavLM, a novel pre-trained model for mixture speech.Specifically, SA-WavLM follows an "extract-merge-predict" pipeline in which the representations of each speaker in the input mixture are first extracted individually and then merged before the final prediction.In this pipeline, SA-WavLM performs speaker-informed extractions with the consideration of the interactions between different speakers.Furthermore, a speaker shuffling strategy is proposed to enhance the robustness towards the speaker absence.Experiments show that SA-WavLM either matches or improves upon the state-of-the-art pre-trained models. Jingru Lin, Meng Ge, Junyi Ao, Liqun Deng, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2024 | WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
Shuai Wang 0016, Shaoxiong Lin, Meng Ge, Jianwei Yu 0001, Yanmin Qian, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2024 | TUT4CRS: Time-aware User-preference Tracking for Conversational Recommendation SystemabstractThe Conversational Recommendation System (CRS) aims to capture user dynamic preferences and provide item recommendations based on multi-turn conversations. However, effectively modeling these dynamic preferences faces challenges due to conversational limitations, which mainly manifests as limited turns in a conversation (quantity aspect) and low compliance with queries (quality aspect). Previous studies often address these challenges in isolation, overlooking their interconnected nature. The fundamental issue underlying both problems lies in the potential abrupt changes in user preferences, to which CRS may not respond promptly. We acknowledge that user preferences are influenced by temporal factors, serving as a bridge between conversation quantity and quality. Therefore, we propose a more comprehensive CRS framework called Time-aware User-preference Tracking for Conversational Recommendation System (TUT4CRS), leveraging time dynamics to tackle both issues simultaneously. Specifically, we construct a global time interaction graph to incorporate rich external information and establish a local time-aware weight graph based on this information to adeptly select queries and effectively model user dynamic preferences. Extensive experiments on two real-world datasets validate that TUT4CRS can significantly improve recommendation performance while reducing the number of conversation turns. Dongxiao He, Jinghan Zhang 0010, Xiaobao Wang, Meng Ge, Zhiyong Feng 0002, Longbiao Wang, Xiaoke Ma 0001 |
ACM Multimedia | 4 |
| 2024 | FUG: Feature-Universal Graph Contrastive Pre-training for Graphs with Diverse Node FeaturesabstractGraph Neural Networks (GNNs), known for their effective graph encoding, are extensively used across various fields. Graph self-supervised pre-training, which trains GNN encoders without manual labels to generate high-quality graph representations, has garnered widespread attention. However, due to the inherent complex characteristics in graphs, GNNs encoders pre-trained on one dataset struggle to directly adapt to others that have different node feature shapes. This typically necessitates either model rebuilding or data alignment. The former results in non-transferability as each dataset need to rebuild a new model, while the latter brings serious knowledge loss since it forces features into a uniform shape by preprocessing such as Principal Component Analysis (PCA). To address this challenge, we propose a new Feature-Universal Graph contrastive pre-training strategy (FUG) that naturally avoids the need for model rebuilding and data reshaping. Specifically, inspired by discussions in existing work on the relationship between contrastive Learning and PCA, we conducted a theoretical analysis and discovered that PCA's optimization objective is a special case of that in contrastive Learning. We designed an encoder with contrastive constraints to emulate PCA's generation of basis transformation matrix, which is utilized to losslessly adapt features in different datasets. Furthermore, we introduced a global uniformity constraint to replace negative sampling, reducing the time complexity from $O(n^2)$ to $O(n)$, and by explicitly defining positive samples, FUG avoids the substantial memory requirements of data augmentation. In cross domain experiments, FUG has a performance close to the re-trained new models. The source code is available at: https://github.com/hedongxiao-tju/FUG. Jitao Zhao, Di Jin 0001, Meng Ge, Lianze Shan, Xin Wang 0030, Dongxiao He, Zhiyong Feng 0002 |
NeurIPS | 3 |
| 2024 | Intelligent event-based lip reading word classification with spiking neural networks using spatio-temporal attention features and triplet loss
Qianhui Liu, Meng Ge, Haizhou Li 0001 |
Inf. Sci. | 2 |
| 2024 | Robust voice activity detection using an auditory-inspired masked modulation encoder based convolutional attention network
Longbiao Wang, Meng Ge, Masashi Unoki, Sheng Li 0010, Jianwu Dang 0001 |
Speech Commun. | 3 |
| 2024 | Selective HuBERT: Self-Supervised Pre-Training for Target Speaker in Clean and Mixture SpeechabstractSelf-supervised pre-trained speech models were shown effective for various downstream speech processing tasks. Since they are mainly pre-trained to map input speech to pseudo-labels, the resulting representations are only effective for the type of pre-train data used, either clean or mixture speech. With the idea of selective auditory attention, we propose a novel pre-training solution called Selective-HuBERT, or SHuBERT, which learns the selective extraction of target speech representations from either clean or mixture speech. Specifically, SHuBERT is trained to predict pseudo labels of a target speaker, conditioned on an enrolment speech from the target speaker. By doing so, SHuBERT is expected to selectively attend to the target speaker in a complex acoustic environment, thus benefiting various downstream tasks. We further introduce a dual-path training strategy and use the cross-correlation constraint between the two branches to encourage the model to generate noise-invariant representation. Experiments on SUPERB benchmark and LibriMix dataset demonstrate the universality and noise-robustness of SHuBERT. Furthermore, we find that our high-quality representation can be easily integrated with conventional supervised learning methods to achieve significant performance, even under extremely low-resource labeled data. Jingru Lin, Meng Ge, Wupeng Wang, Haizhou Li 0001, Mengling Feng |
IEEE Signal Process. Lett. | 2 |
| 2023 | Stream Attention Based U-Net for L3DAS23 ChallengeabstractMachine learning applications of 3D audio are gaining increasing interest in recent years. In this paper, we propose a stream attention based U-Net to remove background noise and reverberation based on ICASSP Signal Processing Grand Challenge 2023: L3DAS23 Challenge1Audio-only track task1 3D Speech Enhancement. Results show that proposed method achieves superior performance than the official baseline model. Honglong Wang, Yanjie Fu, Meng Ge, Longbiao Wang, Xinyuan Qian 0001 |
ICASSP | 4 |
| 2023 | Locate and Beamform: Two-dimensional Locating All-neural Beamformer for Multi-channel Speech Separation
Yanjie Fu, Meng Ge, Honglong Wang, Longbiao Wang, Gaoyan Zhang, Jianwu Dang 0001, Chengyun Deng |
INTERSPEECH | 2 |
| 2023 | Rethinking the Visual Cues in Audio-Visual Speaker Extraction
Meng Ge, Zexu Pan, Longbiao Wang, Jianwu Dang 0001, Shiliang Zhang |
INTERSPEECH | 2 |
| 2023 | PIAVE: A Pose-Invariant Audio-Visual Speaker Extraction NetworkabstractIt is common in everyday spoken communication that we look at the turning head of a talker to listen to his/her voice. Humans see the talker to listen better, so do machines. However, previous studies on audio-visual speaker extraction have not effectively handled the varying talking face. This paper studies how to take full advantage of the varying talking face. We propose a Pose-Invariant Audio-Visual Speaker Extraction Network (PIAVE) that incorporates an additional pose-invariant view to improve audio-visual speaker extraction. Specifically, we generate the pose-invariant view from each original pose orientation, which enables the model to receive a consistent frontal view of the talker regardless of his/her head pose, therefore, forming a multi-view visual input for the speaker. Experiments on the multi-view MEAD and in-the-wild LRS3 dataset demonstrate that PIAVE outperforms the state-of-the-art and is more robust to pose variations. Meng Ge, Zhizheng Wu 0001, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2023 | SDNet: Stream-attention and Dual-feature Learning Network for Ad-hoc Array Speech Separation
Honglong Wang, Chengyun Deng, Yanjie Fu, Meng Ge, Longbiao Wang, Gaoyan Zhang, Jianwu Dang 0001 |
INTERSPEECH | 4 |
| 2023 | Time-Domain Speech Separation Networks With Graph Encoding AuxiliaryabstractEnd-to-end time-domain speech separation with masking strategy has shown its performance advantage, where a 1-D convolutional layer is used as the speech encoder to encode a sliding window of waveform to a latent feature representation, i.e. an embedding vector. A large window leads to low resolution in the speech processing, on the other hand, a small window offers high resolution but at the expense of high computational cost. In this work, we propose a graph encoding technique to model the fine structural knowledge of speech samples in a window of reasonable size. Specifically, we build a graph representation for each latent representation, and encode the structural details with a graph convolutional network encoder. The encoded graph feature representation complements the original latent feature representation and benefits the separation and reconstruction of speech. Experiments on various models and datasets show that our proposed encoding technique significantly improves the speech quality over other time-domain speech encoders. Zexu Pan, Meng Ge, Zhen Yang 0001, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | L-SpEx: Localized Target Speaker ExtractionabstractSpeaker extraction aims to extract the target speaker’s voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction of the target speaker. However, these studies assume that the target speaker’s location is known in advance or detected by an extra visual cue, e.g., face image or video. In this paper, we propose an end-to-end localized target speaker extraction on pure speech cues, that is called L-SpEx. Specifically, we design a speaker localizer driven by the target speaker’s embedding to extract the spatial features, including direction-of-arrival (DOA) of the target speaker and beamforming output. Then, the spatial cues and target speaker’s embedding are both used to form a top-down auditory attention to the target speaker. Experiments on the multi-channel reverberant dataset called MCLibri2Mix show that our L-SpEx approach significantly outperforms the baseline system. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
ICASSP | 1 |
| 2022 | Compressing Transformer-Based ASR Model by Task-Driven Loss and Attention-Based Multi-Level Feature DistillationabstractThe current popular knowledge distillation (KD) methods effectively compress the transformer-based end-to-end speech recognition model. However, existing methods fail to utilize complete information of the teacher model, and they distill only a limited number of blocks of the teacher model. In this study, we first integrate a task-driven loss function into the decoder’s intermediate blocks to generate task-related feature representations. Then, we propose an attention-based multi-level feature distillation to automatically learn the feature representation summarized by all blocks of the teacher model. Under the 1.1M parameters model, the experimental results on the Wall Street Journal dataset reveal that our approach achieves a 12.1% WER reduction compared with the baseline system. Yongjie Lv, Longbiao Wang, Meng Ge, Sheng Li 0010, Chenchen Ding, Lixin Pan, Yuguang Wang 0003, Jianwu Dang 0001, Kiyoshi Honda |
ICASSP | 3 |
| 2022 | RAW-GNN: RAndom Walk Aggregation based Graph Neural NetworkabstractGraph-Convolution-based methods have been successfully applied to representation learning on homophily graphs where nodes with the same label or similar attributes tend to connect with one another. Due to the homophily assumption of Graph Convolutional Networks (GCNs) that these methods use, they are not suitable for heterophily graphs where nodes with different labels or dissimilar attributes tend to be adjacent. Several methods have attempted to address this heterophily problem, but they do not change the fundamental aggregation mechanism of GCNs because they rely on summation operators to aggregate information from neighboring nodes, which is implicitly subject to the homophily assumption. Here, we introduce a novel aggregation mechanism and develop a RAndom Walk Aggregation-based Graph Neural Network (called RAW-GNN) method. The proposed approach integrates the random walk strategy with graph neural networks. The new method utilizes breadth-first random walk search to capture homophily information and depth-first search to collect heterophily information. It replaces the conventional neighborhoods with path-based neighborhoods and introduces a new path-based aggregator based on Recurrent Neural Networks. These designs make RAW-GNN suitable for both homophily and heterophily graphs. Extensive experimental results showed that the new method achieved state-of-the-art performance on a variety of homophily and heterophily graphs. Di Jin 0001, Rui Wang 0102, Meng Ge, Dongxiao He, Xiang Li 0067, Wei Lin 0022, Weixiong Zhang |
IJCAI | 3 |
| 2022 | Dual-stream Speech Dereverberation Network Using Long-term and Short-term CuesabstractFor reverberation, the current speech is usually influenced by the previous frames. Traditional neural network-based speech dereverberation (SD) methods directly map the current speech frame that only has short-term cues to clean speech or learn a mask, which can not utilize long-term information to remove late reverberation and further limit SD's ability. To address this issue, we propose a dual-stream speech dereverberation network (DualSDNet) using long-term and short-term cues. First, we analyze the effectiveness of using a finite impulse response (FIR) based on long-term information recorded filter by reverberation generation progress. Second, to make full use of both long-term and short-term information, we further design a dual-stream network, it can map both long and short speech to high-dimensional representation and pay more attention to a more helpful time index. The results of the REVERB Challenge data show that our DualSDNet consistently outperforms the state-of-the-art SD baselines. Meng Ge, Longbiao Wang, Jianwu Dang 0001 |
IJCNN | 2 |
| 2022 | Iterative Sound Source Localization for Unknown Number of SourcesabstractSound source localization aims to seek the direction of arrival (DOA) of all sound sources from the observed multichannel audio.For the practical problem of unknown number of sources, existing localization algorithms attempt to predict a likelihood-based coding (i.e., spatial spectrum) and employ a pre-determined threshold to detect the source number and corresponding DOA value.However, these threshold-based algorithms are not stable since they are limited by the careful choice of threshold.To address this problem, we propose an iterative sound source localization approach called ISSL, which can iteratively extract each source's DOA without threshold until the termination criterion is met.Unlike threshold-based algorithms, ISSL designs an active source detector network based on binary classifier to accept residual spatial spectrum and decide whether to stop the iteration.By doing so, our ISSL can deal with an arbitrary number of sources, even more than the number of sources seen during the training stage.The experimental results show that our ISSL achieves significant performance improvements in both DOA estimation and source number detection compared with the existing threshold-based algorithms. Yanjie Fu, Meng Ge, Xinyuan Qian 0001, Longbiao Wang, Gaoyan Zhang, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2022 | VCSE: Time-Domain Visual-Contextual Speaker Extraction NetworkabstractSpeaker extraction seeks to extract the target speech in a multitalker scenario given an auxiliary reference.Such reference can be auditory, i.e., a pre-recorded speech, visual, i.e., lip movements, or contextual, i.e., phonetic sequence.References in different modalities provide distinct and complementary information that could be fused to form top-down attention on the target speaker.Previous studies have introduced visual and contextual modalities in a single model.In this paper, we propose a two-stage time-domain visual-contextual speaker extraction network named VCSE, which incorporates visual and selfenrolled contextual cues stage by stage to take full advantage of every modality.In the first stage, we pre-extract a target speech with visual cues and estimate the underlying phonetic sequence.In the second stage, we refine the pre-extracted target speech with the self-enrolled contextual cues.Experimental results on the real-world Lip Reading Sentences 3 (LRS3) database demonstrate that our proposed VCSE network consistently outperforms other state-of-the-art baselines. Meng Ge, Zexu Pan, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2022 | Global Signal-to-noise Ratio Estimation Based on Multi-subband Processing Using Convolutional Neural Network
Meng Ge, Longbiao Wang, Masashi Unoki, Sheng Li 0010, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2022 | A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker ExtractionabstractThe speaker extraction algorithm extracts the target speech from a mixture speech containing interference speech and background noise. The extraction process sometimes over-suppresses the extracted target speech, which not only creates artifacts during listening but also harms the performance of downstream automatic speech recognition algorithms. We propose a hybrid continuity loss function for time-domain speaker extraction algorithms to settle the over-suppression problem. On top of the waveform-level loss used for superior signal quality, i.e., SI-SDR, we introduce a multi-resolution delta spectrum loss in the frequency-domain, to ensure the continuity of an extracted speech signal, thus alleviating the over-suppression. We examine the hybrid continuity loss function using a time-domain audio-visual speaker extraction algorithm on the YouTube LRS2-BBC dataset. Experimental results show that the proposed loss function reduces the over-suppression and improves the word error rate of speech recognition on both clean and noisy two-speakers mixtures, without harming the reconstructed speech quality. Zexu Pan, Meng Ge, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2022 | Language-specific Characteristic Assistance for Code-switching Speech Recognition
Tongtong Song, Meng Ge, Longbiao Wang, Yongjie Lv, Yuqin Lin, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2022 | Self-Distillation Based on High-level Information Supervision for Compressing End-to-End ASR Model
Tongtong Song, Longbiao Wang, Yuqin Lin, Yongjie Lv, Meng Ge, Qiang Yu 0005, Jianwu Dang 0001 |
INTERSPEECH | 7 |
| 2022 | MIMO-DoAnet: Multi-channel Input and Multiple Outputs DoA Network with Unknown Number of Sound SourcesabstractRecent neural network based Direction of Arrival (DoA) estimation algorithms have performed well on unknown number of sound sources scenarios.These algorithms are usually achieved by mapping the multi-channel audio input to the single output (i.e.overall spatial pseudo-spectrum (SPS) of all sources), that is called MISO.However, such MISO algorithms strongly depend on empirical threshold setting and the angle assumption that the angles between the sound sources are greater than a fixed angle.To address these limitations, we propose a novel multi-channel input and multiple outputs DoA network called MIMO-DoAnet.Unlike the general MISO algorithms, MIMO-DoAnet predicts the SPS coding of each sound source with the help of the informative spatial covariance matrix.By doing so, the threshold task of detecting the number of sound sources becomes an easier task of detecting whether there is a sound source in each output, and the serious interaction between sound sources disappears during inference stage.Experimental results show that MIMO-DoAnet achieves relative 18.6% and absolute 13.3%, relative 34.4% and absolute 20.2% F1 score improvement compared with the MISO baseline system in 3, 4 sources scenes.The results also demonstrate MIMO-DoAnet alleviates the threshold setting problem and solves the angle assumption problem effectively. Meng Ge, Yanjie Fu, Gaoyan Zhang, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2022 | USEV: Universal Speaker Extraction With Visual CueabstractA speaker extraction algorithm seeks to extract the target speaker's speech from a multi-talker speech mixture. The prior studies focus mostly on speaker extraction from a highly overlapped multi-talker speech mixture. However, the target-interference speaker overlapping ratios could vary over a wide range from 0% to 100% in natural speech communication, furthermore, the target speaker could be absent in the speech mixture, the speech mixtures in such universal multi-talker scenarios are described asgeneral speech mixtures. The speaker extraction algorithm requires an auxiliary reference, such as a video recording or a pre-recorded speech, to form top-down auditory attention on the target speaker. We advocate that a visual cue, i.e., lip movement, is more informative than an audio cue, i.e., pre-recorded speech, to serve as the auxiliary reference for speaker extraction in disentangling the target speaker from ageneral speech mixture. In this paper, we propose a universal speaker extraction network with a visual cue, that works for all multi-talker scenarios. In addition, we propose a scenario-aware differentiated loss function for network training, to balance the network performance over different target-interference speaker pairing scenarios. The experimental results show that our proposed method outperforms various competitive baselines forgeneral speech mixturesin terms of signal fidelity. Zexu Pan, Meng Ge, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Fractional-Order Control of High Speed Train With Actuator Complete FailureabstractWith considering the uncertainties and nonlinearities, the position/velocity tracking control of high-speed train (HST) with redundancy actuators is investigated in this paper. Unlike the existing methods that focus on the actuator partial failure (APF), the actuator complete failure (ACF) is considered. A fractional-order controller (FOC) without a switching process is developed to achieve stable position/velocity tracking control with considering the ACF. And the established controller cannot only tackle the ACF without any fault detection process, but also compensate for uncertainties and nonlinearities of the system, improve the system in tracking accuracy and suppress disturbances caused by internal and external factors. Besides, the influence of different fractional orders on the control performance is analyzed theoretically, providing a basis for the choice of orders. Lyapunov stability theory and fractional order theory are employed to verify the effectiveness of FOC, and simulation studies show the consistency with theoretical analysis. The influence of different fractional orders is also tested in simulation, and further compared with traditional PID controller. It can be concluded that although either PID or FOC are able to cope with the system nonlinearities and uncertainties as well as ACF, there is great significance in FOC research for the promotion of control precision and robustness as well as anti-jamming ability. Meng Ge |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Finite-Time Control of High-Speed Train With Guaranteed Steady-State and Transient PerformanceabstractWith considering the inevitable factors of input nonlinearity, aerodynamic resistance, in-train force, external disturbance and unknown actuator failures, this paper investigates the position/velocity finite-time tracking control problem of high-speed train (HST) to improve the steady-state (tracking accuracy) and transient performance (overshoot and settling time) of the traction/braking control system. Since the traction/braking control system of HST requires rapid reaction speed to ensure the safe and reliable operation, an integer-order finite-time controller (IO-FTC) is developed firstly by utilizing finite-time control theory to improve the system transient performance. On the basis of ensuring the system transient performance, a fractional-order finite-time controller (FO-FTC) is designed by using the fractional stability principle and fractional integral to further improve the system steady-state performance and its robustness and anti-jamming capability. It should be noted that both IO-FTC and FO-FTC are essentially independent of the system model and can produce proper traction/braking force with only the actual and desired position/velocity of the leading vehicle. And the settling time of the system can be adjusted through the selection of different control parameters. The feasibility and effectiveness of the designed controllers are verified by theoretical analysis and simulation studies. And the developed methods are compared with the PID controller. Simulation results make it clear that the proposed methods are superior to PID. Meng Ge |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference SignalsabstractSpeaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full advantage of short reference speech sample. The extracted speech in early stages is used as the reference speech for late stages. For the first time, we use frame-level sequential speech embedding as the reference for target speaker. This is a departure from the traditional utterance-based speaker embedding reference. In addition, a signal fusion scheme is proposed to combine the decoded signals in multiple scales with automatically learned weights. Experiments on WSJ0-2mix and its noisy versions (WHAM! and WHAMR!) show that SpEx++ consistently outperforms other state-of-the-art baselines. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
ICASSP | 1 |
| 2021 | Robust Voice Activity Detection Using a Masked Auditory Encoder Based Convolutional Neural NetworkabstractVoice activity detection (VAD) based on deep learning has achieved remarkable success. However, when the traditional features (e.g., raw waveforms and MFCCs) are directly fed to the deep neural network model, the performance decreases because of noise interference. Here, we propose a robust VAD approach using a masked auditory encoder based convolutional neural network (M-AECNN). First, we analyze the effectiveness of using auditory features as deep learning encoder. These features can roughly simulate the transmission of sound to human inner-ear hair cells; thus, they are more robust than the raw waveform and frequency domain features designed as encoders. Second, similar to the human ear’s masking effect for different speech frequencies, the proposed auditory encoder can further improve the robustness of VAD by increasing the gain for cleaner speech frequencies. Extensive experimental results demonstrate that this approach achieves about 10.5% absolute improvement in the area under the curve on the AURORA-2J dataset compared with a VAD method based on a CNN and MFCCs. Longbiao Wang, Masashi Unoki, Sheng Li 0010, Rui Wang 0102, Meng Ge, Jianwu Dang 0001 |
ICASSP | 6 |
| 2021 | Speech Dereverberation Based on Scale-Aware Mean Square Error Loss
Luya Qiang, Meng Ge, Longbiao Wang, Sheng Li 0010, Jianwu Dang 0001 |
ICONIP (5) | 3 |
| 2021 | Simultaneous Progressive Filtering-Based Monaural Speech Enhancement
Longbiao Wang, Luya Qiang, Sheng Li 0010, Meng Ge, Gaoyan Zhang, Jianwu Dang 0001 |
ICONIP (5) | 6 |
| 2021 | Neural Speaker Extraction with Speaker-Speech Cross-Attention Network
Wupeng Wang, Chenglin Xu, Meng Ge, Haizhou Li 0001 |
Interspeech | 3 |
| 2021 | A Model and Data Hybrid Driven Detection Scheme for IRS-Assisted Massive MIMO SystemsabstractIn this paper, we study the problem of signal detection for massive multiple-input multiple-output (MIMO) communication systems aided by the intelligent reflecting surface (IRS). In the IRS-assisted massive MIMO systems, the signal between the base station and the users is reflected and hence the signal detection is a challenge problem. This paper proposes a model and data hybrid driven detection method to detect the signal of the receiver through neural network when the channel characteristics change. In the auxiliary transmission of IRS, the signal passes through two hop channels, one hop is the channel between transmitter and the IRS, the other hop is the channel between IRS and receiver. According to the different channel characteristics of two hop channel, a novel neural network detector which is combined the data-driven with model-driven is proposed. Experimental results show that the proposed detection scheme can achieve lower computational complexity and higher detection performance than traditional methods. Meng Ge, Fei Li 0014, Ting Li 0003, Yan Liang 0002 |
VTC Fall | 1 |
| 2020 | Spectrograms Fusion with Minimum Difference Masks Estimation for Monaural Speech DereverberationabstractSpectrograms fusion is an effective method for incorporating complementary speech dereverberation systems. Previous linear spectrograms fusion by averaging multiple spectrograms shows outstanding performance. However, various systems with different features cannot apply this simple method. In this study, we design the minimum difference masks (MDMs) to classify the time-frequency (T-F) bins in spectrograms according to the nearest distances from labels. Then, we propose a two-stage nonlinear spectrograms fusion system for speech dereverberation. First, we conduct a multitarget learning-based speech dereverberation front-end model to get spectrograms simultaneously. Then, MDMs are estimated to take the best parts of different spectrograms. We are using spectrograms in the first stage and MDMs in the second stage to recombine T-F bins. The experiments on the REVERB challenge show that a strong feature complementarity between spectrograms and MDMs. Moreover, the proposed framework can consistently and significantly improve PESQ and SRMR, both real and simulated data, e.g., an average PESQ gain of 0.1 in all simulated data and an average SRMR gain of 1.22 in all real data. Longbiao Wang, Meng Ge, Sheng Li 0010, Jianwu Dang 0001 |
ICASSP | 3 |
| 2020 | SpEx+: A Complete Time Domain Speaker Extraction NetworkabstractSpeaker extraction aims to extract the target speech signal from a multi-talker environment given a target speaker's reference speech.We recently proposed a time-domain solution, SpEx, that avoids the phase estimation in frequency-domain approaches.Unfortunately, SpEx is not fully a time-domain solution since it performs time-domain speech encoding for speaker extraction, while taking frequency-domain speaker embedding as the reference.The size of the analysis window for timedomain and the size for frequency-domain input are also different.Such mismatch has an adverse effect on the system performance.To eliminate such mismatch, we propose a complete time-domain speaker extraction solution, that is called SpEx+.Specifically, we tie the weights of two identical speech encoder networks, one for the encoder-extractor-decoder pipeline, another as part of the speaker encoder.Experiments show that the SpEx+ achieves 0.8dB and 2.1dB SDR improvement over the state-of-the-art SpEx baseline, under different and same gender conditions on WSJ0-2mix-extr database respectively. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2020 | Singing Voice Extraction with Attention-Based Spectrograms Fusion
Longbiao Wang, Sheng Li 0010, Chenchen Ding, Meng Ge, Jianwu Dang 0001, Hiroshi Seki |
INTERSPEECH | 5 |
| 2020 | RBFNN-Based Fractional-Order Control of High-Speed Train With Uncertain Model and Actuator FailuresabstractIn this paper, the position/velocity tracking control problem of high speed train(HST) is investigated with considering some inevitable factors such as the input nonlinearity due to different notches of traction/braking forces, aerodynamic resistance, in-train force, external disturbance and unknown actuator failures, which lead to the uncertainty and nonlinearity of HST. Aiming at the system characteristics of HST, a set of integer-order control methods based on the excellent approximation ability of Radial Basis Function Neural Network(RBFNN) are established firstly, and motivated by them, a kind of RBFNN-based fractional-order control methods are proposed by utilizing the genetic attenuation properties of fractional calculus(FC) in order to improve the control performance in the work. It should be pointed out that all the developed methods are able to deal with uncertainties and nonlinearities as well as actuator failures without the need for any “trail and error” process. The feasibility and effectiveness of the proposed control methods are verified by Lyapunov theoretical analysis and numerical simulation studies. Besides, the control performance of integer-order control system and fractional-order control system is compared and analyzed, and the results show that the fractional-order control system is superior. Meng Ge, Xinrui Hu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2019 | A Fast Convolutional Self-attention Based Speech Dereverberation Method for Robust Speech Recognition
Meng Ge, Longbiao Wang, Jianwu Dang 0001 |
ICONIP (3) | 2 |
| 2019 | Environment-Dependent Attention-Driven Recurrent Convolutional Neural Network for Robust Speech Enhancement
Meng Ge, Longbiao Wang, Jianwu Dang 0001, Xiangang Li |
INTERSPEECH | 1 |
| 2018 | Integrative Network Embedding via Deep Joint ReconstructionabstractNetwork embedding is to learn a low-dimensional representation for a network in order to capture intrinsic features of the network. It has been applied to many applications, e.g., network community detection and user recommendation. One of the recent research topics for network embedding has been focusing on exploitation of diverse information, including network topology and semantic information on nodes of networks. However, such diverse information has not been fully utilized nor adequately integrated in the existing methods, so that the resulting network embedding is far from satisfactory. In this paper, we develop a weight-free multi-component network embedding approach by network reconstruction via a deep Autoencoder. Three key components make our new approach effective, i.e., a uniformed graph representation of network topology and semantic information, enhancement to the graph representation using local network structure (i.e., pairwise relationship on nodes) by sampling with latent space regularization, and integration of the diverse information in graph forms in a deep Autoencoder. Extensive experimental results on seven real-world networks demonstrate a superior performance of our method over nine state-of-the-art methods for embedding. Di Jin 0001, Meng Ge, Liang Yang 0002, Dongxiao He, Longbiao Wang, Weixiong Zhang |
IJCAI | 2 |
| 2017 | Using Deep Learning for Community Discovery in Social NetworksabstractCommunity detection is an important task in social network analysis. Existing methods typically use the topological information alone, and ignore the rich information available in the content data. Recently, some researchers have noticed that user profiles can also benefit to community detection, and hence the combination of topology and node contents has become a new hot topic. Some methods using both topology and content have been proposed. However, they often suffer from two drawbacks: 1) they cannot extract a potential deep representation of the network; 2) they cannot automatically weight different information sources with adequate balance parameters. To overcome these issues, we propose a deep integration representation (DIR) algorithm via deep joint reconstruction, which is motivated by the similarity between deep feedforward auto-encoders and spectral clustering in terms of matrix reconstruction. Thanks to spectral clustering which is one of the best community detection methods, the proposed new method is also good at community discovery task. In addition, DIR has further benefit because it not only provides a nonlinear and deep representation of the network, but also learns the most suitable balance between different components automatically. We compare the proposed new approach with nine state-of-the-art community detection methods on eight real relatively large networks. The experimental results show the definite superiority of this new approach. Di Jin 0001, Meng Ge, Wenhuan Lu, Dongxiao He, Françoise Fogelman-Soulié |
ICTAI | 2 |
| 2014 | EPMLR: Sequence-based linear B-cell epitope prediction method using multiple linear regressionabstractBACKGROUND: B-cell epitopes have been studied extensively due to their immunological applications, such as peptide-based vaccine development, antibody production, and disease diagnosis and therapy. Despite several decades of research, the accurate prediction of linear B-cell epitopes has remained a challenging task. RESULTS: In this work, based on the antigen's primary sequence information, a novel linear B-cell epitope prediction model was developed using the multiple linear regression (MLR). A 10-fold cross-validation test on a large non-redundant dataset was performed to evaluate the performance of our model. To alleviate the problem caused by the noise of negative dataset, 300 experiments utilizing 300 sub-datasets were performed. We achieved overall sensitivity of 81.8%, precision of 64.1% and area under the receiver operating characteristic curve (AUC) of 0.728. CONCLUSIONS: We have presented a reliable method for the identification of linear B cell epitope using antigen's primary sequence information. Moreover, a web server EPMLR has been developed for linear B-cell epitope prediction: http://www.bioinfo.tsinghua.edu.cn/epitope/EPMLR/ . Yao Lian, Meng Ge, Xian-Ming Pan |
BMC Bioinform. | 2 |
| 2009 | Study on Multi-Depots Vehicle Scheduling Problem and Its Two-Phase Particle Swarm Optimization
Suxin Wang, Leizhen Wang, Huilin Yuan, Meng Ge, Ben Niu 0002, Weihong Pang, Yuchuan Liu |
ICIC (2) | 4 |