Meng Sun 0001

dblp:81/1237-1 · DBLP profile ↗
← Back
30ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-7435-3752ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AMID: Audio-visual deepfake detection via adaptive multi-dimensional interaction modeling
Chenlong Xue, Kunyuan Li, Meng Sun 0001, Qiang Zhang 0054, Xiongwei Zhang, Kui Yao
Knowl. Based Syst.3
2025 Bayesian Nonparametric Clustering for Source Counting with a Small Aperture Microphone Array
abstract
Source counting (SC) in an indoor environment is an important problem in computational auditory scene analysis. However, the problem is challenging, especially when reverberation and ambient noise are present in the environment. To address this problem, we propose an augmented Bayesian non-parametric (ABNP) clustering algorithm for source counting based on sound intensity (SI) captured by a small aperture microphone array. The core idea is to incorporate an infinite Gaussian mixture model (IGMM) and a time-frequency (TF) augmented weight selection and update scheme for sound intensity estimation. The use of IGMM enables the exemption of the maximum number of sources assumed in previous methods. Experiments on both simulated and real-world data show the improved performance by the proposed method as compared with the state of the art baseline methods.
Kunkun SongGong, Pufen Zhang, Xiongwei Zhang, Wenwu Wang 0001, Meng Sun 0001, Chong Jia
ICASSP5
2025 Cross-domain redundancy exploration by a deep encoder-decoder network for speech steganography
abstract
The technique of speech steganography involves embedding messages within openly transmitted speech channels without arousing suspicion. Nevertheless, current methods for embedding speech in speech suffer from weak imperceptibility and low message speech intelligibility. In this paper, we introduce a novel approach that explores cross-domain redundancy by leveraging a deep encoder–decoder neural network architecture to embed Mel-spectrograms into magnitude spectrograms. Specifically, the message is transformed into its Mel-spectrogram, while the cover is transformed into its magnitude spectrogram. Subsequently, the Mel-spectrogram is embedded as residuals in the magnitude spectrogram through an encoder known as the spectrogram super-resolution network (SSRN). Upon receiving the stego, a decoder network recoveres the Mel-spectrograms of the messages, and a high-fidelity HiFi-GAN vocoder then recovers the message waveform. The encoder–decoder network’s parameters are optimized to ensure imperceptibility and high quality. To validate the superiority of our proposed method, we compare it with recently proposed baselines using common databases such as the LJ Speech and VCTK datasets. Experimental results demonstrate that our method achieves SNRs of 33.83 dB and 30.28 dB for the cover signals on these two datasets, respectively. Furthermore, both the content and speaker identity of the recovered messages are well preserved, and the experiments also confirm the robustness against noises and the security of our approach.
Xiaoyi Ge, Xiongwei Zhang, Meng Sun 0001, Kunkun SongGong
J. Inf. Secur. Appl.3
2025 A dynamically interactable framework with dual-channel security: GAN-based speech steganography for concealed dialogues
Xiaoyi Ge, Xiongwei Zhang, Meng Sun 0001
Knowl. Based Syst.4
2025 Introducing Euclidean distance optimization into Softmax loss under neural collapse
Qiang Zhang 0054, Xiongwei Zhang, Jibin Yang, Meng Sun 0001, Tieyong Cao
Pattern Recognit.4
2025 A Soft-Contrastive Pseudo Learning Approach Toward Open-World Forged Speech Attribution
abstract
Anti-spoofing of deepfake or forged speech is an important technique for the security usage of generative artificial intelligence. Beyond binary classification of real and forged speech, method attribution of forged speech is becoming a practical solution of interpretable anti-spoofing strategies. However, existing related methods have poor performance on analyzing speech forgery methods unseen in their training data, which is inefficient in open-world scenarios with emerging new forgery methods. In this paper, Open-World Forged Speech Attribution (OW-FSA) is firstly defined towards the attribution of forged speech on the methods generating it, where the recognized methods are not limited to the seen ones in training data and the properties of the unseen methods should also be depicted adequately. A novel algorithm, Soft-contrastive Pseudo Learning (SPL), is proposed to address the challenges outlined in OW-FSA, which introduces two key innovations: 1) Based on similarities between features at different scales, the proposed similarity-based soft filtering module filters and matches utterances from the same forgery class to enhance the intra-class compactness of features through contrastive learning. 2) The proposed similarity-based soft pseudo-labeling module integrates label-smoothing-like and similarity weighting techniques to mitigate possible errors in pseudo-labeling. Besides, an iterative algorithm based on SPL is proposed to predict the number of unseen classes. Extensive experiments have validated the superiority of the proposed algorithm over other recently proposed methods on the task of OW-FSA with or without the knowledge of the number of unseen classes. Intuitive visualization and ablation studies have also been conducted to illustrate the advantages of the proposed algorithm. The newly defined task OW-FSA and the proposed algorithm SPL in this paper will help advance the research in speech anti-spoofing.
Qiang Zhang 0054, Xiongwei Zhang, Meng Sun 0001, Jibin Yang
IEEE Trans. Inf. Forensics Secur.3
2024 Multi-Speaker Localization in the Circular Harmonic Domain on Small Aperture Microphone Arrays Using Deep Convolutional Networks
abstract
Acoustic signal processing in the circular harmonic domain (CHD) is an appealing method for speaker localization, since it inherently supports wideband acoustic sources and provides frequency invariant beampatterns. However, the performance of existing circular harmonic direction-of-arrival (DOA) estimation approaches can be degraded by a variety of factors, including background noise and reverberation in the acoustic environments, small aperture size of the circular array and the presence of multiple active sources. This paper addresses these issues by proposing a novel multi-speaker CHD localization method with small-sized microphone arrays using deep convolutional neural networks (CNN). The core idea is to construct circular harmonic features through joining the selected time-frequency (TF) bins of higher power and the operation of a randomization process by mimicking the sparsity property of speech signals. After that, we implement multi-speaker estimation as a multi-label classification task, and propose to use CNN with binary cross-entropy as the loss function. Experimental results show that our method performs significantly better than the baseline methods, on both simulated and real data, in terms of the accuracy of DOA estimation.
Kunkun SongGong, Pufen Zhang, Xiongwei Zhang, Meng Sun 0001, Wenwu Wang 0001
ICASSP4
2024 Scale-aware dual-branch complex convolutional recurrent network for monaural speech enhancement
Meng Sun 0001, Xiongwei Zhang, Hugo Van hamme
Comput. Speech Lang.2
2024 Noise-robust voice conversion using adversarial training with multi-feature decoupling
Xiongwei Zhang, Meng Sun 0001
Eng. Appl. Artif. Intell.4
2024 Target speaker filtration by mask estimation for source speaker traceability in voice conversion
Xiongwei Zhang, Meng Sun 0001, Xia Zou, Chong Jia
Eng. Appl. Artif. Intell.3
2024 An attack-agnostic defense method against adversarial attacks on speaker verification by fusing downsampling and upsampling of speech signals
Xiongwei Zhang, Meng Sun 0001, Yinan Li 0006
Inf. Sci.3
2023 Self-distillation object segmentation via pyramid knowledge representation and transfer
Meng Sun 0001, Tieyong Cao, Xiongwei Zhang, Lixing Xing, Zheng Fang 0012
Multim. Syst.2
2023 Alternating Direction Method of Multipliers for Convolutive Non-Negative Matrix Factorization
abstract
Non-negative matrix factorization (NMF) has become a popular method for learning interpretable patterns from data. As one of the variants of standard NMF, convolutive NMF (CNMF) incorporates an extra time dimension to each basis, known as convolutive bases, which is well suited for representing sequential patterns. Previously proposed algorithms for solving CNMF use multiplicative updates which can be derived by either heuristic or majorization-minimization (MM) methods. However, these algorithms suffer from problems, such as low convergence rates, difficulty to reach exact zeroes during iterations and prone to poor local optima. Inspired by the success of alternating direction method of multipliers (ADMMs) on solving NMF, we explore variable splitting (i.e., the core idea of ADMM) for CNMF in this article. New closed-form algorithms of CNMF are derived with the commonly used β -divergences as optimization objectives. Experimental results have demonstrated the efficacy of the proposed algorithms on their faster convergence, better optima, and sparser results than state-of-the-art baselines.
Yinan Li 0006, Ruili Wang 0001, Yuqiang Fang, Meng Sun 0001, Zhangkai Luo
IEEE Trans. Cybern.4
2022 Waveform level adversarial example generation for joint attacks against both automatic speaker verification and spoofing countermeasures
Xiongwei Zhang, Wei Liu 0005, Xia Zou, Meng Sun 0001, Jian Zhao 0006
Eng. Appl. Artif. Intell.5
2022 Spatial-temporal slowfast graph convolutional network for skeleton-based action recognition
abstract
Abstract In skeleton‐based action recognition, the graph convolutional network (GCN) has achieved great success. Modelling skeleton data in a suitable spatial‐temporal way and designing the adjacency matrix are crucial aspects for GCN‐based methods to capture joint relationships. In this study, we propose the spatial‐temporal slowfast graph convolutional network (STSF‐GCN) and design the adjacency matrices for the skeleton data graphs in STSF‐GCN. STSF‐GCN contains two pathways: (1) the fast pathway is in a high frame rate, and joints of adjacent frames are unified to build ‘small’ spatial‐temporal graphs. A new spatial‐temporal adjacency matrix is proposed for these ‘small’ spatial‐temporal graphs. Ablation studies verify the effectiveness of the proposed adjacency matrix. (2) The slow pathway is in a low frame rate, and joints from all frames are unified to build one ‘big’ spatial‐temporal graph. The adjacency matrix for the ‘big’ spatial‐temporal graph is obtained by computing self‐attention coefficients of each joint. Finally, outputs from two pathways are fused to predict the action category. STSF‐GCN can efficiently capture both long‐range and short‐range spatial‐temporal joint relationships. On three datasets for skeleton‐based action recognition, STSF‐GCN can achieve state‐of‐the‐art performance with much less computational cost.
Zheng Fang 0012, Xiongwei Zhang, Tieyong Cao, Meng Sun 0001
IET Comput. Vis.5
2021 A Survey on the Development of Self-Organizing Maps for Unsupervised Intrusion Detection
Xiaofei Qu, Linru Ma, Meng Sun 0001, Mingxing Ke
Mob. Networks Appl.5
2021 When Automatic Voice Disguise Meets Automatic Speaker Verification
abstract
The technique of transforming voices in order to hide the real identity of a speaker is called voice disguise, among which automatic voice disguise (AVD) by modifying the spectral and temporal characteristics of voices with miscellaneous algorithms are easily conducted with softwares accessible to the public. AVD has posed great threat to both human listening and automatic speaker verification (ASV). In this paper, we have found that ASV is not only a victim of AVD but could be a tool to beat some simple types of AVD. Firstly, three types of AVD, pitch scaling, vocal tract length normalization (VTLN) and voice conversion (VC), are introduced as representative methods. State-of-the-art ASV methods are subsequently utilized to objectively evaluate the impact of AVD on ASV by equal error rates (EER). Moreover, an approach to restore disguised voice to its original version is proposed by minimizing a function of ASV scores w.r.t. restoration parameters. Experiments are then conducted on disguised voices from Voxceleb, a dataset recorded in real-world noisy scenario. The results have shown that, for the voice disguise by pitch scaling, the proposed approach obtains an EER around 7% comparing to the 30% EER of a recently proposed baseline using the ratio of fundamental frequencies. The proposed approach generalizes well to restore the disguise with nonlinear frequency warping in VTLN by reducing its EER from 34.3% to 18.5%. However, it is difficult to restore the source speakers in VC by our approach, where more complex forms of restoration functions or other paralinguistic cues might be necessary to restore the nonlinear transform in VC. Finally, contrastive visualization on ASV features with and without restoration illustrate the role of the proposed approach in an intuitive way.
Linlin Zheng, Jiakang Li, Meng Sun 0001, Xiongwei Zhang, Thomas Fang Zheng
IEEE Trans. Inf. Forensics Secur.3
2019 Detection of People With Camouflage Pattern Via Dense Deconvolution Network
abstract
In this letter, we explore the detection of people with camouflage pattern in cluttered natural scenes. First, considering the lack of open evaluation and training data on camouflaged people detection, a specific dataset of camouflaged people in natural scenes is constructed by us for the first time to the best of our knowledge. Secon, due to the serious corruption of the discrimination of low-level features by the camouflage patterns and cluttered background, we extract the high-level semantic features in deep convolution network and introduce short connections in deconvolution phase, to construct the dense deconvolution network. In training procedure, we augment and shift the images of camouflaged people to generate the proper training data. Attributing to the usage and fusion of semantic information, the proposed network effectively labels the camouflaged people regions as a whole. Finally, we use the superpixel segmentation and spatial smoothness constraint for further improvement of the detection result. Experimental results demonstrate that the proposed method outperforms the classical camouflaged object detection method and typical CNN-based detection methods.
Xiongwei Zhang, Feng Wang 0032, Tieyong Cao, Meng Sun 0001
IEEE Signal Process. Lett.5
2016 Adaptive extraction of repeating non-negative temporal patterns for single-channel speech enhancement
abstract
Estimating unknown background noise from single-channel noisy speech is a key yet challenging problem for speech enhancement. Given the fact that the background noises typically have the repeating property and the foreground speech is sparse and time-variant, many literatures decompose the noisy spectrogram directly in an unsupervised fashion when there is no isolated training example of the target speaker or particular noise types beforehand. However, recently proposed methods suffer from un-interpretable decomposed patterns, neglecting the temporal structure of the background noise or being constrained by the pre-fixed parameters. To settle these issues, we propose a novel method based on autocorrelation technique and convolutive non-negative matrix factorization. The proposed method can adaptively estimate the underlying non-negative repeating temporal patterns from noisy speech and identify the clean speech spectrogram simultaneously. Experiments on NOIZEUS dataset mixed with various real-world background noises showed that the proposed method performs better than some state-of-the-art methods.
Yinan Li 0006, Xiongwei Zhang, Meng Sun 0001, Gang Min, Jibin Yang
ICASSP3
2016 Joint optimization of audible noise suppression and deep neural networks for single-channel speech enhancement
abstract
Improving the perceptual quality of speech signals is a key yet challenging problem for many real world applications. Taking into account the good performance of deep learning in signal representation, a novel single-channel speech enhancement technique is presented based on joint Deep Neural Networks and audible noise suppression as a whole network architecture. This new deep neural network jointly trains an audible noise suppression function which is used to estimate the magnitude spectrum of the clean speech and shape the spectrum of the audible noise at the same time. Experimental results on TIMIT with 20 noise types at various noise levels demonstrate the superiority of the proposed method over the baselines, no matter whether the noise conditions are included in the training set or not.
Xiongwei Zhang, Gang Min, Meng Sun 0001, Jibin Yang
ICME4
2016 Unseen Noise Estimation Using Separable Deep Auto Encoder for Speech Enhancement
abstract
Unseen noise estimation is a key yet challenging step to make a speech enhancement algorithm work in adverse environments. At worst, the only prior knowledge we know about the encountered noise is that it is different from the involved speech. Therefore, by subtracting the components which cannot be adequately represented by a well defined speech model, the noises can be estimated and removed. Given the good performance of deep learning in signal representation, a deep auto encoder (DAE) is employed in this work for accurately modeling the clean speech spectrum. In the subsequent stage of speech enhancement, an extra DAE is introduced to represent the residual part obtained by subtracting the estimated clean speech spectrum (by using the pre-trained DAE) from the noisy speech spectrum. By adjusting the estimated clean speech spectrum and the unknown parameters of the noise DAE, one can reach a stationary point to minimize the total reconstruction error of the noisy speech spectrum. The enhanced speech signal is thus obtained by transforming the estimated clean speech spectrum back into time domain. The above proposed technique is called separable deep auto encoder (SDAE). Given the under-determined nature of the above optimization problem, the clean speech reconstruction is confined in the convex hull spanned by a pre-trained speech dictionary. New learning algorithms are investigated to respect the non-negativity of the parameters in the SDAE. Experimental results on TIMIT with 20 noise types at various noise levels demonstrate the superiority of the proposed method over the conventional baselines.
Meng Sun 0001, Xiongwei Zhang, Hugo Van hamme, Thomas Fang Zheng
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 SegBOMP: An efficient algorithm for block non-sparse signal recovery
abstract
Block sparse signal recovery methods have attracted great interests which take the block structure of the nonzero coefficients into account when clustering. Compared with traditional compressive sensing methods, it can obtain better recovery performance with fewer measurements by utilizing the block-sparsity explicitly. In this paper we propose a segmented-version of the block orthogonal matching pursuit algorithm in which it divides any vector into several sparse sub-vectors. By doing this, the original method can be significantly accelerated due to the dimension reduction of measurements for each segmented vector. Experimental results showed that with low complexity the proposed method yielded identical or even better reconstruction performance than the conventional methods which treated the signal in the standard block-sparsity fashion. Furthermore, in the specific case, where not all segments contain nonzero blocks, the performance improvement can be interpreted as a gain in “effective SNR” in noisy environment.
Xushan Chen, Xiongwei Zhang, Jibin Yang, Meng Sun 0001
ICME4
2015 Supervised Multi-scale Locality Sensitive Hashing
abstract
LSH is a popular framework to generate compact representations of multimedia data, which can be used for content based search. However, the performance of LSH is limited by its unsupervised nature and the underlying feature scale. In this work, we propose to improve LSH by incorporating two elements - supervised hash bit selection and multi-scale feature representation. First, a feature vector is represented by multiple scales. At each scale, the feature vector is divided into segments. The size of a segment is decreased gradually to make the representation correspond to a coarse-to-fine view of the feature. Then each segment is hashed to generate more bits than the target hash length. Finally the best ones are selected from the hash bit pool according to the notion of bit reliability, which is estimated by bit-level hypothesis testing.
Li Weng, I-Hong Jhuo, Miaojing Shi, Meng Sun 0001, Wen-Huang Cheng, Laurent Amsaleg
ICMR4
2015 Speech enhancement based on robust NMF solved by alternating direction method of multipliers
abstract
A robust version of non-negative matrix factorization (RNMF) with generalized Kullback-Leibler divergence designed for the task of unsupervised monaural speech enhancement is proposed. RNMF tackles unsupervised speech enhancement problem through factorizing the magnitude spectrum of mixture into the sum of a non-negative sparse matrix and a non-negative low-rank matrix. The parameters of nonnegative components are estimated through minimizing the reconstruction error defined by the divergence. The closed-from updating formulae of RNMF are derived using alternating direction method of multipliers. Experimental results demonstrated that the proposed algorithm yields superior results compared with the multiplicative updates at the expense of more computational complexity.
Yinan Li 0006, Xiongwei Zhang, Meng Sun 0001, Jingfeng Pan
MMSP3
2015 A stable approach for model order selection in nonnegative matrix factorization
Meng Sun 0001, Xiongwei Zhang, Hugo Van hamme
Pattern Recognit. Lett.1
2015 Speech Enhancement Under Low SNR Conditions Via Noise Estimation Using Sparse and Low-Rank NMF with Kullback-Leibler Divergence
abstract
A key stage in speech enhancement is noise estimation which usually requires prior models for speech or noise or both. However, prior models can sometimes be difficult to obtain. In this paper, without any prior knowledge of speech and noise, sparse and low-rank nonnegative matrix factorization (NMF) with Kullback-Leibler divergence is proposed to noise and speech estimation by decomposing the input noisy magnitude spectrogram into a low-rank noise part and a sparse speech-like part. This initial unsupervised speech-noise estimation allows us to set a subsequent regularized version of NMF or convolutional NMF to reconstruct the noise and speech spectrogram, either by estimating a speech dictionary on the fly (categorized as unsupervised approaches) or by using a pre-trained speech dictionary on utterances with disjoint speakers (categorized as semi-supervised approaches). Information fusion was investigated by taking the geometric mean of the outputs from multiple enhancement algorithms. The performance of the algorithms were evaluated on five metrics (PESQ, SDR, SNR, STOI, and OVERALL) by making experiments on TIMIT with 15 noise types. The geometric means of the proposed unsupervised approaches outperformed spectral subtraction (SS), minimum mean square estimation (MMSE) under low input SNR conditions. All the proposed semi-supervised approaches showed superiority over SS and MMSE and also obtained better performance than the state-of-the-art algorithms which utilized a prior noise or speech dictionary under low SNR conditions.
Meng Sun 0001, Yinan Li 0006, Jort F. Gemmeke, Xiongwei Zhang
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Joint training of non-negative Tucker decomposition and discrete density hidden Markov models
Meng Sun 0001, Hugo Van hamme
Comput. Speech Lang.1
2012 Tri-factorization learning of sub-word units with application to vocabulary acquisition
abstract
In prior work, we proposed a method for vocabulary acquisition based on a co-occurrence model and non-negative matrix factorization. The vocabulary is described in terms of co-occurrence statistics of frame-level acoustic descriptions and suffers from poor scalability to larger vocabularies. Much like whole-word HMM models, there is no reuse of a sub-word units such as phone models. In this paper, we apply the co-occurrence framework to learn a set of sub-word units unsupervisedly using a matrix tri-factorization and propose a method for computing their posteriorgram and finally show vocabulary acquisition from the posteriorgram. The method outperforms our prior work in that it can learn from a smaller set of labeled data and shows a better recognition accuracy.
Meng Sun 0001, Hugo Van hamme
ICASSP1
2011 Unsupervised vocabulary discovery using non-negative matrix factorization with graph regularization
abstract
In this paper, we present a model for unsupervised pattern discovery using non-negative matrix factorization (NMF) with graph regularization. Though the regularization can be applied to many applications, we illustrate its effectiveness in a task of vocabulary acquisition in which a spoken utterance is represented by its histogram of the acoustic co-occurrences. The regularization expresses that temporally close co-occurrences should tend to end up in the same learned pattern. A novel algorithm that converges to a local optimum of the regularized cost function is proposed. Our experiments show that the graph regularized NMF model always performs better than the primary NMF model on the task of unsupervised acquisition of a small vocabulary.
Meng Sun 0001, Hugo Van hamme
ICASSP1
2011 Image pattern discovery by using the spatial closeness of visual code words
abstract
A graph regularized non-negative matrix factorization (NMF) model is proposed for image pattern discovery. Each image is represented by its histogram of visual words (i.e. bag-of-words) and the image contents are discovered by the NMF model. The graph regularization preserves the spatial closeness of visual code words in the obtained patterns, thus improving the bag-of-words representation against its main shortcoming: the loss of spatial information. Experiments on a subset of the Caltech256 database show the efficacy of the proposed model.
Meng Sun 0001, Hugo Van hamme
ICIP1