EDBT 2026 Demo / reviewers in the wild / expert
Yanxiong Li
dblp:48/4913 · also Yan-Xiong Li
· DBLP profile ↗
41ranked-venue papers
18as first author
27since 2021 · last 2025
0000-0003-4362-1125ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 12 first-author · 22 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 12 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep Enhancement Spotting Network for Low-complexity Keyword Spotting in Noisy EnvironmentsabstractKeyword Spotting (KWS) is crucial for hands-free voice-activated systems, requiring a balance between accuracy and complexity, especially in noisy environments. While Speech Enhancement (SE) can improve KWS accuracy, existing methods often lack the ability to effectively utilize the rich features produced during enhancement. In this paper, we design a low-complexity network to address the challenges of KWS in noisy environments. We integrate the tasks of both SE and KWS into a unified network that learns a shared representation from both tasks. The proposed network features two blocks: a Residual Full-band and Sub-band Fusion (RFSF) block, and a Deformable Transition (DT) block. Our dual-task network surpasses existing KWS models in accuracy with low complexity, making it suitable for deployment on edge devices. Qianhua He, Yanxiong Li, Zunxian Liu, Mingru Yang, Jinxin Huang |
ICASSP | 3 |
| 2025 | Cross-Domain Few-Shot Open-Set Keyword Spotting Using Keyword Adaptation and Prototype ReprojectionabstractPersonalized keyword spotting (KWS) with few enrollment utterances remains an important problem over years. KWS remains a challenging task due to the following factors, including the scarcity of enrollment samples, speech variation in the open-set scenarios, and distributional gap between source and target domains. In this paper, we formulate a KWS task of Cross-Domain Few-Shot Open-Set (CD-FSOS) and propose a dedicated framework Adapt-KWS to bridge the distribution gap between the source domain and target open-set domain with quite limited enrollment data. The proposed Adapt-KWS consists of a set of Custom-Keyword Adapters (CKAs) and a Prototype Reprojection Module (PRM). CKAs enable the efficient adaptation to new target tasks with limited training samples, aiming to improve cross-domain generalization. PRM reprojects the support prototypes into the query embedding space to enhance their alignment, mitigating the potential covariate shift between open-set queries and enrollments. Experimental results demonstrate the effectiveness of our framework and proposed modules on multiple datasets. Code will be available at: https://github.com/Raynaming/CD-FSOS-KWS. Mingru Yang, Qianhua He, Jinxin Huang, Zunxian Liu, Yanxiong Li |
ICASSP | 6 |
| 2025 | Fully Few-shot Class-incremental Audio Classification Using Multi-level Embedding Extractor and Ridge Regression Classifier
Yongjie Si, Yanxiong Li, Jiaxin Tan, Qianhua He, Il-Youp Kwak |
INTERSPEECH | 2 |
| 2025 | Generalizable Audio Deepfake Detection via Hierarchical Structure Learning and Feature Whitening in Poincaré sphere
Mingru Yang, Yanmei Gu, Qianhua He, Yanxiong Li, Peirong Zhang 0001, Huijia Zhu, Weiqiang Wang 0002 |
INTERSPEECH | 4 |
| 2025 | Infant Cry Emotion Recognition Using Improved ECAPA-TDNN with Multi-scale Feature Fusion and Attention Enhancement
Yanxiong Li, Haolin Yu |
INTERSPEECH | 2 |
| 2025 | Infant Cry Detection In Noisy Environment Using Blueprint Separable Convolutions and Time-Frequency Recurrent Neural NetworkabstractInfant cry detection is a crucial component of baby care system. In this paper, we propose a lightweight and robust method for infant cry detection. The method leverages blueprint separable convolutions to reduce computational complexity, and a time-frequency recurrent neural network for adaptive denoising. The overall framework of the method is structured as a multi-scale convolutional recurrent neural network, which is enhanced by efficient spatial attention mechanism and contrast-aware channel attention module, and acquire local and global information from the input feature of log Mel-spectrogram. Multiple public datasets are adopted to create a diverse and representative dataset, and environmental corruption techniques are used to generate the noisy samples encountered in real-world scenarios. Results show that our method exceeds many state-of-the-art methods in accuracy, F1-score, and complexity under various signal-to-noise ratio conditions. The code is at https://github.com/fhfjsd1/ICD_MMSP. Haolin Yu, Yanxiong Li |
MMSP | 2 |
| 2025 | Low-complexity speaker embedding module with feature segmentation, transformation and reconstruction for few-shot speaker identification
Yanxiong Li, Qisheng Huang, Xiaofen Xing, Xiangmin Xu 0001 |
Expert Syst. Appl. | 1 |
| 2024 | Low-Complexity Acoustic Scene Classification Using Parallel Attention-Convolution Network
Yanxiong Li, Jiaxin Tan, Jialong Li 0002, Yongjie Si, Qianhua He |
INTERSPEECH | 1 |
| 2024 | Fully Few-shot Class-incremental Audio Classification Using Expandable Dual-embedding Extractor
Yongjie Si, Yanxiong Li, Jialong Li 0002, Jiaxin Tan, Qianhua He |
INTERSPEECH | 2 |
| 2024 | Detecting video anomalies by jointly utilizing appearance and skeleton information
Wenfeng Pang, Qianhua He, Yanxiong Li, Noman Ahmed |
Expert Syst. Appl. | 3 |
| 2024 | Lightweight Speaker Verification Using Transformation Module With Feature Partition and FusionabstractAlthough many efforts have been made on decreasing the model complexity for speaker verification, it is still challenging to deploy speaker verification systems with satisfactory result on low-resource terminals. We design a transformation module that performs feature partition and fusion to implement lightweight speaker verification. The transformation module consists of multiple simple but effective operations, such as convolution, pooling, mean, concatenation, normalization, and element-wise summation. It works in a plug-and-play way, and can be easily implanted into a wide variety of models to reduce the model complexity while maintaining the model error. First, the input feature is split into several low-dimensional feature subsets for decreasing the model complexity. Then, each feature subset is updated by fusing it with the inter-feature-subsets correlational information to enhance its representational capability. Finally, the updated feature subsets are independently fed into the block (one or several layers) of the model for further processing. The features that are output from current block of the model are processed according to the steps above before they are fed into the next block of the model. Experimental data are selected from two public speech corpora (namely VoxCeleb1 and VoxCeleb2). Results show that implanting the transformation module into three models (namely AMCRN, ResNet34, and ECAPA-TDNN) for speaker verification slightly increases the model error and significantly decreases the model complexity. Our proposed method outperforms baseline methods on the whole in memory requirement and computational complexity with lower equal error rate. It also generalizes well across truncated segments with various lengths. Yanxiong Li, Zhongjie Jiang, Qisheng Huang, Wenchang Cao, Jialong Li 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | Few-Shot Class-Incremental Audio Classification With Adaptive Mitigation of Forgetting and OverfittingabstractFew-shot Class-incremental Audio Classification (FCAC) is a task to continuously identify incremental classes with only few training samples after training the model on base classes with abundant samples. The key to solving the FCAC problem is to ensure that the model has good stability (without forgetting base classes) and strong plasticity (without overfitting incremental classes). In this paper, we propose a FCAC method which is able to adaptively mitigate the model's forgetting of base classes and overfitting of incremental classes. Our model consists of an embedding extractor and an expandable classifier. The former is the backbone of a residual network and is frozen after being trained using sufficient samples of base classes, whereas the latter can be expandable and is updated using few training samples of incremental classes in each incremental session. The expandable classifier consists of two branches and one fusion module. The two branches are designed to mitigate the model's forgetting of base classes and overfitting of incremental classes, respectively. The fusion module is designed to adaptively fuse predictions output by the above two branches. In addition, we define two losses for model training in the base and incremental sessions, respectively. Three experimental datasets (NSynth-100, FSC-89 and LS-100) are created by randomly choosing samples from audio corpora of NSynth, FSD-MIX-CLIP and LibriSpeech, respectively. Experimental results demonstrate that our proposed method outperforms all previous methods in accuracy and has advantage over most previous methods in computational load. The code is available athttps://github.com/Jialongdustin/AMFO. Yanxiong Li, Jialong Li 0002, Yongjie Si, Jiaxin Tan, Qianhua He |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | Few-Shot Class-Incremental Audio Classification Using Dynamically Expanded Classifier With Self-Attention Modified PrototypesabstractMost existing methods for audio classification assume that the vocabulary of audio classes to be classified is fixed. When novel (unseen) audio classes appear, audio classification systems need to be retrained with abundant labeled samples of all audio classes for recognizing base (initial) and novel audio classes. If novel audio classes continue to appear, the existing methods for audio classification will be inefficient and even infeasible. In this work, we propose a method for few-shot class-incremental audio classification, which can continually recognize novel audio classes without forgetting old ones. The framework of our method mainly consists of two parts: an embedding extractor and a classifier, and their constructions are decoupled. The embedding extractor is the backbone of a ResNet based network, which is frozen after construction by a training strategy using only samples of base audio classes. However, the classifier consisting of prototypes is expanded by a prototype adaptation network with few samples of novel audio classes in incremental sessions. Labeled support samples and unlabeled query samples are used to train the prototype adaptation network and update the classifier, since they are informative for audio classification. Three audio datasets, named NSynth-100, FSC-89 and LS-100 are built by choosing samples from audio corpora of NSynth, FSD-MIX-CLIP and LibriSpeech, respectively. Results show that our method exceeds baseline methods in average accuracy and performance dropping rate. In addition, it is competitive compared to baseline methods in computational complexity and memory requirement. Yanxiong Li, Wenchang Cao, Wei Xie 0013, Jialong Li 0002, Emmanouil Benetos |
IEEE Trans. Multim. | 1 |
| 2023 | Clean Sample Guided Self-Knowledge Distillation for Image ClassificationabstractFor two-stage knowledge distillation, the combination with Data Augmentation (DA) is straightforward and effective. Yet, for online Self-knowledge Distillation (SD), DA is not always beneficial because of the absence of a trustworthy teacher model. To address this issue, this paper proposes an SD method named Clean sample guided Self-knowledge Distillation (CleanSD), in which the original clean sample is used as a guide when the model is trained with the augmented samples. The implementation of the CleanSD comes with two DA techniques, namely Mixup (for label-mixing) and Cutout (for label-preserving). Results on CIFAR-100 demonstrate that error rates obtained by the proposed CleanSD are reduced by 2.59%, 1.39%, and 0.47-1.20%, compared to that obtained by the baseline, the vanilla DA techniques, and other peer SD methods, respectively. In addition, the effectiveness and robustness of the CleanSD are verified across multiple DA methods and datasets. Jiyue Wang, Yanxiong Li, Qianhua He, Wei Xie 0013 |
ICASSP | 2 |
| 2023 | Few-shot Class-incremental Audio Classification Using Stochastic Classifier
Yanxiong Li, Wenchang Cao, Jialong Li 0002, Wei Xie 0013, Qianhua He |
INTERSPEECH | 1 |
| 2023 | Few-shot Class-incremental Audio Classification Using Adaptively-refined PrototypesabstractNew classes of sounds constantly emerge with a few samples, making it challenging for models to adapt to dynamic acoustic environments. This challenge motivates us to address the new problem of few-shot class-incremental audio classification. This study aims to enable a model to continuously recognize new classes of sounds with a few training samples of new classes while remembering the learned ones. To this end, we propose a method to generate discriminative prototypes and use them to expand the model's classifier for recognizing sounds of new and learned classes. The model is first trained with a random episodic training strategy, and then its backbone is used to generate the prototypes. A dynamic relation projection module refines the prototypes to enhance their discriminability. Results on two datasets (derived from the corpora of Nsynth and FSD-MIX-CLIPS) show that the proposed method exceeds three state-of-the-art methods in average accuracy and performance dropping rate. Wei Xie 0013, Yanxiong Li, Qianhua He, Wenchang Cao, Tuomas Virtanen |
INTERSPEECH | 2 |
| 2023 | Few-shot class-incremental audio classification via discriminative prototype learning
Wei Xie 0013, Yanxiong Li, Qianhua He, Wenchang Cao |
Expert Syst. Appl. | 2 |
| 2023 | Few-Shot Speaker Identification Using Lightweight Prototypical Network With Feature Grouping and InteractionabstractExisting methods for few-shot speaker identification (FSSI) obtain high accuracy, but their computational complexities and model sizes need to be reduced for lightweight applications. In this work, we propose a FSSI method using a lightweight prototypical network with the final goal to implement the FSSI on intelligent terminals with limited resources, such as smart watches and smart speakers. In the proposed prototypical network, an embedding module is designed to perform feature grouping for reducing the memory requirement and computational complexity, and feature interaction for enhancing the representational ability of the learned speaker embedding. In the proposed embedding module, audio feature of each speech sample is split into several low-dimensional feature subsets that are transformed by a recurrent convolutional block in parallel. Then, the operations of averaging, addition, concatenation, element-wise summation and statistics pooling are sequentially executed to learn a speaker embedding for each speech sample. The recurrent convolutional block consists of a block of bidirectional long short-term memory, and a block of de-redundancy convolution in which feature grouping and interaction are conducted too. Our method is compared to baseline methods on three datasets that are selected from three public speech corpora (VoxCeleb1, VoxCeleb2, and LibriSpeech). The results show that our method obtains higher accuracy under several conditions, and has advantages over all baseline methods in computational complexity and model size. Yanxiong Li, Wenchang Cao, Qisheng Huang, Qianhua He |
IEEE Trans. Multim. | 1 |
| 2023 | Audiovisual Dependency Attention for Violence Detection in VideosabstractViolence detection in videos can help maintain public order, detect crimes, or provide timely assistance. In this paper, we aim to leverage multimodal information to determine whether successive frames contain violence. Specifically, we propose an audiovisual dependency attention (AVD-attention) module modified from the co-attention architecture to fuse visual and audio information, unlike commonly used methods such as the feature concatenation, addition, and score fusion. Because the AVD-attention module’s dependency map contains sufficient fusion information, we argue that it should be applied more sufficiently. A combination pooling method is utilized to convert the dependency map to an attention vector, which can be considered a new feature that includes fusion information or a mask of the attention feature map. Since some information in the input feature might be lost after processing by attention modules, we employ a multimodal low-rank bilinear method that considers all pairwise interactions among two features in each time step to complement the original information for output features of the module. AVD-attention outperformed co-attention in experiments on the XD-Violence dataset. Our system outperforms state-of-the-art systems. Wenfeng Pang, Wei Xie 0013, Qianhua He, Yanxiong Li |
IEEE Trans. Multim. | 4 |
| 2022 | Domestic Activity Clustering from Audio via Depthwise Separable Convolutional Autoencoder NetworkabstractAutomatic estimation of domestic activities from audio can be used to solve many problems, such as reducing the labor cost for nursing the elderly people. This study focuses on solving the problem of domestic activity clustering from audio. The target of domestic activity clustering is to cluster audio clips which belong to the same category of domestic activity into one cluster in an unsupervised way. In this paper, we propose a method of domestic activity clustering using a depthwise separable convolutional autoencoder network. In the proposed method, initial embeddings are learned by the depthwise separable convolutional autoencoder, and a clustering-oriented loss is designed to jointly optimize embedding refinement and cluster assignment. Different methods are evaluated on a public dataset (a derivative of the SINS dataset) used in the challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) in 2018. Our method obtains the normalized mutual information (NMI) score of 54.46%, and the clustering accuracy (CA) score of 63.64%, and outperforms state-of-the-art methods in terms of NMI and CA. In addition, both computational complexity and memory requirement of our method is lower than that of previous deep-model-based methods. Codes: https://github.com/vinceasvp/domestic-activity-clustering-from-audio Yanxiong Li, Wenchang Cao, Konstantinos Drossos, Tuomas Virtanen |
MMSP | 1 |
| 2022 | Predicting skeleton trajectories using a Skeleton-Transformer for video anomaly detection
Wenfeng Pang, Qianhua He, Yanxiong Li |
Multim. Syst. | 3 |
| 2022 | Fall event detection with global and temporal local information in real-world videos
Wenfeng Pang, Qianhua He, Yuanfeng Chen, Yanxiong Li |
Multim. Tools Appl. | 4 |
| 2021 | Domestic Activities Clustering From Audio Recordings Using Convolutional Capsule Autoencoder NetworkabstractRecent efforts have been made on domestic activities classification from audio recordings, especially the works submitted to the challenge of DCASE (Detection and Classification of Acoustic Scenes and Events) since 2018. In contrast, few studies were done on domestic activities clustering, which is a newly emerging problem. Domestic activities clustering from audio recordings aims at merging audio clips which belong to the same class of domestic activity into a single cluster. Domestic activities clustering is an effective way for unsupervised estimation of daily activities performed in home environment. In this study, we propose a method for domestic activities clustering using a convolutional capsule autoencoder network (CCAN). In the method, the deep embeddings are learned by the autoencoder in the CCAN, while the deep embeddings which belong to the same class of domestic activities are merged into a single cluster by a clustering layer in the CCAN. Evaluated on a public dataset adopted in DCASE- 2018 Task 5, the results show that the proposed method outperforms state-of-the-art methods in terms of the metrics of clustering accuracy and normalized mutual information. Ziheng Lin, Yanxiong Li, Zhangjin Huang, Yufeng Tan, Yichun Chen, Qianhua He |
ICASSP | 2 |
| 2021 | Violence Detection in Videos Based on Fusing Visual and Audio InformationabstractDetermining whether given video frames contain violent content is a basic problem in violence detection. Visual and audio information are useful for detecting violence included in a video, and are usually complementary; however, violence detection studies focusing on fusing visual and audio information are relatively rare. Therefore, we explored methods for fusing visual and audio information. We proposed a neural network containing three modules for fusing multimodal information: 1) attention module for utilizing weighted features to generate effective features based on the mutual guidance between visual and audio information; 2) fusion module for integrating features by fusing visual and audio information based on the bilinear pooling mechanism; and 3) mutual Learning module for enabling the model to learn visual information from another neural network with a different architecture. Experimental results indicated that the proposed neural network outperforms existing state-of-the-art methods on the XD-Violence dataset. Wen-Feng Pang, Qianhua He, Yongjian Hu, Yanxiong Li |
ICASSP | 4 |
| 2021 | A Stage Match for Query-by-Example Spoken Term Detection Based On Structure Information of QueryabstractThe state-of-the-art of query-by-example spoken term detection (QbE-STD) strategies are usually based on segmental dynamic time warping (S-DTW). However, the sliding window in S-DTW may separate signal of a word into different segments and produce many illegal candidates required to be compared with the query, which significantly reduce the accuracy and efficiency of detection. In this paper, we propose a stage match strategy based on the structure information of the query, represented with the unvoiced-voiced attribute of the portions in itself. The strategy first locates potential candidates with similar structure against the query in utterances, and further matches the query with Type-Location DTW (TL-DTW), which is a modified DTW with the constraints of pronunciation types and relative positions of paired frames in the voiced sub-segments. Experiments on AISHELL-1 Corpus showed that the proposed approach achieved a relative improvement of 30.51% in AUC against S-DTW and speeded up the retrieval. Junyao Zhan, Qianhua He, Jianbin Su, Yanxiong Li |
ICASSP | 4 |
| 2021 | Domestic Activities Classification from Audio Recordings Using Multi-scale Dilated Depthwise Separable Convolutional NetworkabstractDomestic activities classification (DAC) from audio recordings aims at classifying audio recordings into predefined categories of domestic activities, which is an effective way for estimation of daily activities performed in home environment. In this paper, we propose a method for DAC from audio recordings using a multi-scale dilated depthwise separable convolutional network (DSCN). The DSCN is a lightweight neural network with small size of parameters and thus suitable to be deployed in portable terminals with limited computing resources. To expand the receptive field with the same size of DSCN’s parameters, dilated convolution, instead of normal convolution, is used in the DSCN for further improving the DSCN’s performance. In addition, the embeddings of various scales learned by the dilated DSCN are concatenated as a multi-scale embedding for representing property differences among various classes of domestic activities. Evaluated on a public dataset of the Task 5 of the 2018 challenge on Detection and Classification of Acoustic Scenes and Events (DCASE-2018), the results show that: both dilated convolution and multi-scale embedding contribute to the performance improvement of the proposed method; and the proposed method outperforms the methods based on state-of-the-art lightweight network in terms of classification accuracy. Yufei Zeng, Yanxiong Li, Zhenfeng Zhou, Difeng Lu |
MMSP | 2 |
| 2021 | Speaker Clustering by Co-Optimizing Deep Representation Learning and Cluster EstimationabstractSpeaker clustering is a task to merge speech segments uttered by the same speaker into a single cluster, which is an effective tool for alleviating the management of massive amount of audio documents. In this paper, we present a work for co-optimizing the two main steps of speaker clustering, namely, feature learning and cluster estimation. In our method, the deep representation feature is learned by a deep convolutional autoencoder network (DCAN), while the cluster estimation is realized by a softmax layer that is combined with the DCAN. We devise an integrated loss function to simultaneously minimize the reconstruction loss (for deep representation learning) and the clustering loss (for cluster estimation). Many state-of-the-art audio features and clustering methods are evaluated on experimental datasets selected from two publicly available speech corpora (the AISHELL-2 and the VoxCeleb1). The results show that the proposed method exceeds other speaker clustering methods in regard to the normalized mutual information (NMI) and the clustering accuracy (CA). Additionally, the proposed deep representation feature outperforms other features that were widely used in previous works, in terms of both NMI and CA. Yanxiong Li, Wucheng Wang, Mingle Liu, Zhongjie Jiang, Qianhua He |
IEEE Trans. Multim. | 1 |
| 2020 | Sound Event Detection Via Dilated Convolutional Recurrent Neural NetworksabstractConvolutional recurrent neural networks (CRNNs) have achieved state-of-the-art performance for sound event detection (SED). In this paper, we propose to use a dilated CRNN, namely a CRNN with a dilated convolutional kernel, as the classifier for the task of SED. We investigate the effectiveness of dilation operations which provide a CRNN with expanded receptive fields to capture long temporal context without increasing the amount of CRNN's parameters. Compared to the classifier of the baseline CRNN, the classifier of the dilated CRNN obtains a maximum increase of 1.9%, 6.3% and 2.5% at F1 score and a maximum decrease of 1.7%, 4.1% and 3.9% at error rate (ER), on the publicly available audio corpora of the TUTSED Synthetic 2016, the TUT Sound Event 2016 and the TUT Sound Event 2017, respectively. Yanxiong Li, Mingle Liu, Konstantinos Drossos, Tuomas Virtanen |
ICASSP | 1 |
| 2020 | Sound Event Detection with Depthwise Separable and Dilated ConvolutionsabstractState-of-the-art sound event detection (SED) methods usually employ a series of convolutional neural networks (CNNs) to extract useful features from the input audio signal, and then recurrent neural networks (RNNs) to model longer temporal context in the extracted features. The number of the channels of the CNNs and size of the weight matrices of the RNNs have a direct effect on the total amount of parameters of the SED method, which is to a couple of millions. Additionally, the usually long sequences that are used as an input to an SED method along with the employment of an RNN, introduce implications like increased training time, difficulty at gradient flow, and impeding the parallelization of the SED method. To tackle all these problems, we propose the replacement of the CNNs with depthwise separable convolutions and the replacement of the RNNs with dilated convolutions. We compare the proposed method to a baseline convolutional neural network on a SED task, and achieve a reduction of the amount of parameters by 85% and average training time per epoch by 78%, and an increase the average frame-wise F1score and reduction of the average error rate by 4.6% and 3.8%, respectively. Konstantinos Drossos, Stylianos I. Mimilakis, Shayan Gharib, Yanxiong Li, Tuomas Virtanen |
IJCNN | 4 |
| 2020 | Acoustic Scene Clustering Using Joint Optimization of Deep Embedding Learning and Clustering IterationabstractRecent efforts have been made on acoustic scene classification in the audio signal processing community. In contrast, few studies have been conducted on acoustic scene clustering, which is a newly emerging problem. Acoustic scene clustering aims at merging the audio recordings of the same class of acoustic scene into a single cluster without using prior information and training classifiers. In this study, we propose a method for acoustic scene clustering that jointly optimizes the procedures of feature learning and clustering iteration. In the proposed method, the learned feature is a deep embedding that is extracted from a deep convolutional neural network (CNN), while the clustering algorithm is the agglomerative hierarchical clustering (AHC). We formulate a unified loss function for integrating and optimizing these two procedures. Various features and methods are compared. The experimental results demonstrate that the proposed method outperforms other unsupervised methods in terms of the normalized mutual information and the clustering accuracy. In addition, the deep embedding outperforms many state-of-the-art features. Yanxiong Li, Mingle Liu, Wucheng Wang, Qianhua He |
IEEE Trans. Multim. | 1 |
| 2019 | Acoustic event diarization in TV/movie audios using deep embedding and integer linear programming
Yanxiong Li, Xianku Li, Mingle Liu, Wucheng Wang, Ji-Chen Yang |
Multim. Tools Appl. | 1 |
| 2018 | Frontal Face Generation from Multiple Pose-Variant Faces with CGAN in Real-World Surveillance SceneabstractIt is well known that frontal face is much easier to be recognized than pose-variant face for both human and machine perception. However, it is not easy to acquire a frontal face in real-world video surveillance. This paper proposes a method to synthetize a frontal face for recognition in video surveillance scene, which is based on Conditional Generative Adversarial Networks (cGAN) with input of multiple pose-variant faces from a video. Experimental results show that the proposed approach can generate suitable frontal faces and improve face recognition by around 20% on a dataset of 43276 face images from 19 persons, collected from the real-world video surveillance scene. The effectiveness of multiple frames against single frame as input is demonstrated. Moreover, we investigate the generator with different depth for synthetizing frontal faces, in which an up-down sampling trick is designed for synthetizing higher quality frontal face images and boosts the performance of the generator. Zhu-Liang Chen, Qianhua He, Wen-Feng Pang, Yanxiong Li |
ICASSP | 4 |
| 2018 | Dictionary learning based on M-PCA-N for audio signal sparse representationabstractThe current popular dictionary learning algorithms for sparse representation of signals are K‐means Singular Value Decomposition (K‐SVD) and K‐SVD‐extended. Only rank‐1 approximation is used to update one atom at a time and it is unable to cope with large dictionary efficiently. In order to tackle these two problems, this study proposes M‐Principal Component Analysis‐N (M‐PCA‐N), which is an algorithm for dictionary learning and sparse representation. First, M‐Principal Component Analysis (M‐PCA) utilised information from the top M ranks of SVD decomposition to update M atoms at a time. Then, in order to further utilise the information from remaining ranks, M‐PCA‐N is proposed on the basis of M‐PCA, by transforming information from the following N non‐principal ranks onto the top M principal ranks. The mathematic formula indicates that M‐PCA may be seen as a generalisation of K‐SVD. Experimental results on the BBC Sound Effects Library show that M‐PCA‐N not only lowers the MSE between original signal and approximation signal in audio signal sparse representation, but also obtains higher audio signal classification precision than K‐SVD. Ji-Chen Yang, Qianhua He, Yanxiong Li, Lei-an Liu, Xiaohui Feng |
IET Signal Process. | 3 |
| 2018 | Using multi-stream hierarchical deep neural network to extract deep audio feature for acoustic event detection
Yanxiong Li, Hai Jin 0008, Xianku Li, Qin Wang 0014, Qianhua He |
Multim. Tools Appl. | 1 |
| 2018 | Mobile Phone Clustering From Speech Recordings Using Deep Representation and Spectral ClusteringabstractConsiderable attention has been paid to acquisition device recognition over the past decade in the forensic community, especially in digital image forensics. In contrast, acquisition device clustering from speech recordings is a new problem that aims to merge the recordings acquired by the same device into a single cluster without having prior information about the recordings and training classifiers in advance. In this paper, we propose a method for mobile phone clustering from speech recordings by using a new feature of deep representation and a spectral clustering algorithm. The new feature is learned by a deep auto-encoder network for representing the intrinsic trace left behind by each phone in the recordings, and spectral clustering is used to merge recordings acquired by the same phone into a single cluster. The impacts of the structures of the deep auto-encoder network on the performance of the new feature are discussed. Different features are compared with one another. The proposed method is compared with others and evaluated under special conditions. The results show that the proposed method is effective under these conditions and the new feature outperforms other features. Yanxiong Li, Xianku Li, Ji-Chen Yang, Qianhua He |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Mobile phone clustering from acquired speech recordings using deep Gaussian supervector and spectral clusteringabstractAcquisition device clustering from speech recordings is a new and critical problem in the field of speech forensic, which aims at merging speech recordings acquired by the same device into one cluster without both pre-knowing prior information of the processed data and pre-training classifier. We propose a mobile phone clustering method, in which deep Gaussian supervector learned by deep neural network is used to represent the intrinsic trace left behind by mobile phone in speech recordings, and then spectral clustering technique is adopted to merge speech recordings acquired by the same mobile phone into one cluster. The performance of the proposed method is evaluated on a public corpus of speech recordings acquired by mobile phones. The results show that the proposed method is effective for mobile phone clustering from acquired speech recordings. Yanxiong Li, Xianku Li, Xiaohui Feng, Ji-Chen Yang, Aiwu Chen, Qianhua He |
ICASSP | 1 |
| 2017 | Unsupervised classification of speaker roles in multi-participant conversational speech
Yanxiong Li, Qin Wang 0014, Xinchao Li, Ji-Chen Yang, Xiaohui Feng, Qianhua He |
Comput. Speech Lang. | 1 |
| 2017 | Sparse representation-based quasi-clean speech construction for speech quality assessment under complex environmentsabstractA non‐intrusive speech quality assessment method for complex environments was proposed. In the proposed approach, a new sparse representation‐based speech reconstruction algorithm was presented to acquire the quasi‐clean speech from the noisy degraded signal. Firstly, an over‐complete dictionary of the clean speech power spectrum was learned by the K‐singular value decomposition algorithm. Then in the sparse representation stage, the stopping residue error was adaptively achieved according to the estimated cross‐correlation and the noise spectrum which was adjusted by a posteriori SNR‐weighted factor, and the orthogonal matching pursuit approach was applied to reconstruct the clean speech spectrum from the noisy speech. The quasi‐clean speech was considered as the reference to a modified PESQ perceptual model, and the mean opinion score of the noisy degraded speech was achieved via the distortions estimation between the quasi‐clean speech and the degraded speech. Experimental results show that the proposed approach obtains a correlation coefficient of 0.925 on NOIZEUS complex environment database, which is 99% similar to the performance of the intrusive standard ITU‐T PESQ, and 7.1% outperforms non‐intrusive standard ITU‐T P.563. Weili Zhou, Qianhua He, Yalou Wang, Yanxiong Li |
IET Signal Process. | 4 |
| 2016 | Source cell phone matching from speech recordings by sparse representation and KISS metricabstractSource recording device matching from two speech recordings is a new and important problem of digital media forensics. It aims to answer the question that whether or not two speech recordings are recorded by the same recording device. In this study we propose a source cell phone matching scheme. The Gaussian supervector (GSV) based on Mel-frequency cepstral coefficients (MFCCs) is extracted from the speech recording and is sparse represented with respect to a dictionary learned by K-SVD algorithm. The reduced-dimensional sparse representation coefficient is utilized to characterize the intrinsic fingerprint of the recording device. Then, KISS metric learning based similarity matching is conducted on a pair of fingerprints extracted from the two speech recordings. Evaluation experiments were conducted on a database of speech recordings recorded by 14 cell phones. The experimental results demonstrated the feasibility of the proposed scheme. Ling Zou 0003, Qianhua He, Ji-Chen Yang, Yanxiong Li |
ICASSP | 4 |
| 2014 | Fast speaker clustering using distance of feature matrix mean and adaptive convergence thresholdabstractThe authors propose a method of fast speaker clustering in which a distance (distance of feature matrix mean, DFMM) is first defined for characterising the similarities between any two clusters, and then an adaptive convergence threshold is introduced for terminating the procedure of speaker clustering. If the minimum of the DFMMs between any two clusters is smaller than the threshold, then they are merged. The above mergence of clusters is repeated until the minimum of the DFMMs between any two clusters is larger than the threshold. They conduct experiments on both shorter voice segments (≤ 3 s) and longer voice segments (> 3 s) to compare their method with state‐of‐the‐art methods, agglomerative hierarchical clustering with Bayesian information criterion (AHC + BIC) and vector quantisation with spectral clustering. Experiments show that their method achieves the best results for clustering shorter voice segments, and also obtains satisfactory results for clustering longer voice segments in comparison with other two methods. What is more, their method is faster than other methods in all experimental cases. The initial results show that the hybrid methods by combining their method with the AHC + BIC obtain further improvement in terms of the F score. Yanxiong Li, Hai Jin 0008, Qianhua He, Xiaohui Feng |
IET Signal Process. | 1 |
| 2009 | Characteristics-based effective applause detection for meeting speech
Yanxiong Li, Qianhua He, Sam Kwong, Ji-Chen Yang |
Signal Process. | 1 |