Yan Song 0001

dblp:09/1398-1 · DBLP profile ↗
← Back
87ranked-venue papers
9as first author
29since 2021 · last 2025
0000-0002-5668-9068ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 72 · 9 first-author · 24 since 2021Artificial intelligence and machine learning · 37 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
abstract
Generating reports for computed tomography (CT) images is a challenging task. Although it is related to existing work on medical image report generation, it exhibits several unique characteristics, including the spatial encoding of multiple images and the alignment between image volumes and text. Existing solutions typically use general 2D or 3D image processing techniques to extract features from a CT volume, where they firstly compress the volume and then divide the compressed CT slices into patches for visual encoding. These approaches do not explicitly account for the transformations among CT slices, nor do they effectively integrate multi-level image features, particularly those containing specific organ lesions, to instruct CT report generation (CTRG). In considering the strong correlation among consecutive slices in CT scans, in this paper, we propose a large language model (LLM) based CTRG method with recurrent visual feature extraction and stereo attentions for hierarchical feature modeling. Specifically, we use a vision Transformer to recurrently process each slice in a CT volume, and employ a set of attentions over the encoded slices from different perspectives to selectively obtain important visual information and align it with textual features, so as to better instruct an LLM for CTRG. Experiment results and further analysis on the benchmark M3DCap dataset show that our method outperforms strong baseline models and achieves state-of-the-art results, demonstrating its validity and effectiveness11Code is available at https://github.com/synlp/SA-CTRG
Yuanhe Tian, Yan Song 0001
BIBM3
2025 PNP-RKD: A Positive-Negative Pair based Relational Knowledge Distillation Method for Cross-Domain Speaker Verification
abstract
Existing deep embedding learning based speaker verification (SV) methods suffer from performance degradation under domain shift conditions. This can be alleviated through unsupervised domain adaptation (UDA) techniques. While UDA improves global statistical consistency across domains, discriminative information may be overlooked or misaligned in the process. To combat this, we propose PNP-RKD, a relational knowledge distillation method that utilizes positive and negative pairs from both the source and target domains within a multitask learning framework. Two auxiliary tasks are conducted separately in the source and target domains to support PNP-RKD. Embeddings are learned in a supervised fashion from the labeled source domain, providing a robust foundation of prior knowledge. For the unlabeled target domain, we apply contrastive learning based on swapped prediction, a key component that enhances noise robustness and improves the quality of learned prototypes. More importantly, it facilitates reliable sampling in PNP-RKD, thereby enhancing the alignment of discriminative knowledge across domains. Extensive experiments conducted on the NIST SRE16 and SRE18 datasets demonstrate the superior performance of the proposed PNP-RKD method, achieving EERs of 6.83% and 8.28%, respectively.
Qing Gu 0002, Yan Song 0001, Nan Jiang 0022, Pengfei Cai, Ian McLoughlin 0001
ICASSP2
2025 Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
abstract
A significant challenge in sound event detection (SED) is the effective utilization of unlabeled data, given the limited availability of labeled data due to high annotation costs. Semi-supervised algorithms rely on labeled data to learn from unlabeled data, and the performance is constrained by the quality and size of the former. In this paper, we introduce the Prototype based Masked Audio Model (PMAM) algorithm for self-supervised representation learning in SED, to better exploit unlabeled data. Specifically, semantically rich frame-level pseudo labels are constructed from a Gaussian mixture model (GMM) based prototypical distribution modeling. These pseudo labels supervise the leaning of a Transformer-based masked audio model, in which binary cross-entropy loss is employed instead of the widely used InfoNCE loss, to provide independent loss contributions from different prototypes, which is important in real scenarios in which multiple labels may apply to unsupervised data frames. A final stage of fine-tuning with just a small amount of labeled data yields a very high performing SED model. On like-for-like tests using the DESED task, our method achieves a PSDS1 score of 62.5%, surpassing current state-of-the-art models and demonstrating the superiority of the proposed technique.
Pengfei Cai, Yan Song 0001, Nan Jiang 0022, Qing Gu 0002, Ian McLoughlin 0001
ICASSP2
2025 An Effective Anomalous Sound Detection Method Based on Global and Local Attribute Mining
Nan Jiang 0022, Yan Song 0001, Qing Gu 0002, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH2
2025 Finetune Large Pre-Trained Model Based on Frequency-Wise Multi-Query Attention Pooling for Anomalous Sound Detection
Nan Jiang 0022, Yan Song 0001, Qing Gu 0002, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH2
2025 Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
abstract
Most existing sound event detection (SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED by exploiting audio-language models, their performance is still far from satisfactory due to the lack of fine-grained alignment and cross-modal feature fusion. In this work, we propose the Detect Any Sound Model (DASM), a query-based framework for open-vocabulary SED guided by multi-modal queries. DASM formulates SED as a frame-level retrieval task, where audio features are matched against query vectors derived from text or audio prompts. To support this formulation, DASM introduces a dual-stream decoder that explicitly decouples event recognition and temporal localization: a cross-modality event decoder performs query-feature fusion and determines the presence of sound events at the clip-level, while a context network models temporal dependencies for frame-level localization. Additionally, an inference-time attention masking strategy is proposed to leverage semantic relations between base and novel classes, substantially enhancing generalization to novel classes. Experiments on the AudioSet Strong dataset demonstrate that DASM effectively balances localization accuracy with generalization to novel classes, outperforming CLAP-based methods in open-vocabulary setting (+ 7.8 PSDS) and the baseline in the closed-set setting (+ 6.9 PSDS). Furthermore, in cross-dataset zero-shot evaluation on DESED, DASM achieves a PSDS1 score of 42.2, even exceeding the supervised CRNN baseline. The project page is available at https://cai525.github.io/Transformer4SED/demo_page/DASM/.
Pengfei Cai, Yan Song 0001, Qing Gu 0002, Nan Jiang 0022, Ian McLoughlin 0001
ACM Multimedia2
2024 Large Language Models Are No Longer Shallow Parsers
abstract
The development of large language models (LLMs) brings significant changes to the field of natural language processing (NLP), enabling remarkable performance in various high-level tasks, such as machine translation, questionanswering, dialogue generation, etc., under endto-end settings without requiring much training data.Meanwhile, fundamental NLP tasks, particularly syntactic parsing, are also essential for language study as well as evaluating the capability of LLMs for instruction understanding and usage.In this paper, we focus on analyzing and improving the capability of current state-of-the-art LLMs on a classic fundamental task, namely constituency parsing, which is the representative syntactic task in both linguistics and natural language processing.We observe that these LLMs are effective in shallow parsing but struggle with creating correct full parse trees.To improve the performance of LLMs on deep syntactic parsing, we propose a three-step approach that firstly prompts LLMs for chunking, then filters out low-quality chunks, and finally adds the remaining chunks to prompts to instruct LLMs for parsing, with later enhancement by chain-of-thought prompting.Experimental results on English and Chinese benchmark datasets demonstrate the effectiveness of our approach on improving LLMs' performance on constituency parsing.
Yuanhe Tian, Fei Xia 0004, Yan Song 0001
ACL (1)3
2024 Dialogue Summarization with Mixture of Experts based on Large Language Models
abstract
Dialogue summarization is an important task that requires to generate highlights for a conversation from different aspects (e.g., content of various speakers).While several studies successfully employ large language models (LLMs) and achieve satisfying results, they are limited by using one model at a time or treat it as a black box, which makes it hard to discriminatively learn essential content in a dialogue from different aspects, therefore may lead to anticipation bias and potential loss of information in the produced summaries.In this paper, we propose an LLM-based approach with roleoriented routing and fusion generation to utilize mixture of experts (MoE) for dialogue summarization.Specifically, the role-oriented routing is an LLM-based module that selects appropriate experts to process different information; fusion generation is another LLM-based module to locate salient information and produce finalized dialogue summaries.The proposed approach offers an alternative solution to employing multiple LLMs for dialogue summarization by leveraging their capabilities of in-context processing and generation in an effective manner.We run experiments on widely used benchmark datasets for this task, where the results demonstrate the superiority of our approach in producing informative and accurate dialogue summarization. 1
Yuanhe Tian, Fei Xia 0004, Yan Song 0001
ACL (1)3
2024 Meta Representation Learning Method for Robust Speaker Verification in Unseen Domains
abstract
This paper presents a meta representation learning method for robust speaker verification (SV) in unseen domains. It is known that the existing embedding learning based SV systems may suffer from domain mismatch issues. To address this, we propose an episodic training procedure to compensate domain mismatch conditions at runtime. Specifically, episodes are constructed with domain balanced episodic sampling from two different domains, and a new domain alignment (DA) module is added besides the feature extractor (FE) and classifier to existing network structures. In each episodic training iteration, FE and DA modules are optimized separately with different objectives to improve the robustness of learning. Besides, a cross-domain inter-class alignment (CDICA) loss is proposed for improving the domain generalization ability. Experimental results on CNCeleb and VoxCeleb benchmarks demonstrate significant performance gains for unseen domains in SV.
Jian-Tao Zhang, Yan Song 0001, Wu Guo, Hao-Yu Song, Ian McLoughlin 0001
ICASSP2
2024 MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
Pengfei Cai, Yan Song 0001, Ian McLoughlin 0001
INTERSPEECH2
2024 An Effective Local Prototypical Mapping Network for Speech Emotion Recognition
Yuxuan Xi, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH2
2024 DB-PMAE: Dual-Branch Prototypical Masked AutoEncoder with locality for domain robust speaker verification
Wei-Lin Xie, Yu-Xuan Xi, Yan Song 0001, Jian-Tao Zhang, Hao-Yu Song, Ian McLoughlin 0001
INTERSPEECH3
2024 Improving radiology report generation with multi-grained abnormality prediction
Yuda Jin, Yuanhe Tian, Yan Song 0001
Neurocomputing4
2023 An Effective Anomalous Sound Detection Method Based on Representation Learning with Simulated Anomalies
abstract
In this paper, we propose an effective anomalous sound detection (ASD) method based on representation learning with simulated anomalies. Recently, ASD systems have used Outlier Exposure (OE) strategy to achieve promising performance in DCASE challenges. These exploit deep Convolutional Neural Networks (CNNs) to learn discriminative representations by treating the normal samples from different classes as pseudo anomalies. However, since the anomalous sounds occur only rarely, are diversely distributed and are unseen during training, the OE capability of representations learned from normal samples may be limited. To address this issue, we propose a statistics exchange (StEx) method, which constructs simulated anomalies to improve the effectiveness of representation learning via OE strategy. Specifically, the first and second order statistics are extracted from time or frequency axis of input spectrograms, and the simulated anomaly is then obtained by exchanging the statistics of spectrograms from different classes. Furthermore, an out-of-distribution (OOD) metric is introduced as an importance measure to qualitatively analyze OE capability, which enables appropriate simulated anomaly selection for ASD. Extensive experiments on the DCASE2021 challenge task2 development dataset verify the effectiveness of representation learning with simulated anomaly for OE based ASD.
Yan Song 0001, Zhu Zhuo, Yu-Hong Li, Ian McLoughlin 0001
ICASSP2
2023 Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification
abstract
Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors, such as recording device and speaking style, in real-world applications, which leads to severe performance degradation. Since single-speaker multi-condition (SSMC) data is difficult to collect in practice, existing domain adaptation methods are hard to ensure the feature consistency of the same class but different domains. To this end, we propose a cross-domain data generation method to obtain a domain-invariant ASV system. Inspired by voice conversion (VC) task, a StarGAN based generative model first learns cross-domain mappings from SSMC data, and then generates missing domain data for all speakers, thus increasing the intra-class diversity of the training set. Considering the difference between ASV and VC task, we renovate the corresponding training objectives and network structure to make the adaptation task-specific. Evaluations on achieve a relative performance improvement of about 5-8% over the baseline in terms of minDCF and EER, outperforming the CNSRC winner’s system of the equivalent scale.
Hang-Rui Hu, Yan Song 0001, Jian-Tao Zhang, Li-Rong Dai 0001, Ian McLoughlin 0001, Zhu Zhuo, Yu-Hong Li
ICASSP2
2023 AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer
abstract
In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio tagging (AT) task, termed AST-SED. Pretrained AST models have recently shown promise on DCASE2022 challenge task4 where they help mitigate a lack of sufficient real annotated data. However, mainly due to differences between the AT and SED tasks, it is suboptimal to directly utilize outputs from a pretrained AST model. Hence the proposed AST-SED adopts an encoder-decoder architecture to enable effective and efficient fine-tuning without needing to redesign or retrain the AST model. Specifically, the Frequency-wise Transformer Encoder (FTE) consists of transformers with self attention along the frequency axis to address multiple overlapped audio events issue in a single clip. The Local Gated Recurrent Units Decoder (LGD) consists of nearest-neighbor interpolation (NNI) and Bidirectional Gated Recurrent Units (Bi-GRU) to compensate for temporal resolution loss in the pretrained AST model output. Experimental results on DCASE2022 task4 development set have demonstrated the superiority of the proposed AST-SED with FTE-LGD architecture. Specifically, the Event-Based F1-score (EB-F1) of 59.60% and Polyphonic Sound detection Score scenario1 (PSDS1) of 0.5140 significantly outperform CRNN and other pretrained AST-based systems.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP2
2023 Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection
abstract
In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to perform frame-level prediction. The output of the PAE together with original normal samples, are used for supervised contrastive representative learning in a multi-task framework. Besides cross-entropy loss between classes, contrastive loss is used to separate PAE output and original samples within each class. GeCo aims to better capture context information among frames, thanks to the self-attention mechanism for PAE model. Furthermore, GeCo combines generative and contrastive learning from which we aim to yield more effective and informative representations, compared to existing methods. Extensive experiments have been conducted on the DCASE2020 Task2 development dataset, showing that GeCo outperforms state-of-the-art generative and discriminative methods.
Xiao-Min Zeng, Yan Song 0001, Zhu Zhuo, Yu-Hong Li, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP2
2023 Fine-tuning Audio Spectrogram Transformer with Task-aware Adapters for Sound Event Detection
abstract
In this paper, we present a task-aware fine-tuning method to transfer Patchout faSt Spectrogram Transformer (PaSST) model to sound event detection (SED) task. Pretrained PaSST has shown significant performance on audio tagging (AT) and SED tasks, but it is not optimal to fine-tune the model from a single layer as the local and semantic information have not been well exploited. To address this, we first introduce task-aware adapters including SED-adapter and AT-adapter to fine-tune PaSST for SED and AT task respectively, and then propose task-aware fine-tuning to combine local information from shallower layer with semantic information from deeper layer, based on task-aware adapters. Besides, we propose the self-distillated mean teacher (SdMT) to train a robust student model with soft pseudo labels from teacher. Experiments are conducted on DCASE2022 task4 development set, the EB-F1 of 64.85% and PSDS1 of 0.5548 are achieved which outperform previous state-of-the-art systems.
Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
INTERSPEECH2
2023 Robust Prototype Learning for Anomalous Sound Detection
abstract
In this paper, we present a robust prototype learning framework for anomalous sound detection (ASD), where prototypical loss is exploited to measure the similarity between samples and prototypes. We show that existing generative and discriminative based ASD methods can be unified into this framework from the perspective of prototypical learning. For ASD in recent DCASE challenges, extensions related to imbalanced learning are proposed to improve the robustness of prototypes learned from source and target domains. Specifically, balanced sampling and multiple-prototype expansion (MPE) strategies are proposed to address imbalances across attributes of source and target domains. Furthermore, a novel negative-prototype expansion (NPE) method is used to construct pseudo-anomalies to learn a more compact and effective embedding space for normal sounds. Evaluation on the DCASE2022 Task2 development dataset demonstrates the validity of the proposed prototype learning framework.
Xiao-Min Zeng, Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
INTERSPEECH2
2023 Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023
abstract
In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three different structures based on attention-guided feature gathering (AFG) are designed for deep feature fusion. Then, we introduce a joint decoding structure for emotion classification and valence regression in the decoding stage. A multi-task loss based on uncertainty is also designed to optimize the whole process. Finally, by combining three different structures on the posterior probability level, we obtain the final predictions of discrete and dimensional emotions. When tested on the dataset of multimodal emotion recognition challenge (MER 2023), the proposed framework yields consistent improvements in both emotion classification and valence regression. Our final system achieves state-of-the-art performance and ranks third on the leaderboard on MER-MULTI sub-challenge.
Yuxuan Xi, Hang Chen 0001, Jun Du 0002, Yan Song 0001, Qing Wang 0008, Hengshun Zhou, Jiefeng Ma, Pengfei Hu 0006, Ya Jiang, Shi Cheng 0001, Jie Zhang 0042, Yuzhe Weng
ACM Multimedia5
2022 Self-Supervised Representation Learning for Unsupervised Anomalous Sound Detection Under Domain Shift
abstract
In this paper, a self-supervised representation learning method is proposed for anomalous sound detection (ASD). ASD has received much research attention in recent DCASE challenges. It aims to identify whether a sound emitted from a machine is anomalous or not, given only normal sound data. This is a challenging task due to highly variable time-frequency characteristics of sounds from different machine types, and the fact that many attributes affect machine without being anomalous. This is especially true for domain shift tasks, where only a few training sound clips are available. From the perspective of self-supervised learning, each given sound clip can be considered as a transformation of an original clean sound, where the attribute of each clip may indicate different supervision signals. We propose a unified representation learning framework, equipped with a time-frequency attention mechanism, to perform ASD for different machine types and attributes. For domain shift, a centre imprinting method, which directly sets centres for target domain attributes, is presented. This provides immediate good representation and an initialization for further fine-tuning. Evaluation on DCASE2021 ASD task demonstrates the effectiveness of the proposed method.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP2
2022 Domain Robust Deep Embedding Learning for Speaker Recognition
abstract
This paper presents a domain robust deep embedding learning method for speaker verification (SV) tasks. Most recent methods utilize deep neural networks (DNN) to learn compact and discriminative speaker embeddings from large-scale labeled datasets such as VoxCeleb and the NIST SRE corpus. Despite the success of exiting methods, performance may degrade significantly for new target datasets, mainly due to the distribution discrepancy between training and test domains. Moreover, how corpora are collected, and the languages they contain differ, leading to them spanning multiple, perhaps mismatched, latent domains. To address this, a multi-task end-to-end framework is proposed to learn speaker embeddings from both labeled source and unlabeled target datasets. Motivated by label smoothing, a smoothed knowledge distillation (SKD) based self-supervised learning method is designed to exploit latent structural information from the unlabeled target domain. Furthermore, a domain-aware batch normalization (DABN) module aims to reduce the cross-domain distribution discrepancy, while a domain-agnostic instance normalization (DAIN) module aims to learn features that are robust to within-domain variance. Evaluation on NIST SRE16 demonstrates significant performance gains.
Hang-Rui Hu, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP2
2022 Frontend Attributes Disentanglement for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) with limited size dataset is a challenging task, since a spoken utterance contains various disturbing attributes besides emotion, including speaker, content, and language. However, due to a close relationship between speaker and emotion attributes, simply fine-tuning a linear model is enough to obtain a good SER performance on the utterance-level embeddings (i.e., i-vector and x-vectors) extracted from the pre-trained speaker recognition (SR) frontends. In this paper, we aim to perform frontend attributes disentanglement (AD) for SER task, using a pre-trained SR model. Specifically, the AD module consists of attribute normalization (AN) and attribute reconstruction (AR) phases. The AN filters out the variation information using instance normalization (IN), and AR reconstructs the emotion-relevant features from the residual space to ensure high emotion discrimination. For better disentanglement, a dual space loss is then designed to encourage the separability of emotion-relevant and emotion-irrelevant spaces. To introduce the long-range contextual information for emotion related reconstruction, a time-frequency (TF) attention is further proposed. Different from the style disentanglement of the extracted x-vectors, the proposed AD module can be applied on frontend feature extractor. Experiments on IEMOCAP benchmark demonstrate the effectiveness of the proposed method.
Yuxuan Xi, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
ICASSP2
2022 Class-Aware Distribution Alignment based Unsupervised Domain Adaptation for Speaker Verification
abstract
Existing speaker verification (SV) systems usually suffer from significant performance degradation when applied to a new domain that lies outside the training distribution. Given the unlabeled target-domain dataset, most Unsupervised Domain Adaptation (UDA) methods aim to minimize the distribution divergence between different domains. However, global distribution alignment strategies fail to consider the latent speaker label information and can hardly guarantee the feature discriminative capability in target domain. In this paper, we propose a novel UDA approach called WBDA (Within-class and Between-class Distribution Alignment), which aims to transfer the class-aware information (i.e., within- and between-class distributions) learned from the well-labeled source-domain to unlabeled target-domain. Motivated by the recent progress of self-supervised contrastive learning, the positive and negative pairs are constructed separately for source and target domains, from which the within- and between-class distribution can be estimated. And the SV system can then be learned by jointly optimizing the cross-domain class-aware distribution discrepancy loss and source-domain classification loss in an end-to-end manner. Evaluations on NIST SRE16 and SRE18 achieve a relative performance improvement of about 43.7% and 26.2% over the baseline in terms of Equal Error Rate (EER) separately, significantly outperforming the previous adaption methods based on global distribution alignment.
Hang-Rui Hu, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
INTERSPEECH2
2021 An Effective Deep Embedding Learning Method Based on Dense-Residual Networks for Speaker Verification
abstract
In this paper, we present an effective end-to-end deep embedding learning method based on Dense-Residual networks, which combine the advantages of a densely connected convolutional network (DenseNet) and a residual network (ResNet), for speaker verification (SV). Unlike a model ensemble strategy which merges the results of multiple systems, the proposed Dense-Residual networks perform feature fusion on every basic DenseR building block. Specifically, two types of DenseR blocks are designed. A sequential-DenseR block is constructed by densely connecting stacked basic units in a residual block of ResNet. A parallel-DenseR comprises split and concatenation operations on residual and dense components via corresponding skip connections. These building blocks are stacked into deep networks to exploit the complementary information with different receptive field sizes and growth rates. Extensive experiments have been conducted on the VoxCeleb1 dataset to evaluate the proposed methods. The SV performance achieved by the proposed Dense-Residual networks is shown to outperform corresponding ResNet, DenseNet or fusions of them, with similar model complexity, by a significant margin.
Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
ICASSP2
2021 An Improved Mean Teacher Based Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection
abstract
This paper presents an improved mean teacher (MT) based method for large-scale weakly labeled semi-supervised sound event detection (SED), by focusing on learning a better student model. Two main improvements are proposed based on the authors’ previous perturbation based MT method. Firstly, an event-aware module is de-signed to allow multiple branches with different kernel sizes to be fused via an attention mechanism. By inserting this module after the convolutional layer, each neuron can adaptively adjust its receptive field to suit different sound events. Secondly, instead of using the teacher model to provide a consistency cost term, we propose using a stochastic inference of unlabeled examples to generate high quality pseudo-targets by averaging multiple predictions from the perturbed student model. MixUp of both labeled and unlabeled data is further exploited to improve the effectiveness of student model. Finally, the teacher model can be obtained via exponential moving average (EMA) of the student model, which generates final predictions for SED during inference. Experiments on the DCASE2018 task4 dataset demonstrate the ability of the proposed method. Specifically, an F1-score of 42.1% is achieved, significantly outperforming the 32.4% achieved by the winning system, or the 39.3% by the previous perturbation based method.
Yan Song 0001, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
ICASSP2
2021 A Weight Moving Average Based Alternate Decoupled Learning Algorithm for Long-Tailed Language Identification
abstract
Language identification (LID) research has made tremendous progress in recent years, especially with the introduction of deep learning techniques. However, for real-world applications where the distribution of different language data is highly imbalanced, the performance of existing LID systems is still far from satisfactory. This raises the challenge of long-tailed LID. In this paper, we propose an effective weight moving average (WMA) based alternate decoupled learning algorithm, termed WADCL, for long-tailed LID. The system is divided into two components, a frontend feature extractor and a backend classifier. These are then alternately learned in an end-to-end manner using different sampling schemes to alleviate the distribution mismatch between training and test datasets. Furthermore, our WMA method aims to mitigate the side-effects of re-sampling schemes, by fusing the model parameters learned along the trajectory of stochastic gradient descent (SGD) optimization. To validate the effectiveness of the proposed WADCL algorithm, we evaluate and compare several systems over a language dataset constructed to match a long-tailed distribution based on real world application [1]. The experimental results from the long-tailed language dataset demonstrate that the proposed algorithm is able to achieve significant performance gains over existing state-of-the-art x-vector based LID methods.
Lin Liu 0017, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
Interspeech3
2021 An Effective Mutual Mean Teaching Based Domain Adaptation Method for Sound Event Detection
abstract
In this paper, we present a novel mutual mean teaching based domain adaptation (MMT-DA) method for sound event detection (SED) task, which can effectively exploit synthetic data to improve the SED performance. Existing methods simply treat the synthetic data as strongly-labeled data in semi-supervised learning (SSL) framework. Benefiting from the strong labels of synthetic data, superior SED performance can be achieved. However, a distribution mismatch between synthetic and real data raises an evident challenge for domain adaptation (DA). In MMT-DA, convolutional recurrent neural networks (CRNN) learned from different datasets (i.e. total data:real+synthetic, and real data) are exploited for DA. Specifically, mean teacher method using CRNN is employed for utilizing the unlabeled real data. To compensate the domain diversity, an additional domain classifier with gradient reverse layer(GRL) is used for training a mean teacher for total data. The student CRNNs are mutually taught using the soft predictions of unlabeled data obtained from different teachers. Furthermore, a strip pooling based attention module is exploited to model the inter-dependencies between channels and time-frequency dimensions to exploit the structure information. Experimental results on Task4 of DCASE2020 demonstrate the ability of the proposed method, achieving 52.0% F1-score on the validation dataset, which outperforms the winning system’s 50.6%.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
Interspeech2
2021 Multi-Granularity Sequence Alignment Mapping for Encoder-Decoder Based End-to-End ASR
abstract
Encoder-decoder based automatic speech recognition (ASR) methods are increasingly popular due to their simplified processing stages and low reliance on prior knowledge. Conventional encoder-decoder based approaches usually learn a sequence-to-sequence mapping function from the source speech to target units (e.g., subwords, characters) in an end-to-end manner. However, it is still unclear how to choose the optimal target unit, or granularity of multiple units. In general, as increasing the information available for learning sequence-to-sequence mapping functions can improve modeling effectiveness, we therefore propose a multi-granularity sequence alignment (MGSA) approach. This aims to enhance cross-sequence interactions between different granularity units for both modeling and inference stages in the encoder-decoder based ASR. Specifically, a decoder module is designed to generate multi-granularity sequence predictions. We then exploit the latent alignment mapping among units having different levels of granularity, by utilizing the decoded multi-level sequences as input for model prediction. The cross-sequence interaction can also be employed to re-calibrate output probabilities in the proposed post-inference algorithm. Experimental results on both WSJ-80 hrs and Switchboard-300 hrs datasets show the superiority of the proposed method compared to traditional multi-task methods as well as to single granularity baseline systems.
Jie Zhang 0042, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 An Online Speaker-aware Speech Separation Approach Based on Time-domain Representation
abstract
Despite the significant progress of deep learning based speech separation methods, it remains challenging to extract and track the speech from target speakers, especially in a single-channel multiple speaker situation. Previously, the authors proposed a source-aware context network to exploit the temporal context in mixtures and estimated sources for online speech separation. In this paper, we propose a speaker-aware approach based on the source-aware context network structure, in which the speaker information is explicitly modeled by an auxiliary speaker identification branch. Then speech separation and speaker tracking can be jointly optimized by multi-task learning. Furthermore, we study the effectiveness of time-domain representation by proposing a raw sparse waveform encoder to preserve discriminative information. Experimental results on the WSJ0-2mix benchmark show that the proposed system significantly improves Signal-to-Distortion Ratio (SDR) performance.
Yan Song 0001, Zengxi Li, Ian McLoughlin 0001, Li-Rong Dai 0001
ICASSP2
2020 Task-Aware Mean Teacher Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection
abstract
Weakly labeled semi-supervised learning methods have recently drawn increasing attention from the research community for sound event detection tasks. Due to the weakness of the labelling, neural networks are often designed to perform sound event detection (SED) and audio tagging (AT) at the same time. In this paper, we propose a task-aware mean teacher method using a convolutional recurrent neural network (CRNN) with multi-branch structure to solve the SED and AT tasks differently. Specifically, a branch with coarse-level temporal resolution is designed for the AT task, while a branch with fine-level temporal resolution is designed for the SED task. The mean teacher based semi-supervised learning method is first adopted to improve the performance of the coarse-level AT branch by exploiting unlabeled data. Then the coarse-level AT branch is introduced as a teacher to guide the aggregated AT output of the fine-level SED branch, yielding an improvement in the SED performance. To further improve the AT and SED performance, information from multiple layers is exploited in the form of a multi-resolution feature. Experimental results on Task4 of the DCASE2018 challenge demonstrate the superiority of the proposed method, achieving 37.7% F1-score, which outperforms the winning system's 32.4%.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP2
2020 An Effective Speaker Recognition Method Based on Joint Identification and Verification Supervisions
abstract
Deep embedding learning based speaker verification methods have attracted significant recent research interest due to their superior performance. Existing methods mainly focus on designing frame-level feature extraction structures, utterance-level aggregation methods and loss functions to learn discriminative speaker embeddings. The scores of verification trials are then computed using cosine distance or Probabilistic Linear Discriminative Analysis (PLDA) classifiers. This paper proposes an effective speaker recognition method which is based on joint identification and verification supervisions, inspired by multi-task learning frameworks. Specifically, a deep architecture with convolutional feature extractor, attentive pooling and two classifier branches is presented. The first, an identification branch, is trained with additive margin softmax loss (AM-Softmax) to classify the speaker identities. The second, a verification branch, trains a discriminator with binary cross entropy loss (BCE) to optimize a new triplet-based mutual information. To balance the two losses during different training stages, a ramp-up/ramp-down weighting scheme is employed. Furthermore, an attentive bilinear pooling method is proposed to improve the effectiveness of embeddings. Extensive experiments have been conducted on VoxCeleb1 to evaluate the proposed method, demonstrating results that relatively reduce the equal error rate (EER) by 22% compared to the baseline system using identification supervision only.
Yan Song 0001, Yiheng Jiang, Ian McLoughlin 0001, Lin Liu 0017, Li-Rong Dai 0001
INTERSPEECH2
2020 Semi-Supervised End-to-End ASR via Teacher-Student Learning with Conditional Posterior Distribution
abstract
Encoder-decoder based methods have become popular for automatic speech recognition (ASR), thanks to their simplified processing stages and low reliance on prior knowledge. However, large amounts of acoustic data with paired transcriptions is generally required to train an effective encoder-decoder model, which is expensive, time-consuming to be collected and not always readily available. However unpaired speech data is abundant, hence several semi-supervised learning methods, such as teacher-student (T/S) learning and pseudo-labeling, have recently been proposed to utilize this potentially valuable resource. In this paper, a novel T/S learning with conditional posterior distribution for encoder-decoder based ASR is proposed. Specifically, the 1-best hypotheses and the conditional posterior distribution from the teacher are exploited to provide more effective supervision. Combined with model perturbation techniques, the proposed method reduces WER by 19.2% relatively on the LibriSpeech benchmark, compared with a system trained using only paired data. This outperforms previous reported 1-best hypothesis results on the same task.
Yan Song 0001, Jianshu Zhang 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
INTERSPEECH2
2020 An Effective Perturbation Based Semi-Supervised Learning Method for Sound Event Detection
abstract
Mean teacher based methods are increasingly achieving state-of-the-art performance for large-scale weakly labeled and unlabeled sound event detection (SED) tasks in recent DCASE challenges. By penalizing inconsistent predictions under different perturbations, mean teacher methods can exploit large-scale unlabeled data in a self-ensembling manner. In this paper, an effective perturbation based semi-supervised learning (SSL) method is proposed based on the mean teacher method. Specifically, a new independent component (IC) module is proposed to introduce perturbations for different convolutional layers, designed as a combination of batch normalization and dropblock operations. The proposed IC module can reduce correlation between neurons to improve performance. A global statistics pooling based attention module is further proposed to explicitly model inter-dependencies between the time-frequency domain and channels, using statistics information (e.g. mean, standard deviation, max) along different dimensions. This can provide an effective attention mechanism to adaptively re-calibrate the output feature map. Experimental results on Task 4 of the DCASE2018 challenge demonstrate the superiority of the proposed method, achieving about 39.8% F1-score, outperforming the previous winning system’s 32.4% by a significant margin.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001, Lin Liu 0017
INTERSPEECH2
2019 Topic Detection in Conversational Telephone Speech Using CNN with Multi-stream Inputs
abstract
Topic detection for conversational telephone speech (CTS) is addressed in this paper. The low accuracy of automatic speech recognition (ASR) will cause severe performance deterioration for topic detection. To make up for this, we adopt two ASR systems, HMM-BiLSTM and CTC systems, to provide complementary information for topic detection. After obtaining two sets of different recognized transcriptions, a CNN with multi-stream inputs is trained, and the pooling layer serves as document representations. Finally, element-wise summation of document representations from two streams is used as distributed representations of the documents, which are fed into agglomerative hierarchical clustering (AHC) algorithms to obtain clustering results. The experiments on a Japanese speech corpus demonstrate that the proposed approach can significantly improve the performance of topic detection.
Wu Guo, Yan Song 0001
ICASSP4
2019 A Region Based Attention Method for Weakly Supervised Sound Event Detection and Classification
abstract
Recently, an attention based convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) has achieved state-of-the-art performance for audio tagging (AT) and sound event detection (SED) tasks in the Detection and Classification of Acoustic Scenes and Events (DCASE) challenges. The introduction of GLU and temporal attention-based localization mechanisms plays an important role for both AT and SED tasks. In this paper, we propose a novel region based attention method to further boost the representation power of the existing GLU based CRNN. Specifically, we insert a feature selection (FS) structure after each GLU to create what we term a GLU-F. block, to exploit channel relationships. Furthermore, we extract region features (or the prototypes of certain sound events) from multi-scale sliding windows over higher convolutional layers, which are fed into an attention-based recurrent neural network to model their context information for AT and SED tasks. To evaluate the proposed region based attention method, we conduct extensive experiments on SED and AT tasks in DCASE2017. We achieve 59.5% and 60.1% AT F1-score, 51.3% and 55.1% SED F1-score for development and evaluation sets respectively, significantly outperforming state-of-the-art results.
Yan Song 0001, Wu Guo, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP2
2019 Improving Aggregation and Loss Function for Better Embedding Learning in End-to-End Speaker Verification System
Zhifu Gao, Yan Song 0001, Ian McLoughlin 0001, Yiheng Jiang, Li-Rong Dai 0001
INTERSPEECH2
2019 An Effective Deep Embedding Learning Architecture for Speaker Verification
Yiheng Jiang, Yan Song 0001, Ian McLoughlin 0001, Zhifu Gao, Li-Rong Dai 0001
INTERSPEECH2
2019 Listening and Grouping: An Online Autoregressive Approach for Monaural Speech Separation
abstract
This paper proposes an autoregressive approach to harness the power of deep learning for multi-speaker monaural speech separation. It exploits a causal temporal context in both mixture and past estimated separated signals and performs online separation that is compatible with real-time applications. The approach adopts a learned listening and grouping architecture motivated by computational auditory scene analysis, with a grouping stage that effectively addresses the label permutation problem at both frame and segment levels. Experimental results on the WSJ0-2mix benchmark show that the new approach can achieve better signal-to-distortion ratio and perceptual evaluation of speech quality scores than most of the state-of-the-art methods for both closed-set and open-set evaluations, even methods that exploit whole-utterance statistics for separation. It achieves this while requiring fewer model parameters.
Zengxi Li, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Source-Aware Context Network for Single-Channel Multi-Speaker Speech Separation
abstract
Deep learning based approaches have achieved promising performance in speaker-dependent single-channel multispeaker speech separation. However, partly due to the label permutation problem, they may encounter difficulties in speaker-independent conditions. Recent methods address this problem by some assignment operations. Different from them, we propose a novel source-aware context network, which explicitly inputs speech sources as well as mixture signal. By exploiting the temporal dependency and continuity of the same source signal, the permutation order of outputs can be easily determined without any additional post-processing. Furthermore, a Multi-time-step Prediction Training strategy is proposed to address the mismatch between training and inference stages. Experimental results on benchmark WSJ0-2mix dataset revealed that our network achieved comparable or better results than state-of-the-art methods in both closed-set and open-set conditions, in terms of Signal-to-Distortion Ratio (SDR) improvement.
Zengxi Li, Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP2
2018 An Improved Deep Embedding Learning Method for Short Duration Speaker Verification
abstract
This paper presents an improved deep embedding learning method based on convolutional neural networks (CNN) for short-duration speaker verification (SV). Existing deep learning-based SV methods generally extract frontend embeddings from a feed-forward deep neural network, in which the long-term speaker characteristics are captured via a pooling operation over the input speech. The extracted embeddings are then scored via a backend model, such as Probabilistic Linear Discriminative Analysis (PLDA). Two improvements are proposed for frontend embedding learning based on the CNN structure: (1) Motivated by the WaveNet for speech synthesis, dilated filters are designed to achieve a tradeoff between computational efficiency and receptive-filter size; and (2) A novel cross-convolutional-layer pooling method is exploited to capture $1^{st}$-order statistics for modelling long-term speaker characteristics. Specifically, the activations of one convolutional layer are aggregated with the guidance of the feature maps from the successive layer. To evaluate the effectiveness of our proposed methods, extensive experiments are conducted on the modified female portion of NIST SRE 2010 evaluations, with conditions ranging from 10s-10s to 5s-4s. Excellent performance has been achieved on each evaluation condition, significantly outperforming existing SV systems using i-vector and d-vector embeddings.
Zhifu Gao, Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH2
2018 An Attention Pooling Based Representation Learning Method for Speech Emotion Recognition
abstract
This paper proposes an attention pooling based representation learning method for speech emotion recognition (SER). The emotional representation is learned in an end-to-end fashion by applying a deep convolutional neural network (CNN) directly to spectrograms extracted from speech utterances. Motivated by the success of GoogleNet, two groups of filters with different shapes are designed to capture both temporal and frequency domain context information from the input spectrogram. The learned features are concatenated and fed into the subsequent convolutional layers. To learn the final emotional representation, a novel attention pooling method is further proposed. Compared with the existing pooling methods, such as max-pooling and average-pooling, the proposed attention pooling can effectively incorporate class-agnostic bottom-up, and class-specific top-down, attention maps. We conduct extensive evaluations on benchmark IEMOCAP data to assess the effectiveness of the proposed representation. Results demonstrate a recognition performance of 71.8% weighted accuracy (WA) and 68% unweighted accuracy (UA) over four emotions, which outperforms the state-of-the-art method by about 3% absolute for WA and 4% for UA.
Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH2
2018 Early Detection of Continuous and Partial Audio Events Using CNN
abstract
Sound event detection is an extension of the static auditory classification task into continuous environments, where performance depends jointly upon the detection of overlapping events and their correct classification. Several approaches have been published to date which either develop novel classifiers or employ well-trained static classifiers with a detection front-end. This paper takes the latter approach, by combining a proven CNN classifier acting on spectrogram image features, with time-frequency shaped energy detection that identifies seed regions within the spectrogram that are characteristic of auditory energy events. Furthermore, the shape detector is optimised to allow early detection of events as they are developing. Since some sound events naturally have longer durations than others, waiting until completion of entire events before classification may not be practical in a deployed system. The early detection capability of the system is thus evaluated for the classification of partial events. Performance for continuous event detection is shown to be good, with accuracy being maintained well when detecting partial events. © 2018 International Speech Communication Association. All rights reserved.
Ian McLoughlin 0001, Yan Song 0001, Lam Dang Pham, Ramaswamy Palaniappan, Huy Phan, Yue Lang
INTERSPEECH2
2018 Acoustic Modeling with Densely Connected Residual Network for Multichannel Speech Recognition
abstract
Motivated by recent advances in computer vision research, this paper proposes a novel acoustic model called Densely Connected Residual Network (DenseRNet) for multichannel speech recognition. This combines the strength of both DenseNet and ResNet. It adopts the basic "building blocks" of ResNet with different convolutional layers, receptive field sizes and growth rates as basic components that are densely connected to form so-called denseR blocks. By concatenating the feature maps of all preceding layers as inputs, DenseRNet can not only strengthen gradient back-propagation for the vanishing-gradient problem, but also exploit multi-resolution feature maps. Preliminary experimental results on CHiME-3 have shown that DenseRNet achieves a word error rate (WER) of 7.58% on beamforming-enhanced speech with six channel real test data by cross entropy criteria training while WER is 10.23% for the official baseline. Besides, additional experimental results are also presented to demonstrate that DenseRNet exhibits the robustness to beamforming-enhanced speech as well as near and far-field speech.
Yan Song 0001, Li-Rong Dai 0001, Ian McLoughlin 0001
INTERSPEECH2
2018 Improved Supervised Locality Preserving Projection for I-vector Based Speaker Verification
Lanhua You, Wu Guo, Yan Song 0001
INTERSPEECH3
2018 LID-Senones and Their Statistics for Language Identification
abstract
Recent research on end-to-end training structures for language identification has raised the possibility that intermediate language-sensitive feature units exist which are analogous to phonetically sensitive senones in automatic speech recognition systems. Termed language identification (LID)-senones, the statistics derived from these feature units have been shown to be beneficial in discriminating between languages, particularly for short utterances. This paper examines the evidence for the existence of LID-senones before designing and evaluating LID systems based on low- and high-level statistics of LID-senones with both generative and discriminative models. For the standard NIST LRE 2009 task on 23 languages, LID-senone-based systems are shown to outperform state-of-the-art deep neural network/i-vector methods both when LID-senones are used directly for classification and when LID-senone statistics are used for i-vector formation.
Ma Jin, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Fisher vector based CNN architecture for image classification
abstract
In this paper, we tackle the representation learning problem for small scale fine-grained object recognition and scene classification tasks. Conventional bag of features(BoF) methods exploit hand-crafted frontend local features, and learn the representations via various machine learning techniques. Convolutional neural networks(CNN) directly learn the representation from raw images and benefit from joint optimization of network parameters in an end-to-end manner. However, the performance of existing representation learning methods is still unsatisfactory for the small-scale recognition tasks. To address this issue, we present a FV coding based CNN(FV-CNN) architecture. FV-CNN has three main advantages in that firstly it is able to exploit activations from the intermediate convolutional layer and a probabilistic discriminative model to derive the FV coding. Secondly, it takes advantage of the end-to-end back-propagation of the gradients to jointly optimize the whole learning process. Finally, it can learn a compact representation. When evaluated on benchmark datasets of fine grain object recognition (Caltech-CUB200), and scene classification (MIT67), accuracies of 88.0% and 82.2% are achieved.
Yan Song 0001, Peiseng Wang, Xinhai Hong, Ian McLoughlin 0001
ICIP1
2017 End-to-End Language Identification Using High-Order Utterance Representation with Bilinear Pooling
abstract
A key problem in spoken language identification (LID) is how to design effective representations which are specific to language information. Recent advances in deep neural networks have led to significant improvements in results, with deep end-to-end methods proving effective. This paper proposes a novel network which aims to model an effective representation for high (first and second)-order statistics of LID-senones, defined as being LID analogues of senones in speech recognition. The high-order information extracted through bilinear pooling is robust to speakers, channels and background noise. Evaluation with NIST LRE 2009 shows improved performance compared to current state-of-the-art DBF/i-vector systems, achieving over 33% and 20% relative equal error rate (EER) improvement for 3s and 10s utterances and over 40% relative Cavg improvement for all durations.
Ma Jin, Yan Song 0001, Ian McLoughlin 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH2
2016 Compact convolutional neural network transfer learning for small-scale image classification
abstract
Transfer learning methods have demonstrated state-of-the-art performance on various small-scale image classification tasks. This is generally achieved by exploiting the information from an ImageNet convolution neural network (ImageNet CNN). However, the transferred CNN model is generally with high computational complexity and storage requirement. It raises the issue for real-world applications, especially for some portable devices like phones and tablets without high-performance GPUs. Several approximation methods have been proposed to reduce the complexity by reconstructing the linear or non-linear filters (responses) in convolutional layers with a series of small ones., In this paper, we present a compact CNN transfer learning method for small-scale image classification. Specifically, it can be decomposed into fine-tuning and joint learning stages. In fine-tuning stage, a high-performance target CNN is trained by transferring information from the ImageNet CNN. In joint learning stage, a compact target CNN is optimized based on ground-truth labels, jointly with the predictions of the high-performance target CNN. The experimental results on CIFAR-10 and MIT Indoor Scene demonstrate the effectiveness and efficiency of our proposed method.
Zengxi Li, Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
ICASSP2
2016 Robust Sound Event Detection in Continuous Audio Environments
abstract
Sound event detection in real world environments has attracted significant research interest recently because of it's applications in popular fields such as machine hearing and automated surveillance, as well as in sound scene understanding. This paper considers continuous robust sound event detection, which means multiple overlapped sound events in different types of interfering noise. First, a standard evaluation task is outlined based upon existing testing data sets for the sound event classification of isolated sounds. This paper then proposes and evaluates the use of spectrogram image features employing an energy detector to segment sound events, before developing a novel segmentation method making use of a Bayesian inference criteria. At the back end, a convolutional neural network is used to classify detected regions, and this combination is compared to several alternative approaches. The proposed method is shown capable of achieving very good performance compared with current state-of-the-art techniques.
Haomin Zhang, Ian McLoughlin 0001, Yan Song 0001
INTERSPEECH3
2016 Image classification with CNN-based Fisher vector coding
abstract
Fisher vector coding methods have been demonstrated to be effective for image classification. With the help of convolutional neural networks (CNN), several Fisher vector coding methods have shown state-of-the-art performance by adopting the activations of a single fully-connected layer as region features. These methods generally exploit a diagonal Gaussian mixture model (GMM) to describe the generative process of region features. However, it is difficult to model the complex distribution of high-dimensional feature space with a limited number of Gaussians obtained by unsupervised learning. Simply increasing the number of Gaussians turns out to be inefficient and computationally impractical. To address this issue, we re-interpret a pre-trained CNN as the probabilistic discriminative model, and present a CNN based Fisher vector coding method, termed CNN-FVC. Specifically, activations of the intermediate fully-connected and output soft-max layers are exploited to derive the posteriors, mean and covariance parameters for Fisher vector coding implicitly. To further improve the efficiency, we convert the pre-trained CNN to a fully convolutional one to extract the region features. Extensive experiments have been conducted on two standard scene benchmarks (i.e. SUN397 and MIT67) to evaluate the effectiveness of the proposed method. Classification accuracies of 60.7% and 82.1% are achieved on the SUN397 and MIT67 benchmarks respectively, outperforming previous state-of-the-art approaches. Furthermore, the method is complementary to GMM-FVC methods, allowing a simple fusion scheme to further improve performance to 61.1% and 83.1% respectively.
Yan Song 0001, Xinhai Hong, Ian McLoughlin 0001, Li-Rong Dai 0001
VCIP1
2015 Improved language identification using deep bottleneck network
abstract
Effective representation plays an important role in automatic spoken language identification (LID). Recently, several representations that employ a pre-trained deep neural network (DNN) as the front-end feature extractor, have achieved state-of-the-art performance. However the performance is still far from satisfactory for dialect and short-duration utterance identification tasks, due to the deficiency of existing representations. To address this issue, this paper proposes the improved representations to exploit the information extracted from different layers of the DNN structure. This is conceptually motivated by regarding the DNN as a bridge between low-level acoustic input and high-level phonetic output features. Specifically, we employ deep bottleneck network (DBN), a DNN with an internal bottleneck layer acting as a feature extractor. We extract representations from two layers of this single network, i.e. DBN-TopLayer and DBN-MidLayer. Evaluations on the NIST LRE2009 dataset, as well as the more specific dialect recognition task, show that each representation can achieve an incremental performance gain. Furthermore, a simple fusion of the representations is shown to exceed current state-of-the-art performance.
Yan Song 0001, Ruilian Cui, Xinhai Hong, Ian McLoughlin 0001, Jiong Shi, Li-Rong Dai 0001
ICASSP1
2015 Robust sound event recognition using convolutional neural networks
abstract
Traditional sound event recognition methods based on informative front end features such as MFCC, with back end sequencing methods such as HMM, tend to perform poorly in the presence of interfering acoustic noise. Since noise corruption may be unavoidable in practical situations, it is important to develop more robust features and classifiers. Recent advances in this field use powerful machine learning techniques with high dimensional input features such as spectrograms or auditory image. These improve robustness largely thanks to the discriminative capabilities of the back end classifiers. We extend this further by proposing novel features derived from spectrogram energy triggering, allied with the powerful classification capabilities of a convolutional neural network (CNN). The proposed method demonstrates excellent performance under noise-corrupted conditions when compared against state-of-the-art approaches on standard evaluation tasks. To the author's knowledge this in the first application of CNN in this field.
Haomin Zhang, Ian McLoughlin 0001, Yan Song 0001
ICASSP3
2015 Low frequency ultrasonic voice activity detection using convolutional neural networks
abstract
Low frequency ultrasonic mouth state detection uses reflected audio chirps from the face in the region of the mouth to determine lip state, whether open, closed or partially open. The chirps are located in a frequency range just above the threshold of human hearing and are thus both inaudible as well as unaffected by interfering speech, yet can be produced and sensed using inexpensive equipment. To determine mouth open or closed state, and hence form a measure of voice activity detection, this recently invented technique relies upon the difference in the reflected chirp caused by resonances introduced by the open or partially open mouth cavity. Voice activity is then inferred from lip state through patterns of mouth movement, in a similar way to video-based lip-reading technologies. This paper introduces a new metric based on spectrogram features extracted from the reflected chirp, with a convolutional neural network classification back-end, that yields excellent performance without needing the periodic resetting of the template closed-mouth reflection required by the original technique.
Ian McLoughlin 0001, Yan Song 0001
INTERSPEECH2
2015 Deep bottleneck network based i-vector representation for language identification
abstract
This paper presents a unified i-vector framework for language identification (LID) based on deep bottleneck networks (DBN) trained for automatic speech recognition (ASR). The framework covers both front-end feature extraction and back-end modeling stages.The output from different layers of a DBN are exploited to improve the effectiveness of the i-vector representation through incorporating a mixture of acoustic and phonetic information. Furthermore, a universal model is derived from the DBN with a LID corpus. This is a somewhat inverse process to the GMM-UBM method, in which the GMM of each language is mapped from a GMM-UBM. Evaluations on specific dialect recognition tasks show that the DBN based i-vector can achieve significant and consistent performance gains over conventional GMM-UBM and DNN based i-vector methods. The generalization capability of this framework is also evaluated using DBNs trained on Mandarin and English corpuses. Index Terms: Language Identification, Deep Neural Network, Deep Bottleneck Feature, i-vector representation
Yan Song 0001, Xinhai Hong, Bing Jiang, Ruilian Cui, Ian McLoughlin 0001, Li-Rong Dai 0001
INTERSPEECH1
2015 Deep Bottleneck Feature for Image Classification
abstract
Effective image representation plays an important role for image classification and retrieval. Bag-of-Features (BoF) is well known as an effective and robust visual representation. However, on large datasets, convolutional neural networks (CNN) tend to perform much better, aided by the availability of large amounts of training data. In this paper, we propose a bag of Deep Bottleneck Features (DBF) for image classification, effectively combining the strengths of a CNN within a BoF framework. The DBF features, obtained from a previously well-trained CNN, form a compact and low-dimensional representation of the original inputs, effective for even small datasets. We will demonstrate that the resulting BoDBF method has a very powerful and discriminative capability that is generalisable to other image classification tasks.
Yan Song 0001, Ian McLoughlin 0001, Li-Rong Dai 0001
ICMR1
2015 Robust Sound Event Classification Using Deep Neural Networks
abstract
The automatic recognition of sound events by computers is an important aspect of emerging applications such as automated surveillance, machine hearing and auditory scene understanding. Recent advances in machine learning, as well as in computational models of the human auditory system, have contributed to advances in this increasingly popular research field. Robust sound event classification, the ability to recognise sounds under real-world noisy conditions, is an especially challenging task. Classification methods translated from the speech recognition domain, using features such as mel-frequency cepstral coefficients, have been shown to perform reasonably well for the sound event classification task, although spectrogram-based or auditory image analysis techniques reportedly achieve superior performance in noise. This paper outlines a sound event classification framework that compares auditory image front end features with spectrogram image-based front end features, using support vector machine and deep neural network classifiers. Performance is evaluated on a standard robust classification task in different levels of corrupting noise, and with several system enhancements, and shown to compare very well with current state-of-the-art classification techniques.
Ian McLoughlin 0001, Haomin Zhang, Zhipeng Xie, Yan Song 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Task-aware deep bottleneck features for spoken language identification
abstract
Recently, deep bottleneck features (DBF) extracted from a deep neural network (DNN) containing a narrow bottleneck lay-er, have been applied for language identification (LID), and yield significant performance improvement over state-of-the-art methods on NIST LRE 2009. However, the DNN is trained us-ing a large corpus of specific language which is not directly related to the LID task. More recently, lattice based discrimi-native training methods for extracting more targeted DBF were proposed for ASR. Inspired by this, this paper proposes to tune the post-trained DNN parameters using an LID-specific train-ing corpus, which may make the resulting DBF, termed a Dis-criminative DBF (D2BF), more discriminative and task-aware. Specifically, the maximum mutual information (MMI) criteri-on, with gradient descent, is applied to update the DNN param-eters of the bottleneck layer in an iterative fashion. We evaluate the performance of the proposed D2BF using different back-end models, including GMM-MMI and ivector, over the most con-fused 6-languages selected from NIST LRE 2009. The results show that the proposed D2BF is more appropriate and effective than the original DBF. Index Terms: language identification, deep bottleneck feature, deep neural network, discriminative training, Gaussian mixture model, maximum mutual information 1.
Bing Jiang, Yan Song 0001, Si Wei, Ian McLoughlin 0001, Li-Rong Dai 0001
INTERSPEECH2
2013 Phoneme variation based synthesized speech discrimination for speaker verification
abstract
How to discriminate the synthesized speech from the natural speech for speaker verification is addressed in this paper. With the development of HMM-based speech synthesis, it is easy to obtain high quality synthesized speech which sounds like target speaker, the robustness of synthesized speech become important for speaker verification. In this paper, a method based on the phoneme variation is proposed to discriminate synthesized speech from natural speech, which could be used as front-end module to detect the synthesized speech for speaker verification system. The experimental results show the effectiveness of the proposed method.
LianWu Chen, Wu Guo, Yan Song 0001, Li-Rong Dai 0001
ICASSP3
2013 Exemplar based language recognition method for short-duration speech segments
abstract
This paper proposes a novel exemplar-based language recognition method for short duration speech segments. It is known that language identity is a kind of weak information that can be deduced from the speech content. For short duration speech segments, the limited content also leads to a large intra-language variability. To address this issue, we propose a new method. This borrows a vector quantization based representation from image classification methods, and constructs the exemplar space using the popular i-vector representation of short duration speech segments. A mapping function is then defined to build the new representation. To evaluate the effectiveness of our proposed method, we conduct extensive experiments on the NIST LRE2007 dataset. The experimental results demonstrate improved performance for short duration speech segments.
Meng-Ge Wang, Yan Song 0001, Bing Jiang, Li-Rong Dai 0001, Ian McLoughlin 0001
ICASSP2
2013 Joint spectral distribution modeling using restricted boltzmann machines for voice conversion
Linghui Chen, Zhen-Hua Ling, Yan Song 0001, Li-Rong Dai 0001
INTERSPEECH3
2013 Reconstruction of continuous voiced speech from whispers
abstract
Whispers are an important secondary vocal communications mechanism, that can be necessary for communicating private information and which are an integral aspect of natural human- to-human dialogue. Furthermore, they may be the primary com- munications method of those suffering from certain forms of aphonia, such as laryngectomees. This paper considers the con- version of continuous whispers to natural-sounding speech, and proposes a new reconstruction method based upon the synthesis of individual formants as excitation source, followed by artifi- cial glottal modulation. Early results show that the proposed method can improve quality and intelligibility over the original whispers when evaluated using continuous speech. It requires neither apriori nor speaker-dependent information, is of rela- tively low-complexity and suitable for real-time processing.
Ian McLoughlin 0001, Yan Song 0001
INTERSPEECH3
2012 Exemplar-Based Sparse Representation for Language Recognition on I-Vectors
abstract
In this paper, a new automatic language identification method using sparse representation on i-vectors in low-dimensional total variability space is proposed. It is mainly based on the recently proposed i-vector based language recognition systems. In our proposed method, an over-complete dictionary is first constructed by randomly sampling of the low-dimensional total variability space after Within-Class Covariance Normalization (WCCN) and Linear Discriminate Analysis (LDA). And then for each test sample, the classification score is derived from sparse linear representation with respect to the over-complete dictionary. Furthermore, a random subspace method, which combines different sparse representation classifiers, is introduced to address the possible over-fitting issue. Evaluations on NIST LRE 2007 dataset show that the proposed method outperforms the state-of-the-art i-vector based language recognition system. Especially for 30s test condition, our proposed method achieves relative reduction of 29.6 % on Equal Error Rate (EER) compared with the baseline system.
Bing Jiang, Yan Song 0001, Wu Guo, Li-Rong Dai 0001
INTERSPEECH2
2011 Spatial pooling for transformation invariant image representation
abstract
Spatial Pyramid Matching (SPM) [2] has been proposed to extend the Bag-of-Word (BoW) model for object classification. By re-serving the finer level information, it makes image matching more accurate. However, for not well-aligned images, where the object is rotated, flipped or translated, SPM may lose its discrimination power. To tackle this problem, we propose novel spatial pooling layouts to address various transformations, and generate a more general image representation. To evaluate the effectiveness of the proposed approach, we conduct extensive experiments on three transformation emphasized datasets for object classification task. Experimental results demonstrate its superiority over the state-of-the-arts. Besides, the proposed image representation is compact and consistent with the BoW model, which makes it applicable to image retrieval task as well.
Yan Song 0001, Yijuan Lu, Qi Tian 0001
ACM Multimedia2
2010 Multiple instance learning using visual phrases for object classification
abstract
Recently, bag of words (BoW) model has led to many significant results in visual object classification. However, due to the limited descriptive and discriminative ability of visual words, the resulting performance of visual object classification is still incomparable to its analogy in text domain, i.e. document categorization. Furthermore, for weakly labeled image data, where we only know whether an object is present or not, traditional learning based methods may suffer from background clutters and large appearance variations. To address these issues, we propose a novel visual phrase based Multiple Instance Learning (MIL) method. In this method, the visual phrase is first generated from over-segmented image regions of homogeneous appearance and visual words within each region, which may provide enhanced descriptive ability by enforcing the spatial coherency. Then a MIL algorithm is applied to efficiently learn from the weakly labeled image data. The experiments on benchmark datasets show that our proposed method always significantly outperforms several state-of-the-art algorithms, such as Spatial Pyramid Matching (SPM) and Spatial-LTM.
Yan Song 0001, Qi Tian 0001, Mengyue Wang, Li-Rong Dai 0001
ICME1
2009 An automatic language identification method based on subspace analysis
abstract
Gaussian mixture models (GMM) have become one of the standard acoustic approaches for language identification. Furthermore, the GMM-SVM is proven to work well by introducing the discriminative method into the GMM-based acoustic systems. In these systems, the intersession variability within language has become an important adverse factor that degrades the system performance. To tackle this problem, we propose a subspace analysis method, termed as Intra-language Difference Subspace Estimatio (IDSE), under the GMM-SVM framework. In IDSE method, the difference vector is modeled with three components: Extra-language difference, Intra-language difference and noise difference. Then the Intra-language and noise difference are effectively estimated and eliminated from the difference vector. The experiments on NIST 07 evaluation tasks show effectiveness of the proposed method.
Yan Song 0001, Li-Rong Dai 0001, Renhua Wang
ICME1
2009 Concept representation based video indexing
abstract
This poster introduces a novel concept-based video indexing approach. It is developed based on a rich set of base concepts, of which the models are available. Then, for a given concept with several labeled samples, we combine the base concepts to fit it and its model can thus be obtained accordingly. Empirical results demonstrate that this method can achieve great performance even with very limited labeled data. We have compared different representation approaches including both sparse and non-sparse methods. Our conclusion is that the sparse method will lead to much better performance.
Meng Wang 0001, Yan Song 0001, Xian-Sheng Hua 0001
SIGIR2
2009 Image Fusion Quality Metrics by Directional Projection
abstract
Image fusion has been over-studied recently. Nevertheless, few works aim to how to evaluate the performance of image fusion algorithms. In this paper, we extend the work in image quality evaluation to a novel metric for objective evaluation of image fusion. Firstly the input images and the result image are converted into local sensitive intensity (LSI) by Radon transform. Then we use the sensitive intensity to measure how many information have been transferred from each source into the fused result by the difference of LSI. Finally all the LSI pairs are incorporated into the expression according to Weber-Fechner law. Experimental results demonstrate that our proposed metric is compliant with subjective evaluations and outperforms other recently developed objective metrics of image fusion.
Richang Hong, Yan Song 0001, Jinhui Tang 0001, Jianxin Pang
SMC2
2009 Concept-Dependent Image Annotation via Existence-Based Multiple-Instance Learning
abstract
Conventional multiple-instance learning (MIL) algorithms for image annotation usually neglect concept dependence (i.e., the relationship between positive and negative concepts) and feature selection (i.e., which feature modality is suitable for a specific concept) problems, which have significant influence on the annotation performance. In this paper, we propose a novel concept-dependent algorithm for image annotation, named existence-based MIL (EBMIL), aiming at solving the above two problems in one scheme. In our EBMIL scheme, we give a new MIL formulation, named existence-based MIL, to explore the concept dependence in image annotation. Moreover, we give an optimization procedure in EBMIL, which is able to select different feature modalities for each concept under MIL settings. EBMIL achieves promising experimental results on the benchmark of COREL dataset with comparison to typical MIL algorithms.
Xun Yuan 0001, Meng Wang 0001, Yan Song 0001
SMC3
2009 Semi-supervised kernel density estimation for video annotation
Meng Wang 0001, Xian-Sheng Hua 0001, Tao Mei 0001, Richang Hong, Guo-Jun Qi, Yan Song 0001, Li-Rong Dai 0001
Comput. Vis. Image Underst.6
2009 Unified Video Annotation via Multigraph Learning
abstract
Learning-based video annotation is a promising approach to facilitating video retrieval and it can avoid the intensive labor costs of pure manual annotation. But it frequently encounters several difficulties, such as insufficiency of training data and the curse of dimensionality. In this paper, we propose a method named optimized multigraph-based semi-supervised learning (OMG-SSL), which aims to simultaneously tackle these difficulties in a unified scheme. We show that various crucial factors in video annotation, including multiple modalities, multiple distance functions, and temporal consistency, all correspond to different relationships among video units, and hence they can be represented by different graphs. Therefore, these factors can be simultaneously dealt with by learning with multiple graphs, namely, the proposed OMG-SSL approach. Different from the existing graph-based semi-supervised learning methods that only utilize one graph, OMG-SSL integrates multiple graphs into a regularization framework in order to sufficiently explore their complementation. We show that this scheme is equivalent to first fusing multiple graphs and then conducting semi-supervised learning on the fused graph. Through an optimization approach, it is able to assign suitable weights to the graphs. Furthermore, we show that the proposed method can be implemented through a computationally efficient iterative process. Extensive experiments on the TREC video retrieval evaluation (TRECVID) benchmark have demonstrated the effectiveness and efficiency of our proposed approach.
Meng Wang 0001, Xian-Sheng Hua 0001, Richang Hong, Jinhui Tang 0001, Guo-Jun Qi, Yan Song 0001
IEEE Trans. Circuits Syst. Video Technol.6
2008 Video Annotation Based on Kernel Linear Neighborhood Propagation
abstract
The insufficiency of labeled training data for representing the distribution of the entire dataset is a major obstacle in automatic semantic annotation of large-scale video database. Semi-supervised learning algorithms, which attempt to learn from both labeled and unlabeled data, are promising to solve this problem. In this paper, a novel graph-based semi-supervised learning method namedkernellinearneighborhoodpropagation(KLNP) is proposed and applied to video annotation. This approach combines theconsistencyassumption, which is the basic assumption in semi-supervised learning, and thelocallinearembedding(LLE) method in a nonlinear kernel-mapped space. KLNP improves a recently proposed methodlinearneighborhoodpropagation(LNP) by tackling the limitation of its local linear assumption on the distribution of semantics. Experiments conducted on the TRECVID data set demonstrate that this approach outperforms other popular graph-based semi-supervised learning methods for video semantic annotation.
Jinhui Tang 0001, Xian-Sheng Hua 0001, Guo-Jun Qi, Yan Song 0001, Xiuqing Wu
IEEE Trans. Multim.4
2007 An Interactive Video Annotation Frameowrk with Multiple Modalities
abstract
Active learning and semi-supervised learning methods are frequently applied in multimedia annotation tasks in order to reduce human labeling effort. However, in most of these methods only single modality is applied. This paper presents an interactive video annotation framework, which is based on semi-supervised learning and active learning with multiple multimodalities. In the proposed framework, unlabeled samples are iteratively selected to be annotated manually according to certain strategy which has taken the potentials of different modalities into account, and then a graph-based semi-supervised learning algorithm is conducted on each modality. This process repeats for several rounds, and the results obtained from multiple modalities are then fused to generate final output. The proposed framework is computationally efficient, and the experimental results on TRECVID 2005 benchmark show that the proposed framework considerably outperforms previous approaches.
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Li-Rong Dai 0001, Renhua Wang
ICASSP (1)3
2007 Transductive Inference with Hierarchical Clustering for Video Annotation
abstract
In this paper, we present a novel framework for video semantic detection based on transductive inference and hierarchical clustering, which directly focuses on predicting the available samples in a current unlabeled pool, instead of trying to build a classifier workable for any unavailable data. In this framework, a number of hierarchical clustering results are constructed from theentire video datasetcontaining both labeled and unlabeled examples. We aim to make the clusters aspureas possible, i.e., samples in a same cluster mostly have a same label. To furtherpurifythese hierarchical clustering results, an EM based cluster-tuning algorithm is iteratively employed. Based on these clustering results, several hypotheses are generated byprobability votingamong labeled samples in the obtained clusters. From these hypotheses, one of them is chosen according to theVapnik combined bound, and it is then applied to predict the labels of unlabeled samples. This selected transductive hypothesis, which is only interested in predicting the available unlabeled samples in test set rather than producing a general classifier like inductive inference learning, exploits the structure and distribution of the unlabeled pool to achieve a minimal test error bound. Thus it can have better generalization ability for video annotation both theoretically and experimentally. This is also shown by our experiment results.
Guo-Jun Qi, Xian-Sheng Hua 0001, Yan Song 0001, Jinhui Tang 0001, HongJiang Zhang
ICME3
2007 Lazy Learning Based Efficient Video Annotation
abstract
Eager learning methods, such as SVM, are widely applied in video annotation task for their substantial performance. However, their computational costs are usually prohibitive when a large dataset is faced, especially when annotating a large lexicon of semantic concepts. This paper proposes a video annotation scheme based on lazy learning, and shows that this scheme is much more computationally efficient and flexible. Based on a recently proposed improved Parzen window method, we provide a lazy learning based video annotation scheme. After building the pairwise relationships in dataset, the annotation can be finished rapidly for each concept. Experiments show that the proposed method is much more efficient than SVM while retaining comparable performance.
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Richang Hong, Li-Rong Dai 0001
ICME3
2007 Multi-Graph Semi-Supervised Learning for Video Semantic Feature Extraction
abstract
This paper proposes a video semantic feature extraction approach based on multi-graph semi-supervised learning, which aims to simultaneously deal with the insufficiency of training data and the curse of dimensionality. In contrast to traditional graph-based semi-supervised learning, which generates graph from high-dimensional low-level features, we separate the original low-level features into multiple modalities with minimum correlations, and thus multiple graphs are obtained from these modalities. This way can tackle the curse of dimensionality brought by the high-dimensional feature space. We then propose a criterion to optimally fuse these graphs based on the pairwise relationships among training samples, and implement semi-supervised learning on the fused graph. Experimental results have demonstrated the effectiveness of the proposed approach.
Meng Wang 0001, Xian-Sheng Hua 0001, Xun Yuan 0001, Yan Song 0001, Li-Rong Dai 0001
ICME4
2007 Optimizing multi-graph learning: towards a unified video annotation scheme
abstract
Learning based semantic video annotation is a promising approach for enabling content-based video search. However, severe difficulties, such as insufficiency of training data and curse of dimensionality, are frequently encountered. This paper proposes a novel unified scheme, Optimized Multi-Graph-based Semi-Supervised Learning (OMG-SSL), to simultaneously attack these difficulties. Instead of only using a single graph, OMG-SSL integrates multiple graphs into a regularization and optimization framework to sufficiently explore their complementary nature. We then show that various crucial factors in video annotation, including multiple modalities, multiple distance metrics, and temporal consistency, in fact all correspond to different correlations among samples, and hence they can be represented by different graphs. Therefore, OMG-SSL is able to simultaneously deal with these factors within a unified framework. Experiments on the TRECVID benchmark demonstrate the effectiveness of our proposed approach.
Meng Wang 0001, Xian-Sheng Hua 0001, Xun Yuan 0001, Yan Song 0001, Li-Rong Dai 0001
ACM Multimedia4
2007 Video annotation by graph-based learning with neighborhood similarity
abstract
Graph-based semi-supervised learning methods have been proven effective in tackling the difficulty of training data insufficiency in many practical applications such as video annotation. These methods are all based on an assumption that the labels of similar samples are close. However, as a crucial factor of these algorithms, the estimation of pairwise similarity has not been sufficiently studied. Usually, the similarity of two samples is estimated based on the Euclidean distance between them. But we will show that similarities are not merely related to distances but also related to the structures around the samples. It is shown that distance-based similarity measure may lead to high classification error rates even on several simple datasets. In this paper we propose a novel neighborhood similarity measure, which simultaneously takes into account both thse distance between samples and the difference between the structures around the corresponding samples. Experiments on synthetic dataset and TRECVID benchmark demonstrate that the neighborhood similarity is superior to existing distance based similarity.
Meng Wang 0001, Tao Mei 0001, Xun Yuan 0001, Yan Song 0001, Li-Rong Dai 0001
ACM Multimedia4
2007 An Efficient Automatic Video Shot Size Annotation Scheme
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Li-Rong Dai 0001, Renhua Wang
MMM (1)3
2007 Kernel-Based Linear Neighborhood Propagation for Semantic Video Annotation
Jinhui Tang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Guo-Jun Qi, Xiuqing Wu
PAKDD3
2006 An Automatic Video Semantic Annotation Scheme Based on Combination of Complementary Predictors
abstract
Given a large set of video database, how to connect video segments with a certain set of semantic concepts with least manual labors is an elementary step for video indexing and searching. Due to the large gap between high-level semantics and low-level features, automatic video annotation with high accuracy is a challenging task. In this paper, we propose a novel automatic video annotation framework, which improves the annotation performance by learning from unlabeled samples and exploring temporal relationship in video sequences. To effectively learn from unlabeled data, a sample selection scheme based on combining a set of complementary predictors is proposed, which iteratively refines the performance of the initial predictors. A filtering-based method is applied to further improve the annotation accuracy as well, in which video temporal relationship is sufficiently exploited. Experiment results show that the proposed automatic video annotation method performs superior to general supervised learning methods and co-training
Yan Song 0001, Xian-Sheng Hua 0001, Li-Rong Dai 0001, Meng Wang 0001, Renhua Wang
ICASSP (5)1
2006 Semi-Supervised Kernel Regression
abstract
Insufficiency of training data is a major obstacle in machine learning and data mining applications. Many different semi-supervised learning algorithms have been proposed to tackle this difficulty by leveraging a large amount of unlabeled data. However, most of them focus on semi-supervised classification. In this paper we propose a semi-supervised regression algorithm named semi-supervised kernel regression (SSKR). While classical kernel regression is only based on labeled examples, our approach extends it to all observed examples using a weighting factor to modulate the effect of unlabeled examples. Experimental results prove that SSKR significantly outperforms traditional kernel regression and graph-based semi-supervised regression methods.
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Li-Rong Dai 0001, HongJiang Zhang
ICDM3
2006 Video Annotation by Active Learning and Semi-Supervised Ensembling
abstract
Supervised and semi-supervised learning are frequently applied methods to annotate videos by mapping low-level features into semantic concepts. Due to the large semantic gap, the main constraint of these methods is that the information contained in a limited-size labeled dataset can hardly represent the distributions of the semantic concepts. In this paper, we propose a novel semi-automatic video annotation framework, active learning with semi-supervised ensembling, which tries to tackle the disadvantages of current video annotation solutions. Firstly the initial training set is constructed based on distribution analysis of the entire video dataset and then an active learning scheme is combined into a semi-supervised ensembling framework, which selects the samples to maximize the margin of the ensemble classifier based on both labeled and unlabeled data. Experimental results show that the proposed method performs superior to general semi-supervised learning algorithms and typical active learning algorithms in terms of annotation accuracy and stability
Yan Song 0001, Guo-Jun Qi, Xian-Sheng Hua 0001, Li-Rong Dai 0001, Renhua Wang
ICME1
2006 Enhanced Semi-Supervised Learning for Automatic Video Annotation
abstract
For automatic semantic annotation of large-scale video database, the insufficiency of labeled training samples is a major obstacle. General semi-supervised learning algorithms can help solve the problem but the improvement is limited. In this paper, two semi-supervised learning algorithms, self-training and co-training, are enhanced by exploring the temporal consistency of semantic concepts in video sequences. In the enhanced algorithms, instead of individual shots, time-constraint shot clusters are taken as the basic sample units, in which most mis-classifications can be corrected before they are applied for re-training, thus more accurate statistical models can be obtained. Experiments show that enhanced self-training/co-training significantly improves the performance of video annotation
Meng Wang 0001, Xian-Sheng Hua 0001, Li-Rong Dai 0001, Yan Song 0001
ICME4
2006 Automatic video annotation based on co-adaptation and label correction
abstract
As there is a large gap between high-level semantics and low-level features, it is difficult to obtain high-accuracy video semantic annotation through automatic methods. In this paper, we propose a novel automatic video annotation method, which greatly improves the annotation performance by learning from unlabeled video data, as well as exploring temporal consistency of video sequences. To effectively learn from unlabeled data, a scheme called co-adaptation is proposed to progressively refine two pre-trained complementary classifiers, and then a minimum entropy based method is applied to sufficiently explore the video temporal consistency, which further improves the annotation accuracy. Experiments show that the proposed automatic video annotation method performs superior than both general learning-based and co-training-based methods
Meng Wang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Li-Rong Dai 0001, Shipeng Li 0001
ISCAS3
2006 To construct optimal training set for video annotation
abstract
This paper exploits the criteria to optimize the training set construction for video annotation. Most existing learning-based semantic annotation approaches require a large training set to achieve good generalization capacity, in which a considerable amount of labor-intensively manual labeling is desirable. However, it is observed that the generalization capacity of a classifier highly depends on the geometrical distribution rather than the size of the training data. We argue that a training set which includes most temporal and spatial distribution of the whole data will achieve a satisfying performance even in the case of limited size of training set. In order to capture the geometrical distribution characteristics of a given video collection, we propose the following four metrics for constructing an optimal training set, including Salience Time Dispersiveness Spatial Dispersiveness and Diversity. Moreover, based on these metrics, we propose a set of optimization rules to capture the most distribution information of the whole data for a training set with a given size. Experimental results demonstrate that these rules are effective for training set construction for video annotation, and significantly outperform random training set selection as well.
Jinhui Tang 0001, Yan Song 0001, Xian-Sheng Hua 0001, Tao Mei 0001, Xiuqing Wu
ACM Multimedia2
2006 Automatic video annotation by semi-supervised learning with kernel density estimation
abstract
Insufficiency of labeled training data is a major obstacle for automatically annotating large-scale video databases with semantic concepts. Existing semi-supervised learning algorithms based on parametric models try to tackle this issue by incorporating the information in a large amount of unlabeled data. However, they are based on a assumption that the assumed generative model is correct, which usually cannot be satisfied in automatic video annotation due to the large variations of video semantic concepts. In this paper, we propose a novel semi-supervised learning algorithm, named Semi Supervised Learning by Kernel Density Estimation (SSLKDE), which is based on a non-parametric method, and therefore the assumption is avoided. While only labeled data are utilized in the classical Kernel Density Estimation (KDE) approach, in SSLKDE both labeled and unlabeled data are leveraged to estimate class conditional probability densities based on an extended form of KDE. We also investigate the connection between SSLKDE and existing graph-based semi-supervised learning algorithms. Experiments prove that SSLKDE significantly outperforms existing supervised methods for video annotation.
Meng Wang 0001, Yan Song 0001, Xun Yuan 0001, HongJiang Zhang, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia2