EDBT 2026 Demo / reviewers in the wild / expert
Qirong Mao
dblp:88/1633
· DBLP profile ↗
109ranked-venue papers
10as first author
56since 2021 · last 2026
0000-0002-0616-4431ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 1 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 5 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-Guided Visual Byte-Pair Encoding for Unified Autoregressive Multimodal Modeling
Wenlong Dong, Lijian Gao, Qirong Mao |
ICIC (12) | 4 |
| 2026 | Keyword Mamba: Spoken keyword spotting with state space models
Hanyu Ding, Wenlong Dong, Qirong Mao |
Comput. Speech Lang. | 3 |
| 2026 | Hierarchical Temporal Sequence Segmentation for weakly supervised video anomaly detection
Nuku Atta Kordzo Abiew, Lijian Gao, Qirong Mao |
Expert Syst. Appl. | 3 |
| 2026 | HCATRE-AVAD: Hierarchical cross-alignment and temporal relational encoding for weakly supervised audio-visual anomaly detection
Nuku Atta Kordzo Abiew, Lijian Gao, Godbless Mensah, Wenlong Dong, Qirong Mao |
Image Vis. Comput. | 5 |
| 2026 | Adaptive Key Role Guided Hierarchical Relation Inference for Enhanced Group-Level Emotion RecognitionabstractIn this paper, we propose a novel hierarchical relational network, termed Key Role Guided Hierarchical Relation Inference (KR-HRI), for enhanced group-level emotion recognition (GER). Unlike existing methods that adopt a coarse-grained approach to model interactions among all individuals, our approach adaptively identifies and emphasizes key individuals who play a crucial role in conveying group-level emotions. By integrating coarse-grained relationship modeling with fine-grained key individual enhancement and leveraging global scene information, our method effectively refines discriminative feature generation while minimizing irrelevant interference. We introduce a Multi-branch Interaction Module (MIM) to dynamically fuse features from both the global scene and local individual branches using a localized mask integration strategy. This comprehensive approach enhances the interaction between global and local features, resulting in robust group-level emotion representations. Extensive experiments on three widely adopted GER datasets demonstrate that our framework consistently outperforms state-of-the-art methods, validating the effectiveness and robustness of our proposed approach. Qing Zhu 0002, Qirong Mao, Wenlong Dong, Xiuyan Shao, Xiaohua Huang 0003, Wenming Zheng |
IEEE Trans. Affect. Comput. | 2 |
| 2026 | A Survey on Deep Learning for Group-Level Emotion RecognitionabstractWith the rapid advancement of artificial intelligence, group-level emotion recognition (GER) has emerged as an important domain in human behavior analysis. Early GER methods primarily relied on handcrafted features. However, the recent success of deep learning has shifted the focus toward neural network-based solution, enabling more effective exploitation of the rich visual and contextual cues in group images and videos. Unlike individual-level emotion recognition, GER must account for the diversity and dynamics of multiple individuals within varied social contexts. Over the past decade, numerous deep learning-based methods have been proposed, achieving substantial performance gains. This survey provides a comprehensive review of deep learning-centric review of GER, introducing a new taxonomy that spans representation learning, graph-based modeling, attention and transformer architectures, and multimodal fusion strategies. We summarize benchmark datasets, outline prevailing GER pipelines, and consolidate performance trends from recent state-of-the-art approaches. In addition, we discuss the integration of foundation models and large language model-guided multimodal reasoning into GER. Key challenges are identified, and potential research directions are proposed to support the development of robust, real-world GER systems. This work aims to serve as a pivotal reference for future research in this evolving field. Xiaohua Huang 0003, Xiaopeng Hong, Qirong Mao, Wenming Zheng, Abhinav Dhall |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2026 | Enhanced Face Clustering With Neighbor Structure RefinementabstractFace clustering is crucial for applications such as facial recognition, identity verification, and surveillance in social systems. However, it is often hindered by noise and false-positive connections at cluster boundaries, which significantly degrade performance. To address these challenges, we propose a novel face clustering framework with neighbor structure refinement (NSR-FC), which enhances clustering effectiveness by refining neighbor structures across three dimensions: features, density, and connectivity. Within NSR-FC, cascaded graph convolutional networks (C-GCN) improve feature extraction while simultaneously optimizing the graph structure to reduce noise and generate discriminative features. Leveraging these refined features, we construct a local density graph using updated affinities from the k-nearest neighbor (KNN) graph, effectively eliminating negative pairs while preserving positive ones. Furthermore, we assess edge connectivity within the density graph to construct a global connectivity graph. Finally, the clustering results are obtained by applying breadth-first search (BFS) to the union graph edge set. Extensive experiments on benchmark datasets such as MS-Celeb-1M demonstrate that NSR-FC achieves state-of-the-art (SOTA) performance, underscoring its effectiveness in advancing face clustering. Hongjie Jia, Ying Zhi, Qirong Mao, Heping Song |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | EDSep: An Effective Diffusion-Based Method for Speech Source SeparationabstractGenerative models have attracted considerable attention for speech separation tasks, and among these, diffusion-based methods are being explored. Despite the notable success of diffusion techniques in generation tasks, their adaptation to speech separation has encountered challenges, notably slow convergence and suboptimal separation outcomes. To address these issues and enhance the efficacy of diffusion-based speech separation, we introduce EDSep, a novel single-channel method grounded in score matching via stochastic differential equation (SDE). This method enhances generative modeling for speech source separation by optimizing training and sampling efficiency. Specifically, a novel denoiser function is proposed to approximate data distributions, which obtains ideal denoiser outputs. Additionally, a stochastic sampler is carefully designed to resolve the reverse SDE during the sampling process, gradually separating speech from mixtures. Extensive experiments on databases such as WSJ0-2mix, LRS2-2mix, and VoxCeleb2-2mix demonstrate our proposed method’s superior performance over existing diffusion and discriminative models, validating its efficacy. Jinwei Dong, Qirong Mao |
ICASSP | 3 |
| 2025 | Key Clues Guided Video Character Social Relationship Recognition Enhanced by LLMabstractVideo Character Social Relationship Recognition (VCSRR) requires a comprehensive consideration about spatio-temporal and multi-modal clues in videos. Most existing methods mainly focus on integrating multi-modal clues and modeling interactions among characters. However, they fail to discover key clues in the complex video data or fully understand the clues related to social relationships. In this article, we propose a novel Large Language Model Enhanced Key Clues Selection (LE-KCS) framework to address the aforementioned issues. The core of LE-KCS is to mine multi-scale key clues from the perspectives of time, space and multi-modality, then transfer the knowledge about social relationships of the Large Language Model to VCSRR for understanding the selected clues. We evaluated LE-KCS on the MovieGraphs dataset and the experimental results indicate that our proposed LE-KCS achieves state-of-the-art performance. Wenlong Dong, Qing Zhu 0002, Qirong Mao |
ICASSP | 3 |
| 2025 | Joint Multi-Scale Contextual and Noise Suppression for Group Emotion RecognitionabstractGroup Emotion Recognition (GER) seeks to identify emotional states within multi-person groups. The complexity of the environment and the reliance on a single group emotion label often result in noisy individuals with inconsistent emotional expressions, hampering accurate emotion classification. Mainstream GER networks attempt to downweight noisy individuals, but when their numbers are high and over-relying on a single group label, the negative impact of noise on individuals remains significant, further complicating classification and diminishing accuracy. To address these limitations, we propose the Multi-Scale Contextual and Noise Suppression Model (MCon-NSM), a novel framework that enhances GER by capturing fine-grained interaction contexts and collaboratively suppressing noise at both the individual and group label levels. Extensive experiments on the GAF series datasets demonstrate that our method achieves results comparable to state-of-the-art techniques, validating its effectiveness in mitigating noise in GER. Wangdong Guo, Qing Zhu 0002, Qirong Mao |
ICASSP | 3 |
| 2025 | Global Enhanced Frame Prompt Tuning for Sound Event DetectionabstractSound Event Detection (SED) often employs pre-trained models to address data scarcity issues. However, existing systems usually treat the pretrained models as frozen feature extractors, resulting in suboptimal efficiency, or fully fine-tune the pretrained models, which requires substantial computational resources. To fully leverage the knowledge from pretrained models, we propose a novel Global Enhanced Frame Prompt Tuning (GE-FPT) framework, providing global and local insights tailored for SED tasks. Additionally, Frame Prompt Tuning (FPT) is proposed in our GE-FPT to effectively explore local temporal information, i.e., temporal details and context, which is essential for SED tasks, and in particular, for precise event boundary detection. Extensive experiments claim that our approach significantly outperforms full fine-tuning methods while substantially reducing computational costs. Our system achieves new state-of-the-art results, with PSDS1/PSDS2 scores of 0.628/0.845 on the DCASE2023 Challenge Task4 dataset. The source code is publicly available1. Shiyu Yu, Lijian Gao, Qirong Mao |
ICASSP | 3 |
| 2025 | Scattering-Conditioned Diffusion Models for Multiple Appropriate Facial Reaction GenerationabstractAs embodied intelligence has become a new hot topic in current artificial intelligence research, facial reaction generation has increasingly become a key technology for achieving natural human-computer interaction. Existing methods typically rely on bimodal inputs of audio and visual signals, but they still suffer from poor cross-modal consistency and insufficient feature fusion, making it difficult to effectively capture complex facial features. To enhance the representational capacity of fused features, this paper proposes a Scattering-Conditioned Diffusion Model (SC-Diff), which extracts stable multi-scale structural features in the frequency domain via wavelet scattering transform and injects them into the diffusion generation process as conditional information, thereby enhancing the modeling ability of representing local facial variations. Furthermore, considering that different prediction tasks exhibit varying sensitivity to target changes during training, we introduce an uncertainty-based adaptive loss weighting strategy to dynamically balance three types of supervision targets: facial action units, facial affect, and facial expressions. Experimental results on the REACT 2025 dataset demonstrate that the proposed method outperforms existing state-of-the-art approaches across multiple evaluation metrics. Qirong Mao, Qiwei Wu 0002, Yakui Ding, Lijian Gao |
ACM Multimedia | 1 |
| 2025 | StyU-STD: Style-Diverse Sample Generation from Unlabeled Data for Query-by-Example Spoken Term DetectionabstractIn recent years, query-by-example spoken term detection (QbE-STD) techniques have made significant progress in detection accuracy and speed. However, this task also encounters situations where labeled data is scarce or even nonexistent, with only unlabeled data available. Although some solutions exist, they still struggle to effectively handle highly variable speech, especially when it comes to differing styles. To address this issue, we propose a self-supervised learning method named Style-diverse sample generation from Unlabeled data for query-by-example Spoken Term Detection (StyU-STD). The core idea is to generate samples with the same content but different styles for learning. Specifically, we randomly extract segments from the speech to be tested as positive samples, while segments randomly extracted from other speech data are labeled as negative samples of the speech to be tested. In addition, various transformations are applied to alter the style of both positive and negative samples while preserving their original content. Then, the generated sample pairs are used to train the Style Suppressed Convolutional Network, which focuses more on content-related information in speech and effectively reduces the interference caused by style differences. The experimental results show that, across multiple datasets, our method outperforms existing methods, achieving higher accuracy and robustness. Hanyu Ding, Lijian Gao, Wenlong Dong, Xiangrui Li, Qirong Mao |
SMC | 5 |
| 2025 | Dynamic prompting class distribution optimization for semi-supervised sound event detectionabstractSemi-supervised sound event detection (SSED) tasks typically leverage a large amount of unlabeled and synthetic data to facilitate model generalization during training, reducing overfitting on a limited set of labeled data. However, the generalization training process often encounters challenges from noisy interference introduced by pseudo-labels or domain knowledge gaps. To alleviate noisy interference in class distribution learning, we propose an efficient semi-supervised class distribution learning method through dynamic prompt tuning, named prompting class distribution optimization (PADO). Specifically, when modeling real labeled data, PADO dynamically incorporates independent learnable prompt tokens to explore prior knowledge about the true distribution. Then, the prior knowledge serves as prompt information, dynamically interacting with the posterior noisy-class distribution information. In this case, PADO achieves class distribution optimization while maintaining model generalization, leading to a significant improvement in the efficiency of class distribution learning. Compared with state-of-the-art methods on the SSED datasets from DCASE 2019, 2020, and 2021 challenges, PADO achieves significant performance improvements. Furthermore, it is readily extendable to other benchmark models. Lijian Gao, Qing Zhu 0002, Yaxin Shen, Qirong Mao, Yongzhao Zhan 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2025 | Label correlation preserving visual-semantic joint embedding for multi-label zero-shot learning
Zhongchen Ma, Guangchen Wang, Qirong Mao, Ming Dong 0001 |
Multim. Tools Appl. | 4 |
| 2025 | An Empirical Study of Super-Resolution on Low-Resolution Micro-Expression Recognition
Ling Zhou 0005, Mingpei Wang, Xiaohua Huang 0003, Wenming Zheng, Qirong Mao, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | Towards a Robust Group-Level Emotion Recognition via Uncertainty-Aware LearningabstractGroup-level emotion recognition (GER) is an inseparable part of human behavior analysis, aiming to recognize an overall emotion in a multi-person scene. However, the existing methods are devoted to combing diverse emotion cues while ignoring the inherent uncertainties under unconstrained environments, such as congestion and occlusion occurring within a group. Additionally, since only group-level labels are available, inconsistent emotion predictions among individuals in one group can confuse the network. In this paper, we propose an uncertainty-aware learning (UAL) method to extract more robust representations for GER. By explicitly modeling the uncertainty, we adopt stochastic embedding sourced from a Gaussian distribution instead of deterministic point embedding. It helps capture the probabilities of emotions and facilitates diverse inferences. Additionally, we adaptively assign uncertainty-sensitive scores as the fusion weights for individuals’ faces within a group. Moreover, we developed an image enhancement module to evaluate and filter samples, strengthening the model’s data-level robustness against uncertainties. The overall three-branch model, encompassing face, object, and scene components, is guided by a proportional-weighted fusion strategy and integrates the proposed uncertainty-aware method to produce the final group-level output. Experimental results demonstrate the effectiveness and generalization ability of our method across three widely used databases. Qing Zhu 0002, Qirong Mao, Xiaohua Huang 0003, Wenming Zheng |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Infrared-Visible Image Fusion Using Dual-Branch Auto-Encoder With Invertible High-Frequency EncodingabstractIn the field of Infrared-Visible Image Fusion (IVIF), the preservation of details, edges, and texture is crucial for generating high-quality fused images. However, a major challenge arises due to the inevitable loss of high-frequency information during feature extraction, resulting in fused images that lack significant details. In this paper, we propose a dual-branch auto-encoder by exploiting an invertible high-frequency branch for detailed feature preservation and a transformer-based low-frequency branch for global dependencies modeling. First, the high-frequency branch employs the wavelet transforms and an Invertible Neural Networks (INN)-based encoder to model high-frequency features through an invertible transformation, including a forward process for image fusion and an inverse process for original image reconstruction. Additionally, a high-frequency loss is designed to enhance the high-frequency feature representation for high-quality image fusion. Second, a low-frequency branch based on a transformer encoder and an adaptive fusion module is introduced to capture the global contextual features of the infrared and visible images. Finally, the decoder integrates the low- and high-frequency features from both branches to generate the final fused image. Image fusion, object detection, and semantic segmentation experiments conducted on public datasets such as TNO, MFNet, and M3FD, show that our method outperforms the state-of-the-art (SOTA) image fusion methods. Qirong Mao, Ming Dong 0001, Yongzhao Zhan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | On Learning Frequency-Instance Correlations by Model-Agnostic Training for Synthetic Speech Detection
Lijian Gao, Qirong Mao |
ACML | 4 |
| 2024 | PL-TTS: A Generalizable Prompt-based Diffusion TTS Augmented by Large Language Model
Qirong Mao, Jiatong Shi |
INTERSPEECH | 2 |
| 2024 | Leveraging Contrastive Language-Image Pre-Training and Bidirectional Cross-attention for Multimodal Keyword Spotting
Dong Liu 0037, Qirong Mao, Lijian Gao, Gang Wang 0023 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | A post-processing framework for class-imbalanced learning in a transductive setting
Yu Lu 0017, Yongzhao Zhan 0001, Qirong Mao |
Expert Syst. Appl. | 5 |
| 2024 | A novel conversational hierarchical attention network for speech emotion recognition in dyadic conversation
Mohammed Tellai, Lijian Gao, Qirong Mao, Mounir Abdelaziz |
Multim. Tools Appl. | 3 |
| 2024 | On Local Temporal Embedding for Semi-Supervised Sound Event DetectionabstractSemi-supervised sound event detection (SSED) task requires recognizing the categories of events and marking each event's onset and offset times in a mixed audio recording using a small amount of weakly labeled and a large scale of unlabeled data. So, exploring local temporal information, i.e., local discrimination and local correlations in the time domain, is essential for SSED, and in particular, for precise event boundary detection. Besides, as manual-labeled datasets are scarce, SSED tasks require effectively exploiting unlabelled data to reduce overfitting, typically through regularization techniques. Recently, self-supervised learning provided a viable solution to leverage unlabeled data for effective feature learning in various downstream tasks. In this paper, we propose LTE-Net, a novel multitask framework, to learn the Local Temporal Embedding for SSED. Specifically, LTE-Net first locally down-samples the input spectrogram and learns the token embeddings with a high temporal resolution (i.e., local discrimination). Then, LTE-Net effectively models the local correlations among the token embeddings through self-supervised masked spectrogram modeling. Finally, a novel joint (self- and semi-supervision) regularization framework is employed for the training of LTE-Net to effectively leverage unlabeled data in SSED. Extensive experiments on DCASE 2019, 2020 and 2021 SSED datasets show that LTE-Net significantly outperformed existing methods and achieved 2.1% to 8.7%, 2.1% to 3.9% and 1.2% to 6.1% performance gains on the evaluation set in 2019, 2020 and 2021 datasets, respectively. Lijian Gao, Qirong Mao, Ming Dong 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Adaptive Density Subgraph ClusteringabstractDensity peak clustering (DPC) has garnered growing interest over recent decades due to its capability to identify clusters with diverse shapes and its resilience to the presence of noisy data. Most DPC-based methods exhibit high computational complexity. One approach to mitigate this issue involves utilizing density subgraphs. Nevertheless, the utilization of density subgraphs may impose restrictions on cluster sizes and potentially lead to an excessive number of small clusters. Furthermore, effectively handling these small clusters, whether through merging or separation, to derive accurate results poses a significant challenge, particularly in scenarios where the number of clusters is unknown. To address these challenges, we propose an adaptive density subgraph clustering algorithm (ADSC). ADSC follows a systematic three-step procedure. First, the highdensity regions in the dataset are recognized as density subgraphs based on k-nearest neighbor (KNN) density. Second, the initial clustering is carried out by utilizing an automated mechanism to identify the important density subgraphs and allocate outliers. Last, the obtained initial clustering results are further refined in an adaptive manner using the cluster self-ensemble technique, ultimately yielding the final clustering outcomes. The clustering performance of the proposed ADSC algorithm is evaluated on nineteen benchmark datasets. The experimental results demonstrate that ADSC possesses the ability to automatically determine the optimal number of clusters from intricate density data, all while maintaining high clustering efficiency. Comparative analysis against other well-known density clustering algorithms that require prior knowledge of cluster numbers reveals that ADSC consistently achieves comparable or superior clustering results. Hongjie Jia, Yuhao Wu 0011, Qirong Mao, Yang Li 0231, Heping Song |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Tiny Object Detection via Regional Cross Self-Attention NetworkabstractAs vision sensor technology continues to evolve, the requirements for detecting targets of interest in the images captured by the sensors are increasing. Considering fast detection and high accuracy, the industry favors geometric key point-based solutions. However, there are a large number of small and fuzzy objects in the real world. Geometric key point detectors do not effectively utilize the contextual features of the region of interest, leading to excessive false positive and false negative results. In this work, a simple, effective, and interpretable tiny object detection method called Regional Cross Self-Attention Object Detection Network (RCSANet) is proposed. It adopts Region Proposal Networks and transformers to capture regional background relations and uses regional background relations to generate key point sequences. The regional cross self-attention mechanism is introduced to curtail computation redundancy and minimize the interference of redundant information to the target region. Additionally, a position coding called dynamic implicit position coding is proposed to cooperate with regional cross self-attentiveness. Dynamic implicit location coding can encode arbitrarily long input sequences. The computational cost of RCSANet is significantly lower than that of state-of-the-art object detection solutions. Moreover, RCSANet improves the performance on the four benchmark datasets, of MSCOCO, Tinyperson, DOTA, and AI-TOD, by about 3.0%AP. Keyang Cheng, Honggang Cui, Humaira abdul Ghafoor, Qirong Mao, Yongzhao Zhan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Exploring Prototype-Anchor Contrast for Semantic SegmentationabstractPixel-wise contrastive learning recently offers a new training paradigm in semantic segmentation by directly shaping the pixel embedding space. Compared with pixel-pixel contrast that often requires large memory and high computation cost, pixel-prototype contrast exploits the semantic correlations among pixels in a more efficient way by pulling positive pixel-prototype pairs close and pushing negative pairs apart. However, most existing work treats pixels as anchors to form contrast, either failing to capture the intra-class variance or introducing extra computational overhead. In this work, we propose Prototype-Anchor Contrast (ProAC), a novel prototypical contrastive learning paradigm that strengthens pixel-prototype associations in a simple yet effective fashion. First, ProAC pre-defines class prototypes (serving as cluster centroids) by exploiting the uniformity on the hypersphere in the feature space and thus requires no prototype updating during network optimization, which greatly simplifies the network training process. Second, by treating prototypes as anchors, ProAC builds a novel prototype-to-pixel learning path, where a large amount of negative pixels can naturally be generated to describe rich semantic information without relying on auxiliary sample augmentation techniques. Finally, as a plug-and-play regularization term, ProAC can be attached to most existing segmentation models and assist the network optimization by directly shaping the pixel embedding space. Extensive experiments on different benchmarks show that our ProAC brings an mIoU increase from 1.4% to 2.0% for fully-supervised models and from 0.9% to 6.0% for domain-adaptive models, respectively. It also leads to a gain of mIoU, ranging from 1.8% to 2.7% in more challenging cases, including different resolutions, diverse illuminations and masked scenarios. Qinghua Ren, Shijian Lu, Qirong Mao, Ming Dong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Prototypical Bidirectional Adaptation and Learning for Cross-Domain Semantic SegmentationabstractCross-domain semantic segmentation, which aims to address the distribution shift while adapting from a labeled source domain to an unlabeled target domain, has achieved great progress in recent years. However, most existing work adopts a source-to-target adaptation path, which often suffers from clear class mismatching or class imbalance issues. We design PBAL, a prototypical bidirectional adaptation and learning technique that introduces bidirectional prototype learning and prototypical self-training for optimal inter-domain alignment and adaptation. We perform bidirectional alignments in a complementary and cooperative manner which balances both dominant and tail categories as well as easy and hard samples effectively. In addition, We derive prototypes efficiently from a source-trained classifier, which enables class-aware adaptation as well as synchronous prototype updating and network optimization. Further, we re-examine self-training and introduce prototypical contrast above it which greatly improves inter-domain alignment by promoting better intra-class compactness and inter-class separability in the feature space. Extensive experiments over two widely studied benchmarks show that the proposed PBAL achieves superior domain adaptation performance as compared with the state-of-the-art. Qinghua Ren, Qirong Mao, Shijian Lu |
IEEE Trans. Multim. | 2 |
| 2023 | Joint-Former: Jointly Regularized and Locally Down-sampled Conformer for Semi-supervised Sound Event Detection
Lijian Gao, Qirong Mao, Ming Dong 0001 |
INTERSPEECH | 2 |
| 2023 | TE-KWS: Text-Informed Speech Enhancement for Noise-Robust Keyword SpottingabstractKeyword spotting (KWS) presents a formidable challenge, particularly in high-noise environments. Traditional denoising algorithms that rely solely on speech have difficulty recovering speech that has been severely corrupted by noise. In this investigation, we develop an adaptive text-informed denoising model to bolster reliable keyword identification in the presence of considerable noise degradation. The whole proposed TE-KWS incorporates a tripartite branch structure, where the speech branch (SB) takes noisy speech as input which provides the raw speech information, the alignment branch (AB) accommodates aligned text input which facilitates accurate restoration of the corresponding speech when text with alignment is preserved, and the text branch (TB) handles unaligned text which prompts the model to autonomously learn the alignment between speech and text. To make the proposed denoising model more beneficial for KWS, following the training of the whole model,the alignment branch (AB) is frozen, and the model is fine-tuned by leveraging its speech restoration and forced alignment capabilities. Subsequently, the input for the text branch (TB) is supplanted with designated keywords, and a heavier denoising penalty is applied on the keywords period, thereby explicitly intensifying the speech restoration ability of the model for keywords. Finally, the Combined Adversarial Domain Adaptation (CADA) is implemented to enhance the robustness of KWS with regard to data pre-and post-speech enhancement (SE). Experimental results indicate that our approach not only markedly ameliorates highly corrupted speech, achieving SOTA performance for marginally corrupted speech, but also bolsters the efficacy and generalizability of prevailing mainstream KWS models. Dong Liu 0037, Qirong Mao, Lijian Gao, Qinghua Ren, Zhenghan Chen, Ming Dong 0001 |
ACM Multimedia | 2 |
| 2023 | Multi-branch feature aggregation based on multiple weighting for speaker verification
You-cai Qin, Qinghua Ren, Qirong Mao |
Comput. Speech Lang. | 3 |
| 2023 | A semi-supervised resampling method for class-imbalanced learning
Yu Lu 0017, Yongzhao Zhan 0001, Qirong Mao |
Expert Syst. Appl. | 5 |
| 2023 | Large-scale non-negative subspace clustering based on Nyström approximation
Hongjie Jia, Qize Ren, Longxia Huang, Qirong Mao, Liangjun Wang, Heping Song |
Inf. Sci. | 4 |
| 2023 | Multi-level distance embedding learning for robust acoustic scene classification with unseen devices
Gang Jiang, Zhongchen Ma, Qirong Mao |
Pattern Anal. Appl. | 3 |
| 2023 | Global and local structure preserving nonnegative subspace clustering
Hongjie Jia, Dongxia Zhu, Longxia Huang, Qirong Mao, Liangjun Wang, Heping Song |
Pattern Recognit. | 4 |
| 2023 | Semi-Supervised Clustering Under a "Compact-Cluster" AssumptionabstractSemi-supervised clustering (SSC) aims to improve clustering performance with the support of prior knowledge (i.e., side information). Compared with pairwise constraints, the partial labeling information is more natural to characterize the data distribution in a high level. However, the natural gap between the class information and the clustering is not adequately taken into account in exiting SSC methods when utilizing partial labeling information to guide the clustering procedure. In order to address this problem, we present a compact-cluster assumption for SSC to utilize the partial labeling information via a cluster-splitting technique. Based on this assumption, a general framework, CSSC, is proposed to supervise the traditional clustering with an objective function which is defined by incorporating an item to measure the compact degree of clusters. Furthermore, we provide two effective solutions for Kmeans and spectral clustering within the CSSC framework and derive the corresponding algorithms to seek the optimum number of clusters and their centroids. Corresponding theoretical analyses demonstrate the feasibility and effectivity of the proposed method. Finally, the extensive experiments on eight real-world datasets demonstrate the superiority of our method over other state-of-the-art SSC methods. Yongzhao Zhan 0001, Qirong Mao |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Weighted contrastive learning using pseudo labels for facial expression recognition
Yan Xi, Qirong Mao |
Vis. Comput. | 2 |
| 2022 | Efficient Monaural Speech Separation with Multiscale Time-Delay SamplingabstractRecently, the segmented sample-level modeling approach based on Dual-Path Recurrent Neural Network (DPRNN) has been proved to be effective in Monaural Speech Separation (MSS). Many dual-path networks such as Dual-Path Transformer Network (DPTNet), with a series of improvements to DPRNN, have also improved the separation performance since these methods are effective to process long sequences. However, the receptive fields of these methods are fixed during local and global features learning, which makes it difficult to capture different scale local and global information in long sequences. In this paper, we propose a novel Multiscale Time-Delay Sampling method (MTDS) for the dual-path networks in MSS to learn sequence features from fine to coarse by multiscale time-delay sampling, which effectively integrates different scale local and global information for long sequences. Our experiments on the notable benchmark WSJ0-2mix data corpus result in 21.7dB SDRi and 21.5dB SI-SNRi, which obviously outperforms the state-of-the-arts without data augmentation. Shuang-qing Qian, Lijian Gao, Hongjie Jia, Qirong Mao |
ICASSP | 4 |
| 2022 | Statistical Pyramid Dense Time Delay Neural Network for Speaker VerificationabstractRecently, speaker verification (SV) techniques relay on deep learning frameworks to extract more informative embedding vectors, which greatly improves the accuracy compared with traditional machine learning methods. The well-known x-vector architecture, a time delay neural network (TDNN), is widely adapted for SV tasks. However, most of existing variants rarely combines the global and sub-region context information and suffer from the local receptive field that is engendered by the standard convolutional operation. In this paper, we propose statistical pyramid dense TDNN (SPD-TDNN) with the statistical pyramid pooling module which captures the context information. Specifically, the developed module adaptively exchanges information among contextual regions from different perspectives, which correspond to multiple parallel branches. The statistics collected by the global-region branch are comprised of mean and standard deviation across the time domain to acquire the more global context information. Extensive experiments on the VoxCeleb1&2 datasets demonstrate that the proposed PSD-TDNN outperforms corresponding D-TDNN, D-TDNN-SS and ECAPA-TDNN which achieve the state-of-the-art performances on the SV task, with similar model complexity. Zi-Kai Wan, Qinghua Ren, You-cai Qin, Qirong Mao |
ICASSP | 4 |
| 2022 | DCTCN: Deep Complex Temporal Convolutional Network for Long Time Speech Enhancement
Jigang Ren, Qirong Mao |
INTERSPEECH | 2 |
| 2022 | Adaptive Hierarchical Pooling for Weakly-supervised Sound Event DetectionabstractIn Weakly-supervised Sound Event Detection (WSED), the ground truth of training data contains the presence or absence of each sound event only at the clip-level (i.e., no frame-level annotations). Recently, WSED has been formulated under the multi-instance learning framework, and a critical component within this formulation is the design of the temporal pooling function. In this paper, we propose an adaptive hierarchical pooling (HiPool) for WSED, which combines the advantages of max pooling in audio tagging and weighted average pooling in audio localization through a novel hierarchical structure and learns event-wise optimal pooling functions through continuous relaxation-based joint optimization. Extensive experiments on benchmark datasets show that HiPool outperforms the current pooling methods and greatly improves the performance of WSED. HiPool also has great generality - ready to be plugged into any WSED models. Lijian Gao, Qirong Mao, Ming Dong 0001 |
ACM Multimedia | 3 |
| 2022 | Sparse signal reconstruction via generalized two-stage thresholding
Heping Song, Zehong Ai, Yuping Lai, Hongying Meng, Qirong Mao |
Sci. China Inf. Sci. | 5 |
| 2022 | Weakly Supervised Sentiment-Specific Region Discovery for VSAabstractAbstract Local information has significant contributions to visual sentiment analysis (VSA). Recent studies about local region discovery need manually annotate region location. Affective local information learning and automatic discovery of sentiment-specific region are still the challenges in VSA. In this paper, we propose an end-to-end VSA method for weakly supervised sentiment-specific region discovery. Our method contains two branches: an automatic sentiment-specific region discovery branch and a sentiment analysis branch. In the sentiment-specific region discovery branch, a region proposal network with multiple convolution kernels is proposed to generate candidate affective regions. Then, we design the multiple instance learning (MIL) loss to remove redundant and noisy candidate regions. Finally, the sentiment analysis branch integrates both holistic and localized information obtained in the first branch by feature map coupling for final sentiment classification. Our method automatically discovers sentiment-specific regions by the constraint of MIL loss function without object-level labels. Quantitative and qualitative evaluations on four benchmark affective datasets demonstrate that our proposed method outperforms the state-of-the-art methods. Luoyang Xue, Ang Xu, Qirong Mao, Lijian Gao, Jie Chen 0069 |
Comput. J. | 3 |
| 2022 | Phase sensitive masking-based single channel speech enhancement using conditional generative adversarial network
Sidheswar Routray, Qirong Mao |
Comput. Speech Lang. | 2 |
| 2022 | Convolutional relation network for facial expression recognition in the wild with few-shot learning
Qing Zhu 0002, Qirong Mao, Hongjie Jia, Ocquaye Elias Nii Noi, Juanjuan Tu |
Expert Syst. Appl. | 2 |
| 2022 | A context aware-based deep neural network approach for simultaneous speech denoising and dereverberation
Sidheswar Routray, Qirong Mao |
Neural Comput. Appl. | 2 |
| 2022 | Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition
Ling Zhou 0005, Qirong Mao, Xiaohua Huang 0002, Feifei Zhang 0001, Zhihong Zhang 0001 |
Pattern Recognit. | 2 |
| 2022 | Objective Class-Based Micro-Expression Recognition Under Partial Occlusion Via Region-Inspired Relation Reasoning NetworkabstractMicro-expression recognition (MER) has attracted the attention of many researchers in the past decade. However, occlusion occurs for MER in real-world scenarios. In this paper, a challenging issue in MER that is interesting but unexplored, i.e., occlusion MER, is deeply investigated. First, to research MER under real-world occlusion conditions, synthetic occluded microexpression databases are created by using various community masks. Second, to suppress the influence of occlusion, aRegion-inspiredRelationReasoningNetwork (RRRN) is proposed to model the relations between various facial regions. The RRRN consists of a backbone network, a region-inspired (RI) module and a relation reasoning (RR) module. More specifically, the backbone network aims to extract feature representations from different facial regions, the RI module is designed to compute the adaptive weight from the facial region itself based on the unobstructedness and importance of the region for suppressing the influence of occlusion using an attention mechanism, and the RR module exploits the progressive interactions among these regions by performing graph convolutions. Experiments are conducted on two tasks of MEGC 2018: the holdout-database evaluation task and the composite database evaluation task. Experimental results show that RRRN can be utilized to significantly explore the importance of facial regions and capture the cooperative complementary relationship of facial regions for MER. The results also demonstrate that RRRN outperforms the state-of-the-art approaches, especially with respect to occlusion, where RRRN is more robust. Qirong Mao, Ling Zhou 0005, Wenming Zheng, Xiuyan Shao, Xiaohua Huang 0003 |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | Reproducibility Companion Paper: On Learning Disentangled Representation for Acoustic Event DetectionabstractThis companion paper is provided to describe the major experiments reported in our paper "On Learning Disentangled Representation for Acoustic Event Detection" published in ACM Multimedia 2019. To make the replication of our work easier, we first give an introduction of the computing environment where all of our experiments are conducted. Furthermore, we provide an environmental configuration file to setup the compiling environment and other artifacts including the source code, datasets and the files generated during our experiments. Finally, we summarize the structure and usage of the source code. For more details, please consult the README file in the archive of artifacts on GitHub: https://github.com/mastergofujs/SED_PyTorch. Lijian Gao, Qirong Mao, Ming Dong 0001, Ratna Babu Chinnam, Lucile Sassatelli, Miguel Fabián Romero Rondón, Ujjwal Sharma 0001 |
ACM Multimedia | 2 |
| 2021 | An efficient Nyström spectral clustering algorithm using incomplete Cholesky decomposition
Hongjie Jia, Liangjun Wang, Heping Song, Qirong Mao, Shifei Ding |
Expert Syst. Appl. | 4 |
| 2021 | Cross lingual speech emotion recognition via triple attentive asymmetric convolutional neural networkabstractThe application of cross-corpus for speech emotion recognition (SER) via domain adaptation methods have gain high acknowledgment for developing good robust emotion recognition systems using different corpora or datasets. However, the issue of cross-lingual still remains a challenge in SER and needs more attention to resolve the scenario of applying different language types in both training and testing. In this paper, we propose a triple attentive asymmetric convolutional neural network to address the recognition of emotions for cross-lingual and cross-corpus speech in an unsupervised approach. The proposed method adopts the joint supervision of softmax loss and center loss to learn high power discriminative feature representations for target domain via the use of high quality pseudo-labels. The proposed model uses three attentive convolutional neural networks asymmetrically, where two of the networks are used to artificially label unlabeled target samples as a result of their predictions from training on source labeled samples and the other network is used to obtain salient target discriminative features from the pseudo-labeled target samples. We evaluate our proposed method on three different language types (i.e., English, German, and Italian) data sets. The experimental results indicate that, our proposed method achieves higher prediction accuracy over other state-of-the-art methods. Ocquaye Elias Nii Noi, Qirong Mao, Yanfei Xue, Heping Song |
Int. J. Intell. Syst. | 2 |
| 2021 | Learning to disentangle emotion factors for facial expression recognition in the wildabstractFacial expression recognition (FER) in the wild is a very challenging problem due to different expressions under complex scenario (e.g., large head pose, illumination variation, occlusions, etc.), leading to suboptimal FER performance. Accuracy in FER heavily relies on discovering superior discriminative, emotion-related features. In this paper, we propose an end-to-end module to disentangle latent emotion discriminative factors from the complex factors variables for FER to obtain salient emotion features. The training of proposed method contains two stages. First of all, emotion samples are used to obtain the latent representation using a variational auto-encoder with reconstruction penalization. Furthermore, the latent representation as the input is thrown into a disentangling layer to learn a set of discriminative emotion factors through the attention mechanism (e.g., a Squeeze-and-Excitation block) that encourages to separate emotion-related factors and nonaffective factors. Experimental results on public benchmark databases (RAF-DB and FER2013) show that our approach has remarkable performance in complex scenes than current state-of-the-art methods. Qing Zhu 0002, Lijian Gao, Heping Song, Qirong Mao |
Int. J. Intell. Syst. | 4 |
| 2021 | A survey of micro-expression recognition
Xiuyan Shao, Qirong Mao |
Image Vis. Comput. | 3 |
| 2021 | Latent discriminative representation learning for speaker recognitionabstractExtracting discriminative speaker-specific representations from speech signals and transforming them into fixed length vectors are key steps in speaker identification and verification systems. In this study, we propose a latent discriminative representation learning method for speaker recognition. We mean that the learned representations in this study are not only discriminative but also relevant. Specifically, we introduce an additional speaker embedded lookup table to explore the relevance between different utterances from the same speaker. Moreover, a reconstruction constraint intended to learn a linear mapping matrix is introduced to make representation discriminative. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods based on the Apollo dataset used in the Fearless Steps Challenge in INTERSPEECH2019 and the TIMIT dataset. Duolin Huang, Qirong Mao, Zhongchen Ma, Zhi-shen Zheng, Sidheswar Routray, Ocquaye Elias Nii Noi |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2021 | Erratum to: Latent discriminative representation learning for speaker recognition
Duolin Huang, Qirong Mao, Zhongchen Ma, Zhi-shen Zheng, Sidheswar Routray, Ocquaye Elias Nii Noi |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2021 | Deep face clustering using residual graph convolutional network
Hongjie Jia, Qirong Mao, Liangjun Wang, Heping Song |
Knowl. Based Syst. | 4 |
| 2020 | On Synthesis for Supervised Monaural Speech Separation in Time Domain
Qirong Mao, Dong Liu 0037 |
INTERSPEECH | 2 |
| 2020 | Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech SeparationabstractThe dominant speech separation models are based on complex recurrent or convolution neural network that model speech sequences indirectly conditioning on context, such as passing information through many intermediate states in recurrent neural network, leading to suboptimal separation performance.In this paper, we propose a dual-path transformer network (DPT-Net) for end-to-end speech separation, which introduces direct context-awareness in the modeling for speech sequences.By introduces a improved transformer, elements in speech sequences can interact directly, which enables DPTNet can model for the speech sequences with direct context-awareness.The improved transformer in our approach learns the order information of the speech sequences without positional encodings by incorporating a recurrent neural network into the original transformer.In addition, the structure of dual paths makes our model efficient for extremely long speech sequence modeling.Extensive experiments on benchmark datasets show that our approach outperforms the current state-of-the-arts (20.6 dB SDR on the public WSj0-2mix data corpus). Qirong Mao, Dong Liu 0037 |
INTERSPEECH | 2 |
| 2020 | Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image RetrievalabstractCross-model retrieval has attracted much attention in recent years due to its wide applications. Conventional approaches usually take one modality as query to retrieve relevant data of another modality. In this paper, we devote to an emerging task in cross-modal retrieval, Composing Text and Image to Image Retrieval (CTI-IR), which aims at retrieving images relevant to a query image with text describing desired modifications to the query image. Compared with conventional cross-modal retrieval, the new task is particularly useful for the retrieval that the query image does not perfectly match the user's expectations. Generally, the CTI-IR involves two underlying problems: how to manipulate visual features of the query image specified by the text, and how to model the modality gap between the query and target. Most previous methods focus on solving the second problem. In this paper, we aim to deal with both problems simultaneously in a unified model. Specifically, the proposed method is based on the graph attention network and adversarial learning network, which enjoys several merits. First, the query image and the modification text are constructed in a relation graph for learning text-adaptive representations. Second, semantic contents from the text are injected into the visual features through graph attention. Third, an adversarial loss is incorporated into the conventional cross-modal retrieval loss to learn more discriminative modality invariant representations for CTI-IR. Extensive experiments on three benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art methods. Feifei Zhang 0001, Mingliang Xu 0001, Qirong Mao, Changsheng Xu |
ACM Multimedia | 3 |
| 2020 | Discriminative globality and locality preserving graph embedding for dimensionality reduction
Jianping Gou, Zhang Yi 0001, Jiancheng Lv 0001, Qirong Mao, Yongzhao Zhan 0001 |
Expert Syst. Appl. | 5 |
| 2020 | Latent source-specific generative factor learning for monaural speech separation using weighted-factor autoencoderabstractMuch recent progress in monaural speech separation (MSS) has been achieved through a series of deep learning architectures based on autoencoders, which use an encoder to condense the input signal into compressed features and then feed these features into a decoder to construct a specific audio source of interest. However, these approaches can neither learn generative factors of the original input for MSS nor construct each audio source in mixed speech. In this study, we propose a novel weighted-factor autoencoder (WFAE) model for MSS, which introduces a regularization loss in the objective function to isolate one source without containing other sources. By incorporating a latent attention mechanism and a supervised source constructor in the separation layer, WFAE can learn source-specific generative factors and a set of discriminative features for each source, leading to MSS performance improvement. Experiments on benchmark datasets show that our approach outperforms the existing methods. In terms of three important metrics, WFAE has great success on a relatively challenging MSS case, i.e., speaker-independent MSS. Qirong Mao, You-cai Qin, Shuang-qing Qian, Zhi-shen Zheng |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2020 | NLWSNet: a weakly supervised network for visual sentiment analysis in mislabeled web imagesabstractLarge-scale datasets are driving the rapid developments of deep convolutional neural networks for visual sentiment analysis. However, the annotation of large-scale datasets is expensive and time consuming. Instead, it is easy to obtain weakly labeled web images from the Internet. However, noisy labels still lead to seriously degraded performance when we use images directly from the web for training networks. To address this drawback, we propose an end-to-end weakly supervised learning network, which is robust to mislabeled web images. Specifically, the proposed attention module automatically eliminates the distraction of those samples with incorrect labels by reducing their attention scores in the training process. On the other hand, the special-class activation map module is designed to stimulate the network by focusing on the significant regions from the samples with correct labels in a weakly supervised learning approach. Besides the process of feature learning, applying regularization to the classifier is considered to minimize the distance of those samples within the same class and maximize the distance between different class centroids. Quantitative and qualitative evaluations on well- and mislabeled web image datasets demonstrate that the proposed algorithm outperforms the related methods. Luoyang Xue, Qirong Mao, Xiaohua Huang 0003, Jie Chen 0069 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2020 | Weighted discriminative collaborative competitive representation for robust image classification
Jianping Gou, Lei Wang 0095, Zhang Yi 0001, Yun-Hao Yuan 0001, Weihua Ou, Qirong Mao |
Neural Networks | 6 |
| 2020 | Geometry Guided Pose-Invariant Facial Expression RecognitionabstractDriven by recent advances in human-centered computing, Facial Expression Recognition (FER) has attracted significant attention in many applications. However, most conventional approaches either perform face frontalization on a non-frontal facial image or learn separate classifier for each pose. Different from existing methods, this paper proposes an end-to-end deep learning model that allows to simultaneous facial image synthesis and pose-invariant facial expression recognition by exploiting shape geometry of the face image. The proposed model is based on generative adversarial network (GAN) and enjoys several merits. First, given an input face and a target pose and expression designated by a set of facial landmarks, an identity-preserving face can be generated through guiding by the target pose and expression. Second, the identity representation is explicitly disentangled from both expression and pose variations through the shape geometry delivered by facial landmarks. Third, our model can automatically generate face images with different expressions and poses in a continuous way to enlarge and enrich the training set for the FER task. Our approach is demonstrated to perform well when compared with state-of-the-art algorithms on both controlled and in-the-wild benchmark datasets including Multi-PIE, BU-3DFE, and SFEW. Feifei Zhang 0001, Tianzhu Zhang 0001, Qirong Mao, Changsheng Xu |
IEEE Trans. Image Process. | 3 |
| 2020 | A Unified Deep Model for Joint Facial Expression Recognition, Face Synthesis, and Face AlignmentabstractFacial expression recognition, face synthesis, and face alignment are three coherently related tasks and can be solved in a joint framework. To achieve this goal, in this paper, we propose a novel end-to-end deep learning model by exploiting the expression code, geometry code and generated data jointly for simultaneous pose-invariant facial expression recognition, face image synthesis, and face alignment. The proposed deep model enjoys several merits. First, to the best of our knowledge, this is the first work to address these three tasks jointly in a unified deep model to complement and enhance each other. Second, the proposed model can effectively disentangle the global and local identity representation from different expression and geometry codes. As a result, it can automatically generate facial images with different expressions under arbitrary geometry codes. Third, these three tasks can further boost their performance for each other via our model. Extensive experimental results on three standard benchmarks demonstrate that the proposed deep model performs favorably against state-of-the-art methods on the three tasks. Feifei Zhang 0001, Tianzhu Zhang 0001, Qirong Mao, Changsheng Xu |
IEEE Trans. Image Process. | 3 |
| 2019 | Dual-Inception Network for Cross-Database Micro-Expression RecognitionabstractThis paper presents the technique for our contribution to 2019 Micro-Expression Grand Challenge (MEGC 2019). One sub-challenge of MEGC 2019 named Cross-Database (Cross-DB) challenge aims to classify three main classes (Negative, Positive and Surprise) in the task of Composite Database Evaluation (CDE). Our proposed method utilizes Inception technique to overcome the challenge for the cross-database micro-expression recognition and can be divided into three steps. (1) In the preprocessing stage, onset and mid-position frames of each micro-expression sample are selected for the feature extraction. (2) TV-L1 optical flow features are calculated by the two frames obtained in the first step. (3) The horizontal and vertical components of TV-L1 optical flow features are fed to a Dual-Inception network for the micro-expression recognition. Our experiment results on three benchmark databases show that our proposed mechanism archives the overall unweighted F1 score (UF1) of 0.7322 and unweighted average recall (UAR) of 0.7278, which significantly outperform those metrics of the baseline method ( UF1: 0.5882, UAR: 0.5785). Code is publicly available on GitHub: https://github.com/xly135846/MEGC2019. Ling Zhou 0004, Qirong Mao, Luoyang Xue |
FG | 2 |
| 2019 | Discriminative Group Collaborative Competitive Representation for Visual ClassificationabstractIn pattern recognition, the representation-based classification (RBC) has attracted much attention recently. As a representative one of RBC, collaborative representation-based classification (CRC) and its variants have achieved promising classification performance in many visual classification tasks. However, most of the CRC methods cannot directly consider the class discrimination information of data that is very important for classification. To fully use the class discrimination information, we propose a novel discriminative group collaborative competitive representation-based classification method (DGCCR) in this paper. In the designed DGCCR model, the discriminative competitive relationships of classes, the discriminative decorrelations among classes and the weighted class-specific group constraints are simultaneously taken into account for strengthening the power of pattern discrimination. Experiments on three visual classification data sets demonstrate that the proposed DGCCR out-performs state-of-the-art RBC methods. Jianping Gou, Lei Wang 0095, Zhang Yi 0001, Yun-Hao Yuan 0001, Weihua Ou, Qirong Mao |
ICME | 6 |
| 2019 | On Learning Disentangled Representation for Acoustic Event DetectionabstractPolyphonic Acoustic Event Detection (AED) is a challenging task as the sounds are mixed with the signals from different events, and the features extracted from the mixture do not match well with features calculated from sounds in isolation, leading to suboptimal AED performance. In this paper, we propose a supervised β-VAE model for AED, which adds a novel event-specific disentangling loss in the objective function of disentangled learning. By incorporating either latent factor blocks or latent attention in disentangling, supervised β-VAE learns a set of discriminative features for each event. Extensive experiments on benchmark datasets show that our approach outperforms the current state-of-the-arts (top-1 performers in the Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 AED challenge). Supervised β-VAE has great success in challenging AED tasks with a large variety of events and imbalanced data. Lijian Gao, Qirong Mao, Ming Dong 0001, Yu Jing, Ratna Babu Chinnam |
ACM Multimedia | 2 |
| 2019 | Triple attention network for sentimental visual question answering
Nelson Ruwa, Qirong Mao, Heping Song, Hongjie Jia, Ming Dong 0001 |
Comput. Vis. Image Underst. | 2 |
| 2019 | Two-phase probabilistic collaborative representation-based classification
Jianping Gou, Lei Wang 0095, Bing Hou, Jiancheng Lv 0001, Yun-Hao Yuan 0001, Qirong Mao |
Expert Syst. Appl. | 6 |
| 2019 | Dictionary-induced least squares framework for multi-view dimensionality reduction with multi-manifold embeddingsabstractThis study proposes a novel dimensionality reduction (DR) method for multi‐view datasets. The principal component analysis (PCA) idea of minimising least squares reconstruction errors is extended to consider both data distribution and penalty weights called dictionary to recover outliers free global structures from missing and noisy data points. In this way, PCA is viewed as a special instance of the authors’ proposed dictionary induced least squares framework (DLS). Furthermore, to appropriately handle multi‐view DR, we combine the DLS with multiple manifold embeddings (DLSME). Therefore it can obtain lower projections while maintaining a balance between preserving global structures with DLS and local structures with multi‐manifold embeddings. Extensive experiments on object and face recognition datasets verify that the DLS achieves better classification results with lower dimensional projections than PCA. Also, on many multi‐view datasets of visual recognition and web image annotation, the DLSME method demonstrates more effectiveness than Graph‐Laplacian PCA (gLPCA), robust PCA‐optimal mean, canonical correlation analysis (CCA), bilinear models (BLM), neighbourhood preserving embedding, locality preserving projections, and locality sensitive discriminant analysis. Timothy Apasiba Abeo, Xiangjun Shen, Jianping Gou, Qirong Mao, Bing-Kun Bao |
IET Comput. Vis. | 4 |
| 2019 | Several robust extensions of collaborative representation for image classification
Jianping Gou, Bing Hou, Weihua Ou, Qirong Mao, Hebiao Yang |
Neurocomputing | 4 |
| 2019 | Affective question answering on video
Nelson Ruwa, Qirong Mao, Liangjun Wang, Jianping Gou |
Neurocomputing | 2 |
| 2019 | Mood-aware visual question answering
Nelson Ruwa, Qirong Mao, Liangjun Wang, Jianping Gou, Ming Dong 0001 |
Neurocomputing | 2 |
| 2019 | Multimodal shared features learning for emotion recognition by enhanced sparse local discriminative canonical correlation analysis
Jiamin Fu, Qirong Mao, Juanjuan Tu, Yongzhao Zhan 0001 |
Multim. Syst. | 2 |
| 2019 | A Local Mean Representation-based K-Nearest Neighbor ClassifierabstractK -nearest neighbor classification method (KNN), as one of the top 10 algorithms in data mining, is a very simple and yet effective nonparametric technique for pattern recognition. However, due to the selective sensitiveness of the neighborhood size k , the simple majority vote, and the conventional metric measure, the KNN-based classification performance can be easily degraded, especially in the small training sample size cases. In this article, to further improve the classification performance and overcome the main issues in the KNN-based classification, we propose a local mean representation-based k -nearest neighbor classifier (LMRKNN). In the LMRKNN, the categorical k -nearest neighbors of a query sample are first chosen to calculate the corresponding categorical k -local mean vectors, and then the query sample is represented by the linear combination of the categorical k -local mean vectors; finally, the class-specific representation-based distances between the query sample and the categorical k -local mean vectors are adopted to determine the class of the query sample. Extensive experiments on many UCI and KEEL datasets and three popular face databases are carried out by comparing LMRKNN to the state-of-art KNN-based methods. The experimental results demonstrate that the proposed LMRKNN outperforms the related competitive KNN-based methods with more robustness and effectiveness. Jianping Gou, Wenmo Qiu, Zhang Yi 0001, Yong Xu 0001, Qirong Mao, Yongzhao Zhan 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2019 | An emotion-based responding model for natural language conversation
Qirong Mao, Liangjun Wang, Nelson Ruwa, Jianping Gou, Yongzhao Zhan 0001 |
World Wide Web | 2 |
| 2018 | Joint Pose and Expression Modeling for Facial Expression RecognitionabstractFacial expression recognition (FER) is a challenging task due to different expressions under arbitrary poses. Most conventional approaches either perform face frontalization on a non-frontal facial image or learn separate classifiers for each pose. Different from existing methods, in this paper, we propose an end-to-end deep learning model by exploiting different poses and expressions jointly for simultaneous facial image synthesis and pose-invariant facial expression recognition. The proposed model is based on generative adversarial network (GAN) and enjoys several merits. First, the encoder-decoder structure of the generator can learn a generative and discriminative identity representation for face images. Second, the identity representation is explicitly disentangled from both expression and pose variations through the expression and pose codes. Third, our model can automatically generate face images with different expressions under arbitrary poses to enlarge and enrich the training set for FER. Quantitative and qualitative evaluations on both controlled and in-the-wild datasets demonstrate that the proposed algorithm performs favorably against state-of-the-art methods. Feifei Zhang 0001, Tianzhu Zhang 0001, Qirong Mao, Changsheng Xu |
CVPR | 3 |
| 2018 | Facial Expression Recognition in the Wild: A Cycle-Consistent Adversarial Attention Transfer ApproachabstractFacial expression recognition (FER) is a very challenging problem due to different expressions under arbitrary poses. Most conventional approaches mainly perform FER under laboratory controlled environment. Different from existing methods, in this paper, we formulate the FER in the wild as a domain adaptation problem, and propose a novel auxiliary domain guided Cycle-consistent adversarial Attention Transfer model (CycleAT) for simultaneous facial image synthesis and facial expression recognition in the wild. The proposed model utilizes large-scale unlabeled web facial images as an auxiliary domain to reduce the gap between source domain and target domain based on generative adversarial networks (GAN) embedded with an effective attention transfer module, which enjoys several merits. First, the GAN-based method can automatically generate labeled facial images in the wild through harnessing information from labeled facial images in source domain and unlabeled web facial images in auxiliary domain. Second, the class-discriminative spatial attention maps from the classifier in source domain are leveraged to boost the performance of the classifier in target domain. Third, it can effectively preserve the structural consistency of local pixels and global attributes in the synthesized facial images through pixel cycle-consistency and discriminative loss. Quantitative and qualitative evaluations on two challenging in-the-wild datasets demonstrate that the proposed model performs favorably against state-of-the-art methods. Feifei Zhang 0001, Tianzhu Zhang 0001, Qirong Mao, Ling-Yu Duan, Changsheng Xu |
ACM Multimedia | 3 |
| 2018 | Cascaded Multi-level Transformed Dirichlet Process for Multi-pose Facial Expression RecognitionabstractAs an essential way of human emotional behavior understanding, facial expression recognition (FER) has been studied extensively in recent years. However, the existing methods of FER are typically based on near-frontal face data. High-recognition accuracy for multi-pose FER continues to be a challenge. In this paper, we present a novel cascaded multi-level Transformed Dirichlet Process (cml-TDP) model for multi-pose FER. The top-level structure of the cml-TDP model has been carefully designed to make coarse-to-fine prediction, and the outputs of the model are fused for robust and accurate estimation at each level. There are three primary merits to cml-TDP. First, pose is explicitly introduced into cml-TDP so that separate training and parameter tuning for each pose is not required. Second, cml-TDP describes an image by its detected positions and appearance features to implicitly construct geometric constraints. Third, cml-TDP can learn an intermediate facial expression representation subject to geometric constraints. By sharing the pool of spatially coherent features over expressions and poses, we provide a scalable solution for multi-pose FER. The proposed model has been evaluated on two benchmark databases, BU-3DFE and RAFD, and achieved 79.33% and 75.00% FER accuracy on these two datasets, respectively, which has outperformed current state-of-the-art FER methods. Qirong Mao, Feifei Zhang 0001, Liangjun Wang, Sidian Luo, Ming Dong 0001 |
Comput. J. | 1 |
| 2018 | Two-phase linear reconstruction measure-based classification for face recognition
Jianping Gou, Yong Xu 0001, David Zhang 0001, Qirong Mao, Lan Du 0002, Yongzhao Zhan 0001 |
Inf. Sci. | 4 |
| 2018 | Affective rating ranking based on face images in arousal-valence dimensional spaceabstractIn dimensional affect recognition, the machine learning methods, which are used to model and predict affect, are mostly classification and regression. However, the annotation in the dimensional affect space usually takes the form of a continuous real value which has an ordinal property. The aforementioned methods do not focus on taking advantage of this important information. Therefore, we propose an affective rating ranking framework for affect recognition based on face images in the valence and arousal dimensional space. Our approach can appropriately use the ordinal information among affective ratings which are generated by discretizing continuous annotations. Specifically, we first train a series of basic cost-sensitive binary classifiers, each of which uses all samples relabeled according to the comparison results between corresponding ratings and a given rank of a binary classifier. We obtain the final affective ratings by aggregating the outputs of binary classifiers. By comparing the experimental results with the baseline and deep learning based classification and regression methods on the benchmarking database of the AVEC 2015 Challenge and the selected subset of SEMAINE database, we find that our ordinal ranking method is effective in both arousal and valence dimensions. Guopeng Xu, Haitang Lu, Feifei Zhang 0001, Qirong Mao |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2018 | Discriminative self-adapted locality-sensitive sparse representation for video semantic analysis
Jianping Gou, Yongzhao Zhan 0001, Qirong Mao |
Multim. Tools Appl. | 4 |
| 2018 | Spatially Coherent Feature Learning for Pose-Invariant Facial Expression RecognitionabstractFeature learning has enjoyed much attention and achieved good performance in recent studies of image processing. Unlike the required training conditions often assumed there, far less labeled data is available for training emotion classification systems. In addition, current feature learning is typically performed on an entire face image without considering the dependency between features. These approaches ignore the fact that faces are structured and the neighboring features are dependent. Thus, the learned features lack the power to describe visually coherent facial images. Our method is therefore designed with the goal of simplifying the problem domain by removing expression-irrelevant factors from the input images, with a key region-based mechanism, which is an effort to reduce the amount of data required to effectively train the feature-learning methods. Meanwhile, we can construct geometric constraints between the key regions and its detected positions. To this end, we introduce a Spatially Coherent featurelearning method for Pose-invariant Facial Expression Recognition (SC-PFER). In our model, we first perform face frontalization through a 3D pose-normalization technique, which could normalize poses while preserving the identity information through synthesizing frontal faces for facial images with arbitrary views. Subsequently, we select a sequence of key regions around 51 key points in the synthetic frontal face images for efficient unsupervised feature learning. Finally, we introduce a linkage structure over the learning-based features and the corresponding geometry information of each key region to encode the dependencies of the regions. Our method, on the whole, does not require training multiple models for each specific pose and avoids separating training and parameter tuning for each pose. The proposed framework has been evaluated on two benchmark databases, BU-3DFE and SFEW, for pose-invariant Facial Expression Recognition (FER). The experimental results demonstrate that our algorithm outperforms current state-of-the-art FER methods. Specifically, our model achieves an improvement of 1.72% and 1.11% FER accuracy, on average, on BU-3DFE and SFEW, respectively. Feifei Zhang 0001, Qirong Mao, Xiangjun Shen, Yongzhao Zhan 0001, Ming Dong 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | A Multi-local Means Based Nearest Neighbor ClassifierabstractIn this paper, we propose a multi-local means based nearest neighbor classifier (MLMNN). In the MLMNN, k categorical nearest neighbors of a query sample are first found and used to calculate the corresponding k categorical multi-local mean vectors which can represent different local class-specific sample distributions. Then, the query sample is represented by a linear combination of k categorical local mean vectors and the representation coefficient of each local mean vector as the contribution to representing and classifying the query sample is obtained. Finally, the class-specific representation-based distance (i.e. reconstruction residual) between the query sample and k categorical multi-local mean vectors is adopted to determine the class label of the query sample. The experimental results on three popular face databases show that the proposed MLMNN method outperforms the related competitive KNN-based methods. Jianping Gou, Wenmo Qiu, Qirong Mao, Yongzhao Zhan 0001, Xiangjun Shen, Yunbo Rao |
ICTAI | 3 |
| 2017 | Unsupervised domain adaptation for speech emotion recognition using PCANet
Zhengwei Huang, Wentao Xue, Qirong Mao, Yongzhao Zhan 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Learning emotion-discriminative and domain-invariant features for domain adaptation in speech emotion recognition
Qirong Mao, Guopeng Xu, Wentao Xue, Jianping Gou, Yongzhao Zhan 0001 |
Speech Commun. | 1 |
| 2017 | Hierarchical Bayesian Theme Models for Multipose Facial Expression RecognitionabstractAs an essential way of human emotional behavior understanding, facial expression recognition (FER) has attracted a great deal of attention in multimedia research. Most of studies are conducted in a “lab-controlled” environment, and their real-world performance degenerates greatly due to factors such as head pose variations. In this paper, we propose a pose-based hierarchical Bayesian theme model to address challenging issues in multipose FER. Local appearance features and global geometry information are combined in our model to learn an intermediate face representation before recognizing expressions. By sharing a pool of features with various poses, our model provides a unified solution for multipose FER, bypassing the separate training and parameter tuning for each pose, and thus is scalable to a large number of poses. Experiments on both benchmark facial expression databases and Internet images show the superior/highly competitive performance of our system when compared with the current state of the art. Qirong Mao, Qiyu Rao, Ming Dong 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Domain adaptation for speech emotion recognition by sharing priors between related source and target classesabstractIn speech emotion recognition (SER), speech data is usually captured from different scenarios, which often leads to significant performance degradation due to the inherent mismatch between training and test set. To cope with this problem, we propose a domain adaptation method called Sharing Priors between Related Source and Target classes (SPRST) based on a two-layer neural network. The classifier parameters, namely the weights of the second layer, are imposed the common priors between the related classes, so that the classes with few labeled data in target domain can borrow knowledge from the related classes in source domain. The method is evaluated on the INTERSPEECH 2009 Emotion Challenge two-class task. Experimental results show that our approach significantly improves the performance when only a small number of target labeled instances are available. Qirong Mao, Wentao Xue, Qiyu Rao, Feifei Zhang 0001, Yongzhao Zhan 0001 |
ICASSP | 1 |
| 2016 | Multi-pose Facial Expression Recognition Using Transformed Dirichlet ProcessabstractDriven by recent advances in human-centered computing, Facial Expression Recognition (FER) has attracted significant attention in many applications. In this paper, we propose a novel graphical model, multi-level Transformed Dirichlet Process (ml-TDP), for multi-pose FER. In our approach, pose is explicitly introduced into ml-TDP so that separate training and parameter tuning for each pose is not required. In addition, ml-TDP can learn an intermediate facial expression representation subject to geometric constraints. By sharing the pool of spatially-coherent features over expressions and poses, we provide a scalable solution for multi-pose FER. Extensive experimental result on benchmark facial expression databases shows the superior performance of ml-TDP. Feifei Zhang 0001, Qirong Mao, Ming Dong 0001, Yongzhao Zhan 0001 |
ACM Multimedia | 2 |
| 2016 | Collaborative Q-Learning Based Routing Control in Unstructured P2P Networks
Xiangjun Shen, Jianping Gou, Qirong Mao, Zhengjun Zha, Ke Lu 0002 |
MMM (1) | 4 |
| 2016 | Pose-robust feature learning for facial expression recognition
Feifei Zhang 0001, Qirong Mao, Jianping Gou, Yongzhao Zhan 0001 |
Frontiers Comput. Sci. | 3 |
| 2015 | Multi-pose facial expression recognition based on SURF boostingabstractToday Human Computer Interaction (HCI) is one of the most important topics in machine vision and image processing fields. The ability to handle multi-pose facial expressions is important for computers to understand affective behavior under less constrained environment. In this paper, we propose a SURF (Speeded-Up Robust Features) boosting framework to address challenging issues in multi-pose facial expression recognition (FER). Local SURF features from different overlapping patches are selected by boosting in our model to focus on more discriminable representations of facial expression. And this paper proposes a novel training step during boosting. The experiments using the proposed method demonstrate favorable results on RaFD and KDEF databases. Qiyu Rao, Xing Qu, Qirong Mao, Yongzhao Zhan 0001 |
ACII | 3 |
| 2015 | Learning speech emotion features by joint disentangling-discriminationabstractSpeech plays an important part in human-computer interaction. As a major branch of speech processing, speech emotion recognition (SER) has drawn much attention of researchers. Excellent discriminant features are of great importance in SER. However, emotion-specific features are commonly mixed with some other features. In this paper, we introduce an approach to pull apart these two parts of features as much as possible. First we employ an unsupervised feature learning framework to achieve some rough features. Then these rough features are further fed into a semi-supervised feature learning framework. In this phase, efforts are made to disentangle the emotion-specific features and some other features by using a novel loss function, which combines reconstruction penalty, orthogonal penalty, discriminative penalty and verification penalty. Orthogonal penalty is utilized to disentangle emotion-specific features and other features. The discriminative penalty enlarges inter-emotion variations, while the verification penalty reduces the intra-emotion variations. Evaluations on the FAU Aibo emotion database show that our approach can improve the speech emotion classification performance. Wentao Xue, Zhengwei Huang, Qirong Mao |
ACII | 4 |
| 2015 | A Video Semantic Analysis Method Based on Kernel Discriminative Sparse Representation and Weighted KNNabstractTo improve the video semantic analysis for video surveillance, a new video semantic analysis method based on the kernel discriminative sparse representation (KSVD) and weighted K nearest neighbors (KNN) is proposed in this paper. A discriminative model is built by introducing a kernel discriminative function to the KSVD dictionary optimization algorithm, mapping the sparse representation features into a high-dimensional space. The optimal dictionary is then generated and applied to compute the sparse representations of video features. For video semantic analysis, a weighted KNN algorithm based on the optimal sparse representation is proposed. In the algorithm, a kernel function is introduced to establish discrimination about sparse representation features and the classification vote result is weighted, the purpose of which is to improve the accuracy and rationality for video semantic analysis. The experimental results show that the proposed method significantly improves the discrimination of sparse representation features when compared with the traditional KSVD-based support vector machine method. The method can effectively detect the concept and event, which can be potentially useful for improving the video surveillance. Yongzhao Zhan 0001, Shan Dai, Qirong Mao, Lu Liu 0001, Wei Sheng |
Comput. J. | 3 |
| 2015 | Speech emotion recognition with unsupervised feature learningabstractEmotion-based features are critical for achieving high performance in a speech emotion recognition (SER) system. In general, it is difficult to develop these features due to the ambiguity of the ground-truth. In this paper, we apply several unsupervised feature learning algorithms (including K -means clustering, the sparse auto-encoder, and sparse restricted Boltzmann machines), which have promise for learning task-related features by using unlabeled data, to speech emotion recognition. We then evaluate the performance of the proposed approach and present a detailed analysis of the effect of two important factors in the model setup, the content window size and the number of hidden layer nodes. Experimental results show that larger content windows and more hidden nodes contribute to higher performance. We also show that the two-layer network cannot explicitly improve performance compared to a single-layer network. Zhengwei Huang, Wentao Xue, Qirong Mao |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2015 | Using Kinect for real-time emotion recognition via facial expressionsabstractEmotion recognition via facial expressions (ERFE) has attracted a great deal of interest with recent advances in artificial intelligence and pattern recognition. Most studies are based on 2D images, and their performance is usually computationally expensive. In this paper, we propose a real-time emotion recognition approach based on both 2D and 3D facial expression features captured by Kinect sensors. To capture the deformation of the 3D mesh during facial expression, we combine the features of animation units (AUs) and feature point positions (FPPs) tracked by Kinect. A fusion algorithm based on improved emotional profiles (IEPs) and maximum confidence is proposed to recognize emotions with these real-time facial expression features. Experiments on both an emotion dataset and a real-time video show the superior performance of our method. Qirong Mao, Yongzhao Zhan 0001, Xiangjun Shen |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2015 | A semi-supervised incremental learning method based on adaptive probabilistic hypergraph for video semantic detection
Yongzhao Zhan 0001, Jiayao Sun, DeJiao Niu, Qirong Mao, Jianping Fan 0001 |
Multim. Tools Appl. | 4 |
| 2014 | Speech Emotion Recognition Using CNNabstractDeep learning systems, such as Convolutional Neural Networks (CNNs), can infer a hierarchical representation of input data that facilitates categorization. In this paper, we propose to learn affect-salient features for Speech Emotion Recognition (SER) using semi-CNN. The training of semi-CNN has two stages. In the first stage, unlabeled samples are used to learn candidate features by contractive convolutional neural network with reconstruction penalization. The candidate features, in the second step, are used as the input to semi-CNN to learn affect-salient, discriminative features using a novel objective function that encourages the feature saliency, orthogonality and discrimination. Our experiment results on benchmark datasets show that our approach leads to stable and robust recognition performance in complex scenes (e.g., with speaker and environment distortion), and outperforms several well-established SER features. Zhengwei Huang, Ming Dong 0001, Qirong Mao, Yongzhao Zhan 0001 |
ACM Multimedia | 3 |
| 2014 | An SVM-AdaBoost facial expression recognition system
Ebenezer Owusu, Yongzhao Zhan 0001, Qirong Mao |
Appl. Intell. | 3 |
| 2014 | A neural-AdaBoost based facial expression recognition systemabstractThis study improves the recognition accuracy and execution time of facial expression recognition system. Various techniques were utilized to achieve this. The face detection component is implemented by the adoption of Viola–Jones descriptor. The detected face is down-sampled by Bessel transform to reduce the feature extraction space to improve processing time then. Gabor feature extraction techniques were employed to extract thousands of facial features which represent various facial deformation patterns. An AdaBoost-based hypothesis is formulated to select a few hundreds of the numerous extracted features to speed up classification. The selected features were fed into a well designed 3-layer neural network classifier that is trained by a back-propagation algorithm. The system is trained and tested with datasets from JAFFE and Yale facial expression databases. An average recognition rate of 96.83% and 92.22% are registered in JAFFE and Yale databases, respectively. The execution time for a 100 × 100 pixel size is 14.5 ms. The general results of the proposed techniques are very encouraging when compared with others. Ebenezer Owusu, Yongzhao Zhan 0001, Qirong Mao |
Expert Syst. Appl. | 3 |
| 2014 | An SVM-AdaBoost-based face detection systemabstractFace detection is the first significant step in face recognition and many computer vision applications. The goal of this work was to improve detection accuracy as well as reducing the execution time. Images are pre-processed, scaled and normalised with the discrete cosine transform. Gabor feature extraction techniques were employed to extract thousands of facial vectors. An AdaBoost-based feature selection tool was formulated to select a few hundreds of the Gabor wavelets. These vectors representing significant salient local features are used as input vectors to a support vector machine classifier. The classifier is trained and becomes capable of detecting faces. A detection rate of 97.6% with acceptable false positives was registered with a test set of 507 faces. The execution time of a pixel of size 320 × 240 is 0.0285 s, which is very promising. A comparative evaluation of receiver operating characteristic (ROC) curves of different detectors on FDDB set shows that the proposed method is very effective. Ebenezer Owusu, Yongzhao Zhan 0001, Qirong Mao |
J. Exp. Theor. Artif. Intell. | 3 |
| 2014 | Learning Salient Features for Speech Emotion Recognition Using Convolutional Neural NetworksabstractAs an essential way of human emotional behavior understanding, speech emotion recognition (SER) has attracted a great deal of attention in human-centered signal processing. Accuracy in SER heavily depends on finding good affect- related , discriminative features. In this paper, we propose to learn affect-salient features for SER using convolutional neural networks (CNN). The training of CNN involves two stages. In the first stage, unlabeled samples are used to learn local invariant features (LIF) using a variant of sparse auto-encoder (SAE) with reconstruction penalization. In the second step, LIF is used as the input to a feature extractor, salient discriminative feature analysis (SDFA), to learn affect-salient, discriminative features using a novel objective function that encourages feature saliency, orthogonality, and discrimination for SER. Our experimental results on benchmark datasets show that our approach leads to stable and robust recognition performance in complex scenes (e.g., with speaker and language variation, and environment distortion) and outperforms several well-established SER features. Qirong Mao, Ming Dong 0001, Zhengwei Huang, Yongzhao Zhan 0001 |
IEEE Trans. Multim. | 1 |
| 2013 | Regularized least squares fisher linear discriminant with applications to image recognition
Xiaobo Chen 0001, Jian Yang 0003, Qirong Mao, Fei Han 0001 |
Neurocomputing | 3 |
| 2013 | Speaker-independent speech emotion recognition by fusion of functional and accompanying paralanguage featuresabstractFunctional paralanguage includes considerable emotion information, and it is insensitive to speaker changes. To improve the emotion recognition accuracy under the condition of speaker-independence, a fusion method combining the functional paralanguage features with the accompanying paralanguage features is proposed for the speaker-independent speech emotion recognition. Using this method, the functional paralanguages, such as laughter, cry, and sigh, are used to assist speech emotion recognition. The contributions of our work are threefold. First, one emotional speech database including six kinds of functional paralanguage and six typical emotions were recorded by our research group. Second, the functional paralanguage is put forward to recognize the speech emotions combined with the accompanying paralanguage features. Third, a fusion algorithm based on confidences and probabilities is proposed to combine the functional paralanguage features with the accompanying paralanguage features for speech emotion recognition. We evaluate the usefulness of the functional paralanguage features and the fusion algorithm in terms of precision, recall, and F1-measurement on the emotional speech database recorded by our research group. The overall recognition accuracy achieved for six emotions is over 67% in the speaker-independent condition using the functional paralanguage features. Qirong Mao, Xiao-lei Zhao, Zhengwei Huang, Yongzhao Zhan 0001 |
J. Zhejiang Univ. Sci. C | 1 |
| 2008 | Application Research of Ontology in E-Learning EnvironmentabstractIn this paper, awareness and situation ontology model is presented to process awareness and learning situations in E-Learning environment. By this model, knowledge domain, awareness information and learning situations can be described reasonably and efficiently. In addition, reasoning rules for awareness and situation are given in this model. According to these reasoning rules, the learnerpsilas awareness information and learning situations can be concluded. We design and realize an E-Learning system based on this ontology model. The experiment results show that learner's learning efficiency is raised greatly after awareness and situation ontology model is adopted. Qirong Mao, Yongzhao Zhan 0001 |
CW | 2 |
| 2005 | The shared knowledge space model in Web-based cooperative learning coalitionabstractIn order to make learners share learning data in multi-Web sites, in this paper, based on nested knowledge space model, Web-based collaborative learning coalition is defined, and the shared knowledge space model is put forward, by which the learning data in multi-Web sites can be organized together and formed the shared knowledge space. Each Web site in the coalition is able to customize the shared knowledge space according to its needs, then the nested customized knowledge space is formed in the local site integrated with the local learning data. Furthermore, the management and consistency maintenance strategy for the shared knowledge space and the nested customized knowledge space is proposed. In the end, the Web-based collaborative learning coalition is implemented according to the model and strategy mentioned above, and the application effect shows that the data in multi-Web sites in the coalition is shared and their different needs for learning data are met. Qirong Mao, Yongzhao Zhan 0001, Shunlin Song |
CSCWD (2) | 1 |
| 2004 | Design and Simulation of Multicast Routing Protocol for Mobile Internet
Guangsheng Li, Qirong Mao, Yongzhao Zhan 0001, Yibin Hou |
APWeb | 3 |
| 2004 | DRMR: Dynamic-Ring-Based Multicast Routing Protocol for Ad Hoc Networks
Guangsheng Li, Yongzhao Zhan 0001, Qirong Mao, Yibin Hou |
J. Comput. Sci. Technol. | 4 |