Qingfeng Liu

dblp:95/492 · also Qing-Feng Liu · DBLP profile ↗
← Back
45ranked-venue papers
10as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 22 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-authorSecurity and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
abstract
The introduction of diffusion models has brought significant advances to the field of audio-driven talking head generation. However, the extremely slow inference speed severely limits the practical implementation of diffusion-based talking head generation models. In this study, we propose READ, a real-time diffusion-transformer-based talking head generation framework. Our approach first learns a spatiotemporal highly compressed video latent space via a temporal VAE, significantly reducing the token count to accelerate generation. To achieve better audio-visual alignment within this compressed latent space, a pre-trained Speech Autoencoder (SpeechAE) is proposed to generate temporally compressed speech latent codes corresponding to the video latent space. These latent representations are then modeled by a carefully designed Audio-to-Video Diffusion Transformer (A2V-DiT) backbone for efficient talking head synthesis. Furthermore, to ensure temporal consistency and accelerated inference in extended generation, we propose a novel asynchronous noise scheduler (ANS) for both the training and inference processes of our framework. The ANS leverages asynchronous add-noise and asynchronous motion-guided generation in the latent space, ensuring consistency in generated video clips. Experimental results demonstrate that READ outperforms state-of-the-art methods by generating competitive talking head videos with significantly reduced runtime, achieving an optimal balance between quality and speed while maintaining robust metric stability in long-time generation.
Yuzhe Weng, Jun Du 0002, Cong Liu 0006, Jianqing Gao, Qingfeng Liu
AAAI10
2026 Three-stage modular speaker diarization collaborating with front-end techniques in the CHiME-8 NOTSOFAR-1 challenge
abstract
We propose a modular speaker diarization framework that collaborates with front-end techniques in a three-stage process, designed for the challenging CHiME-8 NOTSOFAR-1 acoustic environment. The framework leverages the strengths of deep learning based speech separation systems and traditional speech signal processing techniques to provide more accurate initializations for the Neural Speaker Diarization (NSD) system at each stage, thereby enhancing the performance of a single-channel NSD system. Firstly, speaker overlap detection and Continuous Speech Separation (CSS) are applied to the multichannel speech to obtain clearer single-speaker speech segments for the Clustering-based Speaker Diarization (CSD), followed by the first NSD decoding. Next, the binary speaker masks from the first decoding are used to initialize a complex Angular Center Gaussian Mixture Model (cACGMM) to estimate speaker masks on the multi-channel speech. Using Mask-to-VAD post-processing techniques, we achieve per-speaker speech activity with reduced speaker error (SpkErr), followed by a second NSD decoding. Finally, the second decoding results are used to Guide Source Separation (GSS) to produce per-speaker speech segments. Short utterances containing one word or fewer are filtered, and the remaining speech segments are re-clustered for the final NSD decoding. We present evaluation results progressively explored from the CHiME-8 NOTSOFAR-1 challenge, demonstrating the effectiveness of our modular diarization system and its contribution to improving speech recognition performance. The code will be open-sourced at https://github.com/rywang99/USTC-NERCSLIP_CHiME-8 . • We propose a novel three-stage modular speaker diarization framework integrating front-end cues. • CSS streams extend single-speaker segments for CSD clustering and NSD decoding initialization. • Spatial information is utilized to progressively reduce SpkErr and improve ASR performance.
Ruoyu Wang 0029, Jun Du 0002, Shutong Niu, Gaobin Yang, Tian Gao 0005, Qingfeng Liu
Comput. Speech Lang.7
2026 Two-stage decomposition network for handwritten Chinese character error correction
Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001, Jianqing Gao, Qingfeng Liu
Pattern Recognit.8
2025 EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
abstract
Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk framework to address these issues. Firstly, to realize better control over the generation of lip movement and facial expression, a Vision-guided Audio Information Decoupling (V-AID) approach is designed to generate audio-based decoupled representations aligned with lip movements and expression. Specifically, to achieve alignment between audio and facial expression representation spaces, we present a Diffusion-based Co-speech Temporal Expansion (Di-CTE) module within V-AID to generate expression-related representations under multi-source emotion condition constraints. Then we propose a well-designed Emotional Talking Head Diffusion (ETHD) backbone to efficiently generate highly expressive talking head videos, which contains an Expression Decoupling Injection (EDI) module to automatically decouple the expressions from reference portraits while integrating the target expression information, achieving more expressive generation performance. Experimental results show that EmotiveTalk can generate expressive talking head videos, ensuring the promised controllability of emotions and metric stability during long-time generation, yielding state-of-the-art performance compared to existing methods. The main page of our paper can be found in https://emotivetalk.github.io/.
Yuzhe Weng, Zilu Guo, Jun Du 0002, Shutong Niu, Jiefeng Ma, Cong Liu 0006, Qingfeng Liu
CVPR13
2025 Col-OLHTR: A Novel Framework for Multimodal Online Handwritten Text Recognition
abstract
Online Handwritten Text Recognition (OLHTR) has gained considerable attention for its diverse range of applications. Current approaches usually treat OLHTR as a sequence recognition task, employing either a single trajectory or image encoder, or multi-stream encoders, combined with a CTC or attention-based recognition decoder. However, these approaches face several drawbacks: 1) single encoders typically focus on either local trajectories or visual regions, lacking the ability to dynamically capture relevant global features in challenging cases; 2) multi-stream encoders, while more comprehensive, suffer from complex structures and increased inference costs. To tackle this, we propose a Collaborative learning-based OLHTR framework, called Col-OLHTR, that learns multimodal features during training while maintaining a single-stream inference process. Col-OLHTR consists of a trajectory encoder, a Point-to-Spatial Alignment (P2SA) module, and an attention-based decoder. The P2SA module is designed to learn image-level spatial features through trajectory-encoded features and 2D rotary position embeddings. During training, an additional image-stream encoder-decoder is collaboratively trained to provide supervision for P2SA features. At inference, the extra streams are discarded, and only the P2SA module is used and merged before the decoder, simplifying the process while preserving high performance. Extensive experimental results on several OLHTR benchmarks demonstrate the state-of-the-art (SOTA) performance, proving the effectiveness and robustness of our design.
Jinshui Hu, Jun Du 0002, Qingfeng Liu
ICASSP7
2025 Model Averaging Under Flexible Loss Functions
abstract
To address model uncertainty under flexible loss functions in prediction problems, we propose a model averaging method that accommodates various loss functions, including asymmetric linear and quadratic loss functions as well as many other asymmetric/symmetric loss functions as special cases. The flexible loss function allows the proposed method to average a large range of models such as the quantile and expectile regression models. To determine the weights of the candidate models, we establish a J-fold cross-validation criterion. Asymptotic optimality and weight convergence are proved for the proposed method. Simulations and an empirical application show the superior performance of the proposed method compared with other methods of model selection and averaging. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: This work was supported by the Beijing Natural Science Foundation [Grant Z240004], Japan Society for the Promotion of Science (KAKENHI) [Grant 22H00833 to Q. Liu], the CAS Project for Young Scientists in Basic Research [Grant YSBR-008], and the National Natural Science Foundation of China [Grants 71925007, 72091212 and 72495124]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0291 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0291 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
Dieqi Gu, Qingfeng Liu, Xinyu Zhang 0024
INFORMS J. Comput.2
2025 Controllable Conformer for Speech Enhancement and Recognition
abstract
We propose a novel approach to speech enhancement, termed Controllable ConforMer for Speech Enhancement (CCMSE), which leverages a Conformer-based architecture integrated with a control factor embedding module. Our method is designed to optimize speech quality for both human auditory perception and automatic speech recognition (ASR). It is observed that while mild denoising typically preserves speech naturalness, stronger denoising can improve human auditory tasks but often at the cost of ASR accuracy due to increased distortion. To address this, we introduce an algorithm that balances these trade-offs. By utilizing differential equations to interpolate between outputs at varying levels of denoising intensity, our method effectively combines the robustness of mild denoising with the clarity of stronger denoising, resulting in enhanced speech that is well-suited for both human and machine listeners. Experimental results on the CHiME-4 dataset validate the effectiveness of our approach.
Zilu Guo, Jun Du 0002, Sabato Marco Siniscalchi, Qingfeng Liu
IEEE Signal Process. Lett.5
2024 NAMER: Non-autoregressive Modeling for Handwritten Mathematical Expression Recognition
Jinshui Hu, Mingjun Chen, Cong Liu 0006, Jun Du 0002, Qingfeng Liu
ECCV (57)9
2024 A Variance-Preserving Interpolation Approach for Diffusion Models With Applications to Single Channel Speech Enhancement and Recognition
abstract
In this paper, we propose a variance-preserving interpolation framework to improve diffusion models for single-channel speech enhancement (SE) and automatic speech recognition (ASR). This new variance-preserving interpolation diffusion model (VPIDM) approach requires only 25 iterative steps and obviates the need for a corrector, an essential element in the existing variance-exploding interpolation diffusion model (VEIDM). Two notable distinctions between VPIDM and VEIDM are the scaling function of the mean of state variables and the constraint imposed on the variance relative to the mean's scale. We conduct a systematic exploration of the theoretical mechanism underlying VPIDM, and develop insights regarding VPIDM's applications in SE and ASR using VPIDM as a frontend. Our proposed approach, evaluated on two distinct data sets, demonstrates VPIDM's superior performances over conventional discriminative SE algorithms. Furthermore, we assess the performance of the proposed model under varying signal-to-noise ratio (SNR) levels. The investigation reveals VPIDM's improved robustness in target noise elimination when compared to VEIDM. Furthermore, utilizing the mid-outputs of both VPIDM and VEIDM results in enhanced ASR accuracies, thereby highlighting the practical efficacy of our proposed approach. Code and audio examples are available onlinehttps://github.com/zelokuo/VPIDM.
Zilu Guo, Qing Wang 0008, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Space-and-speaker-aware acoustic modeling with effective data augmentation for recognition of multi-array conversational speech
Li Chai 0002, Hang Chen 0001, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001
Speech Commun.4
2023 ANSD-MA-MSE: Adaptive Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding
abstract
In this paper, we propose a neural speaker diarization (NSD) network architecture consisting of three key components. First, a memory-aware multi-speaker embedding (MA-MSE) mechanism is proposed to facilitate a dynamical refinement of speaker embedding to reduce a potential data mismatch between the speaker embedding extraction and the NSD network. Next, a speaker selection procedure is introduced to handle situations where the detected number of speakers is different from the assumed speaker size in the NSD network. Finally, an adaptive procedure is proposed to improve the required prior information for the nonoverlap speech segments in a given utterance during each iteration. We call our proposed framework adaptive neural speaker diarization with memory-aware multi-speaker embedding (ANSD-MA-MSE). Our method improves diarization performance in realistic operating scenarios, such as adverse acoustic environments, domain mismatches, and a varying, rather than fixed, number of speakers. Having been tested on both the AMI corpus and the DIHARD-III evaluation sets, our proposed approach consistently outperforms other state-of-the-art techniques in diarization error rates, including the results reported by the best single-model system in the DIHARD-III challenge. Our code is publicly available athttps://github.com/Maokui-He/NSD-MA-MSE.
Maokui He, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 DeepGBASS: Deep Guided Boundary-Aware Semantic Segmentation
abstract
Image semantic segmentation is ubiquitously used in scene understanding applications, such as AI Camera, which require high accuracy and efficiency. Deep learning has significantly advanced the state-of-the-art in semantic segmentation. However, many of recent semantic segmentation works only consider class accuracy and ignore the accuracies at the boundaries between semantic classes. To improve the semantic boundary accuracy, we propose low complexity Deep Guided Decoder (DGD) networks, trained with a novel Semantic Boundary-Aware Learning (SBAL) strategy. Our ablation studies on Cityscapes and the ADE20K-32 confirm the effectiveness of our approach with network of different complexities. We show that our DeepGBASS approach significantly improves the mIoU by up to 11% relative gain and the mean boundary F1-score (mBF) by up to 39.4% when training MobileNetEdgeTPU DeepLab on ADE20K-32 dataset.
Qingfeng Liu, Hai Su, Mostafa El-Khamy, Kee-Bong Song
ICASSP1
2022 Panoptic-Deeplab-DVA: Improving Panoptic Deeplab with Dual Value Attention and Instance Boundary Aware Regression
abstract
Panoptic DeepLab is a state-of-the-art framework that has showed good tradeoff between performance and complexity. In this paper, we focus on improving it to increase wide deployment of panoptic segmentation on mobile devices with low complexity. Specifically, we first present a novel Dual Value Attention (DVA) module to enable context information exchange between the semantic segmentation branch and the instance segmentation branch. Second, we further propose a new instance Boundary Aware Regression (iBAR) loss that assigns more emphasis on the instance boundary during instance regression. To assess the effectiveness of our proposed approach, we evaluate the performance on MSCOCO dataset for panoptic segmentation task, to show that our approach can improve upon the state-of-the-art Panoptic DeepLab with both the light-weight backbone network MobileNetV3 and the heavy-weight backbone network HRNetV2.
Qingfeng Liu, Mostafa El-Khamy
ICIP1
2021 A Neural-Network-Based Approach to Identifying Speakers in Novels
Zhen-Hua Ling, Qingfeng Liu
Interspeech3
2021 A Cross-Entropy-Guided Measure (CEGM) for Assessing Speech Recognition Performance and Optimizing DNN-Based Speech Enhancement
abstract
A new cross-entropy-guided measure (CEGM) is proposed to indirectly assess accuracies of automatic speech recognition (ASR) of degraded speech with a speech enhancement front-end and without directly performing ASR experiments. The proposed CEGM is calculated in three steps, namely: (1) a low-level representations via feature extraction, (2) a high-level nonlinear mapping using an acoustic model, and (3) a final CEGM calculation between the high-level representations of clean and enhanced speech. Specifically, state posterior probabilities from outputs of conventional hybrid acoustic model of the target ASR system are adopted as the high-level representations and a cross-entropy criterion is used to calculate the CEGM. Due to CEGM's differentiability, it can also be used to replace the conventional minimum mean squared error (MMSE) criterion as an objective function for deep neural network (DNN)-based speech enhancement. Therefore, the front-end enhancement model can be optimized towards improving the accuracies of the back-end ASR system. Experiments on single-channel CHiME-4 Challenge show that CEGM yields consistently the highest correlations with word error rate (WER) which is often costly to calculate, and achieves the most accurate assessment of ASR performance when compared to the perceptual evaluation metrics commonly used for assessing speech enhancement performance. Furthermore, CEGM-optimized speech enhancement could effectively reduce the WER on the CHiME-4 real test set when compared to unprocessed noisy speech and enhanced speech obtained with MMSE-optimized enhancement for ASR systems with fixed multi-condition acoustic models in various deep architectures.
Li Chai 0002, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Information Fusion in Attention Networks Using Adaptive and Multi-Level Factorized Bilinear Pooling for Audio-Visual Emotion Recognition
abstract
Multimodal emotion recognition is a challenging task in emotion computing as it is quite difficult to extract discriminative features to identify the subtle differences in human emotions with abstract concept and multiple expressions. Moreover, how to fully utilize both audio and visual information is still an open problem. In this paper, we propose a novel multimodal fusion attention network for audio-visual emotion recognition based on adaptive and multi-level factorized bilinear pooling (FBP). First, for the audio stream, a fully convolutional network (FCN) equipped with 1-D attention mechanism and local response normalization is designed for speech emotion recognition. Next, a global FBP (G-FBP) approach is presented to perform audio-visual information fusion by integrating self-attention based video stream with the proposed audio stream. To improve G-FBP, an adaptive strategy (AG-FBP) to dynamically calculate the fusion weight of two modalities is devised based on the emotion-related representation vectors from the attention mechanism of respective modalities. Finally, to fully utilize the local emotion information, adaptive and multi-level FBP (AM-FBP) is introduced by combining both global-trunk and intra-trunk data in one recording on top of AG-FBP. Tested on the IEMOCAP corpus for speech emotion recognition with only audio stream, the new FCN method outperforms the state-of-the-art results with an accuracy of 71.40%. Moreover, validated on the AFEW database of EmotiW2019 sub-challenge and the IEMOCAP corpus for audio-visual emotion recognition, the proposed AM-FBP approach achieves the best accuracy of 63.09% and 75.49% respectively on the test set.
Hengshun Zhou, Jun Du 0002, Qing Wang 0008, Qingfeng Liu, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 GSANet: Semantic Segmentation With Global And Selective Attention
abstract
This paper proposes a novel deep learning architecture for semantic segmentation. The proposed Global and Selective Attention Network (GSANet) features Atrous Spatial Pyramid Pooling (ASPP) with a novel sparsemax global attention and a novel selective attention that deploys a condensation and diffusion mechanism to aggregate the multi-scale contextual information from the extracted deep features. A selective attention decoder is also proposed to process the GSA-ASPP outputs for optimizing the softmax volume. We are the first to benchmark the performance of semantic segmentation networks with the low-complexity feature extraction network (FXN) MobileNetEdge, that is optimized for low latency on edge devices. We show that GSANet can result in more accurate segmentation with MobileNetEdge, as well as with strong FXNs, such as Xception. GSANet improves the state-of-art semantic segmentation accuracy on both the ADE20k and the Cityscapes datasets.
Qingfeng Liu, Mostafa El-Khamy, Dongwoon Bai
ICIP1
2019 Using Generalized Gaussian Distributions to Improve Regression Error Modeling for Deep Learning-Based Speech Enhancement
abstract
From a statistical perspective, the conventional minimum mean squared error (MMSE) criterion can be considered as the maximum likelihood (ML) solution under an assumed homoscedastic Gaussian error model. However, in this paper, a statistical analysis reveals the super-Gaussian and heteroscedastic properties of the prediction errors in nonlinear regression deep neural network (DNN)-based speech enhancement when estimating clean log-power spectral (LPS) components at DNN outputs with noisy LPS features in DNN input vectors. Accordingly, we propose treating all dimensions of the prediction error vector as statistically independent random variables and model them with generalized Gaussian distributions (GGDs). Then, the objective function with the GGD error model is derived according to the ML criterion. Experiments on the TIMIT corpus corrupted by simulated additive noises show consistent improvements of our proposed DNN framework over the conventional DNN framework in terms of various objective quality measures under 14 unseen noise types evaluated and at various signal-to-noise ratio levels. Furthermore, the ML optimization objective with GGD outperforms the conventional MMSE criterion, achieving improved generalization and robustness.
Li Chai 0002, Jun Du 0002, Qingfeng Liu, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 A Method for PET-CT Lung Cancer Segmentation based on Improved Random Walk
abstract
Segmentation methods only work for a single imaging modality usually suffer from the low spatial resolution in positron emission tomography (PET) or low contrast in computed tomography (CT) when the tumor region is inhomogeneous or not obvious. To address this problem, we develop a segmentation method combining the advantages and disadvantages of PET and CT. Firstly, the initial contours are obtained by the presegmentation of PET images using region growing and mathematical morphology. The initial contours can be used to automatically obtain the seed points required for random walk on PET and CT images, at the same time, they can be also used as a constraint in the random walk on CT images to solve the shortcoming that the tumor areas are not obvious if the CT images have not been enhanced. For the reason that CT provides essential details on anatomic structures, the anatomic structures of CT can be used to improve the weight of random walk on PET images. Finally, the similarity matrices obtained by random walk on PET and CT images are weighted to obtain identical results on PET and CT images. Our methods achieve an average DSC of 0.8456 ± 0.0703 on 14 patients with lung cancer. Our method has much better performance when the tumors are inhomogeneous on PET images and not obvious on CT images.
Zhe Liu 0004, Yuqing Song 0001, Charlie Maere, Qingfeng Liu, Yan Zhu 0018, Hu Lu, Deqi Yuan
ICPR4
2018 Multiple Anthropological Fisher Kernel Framework and Its Application to Kinship Verification
abstract
This paper presents a novel multiple anthropological Fisher kernel (MAFK) framework for kinship verification. The proposed MAFK framework, which goes beyond the Mahalanobis distance metric learning, integrates multiple anthropology inspired features and derives semantically meaningful similarities between images. The major novelty of this paper comes from the following three aspects. First, three new anthropology inspired features (AIF) are derived by extracting the AIF-SIFT, AIF-WLD and AIF-DAISY features on images that are enhanced by an anthropology inspired similarity enhancement method extended from the SIFT flow method. Second, a novel multiple anthropological Fisher kernel framework (MAFK) is proposed which combines multiple features and their metrics between images in a unified paradigm. The MAFK is optimized as a constrained, non-negative, and weighted variant of the sparse representation problem regularized by the criterion of pushing away the nearby non-kinship samples and pulling close the kinship samples. Third, a novel normalized kernel similarity measure (NKSM) is proposed by normalizing the MAFK with the fractional power transformation and L2 normalization. The feasibility of the proposed MAFK framework is assessed on two representative kinship data sets, namely the KinFaceW-I and the KinFaceW-II data sets. The experimental results show the effectiveness of the proposed method.
Ajit Puthenputhussery, Qingfeng Liu, Chengjun Liu
WACV2
2018 Generative and Discriminative Sparse Coding for Image Classification Applications
abstract
This paper presents an enhanced sparse coding method by exploiting both the generative and discriminative information in sparse representation model. Specifically, the proposed generative and discriminative sparse representation (GDSR) method integrates two new criteria, namely a discriminative criterion and a generative criterion, into the conventional sparse representation criterion. The generative criterion reveals the class conditional probability of each dictionary item by using the dictionary distribution coefficients which are derived by representing each dictionary item as a linear combination of the training samples. To further enhance the discriminative ability of the proposed method, a discriminative criterion is also applied using new localized within-class and between-class scatter matrices. Moreover, a novel GDSR based classification (GDSRc) method is proposed by utilizing both the derived sparse representation and the dictionary distribution coefficients. This hybrid method provides new insights, and leads to an effective representation and classification schema for improving the classification performance. The largest step size for learning the sparse representation is theoretically derived to address the convergence issues in the optimization procedure of the GDSR method. Extensive experimental results and analysis on several public classification datasets show the feasibility and effectiveness of the proposed method.
Ajit Puthenputhussery, Qingfeng Liu, Chengjun Liu
WACV2
2017 A Sparse Representation Model Using the Complete Marginal Fisher Analysis Framework and Its Applications to Visual Recognition
abstract
This paper presents an innovative sparse representation model using the complete marginal Fisher analysis (CMFA) framework for different challenging visual recognition tasks. First, a complete marginal Fisher analysis method is presented by extracting the discriminatory features in both the column space of the local samples based within the class scatter matrix and the null space of its transformed matrix. The rationale of extracting features in both spaces is to enhance the discriminatory power by further utilizing the null space, which is not accounted for in the marginal Fisher analysis method. Second, a discriminative sparse representation model is proposed by integrating a representation criterion such as the sparse representation and a discriminative criterion for improving the classification capability. In this model, the largest step size for learning the sparse representation is derived to address the convergence issues in optimization, and a dictionary screening rule is presented to purge the dictionary items with null coefficients for improving the computational efficiency. Experiments on some challenging visual recognition tasks using representative datasets, such as the Painting-91 dataset, the 15 scene categories dataset, the MIT-67 indoor scenes dataset, the Caltech 101 dataset, the Caltech 256 object categories dataset, the AR face dataset, and the extended Yale B dataset, show the feasibility of the proposed method.
Ajit Puthenputhussery, Qingfeng Liu, Chengjun Liu
IEEE Trans. Multim.2
2017 A Novel Locally Linear KNN Method With Applications to Visual Recognition
abstract
A locally linear K Nearest Neighbor (LLK) method is presented in this paper with applications to robust visual recognition. Specifically, the concept of an ideal representation is first presented, which improves upon the traditional sparse representation in many ways. The objective function based on a host of criteria for sparsity, locality, and reconstruction is then optimized to derive a novel representation, which is an approximation to the ideal representation. The novel representation is further processed by two classifiers, namely, an LLK-based classifier and a locally linear nearest mean-based classifier, for visual recognition. The proposed classifiers are shown to connect to the Bayes decision rule for minimum error. Additional new theoretical analysis is presented, such as the nonnegative constraint, the group regularization, and the computational efficiency of the proposed LLK method. New methods such as a shifted power transformation for improving reliability, a coefficients' truncating method for enhancing generalization, and an improved marginal Fisher analysis method for feature extraction are proposed to further improve visual recognition performance. Extensive experiments are implemented to evaluate the proposed LLK method for robust visual recognition. In particular, eight representative data sets are applied for assessing the performance of the LLK method for various visual recognition applications, such as action recognition, scene recognition, object recognition, and face recognition.
Qingfeng Liu, Chengjun Liu
IEEE Trans. Neural Networks Learn. Syst.1
2016 Sparse Representation Based Complete Kernel Marginal Fisher Analysis Framework for Computational Art Painting Categorization
Ajit Puthenputhussery, Qingfeng Liu, Chengjun Liu
ECCV (8)2
2016 SIFT flow based genetic fisher vector feature for kinship verification
abstract
Anthropology studies show that genetic features are inherited by children from their parents resulting in visual resemblance between them. This paper presents a novel SIFT flow based genetic Fisher vector feature (SF-GFVF) which enhances the facial genetic features for kinship verification. The proposed SF-GFVF feature is derived by applying a novel similarity enhancement method based on SIFT flow and learning an inheritable transformation on the Fisher vector feature so as to enhance and encode the genetic features of parent and child image in kinship relations. In particular, the similarity enhancement method is first presented by applying the SIFT flow algorithm to the densely sampled SIFT features in order to intensify the genetic features. Further analysis shows the relation of the extracted genetic features to anthropological results and discovers interesting patterns in different kinship relations. Finally, an inheritable transformation is applied to the enhanced Fisher vector feature which is learned with the criterion of minimizing the distance between kinship samples and maximizing the distance between non-kinship samples. Experimental results on the two representative kinship databases, namely the KinFace W-I and the Kinship W-II data sets show that the proposed method is able to outperform other popular methods.
Ajit Puthenputhussery, Qingfeng Liu, Chengjun Liu
ICIP2
2016 A novel inheritable color space with application to kinship verification
abstract
Anthropology studies discover that some genetic related facial features, which are inherited by children from their parents, can be used for kinship verification. This paper investigates an important inheritable feature - color and presents a novel inheritable color space (InCS) and a generalized InCS (GInCS) framework with application to kinship verification. Specifically, a novel color similarity measure (CSM) is first defined. Second, based on this similarity measure, a new inheritable color space (InCS) is derived by balancing the criterion of minimizing the distance between kinship pairs and the criterion of maximizing the distance between non-kinship pairs. Unlike conventional color spaces, e.g. the RGB color space, the proposed InCS, which is learned automatically from the data, captures the inheritable information between parent and child. Third, theoretical and empirical analysis show that the proposed InCS exhibits the decorrelation property, which is positively related to the performance of kinship verification. Robustness to the illumination variation is also discussed. Fourth, a generalized InCS framework is presented to extend the InCS from the pixel level to the feature level for improving the performance and the robustness to illumination variation. The proposed InCS is evaluated on several popular datasets, namely the KinFaceW-I dataset, the KinFaceW-II dataset, the UB KinFace dataset, and the Cornell KinFace dataset. Experimental results show that the proposed InCS is able to (i) improve the conventional color spaces such as RGB, YUV, YIQ color spaces by a large margin, (ii) achieve robustness to the illumination variation, and (iii) outperforms other popular methods.
Qingfeng Liu, Ajit Puthenputhussery, Chengjun Liu
WACV1
2016 Color multi-fusion fisher vector feature for fine art painting categorization and influence analysis
abstract
This paper presents a novel set of image features that encode the local, color, spatial, relative intensity information and gradient orientation of the painting image for painting artist classification, style classification as well as artist and style influence analysis. In particular, a new color DAISY Fisher vector (CD-FV) feature is first created by computing Fisher vectors on densely sampled DAISY features. Second, a color WLD-SIFT Fisher vector (CWS-FV) feature is developed by fusing Weber local descriptors (WLD) with Scale Invariant Feature Transform (SIFT) descriptors and Fisher vectors are computed on the fused WLD-SIFT features. Finally, an innovative color multi-fusion Fisher vector (CMFFV) feature is developed by integrating the Principal Component Analysis (PCA) features of CD-FV, CWS-FV and color SIFT-FV features. The effectiveness of the proposed CMFFV feature is assessed on the challenging Painting-91 dataset. Experimental results show that the proposed CMFFV feature is able to (i) achieve the state-of-the-art performance for painting artist classification, (ii) outperform other popular image descriptors, as well as (iii) discover the artist and style influence to understand their connections and evolution in different art movement periods.
Ajit Puthenputhussery, Qingfeng Liu, Chengjun Liu
WACV2
2015 A novel locally linear KNN model for visual recognition
abstract
This paper presents a novel locally linear KNN model with the goal of not only developing efficient representation and classification methods, but also establishing a relation between them so as to approximate some classification rules, e.g. the Bayes decision rule. Towards that end, first, the proposed model represents the test sample as a linear combination of all the training samples and derives a new representation by learning the coefficients considering the reconstruction, locality and sparsity constraints. The theoretical analysis shows that the new representation has the grouping effect of the nearest neighbors, which is able to approximate the “ideal representation”. And then the locally linear KNN model based classifier (LLKNNC), which shows its connection to the Bayes decision rule for minimum error in the view of kernel density estimation, is proposed for classification. Besides, the locally linear nearest mean classifier (LLNMC), whose relation to the LLKNNC is just like the nearest mean classifier to the KNN classifier, is also derived. Furthermore, to provide reliable kernel density estimation, the shifted power transformation and the coefficients cut-off method are applied to improve the performance of the proposed method. The effectiveness of the proposed model is evaluated on several visual recognition tasks such as face recognition, scene recognition, object recognition and action recognition. The experimental results show that the proposed model is effective and outperforms some other representative popular methods.
Qingfeng Liu, Chengjun Liu
CVPR1
2015 Unsupervised speaker adaptation of deep neural network based on the combination of speaker codes and singular value decomposition for speech recognition
abstract
Recently, we have proposed a general adaptation scheme for deep neural network based on discriminant condition codes and applied it to supervised speaker adaptation in speech recognition based on either frame-level cross-entropy or sequence-level maximum mutual information training criterion [1, 2, 3, 4]. In this case, each condition code is associated with one speaker in data, which is thus called speaker code for convenience. Our previous work has shown that speaker code based methods are quite effective in adapting DNNs even when only a very small amount of adaptation data is available. However, we have to use a large speaker code size and complex processes to obtain the best ASR performance since good initializations of speaker codes and connection weights are very important. In this paper, we propose a method using singular value decomposition (SVD) as in [5] to initialize speaker codes and connection weights to obtain a comparable ASR performance as before but with a smaller speaker code size and much less computation complexity. Meanwhile, we have evaluated unsupervised speaker adaptation with the proposed method in large vocabulary speech recognition in the Switchboard task. Experimental results have shown that it is effective for providing well initializations and suitable in adapting large DNN models.
Shaofei Xue, Hui Jiang 0001, Li-Rong Dai 0001, Qingfeng Liu
ICASSP4
2015 Novel general KNN classifier and general nearest mean classifier for visual classification
abstract
This paper presents a novel general k nearest neighbour classifier (GKNNc) and a novel general nearest mean classifier (GNMc) for visual classification. Instead of treating the data equally, both GKNNc and GNMc assign a weight coefficient to each data. To achieve good performance, the conditions and properties of the weight coefficients for GKNNc and GNMc are further analysed. Then a sparse representation based method is proposed to derive the weight coefficients for both GKNNc and GNMc. Experimental results on several representative data sets, such as the Caltech 101 dataset and the MIT-67 indoor scenes dataset demonstrate the feasibility of the proposed methods.
Qingfeng Liu, Ajit Puthenputhussery, Chengjun Liu
ICIP1
2015 Learning the discriminative dictionary for sparse representation by a general fisher regularized model
abstract
This paper presents two novel discriminative dictionary learning models for sparse representation, namely the Fisher discriminative sparse model (FDSM) and the marginal Fisher discriminative sparse model (MFDSM). To learn the FDSM and the MFDSM efficiently and homogeneously, a general Fisher regularized model is further derived so that both of them can be learned without much modification. Experimental results on four popular databases, namely the extended Yale face database B, the AR face database, the 15 scenes dataset and the MIT-67 indoor scenes dataset show that the proposed method can improve upon other popular methods.
Qingfeng Liu, Ajit Puthenputhussery, Chengjun Liu
ICIP1
2015 State-Clustering Based Multiple Deep Neural Networks Modeling Approach for Speech Recognition
abstract
The hybrid deep neural network (DNN) and hidden Markov model (HMM) has recently achieved dramatic performance gains in automatic speech recognition (ASR). The DNN-based acoustic model is very powerful but its learning process is extremely time-consuming. In this paper, we propose a novel DNN-based acoustic modeling framework for speech recognition, where the posterior probabilities of HMM states are computed from multiple DNNs (mDNN), instead of a single large DNN, for the purpose of parallel training towards faster turnaround. In the proposed mDNN method all tied HMM states are first grouped into several disjoint clusters based on data-driven methods. Next, several hierarchically structured DNNs are trained separately in parallel for these clusters using multiple computing units (e.g. GPUs). In decoding, the posterior probabilities of HMM states can be calculated by combining outputs from multiple DNNs. In this work, we have shown that the training procedure of the mDNN under popular criteria, including both frame-level cross-entropy and sequence-level discriminative training, can be parallelized efficiently to yield significant speedup. The training speedup is mainly attributed to the fact that multiple DNNs are parallelized over multiple GPUs and each DNN is smaller in size and trained by only a subset of training data. We have evaluated the proposed mDNN method on a 64-hour Mandarin transcription task and the 320-hour Switchboard task. Compared to the conventional DNN, a 4-cluster mDNN model with similar size can yield comparable recognition performance in Switchboard (only about 2% performance degradation) with a greater than 7 times speed improvement in CE training and a 2.9 times improvement in sequence training, when 4 GPUs are used.
Hui Jiang 0001, Li-Rong Dai 0001, Yu Hu 0003, Qingfeng Liu
IEEE ACM Trans. Audio Speech Lang. Process.5
2014 A new locally linear KNN method with an improved marginal Fisher analysis for image classification
abstract
This paper presents a novel locally linear KNN method with an improved marginal Fisher analysis for image classification. First, the discriminating color space (DCS), which is derived by discriminant analysis of the red, green, and blue primary colors, is integrated into the proposed method. Second, an improved marginal Fisher analysis (IMFA) applies an eigenvalue spectrum analysis to improve the generalization performance of the marginal Fisher analysis method. Third, a new locally linear KNN classifier (LLKNN), which represents the test image as a linear combination of its k nearest training images and assigns it to the class with the largest sum of weights, is presented to improve upon the traditional KNN approach. The effectiveness of the proposed method is evaluated on two representative datasets, namely the AR face image data set and the ETH-80 image data set. Experimental results show that the proposed method performs better than some representative state-of-the-art methods.
Qingfeng Liu, Chengjun Liu
IJCB1
2014 A novel hierarchical interaction model and HITS map for action recognition in static images
abstract
This paper proposes a novel fully automatic method to model the low-level human and object interactions for action recognition in the static images. Specifically, we exploit both the superpixels and the grid patches of an image to construct a hierarchical interaction graph and then develop an HITS map learning algorithm to learn the human-object interactions for recognizing the human actions. The major contributions of the paper are three-fold. First, a novel two-layer hierarchical interaction graph based on the superpixels and the grid patches is presented to model the low-level human-object interactions. Second, the novel HITS map, which is derived by the weighted HITS algorithm on the hierarchical interaction graph, assigns heavy weights to the important superpixels and grid patches that reveal more meaningful interactions. Third, the novel weighted image representation is derived from the learned HITS map for action recognition. Extensive experimental results show the feasibility of the proposed method using three representative datasets, namely, the Willow Action dataset, the UIUC Sports Event dataset and the CMU Sports dataset. In particular, the proposed method is able to (i) automatically model the human-object interactions without extensive manual annotations or numerous error-prone detections, and (ii) improve upon other popular methods in terms of action recognition performance.
Qingfeng Liu, Chengjun Liu
SMC1
2014 Fast adaptation of deep neural network based on discriminant codes for speech recognition
abstract
Fast adaptation of deep neural networks (DNN) is an important research topic in deep learning. In this paper, we have proposed a general adaptation scheme for DNN based on discriminant condition codes, which are directly fed to various layers of a pre-trained DNN through a new set of connection weights. Moreover, we present several training methods to learn connection weights from training data as well as the corresponding adaptation methods to learn new condition code from adaptation data for each new test condition. In this work, the fast adaptation scheme is applied to supervised speaker adaptation in speech recognition based on either frame-level cross-entropy or sequence-level maximum mutual information training criterion. We have proposed three different ways to apply this adaptation scheme based on the so-called speaker codes: i) Nonlinear feature normalization in feature space; ii) Direct model adaptation of DNN based on speaker codes; iii) Joint speaker adaptive training with speaker codes. We have evaluated the proposed adaptation methods in two standard speech recognition tasks, namely TIMIT phone recognition and large vocabulary speech recognition in the Switchboard task. Experimental results have shown that all three methods are quite effective to adapt large DNN models using only a small amount of adaptation data. For example, the Switchboard results have shown that the proposed speaker-code-based adaptation methods may achieve up to 8-10% relative error reduction using only a few dozens of adaptation utterances per speaker. Finally, we have achieved very good performance in Switchboard (12.1% in WER) after speaker adaptation using sequence training criterion, which is very close to the best performance reported in this task (“Deep convolutional neural networks for LVCSR,” T. N. Sainath , Proc. IEEE Acoust., Speech, Signal Process., 2013).
Shaofei Xue, Ossama Abdel-Hamid, Hui Jiang 0001, Li-Rong Dai 0001, Qingfeng Liu
IEEE ACM Trans. Audio Speech Lang. Process.5
2013 A cluster-based multiple deep neural networks method for large vocabulary continuous speech recognition
abstract
Recently a pre-trained context-dependent hybrid deep neural network (DNN) and HMM method has achieved significant performance gain in many large-scale automatic speech recognition (ASR) tasks. However, the error back-propagation (BP) algorithm for training neural networks is sequential in nature and is hard to parallelize into multiple computing threads. Therefore, training a deep neural network is extremely time-consuming even with a modern GPU board. In this paper we have proposed a new acoustic modelling framework to use multiple DNNs instead of a single DNN to compute the posterior probabilities of tied HMM states. In our method, all tied states of context-dependent HMMs are first grouped into several disjoined clusters based on the training data associated with these HMM states. Then, several hierarchically structured DNNs are trained separately for these disjoined clusters of data using multiple GPUs. In decoding, the final posterior probability of each tied HMM state can be calculated based on output posteriors from multiple DNNs. We have evaluated the proposed method on a 64-hour Mandarin transcription task and 309-hour Switchboard Hub5 task. Experimental results have shown that the new method using clusterbased multiple DNNs can achieve over 5 times reduction in total training time with only negligible performance degradation (about 1-2% in average) when using 3 or 4 GPUs respectively.
Cong Liu 0006, Qingfeng Liu, Li-Rong Dai 0001, Hui Jiang 0001
ICASSP3
2012 Construction of gene regulatory networks with colored noise
Xuesong Wang 0001, Lijing Li, Yuhu Cheng 0001, Qingfeng Liu
Neural Comput. Appl.4
2009 Land Use/Cover Characterization with MODIS Time Series Data with Hybrid Classification Mothed over Australia for 2001 and 2003
abstract
Improved and up-to-date land use/land cover (LULC) data sets are needed over the whole country of Australia to support science and policy applications focused on understanding the role and response of the LULC to environmental change. The main goal of this study was to map LULC in Australia using MODIS 250 m Normalized Difference Vegetation Index (NDVI), Land Surface Vegetation Index (LSWI) and reflectance time series data of 2000 and 2003. NDVI time-series were filtered by the Savitzky-Golay algorithm in the present study to smooth out noise. A combination of unsupervised ISODATA and a hierarchical decision tree classification were performed on 2 years 12-month time-series MODIS data. Also, Australian Vegetation Map and other land use/land cover data set were used as labeling reference during the classification process. The MODIS land cover products were evaluated using existing land use/cover data derived from Landsat TM as reference data (AUS-2000), also LULC information derived from 11 scenes of Landsat-5 TM data were used as validation data source. The overall classification accuracy was 76.4%. It turned out that our result is acceptable because the relative high resolution of MODIS data and more prior knowledge was applied.
Kaishan Song, Mohsin Hafeez, Zongming Wang, Dongmei Lu, Lihong Zeng, Dianwei Liu, Bai Zhang, Jia Du, Qingfeng Liu
IGARSS (3)10
2009 Land Use/land Cover (LULC) Characterizaitoin with MODIS Time Series Data in the Amu River Basin
abstract
Improved and up-to-date land use/land cover (LULC) data sets are needed over intensively land use/cover change area in the Amur River Basin (ARB) to support science and policy applications focused on understanding of the role and response of the LULC to environmental change issues. The main goal of this study was to map LULC in the Amur River Basin using MODIS 250 m Normalized Difference Vegetation Index (NDVI) and Land Surface Water Index (LSWI) time series data in 2001 and 2007. A combination of unsupervised ISODATA and hierarchical decision tree classification were performed on 12-month time-series of MODIS NDVI data over the study region. The MODIS land cover result of Northeast China was evaluated using existing land use/cover data, and the rest part was evaluated by LULC information derived from LANDSAT-TM. MODIS 250m NDVI, LSWI and reflectance datasets were found to have sufficient spatial, spectral, and temporal resolutions to detect unique multitemporal signatures of the major land cover types over the region. The overall classification accuracy was 0.81 and the kappa coefficient is 0.64. In conclusion, this method has been used successively for LULC change monitoring in the year 2001 and 2007. The result indicate that MODIS 250 NDVI time series data can derive relatively accurate LULC information for hydrological and climate modeling.
Kaishan Song, Zongming Wang, Qingfeng Liu, Dongmei Lu, Lihong Zeng, Dianwei Liu, Bai Zhang, Jia Du
IGARSS (4)3
2007 Supervised Learning Approach to Optimize Ranking Function for Chinese FAQ-Finder
Qingfeng Liu, Renhua Wang
PAKDD3
2007 DITOP: drug-induced toxicity related protein database
abstract
MOTIVATION: Drug-induced toxicity related proteins (DITRPs) are proteins that mediate adverse drug reactions (ADRs) or toxicities through their binding to drugs or reactive metabolites. Collection of these proteins facilitates better understanding of the molecular mechanisms of drug-induced toxicity and the rational drug discovery. Drug-induced toxicity related protein database (DITOP) is such a database that is intending to provide comprehensive information of DITRPs. Currently, DITOP contains 1501 records, covering 618 distinct literature-reported DITRPs, 529 drugs/ligands and 418 distinct toxicity terms. These proteins were confirmed experimentally to interact with drugs or their reactive metabolites, thus directly or indirectly cause adverse effects or toxicities. Five major types of drug-induced toxicities or ADRs are included in DITOP, which are the idiosyncratic adverse drug reactions, the dose-dependent toxicities, the drug-drug interactions, the immune-mediated adverse drug effects (IMADEs) and the toxicities caused by genetic susceptibility. Molecular mechanisms underlying the toxicity and cross-links to related resources are also provided while available. Moreover, a series of user-friendly interfaces were designed for flexible retrieval of DITRPs-related information. The DITOP can be accessed freely at http://bioinf.xmu.edu.cn/databases/ADR/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jing-Xian Zhang, Wei-Juan Huang, Jing-Hua Zeng, Wen-Hui Huang, Bu-Cong Han, Qingfeng Liu, Yuzong Chen 0002, Zhi Liang Ji
Bioinform.8
2000 KD2000 Chinese Text-To-Speech System
Renhua Wang, Qingfeng Liu, Yu Hu 0003, Xiaoru Wu
ICMI2
2000 Prosody generation in Chinese synthesis using the template of quantified prosodic unit and base intonation contour
Yu Hu 0003, Qingfeng Liu, Renhua Wang
INTERSPEECH2
1998 Towards a Chinese text-to-speech system with higher naturalness
Renhua Wang, Qingfeng Liu, Yongsheng Teng, Deyu Xia
ICSLP2
1996 A new Chinese text-to-speech system with high naturalness
Renhua Wang, Qingfeng Liu, Difei Tang
ICSLP2