VLDB 2026 Research / reviewers in the wild / expert
Chunmei Qing
dblp:33/5818
· DBLP profile ↗
46ranked-venue papers
4as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A dual uncertainty-aware fusion framework for face expression recognition in the wildabstractFacial Expression Recognition(FER) is a key task in the broader landscape of affective computing and human-computer interaction, enabling machines to interpret human emotions. To better learn discriminative features under complex facial variations, recent FER research has increasingly adopted multi-branch fusion architectures that aim to capture complementary features from diverse perspectives. However, existing multi-branch fusion strategies, including static weighting, simple concatenation, or uncertainty-aware modeling, lack the capacity to comprehensively capture and reconcile the reliability variations across both individual instances and structural branches. To overcome these limitations, we propose a novel multi-branch fusion strategy, named Dual Uncertainty-Aware Fusion Framework(DUAFF), which improves the discriminability of integrated features by simultaneously modeling instance-wise uncertainty and inter-branch correlations. Specifically, the proposed method comprises two complementary modules: Instance-Discrepant Uncertainty-Aware Fusion Module (ID-UAFM) and Branch-Discrepant Uncertainty-Aware Fusion Module (BD-UAFM). ID-UAFM is introduced to perform channel-wise entropy analysis between semantically distinct samples to estimate instance-level uncertainty, enabling selective channel-wise fusion that emphasizes reliable representations while suppressing uncertain responses. BD-UAFM is further proposed to capture structural uncertainty by evaluating the relative reliability of features across multiple branches and adaptively weighting their contributions based on inter-branch discrepancies. Experimental results demonstrate that the proposed DUAFF consistently outperforms POSTER across three benchmark datasets, achieving accuracy improvements of 0.23 % on RAF-DB, 0.69 % on FER2013, and 0.29 % on AffectNet (7-class), thereby confirming its effectiveness in enhancing the reliability and discriminability of facial representations. Wenfeng Jiang, Lin Wang 0004, Fang Liu 0030, Chunmei Qing, Xiaofen Xing, Xiangmin Xu 0001, Weiquan Fan, Zhanpeng Jin |
Expert Syst. Appl. | 5 |
| 2026 | Progressive multi-branch video style transfer network via confidence reweighted projection
Kunbo Han, Hongyan Yin, Junpeng Tan, Chong-zhi Gao, Chunmei Qing |
Neural Networks | 5 |
| 2026 | Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression ManipulationabstractSpeech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current works depend on inaccessible paired training samples for the person, where two aligned frames exhibit the same speech content yet differ in emotional expression, limiting the SPFEM applications in real-world scenarios. In this work, we discover that speakers who convey the same content with different emotions exhibit highly correlated local facial animations in both spatial and temporal spaces, providing valuable supervision for SPFEM. To capitalize on this insight, we propose a novel spatial-temporal coherent correlation learning (STCCL) algorithm, which models the aforementioned correlations as explicit metrics and integrates the metrics to supervise manipulating facial expression and meanwhile better preserving the facial animation of spoken content. To this end, it first learns a spatial coherent correlation metric, ensuring that the visual correlations of adjacent local regions within an image linked to a specific emotion closely resemble those of corresponding regions in an image linked to a different emotion. Simultaneously, it develops a temporal coherent correlation metric, ensuring that the visual correlations of specific regions across adjacent image frames associated with one emotion are similar to those in the corresponding regions of frames associated with another emotion. Recognizing that visual correlations are not uniform across all regions, we have also crafted a correlation-aware adaptive strategy that prioritizes regions that present greater challenges. During SPFEM model training, we construct the spatial-temporal coherent correlation metric between corresponding local regions of the input and output image frames as additional loss to supervise the generation process. We conduct extensive experiments on various datasets, and the results demonstrate the effectiveness of the proposed STCCL algorithm. Tianshui Chen, Jianman Lin, Zhijing Yang, Chunmei Qing, Guangrun Wang, Liang Lin 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
Zhihua Xu, Tianshui Chen, Zhijing Yang, Chunmei Qing, Keze Wang, Liang Lin 0004 |
IEEE Trans. Multim. | 4 |
| 2025 | Enhancing fNIRS Signal Classification with Test-Time Training by Improved Spatiotemporal Feature ExtractionabstractFunctional Near-Infrared Spectroscopy (fNIRS) is a convenient brain imaging technology that is adaptable to various complex environments. It can detect the brain's responses to external stimuli across different contexts. However, existing fNIRS processing methods fail to mine the brain's sequential association responses in long-term signals over time. To better capture the hidden spatiotemporal correlations and long-range dependencies inherent in fNIRS signals, we propose fNIRS-TTT, a novel classification architecture leveraging Test-Time Training (TTT). Our framework introduces two core innovations: the Cross-Attention Embedding module (CAE) and the TTT Block. The CAE module combines the wide-kernel Patch Conv (capturing cross-channel spatial patterns) and the Channel Conv (processing temporal features within individual channels) together. Their outputs are fused via Cross-Attention to generate spatiotemporal tokens enriched with multi-scale spatial information for fNIRS. The TTT components dynamically adapt to the sequential nature of fNIRS data, effectively capturing long-range temporal context and mitigating overfitting compared to former models. KFold Cross-Validation (KFold-CV) and Leave-One-Subject Cross-Validation (LOSO-CV) are conducted on three open datasets, which demonstrate the superior performance of the proposed fNIRS-TTT compared to state-of-the-art models. Wanxiang Luo, Chunmei Qing, Junpeng Tan, Yihang Zou, Xiangmin Xu 0001 |
BIBM | 2 |
| 2025 | DP-Net: A 3D Dilated Projection Framework For Precise Fetal Brain Tissue SegmentationabstractPrecise segmentation of fetal brain tissues in MRI is essential for studying brain development and for the early diagnosis and treatment of neurological disorders. However, the complex and variable anatomy of the fetal brain, significant morphological changes at different gestational ages, and the low-quality MRI and inherent noise of fetal acquisition pose significant challenges. To address these, we propose a novel 3D Dilated Projection U-net segmentation framework, DP-Net, which incorporates large kernel convolutions, atrous convolution for receptive field expansion, and dual skip connections mechanism to enhance global semantic consistency. Specifically, we introduce a Dilated Projection Block (DPB) that leverages atrous convolution to capture global context across multiple anatomical regions without additional parameters. Furthermore, we propose a Dual Skip Connection (DSC) mechanism to maintain encoder-decoder global consistency by fusing low-level and projected high-level features, mitigating blind spots introduced by atrous convolution. Extensive experiments show that our method significantly outperforms state-of-the-art methods, demonstrating its robustness and effectiveness in addressing the challenges of fetal brain tissue segmentation. Junpeng Tan, Mingjin Chen, Chunmei Qing, Xin Zhang 0013, Xiangmin Xu 0001 |
ICIP | 3 |
| 2025 | Neural Scene Designer: Self-Styled Semantic Image ManipulationabstractMaintaining stylistic consistency is crucial for the cohesion and aesthetic appeal of images, a fundamental requirement in effective image editing and inpainting. However, existing methods primarily focus on the semantic control of generated content, often neglecting the critical task of preserving this consistency. In this work, we introduce the Neural Scene Designer (NSD), a novel framework that enables photo-realistic manipulation of user-specified scene regions while ensuring both semantic alignment with user intent and stylistic consistency with the surrounding environment. NSD leverages an advanced diffusion model, incorporating two parallel cross-attention mechanisms that separately process text and style information to achieve the dual objectives of semantic control and style consistency. To capture fine-grained style representations, we propose the Progressive Self-style Representational Learning (PSRL) module. This module is predicated on the intuitive premise that different regions within a single image share a consistent style, whereas regions from different images exhibit distinct styles. The PSRL module employs a style contrastive loss that encourages high similarity between representations from the same image while enforcing dissimilarity between those from different images. Furthermore, to address the lack of standardized evaluation protocols for this task, we establish a comprehensive benchmark. This benchmark includes competing algorithms, dedicated style-related metrics, and diverse datasets and settings to facilitate fair comparisons. Extensive experiments conducted on our benchmark demonstrate the effectiveness of the proposed framework. Jianman Lin, Tianshui Chen, Chunmei Qing, Zhijing Yang, Shuangping Huang, Yuheng Ren, Liang Lin 0004 |
IEEE Trans. Image Process. | 3 |
| 2025 | Artistic Style Transfer via Fine-Grained Text Guidance and Contrastive Semantics SimilarityabstractDue to the development of text-image multimodal methods, text is used to guide the style transfer of images, which has attracted growing attention. Notably, The existing text-guided image style methods are limited to expressing specific artistic style through simple text. It can only accept coarse-grained text input such as “Van Gogh” and “White Cloud”, and cannot understand fine-grained text input such as “The Night Café by Vincent van Gogh”. To this end, this paper proposes a novel artistic style transfer network based on the fine-grained text guidance and the contrastive semantics similarity, named as TCStyler. It can accept images or texts as style guidance, which is more suitable for fine-grained content understanding stylization. In this network, to address the issue of text-image cross-modal discrepancy, the residual attention feature mapper (RAFM) is introduced to constrain the differences between different modalities in feature space. Then, the global cascading style-sharing module (GCSM) is proposed for performing content-style feature fusion and image-text modality fusion by adopting a global feature-sharing strategy. Furthermore, the contrastive semantics similarity loss is designed to address the problem of multimodal universality. Quantitative and visualization experiments demonstrate that our TCStyler can handle fine-grained artistic text inputs and maintain consistency in the style transfer results guided by different modalities. Chunmei Qing, Junpeng Tan, Jianxiu Jin, Xiangmin Xu 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Multi-view Panoramic Image Style Transfer with Multi-scale Attention and Global SharingabstractStyle transfer for panoramic images is a challenging task, due to the problems associated with its unique structure, including edge discontinuities, pole distortion, fuzzy details, and memory limitation. In this article, we propose a novel Multi-view Transformation network for Panorama Style Transfer (MuTPST). First, this architecture has a multi-view panoramic transformation mechanism, which includes a multi-view cubic projection and a multi-view equirectangular re-projection of panoramic images. This can address pole distortion and edge discontinuity by skillfully applying multiple types of projections and transformations. To capture different levels of context and structure in the stylization stage, we carefully design a multi-scale attention content encoder, which can coordinate the distribution of visual attention across space and channels. Besides, by the sharing of global style features in thumbnails and patches, MuTPST can process ultra-high-resolution panoramic images (e.g., 10,000 \(\times\) 5,000 pixels) with limited GPU memory. Extensive experiments illustrate that the proposed method outperforms the state-of-the-art with a discernible improvement in panoramic image style transfer. More results and interactive features can be found on https://weiyang001.github.io/MuTPST/ . Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Learning Adaptive Spatial Coherent Correlations for Speech-Preserving Facial Expression ManipulationabstractSpeech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current works depend on inaccessible paired training samples for the person, where two aligned frames exhibit the same speech content yet differ in emotional expression, limiting the SPFEM applications in real-world scenarios. In this work, we discover that speak-ers who convey the same content with different emotions exhibit highly correlated local facial animations, providing valuable supervision for SPFEM. To capitalize on this insight, we propose a novel adaptive spatial coherent correlation learning (ASCCL) algorithm, which models the aforementioned correlation as an explicit metric and integrates the metric to supervise manipulating facial expression and meanwhile better preserving the facial animation of spoken contents. To this end, it first learns a spatial coherent correlation metric, ensuring the visual disparities of adjacent local regions of the image belonging to one emotion are similar to those of the corresponding counterpart of the image belonging to another emotion. Recognizing that visual disparities are not uniform across all regions, we have also crafted a disparity-aware adaptive strategy that prioritizes regions that present greater challenges. During SPFEM model training, we construct the adaptive spatial coherent correlation metric between corresponding local regions of the input and output images as addition loss to supervise the generation process. We conduct extensive experiments on variant datasets, and the results demonstrate the effectiveness of the proposed ASCCL algorithm. Code is publicly available at https://githiub.com/jianmanlincjx/ASCCL Tianshui Chen, Jianman Lin, Zhijing Yang, Chunmei Qing, Liang Lin 0004 |
CVPR | 4 |
| 2024 | Consistent Panoramic Video Style Transfer via Temporal-Spatial Cross Perception
Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001 |
ICIC (6) | 2 |
| 2024 | Fetal MRI Reconstruction by Global Diffusion and Consistent Implicit Representation
Junpeng Tan, Xin Zhang 0013, Chunmei Qing, Chaoxiang Yang, He Zhang 0023, Gang Li 0001, Xiangmin Xu 0001 |
MICCAI (7) | 3 |
| 2024 | Self-Supervised Emotion Representation Disentanglement for Speech-Preserving Facial Expression ManipulationabstractSpeech-preserving Facial Expression Manipulation (SPFEM) aims to alter facial emotions in video content while preserving the facial movements associated with speech. Current works often fall short due to the inadequate representation of emotion as well as the absence of time-aligned paired data-two corresponding frames from the same speaker that showcase the same speech content but differ in emotional expression. In this work, we introduce a novel framework, Self-Supervised Emotion Representation Disentanglement (SSERD), to disentangle emotion representation for accurate emotion transfer while implementing a paired data construction module to facilitate automated, photorealistic facial animations. Specifically, We developed a module for learning emotion latent codes using StyleGAN's latent space, employing a cross-attention mechanism to extract and predict emotion editing codes, with contrastive learning to differentiate emotions. To overcome the lack of strictly paired data in the SPFEM task, we exploit pretrained StyleGAN to generate paired data, focusing on expression vectors unrelated to mouth shape. Additionally, we employed a hybrid training strategy using both synthetic paired and real unpaired data to enhance the realism of SPFEM model's generated images. Extensive experiments conducted on benchmark datasets, including MEAD and RAVDESS, have validated the effectiveness of our framework, demonstrating its superior capability in generating photorealistic and expressive facial animations. Zhihua Xu, Tianshui Chen, Zhijing Yang, Chunmei Qing, Yukai Shi, Liang Lin 0004 |
ACM Multimedia | 4 |
| 2024 | MASANet: Multi-Aspect Semantic Auxiliary Network for Visual Sentiment AnalysisabstractRecently, multi-modal affective computing has demonstrated that introducing multi-modal information can enhance performance. However, multi-modal research faces significant challenges due to its high requirements regarding data acquisition, modal integrity, and feature alignment. The widespread use of multi-modal pre-training methods offers the possibility of aiding visual sentiment analysis by introducing cross-domain knowledge. This paper proposes a Multi-Aspect Semantic Auxiliary Network (MASANet) for visual sentiment analysis. Specifically, MASANet achieves modality expansion through cross-modal generation, making it possible to introduce cross-domain semantic assistance. Then, a cross-modal gating module and an adaptive modal fusion module are proposed for aspect-level and cross-modal interaction, respectively. In addition, a designed semantic polarity constraint loss is presented to improve sentiment multi-classification performance. Evaluations of eight widely-used affective image datasets demonstrate that our proposed method outperforms the state-of-the-art methods. Further ablation experiments and visualization results also confirm the effectiveness of the proposed method and its modules. Jinglun Cen, Chunmei Qing, Haochun Ou, Xiangmin Xu 0001, Junpeng Tan |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Fourier Domain Robust Denoising Decomposition and Adaptive Patch MRI ReconstructionabstractThe sparsity of the Fourier transform domain has been applied to magnetic resonance imaging (MRI) reconstruction in k -space. Although unsupervised adaptive patch optimization methods have shown promise compared to data-driven-based supervised methods, the following challenges exist in MRI reconstruction: 1) in previous k -space MRI reconstruction tasks, MRI with noise interference in the acquisition process is rarely considered. 2) Differences in transform domains should be resolved to achieve the high-quality reconstruction of low undersampled MRI data. 3) Robust patch dictionary learning problems are usually nonconvex and NP-hard, and alternate minimization methods are often computationally expensive. In this article, we propose a method for Fourier domain robust denoising decomposition and adaptive patch MRI reconstruction (DDAPR). DDAPR is a two-step optimization method for MRI reconstruction in the presence of noise and low undersampled data. It includes the low-rank and sparse denoising reconstruction model (LSDRM) and the robust dictionary learning reconstruction model (RDLRM). In the first step, we propose LSDRM for different domains. For the optimization solution, the proximal gradient method is used to optimize LSDRM by singular value decomposition and soft threshold algorithms. In the second step, we propose RDLRM, which is an effective adaptive patch method by introducing a low-rank and sparse penalty adaptive patch dictionary and using a sparse rank-one matrix to approximate the undersampled data. Then, the block coordinate descent (BCD) method is used to optimize the variables. The BCD optimization process involves valid closed-form solutions. Extensive numerical experiments show that the proposed method has a better performance than previous methods in image reconstruction based on compressed sensing or deep learning. Junpeng Tan, Xin Zhang 0013, Chunmei Qing, Xiangmin Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Superpoint Transformer for 3D Scene Instance SegmentationabstractMost existing methods realize 3D instance segmentation by extending those models used for 3D object detection or 3D semantic segmentation. However, these non-straightforward methods suffer from two drawbacks: 1) Imprecise bounding boxes or unsatisfactory semantic predictions limit the performance of the overall 3D instance segmentation framework. 2) Existing method requires a time-consuming intermediate step of aggregation. To address these issues, this paper proposes a novel end-to-end 3D instance segmentation method based on Superpoint Transformer, named as SPFormer. It groups potential features from point clouds into superpoints, and directly predicts instances through query vectors without relying on the results of object detection or semantic segmentation. The key step in this framework is a novel query decoder with transformers that can capture the instance information through the superpoint cross-attention mechanism and generate the superpoint masks of the instances. Through bipartite matching based on superpoint masks, SPFormer can implement the network training without the intermediate aggregation step, which accelerates the network. Extensive experiments on ScanNetv2 and S3DIS benchmarks verify that our method is concise yet efficient. Notably, SPFormer exceeds compared state-of-the-art methods by 4.3% on ScanNetv2 hidden test set in terms of mAP and keeps fast inference speed (247ms per frame) simultaneously. Code is available at https://github.com/sunjiahao1999/SPFormer. Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001 |
AAAI | 2 |
| 2023 | Text-Guided Generative Adversarial Network for Image Emotion Transfer
Siqi Zhu, Chunmei Qing, Xiangmin Xu 0001 |
ICIC (2) | 2 |
| 2023 | Multi-Scale Transformer Network for Saliency Prediction on 360-Degree ImagesabstractThe latest methods for saliency prediction on 360° images show that better results can be obtained using equirectangular (ERP) images as input. Due to the limitation of the receptive field, existing convolution-based networks cannot capture long-range information in complex 360° images. Although the transformer has the innate ability to capture long-range correlations with self-attention, large dataset requirement limit its application in saliency prediction of 360° images. In this paper, we present a novel Multi-scale Transformer framework for Saliency prediction on 360° images (MTSal360). The Multi-scale Transformer Module (MTM) is designed in the network to aggregate the contextual long-range information, which includes a Convolutional Positional Encoder (CPE) to enable the model could train and test on cubic and ERP format separately to address the insufficient data. Experiments on two public datasets illustrate that MTSal360 achieves better results over the state-of-the-art methods. Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001 |
ICIP | 2 |
| 2023 | Multi-depth Fusion Transformer and Batch Piecewise Loss for Visual Sentiment Analysis
Haochun Ou, Chunmei Qing, Jinglun Cen, Xiangmin Xu 0001 |
PRCV (10) | 2 |
| 2023 | Emotional generative adversarial network for image emotion transfer
Siqi Zhu, Chunmei Qing, Canqiang Chen, Xiangmin Xu 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Context-Based Adaptive Multimodal Fusion Network for Continuous Frame-Level Sentiment PredictionabstractRecently, video sentiment computing has become the focus of research because of its benefits in many applications such as digital marketing, education, healthcare, and so on. The difficulty of video sentiment prediction mainly lies in the regression accuracy of long-term sequences and how to integrate different modalities. In particular, different modalities may express different emotions. In order to maintain the continuity of long time-series sentiments and mitigate the multimodal conflicts, this paper proposes a novel Context-Based Adaptive Multimodal Fusion Network (CAMFNet) for consecutive frame-level sentiment prediction. A Context-based Transformer (CBT) module was specifically designed to embed clip features into continuous frame features, leveraging its capability to enhance the consistency of prediction results. Moreover, to resolve the multi-modal conflict between modalities, this paper proposed an Adaptive multimodal fusion (AMF) method based on the self-attention mechanism. It can dynamically determines the degree of shared semantics across modalities, enabling the model to flexibly adapt its fusion strategy. Through adaptive fusion of multimodal features, the AMF method effectively resolves potential conflicts arising from diverse modalities, ultimately enhancing the overall performance of the model. The proposed CAMFNet for consecutive frame-level sentiment prediction can ensure the continuity of long time-series sentiments. Extensive experiments illustrate the superiority of the proposed method especially in multimodal conflicts videos. Maochun Huang, Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Progressive Transformer Machine for Natural Character ReenactmentabstractCharacter reenactment aims to control a target person’s full-head movement by a driving monocular sequence that is made up of the driving character video. Current algorithms utilize convolution neural networks in generative adversarial networks, which extract historical and geometric information to iteratively generate video frames. However, convolution neural networks can merely capture local information with limited receptive fields and ignore global dependencies that play a crucial role in face synthesis, leading to generating unnatural video frames. In this work, we design a progressive transformer module that introduces multi-head self-attention with convolution refinement to simultaneously capture global-local dependencies. Specifically, we utilize the non-lapping window-based multi-head self-attention mechanism with hierarchical architecture to obtain the larger receptive fields at low-resolution feature map and thus extract global information. To better model local dependencies, we introduce the convolution operation to further refine the attentional weight in the multi-head self-attention mechanism. Finally, we use several stacked progressive transformer modules with the down-sampling operation to encode information of appearance information of previously generated frames and parameterized 3D face information of the current frame. Similarly, we use several stacked progressive transformer modules with the up-sampling operation to iteratively generate video frames. In this way, it can capture global-local information to facilitate generating video frames that are globally natural while preserving sharp outlines and rich detail information. Extensive experiments on several standard benchmarks suggest that the proposed method outperforms current leading algorithms. Yongzong Xu, Zhijing Yang, Tianshui Chen, Chunmei Qing |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | DFAEN: Double-order knowledge fusion and attentional encoding network for texture recognition
Zhijing Yang, Shujian Lai, Yukai Shi, Yongqiang Cheng 0001, Chunmei Qing |
Expert Syst. Appl. | 6 |
| 2022 | Intra- and Inter-Reasoning Graph Convolutional Network for Saliency Prediction on 360° ImagesabstractCubic projection can be utilized to divide 360° images into multiple rectilinear images, with little distortion. However, the existing saliency prediction models fail to integrate semantic information of these images. In this paper, we address this by proposing an intra- and inter-reasoning graph convolutional network for saliency prediction on 360° images (SalReGCN360). The whole framework contains six sub-networks, each of which contains two branches. In the training phase, after utilizing Multiple Cubic Projection (MCP), six rectilinear images are simultaneously put into corresponding sub-networks. In one of the branches, the global features of a single rectilinear image are extracted by the intra-graph inference module to finely predict local saliency of 360° images. In the other branch, the contextual features are extracted by the inter-graph inference module to effectively integrate semantic information of six rectilinear images. Finally, the feature maps are generated by the two branches fusion, and six corresponding rectilinear saliency maps are predicted. Extensive experiments on two popular saliency datasets illustrate the superiority of the proposed model, especially the improvement in KLD metric. Dongwen Chen, Chunmei Qing, Mengtao Ye, Xiangmin Xu 0001, Patrick Dickinson |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Cross Parallax Attention Network for Stereo Image Super-ResolutionabstractStereo super-resolution (SR) aims to enhance the spatial resolution of one camera view using additional information from the other. Previous deep-learning-based stereo SR methods indeed improved the SR performance effectively by employing additional information, but they are unable to super-resolve stereo images where there are large disparities, or different types of epipolar lines. Moreover, in these methods, one model can only super-solve images of a particular view, and for one specific scale factor. This paper proposes a cross parallax attention stereo super-resolution network (CPASSRnet) which can perform stereo SR of multiple scale factors for both views, with a single model. To overcome the difficulties of large disparity and different types of epipolar lines, a cross parallax attention module (CPAM) is presented, which captures the global correspondence of additional information for each view, relative to the other. CPAM allows the two views to exchange additional information with each other according to the generated attention maps. Quantitative and qualitative results compared with the state of the arts illustrate the superiority of CPASSRnet. Ablation experiments demonstrate that the proposed components are effective and noise tests verify the robustness of CPASSRnet. Canqiang Chen, Chunmei Qing, Xiangmin Xu 0001, Patrick Dickinson |
IEEE Trans. Multim. | 2 |
| 2021 | Hierarchical Lifelong Learning by Sharing Representations and Integrating HypothesisabstractIn lifelong machine learning (LML) systems, consecutive new tasks from changing circumstances are learned and added to the system. However, sufficiently labeled data are indispensable for extracting intertask relationships before transferring knowledge in classical supervised LML systems. Inadequate labels may deteriorate the performance due to the poor initial approximation. In order to extend the typical LML system, we propose a novel hierarchical lifelong learning algorithm (HLLA) consisting of two following layers: 1) the knowledge layer consisted of shared representations and integrated knowledge basis at the bottom and 2) parameterized hypothesis functions with features at the top. Unlabeled data is leveraged in HLLA for pretraining of the shared representations. We also have considered a selective inherited updating method to deal with intertask distribution shifting. Experiments show that our HLLA method outperforms many other recent LML algorithms, especially when dealing with higher dimensional, lower correlation, and fewer labeled data problems. Tong Zhang 0015, Guoxi Su, Chunmei Qing, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | SalBiNet360: Saliency Prediction on 360° Images with Local-Global Bifurcated Deep NetworkabstractWith the development of the virtual reality applications, predicting human visual attention on 360° images is valuable to content creators and encoding algorithms, and becomes essential to understand user behaviour. In this paper, we propose a local-global bifurcated deep network for saliency prediction on 360° images, which is named as SalBiNet360. In the global deep sub-network, multiple multi-scale contextual modules and a multilevel decoder are utilized to integrate the features from the middle and deep layers of the network. In the local deep sub-network, only one multi-scale contextual module and a single-level decoder are utilized to reduce the redundancy of local saliency maps. Finally, fused saliency maps are generated by linear combination of the global and local saliency maps. Experiments on two publicly available datasets illustrate that the proposed SalBiNet360 outperforms the tested state-of-the-art methods. Dongwen Chen, Chunmei Qing, Xiangmin Xu 0001, Huansheng Zhu |
VR | 2 |
| 2019 | Dimensionality reduction based on determinantal point process and singular spectrum analysis for hyperspectral imagesabstractDimensionality reduction is of high importance in hyperspectral data processing, which can effectively reduce the data redundancy and computation time for improved classification accuracy. Band selection and feature extraction methods are two widely used dimensionality reduction techniques. By integrating the advantages of the band selection and feature extraction, the authors propose a new method for reducing the dimension of hyperspectral image data. First, a new and fast band selection algorithm is proposed for hyperspectral images based on an improved determinantal point process (DPP). To reduce the amount of calculation, the dual‐DPP is used for fast sampling representative pixels, followed by k‐nearest neighbour‐based local processing to explore more spatial information. These representative pixel points are used to construct multiple adjacency matrices to describe the correlation between bands based on mutual information. To further improve the classification accuracy, two‐dimensional singular spectrum analysis is used for feature extraction from the selected bands. Experiments show that the proposed method can select a low‐redundancy and representative band subset, where both data dimension and computation time can be reduced. Furthermore, it also shows that the proposed dimensionality reduction algorithm outperforms a number of state‐of‐the‐art methods in terms of classification accuracy. Weizhao Chen, Zhijing Yang, Faxian Cao, Yijun Yan, Meilin Wang, Chunmei Qing, Yongqiang Cheng 0001 |
IET Image Process. | 6 |
| 2019 | Spatial-spectral classification of hyperspectral images: a deep learning framework with Markov Random fields based modellingabstractFor the spatial‐spectral classification of hyperspectral images (HSIs), a deep learning framework is proposed in this study, which consists of convolutional neural networks (CNNs) and Markov random fields (MRFs). Firstly, a CNN model to learn the deep spectral feature from the HSI is built and the class posterior probability distribution is estimated. The CNN with a dropout layer can relieve the overfitting in classification. The CNN is utilised as a pixel‐classifier, so it only works in the spectral domain. Then, the spatial information will be encoded by MRF‐based multilevel logistic prior for regularising the classification. To derive the correlation of both spectral and spatial features for improving algorithm performance, the marginal probability distribution in HSI is learned using MRF‐based loopy belief propagation. In comparison with several state‐of‐the‐art approaches for data classification on three publicly available HSI datasets, experimental results have demonstrated the superior performance of the proposed methodology. Chunmei Qing, Jiawei Ruan, Xiangmin Xu 0001, Jinchang Ren, Jaime Zabalza |
IET Image Process. | 1 |
| 2019 | Accelerating Flexible Manifold Embedding for Scalable Semi-Supervised LearningabstractIn this paper, we address the problem of large-scale graph-based semi-supervised learning for multi-class classification. Most existing scalable graph-based semi-supervised learning methods are based on the hard linear constraint or cannot cope with the unseen samples, which limits their applications and learning performance. To this end, we build upon our previous work flexible manifold embedding (FME) [1] and propose two novel linear-complexity algorithms called fast flexible manifold embedding (f-FME) and reduced flexible manifold embedding (r-FME). Both of the proposed methods accelerate FME and inherit its advantages. Specifically, our methods address the hard linear constraint problem by combining a regression residue term and a manifold smoothness term jointly, which naturally provides the prediction model for handling unseen samples. To reduce computational costs, we exploit the underlying relationship between a small number of anchor points and all data points to construct the graph adjacency matrix, which leads to simplified closed-form solutions. The resultant f-FME and r-FME algorithms not only scale linearly in both time and space with respect to the number of training samples but also can effectively utilize information from both labeled and unlabeled data. Experimental results show the effectiveness and scalability of the proposed methods. Suo Qiu, Feiping Nie 0001, Xiangmin Xu 0001, Chunmei Qing, Dong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Cartoon-to-Photo Facial Translation with Generative Adversarial NetworksabstractCartoon-to-photo facial translation could be widely used in different applications, such as law enforcement and anime remaking. Nevertheless, current general-purpose image-to-image models \ygyan{usually} %can only produce blurry or unrelated results in this task. In this paper, we propose a Cartoon-to-Photo facial translation with Generative Adversarial Networks (\name) for inverting cartoon faces to generate photo-realistic and related face images. In order to produce convincing faces with intact facial parts, we exploit global and local discriminators to capture global facial features and three local facial regions, respectively. Moreover, we use a specific content network to capture and preserve face characteristic and identity between cartoons and photos. As a result, the proposed approach can generate convincing high-quality faces that satisfy both the characteristic and identity constraints of input cartoon faces. Compared with recent works on unpaired image-to-image translation, our proposed method is able to generate more realistic and correlative images. Junhong Huang, Mingkui Tan, Yuguang Yan, Chunmei Qing, Qingyao Wu, Zhu Liang Yu |
ACML | 4 |
| 2018 | Camera identification based on very low bit rate videos with overall noise pattern having time varying statistics
Nili Tian, Bingo Wing-Kuen Ling, Chunmei Qing, Zhijing Yang |
Multim. Tools Appl. | 3 |
| 2017 | Robust object tracking based on sparse representation and incremental weighted PCA
Xiaofen Xing, Fuhao Qiu, Xiangmin Xu 0001, Chunmei Qing, Yinrong Wu |
Multim. Tools Appl. | 4 |
| 2016 | Efficient action recognition from compressed depth mapsabstractWe propose an efficient action recognition scheme based solely on compressed depth maps. Each depth map is coded by a recently proposed scalable encoder that employs multi-scale breakpoints and an adaptive discrete wavelet transform (DWT). DWT coefficients describe smooth variations in depth while breakpoints communicate sharp boundaries. Both of these attributes are extracted from the bit-stream and utilized to construct features which are subject to a classification scheme for human action recognition. By extracting features from the compressed bit-stream computational complexity is significantly reduced thereby making the proposed scheme suitable for real-time applications. A L2-regularized collaborative representation classifier is employed for classification. The proposed scheme is computationally more efficient when compared with conventional approaches. Experimental results on the MSR 3D action dataset validate the effectiveness and efficiency of our proposed scheme. Jie Miao, Xiaoyi Jia, Reji Mathew, Xiangmin Xu 0001, David S. Taubman, Chunmei Qing |
ICIP | 6 |
| 2016 | Image and video dehazing using view-based cluster segmentationabstractTo avoid distortion in sky regions and make the sky and white objects clear, in this paper we propose a new image and video dehazing method utilizing the view-based cluster segmentation. Firstly, GMM(Gaussian Mixture Model)is utilized to cluster the depth map based on the distant view to estimate the sky region and then the transmission estimation is modified to reduce distortion. Secondly, we present to use GMM based on Color Attenuation Prior to divide a single hazy image into K classifications, so that the atmospheric light estimation is refined to improve global contrast. Finally, online GMM cluster is applied to video dehazing. Extensive experimental results demonstrate that the proposed algorithm can have superior haze removing and color balancing capabilities. Chunmei Qing, Xiangmin Xu 0001, Bolun Cai |
VCIP | 2 |
| 2016 | Novel segmented stacked autoencoder for effective dimensionality reduction and feature extraction in hyperspectral imaging
Jaime Zabalza, Jinchang Ren, Jiangbin Zheng 0001, Huimin Zhao 0001, Chunmei Qing, Zhijing Yang, Peijun Du, Stephen Marshall |
Neurocomputing | 5 |
| 2016 | DehazeNet: An End-to-End System for Single Image Haze RemovalabstractSingle image haze removal is a challenging ill-posed problem. Existing methods use various constraints/priors to get plausible dehazing solutions. The key to achieve haze removal is to estimate a medium transmission map for an input hazy image. In this paper, we propose a trainable end-to-end system called DehazeNet, for medium transmission estimation. DehazeNet takes a hazy image as input, and outputs its medium transmission map that is subsequently used to recover a haze-free image via atmospheric scattering model. DehazeNet adopts convolutional neural network-based deep architecture, whose layers are specially designed to embody the established assumptions/priors in image dehazing. Specifically, the layers of Maxout units are used for feature extraction, which can generate almost all haze-relevant features. We also propose a novel nonlinear activation function in DehazeNet, called bilateral rectified linear unit, which is able to improve the quality of recovered haze-free image. We establish connections between the components of the proposed DehazeNet and those used in existing methods. Experiments on benchmark images show that DehazeNet achieves superior performance over existing methods, yet keeps efficient and easy to use. Bolun Cai, Xiangmin Xu 0001, Kui Jia, Chunmei Qing, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2016 | Simple to Complex Transfer Learning for Action RecognitionabstractRecognizing complex human actions is very challenging, since training a robust learning model requires a large amount of labeled data, which is difficult to acquire. Considering that each complex action is composed of a sequence of simple actions which can be easily obtained from existing data sets, this paper presents a simple to complex action transfer learning model (SCA-TLM) for complex human action recognition. SCA-TLM improves the performance of complex action recognition by leveraging the abundant labeled simple actions. In particular, it optimizes the weight parameters, enabling the complex actions to be learned to be reconstructed by simple actions. The optimal reconstruct coefficients are acquired by minimizing the objective function, and the target weight parameters are then represented as a combination of source weight parameters. The main advantage of the proposed SCA-TLM compared with existing approaches is that we exploit simple actions to recognize complex actions instead of only using complex actions as training samples. To validate the proposed SCA-TLM, we conduct extensive experiments on two well-known complex action data sets: 1) Olympic Sports data set and 2) UCF50 data set. The results show the effectiveness of the proposed SCA-TLM for complex action recognition. Fang Liu 0030, Xiangmin Xu 0001, Shuoyang Qiu, Chunmei Qing, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2015 | BIT: Bio-inspired trackerabstractVisual tracking is a challenging problem due to various factors such as deformation, rotation and illumination. As is well known, given the superior tracking performance of human vision, bio-inspired model is expected to improve the computer visual tracking. However, the design of bio-inspired tracking framework is challenging, due to the incomplete comprehension and hyper-scale of senior neurons, which will influence the effectiveness and real-time performance of the tracker. According to the ventral stream in visual cortex, a novel bio-inspired tracker (BIT) is proposed, which simulates shallow neurons (S1 and C1) to extract low-level bio-inspired feature for target appearance and imitates senior learning mechanism (S2 and C2) to combine generative and discriminative model for position estimation. In addition, Fast Fourier Transform (FFT) is adopted for real-time learning and detection in this framework. On the recent benchmark[1], extensive experimental results show BIT performs favorably against state-of-the-art methods in terms of accuracy and robustness. Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing, Chunmei Qing |
ICIP | 4 |
| 2015 | Temporal Variance Analysis for Action RecognitionabstractSlow feature analysis (SFA) extracts slowly varying signals from input data and has been used to model complex cells in the primary visual cortex (V1). It transmits information to both ventral and dorsal pathways to process appearance and motion information, respectively. However, SFA only uses slowly varying features for local feature extraction, because they represent appearance information more effectively than motion information. To better utilize temporal information, we propose temporal variance analysis (TVA) as a generalization of SFA. TVA learns a linear transformation matrix that projects multidimensional temporal data to temporal components with temporal variance. Inspired by the function of V1, we learn receptive fields by TVA and apply convolution and pooling to extract local features. Embedded in the improved dense trajectory framework, TVA for action recognition is proposed to: 1) extract appearance and motion features from gray using slow and fast filters, respectively; 2) extract additional motion features using slow filters from horizontal and vertical optical flows; and 3) separately encode extracted local features with different temporal variances and concatenate all the encoded features as final features. We evaluate the proposed TVA features on several challenging data sets and show that both slow and fast features are useful in the low-level feature extraction. Experimental results show that the proposed TVA features outperform the conventional histogram-based features, and excellent results can be achieved by combining all TVA features. Jie Miao, Xiangmin Xu 0001, Shuoyang Qiu, Chunmei Qing, Dacheng Tao |
IEEE Trans. Image Process. | 4 |
| 2014 | LDA based compact and discriminative dictionary learning for sparse codingabstractThe dictionary response usually affects the recognition results directly as it represents the original data and usually serves as the input of the classifier. However, the over-complete dictionary usually results in high dimensional response and redundancy. The application of the linear discriminant analysis (LDA)-based mapping method transforms the original dictionary response to be more discriminative for compact dictionary learning, resulting in high intra-class similarity and high inter-class dissimilarity in the response domain for better classification. By analyzing the recognition rate, the compactness and the purity, the proposed method can learn a small size of compact and discriminative dictionary with global optimization, and it can get a comparable or even better performance than the over-complete dictionary with much less computation cost. Experimental results demonstrate that the proposed approach also outperforms several recently proposed compact dictionary learning methods on human action recognition and object classification. Jiayong Chen, Xiangmin Xu 0001, Chunmei Qing, Jianxiu Jin |
ICIP | 3 |
| 2014 | High speed deep networks based on Discrete Cosine TransformationabstractThe traditional deep networks take raw pixels of data as input, and automatically learn features using unsupervised learning algorithms. In this configuration, in order to learn good features, the networks usually have multi-layer and many hidden units which lead to extremely high training time costs. As a widely used image compression algorithm, Discrete Cosine Transformation (DCT) is utilized to reduce image information redundancy because only a limited number of the DCT coefficients can preserve the most important image information. In this paper, it is proposed that a novel framework by combining DCT and deep networks for high speed object recognition system. The use of a small subset of DCT coefficients of data to feed into a 2-layer sparse auto-encoders instead of raw pixels. Because of the excellent decorrelation and energy compaction properties of DCT, this approach is proved experimentally not only efficient, but also it is a computationally attractive approach for processing high-resolution images in a deep architecture. Xiaoyi Zou, Xiangmin Xu 0001, Chunmei Qing, Xiaofen Xing |
ICIP | 3 |
| 2011 | Automatic nesting seabird detection based on boosted HOG-LBP descriptorsabstractSeabird populations are considered an important and accessible indicator of the health of marine environments: variations have been linked with climate change and pollution [1]. However, manual monitoring of large populations is labour-intensive, and requires significant investment of time and effort. In this paper, we propose a novel detection system for monitoring a specific population of Common Guillemots on Skomer Island, West Wales (UK). We incorporate two types of features, Histograms of Oriented Gradients (HOG) and Local Binary Pattern (LBP), to capture the edge/local shape information and the texture information of nesting seabirds. Optimal features are selected from a large HOG-LBP feature pool by boosting techniques, to calculate a compact representation suitable for the SVM classifier. A comparative study of two kinds of detectors, i.e., whole-body detector, head-beak detector, and their fusion is presented. When the proposed method is applied to the seabird detection, consistent and promising results are achieved. Chunmei Qing, Patrick Dickinson, Shaun W. Lawson, Robin Freeman |
ICIP | 1 |
| 2010 | An EDBoost algorithm towards robust face recognition in JPEG compressed domain
Chunmei Qing, Jianmin Jiang |
Image Vis. Comput. | 1 |
| 2010 | Normalized Co-Occurrence Mutual Information for Facial Pose Detection Inside VideosabstractHuman faces captured inside videos are often presented with variable poses, making it difficult to recognize and thus pose detection becomes crucial for such face recognition under non-controlled environment. While existing mutual in formation (MI) primarily considers the relationship between corresponding individual pixels, we propose a normalized co occurrence mutual information in this letter to capture the information embedded not only in corresponding pixel values but also in their geographical locations. In comparison with the existing Mis, the proposed presents an essential advantage that both marginal entropy and joint entropy can be optimally exploited in measuring the similarity between two given images. When developed into a facial pose detection algorithm inside video sequences, we show, through extensive experiments, that such design is capable of achieving the best performances among all the representative existing techniques compared. Chunmei Qing, Jianmin Jiang, Zhijing Yang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | A Manifolded AdaBoost for Face Recognition
Chunyuan Lu, Jianmin Jiang, Guo-Can Feng, Chunmei Qing |
KES (1) | 4 |