VLDB 2026 Research / reviewers in the wild / expert
Zhong-Qiu Zhao
dblp:17/1876 · also Zhongqiu Zhao
· DBLP profile ↗
98ranked-venue papers
20as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 14 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 33 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 18 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Q-FETA:Question-Guided Fine-Grained Feature Transport Alignment for Robust Visual Question Answering
Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (9) | 3 |
| 2026 | A novel multi-granularity context-adaptive downsampling convolution for low-resolution images
Zejun Gu, Zhong-Qiu Zhao, Hao Shen 0006, Zhao Zhang 0001, De-Shuang Huang |
Inf. Sci. | 2 |
| 2026 | Semantic attention and progressive training for image authenticity assessment and tampered region highlighting
Bo Wang 0072, Qi Si, Yang Zhao 0002, Zhong-Qiu Zhao, Zhao Zhang 0001 |
Inf. Sci. | 5 |
| 2026 | Human-Structure-Aware Token Position Embedding for Tokenized Pose EstimationabstractTokenized pose estimation (TPE) has demonstrated remarkable performance in lightweight human pose estimation (HPE) models. However, existing TPE methods typically initialize keypoint tokens randomly, without explicitly incorporating human structure priors. These priors play a vital role in HPE by effectively mitigating common challenges such as occlusion and ambiguity. To this end, we propose a Structure-Aware Keypoint Position Embedding (SAKPE). This embedding explicitly encodes inherent structural properties of the human body, such as symmetry and order, into the positional coordinates of keypoint tokens. It also employs learnable scale and offset factors to adapt to diverse human poses, thereby fully exploiting the geometric constraints among keypoints. Furthermore, to better leverage the positional relationships among patch tokens, we introduce a Layer-adaptive Hybrid Patch Position Embedding (LHPPE). It dynamically fuses absolute and relative position embeddings of patch tokens based on attention distributions across Transformer layers, enabling the model to learn both absolute and relative positional information adaptively. Taking the two together, we propose a novel position embedding method for pose estimation, named Human-structure-aware Token Position Embedding (HTPE). It significantly improves the performance of various TPE models. Extensive experiments on COCO, CrowdPose, and OCHuman show that HTPE achieves state-of-the-art (SOTA) performance among lightweight methods, with a negligible increase in parameters and FLOPs. Notably, it demonstrates consistent improvements under occlusion,, achieving up to 3.3 AP gains. The source code can be found in https://github.com/guzejungithub/HTPE. Zejun Gu, Zhong-Qiu Zhao, Henghui Ding, Hao Shen 0006, Zhenhua Tang 0001, Zhao Zhang 0001, De-Shuang Huang |
IEEE Trans. Image Process. | 2 |
| 2026 | Cross-Domain Knowledge Distillation for Low-Resolution Human Pose EstimationabstractIn practical applications of human pose estimation, low-resolution inputs frequently occur, and existing state-of-the-art models perform poorly with low-resolution images. This work focuses on boosting the performance of low-resolution models by distilling knowledge from a high-resolution model. However, we face the challenge of feature size mismatch and class number mismatch when applying knowledge distillation to networks with different input resolutions. To address this issue, we propose a novel cross-domain knowledge distillation (CDKD) framework. In this framework, we construct a scale-adaptive projector ensemble (SAPE) module to spatially align feature maps between models of varying input resolutions. It adopts a projector ensemble to map low-resolution features into multiple common spaces and adaptively merges them based on multi-scale information to match high-resolution features. Additionally, we construct a cross-class alignment (CCA) module to solve the problem of the mismatch of class numbers. By combining an easy-to-hard training (ETHT) strategy, the CCA module further enhances the distillation performance. The effectiveness and efficiency of our approach are demonstrated by extensive experiments on three common benchmark datasets: MPII, COCO, and Crowdpose. Zejun Gu, Zhong-Qiu Zhao, Henghui Ding, Hao Shen 0006, Zhao Zhang 0001, De-Shuang Huang |
IEEE Trans. Multim. | 2 |
| 2025 | BadHPE: Backdoor Attack for Precise Manipulation of Human Pose Estimation Models
Yuze Wang, Zhong-Qiu Zhao |
CGI (3) | 4 |
| 2025 | PsSR: Hybrid Path Selection Mechanism for Efficient Image Super-ResolutionabstractIn practical applications, large images are processed in patches, many of which are simple and smooth, making them suitable for lighter network processing. This paper proposes a hybrid path selection mechanism that enables the dynamic routing of patches through the network, achieving a balance between accuracy and computational efficiency. The proposed method develops a dual-branch architecture, in which the two branches keep the same structural design, but differ in the number of channels. An efficient classification module is integrated for directing patches to the appropriate branches within the model. In addition, within each branch, we introduce an early exit strategy to further reduce the calculation cost. Subsequently, we enhance the restoration capabilities of the lightweight branch through channel and spatial feature distillation. Experimental results demonstrate that our Hybrid Path Selection based Super-Resolution (PsSR) mechanism effectively reduces computational costs while preserving the integrity of the backbone performance. Zhong-Qiu Zhao, Yang Zhao 0002 |
ICASSP | 2 |
| 2025 | High-Fidelity Stereoscopic Image Rain Removal with Texture Integrity and Disparity ConsistencyabstractThis paper tackles the challenge of stereoscopic image rain removal by focusing on enhancing texture integrity and disparity consistency. Existing stereoscopic rain removal techniques often fall short due to 1) disruptions in texture coherence caused by complex rain streaks, and 2) inaccuracies in disparity estimation from inadequate feature fusion. To overcome these limitations, we introduce the StereoIRR method, which incorporates: 1) a Long-range and Cross-view Interaction (LCI) framework that preserves texture integrity by mitigating rain’s adverse effects on stereoscopic features, and 2) a Dual-view Mutual Attention mechanism that ensures disparity consistency by generating precise mutual attention maps for cross-view feature fusion. Our approach not only maintains the integrity of stereoscopic textures but also significantly reduces errors in disparity estimation. Extensive experiments demonstrate that StereoIRR consistently outperforms state-of-the-art monocular and stereoscopic methods on multiple benchmark datasets. Yanyan Wei, Zhao Zhang 0001, Zhong-Qiu Zhao, Yang Zhao 0002, Richang Hong, Yi Yang 0001, Meng Wang 0001 |
ICASSP | 3 |
| 2025 | Cross-Resolution Deep Face Recognition via Collaborative Knowledge Distillation
Weidong Tian 0001, Zejun Gu, Zhong-Qiu Zhao |
ICIC (17) | 4 |
| 2025 | SAM2-Cap: Segment Anything 2 with using Parts and Object Spatial Hierarchical Relationships for Image SegmentationabstractImage segmentation has found widespread applications in computer vision, particularly in fields such as medical image analysis, autonomous driving, and video surveillance. However, as the complexity of segmentation tasks increases, new vision foundation models have emerged, notably the Segment Anything Model (SAM) series. By leveraging large-scale training and cross-domain generalization, SAM2 has significantly enhanced the performance of image segmentation. However, despite these advancements, SAM2 still generates class-agnostic segmentation results in the absence of manual prompts. To enhance the SAM2's performance in these scenarios, we propose an effective framework, termed SAM2-Cap. Specifically, we integrate the backbone of SAM2 with a Capsule Autoencoder, exploiting the inherent properties of capsule networks to effectively capture and decode the spatial hierarchical relationships between the parts and objects. This improves the modeling of object shapes and spatial relationships, enhancing the understanding and segmentation accuracy of complex objects while reducing the reliance on prompt information. Extensive experiments on Prostate MRI segmentation, Polyp segmentation and Cityscapes dataset, demonstrate our method achieves competitive results compared to state-of-the-art models. Xiufeng Liu 0005, Zhong-Qiu Zhao, Yi Yang 0001, Donghui Hu, Zhao Zhang 0001 |
ICME | 2 |
| 2025 | SwinCAE: Capsule Autoencoder using Shifted Windows for 3D Human Pose EstimationabstractEstimating 3D human poses from monocular videos is a challenging task, primarily due to self-occlusion. Many existing methods struggle with unseen viewpoints as they rely on large amounts of data rather than enhancing their generalization ability across different viewpoints. To overcome this limitation, we propose a novel approach using a capsule autoencoder integrated with the shifted-windows model (SwinCAE), which can enhance prediction accuracy by effectively capturing the spatial hierarchical relationship between the parts and objects. Furthermore, we build a Parallel Double Attention with Shifted Windows module to enhance computational efficiency and modeling capacity. Additionally, we construct a Multi-Attention Collaborative module to capture diverse information, including both coarse and fine details. Through the collaboration of these modules, the model representation is significantly improved, resulting in a more accurate generated pose. Extensive experiments demonstrate that SwinCAE achieves better or comparable results to state-of-the-art models about 3D human pose estimation task. Xiufeng Liu 0005, Zhong-Qiu Zhao, Yi Yang 0001, Donghui Hu, Zhao Zhang 0001 |
ICME | 2 |
| 2025 | Capsule network with using shifted windows for 3D human pose estimation
Xiufeng Liu 0005, Zhong-Qiu Zhao, Weidong Tian 0001, Hongmei He |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | When low-light meets flares: Towards Synchronous Flare Removal and Brightness Enhancement
Jiahuan Ren, Zhao Zhang 0001, Suiyi Zhao, Jicong Fan 0001, Zhong-Qiu Zhao, Yang Zhao 0002, Richang Hong, Meng Wang 0001 |
Neural Networks | 5 |
| 2025 | Spatial Frequency Modulation Network for Efficient Image DehazingabstractCurrently, two main research lines in efficient context modeling for image dehazing are tailoring effective feature modulation mechanisms and utilizing the Fourier transform more precisely. The former is usually based on self-scale features that ignore complementary cross-scale/level features, and the latter tends to overlook regions with pronounced haze degradation and intricate structures. This paper introduces a novel spatial and frequency modulation perspective to synergistically investigate contextual feature modeling for efficient image dehazing. Specifically, we delicately develop a Spatial Frequency Modulator (SFM) equipped with a Cross-Scale Modulator (CSM) and Frequency Modulator (FM) to implement intra-block feature modulation. The CSM progressively aggregates hierarchical features across different scales, employing them for spatial self-modulation, and the FM subsequently adopts a dual-branch design to focus more on the crucial areas with severe haze and complex structures for reconstruction. Further, we propose a Cross-Level Modulator (CLM) to facilitate inter-block feature mutual modulation, enhancing seamless interaction between features at different depths and layers. Integrating the above-developed modules into the U-Net architecture, we construct a two-stage spatial frequency modulation network (SFMN). Extensive quantitative and qualitative evaluations showcase the superior performance and efficiency of the proposed SFMN over recent state-of-the-art image dehazing methods. The source code can be found in https://github.com/it-hao/SFMN. Hao Shen 0006, Henghui Ding, Yulun Zhang 0001, Zhong-Qiu Zhao, Xudong Jiang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Adaptive branch selection for accelerate image super-resolution
Zhong-Qiu Zhao, Hao Shen 0006, Xiufeng Liu 0005 |
Vis. Comput. | 2 |
| 2024 | High-Fidelity Diffusion Editor for Zero-Shot Text-Guided Video EditingabstractText-guided image generative diffusion models achieve fast development on the generation and editing of high-quality images. To extend such success to video editing, some efforts combining image generation with video editing have been made, which however only achieve inferior performance. We attribute it to two challenges: 1) different from the static image generation, it is tricky for dynamic video information to ensure the temporal fidelity of motion consistency across different frames; 2) the randomness of the frame generation process makes it hard to continuously retain the similar spatial fidelity for the original detailed features. In this paper, we propose a new high-fidelity diffusion model-based zero-shot text-guided video editing network, called HiFiVEditor, which aims to conduct effective video editing with high fidelity of the original video's detailed and dynamic information. Specifically, we propose a Spatial-Temporal Fidelity Block (STFB) that enables the model to restore the spatial features by enlarging the spatial perceptual field to avoid loss of important information, and capture more dynamic information between different frames by using all frames for preserving temporal consistency to achieve better temporal fidelity. In addition, we introduce Null-Text Embedding to create a soft text embedding to optimize the noise learning process, so that the latent noise can be aligned with the prompt. Furthermore, to tune the video style and render it more realistic, we employ a Prior-Guided Perceptual Loss to constrain the prediction results to avoid deviating from the original video style. Extensive experiments demonstrate the superior video editing capability compared to existing works. Yan Luo 0004, Zhichao Zuo, Zhao Zhang 0001, Zhong-Qiu Zhao, Haijun Zhang 0002, Richang Hong |
ICDM | 4 |
| 2024 | Dual Cross-Stage Partial Learning for Detecting Objects in Dehazed ImagesabstractPerforming an object detection task after the restoration of a hazy image, or rather detecting with the network backbone directly, will result in the inclusion of information mixed with dehazing, which tends to interfere with detection performance. To address these issues, we propose a novel framework for detecting objects in dehazed images via Dual Cross-Stage Partial Learning (DCSP). Specifically, we introduce a Cross-Stage Partial (CSP) module for extracting clean feature information after dehazing. Secondly, to enhance data integrity, we employ a skip-input strategy to supplement information related to object detection features that may be lost during the dehazing task, while avoiding the gradient vanishing problem. In addition, CSP is also introduced to facilitate comprehensive learning of multiple feature representations. Finally, to avoid the inclusion of irrelevant dehazing information in detection, we apply a Ground-Truth Flow at detection network (at dark3), for fine feature information calibration. Additionally, we created a synthetic fog dataset to expand the training data for DCSP. Experimental results on both synthetic and real-world datasets demonstrate the effectiveness and accuracy of the proposed method. The code is available at https://github.com/zhaojinbiao/DCSP. Jinbiao Zhao 0001, Zhao Zhang 0001, Jiahuan Ren, Haijun Zhang 0002, Zhong-Qiu Zhao, Meng Wang 0001 |
ICDM | 5 |
| 2024 | Image Captioning with Masked Diffusion Model
Weidong Tian 0001, Wenzheng Xu, Junxiang Zhao, Zhong-Qiu Zhao |
ICIC (8) | 4 |
| 2024 | Dual-Branch Collaborative Learning for Visual Question Answering
Weidong Tian 0001, Junxiang Zhao, Wenzheng Xu, Zhong-Qiu Zhao |
ICIC (3) | 4 |
| 2024 | Contextual Feature Modulation Network for Efficient Super-Resolution
Wandi Zhang, Hao Shen 0006, Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (6) | 5 |
| 2024 | DDNet: Detection-Focused Dehazing Network
Weidong Tian 0001, Wandi Zhang, Zhong-Qiu Zhao |
ICIC (10) | 4 |
| 2024 | Style-ACAE: Adversarial Capsule Autoencoder with StylesabstractCapsule networks get achievements in many computer vision tasks. However, in the field of image generation, they have huge room for improvement compared to the mainstream models. This is because capsule networks can not fully parse useful features and has limited capabilities of modeling the hierarchical and geometrical structure of the object in background noise. To tackle these issues, we propose a novel capsule autoencoder that can learn the part-object spatial hierarchical features, and we dub it Adversarial Capsule Autoencoder with Styles (Style-ACAE). Specifically, Style-ACAE decomposes the object into a set of semantic-consistent part-level descriptions and then assembles them into object-level descriptions to build the hierarchy. Furthermore, we effectively apply the modified generator structure, which introduces novel style modulation and demodulation. The new generator handles long-range dependency of part-object and captures the global structure of the object. This is the first case of the capsule network for image generation on commonly used benchmarks. The experimental results show that Style-ACAE can generate high-quality images and has competitive performance to the state-of-the-art generative models. Xiufeng Liu 0005, Zhong-Qiu Zhao |
ICME | 2 |
| 2024 | GSLip: A Global Lip-Reading Framework with Solid Dilated ConvolutionsabstractThe mainstream lip-reading framework employs Residual Network (ResNet) for spatial feature extraction and Temporal Convolutional Network (TCN) for constructing the temporal model. Addressing the modes of single-frame spatial and multi-frame temporal information extraction, we analyze and propose respective improvements. When extracting single-frame lip information, it is crucial to consider not only local fine-grained information but also global coarse-grained information. Global features such as lip contour, size, and movement amplitude significantly contribute to extracting local fine-grained lip features. However, traditional convolution, a key component of ResNet, is primarily utilized to extract features based on sliding windows in small regions, making ResNet less effective at capturing global features. Additionally, we observe key motion changes appearing in only a few consecutive frames in lipreading task. For multi-frame temporal information extraction, TCN is found to be inadequate in handling local continuous temporal dependencies. Therefore, we develop GSLip, a global lip-reading framework with solid dilated convolutions. The key ideas of GSLip include: 1) The Residual Global Context Network (ResGNet), which addresses the deficiency of traditional convolution in extracting global features; 2) the Continuous Temporal Convolutional Network (C-TCN), that enhances the ability to extract continuous temporal information. With the incorporation of these modules, our method has demonstrated excellent results on the Lip Reading in the Wild (LRW) and LRW-1000 datasets. Junxia Jiang, Zhong-Qiu Zhao, Weidong Tian 0001 |
IJCNN | 2 |
| 2024 | Global routing between capsules
Hao Shen 0006, Zhong-Qiu Zhao, Yi Yang 0036, Zhao Zhang 0001 |
Pattern Recognit. | 3 |
| 2024 | Spatial-Frequency Adaptive Remote Sensing Image Dehazing With Mixture of ExpertsabstractThe feature modulation mechanism has been demonstrated to be particularly well-suited for efficient network design and is rarely explored in remote sensing dehazing tasks. Moreover, we observe distinct patterns in haze distribution across the low-frequency (LF) and high-frequency (HF) components of haze images from various datasets. However, existing research rarely investigated the potential solution in the frequency domain. In response, we propose a novel spatial-frequency adaptive network (SFAN), which is mainly built by the proposed mixture of modulation experts (MoME) and decoupled frequency learning block (DFLB). Different from the fixed feature modulation design used in other tasks, the MoME adopts the mixture-of-expert mechanism to dynamically learn diverse contextual features of various granularities and scales in a sample-adaptive manner and then utilize them to perform elementwise local feature modulation. This pure convolution architecture enables our network to have superior performance and efficiency tradeoffs. Furthermore, the DFLB is devised to facilitate the LF global haze removal and reconstruction of HF local texture information. At the micro level, we first utilize a mask extractor (ME) to generate the frequency mask from the input hazy image, then employ a dual-branch decoupled learning unit to boost frequency learning, and finally develop a mixture of fusion experts (MoFE) to achieve HF and LF feature interaction. Extensive experiments on publicly available dehazing datasets demonstrate that our network performs superior performance while incurring lower computational costs. Compared to the state-of-the-art approach (DEA-Net), SFAN achieves, an average, 0.83-dB PSNR improvement on five remote sensing datasets but consumes only 51% of the FLOPs. The code will be available athttps://github.com/it-hao/SFAN. Hao Shen 0006, Henghui Ding, Yulun Zhang 0001, Xiaofeng Cong, Zhong-Qiu Zhao, Xudong Jiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | A Lightweight, Effective, and Efficient Model for Label Aggregation in CrowdsourcingabstractDue to the presence of noise in crowdsourced labels, label aggregation (LA) has become a standard procedure for post-processing these labels. LA methods estimate true labels from crowdsourced labels by modeling worker quality. However, most existing LA methods are iterative in nature. They require multiple passes through all crowdsourced labels, jointly and iteratively updating true labels and worker qualities until a termination condition is met. As a result, these methods are burdened with high space and time complexities, which restrict their applicability in scenarios where scalability and online aggregation are essential. Furthermore, defining a suitable termination condition for iterative algorithms can be challenging. In this article, we view LA as a dynamic system and represent it as a Dynamic Bayesian Network. From this dynamic model, we derive two lightweight and scalable algorithms: LAonepassand LAtwopass. These algorithms can efficiently and effectively estimate worker qualities and true labels by traversing all labels at most twice, thereby eliminating the need for explicit termination conditions and multiple traversals over the crowdsourced labels. Due to their dynamic nature, the proposed algorithms are also capable of performing label aggregation online. We provide theoretical proof of the convergence property of the proposed algorithms and bound the error of the estimated worker qualities. Furthermore, we analyze the space and time complexities of our proposed algorithms, demonstrating their equivalence to those of majority voting. Through experiments conducted on 20 real-world datasets, we demonstrate that our proposed algorithms can effectively and efficiently aggregate labels in both offline and online settings, even though they traverse all labels at most twice. The code is on https://github.com/yyang318/LA_onepass . Yi Yang 0036, Zhong-Qiu Zhao, Gong-Qing Wu, Xingrui Zhuo, Qing Liu 0001, Quan Bai 0001, Weihua Li 0007 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2023 | Adaptive Dynamic Filtering Network for Image DenoisingabstractIn image denoising networks, feature scaling is widely used to enlarge the receptive field size and reduce computational costs. This practice, however, also leads to the loss of high-frequency information and fails to consider within-scale characteristics. Recently, dynamic convolution has exhibited powerful capabilities in processing high-frequency information (e.g., edges, corners, textures), but previous works lack sufficient spatial contextual information in filter generation. To alleviate these issues, we propose to employ dynamic convolution to improve the learning of high-frequency and multi-scale features. Specifically, we design a spatially enhanced kernel generation (SEKG) module to improve dynamic convolution, enabling the learning of spatial context information with a very low computational complexity. Based on the SEKG module, we propose a dynamic convolution block (DCB) and a multi-scale dynamic convolution block (MDCB). The former enhances the high-frequency information via dynamic convolution and preserves low-frequency information via skip connections. The latter utilizes shared adaptive dynamic kernels and the idea of dilated convolution to achieve efficient multi-scale feature extraction. The proposed multi-dimension feature integration (MFI) mechanism further fuses the multi-scale features, providing precise and contextually enriched feature representations. Finally, we build an efficient denoising network with the proposed DCB and MDCB, named ADFNet. It achieves better performance with low computational complexity on real-world and synthetic Gaussian noisy datasets. The source code is available at https://github.com/it-hao/ADFNet. Hao Shen 0006, Zhong-Qiu Zhao, Wandi Zhang |
AAAI | 2 |
| 2023 | Lightweight Human Pose Estimation Based on Densely Guided Self-Knowledge Distillation
Mingyue Wu, Zhong-Qiu Zhao, Weidong Tian 0001 |
ICANN (2) | 2 |
| 2023 | Dynamic Attention Filter Capsule Network for Medical Images Segmentation
Zhong-Qiu Zhao |
ICIC (2) | 3 |
| 2023 | Cross-Scale Dynamic Alignment Network for Reference-Based Super-Resolution
Zhong-Qiu Zhao |
ICIC (2) | 3 |
| 2023 | Hierarchical Graph Neural Network for Human Pose EstimationabstractThe human bodies are hierarchical structure and human pose estimation is highly dependent on the correlation between the keypoints of human body. Most existing CNN-based and Transformer-based methods perform well in the visual representation but lack the ability to learn the correlation between keypoints explicitly. To address this problem, we propose a human pose estimation method based on a hierarchical graph neural network with dynamic keypoint weight. The hierarchical graph neural network consists of three parts: the inter-layer process, hierarchical process, and iterative inference. The inter-layer process and hierarchical process learn the correlation between similar keypoints and further keypoints, respectively. And iterative inference is used to refine and aggregate more useful information between keypoints. Furthermore, the dynamic keypoint weight network gives each keypoint a different degree of importance. Extensive experiments demonstrate that our HGNPose-M and HGNPose-L achieve 76.5 AP $(\color{Red}{\uparrow 2.1})$ and 77.4 AP $(\color{Red}{\uparrow 1.6})$ on the COCO dataset respectively. Guanghua Zheng, Zhong-Qiu Zhao, Zhao Zhang 0001, Yi Yang 0036 |
ICME | 2 |
| 2023 | Semantic Object Alignment and Region-Aware Learning for Change CaptioningabstractThe change captioning task, a downstream task of the image captioning task, is an emerging deep learning task. It aims to output sentences to describe the differences between two images, one of which is the original image and the other is the image changed from the original. Many current methods for generating disparity descriptions are based on the encoder-decoder model. The process is to first obtain the grid features using pre-trained ResNet on the image and encode the features, and then to obtain the sentences via a decoder. However, using traditional methods of encoding area features and applying them to grid features regardless can lead to performance degradation. Therefore, we propose a new model incorporating a novel design that is more robust to describe variations of image pairs, and a new addition of pseudo-region features that reduce the information missing from the grid features. Intensive experiments we have done can adequately demonstrate that our proposed model achieves state-of-the-art performance. Weidong Tian 0001, Quan Ren, Zhong-Qiu Zhao, Ruihua Tian |
IJCNN | 3 |
| 2023 | An Answer FeedBack Network for Visual Question AnsweringabstractRecent advances have explored the power of transformer architecture in Visual Question Answering(VQA). However, most of the models suffer from misalignment of multimodal features, and they focus on unimportant image regions when answering the given questions. To address this, in this paper, we propose an Answer FeedBack Network (AFBN) to focus on image region features that are more beneficial for answering questions. The generate answers of the backbone network are again inputted into the network as feedback information. Then, we propose a FeedBack module (FB) to control the answer feedback. Additionally, we adopt the consistency loss function to reconstruct the image region features. By this function, the model can ensure the same of the image region features related to the question or answer. Extensive experiments on VQA-v2 benchmark dataset show that our method achieves better performance than the state-of-the-art methods. Weidong Tian 0001, Ruihua Tian, Zhong-Qiu Zhao, Quan Ren |
IJCNN | 3 |
| 2023 | Dynamic Grouped Interaction Network for Low-Light Stereo Image EnhancementabstractLow-Light Stereo Image Enhancement (LLSIE) tackles the challenge of improving the illumination and restoring the details in stereo images. However, existing deep learning-based LLSIE methods trained on high-resolution low-light images often exhibit sub-optimal performance when interacting with information from the left and right views. We find that this is because of: (1) the high computational cost arising from quadratic complexity, which hinders the enhancement model's ability to process high-resolution images; and (2) the limitations of conventional fusion strategies in previous work, which inadequately capture cross-view cues, resulting in weak feature representation and compromised detail recovery. To address these limitations, we propose a novel Dynamic Grouped Interaction Network (DGI-Net) to enhance illumination and recover more details while reducing the computational cost. Specifically, DGI-Net employs the U-Net structure, which effectively mitigates noise during the low-light enhancement. Furthermore, we design a Grouped Stereo Interaction Module (GSIM) with a grouping strategy to efficiently discover cross-view cues while minimizing computations. To dynamically fuse stereo information and fully exploit cross-view correlations, we also introduce a Dynamic Embedding Module (DEM) to establish dynamic connections between inter-view cues and intra-view features, which performs dynamic weight processing on cross-view cues to eliminate noise during fusion. For intra-view processing, we present a Diversity Enhanced Block (DEB) to extract multi-scale features, thereby improving diversity and feature representation. This multi-scale feature extraction also addresses low image contrast in dark lighting conditions. Experimental results demonstrate that DGI-Net outperforms current state-of-the-art methods in low-light stereo image enhancement. Baiang Li, Zhao Zhang 0001, Yang Zhao 0002, Zhong-Qiu Zhao, Haijun Zhang 0002 |
ACM Multimedia | 5 |
| 2023 | Mutual Information-driven Triple Interaction Network for Efficient Image DehazingabstractMulti-stage architectures have exhibited efficacy in image dehazing, which usually decomposes a challenging task into multiple more tractable sub-tasks and progressively estimates latent hazy-free images. Despite the remarkable progress, existing methods still suffer from the following shortcomings: (1) limited exploration of frequency domain information; (2) insufficient information interaction; (3) severe feature redundancy. To remedy these issues, we propose a novel Mutual Information-driven Triple interaction Network (MITNet) based on spatial-frequency dual domain information and two-stage architecture. To be specific, the first stage, named amplitude-guided haze removal, aims to recover the amplitude spectrum of the hazy images for haze removal. And the second stage, named phase-guided structure refined, devotes to learning the transformation and refinement of the phase spectrum. To facilitate the information exchange between two stages, an Adaptive Triple Interaction Module (ATIM) is developed to simultaneously aggregate cross-domain, cross-scale, and cross-stage features, where the fused features are further used to generate content-adaptive dynamic filters so that applying them to enhance global context representation. In addition, we impose the mutual information minimization constraint on paired scale encoder and decoder features from both stages. Such an operation can effectively reduce information redundancy and enhance cross-stage feature complementarity. Extensive experiments on multiple public datasets exhibit that our MITNet performs superior performance with lower model complexity. The code and models are available at https://github.com/it-hao/MITNet. Hao Shen 0006, Zhong-Qiu Zhao, Yulun Zhang 0001, Zhao Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | Dual Capsule Attention Mask Network with Mutual Learning for Visual Question AnsweringabstractA Visual Question Answering (VQA) model processes images and questions simultaneously with rich semantic information. The attention mechanism can highlight fine-grained features with critical information, thus ensuring that feature extraction emphasizes the objects related to the questions. However, unattended coarse-grained information is also essential for questions involving global elements. We believe that global coarse-grained information and local fine-grained information can complement each other to provide richer comprehensive information. In this paper, we propose a dual capsule attention mask network with mutual learning for VQA. Specifically, it contains two branches processing coarse-grained features and fine-grained features, respectively. We also design a novel stackable dual capsule attention module to fuse features and locate evidence. The two branches are combined to make final predictions for VQA. Experimental results show that our method outperforms the baselines in terms of VQA performance and interpretability and achieves new SOTA performance on the VQA-v2 dataset. Weidong Tian 0001, Zhong-Qiu Zhao |
COLING | 3 |
| 2022 | Lip Reading Using Deformable 3D Convolution and Channel-Temporal Attention
Jie Chai, Zhong-Qiu Zhao, Housen Zhang, Weidong Tian 0001 |
ICANN (4) | 4 |
| 2022 | Lipreading Model Based On Whole-Part Collaborative LearningabstractLipreading is a task to recognize speech content from visual information of the speaker’s lips movements. Recently, some work has focused more on how to adequately extract temporal information, and spatial information is simply used after extraction. In this paper, we focus on the full use of spatial information in lipreading tasks. The whole lip represents global spatial information, while the parts of the lip contain fine-grained spatial information. We propose the lipreading model based on whole-part collaborative learning (WPCL), which can help this model make full use of both global and fine-grained spatial information of the lip. WPCL contains two branches, which deal with the whole and the part features respectively and are trained jointly by collaborative learning. Further, in order to highlight the different importance of part features when fusing them, we propose an adaptive part features fusion module (APFF) to fusion part features. Finally, we prove our viewpoints and evaluate our WPCL by severed experiments. Experiments on LRW and CAS-VSR-W1k datasets demonstrate that our approach achieves state-of-the-art performance. Weidong Tian 0001, Housen Zhang, Zhong-Qiu Zhao |
ICASSP | 4 |
| 2022 | A Sub-captions Semantic-Guided Network for Image Captioning
Weidong Tian 0001, Jun-jun Zhu, Zhong-Qiu Zhao, Yu-Zheng Zhang |
ICIC (3) | 4 |
| 2022 | Improved YOLOv5 Network with Attention and Context for Small Object Detection
Jie Chai, Zhong-Qiu Zhao, Weidong Tian 0001 |
ICIC (3) | 4 |
| 2022 | Training Super-Resolution Network with Difficulty-Based Adaptive SamplingabstractThe performance of super-resolution (SR) networks is highly dependent on the quality and size of the training data. However, research on better use of the available SR dataset remains unexplored. In this work, we propose Difficulty-based Adaptive Sampling (DAS) strategy to fill this gap. Specifically, to further exploit the input samples, DAS first uses a Calculating and Sorting module (CS module) to calculate the upsampling difficulty of the input samples and sorts them. The CS module makes DAS be efficient with an online form and using a fast method to calculate the difficulty degrees of samples. Then it uses a Sampler module to select appropriate samples. Finally, with a Recorder module to record the training states, DAS can dynamically select the samples suitable for different training stages. Extensive experiments demonstrate that DAS selecting the appropriate samples for each iteration can effectively improve the performance of SR networks. Zhong-Qiu Zhao |
ICME | 2 |
| 2022 | Optimization algorithms in wireless monitoring networks: A survey
Na Xia, Huaizhen Peng, Zhong-Qiu Zhao, Huazheng Du, Yongtang Yu |
Neurocomputing | 4 |
| 2022 | Joint operation and attention block search for lightweight image restoration
Hao Shen 0006, Zhong-Qiu Zhao, Wenrui Liao, Weidong Tian 0001, De-Shuang Huang |
Pattern Recognit. | 2 |
| 2021 | Multi-Branch Network for Small Human Pose Estimation
Yuchen Ge, Zhong-Qiu Zhao, Weidong Tian 0001, Hai Min |
ICANN (3) | 2 |
| 2021 | Visual-Textual Semantic Alignment Network for Visual Question Answering
Weidong Tian 0001, Yuzheng Zhang, Junjun Zhu, Zhong-Qiu Zhao |
ICANN (5) | 5 |
| 2021 | VISFF: An Approach for Video Summarization Based on Feature Fusion
Weidong Tian 0001, Xiao-Yu Cheng, Zhong-Qiu Zhao |
ICIC (2) | 4 |
| 2021 | Residual Attention Block Search for Lightweight Image Super-ResolutionabstractRecently, lightweight neural networks with different manual designs have presented a promising performance in single image super-resolution (SR). However, these designs rely on too much expert experience. To address this issue, we focus on searching a lightweight block for efficient and accurate image SR. Due to the frequent use of various residual blocks and attention mechanisms in SR methods, we propose the residual attention search block (RASB) which combines an operation search block (OSB) with an attention search block (ASB). The former is used to explore the suitable operation at the proper position, and the latter is applied to discover the optimal connection of various attention mechanisms. Moreover, we build the modified residual attention network (MRAN) with stacked found blocks and a refinement module. Extensive experiments demonstrate that our MRAN achieves a better trade-off against the state-of-the-art methods in terms of accuracy and model complexity. Wenrui Liao, Zhong-Qiu Zhao, Hao Shen 0006, Weidong Tian 0001 |
ICME | 2 |
| 2021 | An improved steganography without embedding based on attention GAN
Cong Yu 0014, Donghui Hu, Shuli Zheng, Wenjie Jiang 0001, Meng Li 0006, Zhong-Qiu Zhao |
Peer-to-Peer Netw. Appl. | 6 |
| 2020 | Mid-Weight Image Super-Resolution with Bypass Connection Attention NetworkabstractDeeper networks have limited improvements for image super-resolution (SR), and are much more difficult to train. The main reason is that these networks consist of many stacked building blocks which can produce many redundant features. Besides, most of SR methods neglect the fact that different features contain various types of information with varying degrees of contributions to image reconstruction, and thus lack sufficient representational capability. Taking these issues into account, we propose a mid-weight bypass connection attention network (BCAN) with more powerful representational capability but fewer parameters. In detail, we design a novel bypass connection attention module (BCAM), which consists of several bypass connection attention blocks (BCABs), enhancing high contribution information and suppressing redundant information. Further, we embed a mixed residual attention unit (MRAU) in each BCAB, which is composed of a channel attention unit and a spatial attention unit. After obtaining all hierarchical features, we propose an adaptive feature fusion module (AFFM), which can effectively combine hierarchical features based on different contributions of each BCAM. Experiments on benchmark datasets with various degradation models show that our BCAN can achieve better performance than existing state-of-the-art methods. Hao Shen 0006, Zhong-Qiu Zhao |
ECAI | 2 |
| 2020 | Real-Time Object Detection Based on Convolutional Block Attention Module
Ming-Yang Ban, Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (3) | 3 |
| 2020 | Image Super-Resolution Network Based on Prior Information Fusion
Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (3) | 3 |
| 2020 | TFPGAN: Tiny Face Detection with Prior Information and GAN
Dian Liu, Zhong-Qiu Zhao, Weidong Tian 0001 |
ICIC (3) | 2 |
| 2020 | A Novel Approach of Steganalysis to Deal with Steganographic Algorithm Mismatch
Donghui Hu, Shuli Zheng, Zhong-Qiu Zhao |
ICIC (1) | 5 |
| 2020 | Regenerating Image Caption with High-Level Semantics
Weidong Tian 0001, Nan-Xun Wang, Yue-Lin Sun, Zhong-Qiu Zhao |
ICIC (3) | 4 |
| 2020 | Plant Leaf Recognition Network Based on Feature Learning and Metric Learning
Di Wu 0030, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao, Zhong-Qiu Zhao |
ICIC (1) | 7 |
| 2020 | Multi-Channel Co-Attention Network for Visual Question AnsweringabstractVisual Question Answering (VQA) is to reason out correct answers based on input questions and images. Significant progresses have been made by learning rich embedding features from images and questions by bilinear models. Attention mechanisms are widely used to focus on specific visual and textual information in VQA reasoning process. However, most state-of-the-art methods concentrate on fusing the global multi-modal features, while neglect local features. Besides, the dimension is reduced excessively (from K×2048 to 2048) in general visual attention, which causes a mass of visual information loss. In this paper, we propose a novel multi-channel co-attention network (MC-CAN), which integrates multi-modal features from global level to local level. We design different multi-channel attention mechanisms separately for visual (from K×2048 to M×2048) and textual features at different level of integrations. Additionally, we further improve our proposed approach by combining it with the complementary modules such as the MLB and the Count modules. Experiments on benchmark datasets show that our approach achieves better VQA performance than other state-of-the-art methods. Weidong Tian 0001, Nanxun Wang, Zhong-Qiu Zhao |
IJCNN | 4 |
| 2020 | Cascading Top-Down Attention for Visual Question AnsweringabstractFor solving Visual Question Answering (VQA), we commonly employ images and questions simultaneously to predict answers. Some attention mechanisms should be used to focus on the most valuable information, because there are too much information extracted from images and questions. Top-Down Attention (TDA) is one of the famous attention mechanisms. For standard TDA, only important regions of the image associated with the question are highlighted. In this work, we propose a Cascading Top-Down Attention (CTDA) model. CTDA highlights the most important information collected from images and questions by a cascading attention process. First, the key words of the question, associated with the image, are highlighted by using a Question Top-Down Attention (QTDA). Then, important regions of the image, associated with the question, are highlighted by using of Image Top-Down Attention (ITDA), useless information of the images and questions can be ignored effectively. We evaluate our model on two popular VQA data sets. CTDA obtains better results than standard TDA and the other state of the art models. Weidong Tian 0001, Rencai Zhou, Zhong-Qiu Zhao |
IJCNN | 3 |
| 2020 | Extricating from GroundTruth: An Unpaired Learning Based Evaluation Metric for Image CaptioningabstractRecently, instead of pursuing high performance on classical evaluation metrics, the research focus of image captioning has shifted to generating sentences which are more vivid and stylized than human-written ones. However, there are still no applicable metrics which can judge how close the generated captions are to the human-written ones. In this paper, we propose a novel learning-based evaluation metric, namely Unpaired Image Captioning Evaluation (UICE), which can be trained to distinguish between human-written and generated captions. Unlike existing metrics, our UICE consists of two parts: the semantic alignment module measuring the semantic distance between extracted image features and caption meanings, and the syntactic discriminating module syntactically judging how human-like the candidate caption is. The semantic alignment module is implemented by mapping the image features and the word embedding into a unified tensor space. And the syntactic discriminating module is designed to be learning-based, and thereby can be trained to be stylized by users' own, fed with additional personalized corpus during the training process. Extensive experiments indicate that our metric can correctly judge the grammatical correctness of generated captions and the semantic consistency between captions and corresponding images. Zhong-Qiu Zhao, Yue-Lin Sun, Nan-Xun Wang, Weidong Tian 0001 |
IJCNN | 1 |
| 2019 | Person Re-identification Based on Feature Fusion
Qiang-Qiang Ren, Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (3) | 3 |
| 2019 | A Novel Concise Representation of Frequent Subtrees Based on Density
Weidong Tian 0001, Chuang Guo, Hongjuan Zhou, Zhong-Qiu Zhao |
ICIC (3) | 5 |
| 2019 | Improving Object Detection by Deep Networks with Class-Related Features
Shou-tao Xu, Zhong-Qiu Zhao |
ICIC (1) | 2 |
| 2019 | Coarse-to-Fine Supervised Descent Method for Face Alignment
Xijing Zhu, Zhong-Qiu Zhao, Weidong Tian 0001 |
ICIC (1) | 2 |
| 2019 | Deep learning-based methods for person re-identification: A comprehensive review
Di Wu 0030, Si-Jia Zheng, Xiao-Ping Zhang 0002, Chang-an Yuan 0001, Yang Zhao 0002, Yong-Jun Lin, Zhong-Qiu Zhao, Yong-Li Jiang, De-Shuang Huang |
Neurocomputing | 8 |
| 2019 | A review of image set classification
Zhong-Qiu Zhao, Shou-tao Xu, Dian Liu, Weidong Tian 0001, Zhi-Da Jiang |
Neurocomputing | 1 |
| 2019 | Object Detection With Deep Learning: A ReviewabstractDue to object detection's close relationship with video analysis and image understanding, it has attracted much research attention in recent years. Traditional object detection methods are built on handcrafted features and shallow trainable architectures. Their performance easily stagnates by constructing complex ensembles that combine multiple low-level image features with high-level context from object detectors and scene classifiers. With the rapid development in deep learning, more powerful tools, which are able to learn semantic, high-level, deeper features, are introduced to address the problems existing in traditional architectures. These models behave differently in network architecture, training strategy, and optimization function. In this paper, we provide a review of deep learning-based object detection frameworks. Our review begins with a brief introduction on the history of deep learning and its representative tool, namely, the convolutional neural network. Then, we focus on typical generic object detection architectures along with some modifications and useful tricks to improve detection performance further. As distinct specific detection tasks exhibit different characteristics, we also briefly survey several specific tasks, including salient object detection, face detection, and pedestrian detection. Experimental analyses are also provided to compare various methods and draw some meaningful conclusions. Finally, several promising directions and tasks are provided to serve as guidelines for future work in both object detection and relevant neural network-based learning systems. Zhong-Qiu Zhao, Shou-tao Xu, Xindong Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Cooperative Adversarial Network for Accurate Super Resolution
Zhong-Qiu Zhao, Weidong Tian 0001, Ning Ling |
ACCV (2) | 1 |
| 2018 | Multi-template Supervised Descent Method for Face Alignment
Zhong-Qiu Zhao, Qinmu Peng |
ICIC (3) | 2 |
| 2018 | Improved Adaptive Incremental Error-Minimization-Based Extreme Learning Machine with Localized Generalization Error Model
Wen-wen Han, Zhong-Qiu Zhao, Weidong Tian 0001 |
ICIC (3) | 3 |
| 2018 | GridWall: A Novel Condensed Representation of Frequent Itemsets
Weidong Tian 0001, Jianqiang Mei, Hongjuan Zhou, Zhong-Qiu Zhao |
ICIC (1) | 4 |
| 2018 | A set-level joint sparse representation for image set classification
Zhong-Qiu Zhao, Jun Gao 0006, Xindong Wu 0001 |
Inf. Sci. | 2 |
| 2017 | Image Firewall for Filtering Privacy or Sensitive Image Content Based on Joint Sparse Representation
Ning Ling, Donghui Hu, Xiaoxia Hu, Zhong-Qiu Zhao |
ICIC (3) | 6 |
| 2017 | Pedestrian Detection Based on Fast R-CNN and Batch Normalization
Zhong-Qiu Zhao, Haiman Bian, Donghui Hu, Wenjuan Cheng, Hervé Glotin |
ICIC (1) | 1 |
| 2017 | A Multi-modal SPM Model for Image Classification
Zhong-Qiu Zhao, Jun Gao 0006 |
ICIC (3) | 2 |
| 2017 | Image set classification based on cooperative sparse representation
Zhong-Qiu Zhao, Jun Gao 0006, Xindong Wu 0001 |
Pattern Recognit. | 2 |
| 2016 | A Real-Time Head Pose Estimation Using Adaptive POSIT Based on Modified Supervised Descent Method
Zhong-Qiu Zhao, Kewen Cheng, Qinmu Peng, Xindong Wu 0001 |
ICIC (1) | 1 |
| 2016 | Single Image Super Resolution with Neighbor Embedding and In-place Patch Matching
Zhong-Qiu Zhao, Zhen-Wei Hao, Run Su, Xindong Wu 0001 |
ICIC (2) | 1 |
| 2016 | Corrupted and occluded face recognition via cooperative sparse representation
Zhong-Qiu Zhao, Yiu-Ming Cheung, Haibo Hu 0001, Xindong Wu 0001 |
Pattern Recognit. | 1 |
| 2016 | Inverse-Free Extreme Learning Machine With Optimal Information UpdatingabstractThe extreme learning machine (ELM) has drawn insensitive research attentions due to its effectiveness in solving many machine learning problems. However, the matrix inversion operation involved in the algorithm is computational prohibitive and limits the wide applications of ELM in many scenarios. To overcome this problem, in this paper, we propose an inverse-free ELM to incrementally increase the number of hidden nodes, and update the connection weights progressively and optimally. Theoretical analysis proves the monotonic decrease of the training error with the proposed updating procedure and also proves the optimality in every updating step. Extensive numerical experiments show the effectiveness and accuracy of the proposed algorithm. Shuai Li 0002, Zhu-Hong You, Xin Luo 0001, Zhong-Qiu Zhao |
IEEE Trans. Cybern. | 5 |
| 2015 | Plant identification using triangular representation based on salient points and margin pointsabstractLeaf classification is an important component of living plant identification. A leaf contains important information for plant species identification in spite of its complexity. This paper introduces a method of recognizing leaf images based on triangular representations. A leaf is represented by local descriptors associated with margin sample points and salient sample points. We introduce three new triangular representations - salient triangle area representation (STAR), salient triangle side lengths representation (STSL), and salient triangle area, side lengths and two angles representation (STASLA), and then we combine two local descriptors - one provides a triangular representation of the leaf margin while the other represents the spatial correlation between salient points of the leaf and leaf margin. Experiments on the Image-CLEF 2011 leaf datasets show the effectiveness and the efficiency of the proposed method. Zhong-Qiu Zhao, Xindong Wu 0001 |
ICIP | 1 |
| 2015 | Expanding dictionary for robust face recognition: pixel is not necessary while sparsity isabstractSince sparse representation (SR) was first introduced into robust face recognition, the argument has lasted for several years about whether sparsity can improve robust face recognition or not. Some work argued that the robust sparse representation (RSR) model has a similar recognition rate as non‐sparse solution, while it needs a much higher computational cost due to the larger feature dimensionality in the pixel space. In this study, the authors reveal that the standard RSR model, which expands the dictionary with the identity matrix to reconstruct corruption or occlusion in face images, is essentially a non‐sparse solution with a relatively large residual. The reason why the RSR model underperforms may be its inappropriately expanded bases rather than the sparsity itself. Thereby, this study proposes to design a dictionary with an expanded noise bases set which can precisely reconstructs any corruption or occlusion in face images in a subspace . Experimental results show that the algorithm can greatly improve recognition rates for robust face recognition. In addition, the algorithm can be simply performed in a subspace with a small feature dimensionality, thus efficient enough for real systems. This study makes us come to the conclusion that solving the approximation problem in raw pixel space is not necessary for robust face recognition, while solving in a subspace with a much smaller feature dimensionality is enough when the dictionary is well expanded. Finally, this study also confirms that the sparsity plays an important role in SR based classification. Zhong-Qiu Zhao, Yiu-Ming Cheung, Haibo Hu 0001, Xindong Wu 0001 |
IET Comput. Vis. | 1 |
| 2015 | ApLeaf: An efficient android-based plant leaf identification system
Zhong-Qiu Zhao, Lin-Hai Ma, Yiu-Ming Cheung, Xindong Wu 0001, Yuan Yan Tang, C. L. Philip Chen |
Neurocomputing | 1 |
| 2015 | Online Feature Selection with Group Structure AnalysisabstractOnline selection of dynamic features has attracted intensive interest in recent years. However, existing online feature selection methods evaluate features individually and ignore the underlying structure of a feature stream. For instance, in image analysis, features are generated in groups which represent color, texture, and other visual information. Simply breaking the group structure in feature selection may degrade performance. Motivated by this observation, we formulate the problem as an online group feature selection. The problem assumes that features are generated individually but there are group structures in the feature stream. To the best of our knowledge, this is the first time that the correlation among streaming features has been considered in the online feature selection process. To solve this problem, we develop a novel online group feature selection method named OGFS. Our proposed approach consists of two stages: online intra-group selection and online inter-group selection. In the intra-group selection, we design a criterion based on spectral analysis to select discriminative features in each group. In the inter-group selection, we utilize a linear regression model to select an optimal subset. This two-stage procedure continues until there are no more features arriving or some predefined stopping conditions are met. Finally, we apply our method to multiple tasks including image classification and face verification. Extensive empirical studies performed on real-world and benchmark data sets demonstrate that our method outperforms other state-of-the-art online feature selection methods. Jing Wang 0021, Meng Wang 0001, Pei-Pei Li 0001, Luoqi Liu, Zhong-Qiu Zhao, Xuegang Hu, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2014 | Plant Leaf Identification via a Growing Convolution Neural Network with Progressive Sample Learning
Zhong-Qiu Zhao, Bao-Jian Xie, Yiu-Ming Cheung, Xindong Wu 0001 |
ACCV (2) | 1 |
| 2014 | Optimizing widths with PSO for center selection of Gaussian radial basis function networks
Zhong-Qiu Zhao, Xindong Wu 0001, Canyi Lu, Hervé Glotin, Jun Gao 0006 |
Sci. China Inf. Sci. | 1 |
| 2013 | ApLeafis: An Android-Based Plant Leaf Identification System
Lin-Hai Ma, Zhong-Qiu Zhao, Jing Wang 0021 |
ICIC (1) | 2 |
| 2013 | Online Group Feature Selection
Jing Wang 0021, Zhong-Qiu Zhao, Xuegang Hu, Yiu-Ming Cheung, Meng Wang 0001, Xindong Wu 0001 |
IJCAI | 2 |
| 2013 | Locally linear representation Fisher criterionabstractIn this paper, a novel supervised dimensionality reduction method based on LLE is put forward, which is titled locally linear representation Fisher criterion (LLRFC). In the proposed LLRFC, the class information of the original data has been fully considered, according to which an inter-class graph and an intra-class graph can be well modeled respectively. Meanwhile, the neighborhoods in the inter-class graph consist of samples with various labels and the neighborhoods in the intra-graph are just composed of points sampled from the same class. Then the least locally linear representation technique is introduced to optimize the reconstruction weights in both graphs. At last, the Fisher criterion with maximum inter-class scatter and minimum intra-class scatter is reasoned. Experiments on some benchmark face data sets have been conducted and the results validate the proposed method's performance. Bo Li 0002, Jin Liu 0016, Zhong-Qiu Zhao, Wensheng Zhang 0002 |
IJCNN | 3 |
| 2012 | Robust and Efficient Subspace Segmentation via Least Squares Regression
Canyi Lu, Hai Min, Zhong-Qiu Zhao, Lin Zhu 0008, De-Shuang Huang, Shuicheng Yan |
ECCV (7) | 3 |
| 2012 | Cooperative Sparse Representation in Two Opposite Directions for Semi-Supervised Image AnnotationabstractRecent studies have shown that sparse representation (SR) can deal well with many computer vision problems, and its kernel version has powerful classification capability. In this paper, we address the application of a cooperative SR in semi-supervised image annotation which can increase the amount of labeled images for further use in training image classifiers. Given a set of labeled (training) images and a set of unlabeled (test) images, the usual SR method, which we call forward SR, is used to represent each unlabeled image with several labeled ones, and then to annotate the unlabeled image according to the annotations of these labeled ones. However, to the best of our knowledge, the SR method in an opposite direction, that we call backward SR to represent each labeled image with several unlabeled images and then to annotate any unlabeled image according to the annotations of the labeled images which the unlabeled image is selected by the backward SR to represent, has not been addressed so far. In this paper, we explore how much the backward SR can contribute to image annotation, and be complementary to the forward SR. The co-training, which has been proved to be a semi-supervised method improving each other only if two classifiers are relatively independent, is then adopted to testify this complementary nature between two SRs in opposite directions. Finally, the co-training of two SRs in kernel space builds a cooperative kernel sparse representation (Co-KSR) method for image annotation. Experimental results and analyses show that two KSRs in opposite directions are complementary, and Co-KSR improves considerably over either of them with an image annotation performance better than other state-of-the-art semi-supervised classifiers such as transductive support vector machine, local and global consistency, and Gaussian fields and harmonic functions. Comparative experiments with a nonsparse solution are also performed to show that the sparsity plays an important role in the cooperation of image representations in two opposite directions. This paper extends the application of SR in image annotation and retrieval. Zhong-Qiu Zhao, Hervé Glotin, Zhao Xie, Jun Gao 0006, Xindong Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Real-Time Face Detection Using Integral Histogram of Multi-scale Local Binary Patterns
Sébastien Paris, Hervé Glotin, Zhong-Qiu Zhao |
ICIC (1) | 3 |
| 2010 | A Modified Semi-Supervised Learning Algorithm on Laplacian Eigenmaps
Zhong-Qiu Zhao, Jun-Zhao Li, Jun Gao 0006, Xindong Wu 0001 |
Neural Process. Lett. | 1 |
| 2009 | Efficient image concept indexing by harmonic & arithmetic profiles entropyabstractWe propose new efficient visual features called Profile Entropy Features (PEF), giving information on the structure of the image content, and defined as the entropy of the distribution of a projection of the pixels. We analyse two simple projection operators (arithmetic or harmonic mean), and two orientations (horizontal and vertical). PEF are fast to compute (10 images per sec. on a PentiumIV) and of small dimension. Moreover, we show on High Level Feature task in TrecVid2008 that PEF performs in average better than the features of the state of the art (usual color features, edge direction, Gabor, and Local Binary Pattern). Moreover, we show on another international image retrieval campaign, the Visual Concept Detection of ImageCLEF2008, that the arithmetic and harmonic projections give complementary informations, yielding to the third best rank system in the official run of this campaign. Other properties of the PEF are discussed. Hervé Glotin, Zhong-Qiu Zhao, Stéphane Ayache |
ICIP | 2 |
| 2009 | A novel modular neural network for imbalanced classification problems
Zhong-Qiu Zhao |
Pattern Recognit. Lett. | 1 |
| 2007 | An evolutionary modular neural network for unbalanced pattern classificationsabstractIn this paper, an evolutionary modular neural network is proposed to solve multi-class problems with unbalanced training sets. The proposed model can transform an unbalanced classification problem into a set of symmetrical two-class problems, each of which can be solved by a single simple neural network. The experimental results show that the proposed method reduces time consumption for training and improves the classification performance. Zhong-Qiu Zhao, De-Shuang Huang |
IEEE Congress on Evolutionary Computation | 1 |
| 2007 | Palmprint recognition with 2DPCA+PCA based on modular neural networks
Zhong-Qiu Zhao, De-Shuang Huang, Wei Jia 0001 |
Neurocomputing | 1 |
| 2004 | A novel clustering-neural tree for pattern classificationabstractWhen performing classification of large set of samples, neural tree classifiers (NTs) are preferred. However, the classical NTs have poor generalization properties. So, in this paper we propose a new classification method referred to as clustering-neural tree classifier, combining clustering technique with neural networks. It can be well applied to classifications of large set of samples, while having good generalization properties. The experimental results on the two spirals problem and the iris problem show that our proposed NN-tree classifier is effective and efficient. Zhong-Qiu Zhao, De-Shuang Huang |
IJCNN | 1 |
| 2004 | Support Vector Machine Committee for Classification
Bing-Yu Sun, De-Shuang Huang, Zhong-Qiu Zhao |
ISNN (1) | 4 |
| 2004 | Human face recognition based on multi-features using neural networks committee
Zhong-Qiu Zhao, De-Shuang Huang, Bing-Yu Sun |
Pattern Recognit. Lett. | 1 |