Qin Zou 0001

dblp:45/8496 · DBLP profile ↗
← Back
72ranked-venue papers
12as first author
40since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 4 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 5 since 2021Security and privacy · 6 · 1 first-authorSystems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A cyclic diffusion framework for structure-authentic and annotation-disentangled anomaly generation
Linchun Wu, Qin Zou 0001, Xianbiao Qi, Zhongyuan Wang 0001, Qingquan Li 0001
Neurocomputing2
2026 Physics-inspired pseudo anomaly generation and prototype feature guidance for 3D anomaly detection
Jian Ning, Qin Zou 0001, Linchun Wu, Yuanhao Yue, Kunmo Li, Shoubin Chen, Zhongyuan Wang 0001
Pattern Recognit.2
2026 DiffCrack: A semantic-structural controllable framework for crack image generation in complex scenes
abstract
Robust pavement crack detection in complex scenes remains a significant challenge. This stems not merely from the scarcity of annotated data, but more critically, from the severe lack of pattern diversity within existing datasets. Key variations in morphology, scale, background texture, and imaging conditions are often underrepresented, which fundamentally impedes the generalization capability of recognition models. While generative approaches (e.g., GANs and diffusion models) offer a potential path to augment this diversity synthetically, they commonly suffer from poor background realism, entangled structural-appearance representations, and a lack of precise control over generated defects. This paper presents the Crack Diffusion Generator (DiffCrack), a diffusion-based framework designed for semantic-structural controllability in crack image synthesis. DiffCrack decouples crack geometry and visual appearance through two conditioning inputs: a binary mask to anchor spatial layout, and a Hierarchical Prompt Attention (HPA) module to independently modulate attributes such as width, depth, color, and texture. This design enables the controllable and targeted generation of diverse, photorealistic crack patterns that are often missing from real-world datasets. Extensive experiments on real datasets demonstrate that training with DiffCrack-generated images enhances the F1-score of segmentation models by up to 23% on average under complex scene conditions. This result validates that DiffCrack is a scalable, pattern-aware data augmentation tool. By addressing the critical bottleneck of data diversity, our framework offers a principled pathway to improving the robustness and generalization of infrastructure inspection models.
Lingxi Xie, Qin Zou 0001, Qi Tian 0001, Qingquan Li 0001
Pattern Recognit.3
2026 IDRetracor: Towards Visual Forensics against Malicious Face Swapping
abstract
The deepfake-based face swapping technique poses significant risks to personal identity security. Although many detection methods have been proposed to counter malicious face swapping, they typically provide only binary labels (Fake/Real), lacking reliable and interpretable evidence. To address this limitation, we introduce a novel task called face retracing, which aims to visually trace back the original target face from a given fake one through inverse mapping. This task is based on the observation that current face swapping methods are neither flawless nor entirely random, leaving recoverable traces of the original identity. To this end, we propose IDRetracor, a model designed to recover arbitrary original target identities from fake faces generated by various face swapping techniques. Specifically, we first employ a mapping resolver to estimate the possible solution space of the original face for inverse mapping. Then, we introduce Mapping-Aware Convolutions (MACs), which consist of multiple dynamically combined kernels guided by the mapping resolver to adaptively handle diverse face swapping patterns. Extensive experiments demonstrate that IDRetracor achieves strong performance in retracing original faces, validated by both quantitative metrics and qualitative assessments.
Jikang Cheng, Jiaxin Ai, Zhen Han 0002, Chao Liang 0001, Qin Zou 0001, Zhongyuan Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Stacking Brick by Brick: Aligned Feature Isolation for Incremental Face Forgery Detection
abstract
The rapid advancement of face forgery techniques has introduced a growing variety of forgeries. Incremental Face Forgery Detection (IFFD), involving gradually adding new forgery data to fine-tune the previously trained model, has been introduced as a promising strategy to deal with evolving forgery methods. However, a naively trained IFFD model is prone to catastrophic forgetting when new forgeries are integrated, as treating all forgeries as a single “Fake” class in the Real/Fake classification can cause different forgery types overriding one another, thereby resulting in the forgetting of unique characteristics from earlier tasks and limiting the model’s effectiveness in learning forgery specificity and generality. In this paper, we propose to stack the latent feature distributions of previous and new tasks brick by brick, i.e., achieving aligned feature isolation. In this manner, we aim to preserve learned forgery information and accumulate new knowledge by minimizing distribution overriding, thereby mitigating catastrophic forgetting. To achieve this, we first introduce Sparse Uniform Replay (SUR) to obtain the representative subsets that could be treated as the uniformly sparse versions of the previous global distributions. We then propose a Latent-space Incremental Detector (LID) that leverages SUR data to isolate and align distributions. For evaluation, we construct a more advanced and comprehensive benchmark tailored for IFFD. The leading experimental results validate the superiority of our method. Code is available at https://github.com/beautyremain/SUR-LID .
Jikang Cheng, Zhiyuan Yan 0002, Ying Zhang 0021, Jiaxin Ai, Qin Zou 0001, Chen Li 0031, Zhongyuan Wang 0001
CVPR6
2025 Dual-Frequency Spatio-Temporal Phase Unwrapping
abstract
Phase unwrapping poses a critical challenge in 3D reconstruction, particularly due to the presence of noise and discontinuities that compromise the accuracy of phase extraction. Existing convolutional neural network (CNN)-based methods have struggled to effectively integrate traditional approaches while fully utilizing both multi-frequency, i.e., temporal, and spatial information of the phase. In this paper, we propose a novel phase unwrapping method, STPhaseNet, which incorporates both temporal and spatial phase information into the CNN framework. Specifically, we introduce a temporal feature fusion module and a local attention mechanism to extract and integrate features from different frequency phases. To further leverage spatial phase information, we develop a spatial information extraction module that enlarges the local receptive field of the convolution and assigns weights based on the phase information of horizontal and vertical coordinates. Additionally, we design a globally optimized gradient residual loss function to exploit spatial constraints more effectively. To address the lack of real-world training data, we apply a Random Matrix Enlargement (RME) method to generate high-quality dual-frequency wrapped phase data along with corresponding absolute phases for training purposes. Extensive experiments demonstrate that STPhaseNet outperforms existing methods, achieving superior performance in phase unwrapping tasks.
Shuo Du, Qin Zou 0001, Chi Chen 0002, Bisheng Yang
ICASSP2
2025 CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs
abstract
Object goal navigation (ObjectNav) is a fundamental task in embodied AI, requiring an agent to locate a target object in previously unseen environments. This task is particularly challenging because it requires both perceptual and cognitive processes, including object recognition and decision-making. While substantial advancements in perception have been driven by the rapid development of visual foundation models, progress on the cognitive aspect remains constrained, primarily limited to either implicit learning through simulator rollouts or explicit reliance on predefined heuristic rules. Inspired by neuroscientific findings demonstrating that humans maintain and dynamically update fine-grained cognitive states during object search tasks in novel environments, we propose CogNav, a framework designed to mimic this cognitive process using large language models. Specifically, we model the cognitive process using a finite state machine comprising fine-grained cognitive states, ranging from exploration to identification. Transitions between states are determined by a large language model based on a dynamically constructed heterogeneous cognitive map, which contains spatial and semantic information about the scene being explored. Extensive evaluations on the HM3D, MP3D, and RoboTHOR benchmarks demonstrate that our cognitive process modeling significantly improves the success rate of ObjectNav at least by relative 14% over the state-of-the-arts.
Jiazhao Zhang, Zhinan Yu, Shuzhen Liu, Zheng Qin 0002, Qin Zou 0001, Bo Du 0001, Kai Xu 0004
ICCV6
2025 Taming Transformer Without Using Learning Rate Warmup
abstract
Scaling Transformer to a large scale without using some technical tricks such as learning rate warump and an obviously lower learning rate, is an extremely challenging task, and is increasingly gaining more attention. In this paper, we provide a theoretical analysis for the process of training Transformer and reveal a key problem behind model crash phenomenon in the training process, termed *spectral energy concentration* of ${W_q}^{\top} W_k$, which is the reason for a malignant entropy collapse, where ${W_q}$ and $W_k$ are the projection matrices for the query and the key in Transformer, respectively. To remedy this problem, motivated by *Weyl's Inequality*, we present a novel optimization strategy, \ie, making the weight updating in successive steps steady---if the ratio $\frac{\sigma_{1}(\nabla W_t)}{\sigma_{1}(W_{t-1})}$ is larger than a threshold, we will automatically bound the learning rate to a weighted multiple of $\frac{\sigma_{1}(W_{t-1})}{\sigma_{1}(\nabla W_t)}$, where $\nabla W_t$ is the updating quantity in step $t$. Such an optimization strategy can prevent spectral energy concentration to only a few directions, and thus can avoid malignant entropy collapse which will trigger the model crash. We conduct extensive experiments using ViT, Swin-Transformer and GPT, showing that our optimization strategy can effectively and stably train these (Transformer) models without using learning rate warmup.
Xianbiao Qi, Yelin He, Jiaquan Ye, Chun-Guang Li, Bojia Zi, Xili Dai, Qin Zou 0001, Rong Xiao 0003
ICLR7
2025 V2X-DGW: Domain Generalization for Multi-Agent Perception Under Adverse Weather Conditions
abstract
Current LiDAR-based Vehicle-to-Everything (V2X) multi-agent perception systems have shown the significant success on 3D object detection. While these models perform well in the trained clean weather, they struggle in unseen adverse weather conditions with the domain gap. In this paper, we propose a Domain Generalization based approach, named V2X-DGW, for LiDAR-based 3D object detection on multi-agent perception system under adverse weather conditions. Our research aims to not only maintain favorable multi-agent performance in the clean weather but also promote the performance in the unseen adverse weather conditions by learning only on the clean weather data. To realize the Domain Generalization, we first introduce the Adaptive Weather Augmentation (AWA) to mimic the unseen adverse weather conditions, and then propose two alignments for generalizable representation learning: Trust-region Weatherinvariant Alignment (TWA) and Agent-aware Contrastive Alignment (ACA). To evaluate this research, we add Fog, Rain, Snow conditions on two publicized multi-agent datasets based on physics-based models, resulting in two new datasets: OPV2V-w and V2XSet-w. Extensive experiments demonstrate that our V2X-DGW achieved significant improvements in the unseen adverse weathers. The code is available at https://github.com/Baolu1998/V2X-DGW.
Xinyu Liu 0009, Runsheng Xu, Zhengzhong Tu, Jiacheng Guo, Qin Zou 0001, Xiaopeng Li 0020, Hongkai Yu
ICRA7
2025 Adjacent-view Transformers for Supervised Surround-view Depth Estimation
abstract
Depth estimation has been widely studied and serves as the fundamental step of 3D perception for robotics and autonomous driving. Though significant progress has been made in monocular depth estimation in the past decades, these attempts are mainly conducted on the KITTI benchmark with only front-view cameras, which ignores the correlations across surround-view cameras. In this paper, we propose an Adjacent-View Transformer for Supervised Surround-view Depth estimation (AVT-SSDepth), to jointly predict the depth maps across multiple surrounding cameras. Specifically, we employ a global-to-local feature extraction module that combines CNN with transformer layers for enriched representations. Further, the adjacent-view attention mechanism is proposed to enable the intra-view and inter-view feature propagation. The former is achieved by the self-attention module within each view, while the latter is realized by the adjacent attention module, which computes the attention across multi-cameras to exchange the multi-scale representations across surround-view feature maps. In addition, AVT-SSDepth has strong cross-dataset generalization. Extensive experiments show that our method achieves superior performance over existing state-of-the-art methods on both DDAD and nuScenes datasets. Code is available at https://github.com/XiandaGuo/SSDepth.
Xianda Guo, Wenjie Yuan 0004, Chenming Zhang, Qin Zou 0001, Long Chen 0005
IROS7
2025 A Robust 3D CNN with Pyramidal Attention for Spatiotemporal Gait Recognition
abstract
Gait recognition has become an increasingly important biometric technique for identifying individuals from a distance without requiring their active cooperation. Since gait involves a sequence of motion patterns, effectively capturing temporal dynamics is essential for accurate recognition. Traditional methods that extract temporal features independently and fuse them at a later stage often fail to model the continuity and interdependence of motion across frames. To overcome this limitation, we propose a novel three-dimensional convolutional architecture named Robust Spatiotemporal 3D Convolutional Neural Network (RST3D), which jointly captures spatial and temporal correlations throughout gait sequences. The proposed architecture incorporates a comprehensive 3D convolutional block that operates along the temporal, height, and width dimensions, enabling the network to learn more expressive and coherent spatiotemporal representations. In addition, we introduce a Temporal Pyramidal Attention (TPA) block to enhance the network’s ability to model temporal dependencies by capturing discriminative motion patterns across multiple temporal scales. We evaluate our method on four large-scale gait recognition datasets: CASIA-B, OUMVLP, GREW, and Gait3D. Experimental results show that our approach consistently achieves superior performance compared to existing 3D CNN-based methods, particularly under challenging conditions such as view variation, clothing changes, and occlusion.
Jianyu Chen 0008, Qian Zhou 0001, Qin Zou 0001, Chao Liang 0001, Zengmin Xu, Gang Wu 0010, Zhongyuan Wang 0001
MMAsia3
2025 Multi-Modal Gait Recognition via Collaborative Feature Learning from Silhouettes and Skeletons
Jianyu Chen 0008, Zhongyuan Wang 0001, Qian Zhou 0001, Qin Zou 0001, Chao Liang 0001, Gang Wu 0010
PRCV (15)4
2025 Learning 3D Proposals in Spatio-Temporal Transformer for Multi-camera Driving Scene Object Detection
Qin Zou 0001
PRICAI (5)2
2025 Occluded person re-identification with feature complement and dual attention
Wei Xiong 0004, Zixin Tian, Lirong Li, Qin Zou 0001, Song Wang 0002
Expert Syst. Appl.4
2025 BEVFix: Deep feature enhancement for robust 3D object detection
Jian Zhou 0011, Chi Chen 0002, Hongkai Yu, Bo Du 0001, Qin Zou 0001
Neural Networks6
2025 ATCM: Aerial-Terrestrial LiDAR-Based Collaborative Simultaneous Localization and Mapping
abstract
Multi-robot collaborative simultaneous localization and mapping (C-SLAM) offers precise scene reconstruction over single-robot SLAM and enables the data fusion from heterogeneous robots. However, heterogeneous C-SLAM faces challenges in both accurate inter-robot loop closure detection and globally consistent data fusion due to inherent viewpoint disparities and heterogeneous data characteristics. This paper introduces ATCM, an Aerial-Terrestrial LiDAR-based C-SLAM method designed for heterogeneous robots without priori initial relative position. ATCM comprises three modules: single-robot front-end employing diverse SLAM methods, multi-robot loop closure detection, and global pose graph optimization. A novel LiDAR-based cross-view global loop descriptor is proposed for scan-to-scan heterogeneous inter-robot loop closure detection. By uniformly mapping cross-view information into the height domain and integrating dynamic height, the loop descriptor automatically achieves viewpoint correction. Additionally, we introduce a bidirectional loop detection algorithm that validates inter-robot loop closures through both forward and reverse detections. Finally, the two-stage global pose graph optimization integrates multi-source measurements, ensuring globally consistent mapping and localization with cross-view data. We have validated the effectiveness of ATCM on campus scenario datasets and the KITTI dataset, achieving a remarkable 21.95% improvement in trajectory accuracy and a 17.00% enhancement in map precision compared to high-precision point cloud maps, surpassing state-of-the-art LiDAR-based odometry methods. In the ablation experiments, the proposed loop descriptor achieved 97% accuracy in recognizing heterogeneous inter-robot loop closures. Moreover, compared to the traditional unidirectional method, the bidirectional loop detection method demonstrates up to a 31.2% improvement in loop closure accuracy.
Chi Chen 0002, Bisheng Yang, Weitong Wu 0002, Shangzhe Sun, Zhiye Wang, Liuchun Li, Qin Zou 0001
IEEE Trans. Geosci. Remote. Sens.8
2025 ED4: Explicit Data-Level Debiasing for Deepfake Detection
abstract
Learning intrinsic bias from limited data has been considered the main reason for the failure of deepfake detection with generalizability. Apart from the discovered content and specific-forgery bias, we reveal a novel spatial bias, where detectors inertly anticipate observing structural forgery clues appearing at the image center, also can lead to the poor generalization of existing methods. We present ED4, a simple and effective strategy, to address aforementioned biases explicitly at the data level in a unified framework rather than implicit disentanglement via network design. In particular, we develop ClockMix to produce facial structure preserved mixtures with arbitrary samples, which allows the detector to learn from an exponentially extended data distribution with much more diverse identities, backgrounds, local manipulation traces, and the co-occurrence of multiple forgery artifacts. We further propose the Adversarial Spatial Consistency Module (AdvSCM) to prevent extracting features with spatial bias, which adversarially generates spatial-inconsistent images and constrains their extracted feature to be consistent. As a model-agnostic debiasing strategy, ED4 is plug-and-play: it can be integrated with various deepfake detectors to obtain significant benefits. We conduct extensive experiments to demonstrate its effectiveness and superiority over existing deepfake detection approaches. Code is available at https://github.com/beautyremain/ED4.
Jikang Cheng, Ying Zhang 0021, Qin Zou 0001, Zhiyuan Yan 0002, Chao Liang 0001, Zhongyuan Wang 0001, Chen Li 0031
IEEE Trans. Image Process.3
2024 XFusion: Cross-Attention Transformer for Multi-focus Image Fusion
Shouxi Zhao, Qin Zou 0001, Chi Chen 0002, Zhongyuan Wang 0001
ICONIP (7)3
2024 S2R-ViT for Multi-Agent Cooperative Perception: Bridging the Gap from Simulation to Reality
abstract
Due to the lack of enough real multi-agent data and time-consuming of labeling, existing multi-agent cooperative perception algorithms usually select the simulated sensor data for training and validating. However, the perception performance is degraded when these simulation-trained models are deployed to the real world, due to the significant domain gap between the simulated and real data. In this paper, we propose the first Simulation-to-Reality transfer learning framework for multi-agent cooperative perception using a novel Vision Transformer, named as S2R-ViT, which considers both the Deployment Gap and Feature Gap between simulated and real data. We investigate the effects of these two types of domain gaps and propose a novel uncertainty-aware vision transformer to effectively relief the Deployment Gap and an agent-based feature adaptation module with inter-agent and ego-agent discriminators to reduce the Feature Gap. Our intensive experiments on the public multi-agent cooperative perception datasets OPV2V and V2V4Real demonstrate that the proposed S2R-ViT can effectively bridge the gap from simulation to reality and outperform other methods significantly for point cloud-based 3D object detection.
Runsheng Xu, Xinyu Liu 0009, Qin Zou 0001, Jiaqi Ma 0003, Hongkai Yu
ICRA5
2024 SF-Gait: Two-Stage Temporal Compression Network for Learning Gait Micro-Motions and Cycle Patterns
Yuanhao Yue, Yunhe Wang 0011, Laixiang Shi, Zhongyuan Wang 0001, Qin Zou 0001
PRCV (15)5
2024 Re-decoupling the classification branch in object detectors for few-class scenes
Jie Hua 0005, Zhongyuan Wang 0001, Qin Zou 0001, Jinsheng Xiao, Xin Tian 0006
Pattern Recognit.3
2024 Deep motion estimation through adversarial learning for gait recognition
Yuanhao Yue, Laixiang Shi, Long Chen 0005, Zhongyuan Wang 0001, Qin Zou 0001
Pattern Recognit. Lett.6
2024 GPR-Former: Detection and Parametric Reconstruction of Hyperbolas in GPR B-Scan Images With Transformers
abstract
Ground Penetrating Radar (GPR) enables the non-invasive detection of various subsurface objects such as pipes, stones, etc. The location and size of the object in the medium could be obtained by fitting the generated hyperbolic signatures within the GPR B-scan and analyzing its parameters. In this paper, GPR-Former is proposed for automatic target detection and hyperbola fitting on GPR B-scan images. We have designed a transformer-based neural network to extract features to directly regress the parameters of hyperbolic signatures in the GPR B-scan data to detect targets beneath the ground automatically. A symmetry-constrained analytical solution for the hyperbolic parameters is proposed to refine the parameters derived from the transformer network, serving the extraction and analysis of buried objects in underground opaque spaces. Experiments are conducted on three datasets for the qualitative and quantitative validation of the GPR-Former, including ground-penetrating radar detection of submarine pipelines and land pipelines. Results show that the proposed method is able to automatically and efficiently extract hyperbolas from GPR B-scan images. True hyperbola-point precision (TP_Pre) and true hyperbola-point recall (TP_Rec) metrics are introduced to evaluate performances in parametric hyperbola extraction and fitting. The results show that the TP_Pre and TP_Rec of the proposed method reach 0.867, 0.402, 0.744 and 0.762, 0.736, 0.723, with an improvement of 6%, 22%, 4% compared with the state-of-the-art methods (C3 algorithm and migration learning-based method proposed by Yang), respectively.
Ang Jin, Chi Chen 0002, Bisheng Yang, Qin Zou 0001, Zhiye Wang, Zhengfei Yan, Shaolong Wu, Jian Zhou 0011
IEEE Trans. Geosci. Remote. Sens.4
2023 Implicit Identity Driven Deepfake Face Swapping Detection
abstract
In this paper, we consider the face swapping detection from the perspective of face identity. Face swapping aims to replace the target face with the source face and generate the fake face that the human cannot distinguish between real and fake. We argue that the fake face contains the explicit identity and implicit identity, which respectively corresponds to the identity of the source face and target face during face swapping. Note that the explicit identities of faces can be extracted by regular face recognizers. Particularly, the implicit identity of real face is consistent with the its explicit identity. Thus the difference between explicit and implicit identity of face facilitates face swapping detection. Following this idea, we propose a novel implicit identity driven framework for face swapping detection. Specifically, we design an explicit identity contrast (EIC) loss and an implicit identity exploration (IIE) loss, which supervises a CNN backbone to embed face images into the implicit identity space. Under the guidance of EIC, real samples are pulled closer to their explicit identities, while fake samples are pushed away from their explicit identities. More-over, IIE is derived from the margin-based classification loss function, which encourages the fake faces with known target identities to enjoy intra-class compactness and inter-class diversity. Extensive experiments and visualizations on several datasets demonstrate the generalization of our method against the state-of-the-art counterparts.
Baojin Huang, Zhongyuan Wang 0001, Jifan Yang, Jiaxin Ai, Qin Zou 0001, Qian Wang 0002, Dengpan Ye
CVPR5
2023 Deepfake Face Provenance for Proactive Forensics
abstract
Malicious deepfake face not only violates the privacy of personal identities, but also confuses the public and causes huge social harm. The current deepfake detection only stays at the level of distinguishing between true and false, but cannot trace the original genuine face corresponding to the fake face, that is, it does not have the ability to trace the source of evidence. The deepfake countermeasure technology for judicial forensics urgently calls for deepfake inversion. This paper pioneers an interesting question about face deepfake, active forensics that "know what it is and how it happened". Given that deepfake faces do not completely discard the features of original faces, especially facial expressions and poses, we argue that original faces can be approximately speculated from their deepfake counterparts. Correspondingly, we design a disentangling reversing network that decouples latent space features of deepfake faces under the supervision of real-fake face pair samples to infer original faces in reverse.
Jiaxin Ai, Zhongyuan Wang 0001, Baojin Huang, Zhen Han 0002, Qin Zou 0001
ICIP5
2023 Domain Adaptive Object Detection for Autonomous Driving under Foggy Weather
abstract
Most object detection methods for autonomous driving usually assume a consistent feature distribution between training and testing data, which is not always the case when weathers differ significantly. The object detection model trained under clear weather might be not effective enough on the foggy weather because of the domain gap. This paper proposes a novel domain adaptive object detection framework for autonomous driving under foggy weather. Our method leverages both image-level and object-level adaptation to diminish the domain discrepancy in image style and object appearance. To further enhance the model’s capabilities under challenging samples, we also come up with a new adversarial gradient reversal layer to perform adversarial mining for the hard examples together with domain adaptation. Moreover, we propose to generate an auxiliary domain by data augmentation to enforce a new domain-level metric regularization. Experimental results on public benchmarks show the effectiveness and accuracy of the proposed method. The code is available at https://github.com/jinlong17/DA-Detect.
Runsheng Xu, Jin Ma 0005, Qin Zou 0001, Jiaqi Ma 0003, Hongkai Yu
WACV4
2023 An end-to-end network for co-saliency detection in one single image
Yuanhao Yue, Qin Zou 0001, Hongkai Yu, Qian Wang 0002, Zhongyuan Wang 0001, Song Wang 0002
Sci. China Inf. Sci.2
2023 Extreme Low-Resolution Action Recognition with Confident Spatial-Temporal Attention Transfer
Yucai Bai, Qin Zou 0001, Xieyuanli Chen, Lingxi Li 0001, Zhengming Ding, Long Chen 0005
Int. J. Comput. Vis.2
2023 GaitAMR: Cross-view gait recognition via aggregated multi-feature representation
Jianyu Chen 0008, Zhongyuan Wang 0001, Caixia Zheng, Kangli Zeng, Qin Zou 0001, Laizhong Cui
Inf. Sci.5
2023 Depth map guided triplet network for deepfake face detection
Buyun Liang 0002, Zhongyuan Wang 0001, Baojin Huang, Qin Zou 0001, Qian Wang 0002
Neural Networks4
2023 Deep learning for image inpainting: A survey
Hanyu Xiang, Qin Zou 0001, Muhammad Ali Nawaz, Fan Zhang 0006, Hongkai Yu
Pattern Recognit.2
2023 Coarse-to-Fine Task-Driven Inpainting for Geoscience Images
abstract
The processing and recognition of geoscience images have wide applications. Most of existing researches focus on understanding the high-quality geoscience images by assuming that all the images are clear. However, in many real-world cases, the geoscience images might contain occlusions during the image acquisition. This problem actually implies the image inpainting problem in computer vision and multimedia. As far as we know, all the existing image inpainting algorithms learn to repair the occluded regions for a better visualization quality, they are excellent for natural images but not good enough for geoscience images, and they never consider the following gescience task when developing inpainting methods. This paper aims to repair the occluded regions for a better geoscience task performance and advanced visualization quality simultaneously, without changing the current deployed deep learning based geoscience models. Because of the complex context of geoscience images, we propose a coarse-to-fine encoder-decoder network with the help of designed coarse-to-fine adversarial context discriminators to reconstruct the occluded image regions. Due to the limited data of geoscience images, we propose a MaskMix based data augmentation method, which augments inpainting masks instead of augmenting original images, to exploit the limited geoscience image data. The experimental results on three public geoscience datasets for remote sensing scene recognition, cross-view geolocation and semantic segmentation tasks respectively show the effectiveness and accuracy of the proposed method. The code is available at:https://github.com/HMS97/Task-driven-Inpainting.
Huiming Sun, Jin Ma 0005, Qing Guo 0005, Qin Zou 0001, Shaoyue Song, Yuewei Lin, Hongkai Yu
IEEE Trans. Circuits Syst. Video Technol.4
2023 Joint Segmentation and Identification Feature Learning for Occlusion Face Recognition
abstract
The existing occlusion face recognition algorithms almost tend to pay more attention to the visible facial components. However, these models are limited because they heavily rely on existing face segmentation approaches to locate occlusions, which is extremely sensitive to the performance of mask learning. To tackle this issue, we propose a joint segmentation and identification feature learning framework for end-to-end occlusion face recognition. More particularly, unlike employing an external face segmentation model to locate the occlusion, we design an occlusion prediction module supervised by known mask labels to be aware of the mask. It shares underlying convolutional feature maps with the identification network and can be collaboratively optimized with each other. Furthermore, we propose a novel channel refinement network to cast the predicted single-channel occlusion mask into a multi-channel mask matrix with each channel owing a distinct mask map. Occlusion-free feature maps are then generated by projecting multi-channel mask probability maps onto original feature maps. Thus, it can suppress the representation of occlusion elements in both the spatial and channel dimensions under the guidance of the mask matrix. Moreover, in order to avoid misleading aggressively predicted mask maps and meanwhile actively exploit usable occlusion-robust features, we aggregate the original and occlusion-free feature maps to distill the final candidate embeddings by our proposed feature purification module. Lastly, to alleviate the scarcity of real-world occlusion face recognition datasets, we build large-scale synthetic occlusion face datasets, totaling up to 980193 face images of 10574 subjects for the training dataset and 36721 face images of 6817 subjects for the testing dataset, respectively. Extensive experimental results on the synthetic and real-world occlusion face datasets show that our approach significantly outperforms the state-of-the-art in both 1:1 face verification and 1:N face identification.
Baojin Huang, Zhongyuan Wang 0001, Kui Jiang, Qin Zou 0001, Xin Tian 0006, Tao Lu 0001, Zhen Han 0002
IEEE Trans. Neural Networks Learn. Syst.4
2022 Deepfake Video Detection Exploiting Binocular Synchronization
Zhongyuan Wang 0001, Guangcheng Wang, Qin Zou 0001
ICANN (3)4
2022 A Comparative Analysis of LiDAR SLAM-Based Indoor Navigation for Autonomous Vehicles
abstract
Simultaneous localization and mapping (SLAM) is a fundamental technique block in the indoor-navigation system for most autonomous vehicles and robots. SLAM aims at building a global consistent map of the environment while simultaneously determining the position and orientation of the robot in this map. Significant advances have been made in visual SLAM techniques in the past several years. However, due to the fragile performance in tracking feature points in environments that lack texture, e.g., a warehouse with blank white walls, visual SLAM can hardly provide a reliable localization. Compared with visual SLAM, LiDAR SLAM can often provide more robust localization in indoor environments by using 3D spatial information directly captured by LiDAR point clouds. Thus, LiDAR SLAM techniques are often employed in industrial applications such as automated guided vehicles (AGVs). In the past decades, a number of LiDAR SLAM methods have been proposed. However, the strength and weakness points of various LiDAR SLAMs are not clear, which may perplex the researchers and engineers. In this article, analysis and comparisons are made on different LiDAR SLAM-based indoor navigation methods, and extensive experiments are conducted to evaluate their performances in real environments. The comparative analysis and results can help researchers in academia and industry in constructing a suitable LiDAR SLAM system for indoor navigation for their own usage scenarios.
Qin Zou 0001, Qin Sun, Long Chen 0005, Bu Nie, Qingquan Li 0001
IEEE Trans. Intell. Transp. Syst.1
2022 Automatic Tunnel Crack Inspection Using an Efficient Mobile Imaging Module and a Lightweight CNN
abstract
Cracks in tunnel linings are the most common tunnel defects. As early indicators of structural deterioration, cracks represent critical problems for the safety of tunnels. Several mobile tunnel inspection systems (MTISs) have been developed for tunnel crack inspection. However, due to the weak signals of cracks, these MTISs require considerable exposure time to capture high-quality tunnel images, necessitating a low travel speed. Meanwhile, traditional crack detection methods encounter difficulties in processing tunnel crack images because of their low contrast and poor continuity. To overcome these challenges, this study presents a new MTIS for fast tunnel crack inspection that consists of a novel mobile imaging module and an automatic crack detection module. The imaging module is composed of an array of high-resolution charge-coupled device (CCD) cameras, a mobile laser scanner, and a lighting array. The core of the crack detection module is a novel lightweight convolutional neural network (CNN) designed for efficient tunnel crack detection, with an effective spatial constraint strategy to guarantee crack continuity. We collected a new tunnel crack dataset consisting of 1,218 images using our mobile imaging module at a driving speed of 80 km/h. Comprehensive experiments were conducted on this dataset to evaluate the performance of our proposed network. The results demonstrate that the presented CNN can effectively detect tunnel cracks with state-of-the-art performance, achieving an F1-score greater than 0.88 and an inference speed of 17 FPS with only 3.4M model parameters. The code and data are available athttps://github.com/urban-informatics/LinkCrack.
Jianghai Liao, Yuanhao Yue, Dejin Zhang, Wei Tu 0001, Rui Cao 0001, Qin Zou 0001, Qingquan Li 0001
IEEE Trans. Intell. Transp. Syst.6
2022 Transductive Zero-Shot Hashing for Multilabel Image Retrieval
abstract
Hash coding has been widely used in the approximate nearest neighbor search for large-scale image retrieval. Given semantic annotations such as class labels and pairwise similarities of the training data, hashing methods can learn and generate effective and compact binary codes. While some newly introduced images may contain undefined semantic labels, which we call unseen images, zero-shot hashing (ZSH) techniques have been studied for retrieval. However, existing ZSH methods mainly focus on the retrieval of single-label images and cannot handle multilabel ones. In this article, for the first time, a novel transductive ZSH method is proposed for multilabel unseen image retrieval. In order to predict the labels of the unseen/target data, a visual-semantic bridge is built via instance-concept coherence ranking on the seen/source data. Then, pairwise similarity loss and focal quantization loss are constructed for training a hashing model using both the seen/source and unseen/target data. Extensive evaluations on three popular multilabel data sets demonstrate that the proposed hashing method achieves significantly better results than the comparison methods.
Qin Zou 0001, Ling Cao, Zheng Zhang 0036, Long Chen 0005, Song Wang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2021 Object 6DoF Pose Estimation for Power Grid Manipulating Robots
Shan Du 0003, Xiaoye Zhang, Jingpeng Yue, Qin Zou 0001
ICIG (2)5
2021 Metric Learning for Anti-Compression Facial Forgery Detection
abstract
Detecting facial forgery images and videos is an increasingly important topic in multimedia forensics. As forgery images and videos are usually compressed into different formats such as JPEG and H264 when circulating on the Internet, existing forgery-detection methods trained on uncompressed data often suffer from significant performance degradation in identifying them. To solve this problem, we propose a novel anti-compression facial forgery detection framework, which learns a compression-insensitive embedding feature space utilizing both original and compressed forgeries. Specifically, our approach consists of three ideas: (i) extracting compression-insensitive features from both uncompressed and compressed forgeries using an adversarial learning strategy; (ii) learning a robust partition by constructing a metric loss that can reduce the distance of the paired original and compressed images in the embedding space; (iii) improving the accuracy of tampered localization with an attention-transfer module. Experimental results demonstrate that, the proposed method is highly effective in handling both compressed and uncompressed facial forgery images.
Shenhao Cao, Qin Zou 0001, Xiuqing Mao, Dengpan Ye, Zhongyuan Wang 0001
ACM Multimedia2
2021 Luminance Attentive Networks for HDR Image and Panorama Reconstruction
abstract
Abstract It is very challenging to reconstruct a high dynamic range (HDR) from a low dynamic range (LDR) image as an ill‐posed problem. This paper proposes a luminance attentive network named LANet for HDR reconstruction from a single LDR image. Our method is based on two fundamental observations: (1) HDR images stored in relative luminance are scale‐invariant, which means the HDR images will hold the same information when multiplied by any positive real number. Based on this observation, we propose a novel normalization method called “HDR calibration“for HDR images stored in relative luminance, calibrating HDR images into a similar luminance scale according to the LDR images. (2) The main difference between HDR images and LDR images is in under‐/over‐exposed areas, especially those highlighted. Following this observation, we propose a luminance attention module with a two‐stream structure for LANet to pay more attention to the under‐/over‐exposed areas. In addition, we propose an extended network called panoLANet for HDR panorama reconstruction from an LDR panorama and build a dualnet structure for panoLANet to solve the distortion problem caused by the equirectangular panorama. Extensive experiments show that our proposed approach LANet can reconstruct visually convincing HDR images and demonstrate its superiority over state‐of‐the‐art approaches in terms of all metrics in inverse tone mapping. The image‐based lighting application with our proposed panoLANet also demonstrates that our method can simulate natural scene lighting using only LDR panorama. Our source code is available at https://github.com/LWT3437/LANet .
Hanning Yu, Chengjiang Long, Bo Dong 0004, Qin Zou 0001, Chunxia Xiao
Comput. Graph. Forum5
2020 Learning Metric Features for Writer-Independent Signature Verification using Dual Triplet Loss
abstract
Handwritten signature has long been a widely accepted biometric and applied in many verification scenarios. However, automatic signature verification remains an open research problem, which is mainly due to three reasons. 1) Skilled forgeries generated by persons who imitate the original writing pattern are very difficult to be distinguished from genuine signatures. It is especially so in the case of offline signatures, where only the signature image is captured as a feature for verification. 2) Most state-of-the-art models are writer-dependent, requiring a specific model to be trained whenever a new user is registered in verification, which is quite inconvenient. 3) Writer-independent models often have unsatisfactory performance. To this end, we propose a novel metric learning based method for offline writer-independent signature verification. Specifically, a dual triplet loss is used to train the model, where two different triplets are constructed for random and skilled forgeries, respectively. Experiments on three alphabet datasets - GPDS Synthetic, MCYT and CEDAR - show that the proposed method achieves competitive or superior performance to the state-of-the-art methods. Experiments are also conducted on a new offline Chinese signature dataset - CSIG-WHU, and the results show that the proposed method has a high feasibility on character-based signatures.
Qian Wan 0004, Qin Zou 0001
ICPR2
2020 Low-quality watermarked face inpainting with discriminative residual learning
abstract
Most existing image inpainting methods assume that the location of the repair area (watermark) is known, but this assumption does not always hold. In addition, the actual watermarked face is in a compressed low-quality form, which is very disadvantageous to the repair due to compression distortion effects. To address these issues, this paper proposes a low-quality watermarked face inpainting method based on joint residual learning with cooperative discriminant network. We first employ residual learning based global inpainting and facial features based local inpainting to render clean and clear faces under unknown watermark positions. Because the repair process may distort the genuine face, we further propose a discriminative constraint network to maintain the fidelity of repaired faces. Experimentally, the average PSNR of inpainted face images is increased by 4.16dB, and the average SSIM is increased by 0.08. TPR is improved by 16.96% when FPR is 10% in face verification.
Zheng He 0001, Xueli Wei, Kangli Zeng, Zhen Han 0002, Qin Zou 0001, Zhongyuan Wang 0001
MMAsia5
2020 Deep Domain Adaptation With Differential Privacy
abstract
Nowadays, it usually requires a massive amount of labeled data to train a deep neural network. When no labeled data is available in some application scenarios, domain adaption can be employed to transfer a learner from one or more source domains with labeled data to a target domain with unlabeled data. However, due to the exposure of the trained model to the target domain, the user privacy may potentially be compromised. Nevertheless, the private information may be encoded into the representations in different stages of the deep neural networks, i.e., hierarchical convolutional feature maps, which poses a great challenge for a full-fledged privacy protection. In this paper, we propose a novel differentially private domain adaptation framework called DPDA to achieve domain adaptation with privacy assurance. Specifically, we perform domain adaptation in an adversarial-learning manner and embed the differentially private design into specific layers and learning processes. Although applying differential privacy techniques directly will undermine the performance of deep neural networks, DPDA can increase the classification accuracy for the unlabeled target data compared to the prior arts. We conduct extensive experiments on standard benchmark datasets, and the results show that our proposed DPDA can indeed achieve high accuracy in many domain adaptation tasks with only a modest privacy loss.
Qian Wang 0002, Qin Zou 0001, Lingchen Zhao, Song Wang 0002
IEEE Trans. Inf. Forensics Secur.3
2020 Privacy-Preserving Collaborative Deep Learning With Unreliable Participants
abstract
With powerful parallel computing GPUs and massive user data, neural-network-based deep learning can well exert its strong power in problem modeling and solving, and has archived great success in many applications such as image classification, speech recognition and machine translation etc. While deep learning has been increasingly popular, the problem of privacy leakage becomes more and more urgent. Given the fact that the training data may contain highly sensitive information, e.g., personal medical records, directly sharing them among the users (i.e., participants) or centrally storing them in one single location may pose a considerable threat to user privacy. In this paper, we present a practical privacy-preserving collaborative deep learning system that allows users to cooperatively build a collective deep learning model with data of all participants, without direct data sharing and central data storage. In our system, each participant trains a local model with their own data and only shares model parameters with the others. To further avoid potential privacy leakage from sharing model parameters, we use functional mechanism to perturb the objective function of the neural network in the training process to achieve ε-differential privacy. In particular, for the first time, we consider the existence of unreliable participants, i.e., the participants with low-quality data, and propose a solution to reduce the impact of these participants while protecting their privacy. We evaluate the performance of our system on two well-known real-world datasets for regression and classification tasks. The results demonstrate that the proposed system is robust against unreliable participants, and achieves high accuracy close to the model trained in a traditional centralized manner while ensuring rigorous privacy protection.
Lingchen Zhao, Qian Wang 0002, Qin Zou 0001, Yan Zhang 0002, Yanjiao Chen
IEEE Trans. Inf. Forensics Secur.3
2020 Deep Learning-Based Gait Recognition Using Smartphones in the Wild
abstract
Compared to other biometrics, gait is difficult to conceal and has the advantage of being unobtrusive. Inertial sensors, such as accelerometers and gyroscopes, are often used to capture gait dynamics. These inertial sensors are commonly integrated into smartphones and are widely used by the average person, which makes gait data convenient and inexpensive to collect. In this paper, we study gait recognition using smartphones in the wild. In contrast to traditional methods, which often require a person to walk along a specified road and/or at a normal walking speed, the proposed method collects inertial gait data under unconstrained conditions without knowing when, where, and how the user walks. To obtain good person identification and authentication performance, deep-learning techniques are presented to learn and model the gait biometrics based on walking data. Specifically, a hybrid deep neural network is proposed for robust gait feature representation, where features in the space and time domains are successively abstracted by a convolutional neural network and a recurrent neural network. In the experiments, two datasets collected by smartphones for a total of 118 subjects are used for evaluations. The experiments show that the proposed method achieves higher than 93.5% and 93.7% accuracies in person identification and authentication, respectively.
Qin Zou 0001, Qian Wang 0002, Yi Zhao 0011, Qingquan Li 0001
IEEE Trans. Inf. Forensics Secur.1
2020 Surrounding Vehicle Detection Using an FPGA Panoramic Camera and Deep CNNs
abstract
Surrounding vehicle detection is one of the most important modules for a vision-based driver assistance system (VB-DAS) or an autonomous vehicle. In this paper, we put forward a wireless panoramic camera system for real-time and seamless imaging of the 360-degree driving scene. Using an embedded FPGA design, the proposed panoramic camera system can perform fast image stitching and produce panoramic videos in real-time, which greatly relives the computation and storage burden of a traditional multi-camera-based panoramic system. For surrounding vehicle detection, we present a novel deep convolutional neural network - EZ-Net, which perceives the potential vehicles by using 13 convolutional layers and locates the vehicles by a local non-maximum suppression process. Experimental results demonstrate that, the proposed EZ-Net performs vehicle detection on the panoramic video at a speed of 140 fps while holding a competing accuracy with the state-of-the-art detectors.
Long Chen 0005, Qin Zou 0001, Ziyu Pan, Danyu Lai, Liwei Zhu, Zhoufan Hou, Jun Wang 0015, Dongpu Cao
IEEE Trans. Intell. Transp. Syst.2
2020 Improved Deep Hashing With Soft Pairwise Similarity for Multi-Label Image Retrieval
abstract
Hash coding has been widely used in the approximate nearest neighbor search for large-scale image retrieval. Recently, many deep hashing methods have been proposed and shown largely improved performance over traditional feature-learning methods. Most of these methods examine the pairwise similarity on the semantic-level labels, where the pairwise similarity is generally defined in a hard-assignment way. That is, the pairwise similarity is “1” if they share no less than one class label and “0” if they do not share any. However, such similarity definition cannot reflect the similarity ranking for pairwise images that hold multiple labels. In this paper, an improved deep hashing method is proposed to enhance the ability of multi-label image retrieval. We introduce a pairwise quantified similarity calculated on the normalized semantic labels. Based on this, we divide the pairwise similarity into two situations-“hard similarity” and “soft similarity,” where cross-entropy loss and mean square error loss are adapted respectively for more robust feature learning and hash coding. Experiments on four popular datasets demonstrate that the proposed method outperforms the competing methods and achieves the state-of-the-art performance in multi-label image retrieval.
Zheng Zhang 0036, Qin Zou 0001, Yuewei Lin, Long Chen 0005, Song Wang 0002
IEEE Trans. Multim.2
2019 High-Resolution Driving Scene Synthesis Using Stacked Conditional Gans and Spectral Normalization
abstract
Large-scale dataset plays a key role in the driving scene understanding for deep learning based-autonomous driving tasks. Due to the fact that the annotation for a large number of images is extremely labor-intensive and time-consuming, many researchers turn to using image-synthesis techniques for automatic construction of training data. However, traditional methods often have difficulties in producing high-definition driving scene images. To tackle this problem, in this paper, we propose a novel deep model - hdCGAN - for high-definition image-to-image translation. The hdCGAN is built on a conditional GAN in combination with a spectral normalization. Moreover, we improve the hdCGAN by using a stacked network architecture and the enhanced model is called stack-hdCGAN. With the guidance of multi-scale discriminators and the constraint of spectral normalization in the training procedure, the learned models can generate high-resolution and high-quality driving scene images from corresponding semantic segmentation maps. Quantitative and qualitative evaluations on the Cityscapes dataset demonstrate the effectiveness of the proposed models.
Shaobo Lin, Long Chen 0005, Qin Zou 0001, Wei Tian 0001
ICME3
2019 Deep Integration: A Multi-Label Architecture for Road Scene Recognition
abstract
Deep convolutional neural networks have been applied by automobile industries, Internet giants, and academic institutes to boost autonomous driving technologies; while progress has been witnessed in environmental perception tasks, such as object detection and driver state recognition, the scene-centric understanding and identification still remain a virgin land. This mainly encompasses two key issues: 1) the lack of shared large datasets with comprehensively annotated road scene information and 2) the difficulty to find effective ways to train networks concerning the bias of category samples, image resolutions, scene dynamics, and capturing conditions. In this paper, we make two contributions: 1) we introduce a large-scale dataset with over 110 k images, dubbed DrivingScene, covering traffic scenarios under different weather conditions, road structures, and environmental instances and driving places, which is the first large-scale dataset for multi-class traffic scenes classification and 2) we propose a multi-label neural network for road scene recognition, which incorporates both single- and multi-class classification modes into a multi-level cost function for training with imbalanced categories and utilizes a deep data integration strategy to improve the classification ability on hard samples. The experimental results on DrivingScene and PASCAL VOC demonstrate the effectiveness of the proposed approach in handling the challenge of data imbalance.
Long Chen 0005, Wujing Zhan, Wei Tian 0001, Qin Zou 0001
IEEE Trans. Image Process.5
2019 DeepCrack: Learning Hierarchical Convolutional Features for Crack Detection
abstract
Cracks are typical line structures that are of interest in many computer-vision applications. In practice, many cracks, e.g., pavement cracks, show poor continuity and low contrast, which brings great challenges to image-based crack detection by using low-level features. In this paper, we propose DeepCrack - an end-to-end trainable deep convolutional neural network for automatic crack detection by learning high-level features for crack representation. In this method, multi-scale deep convolutional features learned at hierarchical convolutional stages are fused together to capture the line structures. More detailed representations are made in larger-scale feature maps and more holistic representations are made in smaller-scale feature maps. We build DeepCrack net on the encoder-decoder architecture of SegNet, and pairwisely fuse the convolutional features generated in the encoder network and in the decoder network at the same scale. We train DeepCrack net on one crack dataset and evaluate it on three others. The experimental results demonstrate that DeepCrack achieves F-Measure over 0.87 on the three challenging datasets in average and outperforms the current state-of-the-art methods.
Qin Zou 0001, Zheng Zhang 0036, Qingquan Li 0001, Xianbiao Qi, Qian Wang 0002, Song Wang 0002
IEEE Trans. Image Process.1
2018 Dating ancient paintings of Mogao Grottoes using deeply learnt visual codes
Qingquan Li 0001, Qin Zou 0001, De Ma, Qian Wang 0002, Song Wang 0002
Sci. China Inf. Sci.2
2018 Multiple human tracking in wearable camera videos with informationless intervals
Hongkai Yu, Haozhou Yu, Hao Guo 0002, Jeff P. Simmons, Qin Zou 0001, Wei Feng 0005, Song Wang 0002
Pattern Recognit. Lett.5
2018 Robust Gait Recognition by Integrating Inertial and RGBD Sensors
abstract
Gait has been considered as a promising and unique biometric for person identification. Traditionally, gait data are collected using either color sensors, such as a CCD camera, depth sensors, such as a Microsoft Kinect, or inertial sensors, such as an accelerometer. However, a single type of sensors may only capture part of the dynamic gait features and make the gait recognition sensitive to complex covariate conditions, leading to fragile gait-based person identification systems. In this paper, we propose to combine all three types of sensors for gait data collection and gait recognition, which can be used for important identification applications, such as identity recognition to access a restricted building or area. We propose two new algorithms, namely EigenGait and TrajGait, to extract gait features from the inertial data and the RGBD (color and depth) data, respectively. Specifically, EigenGait extracts general gait dynamics from the accelerometer readings in the eigenspace and TrajGait extracts more detailed subdynamics by analyzing 3-D dense trajectories. Finally, both extracted features are fed into a supervised classifier for gait recognition and person identification. Experiments on 50 subjects, with comparisons to several other state-of-the-art gait-recognition approaches, show that the proposed approach can achieve higher recognition accuracy and robustness.
Qin Zou 0001, Lihao Ni, Qian Wang 0002, Qingquan Li 0001, Song Wang 0002
IEEE Trans. Cybern.1
2018 Searchable Encryption over Feature-Rich Data
abstract
Storage services allow data owners to store their huge amount of potentially sensitive data, such as audios, images, and videos, on remote cloud servers in encrypted form. To enable retrieval of encrypted files of interest, searchable symmetric encryption (SSE) schemes have been proposed. However, many schemes construct indexes based on keyword-file pairs and focus on boolean expressions of exact keyword matches. Moreover, most dynamic SSE schemes cannot achieve forward privacy and reveal unnecessary information when updating the encrypted databases. We tackle the challenge of supporting large-scale similarity search over encrypted feature-rich multimedia data, by considering the search criteria as a high-dimensional feature vector instead of a keyword. Our solutions are built on carefully-designed fuzzy Bloom filters which utilize locality sensitive hashing (LSH) to encode an index associating the file identifiers and feature vectors. Our schemes are proven to be secure against adaptively chosen query attack and forward private in the standard model. We have evaluated the performance of our scheme on real-world high-dimensional datasets, and achieved a search quality of 99 percent recall with only a few number of hash tables for LSH. This shows that our index is compact and searching is not only efficient but also accurate.
Qian Wang 0002, Meiqi He, Minxin Du, Sherman S. M. Chow, Russell W. F. Lai, Qin Zou 0001
IEEE Trans. Dependable Secur. Comput.6
2017 Unsupervised Simplification of Image Hierarchies via Evolution Analysis in Scale-Sets Framework
abstract
Region-based hierarchical image representation is crucial in many computer vision applications. However, in practice, an image hierarchy is usually dense, and contains many less informative branches. It is expected that a hierarchy should be accurate and simplified, which is not only desirable for different applications, but also saves considerable computational load for the further analysis. To achieve this target, this paper proposes a novel approach for unsupervised simplification of region-based image hierarchies, which employs the global and local evolution analyses of a hierarchy. First, we introduce a global evolution analysis in the scale-sets framework, which provides clues for eliminating less informative branches. Moreover, a hybrid unsupervised simplification method is designed, utilizing the information from global and local evolution functions. A number of experiments on various images have shown that the proposed approach is effective and efficient in removing less informative nodes (averagely about 90% of the whole nodes), while preserving salient image details and retaining the accuracy.
Zhongwen Hu, Qingquan Li 0001, Qian Zhang 0046, Qin Zou 0001, Zhaocong Wu
IEEE Trans. Image Process.4
2017 Transforming a 3-D LiDAR Point Cloud Into a 2-D Dense Depth Map Through a Parameter Self-Adaptive Framework
abstract
The 3-D LiDAR scanner and the 2-D charge-coupled device (CCD) camera are two typical types of sensors for surrounding-environment perceiving in robotics or autonomous driving. Commonly, they are jointly used to improve perception accuracy by simultaneously recording the distances of surrounding objects, as well as the color and shape information. In this paper, we use the correspondence between a 3-D LiDAR scanner and a CCD camera to rearrange the captured LiDAR point cloud into a dense depth map, in which each 3-D point corresponds to a pixel at the same location in the RGB image. In this paper, we assume that the LiDAR scanner and the CCD camera are accurately calibrated and synchronized beforehand so that each 3-D LiDAR point cloud is aligned with its corresponding RGB image. Each frame of the LiDAR point cloud is then projected onto the RGB image plane to form a sparse depth map. Then, a self-adaptive method is proposed to upsample the sparse depth map into a dense depth map, in which the RGB image and the anisotropic diffusion tensor are exploited to guide upsampling by reinforcing the RGB-depth compactness. Finally, convex optimization is applied on the dense depth map for global enhancement. Experiments on the KITTI and Middlebury data sets demonstrate that the proposed method outperforms several other relevant state-of-the-art methods in terms of visual comparison and root-mean-square error measurement.
Long Chen 0005, Jianda Chen, Qingquan Li 0001, Qin Zou 0001
IEEE Trans. Intell. Transp. Syst.5
2017 Local Pattern Collocations Using Regional Co-occurrence Factorization
abstract
Human vision benefits a lot from pattern collocations in visual activities such as object detection and recognition. Usually, pattern collocations display as the co-occurrences of visual primitives, e.g., colors, gradients, or textons, in neighboring regions. In the past two decades, many sophisticated local feature descriptors have been developed to describe visual primitives, and some of them even take into account the co-occurrence information for improving their discriminative power. However, most of these descriptors only consider feature co-occurrence within a very small neighborhood, e.g., 8-connected or 16-connected area, which would fall short in describing pattern collocations built up by feature co-occurrences in a wider neighborhood. In this paper, we propose to describe local pattern collocations by using a new and general regional co-occurrence approach. In this approach, an input image is first partitioned into a set of homogeneous superpixels. Then, features in each superpixel are extracted by a variety of local feature descriptors, based on which a number of patterns are computed. Finally, pattern co-occurrences within the superpixel and between the neighboring superpixels are calculated and factorized into a final descriptor for local pattern collocation. The proposed regional co-occurrence framework is extensively tested on a wide range of popular shape, color, and texture descriptors in terms of image and object categorizations. The experimental results have shown significant performance improvements by using the proposed framework over the existing popular descriptors.
Qin Zou 0001, Lihao Ni, Qian Wang 0002, Zhongwen Hu, Qingquan Li 0001, Song Wang 0002
IEEE Trans. Multim.1
2016 SecHOG: Privacy-Preserving Outsourcing Computation of Histogram of Oriented Gradients in the Cloud
abstract
Abundant multimedia data generated in our daily life has intrigued a variety of very important and useful real-world applications such as object detection and recognition etc. Accompany with these applications, many popular feature descriptors have been developed, e.g., SIFT, SURF and HOG. Manipulating massive multimedia data locally, however, is a storage and computation intensive task, especially for resource-constrained clients. In this work, we focus on exploring how to securely outsource the famous feature extraction algorithm--Histogram of Oriented Gradients (HOG) to untrusted cloud servers, without revealing the data owner's private information. For the first time, we investigate this secure outsourcing computation problem under two different models and accordingly propose two novel privacy-preserving HOG outsourcing protocols, by efficiently encrypting image data by somewhat homomorphic encryption (SHE) integrated with single-instruction multiple-data (SIMD), designing a new batched secure comparison protocol, and carefully redesigning every step of HOG to adapt it to the ciphertext domain. Explicit Security and effectiveness analysis are presented to show that our protocols are practically-secure and can approximate well the performance of the original HOG executed in the plaintext domain. Our extensive experimental evaluations further demonstrate that our solutions achieve high efficiency and perform comparably to the original HOG when being applied to human detection.
Qian Wang 0002, Jingjun Wang, Shengshan Hu, Qin Zou 0001, Kui Ren 0001
AsiaCCS4
2016 Geodesic-based pavement shadow removal revisited
abstract
Shadows often incur uneven illumination to pavement images, which brings great challenges to image-based pavement crack detection. Thus, it is desired to remove pavement shadows before detecting pavement cracks. However, due to the large penumbras cast by trees, light poles, etc., it is difficult to locate shadows in a pavement image. In this paper, an automatic pavement shadow removal method is proposed based on geodesic analysis. First, a geodesic shadow model is used to partition a pavement shadow into a number of geodesic regions. Then, an optimal background region is selected for reference by statistic analysis. Finally, a texture-balanced illuminance compensation is applied on all geodesic regions over the image. Experiments demonstrate the effectiveness of the proposed method.
Qin Zou 0001, Zhongwen Hu, Long Chen 0005, Qian Wang 0002, Qingquan Li 0001
ICASSP1
2016 A Bilevel Scale-Sets Model for Hierarchical Representation of Large Remote Sensing Images
abstract
Due to the diversity of geographical objects, it makes great sense to introduce multiscale segmentation/representation into the analysis and interpretation of high-spatial-resolution remote sensing images. However, with the increasing use of high-resolution images, traditional multiscale segmentation methods gradually show their lack in efficiency, particularly when handling large-scale images. In this paper, a novel bilevel scale-sets model (BSM) is proposed for multiscale region-based representation of large-scale remote sensing images. In the BSM, first, an image is divided into blocks with overlapped margins, and a low-level scale-sets model is blockwisely implemented. Second, a segmentation result is obtained by retrieving and mosaicking the blockwise segmentation results, based on which a high-level scale-sets model is implemented covering the whole image. To further improve the efficiency of the BSM, a parallel implementation is presented for the blockwise scale-sets model. In the experiments, first, the effectiveness of the BSM is validated using a WorldView2 image covering a coastal area of Shenzhen, where the BSM obtains accurate multiscale representation results without any mosaic artifacts. Then, the efficiency of the BSM is demonstrated by comparing with the state-of-the-art multiscale segmentation method, i.e., the one integrated in the commercial software eCognition v9.2, where the proposed BSM takes about 7 min to process a 24 000 × 24 000 multispectral ZY3 image and is two to three times faster than the competing method.
Zhongwen Hu, Qingquan Li 0001, Qin Zou 0001, Qian Zhang 0046, Guofeng Wu
IEEE Trans. Geosci. Remote. Sens.3
2015 Walls Have Ears! Opportunistically Communicating Secret Messages Over the Wiretap Channel: from Theory to Practice
abstract
Physical layer (PHY) security has aroused great research interest in recent years, exploiting physical uncertainty of wireless channels to provide communication secrecy without placing any computational restrictions on the adversaries under the information-theoretic security model. Particularly, researches have been focused on investigating Wyner's Wiretap Channel for constructing practical wiretap codes that can achieve simultaneous transmission secrecy and reliability. While theoretically sound, PHY security through the wiretap channel has never been realized in practice, and the feasibility and physical limitations of implementing such channels in the real world are yet to be well understood. In this paper, we design and implement a practical opportunistic secret communication system over the wireless wiretap channel for the first time to our best knowledge. We show that, our system can achieve nearly perfect secrecy given a fixed codeword length by carefully controlling the structure of the parity-check matrix of wiretap codes to strike the proper balance between the transmission rate and secrecy. Our system is implemented and evaluated extensively on a USRP N210-based testbed. The experimental results demonstrate the physical limitations and the feasibility of building practical wiretap channels in both the worst channel case and the case where the sender has only the knowledge of instantaneous channel capacities. Our system design and implementation successfully attempts towards bridging the gap between the theoretical wiretap channel and its practice, alleviating the unrealistic and strong assumptions imposed by the theoretical model.
Qian Wang 0002, Kui Ren 0001, Guancheng Li, Chenbo Xia, Xiaobing Chen, Zhibo Wang 0001, Qin Zou 0001
CCS7
2015 Watershed superpixel
abstract
As a pre-processing tool, superpixel algorithms have been popular used in many computer-vision applications. High efficiency is a desired property of superpixel algorithms, especially in real-time vision systems. In this paper, a novel high-efficient superpixel algorithm is developed based on the watershed algorithm, namely the spatial-constrained watershed (SCoW). SCoW performs watersheding in a marker-controlled manner, with a set of evenly placed markers. To align superpixel boundaries to image edges, an edge-preserving scheme is embedded into the SCoW which makes a balance between the homogeneity and the compactness. Without any complex computing, the proposed superpixel algorithm is found to produce high quality superpixels as traditional superpixel algorithms, while holding much higher efficiency.
Zhongwen Hu, Qin Zou 0001, Qingquan Li 0001
ICIP2
2015 Discriminative regional color co-occurrence descriptor
abstract
Traditional color feature descriptors are focused on color-value distributions in the color space, e.g., color histograms, color bag-of-words, which ignore the spatial location and contextual information of different colors. In this paper, a new regional color co-occurrence feature descriptor (RCC) is proposed to reflect spatial relations of colors in an image. First, we partition an image into a number of disjoint regions using superpixel techniques. Then, we construct a color histogram for each region, based on which we construct a color co-occurrence matrix for each pair of neighboring regions. Finally, all the constructed co-occurrence matrices from an image are summed up and normalized as a color descriptor to represent this image. This new color descriptor reflects the color-collocation patterns in the image. We use this new color descriptor for image/object classification and find that it leads to higher classification accuracies than other competing color descriptors.
Qin Zou 0001, Xianbiao Qi, Qingquan Li 0001, Song Wang 0002
ICIP1
2015 Deep Learning Based Feature Selection for Remote Sensing Scene Classification
abstract
With the popular use of high-resolution satellite images, more and more research efforts have been placed on remote sensing scene classification/recognition. In scene classification, effective feature selection can significantly boost the final performance. In this letter, a novel deep-learning-based feature-selection method is proposed, which formulates the feature-selection problem as a feature reconstruction problem. Note that the popular deep-learning technique, i.e., the deep belief network (DBN), achieves feature abstraction by minimizing the reconstruction error over the whole feature set, and features with smaller reconstruction errors would hold more feature intrinsics for image representation. Therefore, the proposed method selects features that are more reconstructible as the discriminative features. Specifically, an iterative algorithm is developed to adapt the DBN to produce the inquired reconstruction weights. In the experiments, 2800 remote sensing scene images of seven categories are collected for performance evaluation. Experimental results demonstrate the effectiveness of the proposed method.
Qin Zou 0001, Lihao Ni, Tong Zhang 0009, Qian Wang 0002
IEEE Geosci. Remote. Sens. Lett.1
2014 Automatic inpainting by removing fence-like structures in RGBD images
Qin Zou 0001, Yu Cao 0003, Qingquan Li 0001, Qingzhou Mao, Song Wang 0002
Mach. Vis. Appl.1
2014 A kernel support vector machine-based feature selection approach for recognizing Flying Apsaras' streamers in the Dunhuang Grotto Murals, China
Zhong Chen 0003, Shengwu Xiong 0001, Zhixiang Fang, Qingquan Li 0001, Qin Zou 0001
Pattern Recognit. Lett.6
2014 Chronological classification of ancient paintings using appearance and shape features
Qin Zou 0001, Yu Cao 0003, Qingquan Li 0001, Chuanhe Huang, Song Wang 0002
Pattern Recognit. Lett.1
2013 Survey on context-awareness in ubiquitous media
Daqiang Zhang 0001, Hongyu Huang 0001, Chin-Feng Lai, Xuedong Liang, Qin Zou 0001, Minyi Guo
Multim. Tools Appl.5
2012 CrackTree: Automatic crack detection from pavement images
Qin Zou 0001, Yu Cao 0003, Qingquan Li 0001, Qingzhou Mao, Song Wang 0002
Pattern Recognit. Lett.1
2011 A Multichannel Edge-Weighted Centroidal Voronoi Tessellation algorithm for 3D super-alloy image segmentation
abstract
In material science and engineering, the grain structure inside a super-alloy sample determines its mechanical and physical properties. In this paper, we develop a new Multichannel Edge-Weighted Centroidal Voronoi Tessellation (MCEWCVT) algorithm to automatically segment all the 3D grains from microscopic images of a super-alloy sample. Built upon the classical k-means/CVT algorithm, the proposed algorithm considers both the voxel-intensity similarity within each cluster and the compactness of each cluster. In addition, the same slice of a super-alloy sample can produce multiple images with different grain appearances using different settings of the microscope. We call this multichannel imaging and in this paper, we further adapt the proposed segmentation algorithm to handle such multichannel images to achieve higher grain-segmentation accuracy. We test the proposed MCEWCVT algorithm on a 4-channel Ni-based 3D super-alloy image consisting of 170 slices. The segmentation performance is evaluated against the manually annotated ground-truth segmentation and quantitatively compared with other six image segmentation/edge-detection methods. The experimental results demonstrate the higher accuracy of the proposed algorithm than the comparison methods.
Yu Cao 0003, Lili Ju, Qin Zou 0001, Chengzhang Qu, Song Wang 0002
CVPR3
2011 FoSA: F* Seed-growing Approach for crack-line detection from pavement images
Qingquan Li 0001, Qin Zou 0001, Daqiang Zhang 0001, Qingzhou Mao
Image Vis. Comput.2
2010 Block-constraint line scanning method for lane detection
abstract
Considering the plentiful road markings in China, we present a Block-Constraint Line scanning (BCLS) method for lane detection in this paper. In this method, images are firstly pre-processed by a morphological top-hat transform, and then an imaging model is created for building relationship between lane parameters of the image coordinate and the WGS coordinate, from which target points on lane lines could be retained by a block-constraint line scanning algorithm. Finally, lanes could be extracted by a Progressive Probabilistic Hough Transform (PPHT) and the number of lanes is figured out through clustering. Our method is fast enough to meet real-time requirement. Experiments were carried out on the intelligent vehicle SmartV (Fig.1) on the Wuhan urban roads in China and the results show that this method can efficiently and accurately extract lanes in complex environments, even with the presence of non-lane road markings.
Long Chen 0005, Qingquan Li 0001, Qingzhou Mao, Qin Zou 0001
Intelligent Vehicles Symposium4