EDBT 2026 Demo / reviewers in the wild / expert
Qi Zou 0001
dblp:72/6963-1
· DBLP profile ↗
38ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0002-8070-5267ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DR-VAD: Definition-guided reasoning for training-free video anomaly detection
Yanting Pei, Minhao Hao, Qi Zou 0001 |
Neurocomputing | 4 |
| 2026 | Hierarchical Contrastive Consistency for Human Pose Estimation in Images and VideosabstractHuman pose estimation (HPE) is an invaluable task in computer vision with various practical applications. This paper proposes a novel Hierarchical Contrastive Consistensy constraint (HICCON) to improve the HPE in both images and videos, which describes the input into multi-granular representations at spatial and temporal domain and performs multi-level feature consistency by exploring the characteristic of human structure and time sequence. The hierarchical contrast is conducted at four levels: keypoint-level, part-level, instance-level and clip-level. In spatial, we consider keypoint-level and part-level consistency across instances within frame to enhance the fine-grained keypoint robustness. The former conducts the single keypoint feature contrast across instances to improve the category-specific keypoint features. The latter explores the specific pair-wise features for preserving the instructive relation. In temporal, we develop the instance-level and clip-level feature consistency across frames to capture more discriminative temporal representations. The former discriminates instance features across frames within the same video, whereas the clip-level constraint aims to discriminate consistent features from different videos in order to capture more distinctive temporal features. Extensive experiments on kinds of architectures across datasets i.e, PoseTrack2017, PoseTrack2018 and PoseTrack2021 show the HICCON achieves about 1.5% improvement than baseline. Besides, the proposed method unleashes the potential of the contrastive learning in HPE field. Xixia Xu, Qi Zou 0001, Jiamao Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | TrajDiffRefine: refinement of spatio-temporal stochastic trajectory prediction via diffusion
Xiangyun Tan, Qi Zou 0001 |
Appl. Intell. | 2 |
| 2025 | Coarse-to-fine text injecting for realistic image super-resolution
Chao Bai, Zhenyao Wu, Xinyi Wu 0002, Qi Zou 0001, Song Wang 0002 |
Neurocomputing | 5 |
| 2025 | Human-object interaction detection via recycling of ground-truth annotations
Xue Lin 0003, Qi Zou 0001, Xixia Xu |
Pattern Recognit. | 2 |
| 2025 | Multi-Person Pose Estimation with Feature Enhancement and Decoupling Based on Contrastive LearningabstractMost methods of multi-person pose estimation (MPPE) treat the human detection and keypoint localization separately. They need additional supervision like instance bounding boxes, or complex hand-crafted processes like RoI cropping or grouping. In this article, we propose a novel one-stage MPPE method, named COPE, which unifies human detection and keypoint regression into an end-to-end learnable framework. To handle the challenges plague one-stage MPPE, i.e., instance overlapping and misalignment of local and global context, we design contrastive constraints at two levels of semantic granularity and feature sampling strategies. Based on a whole-process differentiable pipeline, COPE establishes a simple yet effective framework for MPPE without additional instance-level supervision and resource-intensive modules like transformer. Benefit from specially designed contrastive constraints and sampling strategies, COPE can better handle occluded scenes and correct keypoint localization errors. Extensive experiments demonstrate COPE’s superiority. It attains 71.3 AP and 18.0 FPS on COCO val2017, effectively balancing accuracy and speed. Particularly in crowded and occluded scenarios, COPE achieves state-of-the-art performance on CrowdPose and OCHuman, surpassing CID by 0.6 AP and 1.7 AP, respectively. Furthermore, COPE strongly improves generalization performance on the Human-Art benchmark, outperforming ED-Pose by 6.7 AP and ClickPose by 3.7 AP. Qi Zou 0001, Xixia Xu, Yanting Pei |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Rethinking the Sparse End-to-End Multiperson Pose EstimationabstractCurrent methods of multiperson pose estimation (MPPE) typically treat the human detection and association of joints separately. They introduce complex hand-crafted pose-processes like RoI cropping, NMS and grouping or rely on dense representations to preserve the spatial features. In this article, we dive a deeper thought into this task and propose a simpler and effective framework, termed SparsePose, which can directly predict multiperson joint coordinates from the full image without any post-processes and dense representations. In SparsePose, the full-body instances are decoupled by exploring spatial-aware feature learning (SFL) without box and classification supervision. For improving the quality of instance map, the instance contrastive constraint (ICC) and center correction (CC) strategy are proposed to make the instance-wise spatial feature more discriminative. Importantly, we propose a visibility-guided weighting mechanism to enable model be confident to the visible joint predictions and insensitive to the occlusions or partial bodies. In general, SparsePose is conceptually simpler and plays favorably against the existing counterparts on three benchmarks in terms of both accuracy and efficiency. Xixia Xu, Qi Zou 0001, Jiamao Li |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | Inter-image Contrastive Consistency for Multi-Person Pose EstimationabstractMulti-person pose estimation (MPPE) has achieved impressive progress in recent years. However, due to the large variance of appearances among images or occlusions, the model can hardly learn consistent patterns enough, which leads to severe location jitter and missing issues. In this study, we propose a novel framework, termed Inter-image Contrastive consistency (ICON), to strengthen the keypoint consistency among images for MPPE. Concretely, we consider two-fold consistency constraints, which include single keypoint contrastive consistency (SKCC) and pair relation contrastive consistency (PRCC). The SKCC learns to strengthen the consistency of individual keypoints across images in the same category to improve the category-specific robustness. Only with SKCC, the model can effectively reduce location errors caused by large appearance variations, but remains challenging with extreme postures (e.g., occlusions) due to lack of relational guidance. Therefore, PRCC is proposed to strengthen the consistency of pair-wise joint relation between images to preserve the instructive relation. Cooperating with SKCC, PRCC further improves structure aware robustness in handling extreme postures. Extensive experiments on kinds of architectures across three datasets (i.e., MS-COCO, MPII, CrowdPose) show the proposed ICON achieves substantial improvements over baselines. Furthermore, ICON under the semi-supervised setup can obtain comparable results with the fully-supervised methods using only 30% labeled data. Xixia Xu, Yingguo Gao, Xingjia Pan, Qi Zou 0001 |
AAAI | 6 |
| 2023 | Structure-Enriched Topology Learning For Cross-Domain Multi-Person Pose EstimationabstractHuman pose estimation has been widely studied with much focus on supervised learning. However, in real applications, a pretrained pose estimation model usually needs be adapted to a novel domain without labels or with sparse labels. Existing domain adaptation methods cannot well deal with it since poses have flexible topological structures and need fine-grained local features. Aiming at the characteristics of human pose, we propose a novel domain adaptation method for multi-person pose estimation (MPPE) to alleviate the human-level shift. Firstly, the training samples of human poses are clustered into groups according to the posture similarity. Within the clustered space, we conduct three adaptation modules: Cross-Attentive Feature Alignment (CAFA), Intra-domain Structure Adaptation (ISA) and Adaptive Human-Topology Adaptation (AHTA). The CAFA adopts a bidirectional spatial attention mechanism to explore fine-grained local feature correlation between two humans, and thus to adaptively aggregate consistent features for adaptation. ISA only works in semi-supervised domain adaptation (SSDA) to exploit semantic relationship of corresponding keypoints for reducing the intra-domain bias. Importantly, we creatively propose an AHTA to enrich human topological knowledge for reducing the inter-domain discrepancy. Specifically, the pose structure and the cross-instance topological relations are modeled via graph networks. This flexible topology learning benefits the occluded or extreme pose inference. Extensive experiments are conducted on two popular benchmarks and additional two challenging datasets. Results demonstrate the competency of our method, which works in unsupervised or semi-supervised modes, compared with the existing supervised approaches. Xixia Xu, Qi Zou 0001, Xue Lin 0003 |
IEEE Trans. Multim. | 2 |
| 2023 | Effects of Motion-Relevant Knowledge From Unlabeled Video to Human-Object Interaction DetectionabstractThe existing works on human-object interaction (HOI) detection usually rely on expensive large-scale labeled image datasets. However, in real scenes, labeled data may be insufficient, and some rare HOI categories have few samples. This poses great challenges for deep-learning-based HOI detection models. Existing works tackle it by introducing compositional learning or word embedding but still need large-scale labeled data or extremely rely on the well-learned knowledge. In contrast, the freely available unlabeled videos contain rich motion-relevant information that can help infer rare HOIs. In this article, we creatively propose a multitask learning (MTL) perspective to assist in HOI detection with the aid of motion-relevant knowledge learning on unlabeled videos. Specifically, we design the appearance reconstruction loss (ARL) and sequential motion mining module in a self-supervised manner to learn more generalizable motion representations for promoting the detection of rare HOIs. Moreover, to better transfer motion-related knowledge from unlabeled videos to HOI images, a domain discriminator is introduced to decrease the domain gap between two domains. Extensive experiments on the HICO-DET dataset with rare categories and the V-COCO dataset with minimum supervision demonstrate the effectiveness of motion-aware knowledge implied in unlabeled videos for HOI detection. Xue Lin 0003, Qi Zou 0001, Xixia Xu, Ding Ding 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Adaptive Hypergraph Neural Network for Multi-Person Pose EstimationabstractThis paper proposes a novel two-stage hypergraph-based framework, dubbed ADaptive Hypergraph Neural Network (AD-HNN) to estimate multiple human poses from a single image, with a keypoint localization network and an Adaptive-Pose Hypergraph Neural Network (AP-HNN) added onto the former network. For providing better guided representations of AP-HNN, we employ a Semantic Interaction Convolution (SIC) module within the initial localization network to acquire more explicit predictions. Build upon this, we design a novel adaptive hypergraph to represent a human body for capturing high-order semantic relations among different joints. Notably, it can adaptively adjust the relations between joints and seek the most reasonable structure for the variable poses to benefit the keypoint localization. These two stages are combined to be trained in an end-to-end fashion. Unlike traditional Graph Convolutional Networks (GCNs) that are based on a fixed tree structure, AP-HNN can deal with ambiguity in human pose estimation. Experimental results demonstrate that the AD-HNN achieves state-of-the-art performance both on the MS-COCO, MPII and CrowdPose datasets. Xixia Xu, Qi Zou 0001, Xue Lin 0003 |
AAAI | 2 |
| 2022 | Location-Free Human Pose EstimationabstractHuman pose estimation (HPE) usually requires large-scale training data to reach high performance. However, it is rather time-consuming to collect high-quality and fine-grained annotations for human body. To alleviate this issue, we revisit HPE and propose a location-free framework without supervision of keypoint locations. We reformulate the regression-based HPE from the perspective of classification. Inspired by the CAM-based weakly-supervised object localization, we observe that the coarse keypoint locations can be acquired through the part-aware CAMs but unsatisfactory due to the gap between the fine-grained HPE and the object-level localization. To this end, we propose a customized transformer framework to mine the fine-grained representation of human context, equipped with the structural relation to capture subtle differences among keypoints. Concretely, we design a Multi-scale Spatial-guided Context Encoder to fully capture the global human context while focusing on the part-aware regions and a Relation-encoded Pose Prototype Generation module to encode the structural relations. All these works together for strengthening the weak supervision from image-level category labels on locations. Our model achieves competitive performance on three datasets when only supervised at a category-level and importantly, it can achieve comparable results with fully-supervised methods with only 25% location labels on MS-COCO and MPII. Xixia Xu, Yingguo Gao, Xue Lin 0003, Qi Zou 0001 |
CVPR | 5 |
| 2022 | Dynamic pedestrian trajectory forecasting with LSTM-based Delaunay triangulation
Qiulin Ma, Qi Zou 0001 |
Appl. Intell. | 2 |
| 2022 | CFENet: Content-aware feature enhancement network for multi-person pose estimation
Xixia Xu, Qi Zou 0001, Xue Lin 0003 |
Appl. Intell. | 2 |
| 2022 | A Stronger Baseline for Seismic Facies Classification With Less DataabstractWith the great success of deep learning in computer vision, the application of convolution neural network (CNN) in seismic facies classification is growing rapidly. However, most of the previous works based on pure state-of-the-art CNN architectures still suffer from coarse segmentation results. In this article, we study the challenges of seismic facies classification and propose a stronger baseline. More specifically, we propose a simple yet effective unsupervised approach named spatial pyramid sampling (SPS) to choose representative samples for training to reduce the labeling costs. Next, we propose a multimodal fusion (M2F) module to extract and fuse the edge and frequency information from selected seismic images to build a stable multimodal representation. Finally, we propose a local-to-global (L2G) module, which improves the recognition power by capturing the local relationship between pixels and enhancing the global context representation. Experimental results demonstrate that the proposed method achieves a superior performance with less labeled training data, especially for small categories. Qi Zou 0001, Xixia Xu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | An online multiple object tracker based on structure keeper net
Qi Zou 0001, Qiulin Ma, Haitao Lou, Huiyong Liu |
Appl. Intell. | 2 |
| 2021 | Video sketch: A middle-level representation for action recognition
Xingyuan Zhang, Yang Mi, Yanting Pei, Qi Zou 0001, Song Wang 0002 |
Appl. Intell. | 5 |
| 2021 | Effects of Image Degradation and Degradation Removal to CNN-Based Image ClassificationabstractJust like many other topics in computer vision, image classification has achieved significant progress recently by using deep learning neural networks, especially the Convolutional Neural Networks (CNNs). Most of the existing works focused on classifying very clear natural images, evidenced by the widely used image databases, such as Caltech-256, PASCAL VOCs, and ImageNet. However, in many real applications, the acquired images may contain certain degradations that lead to various kinds of blurring, noise, and distortions. One important and interesting problem is the effect of such degradations to the performance of CNN-based image classification and whether degradation removal helps CNN-based image classification. More specifically, we wonder whether image classification performance drops with each kind of degradation, whether this drop can be avoided by including degraded images into training, and whether existing computer vision algorithms that attempt to remove such degradations can help improve the image classification performance. In this article, we empirically study those problems for nine kinds of degraded images-hazy images, motion-blurred images, fish-eye images, underwater images, low resolution images, salt-and-peppered images, images with white Gaussian noise, Gaussian-blurred images, and out-of-focus images. We expect this article can draw more interests from the community to study the classification of degraded images. Yanting Pei, Qi Zou 0001, Xingyuan Zhang, Song Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Motion-Aware Feature Enhancement Network for Video PredictionabstractVideo prediction is challenging, due to the pixel-level precision requirement and the difficulty in capturing scene dynamics. Most approaches tackle the problems by pixel-level reconstruction objectives and two decomposed branches, which still suffer from blurry generations or dramatic degradations in long-term prediction. In this paper, we propose a Motion-Aware Feature Enhancement (MAFE) network for video prediction to produce realistic future frames and achieve relatively long-term predictions. First, a Channel-wise and Spatial Attention (CSA) module is designed to extract motion-aware features, which enhances the contribution of important motion details during encoding, and subsequently improves the discriminability of attention map for the frame refinement. Second, a Motion Perceptual Loss (MPL) is proposed to guide the learning of temporal cues, which benefits to robust long-term video prediction. Extensive experiments on three human activity video datasets: KTH, Human3.6M, and PennAction demonstrate the effectiveness of the proposed video prediction model compared with the state-of-the-art approaches. Xue Lin 0003, Qi Zou 0001, Xixia Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Action-Guided Attention Mining and Relation Reasoning Network for Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection is important to understand human-centric scenes and is challenging due to subtle difference between fine-grained actions, and multiple co-occurring interactions. Most approaches tackle the problems by considering the multi-stream information and even introducing extra knowledge, which suffer from a huge combination space and the non-interactive pair domination problem. In this paper, we propose an Action-Guided attention mining and Relation Reasoning (AGRR) network to solve the problems. Relation reasoning on human-object pairs is performed by exploiting contextual compatibility consistency among pairs to filter out the non-interactive combinations. To better discriminate the subtle difference between fine-grained actions, an action-aware attention based on class activation map is proposed to mine the most relevant features for recognizing HOIs. Extensive experiments on V-COCO and HICO-DET datasets demonstrate the effectiveness of the proposed model compared with the state-of-the-art approaches. Xue Lin 0003, Qi Zou 0001, Xixia Xu |
IJCAI | 2 |
| 2020 | Alleviating Human-level Shift: A Robust Domain Adaptation Method for Multi-person Pose EstimationabstractHuman pose estimation has been widely studied with much focus on supervised learning requiring sufficient annotations. However, in real applications, a pretrained pose estimation model usually need be adapted to a novel domain with no labels or sparse labels. Such domain adaptation for 2D pose estimation hasn't been explored. The main reason is that a pose, by nature, has typical topological structure and needs fine-grained features in local keypoints. While existing adaptation methods do not consider topological structure of object-of-interest and they align the whole images coarsely. Therefore, we propose a novel domain adaptation method for multi-person pose estimation to conduct the human-level topological structure alignment and fine-grained feature alignment. Our method consists of three modules: Cross-Attentive Feature Alignment (CAFA), Intra-domain Structure Adaptation (ISA) and Inter-domain Human-Topology Alignment (IHTA) module. The CAFA adopts a bidirectional spatial attention module (BSAM) that focuses on fine-grained local feature correlation between two humans to adaptively aggregate consistent features for adaptation. We adopt ISA only in semi-supervised domain adaptation (SSDA) to exploit the corresponding keypoint semantic relationship for reducing the intra-domain bias. Most importantly, we propose an IHTA to learn more domain-invariant human topological representation for reducing the inter-domain discrepancy. We model the human topological structure via the graph convolution network (GCN), by passing messages on which, high-order relations can be considered. This structure preserving alignment based on GCN is beneficial to the occluded or extreme pose inference. Extensive experiments are conducted on two popular benchmarks and results demonstrate the competency of our method compared with existing supervised approaches. Xixia Xu, Qi Zou 0001, Xue Lin 0003 |
ACM Multimedia | 2 |
| 2020 | Detecting dense text in natural imagesabstractMost existing text detection methods are mainly motivated by deep learning‐based object detection approaches, which may result in serious overlapping between detected text lines, especially in dense text scenarios. It is because text boxes are not commonly overlapped, as different from general objects in natural scenes. Moreover, text detection requires higher localisation accuracy than object detection. To tackle these problems, the authors propose a novel dense text detection network (DTDN) to localise tighter text lines without overlapping. Their main novelties are: (i) propose an intersection‐over‐union overlap loss, which considers correlations between one anchor and GT boxes and measures how many text areas one anchor contains, (ii) propose a novel anchor sample selection strategy, named CMax‐OMin, to select tighter positive samples for training. CMax‐OMin strategy not only considers whether an anchor has the largest overlap with its corresponding GT box (CMax), but also ensures the overlapping between one anchor and other GT boxes as little as possible (OMin). Besides, they train a bounding‐box regressor as post‐processing to further improve text localisation performance. Experiments on scene text benchmark datasets and their proposed dense text dataset demonstrate that the proposed DTDN achieves competitive performance, especially for dense text scenarios. Dianzhuan Jiang, Shengsheng Zhang, Qi Zou 0001, Xingyuan Zhang, Mengyang Pu |
IET Comput. Vis. | 4 |
| 2020 | Looking ahead: Joint small group detection and tracking in crowd scenes
Qiulin Ma, Qi Zou 0001, Qingji Guan, Yanting Pei |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | A Hybrid convolutional neural network for sketch recognition
Xingyuan Zhang, Qi Zou 0001, Yanting Pei, Runsheng Zhang, Song Wang 0002 |
Pattern Recognit. Lett. | 3 |
| 2020 | Integral Knowledge Distillation for Multi-Person Pose EstimationabstractBoth accuracy and efficiency are of equal importance to the human pose estimation. Most of the existing methods simply pursue excellent performance, sacrificing massive computing resources and memory. Out of this consideration, we present a novel compact and lightweight framework to train more efficient estimators using knowledge distillation. Three distillation mechanisms are proposed in our method from different perspectives, including logit distillation, feature distillation and structure distillation. Concretely, the logit distillation regards the output of teacher model as soft target to stimulate the student model. The feature distillation distills the high-level features of the teacher model to assist the student. Unlike the above strategies, the structure distillation considers the problem in a global view, aiming at ensuring the student prediction contains quite abundant structure knowledge like the teacher. We empirically demonstrate the effectiveness and efficiency of our methods on two multi-person pose estimation datasets (COCO and MPII). Specifically, our model can achieve competitive performance with the most state-of-the-art methods and consume only 35% model parameters and GFLOPs of our baseline (SimpleBaseline-ResNet-50) on the COCO dataset. Xixia Xu, Qi Zou 0001, Xue Lin 0003 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Object Discovery From a Single Unlabeled Image by Mining Frequent Itemsets With Multi-Scale FeaturesabstractThe goal of our work is to discover dominant objects in a very general setting where only a single unlabeled image is given. This is far more challenge than typical colocalization or weakly-supervised localization tasks. To tackle this problem, we propose a simple but effective pattern mining-based method, called Object Location Mining (OLM), which exploits the advantages of data mining and feature representation of pretrained convolutional neural networks (CNNs). Specifically, we first convert the feature maps from a pre-trained CNN model into a set of transactions, and then discovers frequent patterns from transaction database through pattern mining techniques. We observe that those discovered patterns, i.e., co-occurrence highlighted regions, typically hold appearance and spatial consistency. Motivated by this observation, we can easily discover and localize possible objects by merging relevant meaningful patterns. Extensive experiments on a variety of benchmarks demonstrate that OLM achieves competitive localization performance compared with the state-of-the-art methods. We also evaluate our approach compared with unsupervised saliency detection methods and achieves competitive results on seven benchmark datasets. Moreover, we conduct experiments on finegrained classification to show that our proposed method can locate the entire object and parts accurately, which can benefit to improving the classification results significantly. Runsheng Zhang, Mengyang Pu, Qingji Guan, Qi Zou 0001, Haibin Ling |
IEEE Trans. Image Process. | 6 |
| 2019 | Learning representative features via constrictive annular loss for image classification
Qi Zou 0001 |
Appl. Intell. | 3 |
| 2018 | Does Haze Removal Help CNN-Based Image Classification?
Yanting Pei, Qi Zou 0001, Song Wang 0002 |
ECCV (10) | 3 |
| 2018 | GraphNet: Learning Image Pseudo Annotations for Weakly-Supervised Semantic SegmentationabstractWeakly-supervised semantic image segmentation suffers from lacking accurate pixel-level annotations. In this paper, we propose a novel graph convolutional network-based method, called GraphNet, to learn pixel-wise labels from weak annotations. Firstly, we construct a graph on the superpixels of a training image by combining the low-level spatial relation and high-level semantic content. Meanwhile, scribble or bounding box annotations are embedded into the graph, respectively. Then, GraphNet takes the graph as input and learns to predict high-confidence pseudo image masks by a convolutional network operating directly on graphs. At last, a segmentation network is trained supervised by these pseudo image masks. We comprehensively conduct experiments on the PASCAL VOC 2012 and PASCAL-CONTEXT segmentation benchmarks. Experimental results demonstrate that GraphNet is effective to predict the pixel labels with scribble or bounding box annotations. The proposed framework yields state-of-the-art results in the community. Mengyang Pu, Qingji Guan, Qi Zou 0001 |
ACM Multimedia | 4 |
| 2018 | Online Multiple Person Tracking Using Fully-Convolutional Neural Networks and Motion Invariance Constraints
Qi Zou 0001, Qiulin Ma |
PRCV (4) | 2 |
| 2018 | Joint Headlight Pairing and Vehicle Tracking by Weighted Set Packing in Nighttime Traffic VideosabstractWe propose a set packing (SP) framework for joint headlight pairing and vehicle tracking. Given headlight detections, traditional nighttime vehicle tracking methods usually first pair headlights and then track these pairs. However, the poor photometric condition often introduces tremendous noises in headlight detection and pairing, which leads to unrecoverable errors for vehicle tracking. To overcome the challenge, we propose to jointly model these two tasks in a weighted SP framework. Specifically, a graph is built which takes candidate pair track hypotheses as nodes and encodes in edges both the disjoint constraints for tracking and the no-sharing-headlight constraints for pairing. Solving a weighted SP problem on such a graph produces vehicle trajectories, and facilitates pairing with temporal context and in turn produces high quality vehicle trajectories. The solution, however, raises the issue of unmanageable graph scale since the number of track hypotheses grows exponentially over time. To address this issue, pruning strategies are developed to solve the joint model efficiently. The proposed system is evaluated on two traffic data sets, including videos under various challenging conditions. Both quantitative and qualitative results show that our method outperforms other tested methods, both in nighttime vehicle tracking and in multi-target tracking, confirming the benefits of jointly modeling the two tasks. Qi Zou 0001, Haibin Ling |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2015 | Robust Nighttime Vehicle Detection by Tracking and Grouping HeadlightsabstractNighttime traffic surveillance is difficult due to insufficient and unstable appearance information and strong background interference. We present in this paper a robust nighttime vehicle detection system by detecting, tracking, and grouping headlights. First, we train AdaBoost classifiers for headlights detection to reduce false alarms caused by reflections. Second, to take full advantage of the complementary nature of grouping and tracking, we alternately optimize grouping and tracking. For grouping, motion features produced by tracking are used by headlights pairing. We use a maximal independent set framework for effective pairing, which is more robust than traditional pairing-by-rules methods. For tracking, context information provided by pairing is employed by multiple object tracking. The experiments on challenging datasets and quantitative evaluation show promising performance of our method. Qi Zou 0001, Haibin Ling, Siwei Luo |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2011 | Nonlinear dimensionality reduction using a temporal coherence principle
Jiali Zhao, Yunhui Liu 0005, Siwei Luo, Qi Zou 0001 |
Inf. Sci. | 5 |
| 2009 | Slow Feature Discriminant Analysis and its application on handwritten digit recognitionabstractSlow Feature Analysis (SFA) is an unsupervised algorithm by extracting the slowly varying features from time series and has been used to pattern recognition successfully. Based on SFA, this paper develops a new algorithm, Slow Feature Discriminant Analysis (SFDA), which can maximize the temporal variation of between-class time series, and minimize the temporal variation of within-class time series simultaneously. Due to adoption of discrimination power, the performance on pattern recognition is improved compared to SFA. The experiments results on MNIST digit handwritten database also show that the proposed algorithm is in particular attractive. Jiali Zhao, Qi Zou 0001, Siwei Luo |
IJCNN | 4 |
| 2008 | Face recognition with Neighboring Discriminant AnalysisabstractThe paper presents a dimensionality reduction method called Neighboring Discriminant Analysis (NDA) and its kernel extension to improve face recognition performance. We take into account both the data distribution and class label information. We describe the connection of two data as neighboring or non-neighboring, together with whether the pair are from the same class or belong to different classes by utilizing the Graph Embedding framework as a tool. The compactness graph is constructed by connecting each data vertex with its neighboring data of the same class, while the penalty graph connects the rest data pairs, i.e. the data pairs which are not from the same class or are non-neighboring. NDA algorithm can map the original high dimensional space to a reduced low dimensional space, which compact the neighboring data from the same class and simultaneously separate the data far away from each other or belong to different classes. Real face recognition experiment shows NDA and its kernel extension outperforms LDA etc. Jiali Zhao, Siwei Luo, Qi Zou 0001 |
FG | 5 |
| 2008 | Contour grouping with shape manifold and distance transformabstractObject detection in clutter or occlusion is a hard problem in computer vision. We propose an object detection method based on contour grouping. Two stages are included: a novel distance transform is applied to match templates to the test image so that candidates and locations of the object are obtained; verification using shape manifold is performed to preclude outliers and identify the prior. We use the prior combined with bottom-up edge information to produce the final grouping result. Our contribution lies in two aspects: one is the novel distance transform saves much searching space; the other is introducing shape manifold in verifying candidates of grouping. Experiments show our method achieves considerable accuracy in occlusion and background clutter. Specially, the only feature used is edge and contour rather than combination of multi features. Qi Zou 0001, Siwei Luo |
ICPR | 1 |
| 2007 | Selective Attention Guided Contour Extraction for Perceptual GroupingabstractSelective attention is an important mechanism in human visual system. Existing attention models concentrate on giving a series of sight transferring. In this paper, we extend selective attention to guide contour extraction prepared for perceptual grouping. Selective attention functions in two aspects. One is to compute each pixel's salient value in an image. The other is to extract salient contours. Compared with other contour extraction algorithms, our model gets better result when the same edge detect cue is used. Experiments and quantitative analysis testify our model's good performance contour extraction and contour grouping. Jingjing Zhong, Siwei Luo, Qi Zou 0001 |
ICTAI (1) | 3 |
| 2007 | Biological Inspired Global Descriptor for Shape Matching
Siwei Luo, Qi Zou 0001 |
ISNN (2) | 3 |