VLDB 2026 Research / reviewers in the wild / expert
Qiong Cao
dblp:22/7733
· DBLP profile ↗
34ranked-venue papers
8as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fault-Tolerant Cooperative Formation Control for Heterogeneous Ships: Application to the Obstacle Avoidance ManeuveringabstractThis study presents an adaptive quantized control algorithm designed to address actuator faults and achieve obstacle avoidance for heterogeneous ships. Within this algorithm, a hysteresis quantizer is employed to minimize communication resource utilization. Novel adaptive compensation mechanism is introduced to counteract actuator faults. Additionally, the radial basis function neural networks (RBF-NNs) are utilized to approximate the uncertainties inherent in the ship model. Further adaptive parameters are incorporated to mitigate perturbation errors arising from mismatches between the quantizer and the ships’ unknown parameters. Stability is rigorously established through the construction of a direct Lyapunov function, demonstrating that all signals within the closed-loop system satisfy Semi-Globally Uniformly Ultimately Bounded (SGUUB). To validate the superiority of the proposed algorithm, simulation experiments are conducted for obstacle avoidance mission of marine heterogeneous system. Guoqing Zhang 0004, Zhu Sun 0004, Jiqiang Li, Qiong Cao, Weidong Zhang 0004 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and SegmentationabstractReferring multi-object tracking (RMOT) is an emerging cross-modal task that aims to locate an arbitrary number of target objects and maintain their identities referred by a language expression in a video. This intricate task involves the reasoning of linguistic and visual modalities, along with the temporal association of target objects. However, the seminal work relies on loose feature fusion and neglects long-term information. In this study, we introduce a compact Transformer-based method, termed TenRMOT. We conduct feature fusion at both encoding and decoding stages to fully exploit the advantages of Transformer architecture. Specifically, we incrementally perform cross-modal fusion layer-by-layer during the encoding phase. In the decoding phase, we utilize language-guided queries to probe memory features for accurate prediction of the desired objects. Moreover, we introduce a query update module that explicitly leverages temporal prior information of the tracked objects to enhance the consistency of their trajectories. In addition, we introduce a novel task called Referring Multi-Object Tracking and Segmentation (RMOTS) and construct a new dataset named Ref-KITTI Segmentation. Our dataset consists of 18 videos with 818 expressions, and each expression averages 10.7 masks, which poses a greater challenge compared to the typical single mask in most existing referring video segmentation datasets. TenRMOT demonstrates superior performance on both the referring multi-object tracking and the segmentation tasks. Changcheng Xiao, Qiong Cao, Xiang Zhang 0008, Tao Wang 0006, Canqun Yang, Long Lan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-EvolutionabstractHuman preference alignment can significantly enhance the capabilities of Multimodal Large Language Models (MLLMs). However, collecting high-quality preference data remains costly. One promising solution is the self-evolution strategy, where models are iteratively trained on data they generate. Current multimodal self-evolution techniques, nevertheless, still need human- or GPT-annotated data. Some methods even require extra models or ground truth answers to construct preference data. To overcome these limitations, we propose a novel multimodal self-evolution framework that empowers the model to autonomously generate high-quality questions and answers using only unannotated images. First, in the question generation phase, we implement an image-driven self-questioning mechanism. This approach allows the model to create questions and evaluate their relevance and answerability based on the image content. If a question is deemed irrelevant or unanswerable, the model regenerates it to ensure alignment with the image. This process establishes a solid foundation for subsequent answer generation and optimization. Second, while generating answers, we design an answer self-enhancement technique to boost the discriminative power of answers. We begin by captioning the images and then use the descriptions to enhance the generated answers. Additionally, we utilize corrupted images to generate rejected answers, thereby forming distinct preference pairs for effective optimization. Finally, in the optimization step, we incorporate an image content alignment loss function alongside the Direct Preference Optimization (DPO) loss to mitigate hallucinations. This function maximizes the likelihood of the above generated descriptions in order to constrain the model's attention to the image content. As a result, model can generate more accurate and reliable outputs. Experiments demonstrate that our framework is competitively compared with previous methods that utilize external information, paving the way for more efficient and scalable MLLMs. Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue 0003, Changxing Ding |
AAAI | 2 |
| 2025 | Individual differences in habituation predict dishabituation magnitude in adults and infants
Anjie Cao, Qiong Cao, Michael C. Frank, Shari Liu |
CogSci | 2 |
| 2025 | Surprise isn't symmetrical: Adults' looking suggests non-perceptual considerations during dishabituation
Qiong Cao, Anjie Cao, Gal Raz, Josh Tenenbaum, Shari Liu |
CogSci | 1 |
| 2025 | Goldilocks Pattern of Learning after Observing Unexpected Physical Events
Qiong Cao, Lisa Feigenson |
CogSci | 1 |
| 2024 | Towards Variable and Coordinated Holistic Co-Speech Motion GenerationabstractThis paper addresses the problem of generating lifelike holistic co-speech motions for 3D avatars, focusing on two key aspects: variability and coordination. Variability allows the avatar to exhibit a wide range of motions even with similar speech content, while coordination ensures a harmonious alignment among facial expressions, hand gestures, and body poses. We aim to achieve both with ProbTalk, a unified probabilistic framework designed to jointly model facial, hand, and body movements in speech. ProbTalk builds on the variational autoencoder (VAE) architecture and incorporates three core designs. First, we introduce product quantization (PQ) to the VAE, which enriches the representation of complex holistic motion. Second, we devise a novel non-autoregressive model that embeds 2D positional encoding into the product-quantized representation, thereby preserving essential structure information of the PQ codes. Last, we employ a secondary stage to refine the preliminary prediction, further sharpening the high-frequency details. Coupling these three designs enables ProbTalk to generate natural and diverse holistic co-speech motions, outperforming several state-of-the-art methods in qualitative and quantitative evaluations, particularly in terms of realism. Our code and model will be released for research purposes at https://feifeifeiliu.github.io/probtalk/. Qiong Cao, Yandong Wen, Huaiguang Jiang, Changxing Ding |
CVPR | 2 |
| 2024 | MuEP: A Multimodal Benchmark for Embodied Planning with Foundation Models
Kanxue Li, Baosheng Yu, Yibing Zhan, Qiong Cao, Li Shen 0008, Lusong Li, Dapeng Tao, Xiaodong He 0001 |
IJCAI | 10 |
| 2024 | MambaTrack: A Simple Baseline for Multiple Object Tracking with State Space ModelabstractTracking by detection has been the prevailing paradigm in the field of Multi-object Tracking (MOT). These methods typically rely on the Kalman Filter to estimate the future locations of objects, assuming linear object motion. However, they fall short when tracking objects exhibiting nonlinear and diverse motion in scenarios like dancing and sports. In addition, there has been limited focus on utilizing learning-based motion predictors in MOT. To address these challenges, we resort to exploring data-driven motion prediction methods. Inspired by the great expectation of state space models (SSMs), such as Mamba, in long-term sequence modeling with near-linear complexity, we introduce a Mamba-based motion model named Mamba moTion Predictor (MTP). MTP is designed to model the complex motion patterns of objects like dancers and athletes. Specifically, MTP takes the spatial-temporal location dynamics of objects as input, captures the motion pattern using a bi-Mamba encoding layer, and predicts the next motion. In real-world scenarios, objects may be missed due to occlusion or motion blur, leading to premature termination of their trajectories. To tackle this challenge, we further expand the application of MTP. We employ it in an autoregressive way to compensate for missing observations by utilizing its own predictions as inputs, thereby contributing to more consistent trajectories. Our proposed tracker, MambaTrack, demonstrates advanced performance on benchmarks such as Dancetrack and SportsMOT, which are characterized by complex motion and severe occlusion. Changcheng Xiao, Qiong Cao, Zhigang Luo, Long Lan |
ACM Multimedia | 2 |
| 2024 | MotionTrack: Learning motion predictor for multiple object tracking
Changcheng Xiao, Qiong Cao, Long Lan, Xiang Zhang 0008, Zhigang Luo, Dacheng Tao |
Neural Networks | 2 |
| 2024 | A Survey on Self-Supervised Learning: Algorithms, Applications, and Future TrendsabstractDeep supervised learning algorithms typically require a large volume of labeled data to achieve satisfactory performance. However, the process of collecting and labeling such data can be expensive and time-consuming. Self-supervised learning (SSL), a subset of unsupervised learning, aims to learn discriminative features from unlabeled data without relying on human-annotated labels. SSL has garnered significant attention recently, leading to the development of numerous related algorithms. However, there is a dearth of comprehensive studies that elucidate the connections and evolution of different SSL variants. This paper presents a review of diverse SSL methods, encompassing algorithmic aspects, application domains, three key trends, and open research questions. First, we provide a detailed introduction to the motivations behind most SSL algorithms and compare their commonalities and differences. Second, we explore representative applications of SSL in domains such as image processing, computer vision, and natural language processing. Lastly, we discuss the three primary trends observed in SSL research and highlight the open questions that remain. Jie Gui, Tuo Chen, Jing Zhang 0037, Qiong Cao, Zhenan Sun, Hao Luo 0004, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Few-Shot Learning With Dynamic Graph Structure PreservingabstractIn recent years, few-shot learning has received increasing attention in the Internet of Things areas. Few-shot learning aims to distinguish unseen classes with a few labeled samples from each class. Most recently transductive few-shot studies highly rely on the static geometry distributions generated on the feature space during the label propagation process between unseen class instances. However, these recent methods fail to guarantee that the generated graph structure preserves the true distributions between data properly. In this article, we propose a novel dynamic graph structure preserving (DGSP) model for few-shot learning. Specifically, we formulate the objective function of DGSP by simultaneously considering the data correlations from the feature space and the label space to update the generated graph structure, which can reasonably revise the inappropriate or mistaken local geometry relationships. Then, we design an efficient alternating optimization algorithm to jointly learn the label prediction matrix and the optimal graph structure, the latter of which can be formulated as a linear programming problem. Moreover, our proposed DGSP can be easily combined with any backbone networks during the learning process. We conduct extensive experimental results across different benchmarks, backbones, and task settings, and our method achieves state-of-the-art performance compared with methods based on transductive few-shot learning. Sichao Fu, Qiong Cao, Yunwen Lei, Yibing Zhan, Xinge You |
IEEE Trans. Ind. Informatics | 2 |
| 2023 | TriDet: Temporal Action Detection with Relative Boundary ModelingabstractIn this paper, we present a one-stage framework TriDet for temporal action detection. Existing methods often suffer from imprecise boundary predictions due to the ambiguous action boundaries in videos. To alleviate this problem, we propose a novel Trident-head to model the action boundary via an estimated relative probability distribution around the boundary. In the feature pyramid of TriDet, we propose an efficient Scalable-Granularity Perception (SGP) layer to mitigate the rank loss problem of self-attention that takes place in the video features and aggregate information across different temporal granularities. Benefiting from the Trident-head and the SGP-based feature pyramid, TriDet achieves state-of-the-art performance on three challenging benchmarks: THUMOS14, HACS and EPIC-KITCHEN 100, with lower computational costs, compared to previous methods. For example, TriDet hits an average mAP of 69.3% on THUMOS14, outperforming the previous best by 2.5%, but with only 74.6% of its latency. The code is released to https://github.com/dingfengshi/TriDet. Dingfeng Shi, Qiong Cao, Lin Ma 0002, Jia Li 0003, Dacheng Tao |
CVPR | 3 |
| 2023 | Generating Holistic 3D Human Motion from SpeechabstractThis work addresses the problem of generating 3D holistic body motions from human speech. Given a speech recording, we synthesize sequences of 3D body poses, hand gestures, and facial expressions that are realistic and diverse. To achieve this, we first build a high-quality dataset of 3D holistic body meshes with synchronous speech. We then define a novel speech-to-motion generation framework in which the face, body, and hands are modeled separately. The separated modeling stems from the fact that face articulation strongly correlates with human speech, while body poses and hand gestures are less correlated. Specifically, we employ an autoencoder for face motions, and a compositional vector-quantized variational autoencoder (VQ- VAE) for the body and hand motions. The compositional VQ-VAE is key to generating diverse results. Additionally, we propose a cross-conditional autoregressive model that generates body poses and hand gestures, leading to coherent and realistic motions. Extensive experiments and user studies demonstrate that our proposed approach achieves state-of-the-art performance both qualitatively and quantitatively. Our dataset and code are released for research purposes at https://talkshow.is.tue.mpg.de/. Hongwei Yi, Hualin Liang, Qiong Cao, Yandong Wen, Timo Bolkart, Dacheng Tao, Michael J. Black |
CVPR | 4 |
| 2023 | Sharper Bounds for Uniformly Stable Algorithms with Stationary Mixing Process
Shi Fu, Yunwen Lei, Qiong Cao, Xinmei Tian 0001, Dacheng Tao |
ICLR | 3 |
| 2023 | GraMMaR: Ground-aware Motion Model for 3D Human Motion ReconstructionabstractDemystifying complex human-ground interactions is essential for accurate and realistic 3D human motion reconstruction from RGB videos, as it ensures consistency between the humans and the ground plane. Prior methods have modeled human-ground interactions either implicitly or in a sparse manner, often resulting in unrealistic and incorrect motions when faced with noise and uncertainty. In contrast, our approach explicitly represents these interactions in a dense and continuous manner. To this end, we propose a novel Ground-aware Motion Model for 3D Human Motion Reconstruction, named GraMMaR, which jointly learns the distribution of transitions in both pose and interaction between every joint and ground plane at each time step of a motion sequence. It is trained to explicitly promote consistency between the motion and distance change towards the ground. After training, we establish a joint optimization strategy that utilizes GraMMaR as a dual-prior, regularizing the optimization towards the space of plausible ground-aware motions. This leads to realistic and coherent motion reconstruction, irrespective of the assumed or learned ground plane. Through extensive evaluation on the AMASS and AIST++ datasets, our model demonstrates good generalization and discriminating abilities in challenging cases including complex and ambiguous human-ground interactions. The code will be available at https://github.com/xymsh/GraMMaR. Sihan Ma, Qiong Cao, Hongwei Yi, Jing Zhang 0037, Dacheng Tao |
ACM Multimedia | 2 |
| 2022 | DearKD: Data-Efficient Early Knowledge Distillation for Vision TransformersabstractTransformers are successfully applied to computer vision due to their powerful modeling capacity with self-attention. However, the excellent performance of transformers heavily depends on enormous training images. Thus, a data-efficient transformer solution is urgently needed. In this work, we propose an early knowledge distillation framework, which is termed as DearKD, to improve the data efficiency required by transformers. Our DearKD is a two-stage framework that first distills the inductive biases from the early intermediate layers of a CNN and then gives the transformer full play by training without distillation. Further, our DearKD can be readily applied to the extreme data-free case where no real images are available. In this case, we propose a boundary-preserving intra-divergence loss based on DeepInversion to further close the performance gap against the full-data counterpart. Extensive experiments on ImageNet, partial ImageNet, data-free setting and other downstream tasks prove the superiority of DearKD over its baselines and state-of-the-art methods. Xianing Chen, Qiong Cao, Jing Zhang 0037, Shenghua Gao, Dacheng Tao |
CVPR | 2 |
| 2022 | ReAct: Temporal Action Detection with Relational Queries
Dingfeng Shi, Qiong Cao, Jing Zhang 0037, Lin Ma 0002, Jia Li 0003, Dacheng Tao |
ECCV (10) | 3 |
| 2022 | View Vertically: A Hierarchical Network for Trajectory Prediction via Fourier Spectrums
Conghao Wong, Beihao Xia, Ziming Hong, Qinmu Peng, Wei Yuan 0001, Qiong Cao, Xinge You |
ECCV (22) | 6 |
| 2022 | Learning Sequence Representations by Non-local Recurrent Neural Memory
Wenjie Pei, Xin Feng 0005, Canmiao Fu, Qiong Cao, Guangming Lu 0002, Yu-Wing Tai |
Int. J. Comput. Vis. | 4 |
| 2021 | Push for Center Learning via Orthogonalization and Subspace Masking for Person Re-IdentificationabstractPerson re-identification aims to identify whether pairs of images belong to the same person or not. This problem is challenging due to large differences in camera views, lighting and background. One of the mainstream in learning CNN features is to design loss functions which reinforce both the class separation and intra-class compactness. In this paper, we propose a novel Orthogonal Center Learning method with Subspace Masking for person re-identification. We make the following contributions: 1) we develop a center learning module to learn the class centers by simultaneously reducing the intra-class differences and inter-class correlations by orthogonalization; 2) we introduce a subspace masking mechanism to enhance the generalization of the learned class centers; and 3) we propose to integrate the average pooling and max pooling in a regularizing manner that fully exploits their powers. Extensive experiments show that our proposed method consistently outperforms the state-of-the-art methods on large-scale ReID datasets including Market-1501, DukeMTMC-ReID, CUHK03 and MSMT17. Weinong Wang, Wenjie Pei, Qiong Cao, Shu Liu 0005, Guangming Lu 0002, Yu-Wing Tai |
IEEE Trans. Image Process. | 3 |
| 2020 | Automated Video Face Labelling for Films and TV MaterialabstractThe objective of this work is automatic labelling of characters in TV video and movies, given weak supervisory information provided by an aligned transcript. We make five contributions: (i) a new strategy for obtaining stronger supervisory information from aligned transcripts; (ii) an explicit model for classifying background characters, based on their face-tracks; (iii) employing new ConvNet based face features, and (iv) a novel approach for labelling all face tracks jointly using linear programming. Each of these contributions delivers a boost in performance, and we demonstrate this on standard benchmarks using tracks provided by authors of prior work. As a fifth contribution, we also investigate the generalisation and strength of the features and classifiers by applying them "in the raw" on new video material where no supervisory information is used. In particular, to provide high quality tracks on those material, we propose efficient track classifiers to remove false positive tracks by the face tracker. Overall we achieve a dramatic improvement over the state of the art on both TV series and film datasets, and almost saturate performance on some benchmarks. Omkar M. Parkhi, Esa Rahtu, Qiong Cao, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | MMFace: A Multi-Metric Regression Network for Unconstrained Face ReconstructionabstractWe propose to address the face reconstruction in the wild by using a multi-metric regression network, MMFace, to align a 3D face morphable model (3DMM) to an input image. The key idea is to utilize a volumetric sub-network to estimate an intermediate geometry representation, and a parametric sub-network to regress the 3DMM parameters. Our parametric sub-network consists of identity loss, expression loss, and pose loss which greatly improves the aligned geometry details by incorporating high level loss functions directly defined in the 3DMM parametric spaces. Our high-quality reconstruction is robust under large variations of expressions, poses, illumination conditions, and even with large partial occlusions. We evaluate our method by comparing the performance with state-of-the-art approaches on latest 3D face dataset LS3D-W and Florence. We achieve significant improvements both quantitatively and qualitatively. Due to our high-quality reconstruction, our method can be easily extended to generate high-quality geometry sequences for video inputs. Hongwei Yi, Chen Li 0031, Qiong Cao, Xiaoyong Shen, Sheng Li 0008, Yu-Wing Tai |
CVPR | 3 |
| 2019 | Non-Local Recurrent Neural Memory for Supervised Sequence ModelingabstractTypical methods for supervised sequence modeling are built upon the recurrent neural networks to capture temporal dependencies. One potential limitation of these methods is that they only model explicitly information interactions between adjacent time steps in a sequence, hence the high-order interactions between nonadjacent time steps are not fully exploited. It greatly limits the capability of modeling the long-range temporal dependencies since one-order interactions cannot be maintained for a long term due to information dilution and gradient vanishing. To tackle this limitation, we propose the Non-local Recurrent Neural Memory (NRNM) for supervised sequence modeling, which performs non-local operations to learn full-order interactions within a sliding temporal block and models the global interactions between blocks in a gated recurrent manner. Consequently, our model is able to capture the long-range dependencies. Besides, the latent high-level features contained in high-order interactions can be distilled by our model. We demonstrate the merits of our NRNM approach on two different tasks: action recognition and sentiment analysis. Canmiao Fu, Wenjie Pei, Qiong Cao, Chaopeng Zhang, Yong Zhao 0010, Xiaoyong Shen, Yu-Wing Tai |
ICCV | 3 |
| 2019 | New stability results for impulsive neural networks with time delays
Chao Liu 0026, Xiaoyang Liu 0001, Guangjian Zhang, Qiong Cao, Junjian Huang |
Neural Comput. Appl. | 5 |
| 2018 | VGGFace2: A Dataset for Recognising Faces across Pose and AgeabstractIn this paper, we introduce a new large-scale face dataset named VGGFace2. The dataset contains 3.31 million images of 9131 subjects, with an average of 362.6 images for each subject. Images are downloaded from Google Image Search and have large variations in pose, age, illumination, ethnicity and profession (e.g. actors, athletes, politicians). The dataset was collected with three goals in mind: (i) to have both a large number of identities and also a large number of images for each identity; (ii) to cover a large range of pose, age and ethnicity; and (iii) to minimise the label noise. We describe how the dataset was collected, in particular the automated and manual filtering stages to ensure a high accuracy for the images of each identity. To assess face recognition performance using the new dataset, we train ResNet-50 (with and without Squeeze-and-Excitation blocks) Convolutional Neural Networks on VGGFace2, on MS-Celeb-1M, and on their union, and show that training on VGGFace2 leads to improved recognition performance over pose and age. Finally, using the models trained on these datasets, we demonstrate state-of-the-art performance on the IJB-A and IJB-B face recognition benchmarks, exceeding the previous state-of-the-art by a large margin. The dataset and models are publicly available. Qiong Cao, Li Shen 0005, Weidi Xie, Omkar M. Parkhi, Andrew Zisserman |
FG | 1 |
| 2018 | Urban Land Use/Land Cover Classification Based on Feature Fusion Fusing Hyperspectral Image and Lidar DataabstractHyperspectral images have been widely used in classification because of the abundant spectral information. But it can't distinguish the objective with similar spectral character but different elevation. However, LiDAR data can obtain elevation information. Therefore, it will obtain better classification maps if fusing the two data. In recent years, CNN has attracted much attention due to its powerful ability to excavate the potential representation and features of the raw data. However, it's difficult to distinguish the objects with different spectral information but similar surface character. Unlike CNN features, the traditional manual features, such as the normalized vegetation index (NDVI), have a certain characteristic expression significance. In order to consider both the semantic information of traditional manual features and the advanced features of CNN features, this paper proposes a fusion algorithm of hyperspectral and LiDAR fusion based on feature fusion. The proposed algorithm has achieved a good fusion classification effect on the MUUFL Gulfport Hyperspectral and LiDAR Data set. Qiong Cao, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2018 | Template adaptation for face verification and identification
Nate Crosswhite, Jeffrey Byrne, Chris Stauffer, Omkar M. Parkhi, Qiong Cao, Andrew Zisserman |
Image Vis. Comput. | 5 |
| 2017 | Template Adaptation for Face Verification and Identification
Nate Crosswhite, Jeffrey Byrne, Chris Stauffer, Omkar M. Parkhi, Qiong Cao, Andrew Zisserman |
FG | 5 |
| 2016 | Generalization bounds for metric and similarity learning
Qiong Cao, Zheng-Chu Guo, Yiming Ying |
Mach. Learn. | 1 |
| 2013 | Similarity Metric Learning for Face RecognitionabstractRecently, there is a considerable amount of efforts devoted to the problem of unconstrained face verification, where the task is to predict whether pairs of images are from the same person or not. This problem is challenging and difficult due to the large variations in face images. In this paper, we develop a novel regularization framework to learn similarity metrics for unconstrained face verification. We formulate its objective function by incorporating the robustness to the large intra-personal variations and the discriminative power of novel similarity metrics. In addition, our formulation is a convex optimization problem which guarantees the existence of its global solution. Experiments show that our proposed method achieves the state-of-the-art results on the challenging Labeled Faces in the Wild (LFW) database [10]. Qiong Cao, Yiming Ying |
ICCV | 1 |
| 2012 | Graph-based dimensionality reduction for KNN-based image annotation
Rujie Liu, Fei Li 0005, Qiong Cao |
ICPR | 4 |
| 2012 | Distance Metric Learning Revisited
Qiong Cao, Yiming Ying |
ECML/PKDD (1) | 1 |
| 2010 | An automatic vehicle detection method based on traffic videosabstractA vision-based vehicle detection method is presented in this paper. The proposed method is composed of two steps, i.e., hypothesis generation and hypothesis verification. An adaptive background modeling and updating method is proposed to detect foreground regions in video sequences. With the prior knowledge of the vehicle appearance, the possible vehicle locations are extracted from the foreground regions and the touched vehicles are separated. Finally, hypothesized regions are verified by comparing their appearances with vehicle model. The performance of the proposed method is verified on videos captured under versatile conditions, and good results are achieved even in heavy traffic conditions. Qiong Cao, Rujie Liu, Fei Li 0005, Yuehong Wang |
ICIP | 1 |