Dong Yi

dblp:60/1827 · DBLP profile ↗
← Back
45ranked-venue papers
6as first author
9since 2021 · last 2026
0009-0000-1282-7310ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 30 · 6 first-author · 4 since 2021Security and privacy · 9 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AnyDesign: Versatile area fashion editing via mask-free diffusion
Yunfang Niu, Dong Yi, Lingxiang Wu, Jinqiao Wang
Neural Networks2
2026 Adversarial discriminant attack on text-to-image diffusion models
Hanxiao Wu, Shengwu Xiong 0001, Dong Yi, Lingxiang Wu, Jianqing Zhu, Guibo Zhu, Jinqiao Wang
Neural Networks3
2026 An Empirical Study of Validating Synthetic Data for Text-Based Person Retrieval
Min Cao 0005, Ziyin Zeng, Dong Yi, Jinqiao Wang, Mang Ye
IEEE Trans. Inf. Forensics Secur.4
2025 StreamWMR: A Streaming Framework for Real-time 3D Whole-body Mesh Recovery
abstract
3D whole-body mesh recovery aims to extract parameters for the human body, hands, and head from a single human image. Most applications related to human mesh recovery, such as physical fitness motion capture and operating room motion capture, necessitate real-time video stream processing. However, existing methods ignore the video processing and often require significant computational resources, making real-time performance unattainable and greatly limiting their practicality. Moreover, noticeable misalignments are often observed when concatenating them back to the body and reprojecting them onto the image. In this paper, we propose a streaming framework for whole-body mesh recovery in the video. First, we simplify pose regression by leveraging the root nodes of the hands and head to locate each component. Second, for temporal optimization, we incorporate attention mechanisms related to keypoint velocity to incorporate information from previous frames and achieve more stable and smooth motions. Finally, we propose a multi-view projection loss to eliminate the ambiguity caused by inaccurate 3D regression and pose estimation in computing reprojection errors. The combination of these enables our method to achieve real-time inference speed while maintaining accuracy and stability.
Xiangyu Zhu 0001, Jinlin Wu, Zidu Wang, Shukai Chen, Dong Yi, Zhen Lei 0001
IJCB8
2025 Dual-Chain Reasoning: Enhancing Multimodal Document VQA Through Positive and Negative Reasoning Paths
Hanxiao Wu, Zhaopeng Gu, Dong Yi, Guibo Zhu, Jinqiao Wang
ICIG (3)4
2025 Enhancing Visual Aligning and Grounding for Aerial Vision-and-Dialog Navigation
abstract
Vision-and-Language Navigation tasks require an agent to navigate to a destination following natural language instructions. We focus on a challenging VLN dataset, Aerial Vision-and-Dialog Navigation, which encompasses a diverse array of environments and includes an additional altitude variable. Significant spatial and scale variations in the aerial agent's view make destination visual grounding a crucial capability for the navigation task. However, the existing frameworks not only have insufficient attention to the vision model, but also lack the correlations between visual and textual modalities. To address this, we propose a model that aligns destination visual images with navigation instructions, featuring three innovative components. Firstly, we propose a multi-stage pre-training pipeline that enhances the model's ability to associate language instructions with top-view images of destinations. Secondly, trajectories are augmented elastically to simulate the noise of the controlling process. Finally, a polygon regression loss function is introduced for rotated object detection, which significantly enhances the accuracy of altitude and orientation estimation. Experiments demonstrate the effectiveness of our approach, which achieves state-of-the-art advancements with improvements of 2.9% in the val unseen dataset and 3.0% in test unseen dataset in success weighted by path length.
Guanhui Qiao, Dong Yi, Lingxiang Wu, Hanxiao Wu, Jinqiao Wang
IEEE Signal Process. Lett.2
2024 PFDM: Parser-Free Virtual Try-On via Diffusion Model
abstract
Virtual try-on can significantly improve the garment shopping experiences in both online and in-store scenarios, attracting broad interest in computer vision. However, to achieve high-fidelity try-on performance, most state-of-the-art methods still rely on accurate segmentation masks, which are often produced by near-perfect parsers or manual labeling. To overcome the bottleneck, we propose a parser-free virtual try-on method based on the diffusion model (PFDM). Given two images, PFDM can "wear" garments on the target person seamlessly by implicitly warping without any other information. To learn the model effectively, we synthesize many pseudo-images and construct sample pairs by wearing various garments on persons. Supervised by the large-scale expanded dataset, we fuse the person and garment features using a proposed Garment Fusion Attention (GFA) mechanism. Experiments demonstrate that our proposed PFDM can successfully handle complex cases, synthesize high-fidelity images, and outperform both state-of-the-art parser-free and parser-based models.
Yunfang Niu, Dong Yi, Lingxiang Wu, Zhiwei Liu 0004, Pengxiang Cai, Jinqiao Wang
ICASSP2
2024 Deep evolutionary fusion neural network: a new prediction standard for infectious disease incidence rates
abstract
BACKGROUND: Previously, many methods have been used to predict the incidence trends of infectious diseases. There are numerous methods for predicting the incidence trends of infectious diseases, and they have exhibited varying degrees of success. However, there are a lack of prediction benchmarks that integrate linear and nonlinear methods and effectively use internet data. The aim of this paper is to develop a prediction model of the incidence rate of infectious diseases that integrates multiple methods and multisource data, realizing ground-breaking research. RESULTS: The infectious disease dataset is from an official release and includes four national and three regional datasets. The Baidu index platform provides internet data. We choose a single model (seasonal autoregressive integrated moving average (SARIMA), nonlinear autoregressive neural network (NAR), and long short-term memory (LSTM)) and a deep evolutionary fusion neural network (DEFNN). The DEFNN is built using the idea of neural evolution and fusion, and the DEFNN + is built using multisource data. We compare the model accuracy on reference group data and validate the model generalizability on external data. (1) The loss of SA-LSTM in the reference group dataset is 0.4919, which is significantly better than that of other single models. (2) The loss values of SA-LSTM on the national and regional external datasets are 0.9666, 1.2437, 0.2472, 0.7239, 1.4026, and 0.6868. (3) When multisource indices are added to the national dataset, the loss of the DEFNN + increases to 0.4212, 0.8218, 1.0331, and 0.8575. CONCLUSIONS: We propose an SA-LSTM optimization model with good accuracy and generalizability based on the concept of multiple methods and multiple data fusion. DEFNN enriches and supplements infectious disease prediction methodologies, can serve as a new benchmark for future infectious disease predictions and provides a reference for the prediction of the incidence rates of various infectious diseases.
Tianhua Yao, Xicheng Chen, Haojia Wang, Chengcheng Gao, Dali Yi, Zeliang Wei, Dong Yi, Yazhou Wu
BMC Bioinform.10
2023 Genetic algorithm optimised Hadamard product method for inconsistency judgement matrix adjustment in AHP and automatic analysis system development
abstract
The analytic hierarchy process (AHP) is an important method to solve the multi-objective decision-making weight problem. However, due to the subjective judgement and selection preference of decision-makers, the consistency of the judgement matrix is inevitably poor. Hence, to improve the reliability of decision-making results, it is necessary to adjust the consistency of the judgement matrix. We propose a genetic algorithm optimised Hadamard product (GAOHP) to judge the consistency of the matrix. This method converts the original matrix consistency adjustment problem into an optimisation solution problem and uses a meta-heuristic algorithm to search the global optimisation quickly. Our method, which significantly improves the computing efficiency and realises that the judgement matrix after adjustment satisfies the basic consistency, has fully retained the judgement intention of the decision-makers. It offers the dual advantages of preserving the original intention of decision-makers and greater efficiency of operation. Finally, we developed an automatic analysis system based on MATLAB app designer to realise the rapid adjustment of the consistency of judgement matrix, which is conducive to providing an analysis system with simple and stable operation for AHP analysis.
Chengcheng Gao, Xicheng Chen, Dong Yi, Yazhou Wu
Expert Syst. Appl.5
2020 Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark Study
abstract
Existing enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions.
Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin
IEEE Trans. Image Process.50
2019 Large-Scale Bisample Learning on ID Versus Spot Face Recognition
Xiangyu Zhu 0001, Zhen Lei 0001, Hailin Shi, Fan Yang 0062, Dong Yi, Guo-Jun Qi, Stan Z. Li
Int. J. Comput. Vis.6
2017 Cross-Modality Face Recognition via Heterogeneous Joint Bayesian
abstract
In many face recognition applications, the modalities of face images between the gallery and probe sets are different, which is known as heterogeneous face recognition. How to reduce the feature gap between images from different modalities is a critical issue to develop a highly accurate face recognition algorithm. Recently, joint Bayesian (JB) has demonstrated superior performance on general face recognition compared to traditional discriminant analysis methods like subspace learning. However, the original JB treats the two input samples equally and does not take into account the modality difference between them and may be suboptimal to address the heterogeneous face recognition problem. In this work, we extend the original JB by modeling the gallery and probe images using two different Gaussian distributions to propose a heterogeneous joint Bayesian (HJB) formulation for cross-modality face recognition. The proposed HJB explicitly models the modality difference of image pairs and, therefore, is able to better discriminate the same/different face pairs accurately. Extensive experiments conducted in the case of visible-near-infrared and ID photo versus spot face recognition problems show the superiority of the HJB over previous methods.
Hailin Shi, Xiaobo Wang 0001, Dong Yi, Zhen Lei 0001, Xiangyu Zhu 0001, Stan Z. Li
IEEE Signal Process. Lett.3
2016 Learning Stacked Image Descriptor for Face Recognition
abstract
Learning-based face descriptors have constantly improved the face recognition performance. Compared with the hand-crafted features, learning-based features are considered to be able to exploit information with better discriminative ability for specific tasks. Motivated by the recent success of deep learning, in this paper, we extend the original shallow face descriptors to deep discriminant face features by introducing a stacked image descriptor (SID). With deep structure, more complex facial information can be extracted and the discriminant and compactness of feature representation can be improved. The SID is learned in a forward optimization way, which is computational efficient compared with deep learning. Extensive experiments on various face databases are conducted to show that SID is able to achieve high face recognition performance with compact face representation, compared with other state-of-the-art descriptors.
Zhen Lei 0001, Dong Yi, Stan Z. Li
IEEE Trans. Circuits Syst. Video Technol.2
2015 High-fidelity Pose and Expression Normalization for face recognition in the wild
abstract
Pose and expression normalization is a crucial step to recover the canonical view of faces under arbitrary conditions, so as to improve the face recognition performance. An ideal normalization method is desired to be automatic, database independent and high-fidelity, where the face appearance should be preserved with little artifact and information loss. However, most normalization methods fail to satisfy one or more of the goals. In this paper, we propose a High-fidelity Pose and Expression Normalization (HPEN) method with 3D Morphable Model (3DMM) which can automatically generate a natural face image in frontal pose and neutral expression. Specifically, we firstly make a landmark marching assumption to describe the non-correspondence between 2D and 3D landmarks caused by pose variations and propose a pose adaptive 3DMM fitting algorithm. Secondly, we mesh the whole image into a 3D object and eliminate the pose and expression variations using an identity preserving 3D transformation. Finally, we propose an inpainting method based on Possion Editing to fill the invisible region caused by self occlusion. Extensive experiments on Multi-PIE and LFW demonstrate that the proposed method significantly improves face recognition performance and outperforms state-of-the-art methods in both constrained and unconstrained environments.
Xiangyu Zhu 0001, Zhen Lei 0001, Dong Yi, Stan Z. Li
CVPR4
2015 High-Performance Video Condensation System
abstract
Video synopsis or condensation is a smart solution for fast video browsing and storage. However, most of the existing methods work offline, where two main phases are required. The first phase is to prepare tubes and background images. The second phase is to rearrange tubes and stitch them into backgrounds. However, with a long video sequence, the first phase is memory consuming for data storage, and the second phase is computationally expensive to rearrange all tubes simultaneously. To overcome these problems, we propose a high-performance video condensation system based on an online content-aware framework. The online framework transforms the optimization problem of tube rearrangement into a stepwise optimization problem. Therefore, it can condense video with much less memory and higher speed than the offline framework. With the aid of this transformation, the proposed system can process input videos and produce condensed videos simultaneously. Thus it is suitable for real-time endless surveillance videos. Meanwhile, the online mechanism allows users to directly visit the condensation video that has been generated. Moreover, the content-aware mechanism makes the proposed system able to automatically determine the duration of a condensed video. Finally, the proposed system uses Graphic Processing Unit (GPU) and multicore techniques to improve the speed. Extensive experiments that validate the high efficiency of the system are presented.
Jianqing Zhu, Shikun Feng, Dong Yi, Shengcai Liao, Zhen Lei 0001, Stan Z. Li
IEEE Trans. Circuits Syst. Video Technol.3
2015 Person-Specific Face Antispoofing With Subject Domain Adaptation
abstract
Face antispoofing is important to practical face recognition systems. In previous works, a generic antispoofing classifier is trained to detect spoofing attacks on all subjects. However, due to the individual differences among subjects, the generic classifier cannot generalize well to all subjects. In this paper, we propose a person-specific face antispoofing approach. It recognizes spoofing attacks using a classifier specifically trained for each subject, which dismisses the interferences among subjects. Moreover, considering the scarce or void fake samples for training, we propose a subject domain adaptation method to synthesize virtual features, which makes it tractable to train well-performed individual face antispoofing classifiers. The extensive experiments on two challenging data sets: 1) CASIA and 2) REPLAY-ATTACK demonstrate the prospect of the proposed approach.
Zhen Lei 0001, Dong Yi, Stan Z. Li
IEEE Trans. Inf. Forensics Secur.3
2014 Age Estimation by Multi-scale Convolutional Network
Dong Yi, Zhen Lei 0001, Stan Z. Li
ACCV (3)1
2014 Multiple Target Tracking Based on Undirected Hierarchical Relation Hypergraph
abstract
Multi-target tracking is an interesting but challenging task in computer vision field. Most previous data association based methods merely consider the relationships (e.g. appearance and motion pattern similarities) between detections in local limited temporal domain, leading to their difficulties in handling long-term occlusion and distinguishing the spatially close targets with similar appearance in crowded scenes. In this paper, a novel data association approach based on undirected hierarchical relation hypergraph is proposed, which formulates the tracking task as a hierarchical dense neighborhoods searching problem on the dynamically constructed undirected affinity graph. The relationships between different detections across the spatiotemporal domain are considered in a high-order way, which makes the tracker robust to the spatially close targets with similar appearance. Meanwhile, the hierarchical design of the optimization process fuels our tracker to long-term occlusion with more robustness. Extensive experiments on various challenging datasets (i.e. PETS2009 dataset, ParkingLot), including both low and high density sequences, demonstrate that the proposed method performs favorably against the state-of-the-art methods.
Longyin Wen, Wenbo Li 0001, Zhen Lei 0001, Dong Yi, Stan Z. Li
CVPR5
2014 Salient Color Names for Person Re-identification
Yang Yang 0062, Jimei Yang, Shengcai Liao, Dong Yi, Stan Z. Li
ECCV (1)5
2014 A benchmark study of large-scale unconstrained face recognition
abstract
Many efforts have been made in recent years to tackle the unconstrained face recognition challenge. For the benchmark of this challenge, the Labeled Faces in theWild (LFW) database has been widely used. However, the standard LFW protocol is very limited, with only 3,000 genuine and 3,000 impostor matches for classification. Today a 97% accuracy can be achieved with this benchmark, remaining a very limited room for algorithm development. However, we argue that this accuracy may be too optimistic because the underlying false accept rate may still be high (e.g. 3%). Furthermore, performance evaluation at low FARs is not statistically sound by the standard protocol due to the limited number of impostor matches. Thereby we develop a new benchmark protocol to fully exploit all the 13,233 LFW face images for large-scale unconstrained face recognition evaluation under both verification and open-set identification scenarios, with a focus at low FARs. Based on the new benchmark, we evaluate 21 face recognition approaches by combining 3 kinds of features and 7 learning algorithms. The benchmark results show that the best algorithm achieves 41.66% verification rates at FAR=0.1%, and 18.07% open-set identification rates at rank 1 and FAR=1%. Accordingly we conclude that the large-scale unconstrained face recognition problem is still largely unresolved, thus further attention and effort is needed in developing effective feature representations and learning algorithms. We thereby release a benchmark tool to advance research in this field.
Shengcai Liao, Zhen Lei 0001, Dong Yi, Stan Z. Li
IJCB3
2014 Multi-camera Trajectory Mining: Database and Evaluation
abstract
In recent years, large-scale video search and mining has been an active research area. Exploring the trajectory of pedestrian of interest in non-overlapping multi-camera network, namely the trajectory mining, is very useful for visual surveillance and criminal investigation. The trajectory mentioned in our work describes the transition of pedestrian among cameras from a macroscopic perspective which is different from the concept in conventional tracking field. In this paper, we collect a database called TMin to promote research and development of trajectory mining. This release of Version 1 contains 1680 images from 30 subjects, all the images are extracted from 6 surveillance videos over two hours, and each subject appears in at least two different cameras. We describe the apparatuses, environments and procedure of the data collection and present baseline performance on the TMin database.
Shengcai Liao, Dong Yi, Zhen Lei 0001, Stan Z. Li
ICPR3
2014 Local Gradient Order Pattern for Face Representation and Recognition
abstract
LBP is an effective descriptor for face recognition. LBP encodes the ordinal relationship between the neighborhood samplings and the central one to obtain robust face representation. However, additional information like the difference among neighboring pixels, which may be helpful for face recognition, is ignored. On the other hand, gradient information which enhances the edge response and suppresses the external noise like illumination variation, is usually useful for face recognition. In this paper, we propose a novel face descriptor, namely local gradient order pattern (LGOP), taking into account the ordinal relationship of gradient responses in local region to obtain robust face representation. After pattern encoding, a 2-D histogram is consequently adopted to calculate the occurrence frequency of different patterns and multi-scale histogram features are extracted to represent the face image. We further adopt whitened principal component analysis (WPCA) to reduce the feature dimensionality and improve the computational efficiency. Extensive experiments on FERET, CAS-PEAL and LFW validates the effectiveness of LGOP for both constrained and unconstrained face recognition problems.
Zhen Lei 0001, Dong Yi, Stan Z. Li
ICPR2
2014 Color Models and Weighted Covariance Estimation for Person Re-identification
abstract
Due to illumination changes, partial occlusions, and object scale differences, person re-identification over disjoint camera views becomes a challenging problem. To address this problem, a variety of image representations have been put forward. In this paper, the illumination invariance and distinctiveness of different color models including the proposed color model are firstly evaluated. Since color distribution is robust to image scales and partial occlusions, color distributions based on different color models are then calculated and fused in the stage of feature extraction. Different color models obtain robustness to different types of illumination and thus fusing them can compensate each other and contribute to better performance. In the stage of feature matching, a weighted KISSME is presented to learn a better distance metric than the original KISSME. Experimental results demonstrate its feasibility and effectiveness. Finally, image pairs are matched based on the learned distance metric. Experiments conducted on two public benchmark datasets (VIPeR and PRID 450S) show that the proposed algorithm outperforms the state-of-the-art methods.
Yang Yang 0062, Shengcai Liao, Zhen Lei 0001, Dong Yi, Stan Z. Li
ICPR4
2014 Deep Metric Learning for Person Re-identification
abstract
Various hand-crafted features and metric learning methods prevail in the field of person re-identification. Compared to these methods, this paper proposes a more general way that can learn a similarity metric from image pixels directly. By using a "siamese" deep neural network, the proposed method can jointly learn the color feature, texture feature and metric in a unified framework. The network has a symmetry structure with two sub-networks which are connected by a cosine layer. Each sub network includes two convolutional layers and a full connected layer. To deal with the big variations of person images, binomial deviance is used to evaluate the cost between similarities and labels, which is proved to be robust to outliers. Experiments on VIPeR illustrate the superior performance of our method and a cross database experiment also shows its good generalization.
Dong Yi, Zhen Lei 0001, Shengcai Liao, Stan Z. Li
ICPR1
2014 Robust 3D Morphable Model Fitting by Sparse SIFT Flow
abstract
3D Morph able Model (3DMM) has been widely used in face analysis for many years. The most challenging part of 3DMM is to find the correspondences between 3D points and 2D pixels. Existing methods only use key points, edges, specular highlights and image pixels to complete the task, which are not accurate or robust. This paper proposes a new algorithm called Sparse SIFT Flow (SSF) to improve the reconstruction accuracy. We mark a set of salient points to control the shape of facial components and use SSF to find their corresponding pixels on the input image. We also incorporate SSF into Multi-Features Framework to construct a robust 3DMM fitting algorithm. Compared with the state-of-the art, our approach significantly improves the fitting results in facial component area.
Xiangyu Zhu 0001, Dong Yi, Zhen Lei 0001, Stan Z. Li
ICPR2
2014 Dynamic Image-to-Class Warping for Occluded Face Recognition
abstract
Face recognition (FR) systems in real-world applications need to deal with a wide range of interferences, such as occlusions and disguises in face images. Compared with other forms of interferences such as nonuniform illumination and pose changes, face with occlusions has not attracted enough attention yet. A novel approach, coined dynamic image-to-class warping (DICW), is proposed in this work to deal with this challenge in FR. The face consists of the forehead, eyes, nose, mouth, and chin in a natural order and this order does not change despite occlusions. Thus, a face image is partitioned into patches, which are then concatenated in the raster scan order to form an ordered sequence. Considering this order information, DICW computes the image-to-class distance between a query face and those of an enrolled subject by finding the optimal alignment between the query sequence and all sequences of that subject along both the time dimension and within-class dimension. Unlike most existing methods, our method is able to deal with occlusions which exist in both gallery and probe images. Extensive experiments on public face databases with various types of occlusions have confirmed the effectiveness of the proposed method.
Xingjie Wei, Chang-Tsun Li, Zhen Lei 0001, Dong Yi, Stan Z. Li
IEEE Trans. Inf. Forensics Secur.4
2014 Robust Online Learned Spatio-Temporal Context Model for Visual Tracking
abstract
Visual tracking is an important but challenging problem in the computer vision field. In the real world, the appearances of the target and its surroundings change continuously over space and time, which provides effective information to track the target robustly. However, enough attention has not been paid to the spatio-temporal appearance information in previous works. In this paper, a robust spatio-temporal context model based tracker is presented to complete the tracking task in unconstrained environments. The tracker is constructed with temporal and spatial appearance context models. The temporal appearance context model captures the historical appearance of the target to prevent the tracker from drifting to the background in a long-term tracking. The spatial appearance context model integrates contributors to build a supporting field. The contributors are the patches with the same size of the target at the key-points automatically discovered around the target. The constructed supporting field provides much more information than the appearance of the target itself, and thus, ensures the robustness of the tracker in complex environments. Extensive experiments on various challenging databases validate the superiority of our tracker over other state-of-the-art trackers.
Longyin Wen, Zhaowei Cai, Zhen Lei 0001, Dong Yi, Stan Z. Li
IEEE Trans. Image Process.4
2013 Towards Pose Robust Face Recognition
abstract
Most existing pose robust methods are too computational complex to meet practical applications and their performance under unconstrained environments are rarely evaluated. In this paper, we propose a novel method for pose robust face recognition towards practical applications, which is fast, pose robust and can work well under unconstrained environments. Firstly, a 3D deformable model is built and a fast 3D model fitting algorithm is proposed to estimate the pose of face image. Secondly, a group of Gabor filters are transformed according to the pose and shape of face image for feature extraction. Finally, PCA is applied on the pose adaptive Gabor features to remove the redundances and Cosine metric is used to evaluate the similarity. The proposed method has three advantages: (1) The pose correction is applied in the filter space rather than image space, which makes our method less affected by the precision of the 3D model, (2) By combining the holistic pose transformation and local Gabor filtering, the final feature is robust to pose and other negative factors in face recognition, (3) The 3D structure and facial symmetry are successfully used to deal with self-occlusion. Extensive experiments on FERET and PIE show the proposed method outperforms state-of-the-art methods significantly, meanwhile, the method works well on LFW.
Dong Yi, Zhen Lei 0001, Stan Z. Li
CVPR1
2012 Online Multiple Instance Joint Model for Visual Tracking
abstract
Although numerous online learning strategies have been proposed to handle the appearance variation in visual tracking, the existing methods just perform well in certain cases since they lack effective appearance learning mechanism. In this paper, a joint model tracker (JMT) is presented, which consists of a generative model based on Multiple Subspaces and a discriminative model based on improved Multiple Instance Boosting (MIBoosting). The generative model utilizes a series of local constructed subspaces to update the Multiple Subspaces model and considers the energy dissipation of dimension reduction in updating step. The discriminative model adopts the Gaussian Mixture Model (GMM) to estimate the posterior probability of the likelihood function. These two parts supervise each other to update in multiple instance way which helps our tracker recover from drift. Extensive experiments on various databases validate the effectiveness of our proposed method over other state-of-the-art trackers.
Longyin Wen, Zhaowei Cai, Menglong Yang, Zhen Lei 0001, Dong Yi, Stan Z. Li
AVSS5
2012 Water Filling: Unsupervised People Counting via Vertical Kinect Sensor
abstract
People counting is one of the key components in video surveillance applications, however, due to occlusion, illumination, color and texture variation, the problem is far from being solved. Different from traditional visible camera based systems, we construct a novel system that uses vertical Kinect sensor for people counting, where the depth information is used to remove the affect of the appearance variation. Since the head is always closer to the Kinect sensor than other parts of the body, people counting task equals to find the suitable local minimum regions. According to the particularity of the depth map, we propose a novel unsupervised water filling method that can find these regions with the property of robustness, locality and scale-invariance. Experimental comparisons with mean shift and random forest on two databases validate the superiority of our water filling algorithm in people counting.
Xucong Zhang, Shikun Feng, Zhen Lei 0001, Dong Yi, Stan Z. Li
AVSS5
2012 Online content-aware video condensation
abstract
Explosive growth of surveillance video data presents formidable challenges to its browsing, retrieval and storage. Video synopsis, an innovation proposed by Peleg and his colleagues, is aimed for fast browsing by shortening the video into a synopsis while keeping activities in video captured by a camera. However, the current techniques are offline methods requiring that all the video data be ready for the processing, and are expensive in time and space. In this paper, we propose an online and efficient solution, and its supporting algorithms to overcome the problems. The method adopts an online content-aware approach in a step-wise manner, hence applicable to endless video, with less computational cost. Moreover, we propose a novel tracking method, called sticky tracking, to achieve high-quality visualization. The system can achieve a faster-than-real-time speed with a multi-core CPU implementation. The advantages are demonstrated by extensive experiments with a wide variety of videos. The proposed solution and algorithms could be integrated with surveillance cameras, and impact the way that surveillance videos are recorded.
Shikun Feng, Zhen Lei 0001, Dong Yi, Stan Z. Li
CVPR3
2012 Discriminant image filter learning for face recognition with local binary pattern like representation
abstract
Local binary pattern (LBP) and its variants are effective descriptors for face recognition. The traditional LBP like features are extracted based on the original pixel or patch values of images. In this paper, we propose to learn the discriminative image filter to improve the discriminant power of the LBP like feature. The basic idea is after the image filtering with the learned filter, the difference of pixel difference vectors (PDVs) between the images from the same person is consistent and the difference between the images from different persons is enlarged. In this way, the LBP like features extracted from the filtered images are considered to be more discriminant than those extracted from the original images. Moreover, a coupled discriminant image filters learning method is proposed to deal with the heterogenous face images matching problem by reducing the feature gap between the heterogeneous images. Experiments on FERET, FRGC and a VIS-NIR heterogeneous face databases validate the effectiveness of our proposed image filter learning method combined with LBP like features.
Zhen Lei 0001, Dong Yi, Stan Z. Li
CVPR2
2012 Multi-pedestrian detection in crowded scenes: A global view
abstract
Recent state-of-the-art algorithms have achieved good performance on normal pedestrian detection tasks. However, pedestrian detection in crowded scenes is still challenging due to the significant appearance variation caused by heavy occlusions and complex spatial interactions. In this paper we propose a unified probabilistic framework to globally describe multiple pedestrians in crowded scenes in terms of appearance and spatial interaction. We utilize a mixture model, where every pedestrian is assumed in a special subclass and described by the sub-model. Scores of pedestrian parts are used to represent appearance and quadratic kernel is used to represent relative spatial interaction. For efficient inference, multi-pedestrian detection is modeled as a MAP problem and we utilize greedy algorithm to get an approximation. For discriminative parameter learning, we formulate it as a learning to rank problem, and propose Latent Rank SVM for learning from weakly labeled data. Experiments on various databases validate the effectiveness of the proposed approach.
Zhen Lei 0001, Dong Yi, Stan Z. Li
CVPR3
2012 Online Spatio-temporal Structural Context Learning for Visual Tracking
Longyin Wen, Zhaowei Cai, Zhen Lei 0001, Dong Yi, Stan Z. Li
ECCV (4)4
2012 Face liveness detection by exploring multiple scenic clues
abstract
Liveness detection is an indispensable guarantee for reliable face recognition, which has recently received enormous attention. In this paper we propose three scenic clues, which are non-rigid motion, face-background consistency and imaging banding effect, to conduct accurate and efficient face liveness detection. Non-rigid motion clue indicates the facial motions that a genuine face can exhibit such as blinking, and a low rank matrix decomposition based image alignment approach is designed to extract this non-rigid motion. Face-background consistency clue believes that the motion of face and background has high consistency for fake facial photos while low consistency for genuine faces, and this consistency can serve as an efficient liveness clue which is explored by GMM based motion detection method. Image banding effect reflects the imaging quality defects introduced in the fake face reproduction, which can be detected by wavelet decomposition. By fusing these three clues, we thoroughly explore sufficient clues for liveness detection. The proposed face liveness detection method achieves 100% accuracy on Idiap print-attack database and the best performance on self-collected face anti-spoofing database.
Zhen Lei 0001, Dong Yi, Stan Z. Li
ICARCV4
2012 Regularized Transfer Boosting for Face Detection Across Spectrum
abstract
This letter addresses the problem of face detection in multispectral illuminations. Face detection in visible images has been well addressed based on the large scale training samples. For the recently emerging multispectral face biometrics, however, the face data is scarce and expensive to collect, and it is usually short of face samples to train an accurate face detector. In this letter, we propose to tackle the issue of multispectral face detection by combining existing large scale visible face images and a few multispectral face images. We cast the problem of face detection across spectrum into the transfer learning framework and try to learn the robust multispectral face detector by exploring relevant knowledge from visible data domain. Specifically, a novel Regularized Transfer Boosting algorithm named R-TrBoost is proposed, with features of weighted loss objective and manifold regularization. Experiments are performed with face images of two spectrums, 850 nm and 365 nm, and the results show significant improvement on multispectral face detection using the proposed algorithm.
Dong Yi, Zhen Lei 0001, Stan Z. Li
IEEE Signal Process. Lett.2
2011 Learning sparse feature for eyeglasses problem in face recognition
abstract
Occlusion of eyeglasses, and strong specular reflections on eyeglasses (especially in near infrared (NIR) images), can deteriorate face recognition performance. In this paper, we present a novel method to overcome these problems. The proposed method applies the sparse representation (SR) technique in a local feature space so as to be more tolerant to mis-alignment and abnormal specular pixel values. The SR face features are further transformed by using discriminant analysis. These lead to a good balance between efficiency and robustness. Extensive experiments on a large NIR face database containing 292 persons with/without eyeglasses show the superiority of the proposed method compared with state-of-the-art methods.
Dong Yi, Stan Z. Li
FG1
2011 Face liveness detection by learning multispectral reflectance distributions
abstract
Existing face liveness detection algorithms adopt behavioural challenge-response methods that require user cooperation. To be verified live, users are expected to obey some user unfriendly requirement. In this paper, we present a multispectral face liveness detection method, which is user cooperation free. Moreover, the system is adaptive to various user-system distances. Using the Lambertian model, we analyze multispectral properties of human skin versus non-skin, and the discriminative wavelengths are then chosen. Reflectance data of genuine and fake faces at multi-distances are selected to form a training set. An SVM classifier is trained to learn the multispectral distribution for a final Genuine-or-Fake classification. Compared with previous works, the proposed method has the following advantages: (a) The requirement on the users' cooperation is no longer needed, making the liveness detection user friendly and fast. (b) The system can work without restricted distance requirement from the target being analyzed. Experiments are conducted on genuine versus planar face data, and genuine versus mask face data. Furthermore a comparison with the visible challenge-response liveness detection method is also given. The experimental results clearly demonstrate the superiority of our method over previous systems.
Dong Yi, Zhen Lei 0001, Stan Z. Li
FG2
2011 Competition on counter measures to 2-D facial spoofing attacks
abstract
Spoofing identities using photographs is one of the most common techniques to attack 2-D face recognition systems. There seems to exist no comparative studies of different techniques using the same protocols and data. The motivation behind this competition is to compare the performance of different state-of-the-art algorithms on the same database using a unique evaluation method. Six different teams from universities around the world have participated in the contest. Use of one or multiple techniques from motion, texture analysis and liveness detection appears to be the common trend in this competition. Most of the algorithms are able to clearly separate spoof attempts from real accesses. The results suggest the investigation of more complex attacks.
Murali Mohan Chakka, André Anjos, Sébastien Marcel, Roberto Tronci, Daniele Muntoni, Gianluca Fadda, Maurizio Pili, Nicola Sirena, Gabriele Murgia, Marco Ristori, Fabio Roli, Dong Yi, Zhen Lei 0001, Stan Z. Li, William Robson Schwartz, Anderson Rocha 0001, Hélio Pedrini, Javier Lorenzo-Navarro, Modesto Castrillón-Santana, Jukka Komulainen, Abdenour Hadid, Matti Pietikäinen
IJCB13
2011 Towards incremental and large scale face recognition
abstract
Linear discriminant analysis with nearest neighborhood classifier (LDA + NN) has been commonly used in face recognition, but it often confronts with two problems in real applications: (1) it cannot incrementally deal with the information of training instances; (2) it cannot achieve fast search against large scale gallery set. In this paper, we use incremental LDA (ILDA) and hashing based search method to deal with these two problems. Firstly two incremental LDA algorithms are proposed under spectral regression framework, namely exact incremental spectral regression discriminant analysis (EI-SRDA) and approximate incremental spectral regression discriminant analysis (AI-SRDA). Secondly we propose a similarity hashing algorithm of sub-linear complexity to achieve quick recognition from large gallery set. Experiments on FRGC and self-collected 100,000 faces database show the effective of our methods.
Zhen Lei 0001, Dong Yi, Stan Z. Li
IJCB3
2011 A robust eye localization method for low quality face images
abstract
Eye localization is an important part in face recognition system, because its precision closely affects the performance of face recognition. Although various methods have already achieved high precision on the face images with high quality, their precision will drop on low quality images. In this paper, we propose a robust eye localization method for low quality face images to improve the eye detection rate and localization precision. First, we propose a probabilistic cascade (P-Cascade) framework, in which we reformulate the traditional cascade classifier in a probabilistic way. The P-Cascade can give chance to each image patch contributing to the final result, regardless the patch is accepted or rejected by the cascade. Second, we propose two extensions to further improve the robustness and precision in the P-Cascade framework. There are: (1) extending feature set, and (2) stacking two classifiers in multiple scales. Extensive experiments on JAFFE, BioID, LFW and a self-collected video surveillance database show that our method is comparable to state-of-the-art methods on high quality images and can work well on low quality images. This work supplies a solid base for face recognition applications under unconstrained or surveillance environments.
Dong Yi, Zhen Lei 0001, Stan Z. Li
IJCB1
2011 Low-resolution face recognition via Simultaneous Discriminant Analysis
abstract
Low resolution (LR) is an important issue when handling real world face recognition problems. The performance of traditional recognition algorithms will drop drastically due to the loss of facial texture information in original high resolution (HR) images. To address this problem, in this paper we propose an effective approach named Simultaneous Discriminant Analysis (SDA). SDA learns two mappings from LR and HR images respectively to a common subspace where discrimination property is maximized. In SDA, (1) the data gap between LR and HR is reduced by mapping into a common space; and (2) the mapping is designed for preserving most discriminative information. After that, the conventional classification method is applied in the common space for final decision. Extensive experiments are conducted on both FERET and Multi-PIE, and the results clearly show the superiority of the proposed SDA over state-of-the-art methods.
Changtao Zhou, Dong Yi, Zhen Lei 0001, Stan Z. Li
IJCB3
2009 Learning mappings for face synthesis from near infrared to visual light images
abstract
This paper deals with a new problem in face recognition research, in which the enrollment and query face samples are captured under different lighting conditions. In our case, the enrollment samples are visual light (VIS) images, whereas the query samples are taken under near infrared (NIR) condition. It is very difficult to directly match the face samples captured under these two lighting conditions due to their different visual appearances. In this paper, we propose a novel method for synthesizing VIS images from NIR images based on learning the mappings between images of different spectra (i.e., NIR and VIS). In our approach, we reduce the inter-spectral differences significantly, thus allowing effective matching between faces taken under different imaging conditions. Face recognition experiments clearly show the efficacy of the proposed approach.
Jie Chen 0001, Dong Yi, Jimei Yang, Guoying Zhao 0001, Stan Z. Li, Matti Pietikäinen
CVPR2
2008 2D-3D face matching using CCA
abstract
In recent years, 3D face recognition has obtained much attention. Using 2D face image as probe and 3D face data as gallery is an alternative method to deal with computation complexity, expensive equipment and fussy pretreatment in 3D face recognition systems. In this paper we propose a learning based 2D-3D face matching method using the CCA to learn the mapping between 2D face image and 3D face data. This method makes it possible to match the on-site 2D face image with enrolled 3D face data. Our 2D-3D face matching method decreased the computation complexity drastically compared to the conventional 3D-3D face matching while keeping relative high recognition rate. Furthermore, to simplify the mapping between 2D face image and 3D face data, a patch based strategy is proposed to boost the accuracy of matching. And the kernel method is also evaluated to reveal the non-linear relationship. The experiment results show that CCA based method has good performance and patch based method has significant improvement compared to the holistic method.
Weilong Yang, Dong Yi, Zhen Lei 0001, Stan Z. Li
FG2
2003 Design of a new grasper having XYZ translational motions
abstract
A new 4 DOF parallel mechanism is proposed in this work. This device consists of four parallel kinematics chains and a foldable parallelogrammic platform that can be used to grasp any large or irregular-shaped object. Thus, out of 4-DOF motion space of the device, the one-DOF is used for gripping an irregular object and the other three-DOF is used for adjusting motion of the grasped object. Particularly, the three-DOF motion is restricted in the decoupled three-dimensional translational motion space in spite of revolute joint-based parallel structure. Thus, it is not only very compact, but also has distinctive feature of both grasping and micro-positioning that are one of important aspects required in real applications. In this work, we carry out the position and kinematic analysis for the mechanism, and develop the mechanism for experimental verification of its performance.
Dong Yi, Byung-Ju Yi, Whee Kuk Kim
ICRA1