EDBT 2026 Demo / reviewers in the wild / expert
Xiujuan Chai
dblp:57/6273
· DBLP profile ↗
43ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0002-2757-9900ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 25 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Video understanding and tracking · 28% 3D vision · 27% Face, body and person analysis · 26% | |
| Computer graphics and multimedia
1 paper |
Geometric modeling and processing · 100% |
Topics — the 20 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
sign language recognition |
1.1 | 3 | 2021 | Visual Alignment Constraint for Continuous Sign Language Recognition · ICCV 2021 A Novel Sign Language Recognition Framework Using Hierarchical Grassmann Covariance Matrix · IEEE Trans. Multim. 2019 Iterative Reference Driven Metric Learning for Signer Independent Isolated Sign Language Recognition · ECCV (7) 2016 |
Computer vision › Face, body and person analysis
face recognition |
1.0 | 5 | 2022 | MC-GCN: A Multi-Scale Contrastive Graph Convolutional Network for Unconstrained Face Recognition With Image Sets · IEEE Trans. Image Process. 2022 Maximal Likelihood Correspondence Estimation for Face Recognition Across Pose · IEEE Trans. Image Process. 2014 Morphable Displacement Field Based Image Matching for Face Recognition across Pose · ECCV (1) 2012 |
Computer vision › 3D vision
3d reconstruction |
1.0 | 1 | 2026 | Slender3D: Curve-Guided Multi-View Reconstruction of Slender Structures · AAAI 2026 |
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction |
1.0 | 1 | 2026 | Slender3D: Curve-Guided Multi-View Reconstruction of Slender Structures · AAAI 2026 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.6 | 1 | 2022 | MC-GCN: A Multi-Scale Contrastive Graph Convolutional Network for Unconstrained Face Recognition With Image Sets · IEEE Trans. Image Process. 2022 |
Computer vision › Face, body and person analysis › face recognition › face matching
image set-based face recognition |
0.6 | 1 | 2022 | MC-GCN: A Multi-Scale Contrastive Graph Convolutional Network for Unconstrained Face Recognition With Image Sets · IEEE Trans. Image Process. 2022 |
Computer vision › Video understanding and tracking
visual sequence learning |
0.6 | 1 | 2022 | Deep Radial Embedding for Visual Sequence Learning · ECCV (6) 2022 |
Computer vision › Video understanding and tracking › sign language recognition
continuous sign language recognition |
0.5 | 1 | 2021 | Visual Alignment Constraint for Continuous Sign Language Recognition · ICCV 2021 |
Computer vision › Video understanding and tracking
gesture recognition |
0.4 | 1 | 2020 | An Efficient PointLSTM for Point Clouds Based Gesture Recognition · CVPR 2020 |
Computer vision › 3D vision
point cloud processing |
0.4 | 1 | 2020 | An Efficient PointLSTM for Point Clouds Based Gesture Recognition · CVPR 2020 |
Computer vision › Face, body and person analysis
face alignment |
0.3 | 2 | 2013 | Cascaded Shape Space Pruning for Robust Facial Landmark Detection · ICCV 2013 Joint Face Alignment: Rescue Bad Alignments with Good Ones by Regularized Re-fitting · ECCV (2) 2012 |
Geometric modeling and processing
surface reconstruction |
0.3 | 1 | 2026 | Slender3D: Curve-Guided Multi-View Reconstruction of Slender Structures · AAAI 2026 |
Computer vision › Face, body and person analysis › face recognition › robust face recognition
pose-invariant face recognition |
0.3 | 2 | 2014 | Maximal Likelihood Correspondence Estimation for Face Recognition Across Pose · IEEE Trans. Image Process. 2014 Locally Linear Regression for Pose-Invariant Face Recognition · IEEE Trans. Image Process. 2007 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.2 | 1 | 2016 | Iterative Reference Driven Metric Learning for Signer Independent Isolated Sign Language Recognition · ECCV (7) 2016 |
Computer vision › 3D vision › 3d face modeling
3d morphable model |
0.2 | 1 | 2014 | Maximal Likelihood Correspondence Estimation for Face Recognition Across Pose · IEEE Trans. Image Process. 2014 |
Computer vision › Face, body and person analysis
face modeling |
0.2 | 1 | 2014 | Maximal Likelihood Correspondence Estimation for Face Recognition Across Pose · IEEE Trans. Image Process. 2014 |
Computer vision › Face, body and person analysis › face alignment
large-pose face alignment |
0.1 | 1 | 2012 | Morphable Displacement Field Based Image Matching for Face Recognition across Pose · ECCV (1) 2012 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.1 | 1 | 2020 | An Efficient PointLSTM for Point Clouds Based Gesture Recognition · CVPR 2020 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › riemannian manifold
grassmann manifold |
0.1 | 1 | 2019 | A Novel Sign Language Recognition Framework Using Hierarchical Grassmann Covariance Matrix · IEEE Trans. Multim. 2019 |
Computer vision › Video understanding and tracking › temporal understanding
temporal segmentation |
0.1 | 1 | 2019 | A Novel Sign Language Recognition Framework Using Hierarchical Grassmann Covariance Matrix · IEEE Trans. Multim. 2019 |
Methods — techniques the papers use, named apart from their topics
structure from motion · 2.0mesh rasterization · 2.0differentiable poisson reconstruction · 2.02d gaussian splatting · 2.0deep metric learning · 1.0sequence embedding · 0.6graph convolutional network · 0.6contrastive learning · 0.6attention mechanism · 0.6alignment supervision · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Slender3D: Curve-Guided Multi-View Reconstruction of Slender StructuresabstractAlthough geometric reconstruction of general objects from images has made remarkable progress in recent years, slender structures remain largely underexplored, despite their critical importance in engineering, biomedical, and agricultural applications. To bridge this gap, we propose a dedicated 2DGS-based geometric reconstruction framework tailored for slender structures, achieving accurate and faithful geometry recovery. Our method first addresses the challenge that most slender objects are texture-less, which hinders reliable feature matching and pose estimation in traditional SfM pipelines. By leveraging the curve-like nature of slender structures, we perform a curve-guided SfM process that provides robust camera poses and accurate 3D curve initialization for Gaussian primitives. To ensure SfM reliability, we introduce a high-precision mask extraction strategy that integrates geometric priors with a segmentation network, effectively handling self-occlusion and thin geometry. Furthermore, to enhance fine geometric recovery, we incorporate a differentiable Poisson reconstruction module to extract an initial mesh during training, which is then refined via image-space iterative optimization using differentiable mesh rasterization. In contrast to conventional approaches that rely on differentiable Gaussian rasterization followed by TSDF-based mesh extraction, our method avoids the additional geometric errors and artifacts introduced during the intermediate TSDF conversion, thereby improving the overall reconstruction quality. Comprehensive experiments on both synthetic and real-world datasets validate that our method achieves superior reconstruction quality compared to state-of-the-art approaches. Suqin Wang, Zeyi Wang, Min Shi 0005, Zhaoxin Li, Qi Wang 0111, Xiujuan Chai, Dengming Zhu |
AAAI | 6 |
| 2022 | Deep Radial Embedding for Visual Sequence Learning
Yuecong Min, Peiqi Jiao, Xiaotao Wang, Xiujuan Chai, Xilin Chen 0001 |
ECCV (6) | 6 |
| 2022 | Library on-shelf book segmentation and recognition based on deep visual features
Tan Sun, Guojian Xian, Xiujuan Chai |
Inf. Process. Manag. | 7 |
| 2022 | MC-GCN: A Multi-Scale Contrastive Graph Convolutional Network for Unconstrained Face Recognition With Image SetsabstractIn this paper, a Multi-scale Contrastive Graph Convolutional Network (MC-GCN) method is proposed for unconstrained face recognition with image sets, which takes a set of media (orderless images and videos) as a face subject instead of single media (an image or video). Due to factors such as illumination, posture, media source, etc., there are huge intra-set variances in a face set, and the importance of different face prototypes varies considerably. How to model the attention mechanism according to the relationship between prototypes or images in a set is the main content of this paper. In this work, we formulate a framework based on graph convolutional network (GCN), which considers face prototypes as nodes to build relations. Specifically, we first present a multi-scale graph module to learn the relationship between prototypes at multiple scales. Moreover, a Contrastive Graph Convolutional (CGC) block is introduced to build attention control model, which focuses on those frames with similar prototypes (contrastive information) between pair of sets instead of simply evaluating the frame quality. The experiments on IJB-A, YouTube Face, and an animal face dataset clearly demonstrate that our proposed MC-GCN outperforms the state-of-the-art methods significantly. Xiao Shi 0004, Xiujuan Chai, Jiake Xie, Tan Sun |
IEEE Trans. Image Process. | 2 |
| 2021 | Visual Alignment Constraint for Continuous Sign Language RecognitionabstractVision-based Continuous Sign Language Recognition (CSLR) aims to recognize unsegmented signs from image streams. Overfitting is one of the most critical problems in CSLR training, and previous works show that the iterative training scheme can partially solve this problem while also costing more training time. In this study, we revisit the iterative training scheme in recent CSLR works and realize that sufficient training of the feature extractor is critical to solving the overfitting problem. Therefore, we propose a Visual Alignment Constraint (VAC) to enhance the feature extractor with alignment supervision. Specifically, the proposed VAC comprises two auxiliary losses: one focuses on visual features only, and the other enforces prediction alignment between the feature extractor and the alignment module. Moreover, we propose two metrics to reflect overfitting by measuring the prediction inconsistency between the feature extractor and the alignment module. Experimental results on two challenging CSLR datasets show that the proposed VAC makes CSLR networks end-to-end trainable and achieves competitive performance. Yuecong Min, Aiming Hao, Xiujuan Chai, Xilin Chen 0001 |
ICCV | 3 |
| 2020 | An Efficient PointLSTM for Point Clouds Based Gesture RecognitionabstractPoint clouds contain rich spatial information, which provides complementary cues for gesture recognition. In this paper, we formulate gesture recognition as an irregular sequence recognition problem and aim to capture long-term spatial correlations across point cloud sequences. A novel and effective PointLSTM is proposed to propagate information from past to future while preserving the spatial structure. The proposed PointLSTM combines state information from neighboring points in the past with current features to update the current states by a weight-shared LSTM layer. This method can be integrated into many other sequence learning approaches. In the task of gesture recognition, the proposed PointLSTM achieves state-of-the-art results on two challenging datasets (NVGesture and SHREC'17) and outperforms previous skeleton-based methods. To show its advantages in generalization, we evaluate our method on MSR Action3D dataset, and it produces competitive results with previous skeleton-based methods. Yuecong Min, Yanxiao Zhang, Xiujuan Chai, Xilin Chen 0001 |
CVPR | 3 |
| 2020 | Deep Cross-Species Feature Learning for Animal Face Recognition via Residual Interspecies Equivariant Network
Xiao Shi 0004, Chenxue Yang, Xue Xia 0004, Xiujuan Chai |
ECCV (27) | 4 |
| 2019 | FlickerNet: Adaptive 3D Gesture Recognition from Sparse Point Clouds
Yuecong Min, Xiujuan Chai, Xilin Chen 0001 |
BMVC | 2 |
| 2019 | Prior Knowledge Guided Small Object Detection on High-Resolution ImagesabstractWhen applying common object detection algorithms to detect small objects on high-resolution images, the down-sampling operation of the input images is inevitable due to the limitation of GPU memory. Accordingly, the details for characterizing small objects are lost. To resolve this contradiction, a small object detection method in a coarse-to-fine manner is presented. Specifically, some rough regions of interest (ROI) are firstly computed from low-resolution images. The prior knowledge of the positions of objects is used to guide the generation of ROIs. Then the features of small ROIs are recomputed from high-resolution images, and the features of large ROIs are obtained from the feature maps used to generate ROIs. The proposed method is validated on two datasets. One is a plant phenotyping dataset and the other is a public traffic sign dataset. Experimental results convincingly show the effectiveness of the proposed method. Xiujuan Chai, Ruiping Wang 0001, Weijun Guo, Li Pu, Xilin Chen 0001 |
ICIP | 2 |
| 2019 | Locality-constrained framework for face alignment
Jie Zhang 0071, Meina Kan, Shiguang Shan, Xiujuan Chai, Xilin Chen 0001 |
Frontiers Comput. Sci. | 5 |
| 2019 | Deep memory and prediction neural network for video prediction
Xiujuan Chai, Xilin Chen 0001 |
Neurocomputing | 2 |
| 2019 | A Novel Sign Language Recognition Framework Using Hierarchical Grassmann Covariance MatrixabstractVisual sign language recognition is an interesting and challenging problem. To create a discriminative representation, a hierarchical Grassmann covariance matrix (HGCM) model is proposed for sign description. Furthermore, a multi-temporal belief propagation (MTBP) based segmentation approach is presented for continuous sequence spotting. Concretely speaking, a sign is represented by multiple covariance matrices, followed by evaluating and selecting their most significant singular vectors. These covariance matrices are transformed into a more compact and discriminative HGCM, which is formulated on the Grassmann manifold. Continuous sign sequences can be recognized frame by frame using the HGCM model, before being optimized by MTBP, which is a carefully designed graphic model. The proposed method is thoroughly evaluated on isolated and synthetic and real continuous sign datasets as well as on HDM05. Extensive experimental results convincingly show the effectiveness of our proposed framework. Hanjie Wang, Xiujuan Chai, Xilin Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | ScoringNet: Learning Key Fragment for Action Quality Assessment with Ranking Loss in Skilled Sports
Yongjun Li 0004, Xiujuan Chai, Xilin Chen 0001 |
ACCV (6) | 2 |
| 2018 | Kinematic Constrained Cascaded Autoencoder for Real-Time Hand Pose EstimationabstractHand pose estimation is an attractive problem in computer vision for its key role in gesture controlled humancomputer interaction (HCI) applications. This problem focuses on revealing the hand skeleton structure from visual information. However, it is very challenging for complicated hand configurations. In this paper, a kinematic constrained cascaded autoencoder regression (KCAE) framework is proposed to estimate the hand pose from a single depth image.We introduce a two-stage cascaded structure to regress the palm direction and the whole hand joints successively. In addition, the edge constraints are first introduced to the loss function with an endto- end manner, which maintains the kinematics of hands and makes the prediction more reasonable. With this framework, different features are evaluated, including the handcrafted features and the features learned from CNN. The experiments widely conducted on our collected dataset and the public MSRA hand gesture database demonstrate the effectiveness of KCAE. Overall, The proposed method achieves comparable performance with state-of-the-arts. Yushun Lin, Xiujuan Chai, Xilin Chen 0001 |
FG | 2 |
| 2016 | Iterative Reference Driven Metric Learning for Signer Independent Isolated Sign Language Recognition
Fang Yin, Xiujuan Chai, Xilin Chen 0001 |
ECCV (7) | 2 |
| 2016 | Two streams Recurrent Neural Networks for Large-Scale Continuous Gesture RecognitionabstractIn this paper, we tackle the continuous gesture recognition problem with a two streams Recurrent Neural Networks (2S-RNN) for the RGB-D data input. In our framework, the spotting-recognition strategy is used, that means the continuous gestures are first segmented into separated gestures, and then each isolated gesture is recognized by using the 2S-RNN. Concretely, the gesture segmentation is based on the accurate hand positions provided by the hand detector trained from Faster R-CNN. While in the recognition module, 2S-RNN is designed to efficiently fuse multi-modal features, i.e. the RGB and depth channels. The experimental results on both the validation and test sets of the Continuous Gesture Dataset (ConGD) have shown promising performance of the proposed framework. We ranked 1st in the ChaLearn LAP Large-scale Continuous Gesture Recognition Challenge with the mean Jaccard Index of 0.286915. Xiujuan Chai, Fang Yin, Xilin Chen 0001 |
ICPR | 1 |
| 2016 | Sparse Observation (SO) Alignment for Sign Language Recognition
Hanjie Wang, Xiujuan Chai, Xilin Chen 0001 |
Neurocomputing | 2 |
| 2015 | Communication tool for the hard of hearings: A large vocabulary sign language recognition systemabstractDeaf person has a large social community around the world. The smooth communication is very difficult for these hard of hearings. Automatic Sign Language Recognition (SLR) can build the bridge between the deaf and the hearings and turn the seamless interaction into reality. This paper presents a visualized communication tool for the hard of hearings, i.e. a large vocabulary sign language recognition system based on the RGB-D data input. A novel Grassmann Covariance Matrix (GCM) representation is used to encode a long-term dynamics of a sign sequence and the discriminative kernel SVM is adopted for the sign classification. For continuous sign language recognition, a probability inference method is used to determine the spotting from the labels of sequential frames. Some basic evaluation and comparison of our recognition algorithms are conducted in our collected datasets. This demo will show the recognition of both isolated sign words and the continuous sign language sentences. Xiujuan Chai, Hanjie Wang, Fang Yin, Xilin Chen 0001 |
ACII | 1 |
| 2015 | Weakly Supervised Metric Learning towards Signer Adaptation for Sign Language RecognitionabstractIn this paper, we introduce metric learning into Sign Language Recognition(SLR) for the first time and propose a signer adaption framework to address signer-independent SLR. For adapting the general model to the new signer, both clustering and manifold constraints are considered in the adaptive distance metric optimization. The contribution of our work mainly lies in three-folds. Firstly, a Weakly Supervised Metric Learning(WSML) framework is proposed, which combines the clustering and manifold constraints simultaneously. Secondly, the general framework is applied to signer adaptation and achieves good performance. Thirdly, a fragment based feature is designed for sign language representation and the effectiveness is verified in large vocabulary datasets. Our proposed WSML framework can be decomposed into two key steps. The first one is to learn a generic metric from the given labeled data. Then the second step is to realize the distance metric adaptation by considering the clustering and manifold constraints with the unlabeled data. To learn a generic distance metric, the labeled data are used under clustering assumption with classical large margin hinge loss. Specifically, the distances between data points within the same cluster(with same label) should be minimized and the distances between data points from different clusters(with different labels) should be maximized. Here we define the index set with same labels as Sg = {(i, j)|yi = y j,xi,x j ∈ Xl} and the index triplet Bg = {(i, j,k)|yi = y j,yi 6= yk,xi,x j,xk ∈ Xl}. The objective function is Fang Yin, Xiujuan Chai, Yu Zhou 0015, Xilin Chen 0001 |
BMVC | 2 |
| 2015 | Semantics constrained dictionary learning for signer-independent sign language recognitionabstractIn this paper, a sparse coding based framework is proposed for sign language recognition (SLR), especially for the signer-independent case. To deal with the inter-signer variation, a dictionary capturing the common features among different signers is learnt by considering the semantic constraint. Thus for a given sign from an unknown signer, the sparse representation, which maintains more information of this specific sign class while neglecting the identity information as much as possible, can be generated. In our implementation, each sign is partitioned into a fixed number of fragments and the features fusing hand shape and moving trajectory are extracted from the fragments. The dictionary learnt from the training fragments can be taken as the basic subunits of signs and each fragment of sign video can be coded by these basis vectors. Finally, the recognition result is achieved through SVM with the concatenated sparse coding features of the fragments. The experiments and comparisons show that our method is more effective for the signer-independent recognition problem than other baseline methods. At the same time, it also performs well for the signer-dependent case. Fang Yin, Xiujuan Chai, Yu Zhou 0015, Xilin Chen 0001 |
ICIP | 2 |
| 2015 | Unsupervised adaptive sign language recognition based on hypothesis comparison guided cross validation and linguistic prior filtering
Yu Zhou 0015, Xiaokang Yang 0001, Yongzheng Zhang 0002, Yipeng Wang 0001, Xiujuan Chai, Weiyao Lin |
Neurocomputing | 6 |
| 2014 | CovGa: A novel descriptor based on symmetry of regions for head pose estimation
Bingpeng Ma, Annan Li, Xiujuan Chai, Shiguang Shan |
Neurocomputing | 3 |
| 2014 | Maximal Likelihood Correspondence Estimation for Face Recognition Across PoseabstractDue to the misalignment of image features, the performance of many conventional face recognition methods degrades considerably in across pose scenario. To address this problem, many image matching-based methods are proposed to estimate semantic correspondence between faces in different poses. In this paper, we aim to solve two critical problems in previous image matching-based correspondence learning methods: 1) fail to fully exploit face specific structure information in correspondence estimation and 2) fail to learn personalized correspondence for each probe image. To this end, we first build a model, termed as morphable displacement field (MDF), to encode face specific structure information of semantic correspondence from a set of real samples of correspondences calculated from 3D face models. Then, we propose a maximal likelihood correspondence estimation (MLCE) method to learn personalized correspondence based on maximal likelihood frontal face assumption. After obtaining the semantic correspondence encoded in the learned displacement, we can synthesize virtual frontal images of the profile faces for subsequent recognition. Using linear discriminant analysis method with pixel-intensity features, state-of-the-art performance is achieved on three multipose benchmarks, i.e., CMU-PIE, FERET, and MultiPIE databases. Owe to the rational MDF regularization and the usage of novel maximal likelihood objective, the proposed MLCE method can reliably learn correspondence between faces in different poses even in complex wild environment, i.e., labeled face in the wild database. Shaoxin Li 0001, Xin Liu 0044, Xiujuan Chai, Haihong Zhang, Shihong Lao, Shiguang Shan |
IEEE Trans. Image Process. | 3 |
| 2013 | VisualComm: a tool to support communication between deaf and hearing persons with the KinectabstractWith the quickly increasing of the deaf community, how to communicate with the hearing persons is becoming a serious social problem. Furthermore, the investigation indicates that the deaf community is more self-enclosed and won't exchange ideas with the hearing. To address this challenge, we develop VisualComm, a tool to support communication between deaf and hearing persons with sign language recognition technology by using the Kinect. The main contribution of the system is a holistic solution of a two-way communication between deaf and hearings, and furthermore it is a seamless experience tailored for this particular activity. Currently we have implemented the basic communication based on 370 daily Chinese words for signer. Xiujuan Chai, Xilin Chen 0001, Ming Zhou 0001, Hanjing Li |
ASSETS | 1 |
| 2013 | Cascaded Shape Space Pruning for Robust Facial Landmark DetectionabstractIn this paper, we propose a novel cascaded face shape space pruning algorithm for robust facial landmark detection. Through progressively excluding the incorrect candidate shapes, our algorithm can accurately and efficiently achieve the globally optimal shape configuration. Specifically, individual landmark detectors are firstly applied to eliminate wrong candidates for each landmark. Then, the candidate shape space is further pruned by jointly removing incorrect shape configurations. To achieve this purpose, a discriminative structure classifier is designed to assess the candidate shape configurations. Based on the learned discriminative structure classifier, an efficient shape space pruning strategy is proposed to quickly reject most incorrect candidate shapes while preserve the true shape. The proposed algorithm is carefully evaluated on a large set of real world face images. In addition, comparison results on the publicly available BioID and LFW face databases demonstrate that our algorithm outperforms some state-of-the-art algorithms. Shiguang Shan, Xiujuan Chai, Xilin Chen 0001 |
ICCV | 3 |
| 2013 | A novel feature descriptor based on biologically inspired feature for head pose estimation
Bingpeng Ma, Xiujuan Chai, Tianjiang Wang |
Neurocomputing | 2 |
| 2012 | Locality-Constrained Active Appearance Model
Shiguang Shan, Xiujuan Chai, Xilin Chen 0001 |
ACCV (1) | 3 |
| 2012 | Morphable Displacement Field Based Image Matching for Face Recognition across Pose
Shaoxin Li 0001, Xin Liu 0044, Xiujuan Chai, Haihong Zhang, Shihong Lao, Shiguang Shan |
ECCV (1) | 3 |
| 2012 | Joint Face Alignment: Rescue Bad Alignments with Good Ones by Regularized Re-fitting
Xiujuan Chai, Shiguang Shan |
ECCV (2) | 2 |
| 2012 | Context modeling for facial landmark detection based on Non-Adjacent Rectangle (NAR) Haar-like feature
Xiujuan Chai, Zhiheng Niu, Cherkeng Heng, Shiguang Shan |
Image Vis. Comput. | 2 |
| 2011 | A novel coarse-to-fine hair segmentation methodabstractSegmenting hair regions from human images facilitates many tasks like hair synthesis and hair style trends forecast. However, hair segmentation is quite challenging due to hair/background confusion and large hair pattern diversity. To address these problems to some extent, this paper proposes a novel coarse-to-fine hair segmentation method. In our approach, firstly, the recently proposed “Active Segmentation with Fixation” (ASF) is used to coarsely define an enclosed candidate region with high-recall (but possibly low-precision) of hair pixels and exclude considerable part of the backgrounds which are easily confused with hair. Then Graph Cuts (GC) method is applied to the candidate regions to remove additional false positives by incorporating hair-specific information. Specifically, Bayesian method is employed to select some reliable hair and background regions (seeds) among the ones over-segmented by Mean Shift. SVM classifier is then learnt online from these seeds and explored to predict hair/background likelihood probability, which is subsequently fed into GC algorithm. The novelty of the proposed approach lies in three folds: 1) an elaborate design of hair segmentation framework, which utilizes ASF to reduce the candidate hair regions and adopts GC to achieve more accurate hair region contours; 2) the region-based strategy for seed selection; 3) the exploration of the discriminative method, SVM, to predict the probability of each pixel belonging to hair and background regions. Extensive experimental results demonstrate the approach outperforms recently proposed methods. Xiujuan Chai, Hongming Zhang 0011, Hong Chang 0001, Wei Zeng 0006, Shiguang Shan |
FG | 2 |
| 2011 | Context constrained facial landmark localization based on discontinuous Haar-like featureabstractAbstract—Facial landmark localization is well known as one of the bottlenecks in face recognition. This paper proposes a novel facial landmark localization method, which introduces facial context constrains into cascaded AdaBoost framework. The motivation of our method lies in the basic human physiology observation that not only the local texture information but also the global context information is used together for human to realize the landmark location task. Therefore, in our solution, a novel type of Haar-like feature, called discontinuous Haar-like feature, is proposed to characterize the facial context, i.e. the cooccurrence relationship between target facial landmark and other local texture patterns within face region (including other landmarks, facial organs and also smoothing regions). For the locating task, traditional Haar-like features (characterizing local texture information) and discontinuous Haar-like features (characterizing context constrains in global sense) are combined together to form more powerful representations. Through Real AdaBoost learning, distinctive features are selected automatically and used for facial landmark detection. Our experiments on BioID and Cohn-Kanade databases have validated the proposed method by comparing with other state-of-the-art results. Keywords-face recognition; facial landmark localization; context constraints; discontinuous Haar-like feature I. Xiujuan Chai, Zhiheng Niu, Cherkeng Heng, Shiguang Shan |
FG | 2 |
| 2011 | Local Regression Model for Automatic Face Sketch GenerationabstractAs one of the important artistic styles of portrait, sketch portrait has wide applications for both digital entertainment and law enforcement. In this paper, an automatic face sketch generation approach is presented by learning from photo-sketch pair examples. Specifically, the relationship between a face photo and its corresponding face sketch is learned on image patch level. By applying this relationship to the input face photo patch, we can infer the output face sketch patch by exploiting some regression techniques such as kNN, the Lasso and so on. Via our local regression model, we can synthesize an appealing sketch portrait from a given face photo in a few minutes. Experiments conducted on CUHK database have shown that our results are more compelling than previous methods especially in two respects: (1) our synthesized sketches preserve more identity information of the original face photo, (2) our synthesized sketches presents more pencil sketch texture. Naye Ji, Xiujuan Chai, Shiguang Shan, Xilin Chen 0001 |
ICIG | 2 |
| 2010 | An incremental Bhattacharyya dissimilarity measure for particle filtering
Anbang Yao, Guijin Wang, Xinggang Lin, Xiujuan Chai |
Pattern Recognit. | 4 |
| 2009 | A fast and effective outlier detection method for matching uncalibrated imagesabstractMany image analysis tasks require an outlier detection procedure to identify the false matches. In this paper, a fast and effective outlier detection method is presented to match images in the uncalibrated case. This method employs a hypothesis test on the consistency of dominant orientations of the feature points to significantly increase the detection speed. Moreover, it can also effectively find the outliers that can not be identified by traditional RANSAC-based methods using epipolar constraint. Note that our method does not require the prior knowledge of camera parameters or the percentage of outliers. The experimental results show that our method outperforms the classical RANSAC-based methods both in speed and in accuracy of the results. Xiujuan Chai, Shiming Ge |
ICIP | 3 |
| 2009 | Robust hand gesture analysis and application in gallery browsingabstractThis paper presents a robust hand gesture analysis method using 3D depth data. Our scheme focuses on accurate hand segmentation by eliminating the negative effect of the forearm part. In the general human computer interaction (HCI) tasks, such an assumption usually holds that the depth of hand is smaller than forearm. Therefore, the precise hand region can be obtained through the fusion of the hand geometric features and the 3D depth information in real-time. Moreover, a robust hand gesture recognition method, which combines the global structure information and the local texture variation, is included in our gesture analysis framework. The elaborate hand segmentation makes the succedent recognition problem much easier and gets more accurate recognition results. Experimental results convincingly show the effectiveness of the proposed gesture analysis strategy. Furthermore, a concrete application scenario, gesture controlled picture gallery browsing, is implemented successfully. Xiujuan Chai, Yikai Fang, Kongqiao Wang |
ICME | 1 |
| 2009 | Hand posture recognition in video using multiple cuesabstractHand posture conveys profound information for computer vision applications, but the articulated hand structure and restraint capture condition cast a tough obstacle on practical implementation, especially in real time video. This paper presents a framework to recognize hand postures in consecutive video frames. Mixture of Gaussian skin/non skin models is constructed for hand region detection, followed by particle filter to track hand. Then a soft-decision scheme based on extended Histogram of Orientated gradient is proposed to refine the best posture region and recognize it from pre-defined posture set. Experimental result shows promising performance under various capture conditions. Liang Sha, Guijin Wang, Anbang Yao, Xinggang Lin, Xiujuan Chai |
ICME | 5 |
| 2008 | Recovering 3D facial shape via coupled 2D/3D space learningabstractThis paper presents a method for recovering 3D facial shape from single image via learning the relationship between the 2D intensity images and the 3D facial shapes. With a coupled training set, the intensity images and their corresponding facial shapes make up two vector spaces respectively. But only the correlated components in both spaces are useful for inference, so there must be embedded hidden subspaces in each space which preserve the inter-space correlation information. Thus by learning the projection onto hidden subspaces based on maximum correlation criteria and optimizing the linear transform between the hidden spaces, 3D facial shape is inferred from the intensity image. The effectiveness of the method is demonstrated on both synthesized and real world data. Annan Li, Shiguang Shan, Xilin Chen 0001, Xiujuan Chai, Wen Gao 0001 |
FG | 4 |
| 2007 | Locally Linear Regression for Pose-Invariant Face RecognitionabstractThe variation of facial appearance due to the viewpoint (/pose) degrades face recognition systems considerably, which is one of the bottlenecks in face recognition. One of the possible solutions is generating virtual frontal view from any given nonfrontal view to obtain a virtual gallery/probe face. Following this idea, this paper proposes a simple, but efficient, novel locally linear regression (LLR) method, which generates the virtual frontal view from a given nonfrontal face image. We first justify the basic assumption of the paper that there exists an approximate linear mapping between a nonfrontal face image and its frontal counterpart. Then, by formulating the estimation of the linear mapping as a prediction problem, we present the regression-based solution, i.e., globally linear regression. To improve the prediction accuracy in the case of coarse alignment, LLR is further proposed. In LLR, we first perform dense sampling in the nonfrontal face image to obtain many overlapped local patches. Then, the linear regression technique is applied to each small patch for the prediction of its virtual frontal patch. Through the combination of all these patches, the virtual frontal view is generated. The experimental results on the CMU PIE database show distinct advantage of the proposed method over Eigen light-field method. Xiujuan Chai, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2003 | Novel example-based shape learning for fast face alignmentabstractA novel example-based shape learning (ESL) strategy is proposed for facial feature alignment. The method is motivated by an intuitive and experimental observation that there exists an approximate linearity relationship between the image difference and the shape difference, that is, similar face images imply similar face shapes. Therefore, given a learning set of face images with their corresponding face landmarks labeled, the shape of any novel face image can be learned by estimating its similarities to the training images in the learning set and applying these similarities to the shape reconstruction of the novel face image. Concretely, if the novel face image is expressed by an optimal linear combination of the training images, the same linear combination coefficients can be directly applied to the linear combination of the training shapes to construct the optimal shape for the novel face image. Our experiments have convincingly shown the effectiveness and efficiency of the proposed approach in both speed and accuracy performance compared with other methods. Xiujuan Chai, Shiguang Shan, Wen Gao 0001 |
ICASSP (3) | 1 |
| 2003 | Virtual face image generation for illumination and pose insensitive face recognitionabstractFace recognition has attracted much attention in the past decades for its wide potential applications. Much progress has been made in the past few years. However, specialized evaluation of the state-of-the-art of both academic algorithms and commercial systems illustrates that the performance of most current recognition technologies degrades significantly due to the variations of illumination and/or pose. To solve these problems, providing multiple training samples to the recognition system is a rational choice. However, enough samples are not always available for many practical applications. It is an alternative to augment the training set by generating virtual views from one single face image, that is, relighting the given face images or synthesize novel views of the given face. Based on this strategy, this paper presents some attempts by presenting a ratio-image based face relighting method and a face re-rotating approach based on linear shape prediction and image warp. To evaluate the effect of the additional virtual face images, primary experiments are conducted using our face specific subspace method as face recognition approach, which shows impressive improvement compared with standard benchmark face recognition methods. Wen Gao 0001, Shiguang Shan, Xiujuan Chai, Xiaowei Fu |
ICASSP (4) | 3 |
| 2003 | Novel example-based shape learning for fast face alignmentabstractIn this paper, a novel example-based shape learning (ESL) strategy has been proposed for facial feature alignment. The method is motivated by an intuitive and experimental observation that there exists an approximate linearity relationship between the image difference and the shape difference, that is, similar face images imply similar face shapes. Therefore, given a learning set of face images with their corresponding face landmarks labeled, the shape of any novel face image can be learned by estimating its similarities to the training images in the learning set and applying these similarities to the shape reconstruction of a novel face image. Concretely, if the novel face image is expressed by an optimal linear combination of the training images, the same linear combination coefficients can be directly applied to the linear combination of the training shapes to construct the optimal shape for the novel face image. Our experiments have convincingly shown the effectiveness and efficiency of the proposed approach in both speed and accuracy performance compared with other methods. Xiujuan Chai, Shiguang Shan, Wen Gao 0001 |
ICME | 1 |
| 2003 | Virtual face image generation for illumination and pose insensitive face recognitionabstractFace recognition has attracted much attention in the past decades for its wide potential applications. Much progress has been made in the past few years. However, specialized evaluation of the state-of-the-art in both academic algorithms and commercial systems illustrates that the performance of most current recognition technologies degrades significantly due to the variations of illumination and/or pose. To solve these problems, providing multiple training samples to the recognition system is a rational choice. However, enough samples are not always available for many practical applications. It is an alternative to augment the training set by generating virtual views from one single face image, that is relighting the given face images of synthesize novel views of the given face. Based on this strategy, this paper presents some attempts by presenting a ratio-image based face relighting method and a face re-rotating approach based on linear shape prediction and image warp. To evaluate the effect of the additional virtual face images, primary experiments are conducted using our specific substance method as face recognition approach, which shows impressive improvement compared with standard benchmark face recognition methods. Wen Gao 0001, Shiguang Shan, Xiujuan Chai, Xiaowei Fu |
ICME | 3 |