VLDB 2026 Research / reviewers in the wild / expert
Meina Kan
dblp:21/9606
· DBLP profile ↗
62ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0001-9483-875XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 9 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 8 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Patching the visual ability of large multimodal models by collaborating with small models
Meina Kan, Shiguang Shan, Xilin Chen 0001 |
Frontiers Comput. Sci. | 3 |
| 2025 | Benchmarking Multimodal Large Language Models Against Image Corruptions
Xinkuan Qiu, Meina Kan, Yongbin Zhou, Shiguang Shan |
ICCV | 2 |
| 2025 | Feature Decomposition-Recomposition in Large Vision-Language Model for Few-Shot Class-Incremental Learning
Zongyao Xue, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2025 | Task-Oriented Token Pruning for Efficient Object Detection and SegmentationabstractRobots rely heavily on visual perception to understand and interact with complex environments. To support this capability, modern perception models have become increasingly large and powerful, resulting in high computational costs that hinder their real-time performance in robotic applications. Existing acceleration techniques, such as model pruning and token pruning, focus on reducing architectural or parameter redundancy but still process all object categories, regardless of task requirements. However, in real-world robotic scenarios, different tasks typically require only a subset of object categories. For instance, a service robot may focus on kitchenware while cooking, but shift to furniture and obstacles while cleaning. This task-dependent variation creates opportunities to reduce computational cost by selectively processing relevant information. Existing methods are not designed to exploit this potential for task-specific efficiency. To address this limitation, we propose TaskTP, a task-oriented token pruning method that dynamically adjusts token pruning based on the target category set. A dynamic gating network is introduced between successive Transformer blocks, which evaluates the relevance of each token to the given task. TaskTP allows for more aggressive pruning when fewer categories are required, optimizing computation without sacrificing performance. After a task-agnostic training phase, it can be flexibly configured at deployment time to support any category subset without retraining, making it both efficient and versatile. TaskTP improves the performance of Mask R-CNN from 31.4 fps to 38.5 fps on the COCO dataset. Furthermore, on the ScanNet dataset, where an object search task was defined to simulate real-world robotic applications, processing time was reduced from 3197 ms to 2437 ms, demonstrating significant efficiency gains. Meina Kan, Shiguang Shan, Xilin Chen 0001 |
IROS | 2 |
| 2025 | Precise Integral in NeRFs: Overcoming the Approximation Errors of Numerical QuadratureabstractNeural Radiance Fields (NeRFs) use neural networks to translate spatial coordinates to corresponding volume density and directional radiance, enabling realistic novel view synthesis through volume rendering. Rendering new viewpoints involves computing volume rendering integrals along rays, usually approximated by numerical quadrature because of lacking closed-form solutions. In this paper, utilizing Taylor expansion, we demonstrate that numerical quadrature causes inevitable approximation error in NeRF integrals due to ignoring the parameter associated with the Lagrange remainder. To mitigate the approximation error, we propose a novel neural field with segment representation as input to implicitly model the remainder parameter. In theory, our proposed method is proven to possess the potential to achieve fully precise rendering integral, as demonstrated by comprehensive experiments on several commonly used datasets with state-of-the-art results. Zhenliang He, Meina Kan, Shiguang Shan |
WACV | 3 |
| 2025 | eLabrador: A Wearable Navigation System for Visually Impaired IndividualsabstractVisually impaired individuals encounter significant challenges when walking and acting in unfamiliar environments, particularly in outdoor scenarios. The complexity of outdoor environments, characterized by diverse obstacles, traffic signals, and societal norms, poses substantial barriers to mobility of visually impaired individuals and makes long-distance walking especially arduous. Although GPS-based navigation systems can facilitate long-distance travel, they often suffer from location inaccuracies in urban areas and even completely fail indoors. Moreover, these systems lack the capability to provide detailed information about walkways and immediate surroundings, which are crucial for safe and efficient walking. To address these limitations, we introduce a proof-of-concept wearable navigation system named eLabrador, designed to assist visually impaired individuals in long-distance walking in unfamiliar outdoor environments. The eLabrador integrates public maps (e.g. Amap or Google Maps) and GPS for global route planning, while leveraging computational visual perception to provide precise and safe local guidance. This hybrid approach enables accurate and safe navigation for visually impaired individuals in outdoor scenarios. Specifically, the eLabrador utilizes a head-mounted RGB-D camera to capture environmental geometric terrain and objects in outdoor urban environments. These inputs are processed into a 3D semantic map, offering a detailed representation of the surrounding environment. The planning module then integrates this 3D semantic map with route information from the global map (i.e. Amap) to generate an optimized walking path. Finally, the interaction module utilizes the audio-haptic dual-channel to relay navigation instructions to visually impaired user. Together, these three modules work seamlessly to facilitate long-distance navigation for visually impaired individuals in outdoor environments. The eLabrador is evaluated with two real-world outdoor scenarios, involving 10 visually impaired and visually masked participants. The experiments show that eLabrador successfully guides visually impaired participants to their destinations in outdoor environments. Additionally, the eLabrador provides descriptive information about landmarks and other navigation cues, helping visually impaired users better understand their surroundings. Subjective evaluations further indicate that most participants felt a sense of safety and reported an acceptable cognitive load during navigation, indicating its usability and effectiveness. Note to Practitioners—Visually impaired individuals almost cannot walk long distance in unfamiliar outdoor environments. Without proper assistance, their mobility and quality of life can be severely impacted. To address this issue, this article presents a wearable navigation system eLabrador to assist visually impaired individuals in walking outdoors, such as traveling from a residential entrance to a nearby park. Experimental results from real-world scenarios involving 10 participants demonstrate that eLabrador safely guides visually impaired users to their destination, significantly enhancing their mobility and independence. Meina Kan, Lixuan Zhang, Minxue Fang, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | HPNet: Dynamic Trajectory Forecasting with Historical Prediction AttentionabstractPredicting the trajectories of road agents is essential for autonomous driving systems. The recent mainstream methods follow a static paradigm, which predicts the future trajectory by using a fixed duration of historical frames. These methods make the predictions independently even at adjacent time steps, which leads to potential instability and temporal inconsistency. As successive time steps have largely overlapping historical frames, their forecasting should have intrinsic correlation, such as overlapping predicted trajectories should be consistent, or be different but share the same motion goal depending on the road situation. Motivated by this, in this work, we introduce HPNet, a novel dynamic trajectory forecasting method. Aiming for stable and accurate trajectory forecasting, our method leverages not only historical frames including maps and agent states, but also historical predictions. Specifically, we newly design a Historical Prediction Attention module to automatically encode the dynamic relationship between successive predictions. Besides, it also extends the attention range beyond the currently visible window benefitting from the use of historical predictions. The proposed Historical Prediction Attention together with the Agent Attention and Mode Attention is further formulated as the Triple Factorized Attention module, serving as the core design of HPNet. Experiments on the Argoverse and INTERACTION datasets show that HP-Net achieves state-of-the-art performance, and generates accurate and stable future trajectories. Our code are available at https://github.com/XiaolongTang23/HPNet. Meina Kan, Shiguang Shan, Zhilong Ji, Jinfeng Bai, Xilin Chen 0001 |
CVPR | 2 |
| 2024 | PreLAR: World Model Pre-training with Learnable Action Representation
Lixuan Zhang, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ECCV (23) | 2 |
| 2024 | A Simple Romance Between Multi-Exit Vision Transformer and Token ReductionabstractVision Transformers (ViTs) are now flourishing in the computer vision area. Despite the remarkable success, ViTs suffer from high computational costs, which greatly hinder their practical usage. Token reduction, which identifies and discards unimportant tokens during forward propagation, has then been proposed to make ViTs more efficient. For token reduction methodologies, a scoring metric is essential to distinguish between important and unimportant tokens. The attention score from the $\mathrm{[CLS]}$ token, which takes the responsibility to aggregate useful information and form the final output, has been established by prior works as an advantageous choice. Nevertheless, whereas the task pressure is applied at the end of the whole model, token reduction generally starts from very early blocks. Given the long distance in between, in the early blocks, $\mathrm{[CLS]}$ token lacks the impetus to gather task-relevant information, causing somewhat arbitrary attention allocation. This phenomenon, in turn, degrades the reliability of token scoring and substantially compromises the effectiveness of token reduction. Inspired by advances in the domain of dynamic neural networks, in this paper, we introduce Multi-Exit Token Reduction (METR), a simple romance between multi-exit architecture and token reduction—two areas previously considered orthogonal. By injecting early task pressure via multi-exit loss, the $\mathrm{[CLS]}$ token is spurred to collect task-related information in even early blocks, thus bolstering the credibility of $\mathrm{[CLS]}$ attention as a token-scoring metric. Additionally, we employ self-distillation to further refine the quality of early supervision. Extensive experiments substantiate both the existence and effectiveness of the newfound chemistry. Comparative assessments also indicate that METR outperforms state-of-the-art token reduction methods on standard benchmarks, especially under aggressive reduction ratios. Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ICLR | 2 |
| 2024 | Collaborative Domain Alignment for Multi-source Domain Adaptation
Meina Kan, Zhilong Ji, Jinfeng Bai, Shiguang Shan, Xilin Chen 0001 |
ICPR (27) | 2 |
| 2024 | Shape-biased CNNs are Not Always Superior in Out-of-Distribution RobustnessabstractIn recent years, Out-of-Distribution (o.o.d) Robustness has garnered increasing attention in Deep Learning, and shape-biased Convolutional Neural Networks (CNNs) are believed to exhibit higher robustness, attributed to the inherent shape-based decision rule of human cognition. In this work, we delve deeper into the intricate relationship between shape/texture information and o.o.d robustness by leveraging a carefully curated "Category-Balanced ImageNet" dataset. We find that shape information is not always superior in distinguishing distinct categories and shape-biased model is not always superior across various o.o.d scenarios. Motivated by these insightful findings, we design a novel method named Shape-Texture Adaptive Recombination (STAR) to achieve higher o.o.d robustness. A category-balanced dataset is firstly used to pretrain a debiased backbone and three specialized heads, each adept at robustly extracting shape, texture, and debiased features. Subsequently, an instance-adaptive recombination head is trained to adaptively adjust the contributions of these distinctive features for each given instance. Through comprehensive experiments, our proposed method achieves state-of-the-art o.o.d robustness across various scenarios such as image corruptions, adversarial attacks, style shifts, and dataset shifts, demonstrating its effectiveness. Xinkuan Qiu, Meina Kan, Yongbin Zhou, Yanchao Bi, Shiguang Shan |
WACV | 2 |
| 2023 | DandelionNet: Domain Composition with Instance Adaptive Classification for Domain GeneralizationabstractDomain generalization (DG) attempts to learn a model on source domains that can well generalize to unseen but different domains. The multiple source domains are innately different in distribution but intrinsically related to each other, e.g., from the same label space. To achieve a generalizable feature, most existing methods attempt to reduce the domain discrepancy by either learning domain-invariant feature, or additionally mining domain-specific feature. In the space of these features, the multiple source domains are either tightly aligned or not aligned at all, which both cannot fully take the advantage of complementary information from multiple domains. In order to preserve more complementary information from multiple domains at the meantime of reducing their domain gap, we propose that the multiple domains should not be tightly aligned but composite together, where all domains are pulled closer but still preserve their individuality respectively. This is achieved by using instance-adaptive classifier specified for each instance’s classification, where the instance-adaptive classifier is slightly deviated from a universal classifier shared by samples from all domains. This adaptive classifier deviation allows all instances from the same category but different domains to be dispersed around the class center rather than squeezed tightly, leading to better generalization for unseen domain samples. In result, the multiple domains are harmoniously composite centered on a universal core, like a dandelion, so this work is referred to as DandelionNet. Experiments on multiple DG benchmarks demonstrate that the proposed method can learn a model with better generalization and experiments on source free domain adaption also indicate the versatility. Lanqing Hu, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2023 | Function-Consistent Feature Distillation
Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ICLR | 2 |
| 2023 | BLPSeg: Balance the Label Preference in Scribble-Supervised Semantic SegmentationabstractScribble-supervised semantic segmentation is an appealing weakly supervised technique with low labeling cost. Existing approaches mainly consider diffusing the labeled region of scribble by low-level feature similarity to narrow the supervision gap between scribble labels and mask labels. In this study, we observe an annotation bias between scribble and object mask, i.e., label workers tend to scribble on the spacious region instead of corners. This label preference makes the model learn well on those frequently labeled regions but poor on rarely labeled pixels. Therefore, we propose BLPSeg to balance the label preference for complete segmentation. Specifically, the BLPSeg first predicts an annotation probability map to evaluate the rarity of labels on each image, then utilizes a novel BLP loss to balance the model training by up-weighting those rare annotations. Additionally, to further alleviate the impact of label preference, we design a local aggregation module (LAM) to propagate supervision from labeled to unlabeled regions in gradient backpropagation. We conduct extensive experiments to illustrate the effectiveness of our BLPSeg. Our single-stage method even outperforms other advanced multi-stage methods and achieves state-of-the-art performance. Yude Wang, Jie Zhang 0071, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | GAN with Multivariate Disentangling for Controllable Hair Editing
Meina Kan, Shiguang Shan |
ECCV (15) | 2 |
| 2022 | Mutual Learning of Joint and Separate Domain Alignments for Multi-Source Domain AdaptationabstractMulti-Source Domain Adaptation (MSDA) aims at transferring knowledge from multiple labeled source domains to benefit the task in an unlabeled target domain. The challenges of MSDA lie in mitigating domain gaps and combining information from diverse source domains. In most existing methods, the multiple source domains can be jointly or separately aligned to the target domain. In this work, we consider that these two types of methods, i.e. joint and separate domain alignments, are complementary and propose a mutual learning based alignment network (MLAN) to combine their advantages. Specifically, our proposed method is composed of three components, i.e. a joint alignment branch, a separate alignment branch, and a mutual learning objective between them. In the joint alignment branch, the samples from all source domains and the target domain are aligned together, with a single domain alignment goal, while in the separate alignment branch, each source domain is individually aligned to the target domain. Finally, by taking advantage of the complementarity of joint and separate domain alignment mechanisms, mutual learning is used to make the two branches learn collaboratively. Compared with other existing methods, our proposed MLAN integrates information of different domain alignment mechanisms and thus can mine rich knowledge from multiple domains for better performance. The experiments on Domain-Net, Office-31, and Digits-five datasets demonstrate the effectiveness of our method. Meina Kan, Shiguang Shan, Xilin Chen 0001 |
WACV | 2 |
| 2022 | Personalized Convolution for Face Recognition
Chunrui Han, Shiguang Shan, Meina Kan, Shuzhe Wu, Xilin Chen 0001 |
Int. J. Comput. Vis. | 3 |
| 2022 | Learning pseudo labels for semi-and-weakly supervised semantic segmentation
Yude Wang, Jie Zhang 0071, Meina Kan, Shiguang Shan |
Pattern Recognit. | 3 |
| 2021 | EigenGAN: Layer-Wise Eigen-Learning for GANsabstractRecent studies on Generative Adversarial Network (GAN) reveal that different layers of a generative CNN hold different semantics of the synthesized images. However, few GAN models have explicit dimensions to control the semantic attributes represented in a specific layer. This paper proposes EigenGAN which is able to unsupervisedly mine interpretable and controllable dimensions from different generator layers. Specifically, EigenGAN embeds one linear subspace with orthogonal basis into each generator layer. Via generative adversarial training to learn a target distribution, these layer-wise subspaces automatically discover a set of "eigen-dimensions" at each layer corresponding to a set of semantic attributes or interpretable variations. By traversing the coefficient of a specific eigen-dimension, the generator can produce samples with continuous changes corresponding to a specific semantic attribute. Taking the human face for example, EigenGAN can discover controllable dimensions for high-level concepts such as pose and gender in the subspace of deep layers, as well as low-level concepts such as hue and color in the subspace of shallow layers. Moreover, in the linear case, we theoretically prove that our algorithm derives the principal components as PCA does. Codes can be found in https://github.com/LynnHo/EigenGAN-Tensorflow. Zhenliang He, Meina Kan, Shiguang Shan |
ICCV | 2 |
| 2021 | Image style disentangling for instance-level facial attribute transfer
Meina Kan, Zhenliang He, Xingguang Song, Shiguang Shan |
Comput. Vis. Image Underst. | 2 |
| 2021 | Learning to Learn Adaptive Classifier-Predictor for Few-Shot LearningabstractFew-shot learning aims to learn a well-performing model from a few labeled examples. Recently, quite a few works propose to learn a predictor to directly generate model parameter weights with episodic training strategy of meta-learning and achieve fairly promising performance. However, the predictor in these works is task-agnostic, which means that the predictor cannot adjust to novel tasks in the testing phase. In this article, we propose a novel meta-learning method to learn how to learn task-adaptive classifier-predictor to generate classifier weights for few-shot classification. Specifically, a meta classifier-predictor module, (MPM) is introduced to learn how to adaptively update a task-agnostic classifier-predictor to a task-specialized one on a novel task with a newly proposed center-uniqueness loss function. Compared with previous works, our task-adaptive classifier-predictor can better capture characteristics of each category in a novel task and thus generate a more accurate and effective classifier. Our method is evaluated on two commonly used benchmarks for few-shot classification, i.e., miniImageNet and tieredImageNet. Ablation study verifies the necessity of learning task-adaptive classifier-predictor and the effectiveness of our newly proposed center-uniqueness loss. Moreover, our method achieves the state-of-the-art performance on both benchmarks, thus demonstrating its superiority. Nan Lai, Meina Kan, Chunrui Han, Xingguang Song, Shiguang Shan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Corrections to "Learning to Learn Adaptive Classifier-Predictor for Few-Shot Learning"abstractIn the above article [1], the results of "Fully-supervised (Upper bound)" in Tables III and IV were inadvertently set to intermediate records that were used as placeholders. This error has no effect on any of the interpretations and conclusions. Tables I and II of this amendment show the corrected results (highlighted in italics) of the original Tables III and IV. Nan Lai, Meina Kan, Chunrui Han, Xingguang Song, Shiguang Shan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Unsupervised Domain Adaptation With Hierarchical Gradient SynchronizationabstractDomain adaptation attempts to boost the performance on a target domain by borrowing knowledge from a well established source domain. To handle the distribution gap between two domains, the prominent approaches endeavor to extract domain-invariant features. It is known that after a perfect domain alignment the domain-invariant representations of two domains should share the same characteristics from perspective of the overview and also any local piece. Inspired by this, we propose a novel method called Hierarchical Gradient Synchronization to model the synchronization relationship among the local distribution pieces and global distribution, aiming for more precise domain-invariant features. Specifically, the hierarchical domain alignments including class-wise alignment, group-wise alignment and global alignment are first constructed. Then, these three types of alignment are constrained to be consistent to ensure better structure preservation. As a result, the obtained features are domain invariant and intrinsically structure preserved. As evaluated on extensive domain adaptation tasks, our proposed method achieves state-of-the-art classification performance on both vanilla unsupervised domain adaptation and partial domain adaptation. Lanqing Hu, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2020 | Self-Supervised Equivariant Attention Mechanism for Weakly Supervised Semantic SegmentationabstractImage-level weakly supervised semantic segmentation is a challenging problem that has been deeply studied in recent years. Most of advanced solutions exploit class activation map (CAM). However, CAMs can hardly serve as the object mask due to the gap between full and weak supervisions. In this paper, we propose a self-supervised equivariant attention mechanism (SEAM) to discover additional supervision and narrow the gap. Our method is based on the observation that equivariance is an implicit constraint in fully supervised semantic segmentation, whose pixel-level labels take the same spatial transformation as the input images during data augmentation. However, this constraint is lost on the CAMs trained by image-level supervision. Therefore, we propose consistency regularization on predicted CAMs from various transformed images to provide self-supervision for network learning. Moreover, we propose a pixel correlation module (PCM), which exploits context appearance information and refines the prediction of current pixel by its similar neighbors, leading to further improvement on CAMs consistency. Extensive experiments on PASCAL VOC 2012 dataset demonstrate our method outperforms state-of-the-art methods using the same level of supervision. The code is released online. Yude Wang, Jie Zhang 0071, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
CVPR | 3 |
| 2020 | Deformable face net for pose invariant face recognition
Jie Zhang 0071, Shiguang Shan, Meina Kan, Xilin Chen 0001 |
Pattern Recognit. | 4 |
| 2020 | Learning deep face representation with long-tail data: An aggregate-and-disperse approach
Meina Kan, Shiguang Shan, Xilin Chen 0001 |
Pattern Recognit. Lett. | 2 |
| 2019 | Fully Learnable Group Convolution for Acceleration of Deep Neural NetworksabstractBenefitted from its great success on many tasks, deep learning is increasingly used on low-computational-cost devices, e.g. smartphone, embedded devices, etc. To reduce the high computational and memory cost, in this work, we propose a fully learnable group convolution module (FLGC for short) which is quite efficient and can be embedded into any deep neural networks for acceleration. Specifically, our proposed method automatically learns the group structure in the training stage in a fully end-to-end manner, leading to a better structure than the existing pre-defined, two-steps, or iterative strategies. Moreover, our method can be further combined with depthwise separable convolution, resulting in 5 times acceleration than the vanilla Resnet50 on single CPU. An additional advantage is that in our FLGC the number of groups can be set as any value, but not necessarily 2^k as in most existing methods, meaning better tradeoff between accuracy and speed. As evaluated in our experiments, our method achieves better performance than existing learnable group convolution and standard group convolution when using the same number of groups. Xijun Wang 0002, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2019 | Deformable Face Net: Learning Pose Invariant Feature with Pose Aware Feature Alignment for Face RecognitionabstractFace recognition plays an important role in computer vision. It still remains a challenging task due to pose, expression, illumination, partial occlusion, etc. In this work, we propose a novel Deformable Face Net (DFN) to handle the pose variations in face recognition. The Deformable Face Net introduces deformable convolution modules to simultaneously learn face recognition oriented alignment and feature extraction. Specifically, two loss functions, namely displacement consistency loss (DCL) and identity consistency loss (ICL) are designed to minimize the intra-class feature variation caused by different poses. These two loss functions jointly learn pose-aware displacement fields for deformable convolutions in the DFN. Different from the existing methods, the DFN focuses on aligning features across different poses rather than frontalizing the input faces. Extensive experiments show that the proposed DFN outperforms the state-of-the-art methods, especially on the datasets with large poses. Jie Zhang 0071, Shiguang Shan, Meina Kan, Xilin Chen 0001 |
FG | 4 |
| 2019 | S2GAN: Share Aging Factors Across Ages and Share Aging Trends Among IndividualsabstractGenerally, we human follow the roughly common aging trends, e.g., the wrinkles only tend to be more, longer or deeper. However, the aging process of each individual is more dominated by his/her personalized factors, including the invariant factors such as identity and mole, as well as the personalized aging patterns, e.g., one may age by graying hair while another may age by receding hairline. Following this biological principle, in this work, we propose an effective and efficient method to simulate natural aging. Specifically, a personalized aging basis is established for each individual to depict his/her own aging factors. Then different ages share this basis, being derived through age-specific transforms. The age-specific transforms represent the aging trends which are shared among all individuals. The proposed method can achieve continuous face aging with favorable aging accuracy, identity preservation, and fidelity. Furthermore, befitted from the effective design, a unique model is capable of all ages and the prediction time is significantly saved. Zhenliang He, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2019 | Weakly Supervised Object Detection With Segmentation CollaborationabstractWeakly supervised object detection aims at learning precise object detectors, given image category labels. In recent prevailing works, this problem is generally formulated as a multiple instance learning module guided by an image classification loss. The object bounding box is assumed to be the one contributing most to the classification among all proposals. However, the region contributing most is also likely to be a crucial part or the supporting context of an object. To obtain a more accurate detector, in this work we propose a novel end-to-end weakly supervised detection approach, where a newly introduced generative adversarial segmentation module interacts with the conventional detection module in a collaborative loop. The collaboration mechanism takes full advantages of the complementary interpretations of the weakly supervised localization task, namely detection and segmentation tasks, forming a more comprehensive solution. Consequently, our method obtains more precise object bounding boxes, rather than parts or irrelevant surroundings. Expectedly, the proposed method achieves an accuracy of 53.7% on the PASCAL VOC 2007 dataset, outperforming the state-of-the-arts and demonstrating its superiority for weakly supervised object detection. Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2019 | Locality-constrained framework for face alignment
Jie Zhang 0071, Meina Kan, Shiguang Shan, Xiujuan Chai, Xilin Chen 0001 |
Frontiers Comput. Sci. | 3 |
| 2019 | Hierarchical Attention for Part-Aware Face Detection
Shuzhe Wu, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
Int. J. Comput. Vis. | 2 |
| 2019 | AttGAN: Facial Attribute Editing by Only Changing What You WantabstractFacial attribute editing aims to manipulate single or multiple attributes on a given face image, i.e., to generate a new face image with desired attributes while preserving other details. Recently, the generative adversarial net (GAN) and encoder-decoder architecture are usually incorporated to handle this task with promising results. Based on the encoder-decoder architecture, facial attribute editing is achieved by decoding the latent representation of a given face conditioned on the desired attributes. Some existing methods attempt to establish an attribute-independent latent representation for further attribute editing. However, such attribute-independent constraint on the latent representation is excessive because it restricts the capacity of the latent representation and may result in information loss, leading to over-smooth or distorted generation. Instead of imposing constraints on the latent representation, in this work, we propose to apply an attribute classification constraint to the generated image to just guarantee the correct change of desired attributes, i.e., to change what you want. Meanwhile, the reconstruction learning is introduced to preserve attribute-excluding details, in other words, to only change what you want. Besides, the adversarial learning is employed for visually realistic editing. These three components cooperate with each other forming an effective framework for high quality facial attribute editing, referred as AttGAN. Furthermore, the proposed method is extended for attribute style manipulation in an unsupervised manner. Experiments on two wild datasets, CelebA and LFW, show that the proposed method outperforms the state-of-the-art on realistic attribute editing with other facial details well preserved. Zhenliang He, Wangmeng Zuo, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Task-Adaptive Feature Reweighting for Few Shot Classification
Nan Lai, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ACCV (4) | 2 |
| 2018 | Duplex Generative Adversarial Network for Unsupervised Domain AdaptationabstractDomain adaptation attempts to transfer the knowledge obtained from the source domain to the target domain, i.e., the domain where the testing data are. The main challenge lies in the distribution discrepancy between source and target domain. Most existing works endeavor to learn domain invariant representation usually by minimizing a distribution distance, e.g., MMD and the discriminator in the recently proposed generative adversarial network (GAN). Following the similar idea of GAN, this work proposes a novel GAN architecture with duplex adversarial discriminators (referred to as DupGAN), which can achieve domain-invariant representation and domain transformation. Specifically, our proposed network consists of three parts, an encoder, a generator and two discriminators. The encoder embeds samples from both domains into the latent representation, and the generator decodes the latent representation to both source and target domains respectively conditioned on a domain code, i.e., achieves domain transformation. The generator is pitted against duplex discriminators, one for source domain and the other for target, to ensure the reality of domain transformation, the latent representation domain invariant and the category information of it preserved as well. Our proposed work achieves the state-of-the-art performance on unsupervised domain adaptation of digit classification and object recognition. Lanqing Hu, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2018 | Real-Time Rotation-Invariant Face Detection With Progressive Calibration NetworksabstractRotation-invariant face detection, i.e. detecting faces with arbitrary rotation-in-plane (RIP) angles, is widely required in unconstrained applications but still remains as a challenging task, due to the large variations of face appearances. Most existing methods compromise with speed or accuracy to handle the large RIP variations. To address this problem more efficiently, we propose Progressive Calibration Networks (PCN) to perform rotation-invariant face detection in a coarse-to-fine manner. PCN consists of three stages, each of which not only distinguishes the faces from non-faces, but also calibrates the RIP orientation of each face candidate to upright progressively. By dividing the calibration process into several progressive steps and only predicting coarse orientations in early stages, PCN can achieve precise and fast calibration. By performing binary classification of face vs. non-face with gradually decreasing RIP ranges, PCN can accurately detect faces with full 360° RIP angles. Such designs lead to a real-time rotation-invariant face detector. The experiments on multi-oriented FDDB and a challenging subset of WIDER FACE containing rotated faces in the wild show that our PCN achieves quite promising performance. Xuepeng Shi, Shiguang Shan, Meina Kan, Shuzhe Wu, Xilin Chen 0001 |
CVPR | 3 |
| 2018 | Face Recognition with Contrastive Convolution
Chunrui Han, Shiguang Shan, Meina Kan, Shuzhe Wu, Xilin Chen 0001 |
ECCV (9) | 3 |
| 2018 | Generative Adversarial Network with Spatial Attention for Face Attribute Editing
Gang Zhang 0005, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ECCV (6) | 2 |
| 2018 | Hierarchical Training for Large Scale Face Recognition with Few Samples Per SubjectabstractRecent progress of face recognition benefits a lot from large-scale face datasets with deep Convoluitonal Neural Networks(CNN). However, when dataset contains a large number of subjects but with few samples for each subject, conventional CNN with softmax loss is heavily prone to overfitting. To address this issue, we propose a hierarchical training schema to optimize CNN with coarse-to-fine class labels, referred to as Hit-CNN. Firstly trained with coarse class labels and then refined with fine class labels, Hit-CNN is enabled the to capture the distribution of data from major variations to fine variations progressively, which can effectively relieve the overfitting and lead to better generalization. In this work, the hierarchical coarse-to-fine class labels are obtained via hierarchical k-means clustering according to the face identities. Evaluated on two face datasets, the proposed Hit-CNN provides better results compared with the conventional CNN under the circumstances of large-scale data with few samples per subject. Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ICIP | 2 |
| 2018 | Face Anti-Spoofing with Multi-Scale InformationabstractFace anti-spoofing has encountered increasing demand as one of the key technologies for reliable and safe authentication with faces. Current face anti-spoofing methods generally take a single crop of face region as input for classification, i.e. exploiting information at only one scale. This single-scale scheme mainly focuses on facial characteristics but not utilize the surrounding information, causing poor generalization for different scenarios with varied means of attacks. Besides, it is tedious or highly empirical to determine an optimal scale of face crops. To overcome the limitations of single-scale methods, in this work we propose to integrate Multi-Scale information for better Face ANti-Spoofing (MS-FANS). Specifically, the proposed MS-FANS method takes multiple face crops at different scales as input followed by a convolutional neural network (CNN) for feature extraction. Then the features from different scales form as a sequence, which are fed into a Long Short-Term Memory (LSTM) network for adaptive fusion of multi-scale information, constructing the final representation for classification. Benefited from this multi-scale design, MS-FANS can adaptively utilize context information from multiple scales, leading to promising performance on two challenging face anti-spoofing datasets, Idiap REPLAY-ATTACK and CASIA-FASD, with significant improvement compared with the existing methods. Shiying Luo, Meina Kan, Shuzhe Wu, Xilin Chen 0001, Shiguang Shan |
ICPR | 2 |
| 2017 | A Fully End-to-End Cascaded CNN for Facial Landmark DetectionabstractFacial landmark detection plays an important role in computer vision. It is a challenging problem due to various poses, exaggerated expressions and partial occlusions. In this work, we propose a Fully End-to-End Cascaded Convolutional Neural Network (FEC-CNN) for more promising facial landmark detection. Specifically, FEC-CNN includes several sub- CNNs, which progressively refine the shape prediction via finer and finer modeling, and the overall network is optimized fully end-to-end. Experiments on three challenging datasets, IBUG, 300W competition and AFLW, demonstrate that the proposed method is robust to large poses, exaggerated expressions and partial occlusions. The proposed FEC-CNN significantly improves the accuracy of landmark prediction. Zhenliang He, Meina Kan, Jie Zhang 0071, Xilin Chen 0001, Shiguang Shan |
FG | 2 |
| 2017 | LDF-Net: Learning a Displacement Field Network for Face Recognition across PoseabstractFace recognition is an important problem in computer vision, however, it is still challenging due to a few wild factors, such as large variations caused by pose, expression, lighting, etc. In this work, we mainly focus on dealing with the pose variations for face recognition. The proposed method attempts to directly transform a non-frontal face image into frontal one by Learning a Displacement Field network (LDFNet) and then recognizes with the transformed images. The existing methods, that follow the same scheme of transforming non-frontal faces into frontal ones, either transform by using 3D-model (3D methods) or transform by using 2D reconstructive methods (2D methods). The 3D methods may lead to the invisibility of some pixels in the transformed frontal images, while the 2D methods may lead to difference between the pixels in the transformed frontal images and the original non-frontal images. Our proposed LDF-Net method can handle these two problems by learning a morphable displacement field for each pixel in the transformed frontal image. Therefore, LDF-Net can achieve a frontal image where all pixels are from the original non-frontal image pixels and no invisible pixels exist, so as to maintain the informative information from the non-frontal images as much as possible. The experiments on MultiPIE dataset show that the proposed LDF-Net achieves state-of-theart performance for face recognition across pose, especially for those large poses. Lanqing Hu, Meina Kan, Shiguang Shan, Xingguang Song, Xilin Chen 0001 |
FG | 2 |
| 2017 | Noisy Face Image Sets Refining Collaborated with Discriminant Feature Space LearningabstractLarge-scale face data together with deep learningtechnology have significantly improved the performance of facerecognition in the wild. Hereinto, the large-scale face data playsa fundamental role, and it is nontrivial to collect a large-scaleface dataset with accurate class labels. No wonder it is quitemoney and effort consuming by collecting manually, howeverit is easy to access large scale face images by using a searchengine with names as keywords. Unfortunately, the retrievedface images from search engine are usually messed up withsome noise images with wrong labels, which forms a greatneed of developing algorithms to refine the retrieved noisyface image set. In this work, we propose a joint frameworkin which multiple noisy face image sets refining collaborateswith the discriminant feature space learning. Specifically, thetwo modules, refining each noisy face image set by conductingone-class classification based on learnt discriminant feature andlearning discriminant feature space based on refined face imagesets, are updated iteratively inducing an effective refinementmodel. To investigate the proposed method, we collect a realworlddataset for the evaluation including 15,515 images of46 subjects with 40% ~ 63.5% noise images per subject. Theexperimental results demonstrate that state-of-the-art one-classclassification methods can be significantly improved whenbeing embedded in the proposed framework, and the proposedframework exhibits strong robustness even when the mean noiseproportion is up to 50% ~ 80%. Xin Liu 0044, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
FG | 2 |
| 2017 | Self-Error-Correcting Convolutional Neural Network for Learning with Noisy LabelsabstractConvolutional Neural Network (CNN) together with large-scale labeled data has achieved the state-of-the-art accuracy in various computer vision tasks. In real-world settings, however, the labels of large scale data can be noisy, which shall seriously degenerate the performance of CNN. In this work, we propose a self-error-correcting CNN (SECCNN) to deal with the noisy labels problem, by simultaneously correcting the improbable labels and optimizing the deep model. Specifically, the SEC-CNN provides an opportunity to correct a wrong label by developing a confidence policy to switch between the label of the sample and the max-activated output neuron of the CNN. Based on the assumption that the deep model is more and more accurate during the training, the confidence policy relies more on the given labels at the beginning stages, but tends to believe that the max-activated neuron of the learned network is reliable. SEC-CNN enables CNN learning to be effective even with 80% noisy labels. Extensive experimental results on MNIST, CIFAR-10, ImageNet and CCFD face dataset demonstrate the effectiveness of the proposed method in dealing with noisy labels. Xin Liu 0044, Shaoxin Li 0001, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
FG | 3 |
| 2017 | Recursive Spatial Transformer (ReST) for Alignment-Free Face RecognitionabstractConvolutional Neural Network (CNN) has led to significant progress in face recognition. Currently most CNN-based face recognition methods follow a two-step pipeline, i.e. a detected face is first aligned to a canonical one predefined by a mean face shape, and then it is fed into a CNN to extract features for recognition. The alignment step transforms all faces to the same shape, which can cause loss of geometrical information which is helpful in distinguishing different subjects. Moreover, it is hard to define a single optimal shape for the following recognition, since faces have large diversity in facial features, e.g. poses, illumination, etc. To be free from the above problems with an independent alignment step, we introduce a Recursive Spatial Transformer (ReST) module into CNN, allowing face alignment to be jointly learned with face recognition in an end-to-end fashion. The designed ReST has an intrinsic recursive structure and is capable of progressively aligning faces to a canonical one, even those with large variations. To model non-rigid transformation, multiple ReST modules are organized in a hierarchical structure to account for different parts of faces. Overall, the proposed ReST can handle large face variations and non-rigid transformation, and is end-to-end learnable and adaptive to input, making it an effective alignment-free face recognition solution. Extensive experiments are performed on LFW and YTF datasets, and the proposed ReST outperforms those two-step methods, demonstrating its effectiveness. Wanglong Wu, Meina Kan, Xin Liu 0044, Yi Yang 0001, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2017 | VIPLFaceNet: an open source deep face recognition SDK
Xin Liu 0044, Meina Kan, Wanglong Wu, Shiguang Shan, Xilin Chen 0001 |
Frontiers Comput. Sci. | 2 |
| 2017 | Funnel-structured cascade for multi-view face detection with alignment-awareness
Shuzhe Wu, Meina Kan, Zhenliang He, Shiguang Shan, Xilin Chen 0001 |
Neurocomputing | 2 |
| 2016 | Multi-view Deep Network for Cross-View ClassificationabstractCross-view recognition that intends to classify samples between different views is an important problem in computer vision. The large discrepancy between different even heterogenous views make this problem quite challenging. To eliminate the complex (maybe even highly nonlinear) view discrepancy for favorable cross-view recognition, we propose a multi-view deep network (MvDN), which seeks for a non-linear discriminant and view-invariant representation shared between multiple views. Specifically, our proposed MvDN network consists of two sub-networks, view-specific sub-network attempting to remove view-specific variations and the following common sub-network attempting to obtain common representation shared by all views. As the objective of MvDN network, the Fisher loss, i.e. the Rayleigh quotient objective, is calculated from the samples of all views so as to guide the learning of the whole network. As a result, the representation from the topmost layers of the MvDN network is robust to view discrepancy, and also discriminative. The experiments of face recognition across pose and face recognition across feature type on three datasets with 13 and 2 views respectively demonstrate the superiority of the proposed method, especially compared to the typical linear ones. Meina Kan, Shiguang Shan, Xilin Chen 0001 |
CVPR | 1 |
| 2016 | Occlusion-Free Face Alignment: Deep Regression Networks Coupled with De-Corrupt AutoEncodersabstractFace alignment or facial landmark detection plays an important role in many computer vision applications, e.g., face recognition, facial expression recognition, face animation, etc. However, the performance of face alignment system degenerates severely when occlusions occur. In this work, we propose a novel face alignment method, which cascades several Deep Regression networks coupled with De-corrupt Autoencoders (denoted as DRDA) to explicitly handle partial occlusion problem. Different from the previous works that can only detect occlusions and discard the occluded parts, our proposed de-corrupt autoencoder network can automatically recover the genuine appearance for the occluded parts and the recovered parts can be leveraged together with those non-occluded parts for more accurate alignment. By coupling de-corrupt autoencoders with deep regression networks, a deep alignment model robust to partial occlusions is achieved. Besides, our method can localize occluded regions rather than merely predict whether the landmarks are occluded. Experiments on two challenging occluded face datasets demonstrate that our method significantly outperforms the state-of-the-art methods. Jie Zhang 0071, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2016 | Multi-View Discriminant AnalysisabstractIn many computer vision systems, the same object can be observed at varying viewpoints or even by different sensors, which brings in the challenging demand for recognizing objects from distinct even heterogeneous views. In this work we propose a Multi-view Discriminant Analysis (MvDA) approach, which seeks for a single discriminant common space for multiple views in a non-pairwise manner by jointly learning multiple view-specific linear transforms. Specifically, our MvDA is formulated to jointly solve the multiple linear transforms by optimizing a generalized Rayleigh quotient, i.e., maximizing the between-class variations and minimizing the within-class variations from both intra-view and inter-view in the common space. By reformulating this problem as a ratio trace problem, the multiple linear transforms are achieved analytically and simultaneously through generalized eigenvalue decomposition. Furthermore, inspired by the observation that different views share similar data structures, a constraint is introduced to enforce the view-consistency of the multiple linear transforms. The proposed method is evaluated on three tasks: face recognition across pose, photo versus. sketch face recognition, and visual light image versus near infrared image face recognition on Multi-PIE, CUFSF and HFB databases respectively. Extensive experiments show that our MvDA achieves significant improvements compared with the best known results. Meina Kan, Shiguang Shan, Haihong Zhang, Shihong Lao, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Bi-Shifting Auto-Encoder for Unsupervised Domain AdaptationabstractIn many real-world applications, the domain of model learning (referred as source domain) is usually inconsistent with or even different from the domain of testing (referred as target domain), which makes the learnt model degenerate in target domain, i.e., the test domain. To alleviate the discrepancy between source and target domains, we propose a domain adaptation method, named as Bi-shifting Auto-Encoder network (BAE). The proposed BAE attempts to shift source domain samples to target domain, and also shift the target domain samples to source domain. The non-linear transformation of BAE ensures the feasibility of shifting between domains, and the distribution consistency between the shifted domain and the desirable domain is constrained by sparse reconstruction between them. As a result, the shifted source domain is supervised and follows similar distribution as target domain. Therefore, any supervised method can be applied on the shifted source domain to train a classifier for classification in target domain. The proposed method is evaluated on three domain adaptation scenarios of face recognition, i.e., domain adaptation across view angle, ethnicity, and imaging sensor, and the promising results demonstrate that our proposed BAE can shift samples between domains and thus effectively deal with the domain discrepancy. Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ICCV | 1 |
| 2015 | Leveraging Datasets with Varying Annotations for Face Alignment via Deep Regression NetworkabstractFacial landmark detection, as a vital topic in computer vision, has been studied for many decades and lots of datasets have been collected for evaluation. These datasets usually have different annotations, e.g., 68-landmark markup for LFPW dataset, while 74-landmark markup for GTAV dataset. Intuitively, it is meaningful to fuse all the datasets to predict a union of all types of landmarks from multiple datasets (i.e., transfer the annotations of each dataset to all other datasets), but this problem is nontrivial due to the distribution discrepancy between datasets and incomplete annotations of all types for each dataset. In this work, we propose a deep regression network coupled with sparse shape regression (DRN-SSR) to predict the union of all types of landmarks by leveraging datasets with varying annotations, each dataset with one type of annotation. Specifically, the deep regression network intends to predict the union of all landmarks, and the sparse shape regression attempts to approximate those undefined landmarks on each dataset so as to guide the learning of the deep regression network for face alignment. Extensive experiments on two challenging datasets, IBUG and GLF, demonstrate that our method can effectively leverage the multiple datasets with different annotations to predict the union of all types of landmarks. Jie Zhang 0071, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2014 | Topic-Aware Deep Auto-Encoders (TDA) for Face Alignment
Jie Zhang 0071, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ACCV (3) | 2 |
| 2014 | Stacked Progressive Auto-Encoders (SPAE) for Face Recognition Across PosesabstractIdentifying subjects with variations caused by poses is one of the most challenging tasks in face recognition, since the difference in appearances caused by poses may be even larger than the difference due to identity. Inspired by the observation that pose variations change non-linearly but smoothly, we propose to learn pose-robust features by modeling the complex non-linear transform from the non-frontal face images to frontal ones through a deep network in a progressive way, termed as stacked progressive auto-encoders (SPAE). Specifically, each shallow progressive auto-encoder of the stacked network is designed to map the face images at large poses to a virtual view at smaller ones, and meanwhile keep those images already at smaller poses unchanged. Then, stacking multiple these shallow auto-encoders can convert non-frontal face images to frontal ones progressively, which means the pose variations are narrowed down to zero step by step. As a result, the outputs of the topmost hidden layers of the stacked network contain very small pose variations, which can be used as the pose-robust features for face recognition. An additional attractiveness of the proposed method is that no pose estimation is needed for the test images. The proposed method is evaluated on two datasets with pose variations, i.e., MultiPIE and FERET datasets, and the experimental results demonstrate the superiority of our method to the existing works, especially to those 2D ones. Meina Kan, Shiguang Shan, Hong Chang 0001, Xilin Chen 0001 |
CVPR | 1 |
| 2014 | Coarse-to-Fine Auto-Encoder Networks (CFAN) for Real-Time Face Alignment
Jie Zhang 0071, Shiguang Shan, Meina Kan, Xilin Chen 0001 |
ECCV (2) | 3 |
| 2014 | Domain Adaptation for Face Recognition: Targetize Source Domain Bridged by Common Subspace
Meina Kan, Junting Wu, Shiguang Shan, Xilin Chen 0001 |
Int. J. Comput. Vis. | 1 |
| 2014 | Semisupervised Hashing via Kernel Hyperplane Learning for Scalable Image SearchabstractHashing methods that aim to seek a compact binary code for each image are demonstrated to be efficient for scalable content-based image retrieval. In this paper, we propose a new hashing method called semisupervised kernel hyperplane learning (SKHL) for semantic image retrieval by modeling each hashing function as a nonlinear kernel hyperplane constructed from an unlabeled dataset. Moreover, a Fisher-like criterion is proposed to learn the optimal kernel hyperplanes and hashing functions, using only weakly labeled training samples with side information. To further integrate different types of features, we also incorporate multiple kernel learning (MKL) into the proposed SKHL (called SKHL-MKL), leading to better hashing functions. Comprehensive experiments on CIFAR-100 and NUS-WIDE datasets demonstrate the effectiveness of our SKHL and SKHL-MKL. Meina Kan, Dong Xu 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Adaptive discriminant learning for face recognition
Meina Kan, Shiguang Shan, Yu Su 0009, Dong Xu 0001, Xilin Chen 0001 |
Pattern Recognit. | 1 |
| 2013 | Learning Prototype Hyperplanes for Face Verification in the WildabstractIn this paper, we propose a new scheme called Prototype Hyperplane Learning (PHL) for face verification in the wild using only weakly labeled training samples (i.e., we only know whether each pair of samples are from the same class or different classes without knowing the class label of each sample) by leveraging a large number of unlabeled samples in a generic data set. Our scheme represents each sample in the weakly labeled data set as a mid-level feature with each entry as the corresponding decision value from the classification hyperplane (referred to as the prototype hyperplane) of one Support Vector Machine (SVM) model, in which a sparse set of support vectors is selected from the unlabeled generic data set based on the learnt combination coefficients. To learn the optimal prototype hyperplanes for the extraction of mid-level features, we propose a Fisher’s Linear Discriminant-like (FLD-like) objective function by maximizing the discriminability on the weakly labeled data set with a constraint enforcing sparsity on the combination coefficients of each SVM model, which is solved by using an alternating optimization method. Then, we use the recent work called Side-Information based Linear Discriminant (SILD) analysis for dimensionality reduction and a cosine similarity measure for final face verification. Comprehensive experiments on two data sets, Labeled Faces in the Wild (LFW) and YouTube Faces, demonstrate the effectiveness of our scheme. Meina Kan, Dong Xu 0001, Shiguang Shan, Wen Li 0001, Xilin Chen 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | Multi-view Discriminant Analysis
Meina Kan, Shiguang Shan, Haihong Zhang, Shihong Lao, Xilin Chen 0001 |
ECCV (1) | 1 |
| 2011 | Side-Information based Linear Discriminant Analysis for Face RecognitionabstractIn recent years, face recognition in the unconstrained environment has attracted increasing attentions, and a few methods have been evaluated on the Labeled Faces in the Wild (LFW) database. In the unconstrained conditions, sometimes we cannot obtain the full class label information of all the subjects. Instead we can only get the weak label information, such as the side-information, i.e., the image pairs from the same or different subjects. In this scenario, many multi-class methods (e.g., the well-known Fisher Linear Discriminant Analysis (FLDA)), fail to work due to the lack of full class label information. To effectively utilize the side-information in such case, we propose Side-Information based Linear Discriminant Analysis (SILD), in which the within-class and between-class scatter matrices are directly calculated by using the side-information. Moreover, we theoretically prove that our SILD method is equivalent to FLDA when the full class label information is available. Experiments on LFW and FRGC databases support our theoretical analysis, and SILD using multiple features also achieve promising performance when compared with the state-of-the-art methods. Meina Kan, Shiguang Shan, Dong Xu 0001, Xilin Chen 0001 |
BMVC | 1 |
| 2011 | Adaptive discriminant analysis for face recognition from single sample per personabstractDiscriminant analysis, especially Fisherface and its numerous variants, have achieved great success in face recognition. However, these methods fail to work for face recognition from Single Sample per Person (SSPP), since they need more than one sample per person to estimate the within-class scatter matrix. To break this inability of traditional discriminant analysis, our paper proposes Adaptive Discriminant Analysis (ADA). In our method, the within-class scatter matrix of each enrolled subject is estimated from his/her single sample, by inferring from a generic training set with multiple samples per person. The inference is inspired by a simple intuition that similar person follows similar within-class variations. Specifically, both kNN regression and Lasso regression are explored for this purpose. We evaluate our method on FERET database and a large real-world face database. The results are very impressive compared with dominant traditional solutions to SSPP problem. Meina Kan, Shiguang Shan, Yu Su 0009, Xilin Chen 0001, Wen Gao 0001 |
FG | 1 |