VLDB 2026 Research / reviewers in the wild / expert
Juwei Lu
dblp:06/827
· DBLP profile ↗
42ranked-venue papers
12as first author
16since 2021 · last 2025
0000-0001-9942-4265ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IMFine: 3D Inpainting via Geometry-guided Multi-view RefinementabstractCurrent 3D inpainting and object removal methods are largely limited to front-facing scenes, facing substantial challenges when applied to diverse, "unconstrained" scenes where the camera orientation and trajectory are unrestricted. To bridge this gap, we introduce a novel approach that produces inpainted 3D scenes with consistent visual quality and coherent underlying geometry across both front-facing and unconstrained scenes. Specifically, we propose a robust 3D inpainting pipeline that incorporates geometric priors and a multi-view refinement network trained via test-time adaptation, building on a pre-trained image inpainting model. Additionally, we develop a novel inpainting mask detection technique to derive targeted inpainting masks from object masks, boosting the performance in handling unconstrained scenes. To validate the efficacy of our approach, we create a challenging and diverse benchmark that spans a wide range of scenes. Comprehensive experiments demonstrate that our proposed method substantially outperforms existing state-of-the-art approaches. Zhihao Shi, Dong Huo, Yuhongze Zhou, Yan Min, Juwei Lu, Xinxin Zuo |
CVPR | 5 |
| 2025 | MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked TransformerabstractGenerative masked transformer have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the animation domain, large datasets are not always available. Applying generative masked modeling to generate diverse instances from a single MoCap reference may lead to overfitting, a challenge that remains unexplored. In this work, we present MotionDreamer, a localized masked modeling paradigm designed to learn motion internal patterns from a given motion with arbitrary topology and duration. By embedding the given motion into quantized tokens with a novel distribution regularization method, MotionDreamer constructs a robust and informative codebook for local motion patterns. Moreover, a sliding window local attention is introduced in our masked transformer, enabling the generation of natural yet diverse animations that closely resemble the reference motion patterns. As demonstrated through comprehensive experiments, MotionDreamer outperforms the state-of-the-art methods that are typically GAN or Diffusion-based in both faithfulness and diversity. Thanks to the consistency and robustness of quantization-based approach, MotionDreamer can also effectively perform downstream tasks such as temporal motion editing, crowd motion synthesis, and beat-aligned dance generation, all using a single reference motion. Our implementation, learned models and results are to be made publicly available upon paper acceptance. Chuan Guo 0002, Yuxuan Mu, Muhammad Gohar Javed, Xinxin Zuo, Juwei Lu, Hai Jiang 0001, Li Cheng 0001 |
ICLR | 6 |
| 2024 | TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling
Dong Huo, Zixin Guo, Xinxin Zuo, Zhihao Shi, Juwei Lu, Peng Dai 0002, Songcen Xu, Li Cheng 0001, Yee-Hong Yang |
ECCV (38) | 5 |
| 2024 | GSD: View-Guided Gaussian Splatting Diffusion for 3D Reconstruction
Yuxuan Mu, Xinxin Zuo, Chuan Guo 0002, Juwei Lu, Songcen Xu, Peng Dai 0002, Youliang Yan, Li Cheng 0001 |
ECCV (79) | 5 |
| 2024 | AQF: Assessing the Quality of Hyperspectral Reconstruction with a Learnable MetricabstractThis paper proposes a learnable metric to measure the reconstruction quality of hyperspectral images obtained by computational hyperspectral imaging. Computational hyperspectral imaging aims to obtain low-cost hyperspectral images through consumer camera. While many hyperspectral reconstruction models have been developed for this purpose, conventional image and spectral quality metrics are insufficient to measure the scientific value of the reconstructed HSI cube. This paper proposes an adaptive quality fusion metric (AQF), adaptively aggregating the quality measures from point-wise, spatial-wise and spectral-wise aspects to assess the scientific value preserved by the reconstructed HSI. The proposed AQF metric uses weight parameters generated by a modified hypernetwork to determine the contribution for the three aspects given paired of groundtruth HSI and reconstructed HSI. Experimental results show its compatibility with existing metrics while accurately measuring the scientific information retained by the reconstructed HSI for hyperspectral applications. Pai Chet Ng, Juwei Lu, Konstantinos N. Plataniotis |
ICASSP | 2 |
| 2024 | Generative Human Motion Stylization in Latent SpaceabstractHuman motion stylization aims to revise the style of an input motion while keeping its content unaltered. Unlike existing works that operate directly in pose space, we leverage the \textit{latent space} of pretrained autoencoders as a more expressive and robust representation for motion extraction and infusion. Building upon this, we present a novel \textit{generative} model that produces diverse stylization results of a single motion (latent) code. During training, a motion code is decomposed into two coding components: a deterministic content code, and a probabilistic style code adhering to a prior distribution; then a generator massages the random combination of content and style codes to reconstruct the corresponding motion codes. Our approach is versatile, allowing the learning of probabilistic style space from either style labeled or unlabeled motions, providing notable flexibility in stylization as well. In inference, users can opt to stylize a motion using style cues from a reference motion or a label. Even in the absence of explicit style input, our model facilitates novel re-stylization by sampling from the unconditional style prior distribution. Experimental results show that our proposed stylization models, despite their lightweight design, outperform the state-of-the-arts in style reeanactment, content preservation, and generalization across various applications and settings. Chuan Guo 0002, Yuxuan Mu, Xinxin Zuo, Peng Dai 0002, Youliang Yan, Juwei Lu, Li Cheng 0001 |
ICLR | 6 |
| 2023 | CLIPPING: Distilling CLIP-Based Models with a Student Base for Video-Language RetrievalabstractPre-training a vision-language model and then fine-tuning it on downstream tasks have become a popular paradigm. However, pre-trained vision-language models with the Transformer architecture usually take long inference time. Knowledge distillation has been an efficient technique to transfer the capability of a large model to a small one while maintaining the accuracy, which has achieved remarkable success in natural language processing. However, it faces many problems when applying KD to the multi-modality applications. In this paper, we propose a novel knowledge distillation method, named CLIPPING11In this paper, CLIPPING means cutting something to make it smaller through distilling., where the plentiful knowledge of a large teacher model that has been fine-tuned for video-language tasks with the powerful pre-trained CLIP can be effectively transferred to a small student only at the fine-tuning stage. Especially, a new layer-wise alignment with the student as the base is proposed for knowledge distillation of the intermediate layers in CLIPPING, which enables the student's layers to be the bases of the teacher, and thus allows the student to fully absorb the knowledge of the teacher. CLIPPING with MobileViT-v2 as the vision encoder without any vision-language pre-training achieves 88.1%-95.3% of the performance of its teacher on three video-language retrieval benchmarks, with its vision encoder being 19.5x smaller. CLIPPING also significantly outperforms a state-of-the-art small baseline (ALL-in-one-B) on the MSR-VTT dataset, obtaining relatively 7.4% performance gain, with 29% fewer parameters and 86.9% fewer flops. Moreover, CLIPPING is comparable or even superior to many large pre-training models. Renjing Pei, Jianzhuang Liu, Weimian Li, Songcen Xu, Peng Dai 0002, Juwei Lu, Youliang Yan |
CVPR | 7 |
| 2023 | HiVLP: Hierarchical Interactive Video-Language Pre-TrainingabstractVideo-Language Pre-training (VLP) has become one of the most popular research topics in deep learning. However, compared to image-language pre-training, VLP has lagged far behind due to the lack of large amounts of video-text pairs. In this work, we train a VLP model with a hybrid of image-text and video-text pairs, which significantly outperforms pre-training with only the video-text pairs. Besides, existing methods usually model the cross-modal interaction using cross-attention between single-scale visual tokens and textual tokens. These visual features are either of low resolutions lacking fine-grained information, or of high resolutions without high-level semantics. To address the issue, we propose Hierarchical interactive Video-Language Pre-training (HiVLP) that efficiently uses a hierarchical visual feature group for multi-modal cross-attention during pre-training. In the hierarchical framework, low-resolution features are learned with focus on more global high-level semantic information, while high-resolution features carry fine-grained details. As a result, HiVLP has the ability to effectively learn both the global and fine-grained representations to achieve better alignment between video and text inputs. Furthermore, we design a hierarchical multi-scale vision contrastive loss for self-supervised learning to boost the interaction between them. Experimental results show that HiVLP establishes new state-of-the-art results in three downstream tasks, text-video retrieval, video-text retrieval, and video captioning. Jianzhuang Liu, Renjing Pei, Songcen Xu, Peng Dai 0002, Juwei Lu, Weimian Li, Youliang Yan |
ICCV | 6 |
| 2023 | Decorate3D: Text-Driven High-Quality Texture Generation for Mesh Decoration in the WildabstractThis paper presents Decorate3D, a versatile and user-friendly method for the creation and editing of 3D objects using images. Decorate3D models a real-world object of interest by neural radiance field (NeRF) and decomposes the NeRF representation into an explicit mesh representation, a view-dependent texture, and a diffuse UV texture. Subsequently, users can either manually edit the UV or provide a prompt for the automatic generation of a new 3D-consistent texture. To achieve high-quality 3D texture generation, we propose a structure-aware score distillation sampling method to optimize a neural UV texture based on user-defined text and empower an image diffusion model with 3D-consistent generation capability. Furthermore, we introduce a few-view resampling training method and utilize a super-resolution model to obtain refined high-resolution UV textures (2048$\times$2048) for 3D texturing. Extensive experiments collectively validate the superior performance of Decorate3D in retexturing real-world 3D objects. Project page: https://decorate3d.github.io/Decorate3D/. Xinxin Zuo, Peng Dai 0002, Juwei Lu, Li Cheng 0001, Youliang Yan, Songcen Xu |
NeurIPS | 4 |
| 2023 | Hyper-Skin: A Hyperspectral Dataset for Reconstructing Facial Skin-Spectra from RGB ImagesabstractWe introduce Hyper-Skin, a hyperspectral dataset covering wide range of wavelengths from visible (VIS) spectrum (400nm - 700nm) to near-infrared (NIR) spectrum (700nm - 1000nm), uniquely designed to facilitate research on facial skin-spectra reconstruction.By reconstructing skin spectra from RGB images, our dataset enables the study of hyperspectral skin analysis, such as melanin and hemoglobin concentrations, directly on the consumer device. Overcoming limitations of existing datasets, Hyper-Skin consists of diverse facial skin data collected with a pushbroom hyperspectral camera. With 330 hyperspectral cubes from 51 subjects, the dataset covers the facial skin from different angles and facial poses.Each hyperspectral cube has dimensions of 1024$\times$1024$\times$448, resulting in millions of spectra vectors per image. The dataset, carefully curated in adherence to ethical guidelines, includes paired hyperspectral images and synthetic RGB images generated using real camera responses. We demonstrate the efficacy of our dataset by showcasing skin spectra reconstruction using state-of-the-art models on 31 bands of hyperspectral data resampled in the VIS and NIR spectrum. This Hyper-Skin dataset would be a valuable resource to NeurIPS community, encouraging the development of novel algorithms for skin spectral reconstruction while fostering interdisciplinary collaboration in hyperspectral skin analysis related to cosmetology and skin's well-being. Instructions to request the data and the related benchmarking codes are publicly available at: https://github.com/hyperspectral-skin/Hyper-Skin-2023. Pai Chet Ng, Zhixiang Chi, Yannick Verdie, Juwei Lu, Konstantinos N. Plataniotis |
NeurIPS | 4 |
| 2022 | Self-Supervised Spatiotemporal Representation Learning by Exploiting Video ContinuityabstractRecent self-supervised video representation learning methods have found significant success by exploring essential properties of videos, e.g. speed, temporal order, etc. This work exploits an essential yet under-explored property of videos, the \textit{video continuity}, to obtain supervision signals for self-supervised representation learning. Specifically, we formulate three novel continuity-related pretext tasks, i.e. continuity justification, discontinuity localization, and missing section approximation, that jointly supervise a shared backbone for video representation learning. This self-supervision approach, termed as Continuity Perception Network (CPNet), solves the three tasks altogether and encourages the backbone network to learn local and long-ranged motion and context representations. It outperforms prior arts on multiple downstream tasks, such as action recognition, video retrieval, and action localization. Additionally, the video continuity can be complementary to other coarse-grained video properties for representation learning, and integrating the proposed pretext task to prior arts can yield much performance gains. Hanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen, Peng Dai 0002, Juwei Lu, Yang Wang 0003 |
AAAI | 6 |
| 2022 | Decompose the Sounds and Pixels, Recompose the EventsabstractIn this paper, we propose a framework centering around a novel architecture called the Event Decomposition Recomposition Network (EDRNet) to tackle the Audio-Visual Event (AVE) localization problem in the supervised and weakly supervised settings. AVEs in the real world exhibit common unraveling patterns (termed as Event Progress Checkpoints(EPC)), which humans can perceive through the cooperation of their auditory and visual senses. Unlike earlier methods which attempt to recognize entire event sequences, the EDRNet models EPCs and inter-EPC relationships using stacked temporal convolutions. Based on the postulation that EPC representations are theoretically consistent for an event category, we introduce the State Machine Based Video Fusion, a novel augmentation technique that blends source videos using different EPC template sequences. Additionally, we design a new loss function called the Land-Shore-Sea loss to compactify continuous foreground and background representations. Lastly, to alleviate the issue of confusing events during weak supervision, we propose a prediction stabilization method called Bag to Instance Label Correction. Experiments on the AVE dataset show that our collective framework outperforms the state-of-the-art by a sizable margin. Varshanth R. Rao, Md Ibrahim Khalil, Haoda Li, Peng Dai 0002, Juwei Lu |
AAAI | 5 |
| 2022 | Dual Perspective Network for Audio-Visual Event Localization
Varshanth R. Rao, Md Ibrahim Khalil, Haoda Li, Peng Dai 0002, Juwei Lu |
ECCV (34) | 5 |
| 2021 | Boosting the Generalization Capability in Cross-Domain Few-shot Learning via Noise-enhanced Supervised AutoencoderabstractState of the art (SOTA) few-shot learning (FSL) methods suffer significant performance drop in the presence of domain differences between source and target datasets. The strong discrimination ability on the source dataset does not necessarily translate to high classification accuracy on the target dataset. In this work, we address this cross-domain few-shot learning (CDFSL) problem by boosting the generalization capability of the model. Specifically, we teach the model to capture broader variations of the feature distributions with a novel noise-enhanced supervised autoencoder (NSAE). NSAE trains the model by jointly reconstructing inputs and predicting the labels of inputs as well as their reconstructed pairs. Theoretical analysis based on intra-class correlation (ICC) shows that the feature embeddings learned from NSAE have stronger discrimination and generalization abilities in the target domain. We also take advantage of NSAE structure and propose a two-step fine-tuning procedure that achieves better adaption and improves classification performance in the target domain. Extensive experiments and ablation studies are conducted to demonstrate the effectiveness of the proposed method. Experimental results show that our proposed method consistently outperforms SOTA methods under various conditions. Hanwen Liang, Peng Dai 0002, Juwei Lu |
ICCV | 4 |
| 2021 | Class Semantics-based Attention for Action Detection
Deepak Sridhar, Niamul Quader, Srikanth Muralidharan, Yaoxin Li, Peng Dai 0002, Juwei Lu |
ICCV | 6 |
| 2021 | Learning Causal Representation for Training Cross-Domain Pose Estimator via Generative Interventionsabstract3D pose estimation has attracted increasing attention with the availability of high-quality benchmark datasets. However, prior works show that deep learning models tend to learn spurious correlations, which fail to generalize beyond the specific dataset they are trained on. In this work, we take a step towards training robust models for cross-domain pose estimation task, which brings together ideas from causal representation learning and generative adversarial networks. Specifically, this paper introduces a novel framework for causal representation learning which explicitly exploits the causal structure of the task. We consider changing domain as interventions on images under the data-generation process and steer the generative model to produce counterfactual features. This help the model learn transferable and causal relations across different domains. Our framework is able to learn with various types of unlabeled datasets. We demonstrate the efficacy of our proposed method on both human and hand pose estimation task. The experiment results show the proposed approach achieves state-of-the-art performance on most datasets for both domain adaptation and domain generalization settings. Xiheng Zhang, Yongkang Wong, Juwei Lu, Mohan Kankanhalli, Weidong Geng |
ICCV | 4 |
| 2020 | All at Once: Temporally Adaptive Multi-frame Interpolation with Advanced Motion Modeling
Zhixiang Chi, Rasoul Mohammadi Nasiri, Juwei Lu, Jin Tang 0005, Konstantinos N. Plataniotis |
ECCV (27) | 4 |
| 2020 | Weight Excitation: Built-in Attention Mechanisms in Convolutional Neural Networks
Niamul Quader, Md Mafijul Islam Bhuiyan, Juwei Lu, Peng Dai 0002, Wei Li 0002 |
ECCV (30) | 3 |
| 2020 | Towards Efficient Coarse-to-Fine Networks for Action and Gesture Recognition
Niamul Quader, Juwei Lu, Peng Dai 0002, Wei Li 0002 |
ECCV (30) | 2 |
| 2016 | Sub-event recognition and summarization for structured scenario photos
Liyan Zhang 0001, Bradley Denney, Juwei Lu |
Multim. Tools Appl. | 3 |
| 2014 | A collaborative approach for face verification and attributes refinement
Liyan Zhang 0001, Bradley Denney, Juwei Lu |
Inf. Sci. | 3 |
| 2009 | Gaussian kernel optimization for pattern classification
Jie Wang 0010, Haiping Lu, Konstantinos N. Plataniotis, Juwei Lu |
Pattern Recognit. | 4 |
| 2008 | Kernel quadratic discriminant analysis for small sample size problem
Jie Wang 0010, Konstantinos N. Plataniotis, Juwei Lu, Anastasios N. Venetsanopoulos |
Pattern Recognit. | 3 |
| 2006 | Restoration of Motion Blurred ImagesabstractIn this paper, we present several algorithms developed for restoration of motion blurred images. We begin with a single-image based deblurring approach in the case of linear constant motion. This approach is a wavelet-based method with a novel lp-norm regularization term. Due to the introduction of the wavelet and regularization techniques, the approach is rather robust against noise amplification during deconvolution. Then, we further developed a general multi-image based deblurring framework improved from recent works of Rav-Acha, A et al., (2005). The proposed framework is able to effectively take advantage of information contained in the multiple input images, even when they are blurred in the same direction-a case hard to be dealt with by traditional solutions. The proposed methods are evaluated on both simulated and real data, and the obtained experimental results indicate promising results Juwei Lu, Eunice Poon, Konstantinos N. Plataniotis |
ICME | 1 |
| 2006 | On solving the face recognition problem with one training sample per subject
Jie Wang 0010, Konstantinos N. Plataniotis, Juwei Lu, Anastasios N. Venetsanopoulos |
Pattern Recognit. | 3 |
| 2006 | Ensemble-based discriminant learning with boosting for face recognitionabstractIn this paper, we propose a novel ensemble-based approach to boost performance of traditional Linear Discriminant Analysis (LDA)-based methods used in face recognition. The ensemble-based approach is based on the recently emerged technique known as "boosting". However, it is generally believed that boosting-like learning rules are not suited to a strong and stable learner such as LDA. To break the limitation, a novel weakness analysis theory is developed here. The theory attempts to boost a strong learner by increasing the diversity between the classifiers created by the learner, at the expense of decreasing their margins, so as to achieve a tradeoff suggested by recent boosting studies for a low generalization error. In addition, a novel distribution accounting for the pairwise class discriminant information is introduced for effective interaction between the booster and the LDA-based learner. The integration of all these methodologies proposed here leads to the novel ensemble-based discriminant learning approach, capable of taking advantage of both the boosting and LDA techniques. Promising experimental results obtained on various difficult face recognition scenarios demonstrate the effectiveness of the proposed approach. We believe that this work is especially beneficial in extending the boosting framework to accommodate general (strong/weak) learners. Juwei Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos, Stan Z. Li |
IEEE Trans. Neural Networks | 1 |
| 2005 | An efficient kernel discriminant analysis method
Juwei Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos, Jie Wang 0010 |
Pattern Recognit. | 1 |
| 2005 | Regularization studies of linear discriminant analysis in small sample size scenarios with application to face recognition
Juwei Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos |
Pattern Recognit. Lett. | 1 |
| 2004 | Regularization studies on LDA for face recognitionabstractIt is well-known that the applicability of linear discriminant analysis (LDA) to high-dimensional pattern classification tasks such as face recognition (FR) often suffers from the so-called "small sample size" (SSS) problem arising from the small number of available training samples compared to the dimensionality of the sample space. In this paper, we propose a new LDA method that effectively addresses the SSS problem using a regularization technique. In addition, a scheme of expanding the representational capacity of the face database is introduced to overcome the limitation that the LDA based algorithms require at least two samples per class available for learning. Extensive experimentation performed on the FERET database indicates that the proposed methodology outperforms traditional methods such as eigenfaces and direct LDA in a number of SSS setting scenarios. Juwei Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos |
ICIP | 1 |
| 2003 | Regularized D-LDA for face recognitionabstractLinear Discriminant Analysis (LDA) is derived from the optimal Bayes classifier when classes are assumed to be Gaussian with identical covariance matrices. However, it is well known that the distribution of face images under a perceivable variation in viewpoint, illumination or facial ex-pression, is highly nonlinear and complex. The Quadratic Discriminant Analysis (QDA) which relaxes the identical covariance assumption and allows for nonlinear discrimi-nant boundaries to be formed, seems to be a better choice. However, the applicability of QDA to problems, such as face recognition, where the number of training samples is much smaller than the dimensionality of the sample space is problematic due to the increased number of parameters to be learned. In this paper, we propose a new regularized discriminant analysis method that effectively solves the so-called “small sample size ” problem in very high-dimensional face image space. Extensive experimentation performed on the FERET database indicates that the proposed method-ology outperforms traditional methods such as Eigenfaces, QDA and Direct LDA in a number of application scenarios. 1. Juwei Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos |
ICASSP (3) | 1 |
| 2003 | Boosting linear discriminant analysis for face recognitionabstractIn this paper, we propose a new algorithm to boost performance of traditional linear discriminant analysis (LDA)-based face recognition (FR) methods in complex FR tasks, where highly nonlinear face pattern distributions are often encountered. The algorithm embodies the principle of "divide and conquer", by which a complex problem, is decomposed into a set of simpler ones, each of which can be conquered by a relatively easy solution. The Ad-aBoost technique is utilized within this framework to: 1) generalize a set of simple FR sub-problems and their corresponding LDA solutions; 2) combine results from the multiple, relatively weak, LDA solutions to form a very strong solution. Experimentation performed on the FERET database indicates that the proposed methodology is able to greatly enhance performance of the traditional LDA-based method with an averaged improvement of correct recognition rate (CRR) up to 9% reported. Juwei Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos |
ICIP (1) | 1 |
| 2003 | Regularized discriminant analysis for the small sample size problem in face recognition
Juwei Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos |
Pattern Recognit. Lett. | 1 |
| 2003 | Face recognition using kernel direct discriminant analysis algorithmsabstractTechniques that can introduce low-dimensional feature representation with enhanced discriminatory power is of paramount importance in face recognition (FR) systems. It is well known that the distribution of face images, under a perceivable variation in viewpoint, illumination or facial expression, is highly nonlinear and complex. It is, therefore, not surprising that linear techniques, such as those based on principle component analysis (PCA) or linear discriminant analysis (LDA), cannot provide reliable and robust solutions to those FR problems with complex face variations. In this paper, we propose a kernel machine-based discriminant analysis method, which deals with the nonlinearity of the face patterns' distribution. The proposed method also effectively solves the so-called "small sample size" (SSS) problem, which exists in most FR tasks. The new algorithm has been tested, in terms of classification error rate performance, on the multiview UMIST face database. Results indicate that the proposed methodology is able to achieve excellent performance with only a very small set of features being used, and its error rate is approximately 34% and 48% of those of two other commonly used kernel FR approaches, the kernel-PCA (KPCA) and the generalized discriminant analysis (GDA), respectively. Juwei Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos |
IEEE Trans. Neural Networks | 1 |
| 2003 | Face recognition using LDA-based algorithmsabstractLow-dimensional feature representation with enhanced discriminatory power is of paramount importance to face recognition (FR) systems. Most of traditional linear discriminant analysis (LDA)-based methods suffer from the disadvantage that their optimality criteria are not directly related to the classification ability of the obtained feature representation. Moreover, their classification accuracy is affected by the "small sample size" (SSS) problem which is often encountered in FR tasks. In this paper, we propose a new algorithm that deals with both of the shortcomings in an efficient and cost effective manner. The proposed method is compared, in terms of classification accuracy, to other commonly used FR methods on two face databases. Results indicate that the performance of the proposed method is overall superior to those of traditional FR approaches, such as the eigenfaces, fisherfaces, and D-LDA methods. Juwei Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos |
IEEE Trans. Neural Networks | 1 |
| 2002 | Boosting face recognition on a large-scale databaseabstractThe performance of many state-of-the-art face recognition (FR) methods deteriorates rapidly when large databases are considered. We propose a novel clustering method based on a linear discriminant analysis methodology which deals with the problem of FR on a large-scale database. Contrary to traditional clustering methods such as K-means, which are based on certain "similarity criteria", the proposed method uses a novel "separability criterion" to partition a training set from the large database into a set of K smaller and simpler subsets or maximal-separability clusters (MSCs). Based on these MSCs, a novel two-stage hierarchical classification framework is proposed. Under the framework, the complex FR problem on a large database is decomposed into a set of simpler ones, where traditional methods can be successfully applied. Experiments with a database containing 1654 face images of 157 subjects indicate that the error rate performance of a traditional method under the proposed framework can be greatly improved without significantly increasing computational complexity. Juwei Lu, Konstantinos N. Plataniotis |
ICIP (2) | 1 |
| 2002 | A kernel machine based approach for multi-view face recognitionabstractTechniques that can introduce low-dimensional feature representation with enhanced discriminatory power is of paramount importance in face recognition applications. It is well known that the distribution of face images, under a perceivable variation in viewpoint, illumination or facial expression, is highly nonlinear and complex. It is therefore, not surprising that linear techniques, such as those based on principle component analysis (PCA) or linear discriminant analysis (LDA) cannot provide reliable and robust solutions to those complex face recognition problems. We propose a kernel machine based discriminant analysis method, which deals with the nonlinearity of the face patterns' distribution. The proposed method also effectively solves the "small sample size" (SSS) problem which exists in most face recognition tasks. The new algorithm has been tested, in terms of error rate performance, on the multi-view UMIST Face Database. Results indicate that the proposed methodology outperform other commonly used approaches, such as the kernel-PCA (KPCA) and the generalized discriminant analysis (GDA). Juwei Lu, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos |
ICIP (1) | 1 |
| 2002 | Face recognition with radial basis function (RBF) neural networksabstractA general and efficient design approach using a radial basis function (RBF) neural classifier to cope with small training sets of high dimension, which is a problem frequently encountered in face recognition, is presented. In order to avoid overfitting and reduce the computational burden, face features are first extracted by the principal component analysis (PCA) method. Then, the resulting features are further processed by the Fisher's linear discriminant (FLD) technique to acquire lower-dimensional discriminant patterns. A novel paradigm is proposed whereby data information is encapsulated in determining the structure and initial parameters of the RBF neural classifier before learning takes place. A hybrid learning algorithm is used to train the RBF neural networks so that the dimension of the search space is drastically reduced in the gradient paradigm. Simulation results conducted on the ORL database show that the system achieves excellent performance both in terms of error rates of classification and learning efficiency. Meng Joo Er, Shiqian Wu, Juwei Lu, Hock Lye Toh |
IEEE Trans. Neural Networks | 3 |
| 2001 | Minutiae data synthesis for fingerprint identification applicationsabstractIn this paper, we address the false rejection problem due to the small solid state sensor area available for fingerprint image capture. We propose a minutiae data synthesis approach to circumvent this problem. The main advantages of this approach over the existing image mosaicing approach include low memory storage requirements and low computational complexity. Moreover, the possible matching search overhead due to data redundancy can be reduced. Extensive experiments are conducted to determine the best transformation suitable for minutiae alignment. Among the three transformations presented, affine transformation is found to be most suited for minutiae alignment. We demonstrate the idea of synthesis with an example using physical fingerprint images. The proposed synthesis system is also shown to reduce the number of false rejects caused by the use of different fingerprint regions for matching. Kar-Ann Toh, Weiyun Yau, Xudong Jiang 0001, Tai Pang Chen, Juwei Lu, Eyung Lim |
ICIP (3) | 5 |
| 1999 | Modeling Bayesian Estimation for Deformable ContoursabstractA novel trainable snake model called EigenSnake, is presented in the Bayesian framework. In the EigenSnake, prior knowledge of a specific object shape, such as that of face outlines and facial features, is derived from a training set of the shape and incorporated into a Bayesian snake model in the form of the prior distribution. Further, a "shape space", which is constructed on the basis of a set of eigenvectors obtained from principle component analysis, is used to restrict and stabilize the search for the optimal solution. The effectiveness is demonstrated by experiments, which shows that the EigenSnake produces more reliable and accurate results than existing models. Stan Z. Li, Juwei Lu |
ICCV | 2 |
| 1999 | Face recognition using the nearest feature line methodabstractIn this paper, we propose a novel classification method, called the nearest feature line (NFL), for face recognition. Any two feature points of the same class (person) are generalized by the feature line (FL) passing through the two points. The derived FL can capture more variations of face images than the original points and thus expands the capacity of the available database. The classification is based on the nearest distance from the query feature point to each FL. With a combined face database, the NFL error rate is about 43.7-65.4% of that of the standard eigenface method. Moreover, the NFL achieves the lowest error rate reported to date for the ORL face database. Stan Z. Li, Juwei Lu |
IEEE Trans. Neural Networks | 2 |
| 1998 | Generalizing Capacity of Face Database for Face Recognition
Stan Z. Li, Juwei Lu |
FG | 2 |
| 1998 | Hierarchical linear combinations for face recognitionabstractA hierarchical representation consisting of two level linear combinations (LC) is proposed for face recognition. At the first level, a face image is represented as a linear combination (LC) of a set of basis vectors, i.e. eigenfaces. Thereby a face image corresponds to a feature vector (prototype) in the eigenface space. Normally several such prototypes are available for a face class, each representing the face under a particular condition such as in viewpoint, illumination, and so on. We propose to use the second level LC, that of the prototypes belonging to the same face class, to treat the prototypes coherently. The purpose is to improve face recognition under a new condition not captured by the prototypes by using a linear combination of them. A new distance measure called nearest LC (NLC) is proposed as opposed to the NN. Experiments show that our method yields significantly better results than the one level eigenface methods. Stan Z. Li, Juwei Lu, Kap Luk Chan, Jun Liu 0020 |
ICPR | 2 |