VLDB 2026 Research / reviewers in the wild / expert
Shan Du 0001
dblp:70/6348-1
· DBLP profile ↗
30ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0002-2281-5150ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ArchitectHead: Continuous Level of Detail Control for 3D Gaussian Head Avatarsabstract3D Gaussian Splatting (3DGS) has enabled photorealistic and real-time rendering of 3D head avatars. Existing 3DGS-based avatars typically rely on tens of thousands of 3D Gaussian points (Gaussians), with the number of Gaussians fixed after training. However, many practical applications require adjustable levels of detail (LOD) to balance rendering efficiency and visual quality. In this work, we propose "ArchitectHead", the first framework for creating 3D Gaussian head avatars that support continuous control over LOD. Our key idea is to parameterize the Gaussians in a 2D UV feature space and propose a UV feature field composed of multi-level learnable feature maps to encode their latent features. A lightweight neural network-based decoder then transforms these latent features into 3D Gaussian attributes for rendering. ArchitectHead controls the number of Gaussians by dynamically resampling feature maps from the UV feature field at the desired resolutions. This method enables efficient and continuous control of LOD without retraining. Experimental results show that ArchitectHead achieves state-of-the-art (SOTA) quality in self and cross-identity reenactment tasks at the highest LOD, while maintaining near SOTA performance at lower LODs. At the lowest LOD, our method uses only 6.2% of the Gaussians while the quality degrades moderately (L1 Loss +7.9%, PSNR −0.97%, SSIM −0.6%, LPIPS Loss +24.1%), and the rendering speed nearly doubles. Project homepage: https://peizhiyan.github.io/docs/architect/. Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001 |
WACV | 4 |
| 2026 | JVLGS: joint vision-language gas leak segmentation
Xinlong Zhao, Qixiang Pang, Shan Du 0001 |
Vis. Comput. | 3 |
| 2025 | Estimating Virtual Camera FOV to Reduce Perspective Shape Distortion in 2D-to-3D Face ReconstructionabstractExisting image-based 3D face reconstruction methods rely on a virtual camera to project the reconstructed 3D face onto the 2D image plane for comparison with the input image, a crucial step for accurate results. To simplify the reconstruction process, these methods often use fixed camera intrinsics and assume minimal perspective distortion, overlooking the varying distortion levels in "in-the-wild" images and leading to inaccuracies in reconstructed 3D face shapes. To address this issue, we propose estimating the virtual camera’s optimal field-of-view (FOV) for a given image, enabling consistent 3D face reconstruction across varying distortion levels. We introduce two synthetic datasets: one to train our FOV estimation network (FOV-Net) and another to evaluate its performance and reconstruction accuracy. We use the FOV-Net predicted FOV to initialize the camera, which is used in the fitting-based reconstruction process. Experiments show that our approach significantly improves reconstruction consistency under different levels of perspective distortion. Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001 |
ICIP | 4 |
| 2025 | Fine-Grained Spatial-Temporal Perception for Gas Leak SegmentationabstractGas leaks pose significant risks to human health and the environment. Despite long-standing concerns, there are limited methods that can efficiently and accurately detect and segment leaks due to their concealed appearance and random shapes. In this paper, we propose a Fine-Grained Spatial-Temporal Perception (FGSTP) algorithm for gas leak segmentation. FGSTP captures critical motion clues across frames and integrates them with refined object features in an end-to-end network. Specifically, we first construct a correlation volume to capture motion information between consecutive frames. Then, the fine-grained perception progressively refines the object-level features using previous outputs. Finally, a decoder is employed to optimize boundary segmentation. Because there is no highly precise labeled dataset for gas leak segmentation, we manually label a gas leak video dataset, GasVid. Experimental results on GasVid demonstrate that our model excels in segmenting non-rigid objects such as gas leaks, generating the most accurate mask compared to other state-of-the-art (SOTA) models. Xinlong Zhao, Shan Du 0001 |
ICIP | 2 |
| 2025 | Gaussian Déjà-vu: Creating Controllable 3D Gaussian Head-Avatars with Enhanced Generalization and Personalization AbilitiesabstractRecent advancements in 3D Gaussian Splatting (3DGS) have unlocked significant potential for modeling 3D head avatars, providing greater flexibility than mesh-based methods and more efficient rendering compared to NeRF-based approaches. Despite these advancements, the creation of controllable 3DGS-based head avatars remains time-intensive, often requiring tens of minutes to hours. To expedite this process, we here introduce the “Gaussian Déjà-vu” framework, which first obtains a generalized model of the head avatar and then personalizes the result. The generalized model is trained on large 2D (synthetic and real) image datasets. This model provides a well-initialized 3D Gaussian head that is further refined using a monocular video to achieve the personalized head avatar. For personalizing, we propose learnable expression-aware rectification blendmaps to correct the initial 3D Gaussians, ensuring rapid convergence without the reliance on neural networks. Experiments demonstrate that the proposed method meets its objectives. It outperforms state-of-the-art 3D Gaussian head avatars in terms of photorealistic quality as well as reduces training time consumption to at least a quarter of the existing methods, producing the avatar in minutes. Project homepage: https://peizhiyan.github.io/docs/dejavu Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001 |
WACV | 4 |
| 2025 | Cross-Modal Progressive Perspective Matching Network for Remote Sensing Image-Text RetrievalabstractCross-modality based on remote sensing (RS) text-image retrieval has gained increasing attention in recent years due to its ability to leverage the rich semantics of images and the understandability of text to provide a more comprehensive description. Existing cross-modal retrieval methods typically apply self-attention or cross-attention mechanisms to identify important information in RS data, but they ignore the multi-view perception characteristic of geographical space in RS images. As a result, these retrieval models fail to locate the correct perspective in images according to the query text, ultimately leading to incorrect matching. In this work, a Cross-modal Progressive Perspective Matching Network (CPPMN) is proposed for remote sensing image-text retrieval by establishing a progressive perspective matching mechanism and semantic alignment to further improve the performance of the retrieval model. Specifically, the CPPMN framework consists of three core modules: the Compensation Network for Full Perspective Modeling (CN_FPM), the Graph Transformation for Individual Perspective Modeling (GT_IPM), and the Cascaded Transformer for Cross-modal Semantic Alignment (CT_CSA). The CN_FPM module utilizes all positive text samples as supervision signals to guide the feature extraction training process, aiming to capture full perspective information from images. Subsequently, the GT_IPM module transforms implicit-perspective feature representations into explicit-perspective cross-modal relationship graphs. This transformation enables the identification of specific perspective locations within the image according to the query sentence by analyzing graph density and connectivity. Finally, the CT_CSA module comprises a cascaded Transformer network that aligns features at the semantic level between cross-modal data The quantitative and qualitative experiments are conducted on four large-scale remote sensing cross-modal retrieval datasets to demonstrate the significant performance of adopting the progressive perspective matching mechanism and semantic alignment strategy. Xiu Li 0006, Lei Huang 0010, Shan Du 0001, Jie Nie, Junyu Dong |
IEEE Trans. Multim. | 5 |
| 2025 | Neural 3D Face Shape Stylization Based on Single Style Template via Weakly Supervised Learningabstract3D Face shape stylization refers to transforming a realistic 3D face shape into a different style, such as a cartoon face style. To solve this problem, this paper proposes modeling this task as a deformation transfer problem. This approach significantly reduces labor costs, as the artists would only need to create a single template for each face style. Realistic facial features of the original 3D face e.g. the nose or chin shape, would thus be automatically transferred to those in the style template. Deformation transfer methods, however, have two drawbacks. They are slow and they require re-optimization for every new input face. To address these weaknesses, we propose a neural network-based 3D face shape stylization method. This method is trained through weakly supervised learning, and its template's structure is preserved using our novel template-guided mesh smoothing regularization. Our method is the first learning-based deformation transfer method for 3D face shape stylization. Its employment offers the useful and practical benefit of not requiring paired training data. The experiments show that the quality of the stylized faces obtained by our method is comparable to that of the traditional deformation transfer method, achieving an average Chamfer Distance of approximately 0.01 mm. However, our approach significantly boosts the processing speed, achieving a rate approximately 3,000 times faster than the traditional deformation transfer. Peizhi Yan, Rabab K. Ward, Qiang Tang 0002, Shan Du 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Hear You Say You: An Efficient Framework for Marine Mammal Sounds' ClassificationabstractMarine mammals and their ecosystem face significant threats from, for example, military active sonar and marine transportation. To mitigate this harm, early detection and classification of marine mammals are essential. While recent efforts have utilized spectrogram analysis and machine learning techniques, there remain challenges in their efficiency. Therefore, we propose a novel knowledge distillation framework, named XCFSMN, for this problem. We construct a teacher model that fuses the features extracted from an X-vector extractor, a DenseNet and Cross-Covariance attended compact Feed-Forward Sequential Memory Network (cFSMN). The teacher model transfers knowledge to a simpler cFSMN model through a temperature-cooling strategy for efficient learning. Compared to multiple convolutional neural network backbones and transformers, the proposed framework achieves state-of-the-art efficiency and performance. The improved model size is approximately 20 times smaller and the inference time can be 10 times shorter without affecting the model’s accuracy. Xiangrui Liu, Xiaoou Liu, Shan Du 0001, Julian Cheng 0001 |
AAAI | 3 |
| 2024 | A Novel Medical Image Fusion Framework Integrating Multi-scale Encoder-Decoder with Discrete Wavelet DecompositionabstractIn recent years, many fusion algorithms based on multi-scale transform or neural networks have been proposed to improve medical image fusion (MIF) performance. However, there is still enormous potential to explore the combination of different fusion theories. In this paper, we propose a novel MIF framework to integrate powerful feature representation abilities of the deep learning model and accurate frequency decomposition characteristics of discrete wavelet transform (DWT). Firstly, a multi-scale encoder-decoder network is well-trained to extract feature information in different scales and achieve efficient image reconstruction. In particular, DWT is introduced into each scale to decompose the extracted features into high- and low-frequency sub-bands for information preservation during down-sampling. An elaborate feature fusion process is designed to achieve multi-scale fusion while merging different frequency sub-bands. Experiment results on benchmark datasets demonstrate that the proposed fusion framework outperforms current state-of-the-art methods with comparable time complexity in both objective and subjective evaluation. Renhe Liu, Yu Liu 0004, Shan Du 0001 |
ICASSP | 5 |
| 2024 | Explored seeds generation for weakly supervised semantic segmentation
Terence Chow, Haojin Deng, Yimin Yang 0001, Zhiping Lin 0001, Huiping Zhuang, Shan Du 0001 |
Neural Comput. Appl. | 6 |
| 2024 | Spatiotemporal multi-scale bilateral motion network for gait recognition
Xinnan Ding, Shan Du 0001, Yu Zhang 0223 |
J. Supercomput. | 2 |
| 2023 | A Visible and Infrared Image Fusion Framework Based on Dual-Path Encoder-Decoder and Multi-Scale Discrete Wavelet TransformabstractIn recent years, extensive research has been conducted on visible and infrared image fusion (VIF) task using traditional multi-scale transform-based and deep learning model-based methods. However, there is still a need to explore the combination of neural networks and multi-scale transform. This paper proposes a novel fusion framework based on a dual-path encoder-decoder and multi-scale transform. A dual-path encoder is trained to extract rich features at different depths from source images, while a shared decoder is trained to efficiently reconstruct images from the extracted feature space. We apply the discrete wavelet transform (DWT) to generate various frequency components from the extracted features. A fusion module is utilized to achieve fusion for low and high-frequency sub-bands, respectively, which is constrained by a gradient-based fusion loss function and an absolute values maximum-selection strategy. Our proposed method is superior to current state-of-the-art fusion methods, as demonstrated through quantitative and qualitative comparisons of publicly available datasets. Renhe Liu, Shan Du 0001, Yu Liu 0004 |
ICIP | 3 |
| 2023 | Learning Disentangled Features for Nerf-Based Face ReconstructionabstractThe 3D-aware parametric face model named HeadNeRF achieved advantages in rendering photo-realistic face images. However, it has two limitations: (1) it uses single-image fitting reconstruction that is slow and prone to overfitting; (2) it lacks explicit 3D geometry information, making using semantic facial-parts-based loss challenging. This paper presents a 3D-aware face reconstruction learning framework tailored for HeadNeRF to address the limitations. We train a face encoder network that can directly learn the disentangled features for facial reconstruction to address the first limitation. For the second limitation, we introduce a lightweight semantic face segmentation network and facial-parts-based loss function to improve the reconstruction accuracy and quality. Our experiments show that the proposed method achieves a low reconstruction time consumption and enhanced reconstruction accuracy. Project page: https://peizhiyan.github.io/docs/headnerf+ Peizhi Yan, Rabab K. Ward, Dan Wang 0011, Qiang Tang 0002, Shan Du 0001 |
ICIP | 5 |
| 2022 | A Multimodal Fusion-Based LNG Detection for Monitoring Energy Facilities (Student Abstract)abstractFossil energy products such as liquefied natural gas (LNG) are among Canada's most important exports. Canadian engineers devote themselves to constructing visual surveillance systems for detecting potential LNG emissions in energy facilities. Beyond the previous infrared (IR) surveillance system, in this paper, a multimodal fusion-based LNG detection (MFLNGD) framework is proposed to enhance the detection quality by the integration of IR and visible (VI) cameras. Besides, a Fourier transformer is developed to fuse IR and VI features better. The experimental results suggest the effectiveness of the proposed framework. Junchi Bin, Choudhury A. Rahman, Shane Rogers, Shan Du 0001, Zheng Liu 0002 |
AAAI | 4 |
| 2022 | NEO-3DF: Novel Editing-Oriented 3D Face Creation and Reconstruction
Peizhi Yan, James Gregson, Qiang Tang 0002, Rabab K. Ward, Shan Du 0001 |
ACCV (1) | 6 |
| 2022 | A Lightweight Self-Supervised Training Framework for Monocular Depth EstimationabstractDepth estimation attracts great interest in various sectors such as robotics, human computer interfaces, intelligent visual surveillance, and wearable augmented reality gear. Monocular depth estimation is of particular interest due to its low complexity and cost. Research in recent years was shifted away from supervised learning towards unsupervised or self-supervised approaches. While there have been great achievements, most of the research has focused on large heavy networks which are highly resource intensive that makes them unsuitable for systems with limited resources. We are particularly concerned about the increased complexity during training that current self-supervised approaches bring. In this paper, we propose a lightweight self-supervised training framework which utilizes computationally cheap methods to compute ground truth approximations. In particular, we utilize a stereo pair of images during training which are used to compute photometric reprojection loss and a disparity ground truth approximation. Due to the ground truth approximation, our framework is able to remove the need of pose estimation and the corresponding heavy prediction networks that current self-supervised methods have. In the experiments, we have demonstrated that our framework is capable of increasing the generator’s performance at a fraction of the size required by the current state-of-the-art self-supervised approach. Tim Heydrich, Yimin Yang 0001, Shan Du 0001 |
ICASSP | 3 |
| 2022 | A Novel Lightweight Network for Fast Monocular Depth EstimationabstractDepth estimation is of growing interest in many sectors, from robotics to wearable augmented reality gears. Monocular depth estimation attracts more attention due to its cost efficiency and low complexity. Most recent research has developed very large and resource intensive networks which are not suitable for small systems with limited resources. In this paper, we propose a lightweight network which leverages the advantages of dimension-wise convolutions and depthwise separable convolutions to reduce complexity in the architecture. In particular, the proposed depth estimation architecture utilizes a novel DICE unit-based encoder, optimized for a lightweight encoder-decoder structure. Furthermore, we propose a DICE unit-based decoder structure as well as an optimized depthwise separable convolution-based decoder. Both decoders follow a similar five-layer architecture. In the experiments, we have demonstrated the effectiveness of the proposed architecture as well as the comparison between the two proposed decoders. Our novel lightweight network has a significant decrease in both size and complexity at a marginal cost to accuracy when compared to other state-of-the-art lightweight networks. Tim Heydrich, Yimin Yang 0001, Yu Liu 0004, Shan Du 0001 |
ICASSP | 5 |
| 2022 | Coverless Information Hiding Based on Probability Graph Learning for Secure Communication in IoT EnvironmentabstractTo securely transmit secret data between Internet of Things (IoT) nodes, it is required to the implement information hiding technique for secure communication in the IoT environment. The traditional information hiding approaches generally select a multimedia file, such as texts, images, and video clips as the cover, and then embed secret information into the cover by slight modification. However, it is not feasible to directly apply these approaches in the IoT environment for the following reasons. First, it is hard for some IoT nodes to effectively and efficiently process and transmit the complex multimedia data. Second, the modification trace left in the cover will cause the presence of hidden secret information to be easily exposed by steganalysis tools. To address the above issues, we propose a coverless information hiding scheme based on probability graph learning for secure communication in the IoT environment. Instead of modifying an existing multimedia cover, we conceal secret information in a generated sequence of IoT data to realize secure communication between different nodes in the IoT environment. According to the node-data interaction relationships, we first learn the transition probability graph (TPG) to describe the transition probabilities between IoT data elements. Then, guided by a given secret message that needs to be hidden, we sequentially select a set of highly correlated data elements from the TPG to generate the sequence. The experimental results and theoretical analysis demonstrate that the proposed information hiding scheme can achieve high hiding capacity with desirable imperceptibility and security performances in the IoT environment. Zhili Zhou 0001, Yuecheng Su, Yulan Zhang, Zhihua Xia, Shan Du 0001, Brij B. Gupta, Lianyong Qi |
IEEE Internet Things J. | 5 |
| 2021 | Non-iterative online sequential learning strategy for autoencoder and classifier
Adhri Nandini Paul, Peizhi Yan, Yimin Yang 0001, Hui Zhang 0023, Shan Du 0001, Q. M. Jonathan Wu |
Neural Comput. Appl. | 5 |
| 2021 | A Novel Defensive Strategy for Facial Manipulation Detection Combining Bilateral Filtering and Joint Adversarial TrainingabstractFacial manipulation enables facial expressions to be tampered with or facial identities to be replaced in videos. The fake videos are so realistic that they are even difficult for human eyes to distinguish. This poses a great threat to social and public information security. A number of facial manipulation detectors have been proposed to address this threat. However, previous studies have shown that the accuracy of these detectors is sensitive to adversarial examples. The existing defense methods are very limited in terms of applicable scenes and defense effects. This paper proposes a new defense strategy for facial manipulation detectors, which combines a passive defense method, bilateral filtering, and a proactive defense method, joint adversarial training, to mitigate the vulnerability of facial manipulation detectors against adversarial examples. The bilateral filtering method is applied in the preprocessing stage of the model without any modification to denoise the input adversarial examples. The joint adversarial training starts from the training stage of the model, which mixes various adversarial examples and original examples to train the model. The introduction of joint adversarial training can train a model that defends against multiple adversarial attacks. The experimental results show that the proposed defense strategy positively helps facial manipulation detectors counter adversarial examples. Bin Weng 0001, Shan Du 0001 |
Secur. Commun. Networks | 4 |
| 2021 | An Efficient Q -Algorithm for RFID Tag AnticollisionabstractIn large‐scale Internet of Things (IoT) applications, tags are attached to items, and users use a radiofrequency identification (RFID) reader to quickly identify tags and obtain the corresponding item information. Since multiple tags share the same channel to communicate with the reader, when they respond simultaneously, tag collision will occur, and the reader cannot successfully obtain the information from the tag. To cope with the tag collision problem, ultrahigh frequency (UHF) RFID standard EPC G1 Gen2 specifies an anticollision protocol to identify a large number of RFID tags in an efficient way. The Q‐algorithm has attracted much more attention as the efficiency of an EPC C1 Gen2‐based RFID system can be significantly improved by only a slight adjustment to the algorithm. In this paper, we propose a novel Q‐algorithm for RFID tag identification, namely, HTEQ, which optimizes the time efficiency of an EPC C1 Gen2‐based RFID system to the utmost limit. Extensive simulations verify that our proposed HTEQ is exceptionally expeditious compared to other algorithms, which promises it to be competitive in large‐scale IoT environments. Lukun Wang, Shan Du 0001 |
Wirel. Commun. Mob. Comput. | 3 |
| 2019 | A Lightweight Neural Network For Crowd Analysis Of Images With Congested ScenesabstractFor images with congested scenes, the task of crowd analysis, including crowd counting and crowd distribution prediction, becomes very difficult. To address these issues, various CNN-based approaches have been proposed. However, those methods usually have a large number of parameters and require huge computing resources. In this paper, we focus on low-complexity approaches and propose a lightweight endto-end network for crowd analysis. Our method utilizes an effective scale-aware module to extract multi-scale features and then regresses these features to density maps. The proposed network is consisted by three parts: multi-scale feature extraction, density map estimation and density map correction, and the network which only contains 0.86 M parameters (Lightweight). According to our experiments, our proposal obtain a better result than other existing methods on several testing sequences. Shan Du 0001, Yu Liu 0004 |
ICIP | 2 |
| 2013 | Automatic License Plate Recognition (ALPR): A State-of-the-Art ReviewabstractAutomatic license plate recognition (ALPR) is the extraction of vehicle license plate information from an image or a sequence of images. The extracted information can be used with or without a database in many applications, such as electronic payment systems (toll payment, parking fee payment), and freeway and arterial monitoring systems for traffic surveillance. The ALPR uses either a color, black and white, or infrared camera to take images. The quality of the acquired images is a major factor in the success of the ALPR. ALPR as a real-life application has to quickly and successfully process license plates under different environmental conditions, such as indoors, outdoors, day or night time. It should also be generalized to process license plates from different nations, provinces, or states. These plates usually contain different colors, are written in different languages, and use different fonts; some plates may have a single color background and others have background images. The license plates can be partially occluded by dirt, lighting, and towing accessories on the car. In this paper, we present a comprehensive review of the state-of-the-art techniques for ALPR. We categorize different ALPR techniques according to the features they used for each stage, and compare them in terms of pros, cons, recognition accuracy, and processing speed. Future forecasts of ALPR are given at the end. Shan Du 0001, Mahmoud Ibrahim, Mohamed S. Shehata, Wael Badawy |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Adaptive Region-Based Image Enhancement Method for Robust Face Recognition Under Variable Illumination ConditionsabstractVariable illumination conditions, especially the side lighting effects in face images, form a main obstacle in face recognition systems. To deal with this problem, this paper presents a novel adaptive region-based image preprocessing scheme that enhances face images and facilitates the illumination invariant face recognition task. The proposed method first segments an image into different regions according to its different local illumination conditions, then both the contrast and the edges are enhanced regionally so as to alleviate the side lighting effect. Different from existing contrast enhancement methods, we apply the proposed adaptive region-based histogram equalization on the low-frequency coefficients to minimize illumination variations under different lighting conditions. Besides contrast enhancement, by observing that under poor illuminations the high-frequency features become more important in recognition, we propose enlarging the high-frequency coefficients to make face images more distinguishable. This procedure is called edge enhancement (EdgeE). The EdgeE is also region-based. Compared with existing image preprocessing methods, our method is shown to be more suitable for dealing with uneven illuminations in face images. Experimental results on the representative databases, the Yale B+Extended Yale B database and the Carnegie Mellon University-Pose, Illumination, and Expression database, show that the proposed method significantly improves the performance of face images with illumination variations. The proposed method does not require any modeling and model fitting steps and can be implemented easily. Moreover, it can be applied directly to any single image without using any lighting assumption, and any prior information on 3-D face geometry. Shan Du 0001, Rabab K. Ward |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Component-wise pose normalization for pose-invariant face recognitionabstractThe pose variation involved in facial images significantly degrades the performance of face recognition systems. In this paper, a component-wise pose normalization method for facilitating pose-invariant face recognition is proposed. The main idea is to normalize a non-frontal facial image to a virtual frontal image component by component. In this method, we first partition the whole non-frontal facial image into different facial components and then the virtual frontal view for each component is estimated separately. The final virtual frontal image is generated by integrating the virtual frontal components. The proposed method relies only on 2D images, therefore complex 3D modeling is not needed. The experimental results using the CMU-PIE database demonstrate the advantages of the proposed method. Shan Du 0001, Rabab K. Ward |
ICASSP | 1 |
| 2009 | Improved Face Representation by Nonuniform Multilevel Selection of Gabor Convolution FeaturesabstractGabor wavelets are widely employed in face representation to decompose face images into their spatial-frequency domains. The Gabor wavelet transform, however, introduces very high dimensional data. To reduce this dimensionality, uniform sampling of Gabor features has traditionally been used. Since uniform sampling equally treats all the features, it can lead to a loss of important features while retaining trivial ones. In this paper, we propose a new face representation method that employs nonuniform multilevel selection of Gabor features. The proposed method is based on the local statistics of the Gabor features and is implemented using a coarse-to-fine hierarchical strategy. Gabor features that correspond to important face regions are automatically selected and sampled finer than other features. The nonuniformly extracted Gabor features are then classified using principal component analysis and/or linear discriminant analysis for the purpose of face recognition. To verify the effectiveness of the proposed method, experiments have been conducted on benchmark face image databases where the images vary in illumination, expression, pose, and scale. Compared with the methods that use the original gray-scale image with 4096-dimensional data and uniform sampling with 2560-dimensional data, the proposed method results in a significantly higher recognition rate, with a substantial lower dimension of around 700. The experimental results also show that the proposed method works well not only when multiple sample images are available for training but also when only one sample image is available for each person. The proposed face representation method has the advantages of low complexity, low dimensionality, and high discriminance. Shan Du 0001, Rabab K. Ward |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2007 | A Robust Approach for Eye Localization Under Variable IlluminationsabstractIllumination variation is a main obstacle in facial feature detection. This paper presents a novel automated approach that localizes eyes in gray-scale face images and that is robust to illumination changes. The approach does not require prior knowledge about face orientation and illumination strength. Other advantages are that no initialization and training process are needed. Based on an edge map obtained via multi-resolution wavelet transform, this approach first segments an image into different inhomogeneously illuminated regions. The illumination of every region is then adjusted so that the features' details are more pronounced. To locate the different facial features, for every region, Gabor-based image is constructed from the re-lit image. The eyes sub-regions are then identified using the edge map of the re-lit image. This method has been applied successfully to the images of the Yale B face database that have different illuminations. Shan Du 0001, Rabab K. Ward |
ICIP (1) | 1 |
| 2006 | Adaptive Region-Based Image Enhancement Method for Face Recognition Under Varying Illumination ConditionsabstractIllumination changes in face images form a main obstacle in face recognition systems. To deal with this problem, this study presents a novel adaptive region-based image preprocessing scheme that enhances face images and facilitates the face recognition task. This method enhances both the edges and the contrast in face images regionally so as to alleviate the side lighting effects. Compared with the conventional global histogram equalization method, our method is shown to be more suitable for dealing with uneven illuminations in face images. This method is evaluated on the Yale B face database. The experimental results show the advantages of the proposed method with an improvement of 16.1% on average over the histogram eualization method. Shan Du 0001, Rabab K. Ward |
ICASSP (2) | 1 |
| 2005 | Statistical Non-Uniform Sampling of Gabor Wavelet Coefficients for Face RecongnitionabstractA statistics based, non-uniform sampling of the Gabor wavelet decomposition coefficients for face recognition is presented in this paper. Gabor wavelets are popularly used to decompose face images into their spatial/frequency domains. The derived Gabor coefficients generate an augmented vector, e.g., 40 times larger than the original gray-scale vector. To reduce the dimensionality, uniform sampling of the Gabor coefficients is normally used. In this paper, we propose a non-uniform sampling method of the Gabor coefficients such that the coefficients corresponding to important face features are sampled much finer than those of the other parts of the image. The non-uniform sampling is based on the local statistics of the Gabor coefficients obtained from a set of training images. This adaptation is implemented in a hierarchical fashion; a coarse-to-fine strategy results in multi-level sampling rates. After the samples are obtained, the traditional principal component analysis (PCA) is used to code the samples for the final classification. The experimental results show that the proposed non-uniform sampling of Gabor coefficients outperforms the uniform one and the popular eigenfaces method. Shan Du 0001, Rabab K. Ward |
ICASSP (2) | 1 |
| 2005 | Wavelet-based illumination normalization for face recognitionabstractThe appearance of a face image is severely affected by illumination conditions that hinder the automatic face recognition process. To recognize faces under varying illuminations, we propose a wavelet-based normalization method so as to normalize illuminations. This method enhances the contrast as well as the edges of face images simultaneously, in the frequency domain using the wavelet transform, to facilitate face recognition tasks. It outperforms the conventional illumination normalization method - the histogram equalization that only enhances image pixel gray-level contrast in the spatial domain. With this method, our face recognition system works effectively under a wide range of illumination conditions. The experimental results obtained by testing on the Yale face database B demonstrate the effectiveness of our method with 15.65% improvement, on average, in the face recognition system. Shan Du 0001, Rabab K. Ward |
ICIP (2) | 1 |