EDBT 2026 Demo / reviewers in the wild / expert
Marios Savvides
dblp:13/3793
· DBLP profile ↗
127ranked-venue papers
4as first author
20since 2021 · last 2026
0009-0003-6534-4269ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 91 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 78 · 1 first-author · 18 since 2021Security and privacy · 12 · 1 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in VisionabstractVision-language models (VLMs) have made significant strides in reasoning, yet they often struggle with complex multimodal tasks and tend to generate overly verbose outputs. A key limitation is their reliance on chain-of-thought (CoT) reasoning, despite many tasks benefiting from alternative topologies like trees or graphs. To address this, we introduce STELAR-Vision, a training framework for topology-aware reasoning. At its core is TopoAug, a synthetic data pipeline that enriches training with diverse topological structures. Using supervised fine-tuning and reinforcement learning, we post-train Qwen2VL models with both accuracy and efficiency in mind. Additionally, we propose Frugal Learning, which reduces output length with minimal accuracy loss. On MATH-V and VLM_S2H, STELAR-Vision improves accuracy by 9.7% over its base model and surpasses the larger Qwen2VL-72B-Instruct by 7.3%. On five out-of-distribution benchmarks, it outperforms Phi-4-Multimodal-Instruct by up to 28.4% and LLaMA-3.2-11B-Vision-Instruct by up to 13.2%, demonstrating strong generalization. Compared to Chain-Only training, our approach achieves 4.3% higher overall accuracy on in-distribution datasets and consistently outperforms across all OOD benchmarks. Han Zhang 0048, Zhantao Yang, Fangyi Chen, Anudeepsekhar Bolimera, Marios Savvides |
AAAI | 7 |
| 2025 | The Second Workshop on Generative AI for E-commerce
Mansi Ranjit Mane, Djordje Gligorijevic, Dingxian Wang, Topojoy Biswas, Evren Körpeoglu, Marios Savvides, Yongfeng Zhang 0003, Julian J. McAuley |
RecSys | 6 |
| 2024 | Workshop on Generative AI for E-commerceabstractThe "Gen AI for E-commerce" workshop explores the role of Generative Artificial Intelligence in transforming e-commerce through enhanced user experience and operational efficiency. E-commerce companies grapple with multiple challenges such as lack of quality content for products, subpar user experience, sparse datasets etc. Gen AI offers significant potential to address these complexities. Yet, deploying these technologies at scale presents challenges such as hallucination in data, excessive costs, increased latency response, and limited generalization in sparse data environments. This workshop will bring together experts from academia and industry to discuss these challenges and opportunities, aiming to showcase case studies, breakthroughs, and insights into practical implementations of Gen AI in e-commerce. Mansi Ranjit Mane, Djordje Gligorijevic, Dingxian Wang, Behzad Shahrasbi, Topojoy Biswas, Evren Körpeoglu, Marios Savvides |
CIKM | 7 |
| 2024 | A Reference-Based 3D Semantic-Aware Framework for Accurate Local Facial Attribute EditingabstractFacial attribute editing plays a crucial role in synthesizing realistic faces with specific characteristics while maintaining realistic appearances. Despite advancements, challenges persist in achieving precise, 3D-aware attribute modifications, which are crucial for consistent and accurate representations of faces from different angles. Current methods struggle with semantic entanglement and lack effective guidance for incorporating attributes while maintaining image integrity. To address these issues, we introduce a novel framework that merges the strengths of latent-based and reference-based editing methods. Our approach employs a 3D GAN inversion technique to embed attributes from the reference image into a tri-plane space, ensuring 3D consistency and realistic viewing from multiple perspectives. We utilize blending techniques and predicted semantic masks to locate precise edit regions, merging them with the contextual guidance from the reference image. A coarse-to-fine inpainting strategy is then applied to preserve the integrity of untargeted areas, significantly enhancing realism. Our evaluations demonstrate superior performance across diverse editing tasks, validating our framework’s effectiveness in realistic and applicable facial attribute editing. Yutong Zheng, Yen-Shuo Su, Anudeepsekhar Bolimera, Han Zhang 0048, Fangyi Chen, Marios Savvides |
IJCB | 7 |
| 2023 | Enhanced Training of Query-Based Object Detection via Selective Query RecollectionabstractThis paper investigates a phenomenon where query-based object detectors mispredict at the last decoding stage while predicting correctly at an intermediate stage. We review the training process and attribute the overlooked phenomenon to two limitations: lack of training emphasis and cascading errors from decoding sequence. We design and present Selective Query Recollection (SQR), a simple and effective training strategy for query-based object detectors. It cumulatively collects intermediate queries as decoding stages go deeper and selectively forwards the queries to the downstream stages aside from the sequential structure. Such-wise, SQR places training emphasis on later stages and allows later stages to work with intermediate queries from earlier stages directly. SQR can be easily plugged into various query-based object detectors and significantly enhances their performance while leaving the inference pipeline unchanged. As a result, we apply SQR on Adamixer, DAB-DETR, and Deformable-DETR across various settings (backbone, number of queries, schedule) and consistently brings 1.4 ~2.8 AP improvement. Code is available at https://github.com/Fangyi-Chen/SQR Fangyi Chen, Han Zhang 0048, Kai Hu 0010, Chenchen Zhu, Marios Savvides |
CVPR | 6 |
| 2023 | Boosting Transductive Few-Shot Fine-tuning with Margin-based Uncertainty Weighting and Probability RegularizationabstractFew-Shot Learning (FSL) has been rapidly developed in recent years, potentially eliminating the requirement for significant data acquisition. Few-shot fine-tuning has been demonstrated to be practically efficient and helpful, especially for out-of-distribution datum [7, 13, 17, 29]. In this work, we first observe that the few-shot fine-tuned methods are learned with the imbalanced class marginal distribution, leading to imbalanced per-class testing accuracy. This observation further motivates us to propose the Transductive Fine-tuning with Margin-based uncertainty weighting and Probability regularization (TF-MP), which learns a more balanced class marginal distribution as shown in Fig. 1. We first conduct sample weighting on unlabeled testing data with margin-based uncertainty scores and fur-ther regularize each test sample's categorical probability. TF-MP achieves state-of-the-art performance on in-/out-of-distribution evaluations of Meta- Dataset [31] and sur-passes previous transductive methods by a large margin. Ran Tao 0013, Hao Chen 0102, Marios Savvides |
CVPR | 3 |
| 2023 | Cov Loss: Covariance-Based Loss for Deep Face RecognitionabstractRecently, deep neural networks (DNNs) have emerged as state-of-the-art approaches for various computer vision areas. In this paper, we propose an optimized approach for large-scale face recognition. Our work is motivated through the recent development of deep convolutional neural networks (CNNs) that use different loss functions to learn deep features from face images to perform face recognition. As opposed to previous works, we model Cov loss to optimize deep features along the covariance matrix to enhance discriminative power. We formulate Cov loss to maximize inter-class variance and minimize intra-class variance by optimizing the distance between the deep features and their corresponding class covariances in the Euclidean space. The proposed Cov loss is evaluated on large-scale face recognition problems and present results on LFW, IJB-A Janus, IJB-C Janus, and Celebrity Frontal-Profile (CFP). By optimizing features along both Euclidean and angular spaces, our novel loss function learns more robust feature representations of faces, and improves general performance results comparable to the state-of-the-art results. Ibrahim Alkanhal, Abdullah Almansour, Lamia Alsalloom, Raied Aljadaany, Marios Savvides |
ICASSP | 5 |
| 2023 | SoftMatch: Addressing the Quantity-Quality Tradeoff in Semi-supervised Learning
Hao Chen 0102, Ran Tao 0013, Yidong Wang 0003, Jindong Wang 0001, Bernt Schiele, Xing Xie 0001, Bhiksha Raj, Marios Savvides |
ICLR | 9 |
| 2023 | FreeMatch: Self-adaptive Thresholding for Semi-supervised Learning
Yidong Wang 0003, Hao Chen 0102, Qiang Heng, Wenxin Hou, Zhen Wu 0002, Jindong Wang 0001, Marios Savvides, Takahiro Shinozaki, Bhiksha Raj, Bernt Schiele, Xing Xie 0001 |
ICLR | 8 |
| 2022 | Un-mix: Rethinking Image Mixtures for Unsupervised Visual Representation LearningabstractThe recently advanced unsupervised learning approaches use the siamese-like framework to compare two "views" from the same image for learning representations. Making the two views distinctive is a core to guarantee that unsupervised methods can learn meaningful information. However, such frameworks are sometimes fragile on overfitting if the augmentations used for generating two views are not strong enough, causing the over-confident issue on the training data. This drawback hinders the model from learning subtle variance and fine-grained information. To address this, in this work we aim to involve the soft distance concept on label space in the contrastive-based unsupervised learning task and let the model be aware of the soft degree of similarity between positive or negative pairs through mixing the input data space, to further work collaboratively for the input and loss spaces. Despite its conceptual simplicity, we show empirically that with the solution -- Unsupervised image mixtures (Un-Mix), we can learn subtler, more robust and generalized representations from the transformed input and corresponding new label space. Extensive experiments are conducted on CIFAR-10, CIFAR-100, STL-10, Tiny ImageNet and standard ImageNet-1K with popular unsupervised methods SimCLR, BYOL, MoCo V1&V2, SwAV, etc. Our proposed image mixture and label assignment strategy can obtain consistent improvement by 1~3% following exactly the same hyperparameters and training procedures of the base methods. Code is publicly available at https://github.com/szq0214/Un-Mix. Zechun Liu, Zhuang Liu 0003, Marios Savvides, Trevor Darrell, Eric P. Xing |
AAAI | 4 |
| 2022 | Powering Finetuning in Few-Shot Learning: Domain-Agnostic Bias Reduction with Selected SamplingabstractIn recent works, utilizing a deep network trained on meta-training set serves as a strong baseline in few-shot learning. In this paper, we move forward to refine novel-class features by finetuning a trained deep network. Finetuning is designed to focus on reducing biases in novel-class feature distributions, which we define as two aspects: class-agnostic and class-specific biases. Class-agnostic bias is defined as the distribution shifting introduced by domain difference, which we propose Distribution Calibration Module(DCM) to reduce. DCM owes good property of eliminating domain difference and fast feature adaptation during optimization. Class-specific bias is defined as the biased estimation using a few samples in novel classes, which we propose Selected Sampling(SS) to reduce. Without inferring the actual class distribution, SS is designed by running sampling using proposal distributions around support-set samples. By powering finetuning with DCM and SS, we achieve state-of-the-art results on Meta-Dataset with consistent performance boosts over ten datasets from different domains. We believe our simple yet effective method demonstrates its possibility to be applied on practical few-shot applications. Ran Tao 0013, Han Zhang 0048, Yutong Zheng, Marios Savvides |
AAAI | 4 |
| 2022 | Unitail: Detecting, Reading, and Matching in Retail Scene
Fangyi Chen, Han Zhang 0048, Zaiwang Li, Jiachen Dou, Shentong Mo, Hao Chen 0102, Uzair Ahmed, Chenchen Zhu, Marios Savvides |
ECCV (7) | 10 |
| 2022 | USB: A Unified Semi-supervised Learning Benchmark for ClassificationabstractSemi-supervised learning (SSL) improves model generalization by leveraging massive unlabeled data to augment limited labeled samples. However, currently, popular SSL evaluation protocols are often constrained to computer vision (CV) tasks. In addition, previous work typically trains deep neural networks from scratch, which is time-consuming and environmentally unfriendly. To address the above issues, we construct a Unified SSL Benchmark (USB) for classification by selecting 15 diverse, challenging, and comprehensive tasks from CV, natural language processing (NLP), and audio processing (Audio), on which we systematically evaluate the dominant SSL methods, and also open-source a modular and extensible codebase for fair evaluation of these SSL methods. We further provide the pre-trained versions of the state-of-the-art neural models for CV tasks to make the cost affordable for further tuning. USB enables the evaluation of a single SSL algorithm on more tasks from multiple domains but with less cost. Specifically, on a single NVIDIA V100, only 39 GPU days are required to evaluate FixMatch on 15 tasks in USB while 335 GPU days (279 GPU days on 4 CV datasets except for ImageNet) are needed on 5 CV tasks with TorchSSL. Yidong Wang 0003, Hao Chen 0102, Wang Sun, Ran Tao 0013, Wenxin Hou, Linyi Yang, Zhi Zhou 0007, Lan-Zhe Guo, Heli Qi, Zhen Wu 0002, Yufeng Li 0008, Satoshi Nakamura 0001, Wei Ye 0004, Marios Savvides, Bhiksha Raj, Takahiro Shinozaki, Bernt Schiele, Jindong Wang 0001, Xing Xie 0001, Yue Zhang 0004 |
NeurIPS | 16 |
| 2021 | Partial Is Better Than All: Revisiting Fine-tuning Strategy for Few-shot LearningabstractThe goal of few-shot learning is to learn a classifier that can recognize unseen classes from limited support data with labels. A common practice for this task is to train a model on the base set first and then transfer to novel classes through fine-tuning or meta-learning. However, as the base classes have no overlap to the novel set, simply transferring whole knowledge from base data is not an optimal solution since some knowledge in the base model may be biased or even harmful to the novel class. In this paper, we propose to transfer partial knowledge by freezing or fine-tuning particular layer(s) in the base model. Specifically, layers will be imposed different learning rates if they are chosen to be fine-tuned, to control the extent of preserved transferability. To determine which layers to be recast and what values of learning rates for them, we introduce an evolutionary search based method that is efficient to simultaneously locate the target layers and determine their individual learning rates. We conduct extensive experiments on CUB and mini-ImageNet to demonstrate the effectiveness of our proposed method. It achieves the state-of-the-art performance on both meta-learning and non-meta based frameworks. Furthermore, we extend our method to the conventional pre-training + fine-tuning paradigm and obtain consistent improvement. Zechun Liu, Jie Qin 0004, Marios Savvides, Kwang-Ting Cheng |
AAAI | 4 |
| 2021 | S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-Bit Neural Networks via Guided Distribution CalibrationabstractPrevious studies dominantly target at self-supervised learning on real-valued networks and have achieved many promising results. However, on the more challenging binary neural networks (BNNs), this task has not yet been fully explored in the community. In this paper, we focus on this more difficult scenario: learning networks where both weights and activations are binary, meanwhile, without any human annotated labels. We observe that the commonly used contrastive objective is not satisfying on BNNs for competitive accuracy, since the backbone network contains relatively limited capacity and representation ability. Hence instead of directly applying existing self-supervised methods, which cause a severe decline in performance, we present a novel guided learning paradigm from real-valued to distill binary networks on the final prediction distribution, to minimize the loss and obtain desirable accuracy. Our proposed method can boost the simple contrastive learning baseline by an absolute gain of 5.5∼15% on BNNs. We further reveal that it is difficult for BNNs to recover the similar predictive distributions as real-valued models when training without labels. Thus, how to calibrate them is key to address the degradation in performance. Extensive experiments are conducted on the large-scale ImageNet and downstream datasets. Our method achieves substantial improvement over the simple contrastive learning baseline, and is even comparable to many mainstream supervised BNN methods. Code is available at https://github.com/szq0214/S2-BNN. Zechun Liu, Jie Qin 0004, Lei Huang 0015, Kwang-Ting Cheng, Marios Savvides |
CVPR | 6 |
| 2021 | Unsupervised Disentanglement of Linear-Encoded Facial SemanticsabstractWe propose a method to disentangle linear-encoded facial semantics from StyleGAN without external supervision. The method derives from linear regression and sparse representation learning concepts to make the disentangled latent representations easily interpreted as well. We start by coupling StyleGAN with a stabilized 3D deformable facial reconstruction method to decompose single-view GAN generations into multiple semantics. Latent representations are then extracted to capture interpretable facial semantics. In this work, we make it possible to get rid of labels for disentangling meaningful facial semantics. Also, we demonstrate that the guided extrapolation along the disentangled representations can help with data augmentation, which sheds light on handling unbalanced data. Finally, we provide an analysis of our learned localized facial representations and illustrate that the semantic information is encoded, which surprisingly complies with human intuition. The overall unsupervised design brings more flexibility to representation learning in the wild. Yutong Zheng, Ran Tao 0013, Marios Savvides |
CVPR | 5 |
| 2021 | Semantic Relation Reasoning for Shot-Stable Few-Shot Object DetectionabstractFew-shot object detection is an imperative and long-lasting problem due to the inherent long-tail distribution of real-world data. Its performance is largely affected by the data scarcity of novel classes. But the semantic relation between the novel classes and the base classes is constant regardless of the data availability. In this work, we investigate utilizing this semantic relation together with the visual information and introduce explicit relation reasoning into the learning of novel object detection. Specifically, we represent each class concept by a semantic embedding learned from a large corpus of text. The detector is trained to project the image representations of objects into this embedding space. We also identify the problems of trivially using the raw embeddings with a heuristic knowledge graph and propose to augment the embeddings with a dynamic relation graph. As a result, our few-shot detector, termed SRR-FSD, is robust and stable to the variation of shots of novel objects. Experiments show that SRR-FSD can achieve competitive results at higher shots, and more importantly, a significantly better performance given both lower explicit and implicit shots. The benchmark protocol with implicit shots removed from the pretrained classification dataset can serve as a more realistic setting for future research. Chenchen Zhu, Fangyi Chen, Uzair Ahmed, Marios Savvides |
CVPR | 5 |
| 2021 | Contrast and Order Representations for Video Self-supervised LearningabstractThis paper studies the problem of learning self-supervised representations on videos. In contrast to image modality that only requires appearance information on objects or scenes, video needs to further explore the relations between multiple frames/clips along the temporal dimension. However, the recent proposed contrastive-based self-supervised frameworks do not grasp such relations explicitly since they simply utilize two augmented clips from the same video and compare their distance without referring to their temporal relation. To address this, we present a contrast-and-order representation (CORP) framework for learning self-supervised video representations that can automatically capture both the appearance information within each frame and temporal information across different frames. In particular, given two video clips, our model first predicts whether they come from the same input video, and then predict the temporal ordering of the clips if they come from the same video. We also propose a novel decoupling attention method to learn symmetric similarity (contrast) and anti-symmetric patterns (order). Such design involves neither extra parameters nor computation, but can speed up the learning process and improve accuracy compared to the vanilla multi-head attention. We extensively validate the representation ability of our learned video features for the downstream action recognition task on Kinetics-400 and Something-something V2. Our method outperforms previous state-of-the-arts by a significant margin. Kai Hu 0010, Jie Shao 0006, Bhiksha Raj, Marios Savvides |
ICCV | 5 |
| 2021 | Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical Study
Zechun Liu, Dejia Xu, Zitian Chen, Kwang-Ting Cheng, Marios Savvides |
ICLR | 6 |
| 2021 | CDTD: A Large-Scale Cross-Domain Benchmark for Instance-Level Image-to-Image Translation and Domain Adaptive Object Detection
Mingyang Huang, Jianping Shi, Zechun Liu, Harsh Maheshwari, Yutong Zheng, Xiangyang Xue 0001, Marios Savvides, Thomas S. Huang |
Int. J. Comput. Vis. | 8 |
| 2020 | Learning Non-Parametric Invariances from Data with Permanent Random Connectomes
Dipan K. Pal, Akshay Chawla, Marios Savvides |
BMVC | 3 |
| 2020 | Towards a Hypothesis on Visual Transformation based Self-Supervision
Dipan K. Pal, Sreena Nallamothu, Marios Savvides |
BMVC | 3 |
| 2020 | Binarizing MobileNet via Evolution-Based SearchingabstractBinary Neural Networks (BNNs), known to be one among the effectively compact network architectures, have achieved great outcomes in the visual tasks. Designing efficient binary architectures is not trivial due to the binary nature of the network. In this paper, we propose a use of evolutionary search to facilitate the construction and training scheme when binarizing MobileNet, a compact network with separable depth-wise convolution. Being inspired by one-shot architecture search frameworks, we manipulate the idea of group convolution to design efficient 1-Bit Convolutional Neural Networks (CNNs), assuming an approximately optimal trade-off between computational cost and model accuracy. Our objective is to come up with a tiny yet efficient binary neural architecture by exploring the best candidates of the group convolution while optimizing the model performance in terms of complexity and latency. The approach is threefold. First, we modify and train strong baseline binary networks with a wide range of random group combinations at each convolutional layer. This set-up gives the binary neural networks a capability of preserving essential information through layers. Second, to find a good set of hyper-parameters for group convolutions we make use of the evolutionary search which leverages the exploration of efficient 1-bit models. Lastly, these binary models are trained from scratch in a usual manner to achieve the final binary model. Various experiments on ImageNet are conducted to show that following our construction guideline, the final model achieves 60.09% Top-1 accuracy and outperforms the state-of-the-art CI-BCNN with the same computational cost. NhatHai Phan, Zechun Liu, Dang Huynh, Marios Savvides, Kwang-Ting Cheng |
CVPR | 4 |
| 2020 | ReActNet: Towards Precise Binary Neural Network with Generalized Activation Functions
Zechun Liu, Marios Savvides, Kwang-Ting Cheng |
ECCV (14) | 3 |
| 2020 | Online Ensemble Model Compression Using Knowledge Distillation
Devesh Walawalkar, Marios Savvides |
ECCV (19) | 3 |
| 2020 | Soft Anchor-Point Object Detection
Chenchen Zhu, Fangyi Chen, Marios Savvides |
ECCV (9) | 4 |
| 2020 | Attentive Cutmix: An Enhanced Data Augmentation Approach for Deep Learning Based Image ClassificationabstractConvolutional neural networks (CNN) are capable of learning robust representation with different regularization methods and activations as convolutional layers are spatially correlated. Based on this property, a large variety of regional dropout strategies have been proposed, such as Cutout [1], DropBlock [2], CutMix [3], etc. These methods aim to promote the network to generalize better by partially occluding the discriminative parts of objects. However, all of them perform this operation randomly, without capturing the most important region(s) within an object. In this paper, we propose Attentive CutMix, a naturally enhanced augmentation strategy based on CutMix [3]. In each training iteration, we choose the most descriptive regions based on the intermediate attention maps from a feature extractor, which enables searching for the most discriminative parts in an image. Our proposed method is simple yet effective, easy to implement and can boost the baseline significantly. Extensive experiments on CIFAR-10/100, ImageNet datasets with various CNN architectures (in a unified setting) demonstrate the effectiveness of our proposed method, which consistently outperforms the baseline CutMix and other methods by a significant margin. Devesh Walawalkar, Zechun Liu, Marios Savvides |
ICASSP | 4 |
| 2020 | Solving Missing-Annotation Object Detection with Background Recalibration LossabstractThis paper focuses on a novel and challenging detection scenario: A majority of true objects/instances is unlabeled in the datasets, so these missing-labeled areas will be regarded as the background during training. Previous art [1] on this problem has proposed to use soft sampling to re-weight the gradients of RoIs based on the overlaps with positive instances, while their method is mainly based on the two-stage detector (i.e. Faster RCNN) which is more robust and friendly for the missing label scenario. In this paper, we introduce a superior solution called Background Recalibration Loss (BRL) that can automatically re-calibrate the loss signals according to the pre-defined IoU threshold and input image. Our design is built on the one-stage detector which is faster and lighter. Inspired by the Focal Loss [2] formulation, we make several significant modifications to fit on the missing-annotation circumstance. We conduct extensive experiments on the curated PASCAL VOC [3] and MS COCO [4] datasets. The results demonstrate that our proposed method outperforms the baseline and other state-of-the-arts by a large margin. Han Zhang 0048, Fangyi Chen, Qiqi Hao, Chenchen Zhu, Marios Savvides |
ICASSP | 6 |
| 2020 | Offset Curves Loss for Imbalanced Problem in Medical SegmentationabstractMedical image segmentation has played an important role in medical analysis and widely developed for many clinical applications. Deep learning-based approaches have achieved high performance in semantic segmentation but they are limited to pixel-wise setting and imbalanced classes data problem. In this paper, we tackle those limitations by developing a new deep learning-based model which takes into account both higher feature level i.e. region inside contour, intermediate feature level i.e. offset curves around the contour and lower feature level i.e. contour. Our proposed Offset Curves (OsC) loss consists of three main fitting terms. The first fitting term focuses on pixel-wise level segmentation whereas the second fitting term acts as attention model which pays attention to the area around the boundaries (offset curves). The third terms plays a role as regularization term which takes the length of boundaries into account. We evaluate our proposed OsC loss on both 2D network and 3D network. Two common medical datasets, i.e. retina DRIVE and brain tumor BRATS 2018 datasets are used to benchmark our proposed loss performance. The experiments have shown that our proposed OsC loss function outperforms other mainstream loss functions such as Cross-Entropy, Dice, Focal on the most common segmentation networks Unet, FCN. T. Hoang Ngan Le, Kashu Yamazaki, Toan Duc Bui, Khoa Luu, Marios Savvides |
ICPR | 6 |
| 2020 | A Multi-task Contextual Atrous Residual Network for Brain Tumor Detection & SegmentationabstractIn recent years, deep neural networks have achieved state-of-the-art performance in a variety of recognition and segmentation tasks in medical imaging including brain tumor segmentation. We investigate that segmenting a brain tumor is facing to the imbalanced data problem where the number of pixels belonging to the background class (non tumor pixel) is much larger than the number of pixels belonging to the foreground class (tumor pixel). To address this problem, we propose a multitask network which is formed as a cascaded structure. Our model consists of two targets, i.e., (i) effectively differentiate the brain tumor regions and (ii) estimate the brain tumor mask. The first objective is performed by our proposed contextual brain tumor detection network, which plays a role of an attention gate and focuses on the region around brain tumor only while ignoring the far neighbor background which is less correlated to the tumor. Different from other existing object detection networks which process every pixel, our contextual brain tumor detection network only processes contextual regions around ground-truth instances and this strategy aims at producing meaningful regions proposals. The second objective is built upon a 3D atrous residual network and under an encode-decode network in order to effectively segment both large and small objects (brain tumor). Our 3D atrous residual network is designed with a skip connection to enables the gradient from the deep layers to be directly propagated to shallow layers, thus, features of different depths are preserved and used for refining each other. In order to incorporate larger contextual information from volume MRI data, our network utilizes the 3D atrous convolution with various kernel sizes, which enlarges the receptive field of filters. Our proposed network has been evaluated on various datasets including BRATS2015, BRATS2017 and BRATS2018 datasets with both validation set and testing set. Our performance has been benchmarked by both region-based metrics and surface-based metrics. We also have conducted comparisons against state-of-the-art approaches.11Code and models will be publicly available. T. Hoang Ngan Le, Kashu Yamazaki, Kha Gia Quach, Dat T. Truong, Marios Savvides |
ICPR | 5 |
| 2020 | MoBiNet: A Mobile Binary Network for Image ClassificationabstractMobileNet and Binary Neural Networks are two among the most widely used techniques to construct deep learning models for performing a variety of tasks on mobile and embedded platforms. In this paper, we present a simple yet efficient scheme to exploit MobileNet binarization at activation function and model weights. However, training a binary network from scratch with separable depth-wise and point-wise convolutions in case of MobileNet is not trivial and prone to divergence. To tackle this training issue, we propose a novel neural network architecture, namely MoBi- Net - Mobile Binary Network in which skip connections are manipulated to prevent information loss and vanishing gradient, thus facilitate the training process. More importantly, while existing binary neural networks often make use of cumbersome backbones such as Alex-Net, ResNet, VGG-16 with float-type pre-trained weights initialization, our MoBi- Net focuses on binarizing the already-compressed neural networks like MobileNet without the need of a pre-trained model to start with. Therefore, our proposal results in an effectively small model while keeping the accuracy comparable to existing ones. Experiments on ImageNet dataset show the potential of the MoBiNet as it achieves 54.40% top-1 accuracy and dramatically reduces the computational cost with binary operators. NhatHai Phan, Dang Huynh, Yihui He, Marios Savvides |
WACV | 4 |
| 2020 | NCMS: Towards accurate anchor free object detection through ℓ2 norm calibration and multi-feature selectionabstractWe present simple and flexible drop-in modules in feature pyramids for general object detection, which can be easily generalized to other anchor-free detectors without introducing extra parameters, and only involves negligible computational cost on training and testing. The proposed detector, called NCMS, inserts a simple norm calibration (NC) operation between the feature pyramids and detection head to alleviate and balance the norm bias caused by feature pyramid network (FPN). Furthermore, the NCMS leverages an enhanced multi-feature selective strategy (MS) during training to assign the ground-truth to particular feature pyramid levels as supervisions, in order to obtain more discriminative representation for objects. By generalizing to the state-of-the-art FSAF module (Zhu et al., 2019), our NCMS improves it by 1.6% on COCO val set without bells and whistles. The resulting best model achieves 44.0% mAP with single-model and single-scale testing, which is a fairly competitive result. Fangyi Chen, Chenchen Zhu, Han Zhang 0048, Marios Savvides |
Comput. Vis. Image Underst. | 5 |
| 2019 | Non-Parametric Transformation Networks for Learning General Invariances from Data
Dipan K. Pal, Marios Savvides |
AAAI | 2 |
| 2019 | Improving Object Detection from Scratch via Gated Feature Reuse
Humphrey Shi, NhatHai Phan, Rogério Feris, Liangliang Cao, Ding Liu 0001, Xinchao Wang, Thomas S. Huang, Marios Savvides |
BMVC | 10 |
| 2019 | Douglas-Rachford Networks: Learning Both the Image Prior and Data Fidelity Terms for Blind Image DeconvolutionabstractBlind deconvolution problems are heavily ill-posed where the specific blurring kernel is not known. Recovering these images typically requires estimates of the kernel. In this paper, we present a method called Dr-Net, which does not require any such estimate and is further able to invert the effects of the blurring in blind image recovery tasks. These image recovery problems typically have two terms, the data fidelity term (for faithful reconstruction) and the image prior (for realistic looking reconstructions). We use the Douglas-Rachford iterations to solve this problem since it is a more generally applicable optimization procedure than methods such as the proximal gradient descent algorithm. Two proximal operators originate from these iterations, one from the data fidelity term and the second from the image prior. It is non-trivial to design a hand-crafted function to represent these proximal operators for the data fidelity and the image prior terms which would work with real-world image distributions. We therefore approximate both these proximal operators using deep networks. This provides a sound motivation for the final architecture for Dr-Net which we find outperforms the state-of-the-art on two mainstream blind deconvolution benchmarks. We also find that Dr-Net is one of the fastest algorithms according to wall-clock times while doing so. Raied Aljadaany, Dipan K. Pal, Marios Savvides |
CVPR | 3 |
| 2019 | Bounding Box Regression With Uncertainty for Accurate Object DetectionabstractLarge-scale object detection datasets (e.g., MS-COCO) try to define the ground truth bounding boxes as clear as possible. However, we observe that ambiguities are still introduced when labeling the bounding boxes. In this paper, we propose a novel bounding box regression loss for learning bounding box transformation and localization variance together. Our loss greatly improves the localization accuracies of various architectures with nearly no additional computation. The learned localization variance allows us to merge neighboring bounding boxes during non-maximum suppression (NMS), which further improves the localization performance. On MS-COCO, we boost the Average Precision (AP) of VGG-16 Faster R-CNN from 23.6% to 29.1%. More importantly, for ResNet-50-FPN Mask R-CNN, our method improves the AP and AP90 by 1.8% and 6.2% respectively, which significantly outperforms previous state-of-the-art bounding box refinement methods. Our code and models are available at github.com/yihui-he/KL-Loss. Yihui He, Chenchen Zhu, Jianren Wang, Marios Savvides, Xiangyu Zhang 0005 |
CVPR | 4 |
| 2019 | Feature Selective Anchor-Free Module for Single-Shot Object DetectionabstractWe motivate and present feature selective anchor-free (FSAF) module, a simple and effective building block for single-shot object detectors. It can be plugged into single-shot detectors with feature pyramid structure. The FSAF module addresses two limitations brought up by the conventional anchor-based detection: 1) heuristic-guided feature selection; 2) overlap-based anchor sampling. The general concept of the FSAF module is online feature selection applied to the training of multi-level anchor-free branches. Specifically, an anchor-free branch is attached to each level of the feature pyramid, allowing box encoding and decoding in the anchor-free manner at an arbitrary level. During training, we dynamically assign each instance to the most suitable feature level. At the time of inference, the FSAF module can work independently or jointly with anchor-based branches. We instantiate this concept with simple implementations of anchor-free branches and online feature selection strategy. Experimental results on the COCO detection track show that our FSAF module performs better than anchor-based counterparts while being faster. When working jointly with anchor-based branches, the FSAF module robustly improves the baseline RetinaNet by a large margin under various settings, while introducing nearly free inference overhead. And the resulting best model can achieve a state-of-the-art 44.6% mAP, outperforming all existing single-shot detectors on COCO. Chenchen Zhu, Yihui He, Marios Savvides |
CVPR | 3 |
| 2019 | A Novel Collaborative Control Strategy for Enhanced Training of Vehicle RecognitionabstractDeep learning methods support vehicular technology in various aspects. How to efficiently and effectively optimize deep learning models remains a challenge. It is known that the learning rate is an important hyper-parameter to optimize models and the batch size is one of the keys to speed up training. This paper empirically studies the principles of scheduling batch size and learning rate during training, and proposes the Collaborative Control Strategy (CCS), which practically improves both classification accuracy and training speed. Instead of stepwise decreasing learning rate and keeping batch size unchanged, we asynchronously adjust batch size and learning rate based on time-division and restart policy. We study and analyze the proposed CCS on general image recognition benchmarks. Compared to traditional training strategy, without bells and whistles CCS decreases error rates by absolute 0.78% and 0.75% on CIFAR-10 and CIFAR-100 respectively, and reduces 40% of training time. For vehicular applications, we demonstrate its advantage on Stanford car-196 dataset with different architectures, showing consistent training speedup and accuracy improvement. Fangyi Chen, Chenchen Zhu, Marios Savvides |
VTC Fall | 3 |
| 2019 | Is Pose Really Solved? A Frontalization Study On Off-Angle Face MatchingabstractRecently, impressive results have been achieved on many large-scale face recognition benchmarks, such as IJB-A Janus and Janus CS3. These datasets were designed to test robustness to nuisance transformations simultaneously such as pose, illumination, expression etc. We present a study paper, where we find that despite this goal in evaluation, there exists a significant frontal bias in yaw pose in these datasets. Therefore, high-performance on these recent datasets is misleading and does not reflect robustness to extreme pose in yaw. Moreover, many real-world applications only allow for a single frontal enrollment in a gallery (law enforcement, immigration etc.). As we show in our study, face recognition in this highly constrained setting with extreme pose variation in the probe images remains a highly challenging problem. Traditional approaches, performing well on datasets such as IJB-A Janus, perform much worse on older but highly controlled datasets such as CMU MPIE. To aid our study, we present a simple and practical method to handle pose variation in face recognition pipelines designed to deal with extremely off-angle faces. Our approach is to ignore the half of the face with any self-occlusion. This method allows our models to be highly robust to pose, and helps us achieve state-of-the-art results on several protocols using the CMU MPIE dataset as well as very accurate results on the CFP dataset, outperforming recent efforts using the same training data. Dipan K. Pal, Chandrasekhar Bhagavatula, Yutong Zheng, Ran Tao 0013, Marios Savvides |
WACV | 5 |
| 2019 | Learning from Longitudinal Face Demonstration - Where Tractable Deep Modeling Meets Inverse Reinforcement Learning
Chi Nhan Duong, Kha Gia Quach, Khoa Luu, T. Hoang Ngan Le, Marios Savvides, Tien D. Bui |
Int. J. Comput. Vis. | 5 |
| 2019 | SSR2: Sparse signal recovery for single-image super-resolution on faces with extreme low resolutions
Ramzi Abiantun, Felix Juefei-Xu, Utsav Prabhu, Marios Savvides |
Pattern Recognit. | 4 |
| 2018 | RankGAN: A Maximum Margin Ranking GAN for Generating Faces
Felix Juefei-Xu, Rahul Dey, Vishnu Naresh Boddeti, Marios Savvides |
ACCV (3) | 4 |
| 2018 | Perturbative Neural NetworksabstractConvolutional neural networks are witnessing wide adoption in computer vision systems with numerous applications across a range of visual recognition tasks. Much of this progress is fueled through advances in convolutional neural network architectures and learning algorithms even as the basic premise of a convolutional layer has remained unchanged. In this paper, we seek to revisit the convolutional layer that has been the workhorse of state-of-the-art visual recognition models. We introduce a very simple, yet effective, module called a perturbation layer as an alternative to a convolutional layer. The perturbation layer does away with convolution in the traditional sense and instead computes its response as a weighted linear combination of non-linearly activated additive noise perturbed inputs. We demonstrate both analytically and empirically that this perturbation layer can be an effective replacement for a standard convolutional layer. Empirically, deep neural networks with perturbation layers, called Perturbative Neural Networks (PNNs), in lieu of convolutional layers perform comparably with standard CNNs on a range of visual datasets (MNIST, CIFAR-10, PASCAL VOC, and ImageNet) with fewer parameters. Felix Juefei-Xu, Vishnu Naresh Boddeti, Marios Savvides |
CVPR | 3 |
| 2018 | Ring Loss: Convex Feature Normalization for Face RecognitionabstractWe motivate and present Ring loss, a simple and elegant feature normalization approach for deep networks designed to augment standard loss functions such as Softmax. We argue that deep feature normalization is an important aspect of supervised classification problems where we require the model to represent each class in a multi-class problem equally well. The direct approach to feature normalization through the hard normalization operation results in a non-convex formulation. Instead, Ring loss applies soft normalization, where it gradually learns to constrain the norm to the scaled unit circle while preserving convexity leading to more robust features. We apply Ring loss to large-scale face recognition problems and present results on LFW, the challenging protocols of IJB-A Janus, Janus CS3 (a superset of IJB-A Janus), Celebrity Frontal-Profile (CFP) and MegaFace with 1 million distractors. Ring loss outperforms strong baselines, matches state-of-the-art performance on IJB-A Janus and outperforms all other results on the challenging Janus CS3 thereby achieving state-of-the-art. We also outperform strong baselines in handling extremely low resolution face matching. Yutong Zheng, Dipan K. Pal, Marios Savvides |
CVPR | 3 |
| 2018 | Seeing Small Faces From Robust Anchor's PerspectiveabstractThis paper introduces a novel anchor design principle to support anchor-based face detection for superior scaleinvariant performance, especially on tiny faces. To achieve this, we explicitly address the problem that anchor-based detectors drop performance drastically on faces with tiny sizes, e.g. less than 16 × 16 pixels. In this paper, we investigate why this is the case. We discover that current anchor design cannot guarantee high overlaps between tiny faces and anchor boxes, which increases the difficulty of training. The new Expected Max Overlapping (EMO) score is proposed which can theoretically explain the low overlapping issue and inspire several effective strategies of new anchor design leading to higher face overlaps, including anchor stride reduction with new network architectures, extra shifted anchors, and stochastic face shifting. Comprehensive experiments show that our proposed method significantly outperforms the baseline anchor-based detector, while consistently achieving state-of-the-art results on challenging face detection datasets with competitive runtime speed. Chenchen Zhu, Ran Tao 0013, Khoa Luu, Marios Savvides |
CVPR | 4 |
| 2018 | Enhancing Interior and Exterior Deep Facial Features for Face Detection in the WildabstractAlthough face detection has been intensely studied for decades, it is still a challenging topic due to numerous conditions, e.g. heavy occlusions, low resolutions, extreme poses, non-face patterns that look like human faces, etc. This paper proposes a novel region-based ConvNet to address these issues. Our approach enhances the interior deep facial features and explicitly incorporates the exterior deep features. The enhanced interior features provide fine details for small faces. The exterior features capture the local information surrounding the face, supporting the detection under challenging conditions. Experiments show that our proposed components improve the baseline method significantly. Additionally, our approach consistently achieves competitive performance in four challenging databases, i.e. Wider Face, AFW, PASCAL Faces, and FDDB. We also introduce a new challenging non-face dataset 1 of 6,000 images to benchmark false positive rates for future research. Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides |
FG | 4 |
| 2018 | Automatic Seizure Detection via an Optimized Image-Based Deep Feature LearningabstractIn this paper, our goal is to find an optimized approach that can learn features from multichannel EEG time-series data to perform automatic seizure detection. In general, it is not easy to learn robust features from EEG signals due to the variations in both intra and inter-patient variability. However, to achieve good generalization, we use an algorithm that tries to capture spectral, temporal and spatial information, in contrast to standard EEG analysis techniques that ignore spatial aspects. The first stage of this algorithm is to transform EEG signals into a sequence of topology-preserving multi-spectral and temporal images. After that, these generated images are fed as inputs to a convolutional neural network. By overcoming the lack of data, especially the positive samples, and creating a process to deal with unbalanced datasets and optimizing the complexity of the network, our convolutional neural network learns a general spatially invariant representation of a seizure in a reasonable time and improves sensitivity, specificity and accuracy result comparable to the state-of-the-art results. Ibrahim Alkanhal, B. V. K. Vijaya Kumar, Marios Savvides |
ICMLA | 3 |
| 2018 | Deep Recurrent Level Set for Segmenting Brain Tumors
T. Hoang Ngan Le, Raajitha Gummadi, Marios Savvides |
MICCAI (3) | 3 |
| 2018 | Deep contextual recurrent residual networks for scene labeling
T. Hoang Ngan Le, Chi Nhan Duong, Ligong Han, Khoa Luu, Kha Gia Quach, Marios Savvides |
Pattern Recognit. | 6 |
| 2018 | Reformulating Level Sets as Deep Recurrent Neural Network Approach to Semantic SegmentationabstractVariational Level Set (LS) has been a widely used method in medical segmentation. However, it is limited when dealing with multi-instance objects in the real world. In addition, its segmentation results are quite sensitive to initial settings and highly depend on the number of iterations. To address these issues and boost the classic variational LS methods to a new level of the learnable deep learning approaches, we propose a novel definition of contour evolution named Recurrent Level Set (RLS) 1 to employ Gated Recurrent Unit under the energy minimization of a variational LS functional. The curve deformation process in RLS is formed as a hidden state evolution procedure and updated by minimizing an energy functional composed of fitting forces and contour length. By sharing the convolutional features in a fully end-to-end trainable framework, we extend RLS to Contextual RLS (CRLS) to address semantic segmentation in the wild. The experimental results have shown that our proposed RLS improves both computational time and segmentation accuracy against the classic variational LS-based method whereas the fully end-to-end system CRLS achieves competitive performance compared to the state-of-the-art semantic segmentation approaches. T. Hoang Ngan Le, Kha Gia Quach, Khoa Luu, Chi Nhan Duong, Marios Savvides |
IEEE Trans. Image Process. | 5 |
| 2017 | Local Binary Convolutional Neural Networks
Felix Juefei-Xu, Vishnu Naresh Boddeti, Marios Savvides |
CVPR | 3 |
| 2017 | Faster than Real-Time Facial Alignment: A 3D Spatial Transformer Network Approach in Unconstrained Poses
Chandrasekhar Bhagavatula, Chenchen Zhu, Khoa Luu, Marios Savvides |
ICCV | 4 |
| 2017 | Temporal Non-volume Preserving Approach to Facial Age-Progression and Age-Invariant Face RecognitionabstractModeling the long-term facial aging process is extremely challenging due to the presence of large and non-linear variations during the face development stages. In order to efficiently address the problem, this work first decomposes the aging process into multiple short-term stages. Then, a novel generative probabilistic model, named Temporal Non-Volume Preserving (TNVP) transformation, is presented to model the facial aging process at each stage. Unlike Generative Adversarial Networks (GANs), which requires an empirical balance threshold, and Restricted Boltzmann Machines (RBM), an intractable model, our proposed TNVP approach guarantees a tractable density function, exact inference and evaluation for embedding the feature transformations between faces in consecutive stages. Our model shows its advantages not only in capturing the non-linear age related variance in each stage but also producing a smooth synthesis in age progression across faces. Our approach can model any face in the wild provided with only four basic landmark points. Moreover, the structure can be transformed into a deep convolutional network while keeping the advantages of probabilistic models with tractable log-likelihood density estimation. Our method is evaluated in both terms of synthesizing age-progressed faces and cross-age face verification and consistently shows the state-of-the-art results in various face aging databases, i.e. FG-NET, MORPH, AginG Faces in the Wild (AGFW), and Cross-Age Celebrity Dataset (CACD). A large-scale face verification on Megaface challenge 1 is also performed to further show the advantages of our proposed approach. Chi Nhan Duong, Kha Gia Quach, Khoa Luu, T. Hoang Ngan Le, Marios Savvides |
ICCV | 5 |
| 2017 | Max-Margin Invariant Features from Transformed Unlabelled DataabstractThe study of representations invariant to common transformations of the data is important to learning. Most techniques have focused on local approximate invariance implemented within expensive optimization frameworks lacking explicit theoretical guarantees. In this paper, we study kernels that are invariant to a unitary group while having theoretical guarantees in addressing the important practical issue of unavailability of transformed versions of labelled data. A problem we call the Unlabeled Transformation Problem which is a special form of semi-supervised learning and one-shot learning. We present a theoretically motivated alternate approach to the invariant kernel SVM based on which we propose Max-Margin Invariant Features (MMIF) to solve this problem. As an illustration, we design an framework for face recognition and demonstrate the efficacy of our approach on a large scale semi-synthetic dataset with 153,000 images and a new challenging protocol on Labelled Faces in the Wild (LFW) while out-performing strong baselines. Dipan K. Pal, Ashwin A. Kannan, Gautam Arakalgud, Marios Savvides |
NIPS | 4 |
| 2017 | Semi self-training beard/moustache detection and segmentation simultaneously
T. Hoang Ngan Le, Khoa Luu, Chenchen Zhu, Marios Savvides |
Image Vis. Comput. | 4 |
| 2017 | Compressed Submanifold Multifactor AnalysisabstractAlthough widely used, Multilinear PCA (MPCA), one of the leading multilinear analysis methods, still suffers from four major drawbacks. First, it is very sensitive to outliers and noise. Second, it is unable to cope with missing values. Third, it is computationally expensive since MPCA deals with large multi-dimensional datasets. Finally, it is unable to maintain the local geometrical structures due to the averaging process. This paper proposes a novel approach named Compressed Submanifold Multifactor Analysis (CSMA) to solve the four problems mentioned above. Our approach can deal with the problem of missing values and outliers via SVD-L1. The Random Projection method is used to obtain the fast low-rank approximation of a given multifactor dataset. In addition, it is able to preserve the geometry of the original data. Our CSMA method can be used efficiently for multiple purposes, e.g. noise and outlier removal, estimation of missing values, biometric applications. We show that CSMA method can achieve good results and is very efficient in the inpainting problem as compared to [1], [2]. Our method also achieves higher face recognition rates compared to LRTC, SPMA, MPCA and some other methods, i.e. PCA, LDA and LPP, on three challenging face databases, i.e. CMU-MPIE, CMU-PIE and Extended YALE-B. Khoa Luu, Marios Savvides, Tien D. Bui, Ching Y. Suen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | DeepSafeDrive: A grammar-aware driver parsing approach to Driver Behavioral Situational Awareness (DB-SAW)
T. Hoang Ngan Le, Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides |
Pattern Recognit. | 5 |
| 2016 | Learning to Invert Local Binary Patterns
Felix Juefei-Xu, Marios Savvides |
BMVC | 2 |
| 2016 | Discriminative Invariant Kernel Features: A Bells-and-Whistles-Free Approach to Unsupervised Face Recognition and Pose EstimationabstractWe propose an explicitly discriminative and 'simple' approach to generate invariance to nuisance transformations modeled as unitary. In practice, the approach works well to handle non-unitary transformations as well. Our theoretical results extend the reach of a recent theory of invariance to discriminative and kernelized features based on unitary kernels. As a special case, a single common framework can be used to generate subject-specific pose-invariant features for face recognition and vice-versa for pose estimation. We show that our main proposed method (DIKF) can perform well under very challenging large-scale semisynthetic face matching and pose estimation protocols with unaligned faces using no landmarking whatsoever. We additionally benchmark on CMU MPIE and outperform previous work in almost all cases on off-angle face matching while we are on par with the previous state-of-the-art on the LFW unsupervised and image-restricted protocols, without any low-level image descriptors other than raw-pixels. Dipan K. Pal, Felix Juefei-Xu, Marios Savvides |
CVPR | 3 |
| 2016 | Simultaneous forgery identification and localization in paintings using advanced correlation filtersabstractWith the availability of high resolution digital technology, there has been increased interest in developing statistical and image processing techniques that can enhance the existing capabilities of analyzing works of art for authenticity. This work explores the merits of using advanced correlation filters in supplementing art experts efforts in identifying forgeries among disputed paintings. We show that by training the optimal trade-off synthetic discriminant function (OTSDF) filter on each section of a coarsely parceled image of an original painting, we are not only able to distinguish between a low-quality digitized representation of a painting and its forgery, but also specifically indicate where the differences occur and where the replica is particularly faithful to the original. This method is also valuable in determining whether an original painting has undergone any modifications, given that a representation of the initial version is available. Paul Buchana, Irina Cazan, Manuel Diaz-Granados, Felix Juefei-Xu, Marios Savvides |
ICIP | 5 |
| 2016 | Pose estimation using Spectral and Singular Value recompositionabstractIn face recognition tasks, the changing pose of the face can cause enough information to be lost to cause the recognition to fail so being able to determine the pose of the face beforehand can allow for some better recognition performance. Many methods used for pose estimation tasks rely on finding some underlying structure of the data given to create a classifier. We propose an alternative method in which the training data itself is the underlying structure of a classifier. This is accomplished through the use of matrix decomposition equations. However, instead of decomposing a matrix, one is created by carefully selecting the terms in the decomposition equation such that the resulting matrix has the desired properties for classification. We show two recomposition methods using the Spectral Decomposition and Singular Value Decomposition equations. We show this method can perform pose estimation with a high accuracy of 85.21% and an accuracy of 98.42% when allowing a ±15° tolerance on the pose estimate on the CUbiC FacePix dataset. We also show results on both yaw and pitch estimation on the Pointing'04 dataset with our methods achieving 77.01% accuracy on yaw estimation. Chandrasekhar Bhagavatula, Raied Aljadaany, Marios Savvides |
ICPR | 3 |
| 2016 | Robust hand detection in VehiclesabstractThe problems of hand detection have been widely addressed in many areas, e.g. human computer interaction environment, driver behaviors monitoring, etc. However, the detection accuracy in recent hand detection systems are still far away from the demands in practice due to a number of challenges, e.g. hand variations, highly occlusions, low-resolution and strong lighting conditions. This paper presents the Multiple Scale Faster Region-based Convolutional Neural Network (MS-FRCNN) to handle the problems of hand detection in given digital images collected under challenging conditions. Our proposed method introduces a multiple scale deep feature extraction approach in order to handle the challenging factors to provide a robust hand detection algorithm. The method is evaluated on the challenging hand database, i.e. the Vision for Intelligent Vehicles and Applications (VIVA) Challenge, and compared against various recent hand detection methods. Our proposed method achieves the state-of-the-art results with 20% of the detection accuracy higher than the second best one in the VIVA challenge. T. Hoang Ngan Le, Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides |
ICPR | 5 |
| 2016 | Towards a Unified Framework for Pose, Expression, and Occlusion Tolerant Automatic Facial AlignmentabstractWe propose a facial alignment algorithm that is able to jointly deal with the presence of facial pose variation, partial occlusion of the face, and varying illumination and expressions. Our approach proceeds from sparse to dense landmarking steps using a set of specific models trained to best account for the shape and texture variation manifested by facial landmarks and facial shapes across pose and various expressions. We also propose the use of a novel l1-regularized least squares approach that we incorporate into our shape model, which is an improvement over the shape model used by several prior Active Shape Model (ASM) based facial landmark localization algorithms. Our approach is compared against several state-of-the-art methods on many challenging test datasets and exhibits a higher fitting accuracy on all of them. Keshav Seshadri, Marios Savvides |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Multi-class Fukunaga Koontz discriminant analysis for enhanced face recognition
Felix Juefei-Xu, Marios Savvides |
Pattern Recognit. | 2 |
| 2016 | A novel Shape Constrained Feature-based Active Contour model for lips/mouth segmentation in the wild
T. Hoang Ngan Le, Marios Savvides |
Pattern Recognit. | 2 |
| 2015 | IRIS super-resolution via nonparametric over-complete dictionary learningabstractThis paper presents a novel iris super-resolution approach using a powerful nonparametric Bayesian modeling technique in the framework of sparse representation and over-complete dictionary. Far apart from previous iris super-resolution methods, our proposed approach has ability to automatically discover optimal parameter sets and optimally adapt from a given training data. Particularly, the Beta Process will be employed to build a nonparametric discriminative over-complete dictionary to represent and discriminate input samples simultaneously. Our proposed method will be evaluated on Casia iris database and compared with the linear interpolation super resolution. The result shows that our approach improves the performance of iris recognition. Raied Aljadaany, Khoa Luu, Shreyas Venugopalan, Marios Savvides |
ICIP | 4 |
| 2015 | Pareto-optimal discriminant analysisabstractIn this work, we have proposed the Pareto-optimal discriminant analysis (PDA), an optimally designed linear subspace learning method that harnesses advantages across many well-known methods such as PCA, LDA, UDP and LPP. By optimizing over the joint objective function and carrying out an alternative coefficients updating scheme, we are able to obtain a linear subspace which is optimized to truly maximize the objective function in discriminant analysis. The proposed method also provides flexibility for formulating the linear transformation matrix in an overcomplete fashion, allowing for a sparse representation. We have shown, in the context of large scale unconstrained face recognition and illumination invariant face recognition, that our proposed PDA significantly outperforms other linear subspace methods. Felix Juefei-Xu, Marios Savvides |
ICIP | 2 |
| 2015 | Single face image super-resolution via solo dictionary learningabstractIn this work, we have proposed a single face image super-resolution approach based on solo dictionary learning. The core idea of the proposed method is to recast the super-resolution task as a missing pixel problem, where the low-resolution image is considered as its high-resolution counterpart with many pixels missing in a structured manner. A single dictionary is therefore sufficient for recovering the super-resolved image by filling the missing pixels. In order to fill in 93.75% of the missing pixels when super-resolving a 16 × 16 low-resolution image to a 64 × 64 one, we adopt a whole image-based solo dictionary learning scheme. The proposed procedure can be easily extended to low-resolution input images with arbitrary dimensions, as well as high-resolution recovery images of arbitrary dimensions. Also, for a fixed desired super-resolution dimension, there is no need to retrain the dictionary when the input low-resolution image has arbitrary zooming factors. Based on a large-scale fidelity experiment on the FRGC ver2 database, our proposed method has outperformed other well established interpolation methods as well as the coupled dictionary learning approach. Felix Juefei-Xu, Marios Savvides |
ICIP | 2 |
| 2015 | Encoding and decoding local binary patterns for harsh face illumination normalizationabstractIn this work, we propose a new illumination normalization technique based on a simple, yet widely used descriptor: local binary patterns (LBP). We capitalize on the fact that LBP retains tolerance to illumination changes and use the LBP mapping to remove illumination variations cast on face images. Through learning a reverse mapping from the LBP domain to the pixel domain, we are able to recover the illumination normalized face with high fidelity. The reverse mapping step is made possible via a joint dictionary learning framework between the LBP domain and the pixel domain. The illumination normalized faces using our proposed LBP encoding and decoding method not only exhibit very high fidelity against neutrally illuminated face, but also allow for a significant improvement in face verification experiments using even the simplest nearest-neighbor classifier. These conclusions are drawn after benchmarking our algorithm against 22 prevailing illumination normalization techniques on Extended YaleB database which has been widely adopted for challenging face illumination problems. Felix Juefei-Xu, Marios Savvides |
ICIP | 2 |
| 2015 | A robust contour sampling and tensor-based approach to facial beard and mustache shape segmentation and matchingabstractIn this paper, we propose a novel system for beard and mustache segmentation and matching in facial images. We first segment out facial hair contours from the image by utilizing a sparse dictionary on self-quotient images to classify regions as either skin or facial hair. We then landmark the shape contour to obtain points around the contour of the image using a combination of two algorithms, a novel non-uniform sampling algorithm, and points obtained from SIFT. We utilize these landmark points to extract inner distance-based shape context features. Finally, these features are used as inputs for a tensor product graph-based matching system. We run experiments on the Multiple Biometric Grand Challenge (MBGC) and the PINELLAS mugshot databases. Our pipeline achieves 90.3% matching accuracy on a subset of the PINELLAS database when divided into four types of facial hair. Karanhaar Singh, Khoa Luu, T. Hoang Ngan Le, Marios Savvides |
ICIP | 4 |
| 2015 | Investigating the feasibility of image-based nose biometricsabstractThe search for new biometrics is never ending. In this work, we investigate the use of image based nasal features as a biometric. In many real-world recognition scenarios, partial occlusions on the face leave the nose region visible (e.g. sunglasses). Face recognition systems often fail or perform poorly in such settings. Furthermore, the nose region naturally contain more invariance to expression than features extracted from other parts of the face. In this study, we extract discriminative nasal features using Kernel Class-Dependence Feature Analysis (KCFA) based on Optimal Trade-off Synthetic Discriminant Function (OTSDF) filters. We evaluate this technique on the FRGC ver2.0 database and the AR Face database, training and testing exclusively on nasal features and have compared the results to the full face recognition using KCFA features. We find that the between-subject discriminability in nasal features is comparable to that found in facial features. This shows that nose biometrics have a potential to support and boost biometric identification, that has largely been under utilized. Moreover, our extracted KCFA nose features have significantly outperformed the PittPatt face matcher which works with the original JPEG images on the AR facial occlusion database. This shows that nose biometrics can be used as a stand-alone biometric trait when the subjects are under occlusions. Niv Zehngut, Felix Juefei-Xu, Rishabh Bardia, Dipan K. Pal, Chandrasekhar Bhagavatula, Marios Savvides |
ICIP | 6 |
| 2015 | Facial aging and asymmetry decomposition based approaches to identification of twins
T. Hoang Ngan Le, Keshav Seshadri, Khoa Luu, Marios Savvides |
Pattern Recognit. | 4 |
| 2015 | Spartans: Single-Sample Periocular-Based Alignment-Robust Recognition Technique Applied to Non-Frontal ScenariosabstractIn this paper, we investigate a single-sample periocular-based alignment-robust face recognition technique that is pose-tolerant under unconstrained face matching scenarios. Our Spartans framework starts by utilizing one single sample per subject class, and generate new face images under a wide range of 3D rotations using the 3D generic elastic model which is both accurate and computationally economic. Then, we focus on the periocular region where the most stable and discriminant features on human faces are retained, and marginalize out the regions beyond the periocular region since they are more susceptible to expression variations and occlusions. A novel facial descriptor, high-dimensional Walsh local binary patterns, is uniformly sampled on facial images with robustness toward alignment. During the learning stage, subject-dependent advanced correlation filters are learned for pose-tolerant non-linear subspace modeling in kernel feature space followed by a coupled max-pooling mechanism which further improve the performance. Given any unconstrained unseen face image, the Spartans can produce a highly discriminative matching score, thus achieving high verification rate. We have evaluated our method on the challenging Labeled Faces in the Wild database and solidly outperformed the state-of-the-art algorithms under four evaluation protocols with a high accuracy of 89.69%, a top score among image-restricted and unsupervised protocols. The advancement of Spartans is also proven in the Face Recognition Grand Challenge and Multi-PIE databases. In addition, our learning method based on advanced correlation filters is much more effective, in terms of learning subject-dependent pose-tolerant subspaces, compared with many well-established subspace methods in both linear and non-linear cases. Felix Juefei-Xu, Khoa Luu, Marios Savvides |
IEEE Trans. Image Process. | 3 |
| 2014 | Distributed class dependent feature analysis - A big data approachabstractBig data has been becoming ubiquitous and applied in numerous fields recently. The challenges to solve a large-scale machine learning problem in big data scenario generally lie in three aspects. Firstly, a proposed machine learning algorithm has to be appropriated for the distributed optimization problem. Secondly, it needs a platform for the distributed implementation. Finally, the communication delays different machines may cause problems in convergence even though the non-distributed algorithm shows a good convergence rate. In order to solve these challenges, we propose a new machine learning approach named Distributed Class-dependent Feature Analysis (DCFA), to combine the advantages of sparse representation in an over-complete dictionary. The classifier is based on the estimation of class-specific optimal filters, by solving an l1-norm optimization problem. We demonstrate how this problem is solved using the Alternating Direction Method of Multipliers and also explore relevant convergency details. More importantly, our proposed framework can be efficiently implemented on a robust distributed framework. Thus, it improves both accuracy and computational time in large-scale databases. Our method achieves very high classification accuracies in face recognition in the presence of occlusions on AR database. It also outperforms the state of the art methods in object recognition on two challenging large-scale object databases, i.e. Caltech101 and Caltech256. It hence shows its applicability to general computer vision and pattern recognition problems. In addition, computational time experiments show our distributed method achieves high speedup of 7.85x on Caltech256 databases with just 10 machine nodes compared to the non-distributed version and can gain even more with more computing resources. Khoa Luu, Chenchen Zhu, Marios Savvides |
IEEE BigData | 3 |
| 2014 | A novel eyebrow segmentation and eyebrow shape-based identificationabstractRecent studies in biometrics have shown that the periocular region of the face is sufficiently discriminative for robust recognition, and particularly effective in certain scenarios such as extreme occlusions, and illumination variations where traditional face recognition systems are unreliable. In this paper, we first propose a fully automatic, robust and fast graph-cut based eyebrow segmentation technique to extract the eyebrow shape from a given face image. We then propose an eyebrow shape-based identification system for periocular face recognition. Our experiments have been conducted over large datasets from the MBGC and AR databases and the resilience of the proposed approach has been evaluated under varying data conditions. The experimental results show that the proposed eyebrow segmentation achieves high accuracy with an F-Measure of 99.4% and the identification system achieves rates of 76.0% on the AR database and 85.0% on the MBGC database. T. Hoang Ngan Le, Utsav Prabhu, Marios Savvides |
IJCB | 3 |
| 2014 | Sparse Feature Extractionfor Pose-Tolerant Face RecognitionabstractAutomatic face recognition performance has been steadily improving over years of research, however it remains significantly affected by a number of factors such as illumination, pose, expression, resolution and other factors that can impact matching scores. The focus of this paper is the pose problem which remains largely overlooked in most real-world applications. Specifically, we focus on one-to-one matching scenarios where a query face image of a random pose is matched against a set of gallery images. We propose a method that relies on two fundamental components: (a) A 3D modeling step to geometrically correct the viewpoint of the face. For this purpose, we extend a recent technique for efficient synthesis of 3D face models called 3D Generic Elastic Model. (b) A sparse feature extraction step using subspace modeling and ℓ1-minimization to induce pose-tolerance in coefficient space. This in return enables the synthesis of an equivalent frontal-looking face, which can be used towards recognition. We show significant performance improvements in verification rates compared to commercial matchers, and also demonstrate the resilience of the proposed method with respect to degrading input quality. We find that the proposed technique is able to match non-frontal images to other non-frontal images of varying angles. Ramzi Abiantun, Utsav Prabhu, Marios Savvides |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Subspace-Based Discrete Transform Encoded Local Binary Patterns Representations for Robust Periocular Matching on NIST's Face Recognition Grand ChallengeabstractIn this paper, we employ several subspace representations (principal component analysis, unsupervised discriminant projection, kernel class-dependence feature analysis, and kernel discriminant analysis) on our proposd discrete transform encoded local binary patterns (DT-LBP) to match periocular region on a large data set such as NIST's face recognition grand challenge (FRGC) ver2 database. We strictly follow FRGC Experiment 4 protocol, which involves 1-to-1 matching of 8014 uncontrolled probe periocular images to 16 028 controlled target periocular images (~128 million pairwise face match comparisons). The performance of the periocular region is compared with that of full face with different illumination preprocessing schemes. The verification results on periocular region show that subspace representation on DT-LBP outperforms LBP significantly and gains a giant leap from traditional subspace representation on raw pixel intensity. Additionally, our proposed approach using only the periocular region is almost as good as full face with only 2.5% reduction in verification rate at 0.1% false accept rate, yet we gain tolerance to expression, occlusion, and capability of matching partial faces in crowds. In addition, we have compared the best standalone DT-LBP descriptor with eight other state-of-the-art descriptors for facial recognition and achieved the best performance. The two general frameworks are our major contribution: 1) a general framework that employs various generative and discriminative subspace modeling techniques for DT-LBP representation and 2) a general framework that encodes discrete transforms with local binary patterns for the creation of robust descriptors. Felix Juefei-Xu, Marios Savvides |
IEEE Trans. Image Process. | 2 |
| 2013 | An Automatic Iris Occlusion Estimation Method Based on High-Dimensional Density EstimationabstractIris masks play an important role in iris recognition. They indicate which part of the iris texture map is useful and which part is occluded or contaminated by noisy image artifacts such as eyelashes, eyelids, eyeglasses frames, and specular reflections. The accuracy of the iris mask is extremely important. The performance of the iris recognition system will decrease dramatically when the iris mask is inaccurate, even when the best recognition algorithm is used. Traditionally, people used the rule-based algorithms to estimate iris masks from iris images. However, the accuracy of the iris masks generated this way is questionable. In this work, we propose to use Figueiredo and Jain's Gaussian Mixture Models (FJ-GMMs) to model the underlying probabilistic distributions of both valid and invalid regions on iris images. We also explored possible features and found that Gabor Filter Bank (GFB) provides the most discriminative information for our goal. Finally, we applied Simulated Annealing (SA) technique to optimize the parameters of GFB in order to achieve the best recognition rate. Experimental results show that the masks generated by the proposed algorithm increase the iris recognition rate on both ICE2 and UBIRIS dataset, verifying the effectiveness and importance of our proposed method for iris occlusion estimation. Yung-hui Li, Marios Savvides |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | SparCLeS: Dynamic 퓁 1 Sparse Classifiers With Level Sets for Robust Beard/Moustache Detection and SegmentationabstractRobust facial hair detection and segmentation is a highly valued soft biometric attribute for carrying out forensic facial analysis. In this paper, we propose a novel and fully automatic system, called SparCLeS, for beard/moustache detection and segmentation in challenging facial images. SparCLeS uses the multiscale self-quotient (MSQ) algorithm to preprocess facial images and deal with illumination variation. Histogram of oriented gradients (HOG) features are extracted from the preprocessed images and a dynamic sparse classifier is built using these features to classify a facial region as either containing skin or facial hair. A level set based approach, which makes use of the advantages of both global and local information, is then used to segment the regions of a face containing facial hair. Experimental results demonstrate the effectiveness of our proposed system in detecting and segmenting facial hair regions in images drawn from three databases, i.e., the NIST Multiple Biometric Grand Challenge (MBGC) still face database, the NIST Color Facial Recognition Technology FERET database, and the Labeled Faces in the Wild (LFW) database. T. Hoang Ngan Le, Khoa Luu, Marios Savvides |
IEEE Trans. Image Process. | 3 |
| 2012 | Automatic segmentation of cardiosynchronous waveforms using cepstral analysis and continuous wavelet transformsabstractThe cardiosynchronous signal obtained through Radio Frequency Impedance Interrogation (RFII) is a non-invasive method for monitoring hemodynamics with potential applications in combat triage and biometric identification. The RFII signal is periodic in nature dominated by the heart beat cycle. The first step in both of these applications is to segment the signal by identifying a fiducial point in each heart beat cycle. A continuous wavelet transform was utilized to locate the fiducial points with high temporal resolution. Cepstral Analysis was used to estimate the average heart rate to focus on the appropriate portion of the time-frequency spectrum. Robust heartbeats from RFII signals collected from four subjects were segmented using this method. Chandrasekhar Bhagavatula, Aaron Jaech, Marios Savvides, B. V. K. Vijaya Kumar, Robert Friedman, Rebecca Blue, Marc O. Griofa |
ICIP | 3 |
| 2012 | Second-degree correlation surface features from Optimal Trade-off Synthetic Discriminant Function filters for subject identification using radio frequency cardiosynchronous waveformsabstractRadio Frequency Impedance Interrogation (RFII) measures hemodynamic function via resonance frequency coupling to a hydrophilic protein molecule. The RFII device generates a cardiosynchronous waveform from the identification of blood movement in the time, frequency, and voltage domains. This paper examines RFII signals with the end goal of allowing confirmation of the identity of a subject in an operational setting. An Optimal Trade-off Synthetic Discriminant Function (OT-SDF) was applied to filter the data stream for subject identification. Preliminary results using the OT-SDF Filters demonstrate 63.3% successful single-heartbeat subject identification. However, each individual's correlation surfaces appear to have a unique waveform morphology that is visually distinct from the other individuals in the data set. Improved identification was seen with second-degree correlation suggesting that a second-degree correlation may hold great potential as a biometric feature extraction identifier. We show that using correlation plane outputs as features actually provide a robust biometric identifier and significant higher identification accuracy. Madhusudan Bhagavatula, Marios Savvides, B. V. K. Vijaya Kumar, Robert Friedman, Rebecca Blue, Marc O. Griofa |
ICIP | 2 |
| 2012 | A novel energy based filter for cross-blink eye detectionabstractBased on the fact that eye regions may be considered as a sequence of consecutive low and high spatial frequency regions, we propose a novel and efficient filter based on Isotropic Gaussian Energy which can be used for eye detection. When applied to facial images, the designed filter highlights the eye region with a prominent pattern, which can then be localized by template matching technique such as MACE correlation filter. The proposed filter is proved useful to detect both open and closed eyes. We demonstrate the effectiveness of the proposed filter by conducting experiments on both FERET and MBGC face databases. T. Hoang Ngan Le, Khoa Luu, Utsav Prabhu, Marios Savvides |
ICIP | 4 |
| 2012 | Beard and mustache segmentation using sparse classifiers on self-quotient imagesabstractIn this paper, we propose a novel system for beard and mustache detection and segmentation in challenging facial images. Our system first eliminates illumination artifacts using the self-quotient algorithm. A sparse classifier is then used on these self-quotient images to classify a region as either containing skin or facial hair. We conduct experiments on the MBGC and color FERET databases to demonstrate the effectiveness of our proposed system. T. Hoang Ngan Le, Khoa Luu, Keshav Seshadri, Marios Savvides |
ICIP | 4 |
| 2012 | Facecut - a robust approach for facial feature segmentationabstractSegmentation of facial features is a key pre-processing step in enabling facial recognition, building of 3D facial models, expression analysis, and pose estimation. Recently, graph cuts based algorithms have been adapted to carry out this task but many of these methods require manual initialization of points in the foreground and background. In this paper, we propose a novel and fully automatic approach, named Face-Cut, to perform accurate facial feature segmentation. FaceCut combines the positive features of the Modified Active Shape Model (MASM) and GrowCut algorithms to ensure highly accurate and completely automatic segmentation of facial features. We demonstrate the effectiveness of FaceCut on images from two challenging databases. Khoa Luu, T. Hoang Ngan Le, Keshav Seshadri, Marios Savvides |
ICIP | 4 |
| 2012 | Compressed Submanifold Multifactor Analysis with adaptive factor structures
Khoa Luu, Marios Savvides, Tien D. Bui, Ching Y. Suen |
ICPR | 2 |
| 2012 | Unconstrained periocular biometric acquisition and recognition using COTS PTZ camera for uncooperative and non-cooperative subjectsabstractWe propose an acquisition and recognition system based only on periocular biometric using the COTS PTZ camera to tackle the difficulty that the full face recognition approach has encountered in highly unconstrained real-world scenario, especially for capturing and recognizing uncooperative and non-cooperative subjects with expression, closed eyes, and facial occlusions. We evaluate our algorithm on the periocular region and compare that to the performance of the full face on the Compass database we have collected. The results have shown that the periocular region, when tackling unconstrained matching, is a much better choice than the full face for face recognition even with less than 2/5 the size of the full face. To be more specific, the periocular matching across all facial manners, i.e., neutral expression, smiling expression, closed eyes, and facial occlusion, is able to achieve 60.7% verification rate at 0.1% false accept rate, a 16.9% performance boost over the full face. Felix Juefei-Xu, Marios Savvides |
WACV | 2 |
| 2012 | Gender and Ethnicity Specific Generic Elastic Models from a Single 2D Image for Novel 2D Pose Face Synthesis and RecognitionabstractIn this paper, we propose a novel method for generating a realistic 3D human face from a single 2D face image for the purpose of synthesizing new 2D face images at arbitrary poses using gender and ethnicity specific models. We employ the Generic Elastic Model (GEM) approach, which elastically deforms a generic 3D depth-map based on the sparse observations of an input face image in order to estimate the depth of the face image. Particularly, we show that Gender and Ethnicity specific GEMs (GE-GEMs) can approximate the 3D shape of the input face image more accurately, achieving a better generalization of 3D face modeling and reconstruction compared to the original GEM approach. We qualitatively validate our method using publicly available databases by showing each reconstructed 3D shape generated from a single image and new synthesized poses of the same person at arbitrary angles. For quantitative comparisons, we compare our synthesized results against 3D scanned data and also perform face recognition using synthesized images generated from a single enrollment frontal image. We obtain promising results for handling pose and expression changes based on the proposed method. Jingu Heo, Marios Savvides |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | 3-D Generic Elastic Models for Fast and Texture Preserving 2-D Novel Pose SynthesisabstractThis paper provides an in-depth analysis on face shape alignment for pose insensitive face recognition. The dissimilarity between two face images can be modeled as the difference in intensity between these two images, obtained by warping these faces onto the same shape. In order to achieve this, we must first align both face images independently to obtain a sparse 2-D shape representation. We achieve this by using a Combination of ASMs and AAMs (CASAAMs). We then exchange these two shapes and obtain new intensity (texture) faces based on these exchanged shapes. This allows us to align the two faces with increased pixel-level correspondence while simultaneously achieving a certain degree of pose correction. In order to account for large pose variation, it becomes necessary to model the underlying 3-D face structure for the synthesis of novel 2-D poses. However, in many real-world scenarios, only a single image of the subject is provided and acquisition of the 3-D model is not always feasible. To tackle this common real-world scenario, we propose a novel approach for modeling faces, called 3D Generic Elastic Model (3D-GEM), which can be deformed from a single 2-D image. Our analysis shows that 3-D depth information of human faces does not dramatically change across people, indicating that precise depth information of a person is not needed to generate useful novel 2-D poses. This is a significant discovery that makes our method feasible. We thus demonstrate that our 3-D face model can be efficiently produced by using a generic depth model which can be elastically deformed based on input facial features. This face model can then be rotated in 3-D in order to synthesize any arbitrary 2-D facial pose. Experimental results show that 3-D faces modeled by our proposed work effectively handle large 3-D pose changes in face alignment and they can be used for achieving pose tolerant face recognition. We also provide comparative results of face synthesis obtained by an actual 3-D face scanner and our approach, showing our proposed modeling approach is both effective and efficient. Jingu Heo, Marios Savvides |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2012 | An Analysis of the Sensitivity of Active Shape Models to Initialization When Applied to Automatic Facial LandmarkingabstractActive Shape Models (ASMs) have recently gained popularity for performing automatic facial landmark fitting. Their demonstrated ability to generalize and fit unseen faces make them ideal candidates for this task unlike the traditional Active Appearance Model (AAM)-based approaches, which have difficulty in accurately landmarking unseen images. Given a test image, a face detector is used to determine the locations, orientations and sizes of faces in the image. Facial landmarking algorithms, such as ASMs, are initialized based on these parameters. In this paper, we conduct a series of experiments to exhaustively evaluate the tolerance of three popular ASMs to initialization perturbations (translation, rotation, and scaling in size) of the face detected, a topic that has not been analyzed in depth to date. Our results are consistent across different databases, provide an understanding of the role initialization plays in the landmark fitting process and serve as a performance gauge that could be considered when comparing facial landmarking algorithms. Keshav Seshadri, Marios Savvides |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2011 | Multifactor analysis based on factor-dependent geometryabstractThis paper proposes a novel method that preserves the geometrical structure created by variation of multiple factors in analysis of multiple factor models, i.e., multifactor analysis. We use factor-dependent submanifolds as constituent elements of the factor-dependent geometry in a multiple factor framework. In this paper, a submanifold is defined as some subset of a manifold in the data space, and factor-dependent submanifolds are defined as the submani-folds created for each factor by varying only this factor. In this paper, we show that MPCA is formulated using factor-dependent submanifolds, as is our proposed method. We show, however, that MPCA loses the original shapes of these submanifolds because MPCA's parameterization is based on averaging the shapes of factor-dependent subman-ifolds for each factor. On the other hand, our proposed multifactor analysis preserves the shapes of individual factor-dependent submanifolds in low-dimensional spaces. Because the parameters obtained by our method do not lose their structures, our method, unlike MPCA, sufficiently covers original factor-dependent submanifolds. As a result of sufficient coverage, our method is appropriate for accurate classification of each sample. Sung Won Park, Marios Savvides |
CVPR | 2 |
| 2011 | Rapid 3D face modeling using a frontal face and a profile face for accurate 2D pose synthesisabstractThis paper proposes an efficient way of modeling 3D faces by using only two - a frontal and a profile - images. Although it is desirable to utilize only one single image for 3D face modeling, more accurate depth information can be obtained if we use a profile face image additionally. Despite this seemly straightforward task, however, no standard solutions for 3D face modeling with two images have yet been reported. To tackle this problem, in our work, we first extract facial shape information from each image and then align these two shapes in order to obtain a sparse 3D face. Then, the observed sparse 3D face is combined into generic dense depth information. By doing so, we reflect both the observed 3D sparse depth information and smooth depth changes around facial areas in our reconstructed 3D shape. Finally, the intensity of the frontal image is texture-mapped onto the reconstructed 3D shape for realistic 3D modeling. Unlike other 3D modeling methods, our proposed work is extremely fast (within a few seconds) and does not require any complex hardware settings or calibration. We illustrate our 3D modeling results by using the MPIE-database and demonstrate the effectiveness of the proposed approach. Jingu Heo, Marios Savvides |
FG | 2 |
| 2011 | The multifactor extension of Grassmann manifolds for face recognitionabstractWe propose the use of a multifactor model that extends Grassmann manifold to multiple factor frameworks. Both manifold learning algorithms and multifactor analysis are state-of-the-art dimension reduction techniques that are suitable to model variations of face images. In this paper, we demonstrate that Grassmann manifold can be extended to Mul-tifactor Grassmann manifold when used in conjunction with Multilinear PCA (MPCA). Indeed, the multifactor manifold learning algorithm proposed in this paper can be interpreted as MPCA's kernel-based extension using a kernel function that is defined in terms of geodesic distance. As a result, we first propose the use of Multifactor Grassmann manifold, which can learn both a multifactor structure and an underlying manifold in a given set of face images. We then demonstrate that our proposed method, Multifactor Grassmann manifold, produces more reliable results in the context of face recognition than the traditional dimension reduction techniques. Sung Won Park, Marios Savvides |
FG | 2 |
| 2011 | Generic 3D face pose estimation using facial shapesabstractGeneric 3D face pose estimation from a single 2D facial image is an extremely crucial requirement for face-related research areas. To meet with the remaining challenges for face pose estimation, suggested Murphy-Chutorian et al. [13], we believe that the first step is to create a large corpus of a 3D facial shape database in which the statistical relationship between projected 2D shapes and corresponding pose parameters can be easily observed. Because fa- cial geometry provides the most essential information for facial pose, understanding the effect of pose parameters in 2D facial shapes is a key step toward solving the remaining challenges. In this paper, we present necessary tasks to reconstruct 3D facial shapes from multiple 2D images and then explain how to generate 2D projected shapes at any rotation interval. To deal with self occlusions, a novel hidden points removal (HPR) algorithm is also proposed. By flexibly changing the number of points in 2D shapes, we evaluate the performance of two different approaches for achieving generic 3D pose estimation in both coarse and fine levels and analyze the importance of facial shapes toward generic 3D pose estimation. Jingu Heo, Marios Savvides |
IJCB | 2 |
| 2011 | Fusion of region-based representations for gender identificationabstractMuch of the current work on gender identification relies on legacy datasets of heavily controlled images with minimal facial appearance variations. As studies explore the effects of adding elements of variation into the data, they have met challenges in achieving granular statistical significance due to the limited size of their datasets. In this study, we aim to create a classification framework that is robust to non-studio, uncontrolled, real-world images. We show that the fusion of separate linear classifiers trained on smart-selected local patches achieves 90% accuracy, which is a 5% improvement over a baseline linear classifier on a straightforward pixel representation. These results are re- ported on our own uncontrolled database of ~26, 700 images collected from the Web. Si Ying Diana Hu, Brendan Jou, Aaron Jaech, Marios Savvides |
IJCB | 4 |
| 2011 | Investigating age invariant face recognition based on periocular biometricsabstractIn this paper, we will present a novel framework of utilizing periocular region for age invariant face recognition. To obtain age invariant features, we first perform preprocessing schemes, such as pose correction, illumination and periocular region normalization. And then we apply robust Walsh-Hadamard transform encoded local binary patterns (WLBP) on preprocessed periocular region only. We find the WLBP feature on periocular region maintains consistency of the same individual across ages. Finally, we use unsupervised discriminant projection (UDP) to build subspaces on WLBP featured periocular images and gain 100% rank-1 identification rate and 98% verification rate at 0.1% false accept rate on the entire FG-NET database. Compared to published results, our proposed approach yields the best recognition and identification results. Felix Juefei-Xu, Khoa Luu, Marios Savvides, Tien D. Bui, Ching Y. Suen |
IJCB | 3 |
| 2011 | Contourlet appearance model for facial age estimationabstractIn this paper we propose a novel Contourlet Appearance Model (CAM) that is more accurate and faster at localizing facial landmarks than Active Appearance Models (AAMs). Our CAM also has the ability to not only extract holistic texture information, as AAMs do, but can also extract local texture information using the Nonsubsampled Contourlet Transform (NSCT). We demonstrate the efficiency of our method by applying it to the problem of facial age estimation. Compared to previously published age estimation techniques, our approach yields more accurate results when tested on various face aging databases. Khoa Luu, Keshav Seshadri, Marios Savvides, Tien D. Bui, Ching Y. Suen |
IJCB | 3 |
| 2011 | Long range iris acquisition system for stationary and mobile subjectsabstractMost iris based biometric systems require a lot of co- operation from the users so that iris images of acceptable quality may be acquired. Features from these may then be used for recognition purposes. Relatively fewer works in literature address the question of less cooperative iris acquisition systems in order to reduce constraints on users. In this paper, we describe our ongoing work in designing and developing such a system. It is capable of capturing images of the iris up to distances of 8 meters with a resolution of 200 pixels across the diameter. If the resolution requirement is decreased to 150 pixels, then the same system may be used to capture images from up to 12 meters. We have incorporated velocity estimation and focus tracking modules so that images may be acquired from subjects on the move as well. We describe the various components that make up the system, including the lenses used, the imaging sensor, our auto-focus function and velocity estimation module. All the hardware components are Commercial Off The Shelf (COTS) with little or no modifications. We also present preliminary iris acquisition results using our system for both stationary and mobile subjects. Shreyas Venugopalan, Unni Prasad, Khalid Harun, Kyle Neblett, Douglas Toomey, Joseph Heyman, Marios Savvides |
IJCB | 7 |
| 2011 | An analysis of facial shape and texture for recognition: A large scale evaluation on FRGC ver2.0abstractTraditional approaches to face recognition have utilized aligned facial images containing both shape and texture information. This paper analyzes the contributions of the individual facial shape and texture components to face recognition. These two components are evaluated independently and we investigate methods to combine the information gained from each of them to enhance face recognition performance. The contributions of this paper are the following: (1) to the best of our knowledge, it is the first large-scale study of how face recognition is influenced by shape and texture as all of our results are benchmarked against traditional approaches on the challenging NIST FRGC ver2.0 experiment 4 dataset, (2) we empirically show that shape information is reasonably discriminative, (3) we demonstrate significant improvement in performance by registering texture with dense shape information, and finally (4) show that fusing shape and texture information consistently boosts recognition results across different subspace-based algorithms. Ramzi Abiantun, Utsav Prabhu, Keshav Seshadri, Jingu Heo, Marios Savvides |
WACV | 5 |
| 2011 | Unconstrained Pose-Invariant Face Recognition Using 3D Generic Elastic ModelsabstractClassical face recognition techniques have been successful at operating under well-controlled conditions; however, they have difficulty in robustly performing recognition in uncontrolled real-world scenarios where variations in pose, illumination, and expression are encountered. In this paper, we propose a new method for real-world unconstrained pose-invariant face recognition. We first construct a 3D model for each subject in our database using only a single 2D image by applying the 3D Generic Elastic Model (3D GEM) approach. These 3D models comprise an intermediate gallery database from which novel 2D pose views are synthesized for matching. Before matching, an initial estimate of the pose of the test query is obtained using a linear regression approach based on automatic facial landmark annotation. Each 3D model is subsequently rendered at different poses within a limited search space about the estimated pose, and the resulting images are matched against the test query. Finally, we compute the distances between the synthesized images and test query by using a simple normalized correlation matcher to show the effectiveness of our pose synthesis method to real-world data. We present convincing results on challenging data sets and video sequences demonstrating high recognition accuracy under controlled as well as unseen, uncontrolled real-world scenarios using a fast implementation. Utsav Prabhu, Jingu Heo, Marios Savvides |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | How to Generate Spoofed Irises From an Iris Code TemplateabstractBiometrics has gained a lot of attention over recent years as a way to identify individuals. Of all biometrics-based techniques, the iris-pattern-based systems have recently shown very high accuracies in verifying an individual's identity. The premise here is that iris patterns are unique across people. Only the iris bit code template specific to an individual need be stored for future identity verification. It is generally accepted that this iris bit code is unidentifiable data. However, in this work, we explore methods to generate alternate iris textures for a given person for the purpose of bypassing a system based on this iris bit code. We show that, if this spoof texture is presented to an iris recognition system, it will generate the same score response as that of the original iris texture. Hence, this approach can bypass filter-based feature extraction systems (such as Daugman style systems) without using the actual texture of the target iris that we want to spoof, by obtaining a hamming distance match score that falls within the authentic score range. This approach assumes we know the feature extraction mechanism of the iris matching scheme. We embed features within a person's natural iris texture to spoof another person's iris. A very convincing preliminary investigation into how one can get by any iris recognition system by synthesizing various levels of “natural” looking irises is presented here and we hope to use this knowledge to build countermeasures into the feature extraction scheme of the recognition module. Shreyas Venugopalan, Marios Savvides |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2010 | An extension of multifactor analysis for face recognition based on submanifold learningabstractLately, Multilinear Principal Component Analysis (MPCA) has been successfully applied to face recognition since MPCA provides analysis of multiple factors of face images such as people's identities, viewpoints, and lighting conditions. MPCA employees multiple linear subspaces constructed by varying factors. In this paper, we propose nonlinear submanifold analysis, which can represent the variation of each factor more accurately than the conventional multilinear subspace analysis. Based on submanifold learning, we propose an extension of the multiple factor analysis. This paper proposes the kernel-based extension of MPCA whose definition of a kernel function and neighbors of each sample is robust for submanifold learning. The experimental results in this paper demonstrate that the proposed methods produce a synergetic advantage for face recognition. This is because our method offers the combined virtues of both multifactor analysis and manifold learning. Sung Won Park, Marios Savvides |
CVPR | 2 |
| 2009 | A pixel-wise, learning-based approach for occlusion estimation of iris images in polar domainabstractOn normalized iris images, there are many kinds of noises, such as eyelids, eyelashes, shadows or specular reflections, that often occlude the true iris texture. If high recognition rate is desired, those occluded areas must be estimated accurately in order for them to be excluded during the matching stage. In this paper, we propose a unified, probabilistic and learning-based approach to estimate all kinds of occlusions within one unified model. Experiments have shown that our method not only estimates occlusion very accurately, but also does it with high speed, which makes it useful for practical iris recognition systems. Yung-hui Li, Marios Savvides |
ICASSP | 2 |
| 2008 | Evaluating Active Shape Models for Eye-Shape ClassificationabstractThis paper explores the goal of applying Active Shape Models (ASMs) on the eye images to classify eye shapes and identify whether the images belong from left or right irises. In many applications, particular to data collected from single eye capture devices (such as the PIER mobile iris image acquisition device), it is of importance to be able to sort and correct mislabeled collected data. ASMs have traditionally been applied for classification or identification of a wide variety of objects ranging from faces, assembly line objects to biomedical objects such as bone structures (like the spine etc). In this paper we apply and evaluate ASM models to fit on the eye shape to determine if the image belongs to a left or right eye. The approach we employ is based on building 2 ASM models, one for the left eye and one for right eye. The best fit model is chosen as the result. Our preliminary evaluation using vanilla ASM shows that preprocessing techniques like illumination compensation, shape normalization, and accurate Iris detection are key steps required to improve the classification performance. Shuvra Bhat, Marios Savvides |
ICASSP | 2 |
| 2008 | Face Recognition Across Pose Using View Based Active Appearance Models (VBAAMs) on CMU Multi-PIE Dataset
Jingu Heo, Marios Savvides |
ICVS | 2 |
| 2008 | Communication-Aware Face Detection Using Noc Architecture
Hung-Chih Lai, Radu Marculescu, Marios Savvides, Tsuhan Chen |
ICVS | 3 |
| 2007 | Graphical Model Approach to Iris Matching Under Deformation and OcclusionabstractTemplate matching of iris images for biometric recognition typically suffers from both local deformations between the template and query images and large occlusions from the eyelid. In this work, we model deformation and occlusion as a set of hidden variables for each iris comparison. We use afield of directional vectors to represent deformation and a field of binary variables to represent occlusion. We impose a probability distribution on these fields using a lattice-type undirected graphical model, in which the graph edges represent interdependencies between neighboring iris regions. Gabor wavelet-based similarity scores and intensity statistics are used as observations in the model. Loopy belief propagation is applied to estimate the conditional distributions on the hidden variables, which are in turn used to compute final match scores. We present underlying theory as well as experimental results from both the CASIA iris database and the database provided for the iris challenge evaluation (ICE). We show that our proposed method significantly improves recognition accuracy on these datasets over existing methods. Ryan A. Kerekes, Jason Thornton, Marios Savvides, B. V. K. Vijaya Kumar |
CVPR | 4 |
| 2007 | Kernel Fukunaga-Koontz Transform Subspaces For Enhanced Face RecognitionabstractTraditional linear Fukunaga-Koontz transform (FKT) (F. Fukunaga and W. Koontz, 1970) is a powerful discriminative subspaces building approach. Previous work has successfully extended FKT to be able to deal with small-sample-size. In this paper, we extend traditional linear FKT to enable it to work in multi-class problem and also in higher dimensional (kernel) subspaces and therefore provide enhanced discrimination ability. We verify the effectiveness of the proposed kernel Fukunaga-Koontz transform by demonstrating its effectiveness in face recognition applications; however the proposed non-linear generalization can be applied to any other domain specific problems. Yung-hui Li, Marios Savvides |
CVPR | 2 |
| 2007 | Generalized Low Dimensional Feature Subspace for Robust Face Recognition on Unseen datasets using Kernel Correlation Feature AnalysisabstractIn this paper we analyze and demonstrate the subspace generalization power of the kernel correlation feature analysis (KCFA) method for producing compact low dimensional subspace that has good representation ability to work on unseen, untrained datasets. Examining the portability of an algorithm across different datasets is an important practical aspect of face recognition applications where the technology cannot be dataset-dependant in real-world practical applications. In most face recognition literature, algorithms are demonstrated on datasets by training on some part of the dataset and testing on the remainder. In general, the training and testing data have the same people but different capture sessions so essentially, some of the expected variation and people are modeled in the training set. In this paper we describe how we efficiently build a compact feature space using kernel correlation filter analysis on the generic training set of the FRGC dataset, and test the built subspace on other well-known face datasets. We show that the feature subspace produced by KCFA has good representation and discrimination to unseen datasets and produces good verification and identification rates compared to other subspace methods such as PCA. Its efficiency, lower dimensionality (the KCFA is only a 222 dimensional subspace) and discriminative power make it more practical and powerful than PCA as a powerful lower dimensionality reduction method for modeling faces and facial variations. Ramzi Abiantun, Marios Savvides, B. V. K. Vijaya Kumar |
ICASSP (1) | 2 |
| 2007 | Analyzing Facial Images using Empirical Mode Decomposition for Illumination Artifact Removal and Improved Face RecognitionabstractA popular modality of biometrics, facial recognition is effective when used in controlled environments as in those situations where factors such as camera position, facial expression, and illumination effects are either completely or partially controlled in a beneficial way. Regulation of such factors has an immediate effect on the performance of facial recognition algorithms, in particular illumination effects which can not be controlled by even the most cooperative of users, in this paper we describe a method to address illumination effects in the biometric modality of face recognition using the signal processing analysis tool of empirical mode decomposition (EMD) to decompose images into their intrinsic mode function that correspond to the dominant illumination factors. Using these illumination modes we reconstruct the facial image without these illumination distortion components to synthesize a more illumination neutral facial image. We then perform verification experiments using algorithms such as principal component analysis (PCA), Fisher linear discriminant analysis (FEDA), and advanced correlation filters (ACF's) to demonstrate the fundamental effectiveness of EMD as an illumination compensation method. Results are reported on the Carnegie Mellon University pose-illumination-expression (CMU PIE) database. Ramamurthy Bhagavatula, Marios Savvides |
ICASSP (1) | 2 |
| 2007 | Breaking the Limitation of Manifold Analysis for Super-Resolution of Facial ImagesabstractA novel method for robust super-resolution of face images is proposed in this paper. Face super-resolution is a particular interest in video surveillance where face images have typically very low-resolution quality and there is a need to apply face enhancement or super-resolution algorithms. In this paper, we apply a manifold learning method which has hardly been used for super-resolution. A manifold is a natural generalization of a Euclidean space to a locally Euclidean space. Manifold learning algorithms are more powerful than other pattern recognition methods which analyze a Euclidean space because they can reveal the underlying nonlinear distribution of the face space; however, there are some practical problems which prevent these algorithms from being applied to super-resolution. Almost all of the manifold learning methods cannot generate mapping functions for new test images which are absent from a training set. Another factor is that super-resolution seeks to recover a high-dimensional image from a lower-dimensional one while manifold learning methods perform the exact opposite as they are applied to dimensionality reduction. In this paper, we break the limitation of applying manifold learning methods for face super-resolution by proposing a novel method using locality preserving projections (LPP). Sung Won Park, Marios Savvides |
ICASSP (1) | 2 |
| 2007 | Statistical Performance Evaluation of Biometric Authentication Systems Using Random Effects ModelsabstractAs biometric authentication systems become more prevalent, it is becoming increasingly important to evaluate their performance. This paper introduces a novel statistical method of performance evaluation for these systems. Given a database of authentication results from an existing system, the method uses a hierarchical random effects model, along with Bayesian inference techniques yielding posterior predictive distributions, to predict performance in terms of error rates using various explanatory variables. By incorporating explanatory variables as well as random effects, the method allows for prediction of error rates when the authentication system is applied to potentially larger and/or different groups of subjects than those originally documented in the database. We also extend the model to allow for prediction of the probability of a false alarm on a "watch-list" as a function of the list size. We consider application of our methodology to three different face authentication systems: a filter-based system, a Gaussian Mixture Model (GMM)-based system, and a system based on frequency domain representation of facial asymmetry. Sinjini Mitra, Marios Savvides, Anthony Brockwell |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Correction to "Statistical Performance Evaluation of Biometric Authentication Systems Using Random Effects Models"abstractThe authors of the above titled paper (ibid., vol. 29, no. 4, pp. 517-530, Apr 07) point out that the caption in Fig. 2 should read differently. The revised caption is provided here: http://csdl.computer.org/comp/trans/tp/2007/09/tp01.pdf Sinjini Mitra, Marios Savvides, Anthony Brockwell |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | A Bayesian Approach to Deformed Pattern Matching of Iris ImagesabstractWe describe a general probabilistic framework for matching patterns that experience in-plane nonlinear deformations, such as iris patterns. Given a pair of images, we derive a maximum a posteriori probability (MAP) estimate of the parameters of the relative deformation between them. Our estimation process accomplishes two things simultaneously: It normalizes for pattern warping and it returns a distortion-tolerant similarity metric which can be used for matching two nonlinearly deformed image patterns. The prior probability of the deformation parameters is specific to the pattern-type and, therefore, should result in more accurate matching than an arbitrary general distribution. We show that the proposed method is very well suited for handling iris biometrics, applying it to two databases of iris images which contain real instances of warped patterns. We demonstrate a significant improvement in matching accuracy using the proposed deformed Bayesian matching methodology. We also show that the additional computation required to estimate the deformation is relatively inexpensive, making it suitable for real-time applications. Jason Thornton, Marios Savvides, B. V. K. Vijaya Kumar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Palmprint Classification Using Multiple Advanced Correlation Filters and Palm-Specific SegmentationabstractWe propose a palmprint classification algorithm with the use of multiple correlation filters per class. Correlation filters are two-class classifiers that produce a sharp peak when filtering a sample of their class and a noisy output otherwise. For every class, we train the filters for a palm at different locations, where the palmprint region has a high degree of line content. With the use of a line detection procedure and a simple line energy measure, any region of the palm can be scored and the top-ranked regions are used to train the filters for each class. Using an enhanced palmprint segmentation algorithm, our proposed classifier achieves an average equal error rate of 1.12 times10-4% on a large database of 385 classes using multiple filters of size 64 times 64 pixels. The average false acceptance rate when the false rejection rate is zero is 2.25 times10-4%. Pablo H. Hennings-Yeomans, B. V. K. Vijaya Kumar, Marios Savvides |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2007 | Individual Kernel Tensor-Subspaces for Robust Face Recognition: A Computationally Efficient Tensor Framework Without Requiring Mode FactorizationabstractFacial images change appearance due to multiple factors such as different poses, lighting variations, and facial expressions. Tensors are higher order extensions of vectors and matrices, which make it possible to analyze different appearance factors of facial variation. Using higher order tensors, we can construct a multilinear structure and model the multiple factors of face variation. In particular, among the appearance factors, the factor of a person's identity modeled by a tensor structure can be used for face recognition. However, this tensor-based face recognition creates difficulty in factorizing the unknown parameters of a new test image and solving for the person-identity parameter. In this paper, to break this limitation of applying the tensor-based methods to face recognition, we propose a novel tensor approach based on an individual-modeling method and nonlinear mappings. The proposed method does not require the problematic tensor factorization and is more efficient than the traditional TensorFaces method with respect to computation and memory. We set up the problem of solving for the unknown factors as a least squares problem with a quadratic equality constraint and solve it using numerical optimization techniques. We show that an individual-multilinear approach reduces the order of the tensor so that it makes face-recognition tasks computationally efficient as well as analytically simpler. We also show that nonlinear kernel mappings can be applied to this optimization problem and provide more accuracy to face-recognition systems than linear mappings. In this paper, we show that the proposed method, Individual Kernel TensorFaces, produces the better discrimination power for classification. The novelty in our approach as compared to previous work is that the Individual Kernel TensorFaces method does not require estimating any factor of a new test image for face recognition. In addition, we do not need to have any a priori knowledge of or assumption about the factors of a test image when using the proposed method. We can apply Individual Kernel TensorFaces even if the factors of a test image are absent from the training set. Based on various experiments on the Carnegie Mellon University Pose, Illumination, and Expression database, we demonstrate that the proposed method produces reliable results for face recognition. Sung Won Park, Marios Savvides |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2006 | Face Recognition with Kernel Correlation Filters on a Large Scale DatabaseabstractRecently, Direct Linear Discriminant Analysis (D-LDA) and Gram-Schmidt LDA methods have been proposed for face recognition. By also utilizing some of the null-space of the within-class scatter matrix, they exhibit better performance compared to Fisherfaces and Eigenfaces. However, these linear subspace methods may not discriminate faces well due to large nonlinear distortions in the face images. Redundant class dependence feature analysis (CFA) method exhibits superior performance compared to other methods by representing nonlinear features well. We show that with a proper choice of kernel parameters used with the proposed Kernel Correlation Filters within the CFA framework, the overall face recognition performance is significantly improved. We present results of this proposed approach on a large scale database from the Face Recognition Grand Challenge (FRGC) which contains over 36,000 images. Jingu Heo, Marios Savvides, Ramzi Abiantun, Chunyan Xie, B. V. K. Vijaya Kumar |
ICASSP (2) | 2 |
| 2006 | Illumination Tolerant Face Recognition Using a Novel Face From Sketch Synthesis Approach and Advanced Correlation FiltersabstractCurrent state-of-the-art approach for performing face sketch recognition transforms all the test face images into sketches, and then performs recognition on sketch domain using the sketch composite. In our approach we propose the opposite; which has advantages in a real-time sysrtem; we propose to generate a realistic face image from the composite sketch using a Hybrid subspace method and then build an illumination tolerant correlation filter which can recognize the person under different illumination variations from a surveillance video footage. We show how effective proposed algorithm works on the CMU PIE (Pose Illumination and Expression) database. Yung-hui Li, Marios Savvides, B. V. K. Vijaya Kumar |
ICASSP (2) | 2 |
| 2006 | Improved Human Face Identification Using Frequency Domain Representation of Facial AsymmetryabstractThis paper explores the role of facial asymmetry in identification tasks using a frequency domain representation. Satisfactory results are obtained for two different tasks, namely, human identification under extreme expression variations and expression classification, using a PCA-type classifier which establishes the robustness of these measures to intra-personal distortions. We next demonstrate that it is possible to even improve upon these results by simple means. In particular, we use two methods, namely, feature set combination and statistical resampling methods like bagging, which attains perfect classification results (0% error rate) in some cases. Both these methods require very few additional resources in terms of computing power, hence they are useful for practical applications as well Sinjini Mitra, Marios Savvides |
ICASSP (2) | 2 |
| 2006 | Class Dependent Kernel Discrete Cosine Transform Features for Enhanced Holistic Face Recognition in FRGC-IIabstractFace recognition is one of the least intrusive biometric modalities that can be used to identify individuals from surveillance video. In such scenarios the users are under the least co-operative conditions and thus the ability to perform robust face recognition in such scenarios is very challenging. In this paper we focus on improving the face recognition performance on a large database with over 36,000 facial images from the Face Recognition Grand Challenge Phase-II data collected by University of Notre Dame. We particularly focus on Experiment 4 which is the most challenging and captured in uncontrolled conditions where the baseline PCA algorithm yields 12% verification rate at 0.1% FAR. We propose a novel approach using class-dependent kernel discrete cosine transform features which improves the performance significantly yielding a 91.33% verification rate at 0.1% FAR, and we also show that by working in the DCT transform domain for obtaining non-linear features is more optimal than working in the original spatial-pixel domain which only yields a verification rate of 85% at 0.1% FAR. Thus our proposed method outperforms the baseline by 79.33% in verification rate @0.1% False Acceptance Rate. Marios Savvides, Jingu Heo, Ramzi Abiantun, Chunyan Xie, B. V. K. Vijaya Kumar |
ICASSP (2) | 1 |
| 2006 | Correlation Pattern Recognition for Face RecognitionabstractTwo-dimensional (2-D) face recognition (FR) is of interest in many verification (1:1 matching) and identification ($1:N$matching) applications because of its nonintrusive nature and because digital cameras are becoming ubiquitous. However, the performance of 2-D FR systems can be degraded by natural factors such as expressions, illuminations, pose, and aging. Several FR algorithms have been proposed to deal with the resulting appearance variability. However, most of these methods employ features derived in the image or the space domain whereas there are benefits to working in the spatial frequency domain (i.e., the 2-D Fourier transforms of the images). These benefits include shift-invariance, graceful degradation, and closed-form solutions. We discuss the use of spatial frequency domain methods (also known as correlation filters or correlation pattern recognition) for FR and illustrate the advantages. However, correlation filters can be computationally demanding due to the need for computing 2-D Fourier transforms and may not match well for large-scale FR problems such as in the Face Recognition Grand Challenge (FRGC) phase-II experiments that require the computation of millions of similarity metrics. We will discuss a new method [called the class-dependence feature analysis (CFA)] that reduces the computational complexity of correlation pattern recognition and show the results of applying CFA to the FRGC phase-II data. B. V. K. Vijaya Kumar, Marios Savvides, Chunyan Xie |
Proc. IEEE | 2 |
| 2006 | Face identification using novel frequency-domain representation of facial asymmetryabstractFace recognition is a challenging task. This paper introduces a novel set of biometrics, defined in the frequency domain and representing a form of "facial asymmetry." A comparison with existing spatial asymmetry measures suggests that the frequency-domain representation provides an efficient approach for performing human identification in the presence of severe expressions and for expression classification. Error rates of less than 5% are observed for human identification and around 25% for expression classification on a database of 55 individuals. Feature analysis indicates that asymmetry of the different face parts helps in these two apparently conflicting classification problems. An interesting connection between asymmetry and the Fourier domain phase spectra is then established. Finally, a compact one-bit frequency-domain representation of asymmetry is introduced, and a simplistic Hamming distance classifier is shown to be more efficient than traditional classifiers from storage and the computation point of view, while producing equivalent human identification results. In addition, the application of these compact measures to verification and a statistical analysis are presented Sinjini Mitra, Marios Savvides, B. V. K. Vijaya Kumar |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2005 | Analyzing asymmetry biometric in the frequency domain for face recognitionabstractThe paper introduces a novel set of facial biometrics based on quantified facial asymmetry measures in the frequency domain. In particular, we show that these biometrics work well for images showing expression variations. A comparison of the recognition rates with those obtained from spatial domain asymmetry measures based on raw intensity values suggests that the frequency domain representation is more robust to intra-personal distortions and, indeed, provides an efficient approach for performing classification or recognition. The role of asymmetry of the different regions (e.g., eyes, mouth, nose) of the face is investigated to determine which regions provide the maximum discrimination among individuals in the presence of different expressions for better classification results in such a scenario. Sinjini Mitra, Marios Savvides |
ICASSP (2) | 2 |
| 2005 | Quaternion Correlation Filters for Face Recognition in Wavelet DomainabstractA new frequency domain face recognition method using wavelet decomposition and quaternion correlation filters is proposed. The wavelet decomposition of the face image leads to a wavelet subband representation, which contains four subband images corresponding to four orthogonal channels. These four subbands can be encoded into a 2D quaternion number array. The quaternion correlation filter method is developed to perform pattern recognition jointly on multichannel 2D signals. The proposed method has been shown to achieve significant improvement in face recognition results compared to the traditional advanced correlation filter method for handling illumination variations of face images. We experimented with the CMU PIE database consisting of 65 people with 21 illumination variations per person, showing that our method can achieve close to 100% recognition accuracy using just a single training image of a person under neutral frontal lighting and testing on all other unseen harsh illumination conditions. Chunyan Xie, Marios Savvides, B. V. K. Vijaya Kumar |
ICASSP (2) | 2 |
| 2004 | "Corefaces" - Robust Shift Invariant PCA Based Correlation Filter for Illumination Tolerant Face Recognition
Marios Savvides, B. V. K. Vijaya Kumar, Pradeep K. Khosla |
CVPR (2) | 1 |
| 2003 | Efficient Design of Advanced Correlation Filters for Robust Distortion-Tolerant Face RecognitionabstractThe paper summarizes new research in performing face recognition using advanced correlation filters. We examine the performance of such filters in the area of biometrics for face authentication. We also compare results when the filters are applied to face identification. Our results are based on the illumination subsets of the CMU PIE database. We also present methods that reduce the memory requirements of these filters to run on limited computational resources, including computationally efficient methods of synthesizing these filters. Finally, we describe an online training algorithm implemented on a face verification system for synthesizing correlation filters from a video stream to handle pose/scale variations. The system also uses an efficient scheme to perform face localization within the current framework during the authentication stage. Marios Savvides, B. V. K. Vijaya Kumar |
AVSS | 1 |
| 2003 | Incremental updating of advanced correlation filters for biometric authentication systemsabstractIn this paper we show mathematical formulation of incrementally building advanced correlation filters used in authentication systems that are based on face and fingerprint images as biometrics for verification. This method is crucial for incorporating such algorithms on small devices with limited memory and computational resources. We also present results that show that these correlation filters perform well for face and fingerprint images. We used the PIE (pose, illumination and expression) database from CMY to test the verification performance using face images. Similarly, for fingerprint images we used the NIST special database 24 to evaluate verification performance. Marios Savvides, Krithika Venkataramani, B. V. K. Vijaya Kumar |
ICME | 1 |
| 2002 | Spatial frequency domain image processing for biometric recognitionabstractBiometric recognition refers to the process of matching an input biometric to stored biometric information. In particular, biometric verification refers to matching the live biometric input from an individual to the stored biometric template about that individual. Examples of biometrics include face images, fingerprint images, iris images, retinal scans, etc. Thus, image processing techniques prove useful in the biometric recognition. We discuss spatial frequency domain image processing methods useful for biometric recognition. B. V. K. Vijaya Kumar, Chunyan Xie, Marios Savvides, Krithika Venkataramani |
ICIP (1) | 3 |