Hazim Kemal Ekenel

dblp:26/1396 · DBLP profile ↗
← Back
68ranked-venue papers
11as first author
20since 2021 · last 2026
0000-0003-3697-8548ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 49 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 39 · 7 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 since 2021Security and privacy · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 On Applicability of Synthetic Datasets for Facial Expression Recognition
Ali Azmoudeh, Erdi Saritas, Ömer Yildirim, Hazim Kemal Ekenel
FG4
2026 Employing Vision-Language Models for Face Image Quality Assessment
Erdi Saritas, Eren Onaran, Vitomir Struc, Hazim Kemal Ekenel
FG4
2026 CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
abstract
We address Embodied Reference Understanding, the task of predicting the object a person in the scene refers to through pointing gesture and language. This requires multimodal reasoning over text, visual pointing cues, and scene context, yet existing methods often fail to fully exploit visual disambiguation signals. We also observe that while the referent often aligns with the head-to-fingertip direction, in many cases it aligns more closely with the wrist-to-fingertip direction, making a single-line assumption overly limiting. To address this, we propose a dual-model framework, where one model learns from the head-to-fingertip direction and the other from the wrist-to-fingertip direction. We introduce a Gaussian ray heatmap representation of these lines and use them as input to provide a strong supervisory signal that encourages the model to better attend to pointing cues. To fuse their complementary strengths, we present the CLIP-Aware Pointing Ensemble module, which performs a hybrid ensemble guided by CLIP features. We further incorporate an auxiliary object center prediction head to enhance referent localization. We validate our approach on YouRefIt, achieving 75.0 mAP at 0.25 IoU, alongside state-of-the-art CLIP and CDscores, and demonstrate its generality on unseen CAESAR and ISL Pointing, showing robust performance across benchmarks.
Fevziye Irem Eyiokur, Doggucan Yaman, Hazim Kemal Ekenel, Alex Waibel
WACV3
2026 A survey on class-agnostic counting: Advancements from reference-based to open-world text-guided approaches
abstract
Visual object counting has recently shifted towards class-agnostic counting (CAC), which addresses the challenge of counting objects across arbitrary categories—a crucial capability for flexible and generalizable counting systems. Unlike humans, who effortlessly identify and count objects from diverse categories without prior knowledge, most existing counting methods are restricted to enumerating instances of known classes, requiring extensive labeled datasets for training and struggling in open-vocabulary settings. In contrast, CAC aims to count objects belonging to classes never seen during training, operating in a few-shot setting. In this paper, we present the first comprehensive review of CAC methodologies. We propose a taxonomy to categorize CAC approaches into three paradigms based on how target object classes can be specified: reference-based, reference-less, and open-world text-guided. Reference-based approaches achieve state-of-the-art performance by relying on exemplar-guided mechanisms. Reference-less methods eliminate exemplar dependency by leveraging inherent image patterns. Finally, open-world text-guided methods use vision-language models, enabling object class descriptions via textual prompts, offering a flexible and promising solution. Based on this taxonomy, we provide an overview of 30 CAC architectures and report their performance on gold-standard benchmarks, discussing key strengths and limitations. Specifically, we present results on the FSC-147 dataset, setting a leaderboard using gold-standard metrics, and on the CARPK dataset to assess generalization capabilities. Finally, we offer a critical discussion of persistent challenges, such as annotation dependency and generalization, alongside future directions. We believe this survey offers a valuable resource for researchers, capturing the evolution of CAC and providing insights to guide future developments in the field. • Provides the first comprehensive survey on class-agnostic object counting (CAC). • Proposes a taxonomy: reference-based, reference-less, and open-world text-guided. • Benchmarks 30 CAC methods on FSC-147 and CARPK, setting a leaderboard. • Discusses trade-offs between accuracy, flexibility, and annotation requirements. • Identifies open challenges and future directions for generalizable counting.
Luca Ciampi, Ali Azmoudeh, Elif Ecem Akbaba, Erdi Saritas, Ziya Ata Yazici, Hazim Kemal Ekenel, Giuseppe Amato 0001, Fabrizio Falchi
Comput. Vis. Image Underst.6
2025 BaMCo: Balanced Multimodal Contrastive Learning for Knowledge-Driven Medical VQA
Ziya Ata Yazici, Hazim Kemal Ekenel
MICCAI (7)2
2025 Bias-aware face mask detection dataset
abstract
Abstract In December 2019, a novel coronavirus (COVID-19) spread so quickly around the world that many countries had to set mandatory face mask rules in public areas to reduce the transmission of the virus. To monitor public adherence, researchers aimed to rapidly develop efficient systems that can detect faces with masks automatically. However, the lack of representative and novel datasets posed challenges for training efficient models. Early attempts to collect face mask datasets did not account for potential race, gender, and age biases. Therefore, the resulting models show inherent biases toward specific race groups, such as Asian or Caucasian. In this work, we present a novel face mask detection dataset that contains images posted on Twitter during the pandemic from around the world. Unlike previous datasets, the proposed Bias-Aware Face Mask Detection (BAFMD) dataset contains more images from underrepresented races and age groups to mitigate the problem of the face mask detection task. We perform experiments to investigate potential biases in widely used face mask detection datasets and illustrate that the BAFMD dataset yields models with better performance and generalization ability. The dataset is publicly available at https://github.com/Alpkant/BAFMD .
Alperen Kantarci, Ferda Ofli, Muhammad Imran 0002, Hazim Kemal Ekenel
Multim. Tools Appl.4
2024 Audio-Driven Talking Face Generation with Stabilized Synchronization Loss
Doggucan Yaman, Fevziye Irem Eyiokur, Leonard Bärmann, Hazim Kemal Ekenel, Alex Waibel
ECCV (19)4
2024 Welcome
abstract
It was our pleasure and privilege to welcome you to Istanbul for the 18th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2024). We hope your experience at FG was rewarding both professionally and personally!
Hazim Kemal Ekenel, Albert Ali Salah, Arun Ross, Vitomir Struc, Lale Akarun, Xilin Chen 0001, Shaun J. Canavan
FG1
2024 Analyzing the Feature Extractor Networks for Face Image Synthesis
abstract
Advancements like Generative Adversarial Networks have attracted the attention of researchers toward face image synthesis to generate ever more realistic images. Thereby, the need for the evaluation criteria to assess the realism of the generated images has become apparent. While FID utilized with InceptionV3 is one of the primary choices for benchmarking, concerns about InceptionV3 's limitations for face images have emerged. This study investigates the behavior of diverse feature extractors - InceptionV3, CLIP, DINOv2, and ArcFace - considering a variety of metrics - FID, KID, Precision&Recall. While the FFHQ dataset is used as the target domain, as the source domains, the CelebA-HQ dataset and the synthetic datasets generated using Style-GAN2 and Projected FastGAN are used. Experiments include deep-down analysis of the features:$L_{2}$normalization, model attention during extraction, and domain distributions in the feature space. We aim to give valuable insights into the behavior of feature extractors for evaluating face image synthesis methodologies. The code is publicly available at https://github.com/ThEnded32/AnalyzingFeatureExtractors.
Erdi Saritas, Hazim Kemal Ekenel
FG2
2024 Analyzing the Effect of Combined Degradations on Face Recognition
abstract
A face recognition model is typically trained on large datasets of images that may be collected from controlled environments. This results in performance discrepancies when applied to real-world scenarios due to the domain gap between clean and in-the-wild images. Therefore, some researchers have investigated the robustness of these models by analyzing synthetic degradations. Yet, existing studies have mostly focused on single degradation factors, which may not fully capture the complexity of real-world degradations. This work addresses this problem by analyzing the impact of both single and combined degradations using a real-world degradation pipeline extended with under/over-exposure conditions. We use the LFW dataset for our experiments and assess the model's performance based on verification accuracy. Results reveal that single and combined degradations show dissimilar model behavior. The combined effect of degradation significantly lowers performance even if its single effect is negligible. This work emphasizes the importance of accounting for real-world complexity to assess the robustness of face recognition models in real-world settings. The code is publicly available at https://github.com/ThEnded32/AnalyzingCombinedDegradations
Erdi Saritas, Hazim Kemal Ekenel
FG2
2024 GLIMS: Attention-guided lightweight multi-scale hybrid network for volumetric semantic segmentation
Ziya Ata Yazici, Ilkay Öksüz, Hazim Kemal Ekenel
Image Vis. Comput.3
2023 Attention Assessment in Children with Autism Using Head Pose and Motion Parameters from Real Videos
abstract
In children with autism spectrum disorders (ASD), attention assessment plays a crucial role in understanding their behavioral and cognitive functioning. Difficulties with attention are a common feature of children with autism and have a significant impact on their ability to learn and socialize. In this paper, we propose a non-invasive and objective method to assess attention in children with autism from real videos by utilizing the head poses and motion parameters. The proposed approach is an ensemble of a deep learning model that extracts head pose parameters, an optical flow approach that extracts motion parameters from consecutive frames, temporal head pose parameters extraction and an autoencoder for attention assessment. The experimental study was conducted on 39 children (ASD = 19, neurotypical children = 20) by giving different attention tasks and capturing their video using an attached webcam. Results are analyzed for participant and task differences, which demonstrate that our approach is successful in measuring a child's attention control and inattention. In particular, the assessment of the head poses and motion parameters will enable the development of real-time attention recognition systems that can be used for both learning and targeted intervention.
Elizabeth B. Varghese, Marwa Qaraqe, Dena Al-Thani, Hazim Kemal Ekenel
GLOBECOM4
2023 The Unconstrained Ear Recognition Challenge 2023: Maximizing Performance and Minimizing Bias
abstract
The paper provides a summary of the 2023 Unconstrained Ear Recognition Challenge (UERC), a benchmarking effort focused on ear recognition from images acquired in uncontrolled environments. The objective of the challenge was to evaluate the effectiveness of current ear recognition techniques on a challenging ear dataset while analyzing the techniques from two distinct aspects, i.e., verification performance and bias with respect to specific demographic factors, i.e., gender and ethnicity. Seven research groups participated in the challenge and submitted a seven distinct recognition approaches that ranged from descriptor-based methods and deep-learning models to ensemble techniques that relied on multiple data representations to maximize performance and minimize bias. A comprehensive investigation into the performance of the submitted models is presented, as well as an in-depth analysis of bias and associated performance differentials due to differences in gender and ethnicity. The results of the challenge suggest that a wide variety of models (e.g., transformers, convolutional neural networks, ensemble models) is capable of achieving competitive recognition results, but also that all of the models still exhibit considerable performance differentials with respect to both gender and ethnicity. To promote further development of unbiased and effective ear recognition models, the starter kit of UERC 2023 together with the baseline model, and training and test data is made available from: http://ears.fri.uni-lj.si/
Ziga Emersic, Tetsushi Ohki, Muku Akasaka, Takahiko Arakawa, Soshi Maeda, Masora Okano, Yuya Sato, Anjith George, Sébastien Marcel, Iyyakutti Iyappan Ganapathi, Syed Sadaf Ali, Sajid Javed, Naoufel Werghi, S. G. Isik, Erdi Saritas, Hazim Kemal Ekenel, V. Hudovernik, Jan Niklas Kolf, Fadi Boutros, Naser Damer, G. Sharma, Aman Kamboj, Aditya Nigam, Deepak Kumar Jain 0001, G. Cámara-Chávez, Peter Peer, Vitomir Struc
IJCB16
2023 A survey on computer vision based human analysis in the COVID-19 era
Fevziye Irem Eyiokur, Alperen Kantarci, Mustafa Ekrem Erakin, Naser Damer, Ferda Ofli, Muhammad Imran 0002, Janez Krizaj, Albert Ali Salah, Alex Waibel, Vitomir Struc, Hazim Kemal Ekenel
Image Vis. Comput.11
2022 OCFR 2022: Competition on Occluded Face Recognition from Synthetically Generated Structure-Aware Occlusions
abstract
This work summarizes the IJCB Occluded Face Recognition Competition 2022 (IJCB-OCFR-2022) embraced by the 2022 International Joint Conference on Biometrics (IJCB 2022). OCFR-2022 attracted a total of 3 participating teams, from academia. Eventually, six valid submissions were submitted and then evaluated by the organizers. The competition was held to address the challenge of face recognition in the presence of severe face occlusions. The participants were free to use any training data and the testing data was built by the organisers by synthetically occluding parts of the face images using a well-known dataset. The submitted solutions presented innovations and performed very competitively with the considered baseline. A major output of this competition is a challenging, realistic, and diverse, and publicly available occluded face recognition benchmark with well defined evaluation protocols.
Pedro C. Neto, Fadi Boutros, João Ribeiro Pinto, Naser Damer, Ana Filipa Sequeira, Jaime S. Cardoso 0001, Messaoud Bengherabi, Abderaouf Bousnat, Sana Boucheta, Nesrine Hebbadj, Mustafa Ekrem Erakin, Ugur Demir, Hazim Kemal Ekenel, Pedro Vidal 0001, David Menotti
IJCB13
2022 Multimodal soft biometrics: combining ear and face biometrics for age and gender classification
Doggucan Yaman, Fevziye Irem Eyiokur, Hazim Kemal Ekenel
Multim. Tools Appl.3
2022 MOCCA: Multilayer One-Class Classification for Anomaly Detection
abstract
Anomalies are ubiquitous in all scientific fields and can express an unexpected event due to incomplete knowledge about the data distribution or an unknown process that suddenly comes into play and distorts the observations. Usually, due to such events’ rarity, to train deep learning (DL) models on the anomaly detection (AD) task, scientists only rely on “normal” data, i.e., nonanomalous samples. Thus, letting the neural network infer the distribution beneath the input data. In such a context, we propose a novel framework, named multilayer one-class classification (MOCCA), to train and test DL models on the AD task. Specifically, we applied our approach to autoencoders. A key novelty in our work stems from the explicit optimization of the intermediate representations for the task at hand. Indeed, differently from commonly used approaches that consider a neural network as a single computational block, i.e., using the output of the last layer only, MOCCA explicitly leverages the multilayer structure of deep architectures. Each layer’s feature space is optimized for AD during training, while in the test phase, the deep representations extracted from the trained layers are combined to detect anomalies. With MOCCA, we split the training process into two steps. First, the autoencoder is trained on the reconstruction task only. Then, we only retain the encoder tasked with minimizing the$L_{2}$distance between the output representation and a reference point, the anomaly-free training data centroid, at each considered layer. Subsequently, we combine the deep features extracted at the various trained layers of the encoder model to detect anomalies at inference time. To assess the performance of the models trained with MOCCA, we conduct extensive experiments on publicly available datasets, namely CIFAR10, MVTec AD, and ShanghaiTech. We show that our proposed method reaches comparable or superior performance to state-of-the-art approaches available in the literature. Finally, we provide a model analysis to give insights regarding the benefits of our training procedure.
Fabio Valerio Massoli, Fabrizio Falchi, Alperen Kantarci, Seymanur Akti, Hazim Kemal Ekenel, Giuseppe Amato 0001
IEEE Trans. Neural Networks Learn. Syst.5
2021 MFR 2021: Masked Face Recognition Competition
abstract
This paper presents a summary of the Masked Face Recognition Competitions (MFR) held within the 2021 International Joint Conference on Biometrics (IJCB 2021). The competition attracted a total of 10 participating teams with valid submissions. The affiliations of these teams are diverse and associated with academia and industry in nine different countries. These teams successfully submitted 18 valid solutions. The competition is designed to motivate solutions aiming at enhancing the face recognition accuracy of masked faces. Moreover, the competition considered the deployability of the proposed solutions by taking the compactness of the face recognition models into account. A private dataset representing a collaborative, multisession, real masked, capture scenario is used to evaluate the submitted solutions. In comparison to one of the topperforming academic face recognition solutions, 10 out of the 18 submitted solutions did score higher masked face verification accuracy.
Fadi Boutros, Naser Damer, Jan Niklas Kolf, Kiran B. Raja, Florian Kirchbuchner, Ramachandra Raghavendra, Arjan Kuijper, Pengcheng Fang, Fei Wang 0032, David Montero 0002, Naiara Aginako, Basilio Sierra, Marcos Nieto Doncel, Mustafa Ekrem Erakin, Ugur Demir, Hazim Kemal Ekenel, Asaki Kataoka, Kohei Ichikawa, Shizuma Kubo, Jie Zhang 0071, Shiguang Shan, Klemen Grm, Vitomir Struc, Sachith Seneviratne, Nuran Kasthuriarachchi, Sanka Rasnayaka, Pedro C. Neto, Ana Filipa Sequeira, João Ribeiro Pinto, Mohsen Saffari, Jaime S. Cardoso 0001
IJCB17
2021 Face Liveness Detection Competition (LivDet-Face) - 2021
abstract
Liveness Detection (LivDet)-Face is an international competition series open to academia and industry. The competition’s objective is to assess and report state-of-the-art in liveness / Presentation Attack Detection (PAD) for face recognition. Impersonation and presentation of false samples to the sensors can be classified as presentation attacks and the ability for the sensors to detect such attempts is known as PAD. LivDet-Face 2021 * will be the first edition of the face liveness competition. This competition serves as an important benchmark in face presentation attack detection, offering (a) an independent assessment of the current state of the art in face PAD, and (b) a common evaluation protocol, availability of Presentation Attack Instruments (PAI) and live face image dataset through the Biometric Evaluation and Testing (BEAT) platform. The competition can be easily followed by researchers after it is closed, in a platform in which participants can compare their solutions against the LivDet-Face winners.
Sandip Purnapatra, Nic Smalt, Keivan Bahmani, Priyanka Das 0004, David Yambay, Amir Mohammadi, Anjith George, Thirimachos Bourlai, Sébastien Marcel, Stephanie Schuckers, Meiling Fang, Naser Damer, Fadi Boutros, Arjan Kuijper, Alperen Kantarci, Basar Demir, Zafer Yildiz, Zabi Ghafoory, Hasan Dertli, Hazim Kemal Ekenel, Ngoc-Son Vu, Vassilis Christophides, Dashuang Liang, Zhanlong Hao, Junfu Liu, Yufeng Jin, Samo Liu, Salieri Kuei, Jag Mohan Singh, Ramachandra Raghavendra
IJCB20
2021 Benefiting from Bicubically Down-Sampled Images for Learning Real-World Image Super-Resolution
abstract
Super-resolution (SR) has traditionally been based on pairs of high-resolution images (HR) and their low-resolution (LR) counterparts obtained artificially with bicubic downsampling. However, in real-world SR, there is a large variety of realistic image degradations and analytically modeling these realistic degradations can prove quite difficult. In this work, we propose to handle real-world SR by splitting this ill-posed problem into two comparatively more well-posed steps. First, we train a network to transform real LR images to the space of bicubically down-sampled images in a supervised manner, by using both real LR/HR pairs and synthetic pairs. Second, we take a generic SR network trained on bicubically downsampled images to super-resolve the transformed LR image. The first step of the pipeline addresses the problem by registering the large variety of degraded images to a common, well understood space of images. The second step then leverages the already impressive performance of SR on bicubically downsampled images, sidestepping the issues of end-to-end training on datasets with many different image degradations. We demonstrate the effectiveness of our proposed method by comparing it to recent methods in real-world SR and show that our proposed approach outperforms the state-of-the-art works in terms of both qualitative and quantitative results, as well as results of an extensive user study conducted on several real image datasets.
Mohammad Saeed Rad, Thomas Yu, Claudiu Cristian Musat, Hazim Kemal Ekenel, Behzad Bozorgtabar, Jean-Philippe Thiran
WACV4
2020 Benefiting from multitask learning to improve single image super-resolution
Mohammad Saeed Rad, Behzad Bozorgtabar, Claudiu Cristian Musat, Urs-Viktor Marti, Max Basler, Hazim Kemal Ekenel, Jean-Philippe Thiran
Neurocomputing6
2019 G2-VER: Geometry Guided Model Ensemble for Video-based Facial Expression Recognition
abstract
This paper addresses the problem of automatic facial expression recognition in videos, where the goal is to predict discrete emotion labels best describing the emotions expressed in short video clips. Building on a pre-trained convolutional neural network (CNN) model dedicated to analyzing the video frames and LSTM network designed to process the trajectories of the facial landmarks, this paper investigates several novel directions. First of all, improved face descriptors based on 2D CNNs and facial landmarks are proposed. Second, the paper investigates fusion methods of the features temporally, including a novel hierarchical recurrent neural network combining facial landmark trajectories over time. In addition, we propose a modification to state-of-the-art expression recognition architectures to adapt them to video processing in a simple way. In both ensemble approaches, the temporal information is integrated. Comparative experiments on publicly available video-based facial expression recognition datasets verified that the proposed framework outperforms state-of-the-art methods. Moreover, we introduce a near-infrared video dataset containing facial expressions from subjects driving their cars, which are recorded in real world conditions.
Tanguy Albrici, Mandana Fasounaki, Saleh Bagher Salimi, Guillaume Vray, Behzad Bozorgtabar, Hazim Kemal Ekenel, Jean-Philippe Thiran
FG6
2019 Using Photorealistic Face Synthesis and Domain Adaptation to Improve Facial Expression Analysis
abstract
Cross-domain synthesizing realistic faces to learn deep models has attracted increasing attention for facial expression analysis as it helps to improve the performance of expression recognition accuracy despite having small number of real training images. However, learning from synthetic face images can be problematic due to the distribution discrepancy between low-quality synthetic images and real face images and may not achieve the desired performance when the learned model applies to real world scenarios. To this end, we propose a new attribute guided face image synthesis to perform a translation between multiple image domains using a single model. In addition, we adopt the proposed model to learn from synthetic faces by matching the feature distributions between different domains while preserving each domain's characteristics. We evaluate the effectiveness of the proposed approach on several face datasets on generating realistic face images. We demonstrate that the expression recognition performance can be enhanced by benefiting from our face synthesis model. Moreover, we also conduct experiments on a near-infrared dataset containing facial expression videos of drivers to assess the performance using in-the-wild data for driver emotion recognition.
Behzad Bozorgtabar, Mohammad Saeed Rad, Hazim Kemal Ekenel, Jean-Philippe Thiran
FG3
2019 SROBB: Targeted Perceptual Loss for Single Image Super-Resolution
abstract
By benefiting from perceptual losses, recent studies have improved significantly the performance of the superresolution task, where a high-resolution image is resolved from its low-resolution counterpart. Although such objective functions generate near-photorealistic results, their capability is limited, since they estimate the reconstruction error for an entire image in the same way, without considering any semantic information. In this paper, we propose a novel method to benefit from perceptual loss in a more objective way. We optimize a deep network-based decoder with a targeted objective function that penalizes images at different semantic levels using the corresponding terms. In particular, the proposed method leverages our proposed OBB (Object, Background and Boundary) labels, generated from segmentation labels, to estimate a suitable perceptual loss for boundaries, while considering texture similarity for backgrounds. We show that our proposed approach results in more realistic textures and sharper edges, and outperforms other state-of-the-art algorithms in terms of both qualitative results on standard benchmarks and results of extensive user studies.
Mohammad Saeed Rad, Behzad Bozorgtabar, Urs-Viktor Marti, Max Basler, Hazim Kemal Ekenel, Jean-Philippe Thiran
ICCV5
2019 Learn to synthesize and synthesize to learn
Behzad Bozorgtabar, Mohammad Saeed Rad, Hazim Kemal Ekenel, Jean-Philippe Thiran
Comput. Vis. Image Underst.3
2019 From 2D to 3D real-time expression transfer for facial animation
Beste Ekmen, Hazim Kemal Ekenel
Multim. Tools Appl.2
2019 Cross-dataset person re-identification using deep convolutional neural networks: effects of context and domain adaptation
Anil Genç, Hazim Kemal Ekenel
Multim. Tools Appl.2
2018 Vision-based game design and assessment for physical exercise in a robot-assisted rehabilitation system
abstract
Engagement is a key factor in gaming. Especially, in gamification applications, users’ engagement levels have to be assessed in order to determine the usability of the developed games. The authors first present computer vision‐based game design for physical exercise. All games are played with gesture controls. The authors conduct user studies in order to evaluate the perception of the games using a game engagement questionnaire. Participants state that the games are interesting and they want to play them again. Next, as a use case, the authors integrate one of these games into a robot‐assisted rehabilitation system. The authors perform additional user studies by employing self‐assessment manikin to assess the difficulty levels that can range from boredom to excitement. The authors observe that with the increasing difficulty level, users’ arousal increases. Additionally, the authors perform psychophysiological signal analysis of the participants during the execution of the game under two distinctive difficulty levels. The authors derive features from the signals obtained from blood volume pulse (BVP), skin conductance, and skin temperature sensors. As a result of analysis of variance and sequential forward selection, the authors find that changes in the temperature and frequency content of BVP provide useful information to estimate the players’ engagement.
Huseyin Erdogan, Yunus Palaska, Engin Masazade, Duygun Erol, Hazim Kemal Ekenel
IET Comput. Vis.5
2017 Combining LiDAR space clustering and convolutional neural networks for pedestrian detection
abstract
Pedestrian detection is an important component for safety of autonomous vehicles, as well as for traffic and street surveillance. There are extensive benchmarks on this topic and it has been shown to be a challenging problem when applied on real use-case scenarios. In purely image-based pedestrian detection approaches, the state-of-the-art results have been achieved with convolutional neural networks (CNN) and surprisingly few detection frameworks have been built upon multi-cue approaches. In this work, we develop a new pedestrian detector for autonomous vehicles that exploits LiDAR data, in addition to visual information. In the proposed approach, LiDAR data is utilized to generate region proposals by processing the three dimensional point cloud that it provides. These candidate regions are then further processed by a state-of-the-art CNN classifier that we have fine-tuned for pedestrian detection. We have extensively evaluated the proposed detection process on the KITTI dataset. The experimental results show that the proposed LiDAR space clustering approach provides a very efficient way of generating region proposals leading to higher recall rates and fewer misses for pedestrian detection. This indicates that LiDAR data can provide auxiliary information for CNN-based approaches.
Damien Matti, Hazim Kemal Ekenel, Jean-Philippe Thiran
AVSS2
2017 The unconstrained ear recognition challenge
abstract
In this paper we present the results of the Unconstrained Ear Recognition Challenge (UERC), a group benchmarking effort centered around the problem of person recognition from ear images captured in uncontrolled conditions. The goal of the challenge was to assess the performance of existing ear recognition techniques on a challenging large-scale dataset and identify open problems that need to be addressed in the future. Five groups from three continents participated in the challenge and contributed six ear recognition techniques for the evaluation, while multiple baselines were made available for the challenge by the UERC organizers. A comprehensive analysis was conducted with all participating approaches addressing essential research questions pertaining to the sensitivity of the technology to head rotation, flipping, gallery size, large-scale recognition and others. The top performer of the UERC was found to ensure robust performance on a smaller part of the dataset (with 180 subjects) regardless of image characteristics, but still exhibited a significant performance drop when the entire dataset comprising 3,704 subjects was used for testing.
Ziga Emersic, Dejan Stepec, Vitomir Struc, Peter Peer, Anjith George, Adil M. Ahmad, Elshibani Omar, Terrance E. Boult, Reza Safdari, Stefanos Zafeiriou, Doggucan Yaman, Fevziye Irem Eyiokur, Hazim Kemal Ekenel
IJCB14
2017 A Computer Vision System to Localize and Classify Wastes on the Streets
Mohammad Saeed Rad, Andreas von Kaenel, Andre Droux, François Tièche, Nabil Ouerhani, Hazim Kemal Ekenel, Jean-Philippe Thiran
ICVS6
2017 Face deidentification with generative deep neural networks
abstract
Face deidentification is an active topic amongst privacy and security researchers. Early deidentification methods relying on image blurring or pixelisation have been replaced in recent years with techniques based on formal anonymity models that provide privacy guaranties and retain certain characteristics of the data even after deidentification. The latter aspect is important, as it allows the deidentified data to be used in applications for which identity information is irrelevant. In this work, the authors present a novel face deidentification pipeline, which ensures anonymity by synthesising artificial surrogate faces using generative neural networks (GNNs). The generated faces are used to deidentify subjects in images or videos, while preserving non‐identity‐related aspects of the data and consequently enabling data utilisation. Since generative networks are highly adaptive and can utilise diverse parameters (pertaining to the appearance of the generated output in terms of facial expressions, gender, race etc.), they represent a natural choice for the problem of face deidentification. To demonstrate the feasibility of the authors’ approach, they perform experiments using automated recognition tools and human annotators. Their results show that the recognition performance on deidentified images is close to chance, suggesting that the deidentification process based on GNNs is effective.
Blaz Meden, Refik Can Malli, Sebastjan Fabijan, Hazim Kemal Ekenel, Vitomir Struc, Peter Peer
IET Signal Process.4
2016 Automatic emotion recognition in the wild using an ensemble of static and dynamic representations
abstract
Automatic emotion recognition in the wild video datasets is a very challenging problem because of the inter-class similarities among different facial expressions and large intra-class variabilities due to the significant changes in illumination, pose, scene, and expression. In this paper, we present our proposed method for video-based emotion recognition in the EmotiW 2016 challenge. The task considers the unconstrained emotion recognition problem by training from short video clips extracted from movies and testing on short movie clips and spontaneous video clips of the reality TV data. Four different methods are employed to extract both static and dynamic emotion representations from the videos. First, local binary patterns of three orthogonal planes are used to describe spatiotemporal features of the video frames. Second, principal component analysis is applied to the image patches in a two-stage convolutional network to learn weights and extract facial features from the aligned faces. Third, the deep convolutional neural network model of VGG-Face is deployed to extract deep facial representations from aligned faces. Fourth, a bag of visual words is computed based on dense scale-invariant feature transform descriptors from aligned face images to form hand-crafted representations. Support vector machines are then utilized to train and classify the obtained spatiotemporal representations and facial features. Finally, score-level fusion is applied to combine the classification results and predict the emotion labels of the video clips. The results show that the proposed combined method has outperformed all the utilized techniques with the overall validations and test accuracies of 43.13% and 40.13%, respectively. This system, is relatively a good classifier in Happy and Angry emotion categories and is unsuccessful in detecting Surprise, Disgust, and Fear.
Mostafa Mehdipour-Ghazi, Hazim Kemal Ekenel
ICMI2
2016 The CAMOMILE Collaborative Annotation Platform for Multi-modal, Multi-lingual and Multi-media Documents
Johann Poignant, Mateusz Budnik, Hervé Bredin, Claude Barras, Mickaël Stefas, Pierrick Bruneau, Gilles Adda, Laurent Besacier, Hazim Kemal Ekenel, Gil Francopoulo, Javier Hernando, Joseph Mariani, Ramon Morros, Georges Quénot, Sophie Rosset, Thomas Tamisier
LREC9
2015 Accio: A Data Set for Face Track Retrieval in Movies Across Age
abstract
Video face recognition is a very popular task and has come a long way. The primary challenges such as illumination, resolution and pose are well studied through multiple data sets. However there are no video-based data sets dedicated to study the effects of aging on facial appearance. We present a challenging face track data set, Harry Potter Movies Aging Data set (Accio1), to study and develop age invariant face recognition methods for videos. Our data set not only has strong challenges of pose, illumination and distractors, but also spans a period of ten years providing substantial variation in facial appearance. We propose two primary tasks: within and across movie face track retrieval; and two protocols which differ in their freedom to use external data. We present baseline results for the retrieval performance using a state-of-the-art face track descriptor. Our experiments show clear trends of reduction in performance as the age gap between the query and database increases. We will make the data set publicly available for further exploration in age-invariant video face recognition.
Esam Ghaleb, Makarand Tapaswi, Ziad Al-Halah, Hazim Kemal Ekenel, Rainer Stiefelhagen
ICMR4
2014 Automatic analysis of facial attractiveness from video
abstract
There has been a growing interest in the computer science field for automatic analysis and recognition of facial beauty and attractiveness. Most of the proposed studies attempt to model and predict facial attractiveness using a single static facial image. While a static image provides limited information about facial attractiveness, using a video clip that contains information about the motion and the dynamic behaviour of the face provides a richer understanding and valuable insights into analysing facial attractiveness. With this motivation, we propose to use dynamic features obtained from video clips along with static features obtained from static frames for automatic analysis of facial attractiveness. Support Vector Machine (SVM) and Random Forest (RF) are utilised to create and train models of attractiveness using the features extracted. Experimental results show that combining static and dynamic features improve performance over using either of these feature sets alone, and SVM provides the best prediction performance.
Sacide Kalayci, Hazim Kemal Ekenel, Hatice Gunes
ICIP2
2014 Cleaning up after a face tracker: False positive removal
abstract
Automatic person identification in TV series has gained popularity over the years. While most of the works rely on using face-based recognition, errors during tracking such as false positive face tracks are typically ignored. We propose a variety of methods to remove false positive face tracks and categorize the methods into confidence- and context-based. We evaluate our methods on a large TV series data set and show that up to 75% of the false positive face tracks are removed at the cost of 3.6% true positive tracks. We further show that the proposed method is general and applicable to other detectors or trackers.
Makarand Tapaswi, Cemal Cagn Corez, Martin Bäuml, Hazim Kemal Ekenel, Rainer Stiefelhagen
ICIP4
2014 Extending explicit shape regression with mixed feature channels and pose priors
abstract
Facial feature detection offers a wide range of applications, e.g. in facial image processing, human computer interaction, consumer electronics, and the entertainment industry. These applications impose two antagonistic key requirements: high processing speed and high detection accuracy. We address both by expanding upon the recently proposed explicit shape regression [1] to (a) allow usage and mixture of different feature channels, and (b) include head pose information to improve detection performance in non-cooperative environments. Using the publicly available “wild” datasets LFW [10] and AFLW [11], we show that using these extensions outperforms the baseline (up to 10% gain in accuracy at 8% IOD) as well as other state-of-the-art methods.
Matthias Richter 0003, Hua Gao, Hazim Kemal Ekenel
WACV3
2013 Multimodal genre classification of TV programs and YouTube videos
Hazim Kemal Ekenel, Tomas Semela
Multim. Tools Appl.1
2013 Combining texture and stereo disparity cues for real-time face detection
Feijun Jiang, Mika Fischer, Hazim Kemal Ekenel, Bertram E. Shi
Signal Process. Image Commun.3
2012 Face Alignment Using a Ranking Model based on Regression Trees
abstract
In this work, we exploit the regression trees-based ranking model, which has been successfully applied in the domain of web-search ranking, to build appearance models for face alignment. The model is an ensemble of regression trees which is learned with gradient boosting. The MCT (Modified Census Transform) as well as its unbinarized version PCT (Pseudo Census Transform) are used as features due to their robustness to illumination changes. To avoid the overfitting problem in gradient boosting, we use random trees to initialize the boosting. The Nelder Mead’s simplex method is applied for fitting the learned model. We compare the proposed regression trees-based pointwise ranking model to pairwise ranking model. Experiments show that the proposed model improves both robustness and accuracy for face alignment.
Hua Gao, Hazim Kemal Ekenel, Rainer Stiefelhagen
BMVC2
2012 A ranking model for face alignment with Pseudo Census Transform
Hua Gao, Hazim Kemal Ekenel, Rainer Stiefelhagen
ICPR2
2012 Multi-view facial expression recognition using local appearance features
Nikolas Hesse, Tobias Gehrig, Hua Gao, Hazim Kemal Ekenel
ICPR4
2012 Efficient and robust integration of face detection and head pose estimation
Feijun Jiang, Hazim Kemal Ekenel, Bertram E. Shi
ICPR2
2012 Facial expression classification on web images
Matthias Richter 0003, Tobias Gehrig, Hazim Kemal Ekenel
ICPR3
2011 Boosting Pseudo Census Transform Features for Face Alignment
abstract
Face alignment using deformable face model has attracted broad interest in recent years for its wide range of applications in facial analysis. Previous work has shown that discriminative deformable models have better generalization capacity compared to generative models [8, 9]. In this paper, we present a new discriminative face model based on boosting pseudo census transform features. This feature is considered to be less sensitive to illumination changes, which yields a more robust alignment algorithm. The alignment is based on maximizing the scores of boosted strong classifier, which indicate whether the current alignment is a correct or incorrect one. The proposed approach has been evaluated extensively on several databases. The experimental results show that our approach generalizes better on unseen data compared to the Haar feature-based approach. Moreover, its training procedure is much faster due to the low dimensionality of the configuration space of the proposed feature.
Hua Gao, Hazim Kemal Ekenel, Mika Fischer, Rainer Stiefelhagen
BMVC2
2011 Effective discretization of Gabor features for real-time face detection
abstract
We describe a real-time face detector based on Gabor features. While Gabor features often lead to improved performance, they are often avoided as they are perceived as being computationally expensive. We address this in two ways. First, we propose an efficient discrete encoding method for the Gabor feature vector. This enables us to use a computationally efficient multi-stage classifier based on boosting and winnowing. Second, we accelerate computationally complex computations using the parallelization provided by graphics processing units (GPUs). With these innovations, the resulting detector runs at 16.8 fps for 640 × 480 images on a PC equipped with an i5 CPU and a GTX 465 graphic card.
Feijun Jiang, Bertram E. Shi, Mika Fischer, Hazim Kemal Ekenel
ICIP4
2011 Person re-identification in TV series using robust face recognition and user feedback
Mika Fischer, Hazim Kemal Ekenel, Rainer Stiefelhagen
Multim. Tools Appl.2
2010 Multi-pose Face Recognition for Person Retrieval in Camera Networks
abstract
In this paper, we study the use of facial appearance features for the re-identification of persons using distributed camera networks in a realistic surveillance scenario. In contrast to features commonly used for person reidentification, such as whole body appearance, facial features offer the advantage of remaining stable over much larger intervals of time. The challenge in using faces for such applications, apart from low captured face resolutions, is that their appearance across camera sightings is largely influenced by lighting and viewing pose. Here, a number of techniques to address these problems are presented and evaluated on a database of surveillance-type recordings. A system for online capture and interactive retrieval is presented that allows to search for sightings of particular persons in the video database. Evaluation results are presented on surveillance data recorded with four cameras over several days. A mean average precision of 0.60 was achieved for inter-camera retrieval using just a single track as query set, and up to 0.86 after relevance feedback by an operator.
Martin Bäuml, Keni Bernardin, Mika Fischer, Hazim Kemal Ekenel, Rainer Stiefelhagen
AVSS4
2010 Automatic Frequency Band Selection for Illumination Robust Face Recognition
abstract
Varying illumination conditions cause a dramatic change in facial appearance that leads to a significant drop in face recognition algorithms' performance. In this paper, to overcome this problem, we utilize an automatic frequency band selection scheme. The proposed approach is incorporated to a local appearance-based face recognition algorithm, which employs discrete cosine transform (DCT) for processing local facial regions. From the extracted DCT coefficients, the approach determines to the ones that should be used for classification. Extensive experiments conducted on the extended Yale face database B have shown that benefiting from frequency information provides robust face recognition under changing illumination conditions.
Hazim Kemal Ekenel, Rainer Stiefelhagen
ICPR1
2010 Multi-resolution Local Appearance-Based Face Verification
abstract
Facial analysis based on local regions/blocks usually outperforms holistic approaches because it is less sensitive to local deformations and occlusions. Moreover, modeling local features enables us to avoid the problem of high dimensionality of feature space. In this paper, we model the local face blocks with Gabor features and project them into a discriminant identity space. The similarity score of a face pair is determined by fusion of the local classifiers. To acquire complementary information in different scales of face images, we integrate the local decisions from various image resolutions. The proposed multi-resolution block based face verification system is evaluated on the experiment 4 of Face Recognition Grand Challenge (FRGC) version 2.0. We obtained 92.5% verification [email protected]% FAR, which is the highest performance reported on this experiment so far in the literature.
Hua Gao, Hazim Kemal Ekenel, Mika Fischer, Rainer Stiefelhagen
ICPR2
2010 Multi-view Based Estimation of Human Upper-Body Orientation
abstract
The knowledge about the body orientation of humans can improve speed and performance of many service components of a smart-room. Since many of such components run in parallel, an estimator to acquire this knowledge needs a very low computational complexity. In this paper we address these two points with a fast and efficient algorithm using the smart-room's multiple camera output. The estimation is based on silhouette information only and is performed for each camera view separately. The single view results are fused within a Bayesian filter framework. We evaluate our system on a subset of videos from the CLEAR 2007 dataset and achieve an average correct classification rate of 87.8%, while the estimation itself just takes 12 ms when four cameras are used.
Lukas Rybok, Michael Voit, Hazim Kemal Ekenel, Rainer Stiefelhagen
ICPR3
2010 Interactive person-retrieval in TV series and distributed surveillance video
abstract
Tracking and identifying persons in videos are important building blocks in many applications. For browsing of multimedia data or interactive investigation of surveillance footage it is not even necessary to uniquely identify a person. Rather it often suffices to find occurrences of a person indicated by the user with an exemplary image sequence. We present two systems in which the search for a specific person can be initiated by a sample image sequence and then be further refined by interactive feedback by the operator. In the first system, episodes of TV series have been processed offline and can be searched for occurrences of the different characters. The second system tracks people online in multiple cameras and makes the sequences immediately searchable from a central station
Martin Bäuml, Mika Fischer, Keni Bernardin, Hazim Kemal Ekenel, Rainer Stiefelhagen
ACM Multimedia4
2010 A video-based door monitoring system using local appearance-based face models
Hazim Kemal Ekenel, Johannes Stallkamp, Rainer Stiefelhagen
Comput. Vis. Image Underst.1
2009 Open-Set Face Recognition-Based Visitor Interface System
Hazim Kemal Ekenel, Lorant Szasz-Toth, Rainer Stiefelhagen
ICVS1
2009 Multimodal identity tracking in a smart room
Keni Bernardin, Hazim Kemal Ekenel, Rainer Stiefelhagen
Pers. Ubiquitous Comput.2
2009 Multimodal identity tracking in a smart room
Keni Bernardin, Hazim Kemal Ekenel, Rainer Stiefelhagen
Pers. Ubiquitous Comput.2
2008 Face recognition for smart interactions
abstract
In this paper, face recognition systems that have been developed for smart interactions at the interACT Research Center is presented. The face recognition efforts at the interACT Research Center consist of development of a fast and robust face recognition algorithm and fully automatic face recognition systems that can be deployed for real-life smart interaction applications. The face recognition algorithm is based on appearances of local facial regions that are represented with discrete cosine transform coefficients. Many fully automatic face recognition systems have been developed based on this algorithm. Among these systems two of the portable ones will be shown as interactive demos. Moreover, demo videos will be shown for the other systems.
Hazim Kemal Ekenel, Mika Fischer, Hua Gao, Lorant Szasz-Toth, Rainer Stiefelhagen
FG1
2008 Tracking identities and attention in smart environments - contributions and progress in the CHIL project
abstract
To provide intelligent services in a smart environments it is necessary to acquire information about the room, the people in it and their interactions. This includes, for example, the number of people, their identities, locations, postures, body and head orientations, among others. This paper gives an overview of the perceptual technology evaluations that were conducted in the CHIL project, specifically those held in the CLEAR 2006 and 2007 evaluation workshops. We then summarize the main achievements and lessons learnt in the project in the areas of person tracking, person identification and head pose estimation, all of which are critical perception components in order to build perceptive smart environments.
Rainer Stiefelhagen, Keni Bernardin, Hazim Kemal Ekenel, Michael Voit
FG3
2007 Multi-modal Person Identification in a Smart Environment
abstract
In this paper, we present a detailed analysis of multimodal fusion for person identification in a smart environment. The multi-modal system consists of a video-based face recognition system and a speaker identification system. We investigated different score normalization, modality weighting and modality combination schemes during the fusion of the individual modalities. We introduced two new modality weighting schemes, namely, the cumulative ratio of correct matches (CRCM) and distance-to-second-closest (DT2ND) measures. In addition, we also assessed the effects of the well-known score normalization and classifier combination methods on the identification performance. Experimental results obtained on the CLEAR 2007 evaluation corpus, which contains audio-visual recordings from different smart rooms, show that CRCM-based modality weighting improves the correct identification rates significantly.
Hazim Kemal Ekenel, Mika Fischer, Qin Jin, Rainer Stiefelhagen
CVPR1
2007 Video-based Face Recognition on Real-World Data
abstract
In this paper, we present the classification sub-system of a real-time video-based face identification system which recognizes people entering through the door of a laboratory. Since the subjects are not asked to cooperate with the system but are allowed to behave naturally, this application scenario poses many challenges. Continuous, uncontrolled variations of facial appearance due to illumination, pose, expression, and occlusion need to be handled to allow for successful recognition. Faces are classified by a local appearance-based face recognition algorithm. The obtained confidence scores from each classification are progressively combined to provide the identity estimate of the entire sequence. We introduce three different measures to weight the contribution of each individual frame to the overall classification decision. They are distance- to-model (DTM), distance-to-second-closest (DT2ND), and their combination. Both a k-nearest neighbor approach and a set of Gaussian mixtures are evaluated to produce individual frame scores. We have conducted closed set and open set identification experiments on a database of 41 subjects. The experimental results show that the proposed system is able to reach high correct recognition rates in a difficult scenario.
Johannes Stallkamp, Hazim Kemal Ekenel, Rainer Stiefelhagen
ICCV2
2007 Face Recognition for Smart Interactions
abstract
In this paper an overview of face recognition research activities at the interACT Research Center is given. The face recognition efforts at the interACT Research Center consist of development of a fast and robust face recognition algorithm and fully automatic face recognition systems that can be deployed for real-life smart interaction applications. The face recognition algorithm is based on appearances of local facial regions that are represented with discrete cosine transform coefficients. Three fully automatic face recognition systems have been developed that are based on this algorithm. The first one is the "door monitoring system" that observes the entrance of a room and identifies the subjects while they are entering the room. The second one is the "portable face recognition system" that aims at environment-free face recognition and recognizes the user of a machine. The third system, "3D face recognition system", performs fully automatic face recognition on 3D range data.
Hazim Kemal Ekenel, Johannes Stallkamp, Hua Gao, Mika Fischer, Rainer Stiefelhagen
ICME1
2007 3-D Face Recognition Using Local Appearance-Based Models
abstract
In this paper, we present a local appearance-based approach for 3-D face recognition. In the proposed algorithm, we first register the 3-D point clouds to provide a dense correspondence between faces. Afterwards, we analyze two mapping techniques—the closest-point mapping and the ray-casting mapping, to construct depth images from the corresponding well-registered point clouds. The depth images that are obtained are then divided into local regions where the discrete cosine transformation is performed to extract local information. The local features are combined at the feature level for classification. Experimental results on the FRGC version 2.0 face database show that the proposed algorithm performs superior to the well-known face recognition algorithms.
Hazim Kemal Ekenel, Hua Gao, Rainer Stiefelhagen
IEEE Trans. Inf. Forensics Secur.1
2007 Enabling Multimodal Human-Robot Interaction for the Karlsruhe Humanoid Robot
abstract
In this paper, we present our work in building technologies for natural multimodal human-robot interaction. We present our systems for spontaneous speech recognition, multimodal dialogue processing, and visual perception of a user, which includes localization, tracking, and identification of the user, recognition of pointing gestures, as well as the recognition of a person's head orientation. Each of the components is described in the paper and experimental results are presented. We also present several experiments on multimodal human-robot interaction, such as interaction using speech and gestures, the automatic determination of the addressee during human-human-robot interaction, as well on interactive learning of dialogue strategies. The work and the components presented here constitute the core building blocks for audiovisual perception of humans and multimodal human-robot interaction used for the humanoid robot developed within the German research project (Sonderforschungsbereich) on humanoid cooperative robots.
Rainer Stiefelhagen, Hazim Kemal Ekenel, Christian Fügen, Petra Gieselmann, Hartwig Holzapfel, Florian Kraft, Kai Nickel, Michael Voit, Alex Waibel
IEEE Trans. Robotics2
2006 Audio-visual perception of a lecturer in a smart seminar room
Rainer Stiefelhagen, Keni Bernardin, Hazim Kemal Ekenel, John W. McDonough, Kai Nickel, Michael Voit, Matthias Wölfel
Signal Process.3
2005 The connector: facilitating context-aware communication
abstract
We present the Connector, a context-aware service that intelligently connects people. It maintains an awareness of its users' activities, preoccupations and social relationships to mediate a proper connection at the right time between them. In addition to providing users with important contextual cues about the availability of potential callees, the Connector adapts the behavior of the contactee's device automatically in order to avoid inappropriate interruptions.To acquire relevant context information, perceptual components analyze sensor input obtained from a smart mobile phone and --- if available --- from a variety of audio-visual sensors built into a smart meeting room environment. The Connector also uses any available multimodal interface (e.g. a speech interface to the smart phone, steerable camera-projector, targeted loudspeakers) in the smart meeting room, to deliver information to users in the most unobtrusive way possible.
Maria Danninger, G. Flaherty, Keni Bernardin, Hazim Kemal Ekenel, Thilo Köhler, Robert G. Malkin, Rainer Stiefelhagen, Alex Waibel
ICMI4
2005 Multiresolution face recognition
Hazim Kemal Ekenel, Bülent Sankur
Image Vis. Comput.1
2004 Feature selection in the independent component subspace for face recognition
Hazim Kemal Ekenel, Bülent Sankur
Pattern Recognit. Lett.1