EDBT 2026 Demo / reviewers in the wild / expert
Kamal Nasrollahi
dblp:02/4861
· DBLP profile ↗
51ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-1953-0429ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 24 · 2 first-author · 10 since 2021Security and privacy · 2Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedSFA: Federating Spikes Fired, Approximately
Alice Evelyn, Kamal Nasrollahi, Thomas B. Moeslund |
ICPR (12) | 2 |
| 2025 | Privacy Aware Human-Object Interaction in the Wild Novel Dataset
Luna Lux Fredenslund, Vasiliki Ismiroglou, Thomas B. Moeslund, Kamal Nasrollahi |
ACIVS | 4 |
| 2025 | Harborfront Anomaly DetectionabstractAbstract Creating high-quality datasets for the task of video anomaly detection is challenging due to a subjective anomaly definition and the rarity of anomalies, which oust the possibility of obtaining statistically significant data. This results in datasets where anomalies are placed in a single category, and are often considered less relevant from a security standpoint. Instead, we propose to create video anomaly datasets based on a framework utilizing object annotations to ease the annotation process and allow users to decide on the anomaly definition. Furthermore, this allows for a fine-grained evaluation w.r.t. anomaly types, which represents a novelty in the area of video anomaly detection. The framework is demonstrated using the existing thermal long-term drift (LTD) dataset, identifying and evaluating five different types of anomalies (appearance, motion, localization, density, and tampering) on six test sets. State-of-the-art anomaly detection methods are evaluated and found to underperform on the thermal anomaly detection dataset, which emphasizes a need for an adjustable anomaly definition in order to produce better anomaly datasets and models that generalize towards practical use. We share the code of the proposed framework to extract anomaly types along with object annotations for the LTD dataset at https://github.com/jagob/harborfront-vad . Jacob V. Dueholm, Mia Sandra Nicole Siemon, Radu Tudor Ionescu, Thomas B. Moeslund, Kamal Nasrollahi |
Neural Process. Lett. | 5 |
| 2024 | Visual Context-Aware Person Fall Detection
Aleksander Nagaj, Zenjie Li, Dim P. Papadopoulos, Kamal Nasrollahi |
KES-IDT | 4 |
| 2024 | PDA-RWSR: Pixel-Wise Degradation Adaptive Real-World Super-ResolutionabstractWhile many methods have been proposed to solve the Super-Resolution (SR) problem of Low-Resolution (LR) images with complex unknown degradations, their performance still drops significantly when evaluated on images with challenging real-world degradations. One often overlooked factor contributing to this, is the presence of spatially varying degradations in real LR images. To address this issue, we propose a novel degradation pipeline capable of generating paired LR/High-Resolution (HR) images with spatially varying noise, a key contributor to reduced image quality. Furthermore, to fully leverage such training data, we novelly propose a Pixel-Wise Degradation Adaptive Real-World Super-Resolution (PDA-RWSR) framework. Specifically, we design a new Restormer-based Real-World Super-Resolution (RWSR) model capable of adapting the reconstruction process based on pixel-wise degradation features extracted by a new supervised degradation estimation model. Along with our proposed method, we also introduce a new challenging real-world Spatially Variant Super-Resolution (SVSR) benchmarking dataset, where the images are degraded by complex noise of varying intensity and type, to evaluate the robustness of existing RWSR methods. Comprehensive experiments on synthetic and the proposed challenging real dataset demonstrates the superiority of our method over the current State-of-The-Art (SoTA). The SVSR dataset is available at https://doi.org/10.5281/zenodo.10044260. Andreas Aakerberg, Majed El Helou, Kamal Nasrollahi, Thomas B. Moeslund |
WACV | 3 |
| 2024 | CL-MAE: Curriculum-Learned Masked AutoencodersabstractMasked image modeling has been demonstrated as a powerful pretext task for generating robust representations that can be effectively generalized across multiple downstream tasks. Typically, this approach involves randomly masking patches (tokens) in input images, with the masking strategy remaining unchanged during training. In this paper, we propose a curriculum learning approach that updates the masking strategy to continually increase the complexity of the self-supervised reconstruction task. We conjecture that, by gradually increasing the task complexity, the model can learn more sophisticated and transferable representations. To facilitate this, we introduce a novel learnable masking module that possesses the capability to generate masks of different complexities, and integrate the proposed module into masked autoencoders (MAE). Our module is jointly trained with the MAE, while adjusting its behavior during training, transitioning from a partner to the MAE (optimizing the same reconstruction loss) to an adversary (optimizing the opposite loss), while passing through a neutral state. The transition between these behaviors is smooth, being regulated by a factor that is multiplied with the reconstruction loss of the masking module. The resulting training procedure generates an easy-to-hard curriculum. We train our Curriculum-Learned Masked Autoencoder (CL-MAE) on ImageNet and show that it exhibits superior representation learning capabilities compared to MAE. The empirical results on five downstream tasks confirm our conjecture, demonstrating that curriculum learning can be successfully used to self-supervise masked autoencoders. We release our code at https://github.com/ristea/cl-mae. Neelu Madan, Nicolae-Catalin Ristea, Kamal Nasrollahi, Thomas B. Moeslund, Radu Tudor Ionescu |
WACV | 3 |
| 2024 | Self-Supervised Masked Convolutional Transformer Block for Anomaly DetectionabstractAnomaly detection has recently gained increasing attention in the field of computer vision, likely due to its broad set of applications ranging from product fault detection on industrial production lines and impending event detection in video surveillance to finding lesions in medical scans. Regardless of the domain, anomaly detection is typically framed as a one-class classification task, where the learning is conducted on normal examples only. An entire family of successful anomaly detection methods is based on learning to reconstruct masked normal inputs (e.g. patches, future frames, etc.) and exerting the magnitude of the reconstruction error as an indicator for the abnormality level. Unlike other reconstruction-based methods, we present a novel self-supervised masked convolutional transformer block (SSMCTB) that comprises the reconstruction-based functionality at a core architectural level. The proposed self-supervised block is extremely flexible, enabling information masking at any layer of a neural network and being compatible with a wide range of neural architectures. In this work, we extend our previous self-supervised predictive convolutional attentive block (SSPCAB) with a 3D masked convolutional layer, a transformer for channel-wise attention, as well as a novel self-supervised objective based on Huber loss. Furthermore, we show that our block is applicable to a wider variety of tasks, adding anomaly detection in medical images and thermal videos to the previously considered tasks based on RGB images and surveillance videos. We exhibit the generality and flexibility of SSMCTB by integrating it into multiple state-of-the-art neural models for anomaly detection, bringing forth empirical results that confirm considerable performance improvements on five benchmarks: MVTec AD, BRATS, Avenue, ShanghaiTech, and Thermal Rare Event. Neelu Madan, Nicolae-Catalin Ristea, Radu Tudor Ionescu, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B. Moeslund, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | SSMTL++: Revisiting self-supervised multi-task learning for video anomaly detection
Antonio Barbalau, Radu Tudor Ionescu, Mariana-Iuliana Georgescu, Jacob V. Dueholm, Bharathkumar Ramachandra, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B. Moeslund, Mubarak Shah |
Comput. Vis. Image Underst. | 6 |
| 2023 | Video Transformers: A SurveyabstractTransformer models have shown great success handling long-range interactions, making them a promising tool for modeling video. However, they lack inductive biases and scale quadratically with input length. These limitations are further exacerbated when dealing with the high dimensionality introduced by the temporal dimension. While there are surveys analyzing the advances of Transformers for vision, none focus on an in-depth analysis of video-specific designs. In this survey, we analyze the main contributions and trends of works leveraging Transformers to model video. Specifically, we delve into how videos are handled at the input level first. Then, we study the architectural changes made to deal with video more efficiently, reduce redundancy, re-introduce useful inductive biases, and capture long-term temporal dynamics. In addition, we provide an overview of different training regimes and explore effective self-supervised learning strategies for video. Finally, we conduct a performance comparison on the most common benchmark for Video Transformers (i.e., action classification), finding them to outperform 3D ConvNets even with less computational complexity. Javier Selva, Anders Skaarup Johansen, Sergio Escalera, Kamal Nasrollahi, Thomas B. Moeslund, Albert Clapés |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Self-Supervised Predictive Convolutional Attentive Block for Anomaly DetectionabstractAnomaly detection is commonly pursued as a one-class classification problem, where models can only learn from normal training samples, while being evaluated on both normal and abnormal test samples. Among the successful approaches for anomaly detection, a distinguished category of methods relies on predicting masked information (e.g. patches, future frames, etc.) and leveraging the reconstruction error with respect to the masked information as an abnormality score. Different from related methods, we propose to integrate the reconstruction-based functionality into a novel self-supervised predictive architectural building block. The proposed self-supervised block is generic and can easily be incorporated into various state-of-the-art anomaly detection methods. Our block starts with a convolutional layer with dilated filters, where the center area of the receptive field is masked. The resulting activation maps are passed through a channel attention module. Our block is equipped with a loss that minimizes the reconstruction error with respect to the masked area in the receptive field. We demonstrate the generality of our block by integrating it into several state-of-the-art frameworks for anomaly detection on image and video, providing empirical evidence that shows considerable performance improvements on MVTec AD, Avenue, and ShanghaiTech. We release our code as open source at: https://github.com/ristea/sspcab. Nicolae-Catalin Ristea, Neelu Madan, Radu Tudor Ionescu, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B. Moeslund, Mubarak Shah |
CVPR | 4 |
| 2022 | A graph-based approach to video anomaly detection from the perspective of superpixelsabstractVideo Anomaly Detection refers to the concept of discovering activities in a video feed that deviate from the usual visible pattern. It is a very well-studied and explored field in the domain of Computer Vision and Deep Learning, in which automated learning-based systems are capable of detecting certain kinds of anomalies at an accuracy greater than 90%. Deep Learning based Artificial Neural Network models, however, suffer from very low interpretability. In order to address and design a possible solution for this issue, this work proposes to shape the given problem by means of graphical models. Given the high flexibility of compositing easily interpretable graphs, a great variety of techniques exist to build a model representing spatial as well as temporal relationships occurring in the given video sequence. The experiments conducted on common anomaly detection benchmark datasets show that significant performance gains can be achieved through simple re-modelling of individual graph components. In contrast to other video anomaly detection approaches, the one presented in this work focuses primarily on the exploration of the possibility to shift the way we currently look at and process videos when trying to detect anomalous events. Mia Sandra Nicole Siemon, Kamal Nasrollahi, Thomas B. Moeslund |
ICMV | 2 |
| 2022 | Real-world super-resolution of face-images from surveillance camerasabstractAbstract Most existing face image Super‐Resolution (SR) methods assume that the Low‐Resolution (LR) images were artificially downsampled from High‐Resolution (HR) images with bicubic interpolation. This operation changes the natural image characteristics and reduces noise. Hence, SR methods trained on such data most often fail to produce good results when applied to real LR images. To solve this problem, a novel framework for the generation of realistic LR/HR training pairs is proposed. The framework estimates realistic blur kernels, noise distributions, and JPEG compression artifacts to generate LR images with similar image characteristics as the ones in the source domain. This allows to train an SR model using high‐quality face images as Ground‐Truth (GT). For better perceptual quality, a Generative Adversarial Network (GAN) based SR model is used, where the commonly used VGG‐loss [1] is exchanged with LPIPS‐loss [2]. Experimental results on both real and artificially corrupted face images show that our method results in more detailed reconstructions with less noise compared to the existing State‐of‐the‐Art (SoTA) methods. In addition, it is shown that the traditional non‐reference Image Quality Assessment (IQA) methods fail to capture this improvement and demonstrate that the more recent NIMA metric [3] correlates better with human perception via Mean Opinion Rank (MOR). Andreas Aakerberg, Kamal Nasrollahi, Thomas B. Moeslund |
IET Image Process. | 2 |
| 2022 | Deep transfer learning in human-robot interaction for cognitive and physical rehabilitation purposes
Chaudhary Muhammad Aqdus Ilyas, Matthias Rehm, Kamal Nasrollahi, Yeganeh Madadi, Thomas B. Moeslund, Vahid Seydi |
Pattern Anal. Appl. | 3 |
| 2022 | Deep Pain: Exploiting Long Short-Term Memory Networks for Facial Expression ClassificationabstractPain is an unpleasant feeling that has been shown to be an important factor for the recovery of patients. Since this is costly in human resources and difficult to do objectively, there is the need for automatic systems to measure it. In this paper, contrary to current state-of-the-art techniques in pain assessment, which are based on facial features only, we suggest that the performance can be enhanced by feeding the raw frames to deep learning models, outperforming the latest state-of-the-art results while also directly facing the problem of imbalanced data. As a baseline, our approach first uses convolutional neural networks (CNNs) to learn facial features from VGG_Faces, which are then linked to a long short-term memory to exploit the temporal relation between video frames. We further compare the performances of using the so popular schema based on the canonically normalized appearance versus taking into account the whole image. As a result, we outperform current state-of-the-art area under the curve performance in the UNBC-McMaster Shoulder Pain Expression Archive Database. In addition, to evaluate the generalization properties of our proposed methodology on facial motion recognition, we also report competitive results in the Cohn Kanade+ facial expression database. Pau Rodríguez, Guillem Cucurull, Jordi Gonzàlez 0001, Josep M. Gonfaus, Kamal Nasrollahi, Thomas B. Moeslund, F. Xavier Roca |
IEEE Trans. Cybern. | 5 |
| 2022 | Effective fusion of deep multitasking representations for robust visual tracking
Seyed Mojtaba Marvasti-Zadeh, Hossein Ghanei-Yakhdan, Shohreh Kasaei, Kamal Nasrollahi, Thomas B. Moeslund |
Vis. Comput. | 4 |
| 2021 | Single-Loss Multi-task Learning For Improving Semantic Segmentation Using Super-Resolution
Andreas Aakerberg, Anders Skaarup Johansen, Kamal Nasrollahi, Thomas B. Moeslund |
CAIP (2) | 3 |
| 2021 | Object-Centric Anomaly Detection Using Memory Augmentation
Jacob V. Dueholm, Kamal Nasrollahi, Thomas B. Moeslund |
CAIP (1) | 2 |
| 2021 | Memory- and time-efficient dense network for single-image super-resolutionabstractAbstract Dense connections in convolutional neural networks (CNNs), which connect each layer to every other layer, can compensate for mid/high‐frequency information loss and further enhance high‐frequency signals. However, dense CNNs suffer from high memory usage due to the accumulation of concatenating feature‐maps stored in memory. To overcome this problem, a two‐step approach is proposed that learns the representative concatenating feature‐maps. Specifically, a convolutional layer with many more filters is used before concatenating layers to learn richer feature‐maps. Therefore, the irrelevant and redundant feature‐maps are discarded in the concatenating layers. The proposed method results in 24% and 6% less memory usage and test time, respectively, in comparison to single‐image super‐resolution (SISR) with the basic dense block. It also improves the peak signal‐to‐noise ratio by 0.24 dB. Moreover, the proposed method, while producing competitive results, decreases the number of filters in concatenating layers by at least a factor of 2 and reduces the memory consumption and test time by 40% and 12%, respectively. These results suggest that the proposed approach is a more practical method for SISR. Nasrin Imanpour, Ahmad Reza Naghsh-Nilchi, S. Amirhassan Monadjemi, Hossein Karshenas, Kamal Nasrollahi, Thomas B. Moeslund |
IET Signal Process. | 5 |
| 2020 | One-To-One Person Re-Identification For Queue Time EstimationabstractQueue time measurements in airport check-points are essential in order to regulate staff allocation and maintain a low queue time. In this paper, we propose a method to measure queue times based on person re-identification. Specifically, we capture passenger features from entrance and exit points of an airport check-point and match features to determine queue times of passengers. However, a passenger seen at the exit, potentially, can be matched to multiple passengers seen at the entrance. Therefore, to increase the precision of re-identification, we propose a simple, yet effective, algorithm that assigns each passenger seen at the exit to only a single passenger seen at the entrance. Through experiments on a dataset collected from an airport immigration check-point, we show that our proposed assignment increases precision and recall by 22 % and 16 %, respectively, to naively assigning the best match. Aske R. Lejbølle, Benjamin Krogh, Kamal Nasrollahi, Thomas B. Moeslund |
ICIP | 3 |
| 2020 | Deep visual unsupervised domain adaptation for classification tasks: a surveyabstractLearning methods are challenged when there is not enough labelled data. It gets worse when the existing learning data have different distributions in different domains. To deal with such situations, deep unsupervised domain adaptation techniques have newly been widely used. This study surveys such domain adaptation methods that have been used for classification tasks in computer vision. The survey includes the very recent papers on this topic that have not been included in the previous surveys and introduces a taxonomy by grouping methods published on unsupervised domain adaptation into five groups of discrepancy‐, adversarial‐, reconstruction‐, representation‐, and attention‐based methods. Yeganeh Madadi, Vahid Seydi, Kamal Nasrollahi, Reshad Hosseini, Thomas B. Moeslund |
IET Image Process. | 3 |
| 2020 | Cycle-consistent generative adversarial neural networks based low quality fingerprint enhancement
Dogus Karabulut, Pavlo Tertychnyi, Hasan Sait Arslan, Cagri Ozcinar, Kamal Nasrollahi, Joan Valls, Joan Vilaseca, Thomas B. Moeslund, Gholamreza Anbarjafari |
Multim. Tools Appl. | 5 |
| 2020 | A Double-Deep Spatio-Angular Learning Framework for Light Field-Based Face RecognitionabstractFace recognition has attracted increasing attention due to its wide range of applications, but it is still challenging when facing large variations in the biometric data characteristics. Lenslet light field cameras have recently come into prominence to capture rich spatio-angular information, thus offering new possibilities for advanced biometric recognition systems. This paper proposes a double-deep spatio-angular learning framework for light field-based face recognition, which is able to model both the intra-view/spatial and inter-view/angular information using two deep networks in sequence. This is a novel recognition framework that has never been proposed in the literature for face recognition or any other visual recognition task. The proposed double-deep learning framework includes a long short-term memory (LSTM) recurrent network, whose inputs are VGG-Face descriptions, computed using a VGG-16 convolutional neural network (CNN). The VGG-Face spatial descriptions are extracted from a selected set of 2D sub-aperture (SA) images rendered from the light field image, corresponding to different observation angles. A sequence of the VGG-Face spatial descriptions is then analyzed by the LSTM network. A comprehensive set of experiments has been conducted using the IST-EURECOM light field face database, addressing varied and challenging recognition tasks. The results show that the proposed framework achieves superior face recognition performance when compared to the state of the art. Alireza Sepas-Moghaddam, Mohammad A. Haque, Paulo Lobato Correia, Kamal Nasrollahi, Thomas B. Moeslund, Fernando Pereira 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Person Re-Identification Using Spatial and Layer-Wise AttentionabstractPerson re-identification requires extraction of discriminative features to ensure a correct match; this must be done independent of challenges, such as occlusion, view, or lighting changes. While occlusion can be eliminated by changing the camera setup from a horizontal to a vertical (overhead) viewpoint, other challenges arise as the total visible surface area of persons is decreased. As a result, methods that focus on the most discriminative regions of persons must be applied, while different domains should also be considered to extract different semantics. To further increase feature discriminability, complementary features extracted at different abstraction levels should be fused. To emphasize features at certain abstraction levels depending on the input, fusion should be done intelligently. This work considers multiple domains and feature discrimination, where a multimodal convolution neural network is applied to fuse RGB and depth information. To extract multi-local discriminative features, two different attention modules are proposed: (1) a spatial attention module, which is able to capture local information at different abstraction levels, and (2) a layer-wise attention module, which works as a dynamic weighting scheme to assign weights and fuse local abstraction-level features intelligently, depending on the input image. By fusing local and global features in a multimodal context, we show state-of-the-art accuracies on two publicly available datasets, DPI-T and TVPR, while increasing the state-of-the-art accuracy on a third dataset, OPR. Finally, through both visual and quantitative analysis, we show the ability of the proposed system to leverage multiple frames, by adapting feature weighting depending on the input. Aske R. Lejbølle, Kamal Nasrollahi, Benjamin Krogh, Thomas B. Moeslund |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | Teaching Pepper Robot to Recognize Emotions of Traumatic Brain Injured Patients Using Deep Neural NetworksabstractSocial signal extraction from the facial analysis is a popular research area in human-robot interaction. However, recognition of emotional signals from Traumatic Brain Injured (TBI) patients with the help of robots and non-intrusive sensors is yet to be explored. Existing robots have limited abilities to automatically identify human emotions and respond accordingly. Their interaction with TBI patients could be even more challenging and complex due to unique, unusual and diverse ways of expressing their emotions. To tackle the disparity in a TBI patient's Facial Expressions (FEs), a specialized deep-trained model for automatic detection of TBI patients' emotions and FE (TBI-FER model) is designed, for robot-assisted rehabilitation activities. In addition, the Pepper robot's built-in model for FE is investigated on TBI patients as well as on healthy people. Variance in their emotional expressions is determined by comparative studies. It is observed that the customized trained system is highly essential for the deployment of Pepper robot as a Socially Assistive Robot (SAR). Chaudhary Muhammad Aqdus Ilyas, Viktor Schmuck, Mohammad A. Haque, Kamal Nasrollahi, Matthias Rehm, Thomas B. Moeslund |
RO-MAN | 4 |
| 2019 | A novel deep network architecture for reconstructing RGB facial images from thermal for face recognition
Andre Litvin, Kamal Nasrollahi, Sergio Escalera, Cagri Ozcinar, Thomas B. Moeslund, Gholamreza Anbarjafari |
Multim. Tools Appl. | 2 |
| 2019 | Guest editorial: special issue on human abnormal behavioural analysis
Gholamreza Anbarjafari, Sergio Escalera, Kamal Nasrollahi, Hugo Jair Escalante, Xavier Baró, Jun Wan 0001, Thomas B. Moeslund |
Mach. Vis. Appl. | 3 |
| 2018 | Changes in Facial Expression as Biometric: A Database and Benchmarks of IdentificationabstractFacial dynamics can be considered as unique signatures for discrimination between people. These have started to become important topic since many devices have the possibility of unlocking using face recognition or verification. In this work, we evaluate the efficacy of the transition frames of video in emotion as compared to the peak emotion frames for identification. For experiments with transition frames we extract features from each frame of the video from a fine-tuned VGG-Face Convolutional Neural Network (CNN) and geometric features from facial landmark points. To model the temporal context of the transition frames we train a Long-Short Term Memory (LSTM) on the geometric and the CNN features. Furthermore, we employ two fusion strategies: first, an early fusion, in which the geometric and the CNN features are stacked and fed to the LSTM. Second, a late fusion, in which the prediction of the LSTMs, trained independently on the two features, are stacked and used with a Support Vector Machine (SVM). Experimental results show that the late fusion strategy gives the best results and the transition frames give better identification results as compared to the peak emotion frames. Rain Eric Haamer, Kaustubh Kulkarni, Nasrin Imanpour, Mohammad A. Haque, Egils Avots, Michelle Breisch, Kamal Nasrollahi, Sergio Escalera, Cagri Ozcinar, Xavier Baró, Ahmad Reza Naghsh-Nilchi, Thomas B. Moeslund, Gholamreza Anbarjafari |
FG | 7 |
| 2018 | Deep Multimodal Pain Recognition: A Database and Comparison of Spatio-Temporal Visual ModalitiesabstractPain is a symptom of many disorders associated with actual or potential tissue damage in human body. Managing pain is not only a duty but also highly cost prone. The most primitive state of pain management is the assessment of pain. Traditionally it was accomplished by self-report or visual inspection by experts. However, automatic pain assessment systems from facial videos are also rapidly evolving due to the need of managing pain in a robust and cost effective way. Among different challenges of automatic pain assessment from facial video data two issues are increasingly prevalent: first, exploiting both spatial and temporal information of the face to assess pain level, and second, incorporating multiple visual modalities to capture complementary face information related to pain. Most works in the literature focus on merely exploiting spatial information on chromatic (RGB) video data on shallow learning scenarios. However, employing deep learning techniques for spatio-temporal analysis considering Depth (D) and Thermal (T) along with RGB has high potential in this area. In this paper, we present the first state-of-the-art publicly available database, 'Multimodal Intensity Pain (MIntPAIN)' database, for RGBDT pain level recognition in sequences. We provide a first baseline results including 5 pain levels recognition by analyzing independent visual modalities and their fusion with CNN and LSTM models. From the experimental evaluation we observe that fusion of modalities helps to enhance recognition performance of pain levels in comparison to isolated ones. In particular, the combination of RGB, D, and T in an early fusion fashion achieved the best recognition rate. Mohammad A. Haque, Rubén Ballester, Fatemeh Noroozi, Kaustubh Kulkarni, Christian B. Laursen, Ramin Irani, Marco Bellantonio, Sergio Escalera, Gholamreza Anbarjafari, Kamal Nasrollahi, Ole Kæseler Andersen, Erika G. Spaich, Thomas B. Moeslund |
FG | 10 |
| 2018 | Rehabilitation of Traumatic Brain Injured Patients: Patient Mood Analysis from Multimodal VideoabstractRehabilitation after traumatic brain injury (TBI) is very critical as it is largely unpredictable depending upon the nature of the injury. Rehabilitation process and recovery time also varies, as it takes months and years, depending upon the assessment of treatment, mental and physical conditions and strategies. Due to non-cooperative behaviour of patients, and increase in negative emotional expressions it is very beneficial to evaluate these expressions in a contactless way, and perform a rehabilitation physiotherapy, cognitive or other behavioral activities when the patient is in a positive mood. In this paper we have analyzed the methods for facial features extraction for TBI patients to determine optimal time to have aforementioned rehabilitation process on the basis of positive and negative facial expressions. We have employed a deep learning architecture based on convolutional neural network and long short term memory on RGB and thermal data that were collected in challenging scenarios from real patients. It automatically identifies the patient's facial expressions, and inform experts or trainers that “it is the time” to start rehabilitation session. Chaudhary Muhammad Aqdus Ilyas, Kamal Nasrollahi, Matthias Rehm, Thomas B. Moeslund |
ICIP | 2 |
| 2018 | Occlusion-aware pedestrian detectionabstractFailure in pedestrian detection systems can be extremely crucial, specifically in driverless driving. In this paper, failures in pedestrian detectors are refined by re-evaluating the results of state of the art pedestrian detection systems, via a fully convolutional neural network. The network is trained on a number of datasets which include a custom designed occluded pedestrian dataset to address the problem of occlusion. Results show that when applying the proposed network, detectors can not only maintain their state of the art performance, but they even decrease average false positives rate per image, especially in the case where pedestrians are occluded. Christos Apostolopoulos, Kamal Nasrollahi, M. Hsuan Yang, Mohammad Naser Sabet Jahromi, Thomas B. Moeslund |
ICMV | 2 |
| 2018 | Multimodal heartbeat rate estimation from the fusion of facial RGB and thermal videosabstractMeasuring Heartbeat Rate (HR) is an important tool for monitoring the health of a person. When the heart beats the influx of blood to the head causes slight involuntary movement and subtle skin color changes, which cannot be seen by the naked eye but can be tracked from facial videos using computer vision techniques and can be analyzed to estimate the HR. However, the current state of the art solutions encounter an increasing amount of complications when the subject has voluntary motion on the face or when the lighting conditions change in the video. Thus the accuracy of the HR estimation using computer vision is still inferior to that of a physical Electrocardiography (ECG) based system. The aim of this work is to improve the current non-invasive HR measurement by fusing the motion-based and color-based HR estimation methods and using them on multiple input modalities, e.g., RGB and thermal imaging. Our experiments indicate that late-fusion of the results of these methods (motion and color-based) applied to these different modalities, produces more accurate results compared to the existing solutions Anders Skaarup Johansen, Jesper W. Henriksen, Mohammad A. Haque, Mohammad Naser Sabet Jahromi, Kamal Nasrollahi, Thomas B. Moeslund |
ICMV | 5 |
| 2017 | Depth Value Pre-Processing for Accurate Transfer Learning based RGB-D Object RecognitionabstractObject recognition is one of the important tasks in computer vision which has found enormous applications.Depth modality is proven to provide supplementary information to the common RGB modality for objectrecognition. In this paper, we propose methods to improve the recognition performance of an existing deeplearning based RGB-D object recognition model, namely the FusionNet proposed by Eitel et al. First, we showthat encoding the depth values as colorized surface normals is beneficial, when the model is initialized withweights learned from training on ImageNet data. Additionally, we show that the RGB stream of the FusionNetmodel can benefit from using deeper network architectures, namely the 16-layered VGGNet, in exchange forthe 8-layered CaffeNet. In combination, these changes improves the recognition performance with 2.2% incomparison to the original FusionNet, when evaluating on the Washington RGB-D Object Dataset. Andreas Aakerberg, Kamal Nasrollahi, Christoffer B. Rasmussen, Thomas B. Moeslund |
IJCCI | 2 |
| 2017 | Real-Time Barcode Detection and Classification using Deep LearningabstractBarcodes, in their different forms, can be found on almost any packages available in the market. Detecting and then decoding of barcodes have therefore great applications. We describe how to adapt the state-of-the- art deep learning-based detector of You Only Look Once (YOLO) for the purpose of detecting barcodes in a fast and reliable way. The detector is capable of detecting both 1D and QR barcodes. The detector achieves state-of-the-art results on the benchmark dataset of Muenster BarcodeDB with a detection rate of 0.991. The developed system can also find the rotation of both the 1D and QR barcodes, which gives the opportunity of rotating the detection accordingly which is shown to benefit the decoding process in a positive way. Both the detection and the rotation prediction shows real-time performance. Daniel Kold Hansen, Kamal Nasrollahi, Christoffer B. Rasmussen, Thomas B. Moeslund |
IJCCI | 2 |
| 2017 | R-FCN Object Detection Ensemble based on Object Resolution and Image QualityabstractObject detection can be difficult due to challenges such as variations in objects both inter- and intra-class. Additionally, variations can also be present between images. Based on this, research was conducted into creating an ensemble of Region-based Fully Convolutional Networks (R-FCN) object detectors. Ensemble strategies explored were firstly data sampling and selection and secondly combination strategies. Data sampling and selection aimed to create different subsets of data with respect to object size and image quality such that expert R-FCN ensemble members could be trained. Two combination strategies were explored for combining the individual member detections into an ensemble result, namely average and a weighted average. R-FCNs were trained and tested on the PASCAL VOC benchmark object detection dataset. Results proved positive with an increase in Average Precision (AP), compared to state-of-the-art similar systems, when ensemble members were combined appropriately. Christoffer B. Rasmussen, Kamal Nasrollahi, Thomas B. Moeslund |
IJCCI | 2 |
| 2017 | A new low-complexity patch-based image super-resolutionabstractIn this study, a novel single image super‐resolution (SR) method, which uses a generated dictionary from pairs of high‐resolution (HR) images and their corresponding low‐resolution (LR) representations, is proposed. First, HR and LR dictionaries are created by dividing HR and LR images into patches Afterwards, when performing SR, the distance between every patch of the input LR image and those of available LR patches in the LR dictionary are calculated. The minimum distance between the input LR patch and those in the LR dictionary is taken, and its counterpart from the HR dictionary will be passed through an illumination enhancement process resulting in consistency of illumination between neighbour patches. This process is applied to all patches of the LR image. Finally, in order to remove the blocking effect caused by merging the patches, an average of the obtained HR image and the interpolated image is calculated. Furthermore, it is shown that the stabe of dictionaries is reducible to a great degree. The speed of the system is improved by 62.5%. The quantitative and qualitative analyses of the experimental results show the superiority of the proposed technique over the conventional and state‐of‐the‐art methods. Pejman Rasti, Kamal Nasrollahi, Olga Orlova, Gert Tamberg, Cagri Ozcinar, Thomas B. Moeslund, Gholamreza Anbarjafari |
IET Comput. Vis. | 2 |
| 2016 | Facial video-based detection of physical fatigue for maximal muscle activityabstractPhysical fatigue reveals the health condition of a person at, for example, health checkup, fitness assessment, or rehabilitation training. This study presents an efficient non‐contact system for detecting non‐localised physical fatigue from maximal muscle activity using facial videos acquired in a realistic environment with natural lighting where subjects were allowed to voluntarily move their head, change their facial expression, and vary their pose. The proposed method utilises a facial feature point tracking method by combining a ‘good feature to track’ and a ‘supervised descent method’ to address the challenges that originate from realistic scenario. A face quality assessment system was also incorporated in the proposed system to reduce erroneous results by discarding low quality faces that occurred in a video sequence due to problems in realistic lighting, head motion, and pose variation. Experimental results show that the proposed system outperforms video‐based existing system for physical fatigue detection. Mohammad A. Haque, Ramin Irani, Kamal Nasrollahi, Thomas B. Moeslund |
IET Comput. Vis. | 3 |
| 2016 | Complementary Cohort Strategy for Multimodal Face Pair MatchingabstractFace pair matching is the task of determining whether two face images represent the same person. Due to the limited expressive information embedded in the two face images as well as various sources of facial variations, it becomes a quite difficult problem. Toward the issue of few available images provided to represent each face, we propose to exploit an extra cohort set (identities in the cohort set are different from those being compared) by a series of cohort list comparisons. Useful cohort coefficients are then extracted from both sorted cohort identities and sorted cohort images for complementary information. To augment its robustness to complicated facial variations, we further employ multiple face modalities owing to their complementary value to each other for the face pair matching task. The final decision is made by fusing the extracted cohort coefficients with the direct matching score for all the available face modalities. To investigate the capacity of each individual modality on matching faces, the cohort behavior, and the performance achieved using our complementary cohort strategy, we conduct a set of experiments on two recently collected multimodal face databases. It is shown that using different modalities leads to different face pair matching performance. For each modality, employing our cohort scheme significantly reduces the equal error rate. By applying the proposed multimodal complementary cohort strategy, we achieve the best performance on our face pair matching task. Yunlian Sun, Kamal Nasrollahi, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2015 | EREL: Extremal regions of extremum levelsabstractExtremal Regions of Extremum Levels (EREL) are regions detected from a set of all extremal regions of an image. Maximally Stable Extremal Regions (MSER) which is a novel affine covariant region detector, detects regions from a same set of extremal regions as well. Although MSER results in regions with almost high repeatability, it is heavily dependent on the union-find approach which is a fairly complicated algorithm, and should be completed sequentially. Furthermore, it detects regions with low repeatability under the blur transformations. The reason for the latter shortcoming is the absence of boundaries information in stability criterion. To tackle these problems we propose to employ prior information about boundaries of regions, which results in a novel region detector algorithm that not only outperforms MSER, but avoids the MSER's rather complicated steps of union-finding. To achieve that, we introduce Maxima of Gradient Magnitudes (MGMs) and use them to find handful of Extremum Levels (ELs). The chosen ELs are then scanned to detect their Extremal Regions (ER). The proposed algorithm which is called Extremal Regions of Extremum Levels (EREL) has been tested on the public benchmark dataset of Mikolajczyk [1]. Our experimental evaluations illustrate that, in many cases EREL achieves higher repeatability scores than MSER even for very low overlap errors. Mehdi Faraji, Jamshid Shanbehzadeh, Kamal Nasrollahi, Thomas B. Moeslund |
ICIP | 3 |
| 2015 | Circular Hough Transform and Local Circularity Measure for Weight Estimation of a Graph-Cut Based Wood Stack MeasurementabstractOne of the time consuming tasks in the timber industry is the manually measurement of features of wood stacks. Such features include, but are not limited to, the number of the logs in a stack, their diameters distribution, and their volumes. Computer vision techniques have recently been used for solving this real-world industrial application. Such techniques are facing many challenges as the task is usually performed in outdoor, uncontrolled, environments. Furthermore, the logs can vary in texture and they can be occluded by different obstacles. These all make the segmentation of the wood logs a difficult task. Graph-cut has shown to be good enough for such a segmentation. However, it is hard to find proper graph weights. This is exactly the contribution of this paper to propose a method for setting the weights of the graph. To do so, we use Circular Hough Transform (CHT) for obtaining information about the fore and background regions of a stack image, and then use this together with a Local Circularity Measure (LCM) to modify the weights of the graph to segment the wood logs from the rest of the image. We further improve the segmentation by separating overlapping logs. These segmented wood logs are finally scaled and used to acquire the necessary wood stack measurements in real-world scale (in cm). The proposed system, which works automatically, has been tested on two different datasets, containing real outdoor images of logs which vary in shapes and sizes. The experimental results show that the proposed approach not only achieves the same results as the state-of-the-art systems, it produces more stable results. Bo Galsgaard, Dennis H. Lundtoft, Ivan A. Nikolov, Kamal Nasrollahi, Thomas B. Moeslund |
WACV | 4 |
| 2015 | Quality-Aware Estimation of Facial Landmarks in Video SequencesabstractFace alignment in video is a primitive step for facial image analysis. The accuracy of the alignment greatly depends on the quality of the face image in the video frames and low quality faces are proven to cause erroneous alignment. Thus, this paper proposes a system for quality aware face alignment by using a Supervised Decent Method (SDM) along with a motion based forward extrapolation method. The proposed system first extracts faces from video frames. Then, it employs a face quality assessment technique to measure the face quality. If the face quality is high, the proposed system uses SDM for facial landmark detection. If the face quality is low the proposed system corrects the facial landmarks that are detected by SDM. Depending upon the face velocity in consecutive video frames and face quality measure, two algorithms are proposed for correction of landmarks in low quality faces by using an extrapolation polynomial. Experimental results illustrate the competency of the proposed method while comparing with the state-of-the art methods including an SDM-based method (from CVPR-2013) and a very recent method (from CVPR-2014) that uses parallel cascade of linear regression (Par-CLR). Mohammad A. Haque, Kamal Nasrollahi, Thomas B. Moeslund |
WACV | 2 |
| 2015 | On soft biometrics
Mark S. Nixon, Paulo Lobato Correia, Kamal Nasrollahi, Thomas B. Moeslund, Abdenour Hadid, Massimo Tistarelli |
Pattern Recognit. Lett. | 3 |
| 2015 | Extremal Regions Detection Guided by Maxima of Gradient MagnitudeabstractA problem of computer vision applications is to detect regions of interest under different imaging conditions. The state-of-the-art maximally stable extremal regions (MSERs) detects affine covariant regions by applying all possible thresholds on the input image, and through three main steps including: (1) making a component tree of extremal regions' evolution; (2) obtaining region stability criterion; and (3) cleaning up. The MSER performs very well, but, it does not consider any information about the boundaries of the regions, which are important for detecting repeatable extremal regions. We have shown in this paper that employing prior information about boundaries of regions results in a novel region detector algorithm that not only outperforms MSER, but avoids the MSER's rather complicated steps of enumeration and the cleaning up. To employ the information about the region boundaries, we introduce maxima of gradient magnitudes (MGMs) which are shown to be points that are mostly around the boundaries of the regions. Having found the MGMs, the method obtains a global criterion for each level of the input image which is used to find extremum levels (ELs). The found ELs are then used to detect extremal regions. The proposed algorithm which is called extremal regions of extremum levels (EREL) has been tested on the public benchmark data set of Mikolajczyk. The obtained experimental results show that the inclusion of region boundaries through MGMs, results in a detector that detects regions with high repeatability scores and is more robust against noise compared with MSER. Mehdi Faraji, Jamshid Shanbehzadeh, Kamal Nasrollahi, Thomas B. Moeslund |
IEEE Trans. Image Process. | 3 |
| 2014 | Contactless measurement of muscles fatigue by tracking facial feature points in a videoabstractPhysical exercise may result in muscle tiredness which is known as muscle fatigue. This occurs when the muscles cannot exert normal force, or when more than normal effort is required. Fatigue is a vital sign, for example, for therapists to assess their patient's progress or to change their exercises when the level of the fatigue might be dangerous for the patients. The current technology for measuring tiredness, like Electromyography (EMG), requires installing some sensors on the body. In some applications, like remote patient monitoring, this however might not be possible. To deal with such cases, in this paper we present a contactless method based on computer vision techniques to measure tiredness by detecting, tracking, and analyzing some facial feature points during the exercise. Experimental results on several test subjects and comparing them against ground truth data show that the proposed system can properly find the temporal point of tiredness of the muscles when the test subjects are doing physical exercises. Ramin Irani, Kamal Nasrollahi, Thomas B. Moeslund |
ICIP | 2 |
| 2014 | Stereoscopic roadside curb height measurement using V-disparityabstractManaging road assets, such as roadside curbs, is one of the interests of municipalities. As an interesting application of computer vision, this paper proposes a system for automated measurement of the height of the roadside curbs. The developed system uses the spatial information available in the disparity image obtained from a stereo setup. Data about the geometry of the scene is extracted in the form of a row-wise histogram of the disparity map. From parameterizing the two strongest lines, each pixel can be labeled as belonging to one plane, either ground, sidewalk or curb candidates. Experimental results show that the system can measure the height of the roadside curb with good accuracy and precision. Florin Octavian Matu, Iskren Vlaykov, Mikkel Thøgersen, Kamal Nasrollahi, Thomas B. Moeslund |
ICMV | 4 |
| 2014 | RGB-D-T Based Face RecognitionabstractFacial images are of critical importance in many real-world applications from gaming to surveillance. The current literature on facial image analysis, from face detection to face and facial expression recognition, are mainly performed in either RGB, Depth (D), or both of these modalities. But, such analyzes have rarely included Thermal (T) modality. This paper paves the way for performing such facial analyzes using synchronized RGB-D-T facial images by introducing a database of 51 persons including facial images of different rotations, illuminations, and expressions. Furthermore, a face recognition algorithm has been developed to use these images. The experimental results show that face recognition using such three modalities provides better results compared to face recognition in any of such modalities in most of the cases. Olegs Nikisins, Kamal Nasrollahi, Modris Greitans, Thomas B. Moeslund |
ICPR | 2 |
| 2014 | Adaptive Non-local Means for Cost Aggregation in a Local Disparity Estimation AlgorithmabstractThe overall method used for determining disparity in a stereo setup is a widely recognized framework consisting of four steps of cost space computation, cost aggregation, disparity selection, and post-processing. In this paper a cost aggregation approach for a typical local disparity estimation method is introduced. The method introduced is built on top of an existing method called Adaptive Support-Weight using this known framework. The introduced method improves Adaptive Support-Weight method by utilizing a larger amount of data inspired by the method of Non-Local Means. The extra data is handled in a way that tries to preserve the location of depth discontinuities in the final disparity map. Experimental results on Middlebury benchmark database show that the proposed method suffers from less artifacts compared to state-of-the-art disparity estimation methods. Casper Pedersen, Kamal Nasrollahi, Thomas B. Moeslund |
ICPR | 2 |
| 2014 | Super-resolution: a comprehensive survey
Kamal Nasrollahi, Thomas B. Moeslund |
Mach. Vis. Appl. | 1 |
| 2013 | Real-time acquisition of high quality face sequences from an active pan-tilt-zoom cameraabstractTraditional still camera-based facial image acquisition systems in surveillance applications produce low quality face images. This is mainly due to the distance between the camera and subjects of interest. Furthermore, people in such videos usually move around, change their head poses, and facial expressions. Moreover, the imaging conditions like illumination, occlusion, and noise may change. These all aggregate the quality of most of the detected face images in terms of measures like resolution, pose, brightness, and sharpness. To deal with these problems this paper presents an active camera-based realtime high-quality face image acquisition system, which utilizes pan-tilt-zoom parameters of a camera to focus on a human face in a scene and employs a face quality assessment method to log the best quality faces from the captured frames. The system consists of four modules: face detection, camera control, face tracking, and face quality assessment before logging. Experimental results show that the proposed system can effectively log the high quality faces from the active camera in real-time (an average of 61.74ms was spent per frame) with an accuracy of 85.27% compared to human annotated data. Mohammad A. Haque, Kamal Nasrollahi, Thomas B. Moeslund |
AVSS | 2 |
| 2013 | Are Haar-Like Rectangular Features for Biometric Recognition Reducible?
Kamal Nasrollahi, Thomas B. Moeslund |
CIARP (2) | 1 |
| 2013 | Haar-like features for robust real-time face recognitionabstractFace recognition is still a very challenging task when the input face image is noisy, occluded by some obstacles, of very low-resolution, not facing the camera, and not properly illuminated. These problems make the feature extraction and consequently the face recognition system unstable. The proposed system in this paper introduces the novel idea of using Haar-like features, which have commonly been used for object detection, along with a probabilistic classifier for face recognition. The proposed system is simple, real-time, effective and robust against most of the mentioned problems. Experimental results on public databases show that the proposed system indeed outperforms the state-of-the-art face recognition systems. Kamal Nasrollahi, Thomas B. Moeslund |
ICIP | 1 |
| 2011 | Extracting a Good Quality Frontal Face Image From a Low-Resolution Video SequenceabstractFeeding low-resolution and low-quality images, from inexpensive surveillance cameras, to systems like, e.g., face recognition, produces erroneous and unstable results. Therefore, there is a need for a mechanism to bridge the gap between on one hand low-resolution and low-quality images and on the other hand facial analysis systems. The proposed system in this paper deals with exactly this problem. Our approach is to apply a reconstruction-based super-resolution algorithm. Such an algorithm, however, has two main problems: first, it requires relatively similar images with not too much noise and second is that its improvement factor is limited by a factor close to two. To deal with the first problem we introduce a three-step approach, which produces a face-log containing images of similar frontal faces of the highest possible quality. To deal with the second problem, limited improvement factor, we use a learning-based super-resolution algorithm applied to the result of the reconstruction-based part to improve the quality by another factor of two. This results in an improvement factor of four for the entire system. The proposed system has been tested on 122 low-resolution sequences from two different databases. The experimental results show that the proposed system can indeed produce a high-resolution and good quality frontal face image from low-resolution video sequences. Kamal Nasrollahi, Thomas B. Moeslund |
IEEE Trans. Circuits Syst. Video Technol. | 1 |