VLDB 2026 Research / reviewers in the wild / expert
Thomas B. Moeslund
dblp:12/6781 · also Thomas Baltzer Moeslund
· DBLP profile ↗
128ranked-venue papers
8as first author
33since 2021 · last 2026
0000-0001-7584-5209ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 75 · 3 first-author · 21 since 2021Artificial intelligence and machine learning · 68 · 6 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 since 2021Human-computer interaction and ubiquitous computing · 9 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedSFA: Federating Spikes Fired, Approximately
Alice Evelyn, Kamal Nasrollahi, Thomas B. Moeslund |
ICPR (12) | 3 |
| 2025 | Privacy Aware Human-Object Interaction in the Wild Novel Dataset
Luna Lux Fredenslund, Vasiliki Ismiroglou, Thomas B. Moeslund, Kamal Nasrollahi |
ACIVS | 3 |
| 2025 | The Challenge of Computing Responsible AI
Thomas B. Moeslund |
ICPRAM | 1 |
| 2025 | 8th ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'25)abstractThe 8th ACM International Workshop on Multimedia Content Analysis in Sports is held in Dublin, Ireland on October 28th, 2025. It is co-located with ACM Multimedia 2025. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing the multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation as well as understanding, statistical analysis, and evaluation in amateur and professional sports. There is a lack of research communities focusing on the fusion of multiple modalities. Thus, this workshop series on multimedia content analysis in sports aims to contribute to the closure of this research gap by bringing together the breadth and depth of these diverse approaches to stimulate each other with new ideas and foster research progress. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 2 |
| 2025 | Harborfront Anomaly DetectionabstractAbstract Creating high-quality datasets for the task of video anomaly detection is challenging due to a subjective anomaly definition and the rarity of anomalies, which oust the possibility of obtaining statistically significant data. This results in datasets where anomalies are placed in a single category, and are often considered less relevant from a security standpoint. Instead, we propose to create video anomaly datasets based on a framework utilizing object annotations to ease the annotation process and allow users to decide on the anomaly definition. Furthermore, this allows for a fine-grained evaluation w.r.t. anomaly types, which represents a novelty in the area of video anomaly detection. The framework is demonstrated using the existing thermal long-term drift (LTD) dataset, identifying and evaluating five different types of anomalies (appearance, motion, localization, density, and tampering) on six test sets. State-of-the-art anomaly detection methods are evaluated and found to underperform on the thermal anomaly detection dataset, which emphasizes a need for an adjustable anomaly definition in order to produce better anomaly datasets and models that generalize towards practical use. We share the code of the proposed framework to extract anomaly types along with object annotations for the LTD dataset at https://github.com/jagob/harborfront-vad . Jacob V. Dueholm, Mia Sandra Nicole Siemon, Radu Tudor Ionescu, Thomas B. Moeslund, Kamal Nasrollahi |
Neural Process. Lett. | 4 |
| 2024 | A Noisy Elephant in the Room: Is Your out-of-Distribution Detector Robust to Label Noise?abstractThe ability to detect unfamiliar or unexpected images is essential for safe deployment of computer vision systems. In the context of classification, the task of detecting images outside of a model's training domain is known as out-of-distribution (OOD) detection. While there has been a growing research interest in developing post-hoc OOD detection methods, there has been comparably little discussion around how these methods perform when the underlying classifier is not trained on a clean, carefully curated dataset. In this work, we take a closer look at 20 state-of-the-art OOD detection methods in the (more realistic) scenario where the labels used to train the underlying classifier are unreliable (e.g. crowd-sourced or web-scraped labels). Extensive experiments across different datasets, noise types & levels, architectures and checkpointing strategies provide insights into the effect of class label noise on OOD detection, and show that poor separation between incorrectly classified ID samples vs. OOD samples is an overlooked yet important limitation of existing methods. Code: https://github.com/glhr/ood-labelnoise Galadrielle Humblot-Renaux, Sergio Escalera, Thomas B. Moeslund |
CVPR | 3 |
| 2024 | Agglomerative Token Clustering
Joakim Bruslund Haurum, Sergio Escalera, Graham W. Taylor, Thomas B. Moeslund |
ECCV (57) | 4 |
| 2024 | PDA-RWSR: Pixel-Wise Degradation Adaptive Real-World Super-ResolutionabstractWhile many methods have been proposed to solve the Super-Resolution (SR) problem of Low-Resolution (LR) images with complex unknown degradations, their performance still drops significantly when evaluated on images with challenging real-world degradations. One often overlooked factor contributing to this, is the presence of spatially varying degradations in real LR images. To address this issue, we propose a novel degradation pipeline capable of generating paired LR/High-Resolution (HR) images with spatially varying noise, a key contributor to reduced image quality. Furthermore, to fully leverage such training data, we novelly propose a Pixel-Wise Degradation Adaptive Real-World Super-Resolution (PDA-RWSR) framework. Specifically, we design a new Restormer-based Real-World Super-Resolution (RWSR) model capable of adapting the reconstruction process based on pixel-wise degradation features extracted by a new supervised degradation estimation model. Along with our proposed method, we also introduce a new challenging real-world Spatially Variant Super-Resolution (SVSR) benchmarking dataset, where the images are degraded by complex noise of varying intensity and type, to evaluate the robustness of existing RWSR methods. Comprehensive experiments on synthetic and the proposed challenging real dataset demonstrates the superiority of our method over the current State-of-The-Art (SoTA). The SVSR dataset is available at https://doi.org/10.5281/zenodo.10044260. Andreas Aakerberg, Majed El Helou, Kamal Nasrollahi, Thomas B. Moeslund |
WACV | 4 |
| 2024 | CL-MAE: Curriculum-Learned Masked AutoencodersabstractMasked image modeling has been demonstrated as a powerful pretext task for generating robust representations that can be effectively generalized across multiple downstream tasks. Typically, this approach involves randomly masking patches (tokens) in input images, with the masking strategy remaining unchanged during training. In this paper, we propose a curriculum learning approach that updates the masking strategy to continually increase the complexity of the self-supervised reconstruction task. We conjecture that, by gradually increasing the task complexity, the model can learn more sophisticated and transferable representations. To facilitate this, we introduce a novel learnable masking module that possesses the capability to generate masks of different complexities, and integrate the proposed module into masked autoencoders (MAE). Our module is jointly trained with the MAE, while adjusting its behavior during training, transitioning from a partner to the MAE (optimizing the same reconstruction loss) to an adversary (optimizing the opposite loss), while passing through a neutral state. The transition between these behaviors is smooth, being regulated by a factor that is multiplied with the reconstruction loss of the masking module. The resulting training procedure generates an easy-to-hard curriculum. We train our Curriculum-Learned Masked Autoencoder (CL-MAE) on ImageNet and show that it exhibits superior representation learning capabilities compared to MAE. The empirical results on five downstream tasks confirm our conjecture, demonstrating that curriculum learning can be successfully used to self-supervise masked autoencoders. We release our code at https://github.com/ristea/cl-mae. Neelu Madan, Nicolae-Catalin Ristea, Kamal Nasrollahi, Thomas B. Moeslund, Radu Tudor Ionescu |
WACV | 4 |
| 2024 | Self-Supervised Masked Convolutional Transformer Block for Anomaly DetectionabstractAnomaly detection has recently gained increasing attention in the field of computer vision, likely due to its broad set of applications ranging from product fault detection on industrial production lines and impending event detection in video surveillance to finding lesions in medical scans. Regardless of the domain, anomaly detection is typically framed as a one-class classification task, where the learning is conducted on normal examples only. An entire family of successful anomaly detection methods is based on learning to reconstruct masked normal inputs (e.g. patches, future frames, etc.) and exerting the magnitude of the reconstruction error as an indicator for the abnormality level. Unlike other reconstruction-based methods, we present a novel self-supervised masked convolutional transformer block (SSMCTB) that comprises the reconstruction-based functionality at a core architectural level. The proposed self-supervised block is extremely flexible, enabling information masking at any layer of a neural network and being compatible with a wide range of neural architectures. In this work, we extend our previous self-supervised predictive convolutional attentive block (SSPCAB) with a 3D masked convolutional layer, a transformer for channel-wise attention, as well as a novel self-supervised objective based on Huber loss. Furthermore, we show that our block is applicable to a wider variety of tasks, adding anomaly detection in medical images and thermal videos to the previously considered tasks based on RGB images and surveillance videos. We exhibit the generality and flexibility of SSMCTB by integrating it into multiple state-of-the-art neural models for anomaly detection, bringing forth empirical results that confirm considerable performance improvements on five benchmarks: MVTec AD, BRATS, Avenue, ShanghaiTech, and Thermal Rare Event. Neelu Madan, Nicolae-Catalin Ristea, Radu Tudor Ionescu, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B. Moeslund, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | What is Proactive Human-Robot Interaction? - A Review of a Progressive Field and Its DefinitionsabstractDuring the past 15 years, an increasing amount of works have investigated proactive robotic behavior in relation to Human–Robot Interaction (HRI). The works engage with a variety of research topics and technical challenges. In this article, a review of the related literature identified through a structured block search is performed. Variations in the corpus are investigated, and a definition of Proactive HRI is provided. Furthermore, a taxonomy is proposed based on the corpus and exemplified through specific works. Finally, a selection of noteworthy observations is discussed. Marike K. van den Broek, Thomas B. Moeslund |
ACM Trans. Hum. Robot Interact. | 2 |
| 2023 | MMSports '23: 6th International Workshop on Multimedia Content Analysis in SportsabstractThe sixth ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'23) is part of the ACM International Conference on Multimedia 2023 (ACM Multimedia 2023). The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing multimedia/multimodal data in sports, sports broadcasts, sports games and sports medicine. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation and understanding, for statistical analysis and evaluation, and for sensor fusion during workouts as well as competitions. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Hideo Saito 0001, Thomas B. Moeslund, Rainer Lienhart |
ACM Multimedia | 2 |
| 2023 | SSMTL++: Revisiting self-supervised multi-task learning for video anomaly detection
Antonio Barbalau, Radu Tudor Ionescu, Mariana-Iuliana Georgescu, Jacob V. Dueholm, Bharathkumar Ramachandra, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B. Moeslund, Mubarak Shah |
Comput. Vis. Image Underst. | 8 |
| 2023 | Video Transformers: A SurveyabstractTransformer models have shown great success handling long-range interactions, making them a promising tool for modeling video. However, they lack inductive biases and scale quadratically with input length. These limitations are further exacerbated when dealing with the high dimensionality introduced by the temporal dimension. While there are surveys analyzing the advances of Transformers for vision, none focus on an in-depth analysis of video-specific designs. In this survey, we analyze the main contributions and trends of works leveraging Transformers to model video. Specifically, we delve into how videos are handled at the input level first. Then, we study the architectural changes made to deal with video more efficiently, reduce redundancy, re-introduce useful inductive biases, and capture long-term temporal dynamics. In addition, we provide an overview of different training regimes and explore effective self-supervised learning strategies for video. Finally, we conduct a performance comparison on the most common benchmark for Video Transformers (i.e., action classification), finding them to outperform 3D ConvNets even with less computational complexity. Javier Selva, Anders Skaarup Johansen, Sergio Escalera, Kamal Nasrollahi, Thomas B. Moeslund, Albert Clapés |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Self-Supervised Predictive Convolutional Attentive Block for Anomaly DetectionabstractAnomaly detection is commonly pursued as a one-class classification problem, where models can only learn from normal training samples, while being evaluated on both normal and abnormal test samples. Among the successful approaches for anomaly detection, a distinguished category of methods relies on predicting masked information (e.g. patches, future frames, etc.) and leveraging the reconstruction error with respect to the masked information as an abnormality score. Different from related methods, we propose to integrate the reconstruction-based functionality into a novel self-supervised predictive architectural building block. The proposed self-supervised block is generic and can easily be incorporated into various state-of-the-art anomaly detection methods. Our block starts with a convolutional layer with dilated filters, where the center area of the receptive field is masked. The resulting activation maps are passed through a channel attention module. Our block is equipped with a loss that minimizes the reconstruction error with respect to the masked area in the receptive field. We demonstrate the generality of our block by integrating it into several state-of-the-art frameworks for anomaly detection on image and video, providing empirical evidence that shows considerable performance improvements on MVTec AD, Avenue, and ShanghaiTech. We release our code as open source at: https://github.com/ristea/sspcab. Nicolae-Catalin Ristea, Neelu Madan, Radu Tudor Ionescu, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B. Moeslund, Mubarak Shah |
CVPR | 6 |
| 2022 | MOTCOM: The Multi-Object Tracking Dataset Complexity Metric
Malte Pedersen, Joakim Bruslund Haurum, Patrick Dendorfer, Thomas B. Moeslund |
ECCV (8) | 4 |
| 2022 | A graph-based approach to video anomaly detection from the perspective of superpixelsabstractVideo Anomaly Detection refers to the concept of discovering activities in a video feed that deviate from the usual visible pattern. It is a very well-studied and explored field in the domain of Computer Vision and Deep Learning, in which automated learning-based systems are capable of detecting certain kinds of anomalies at an accuracy greater than 90%. Deep Learning based Artificial Neural Network models, however, suffer from very low interpretability. In order to address and design a possible solution for this issue, this work proposes to shape the given problem by means of graphical models. Given the high flexibility of compositing easily interpretable graphs, a great variety of techniques exist to build a model representing spatial as well as temporal relationships occurring in the given video sequence. The experiments conducted on common anomaly detection benchmark datasets show that significant performance gains can be achieved through simple re-modelling of individual graph components. In contrast to other video anomaly detection approaches, the one presented in this work focuses primarily on the exploration of the possibility to shift the way we currently look at and process videos when trying to detect anomalous events. Mia Sandra Nicole Siemon, Kamal Nasrollahi, Thomas B. Moeslund |
ICMV | 3 |
| 2022 | MMSports'22: 5th International ACM Workshop on Multimedia Content Analysis in SportsabstractThe fifth ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'22) is part of the ACM International Conference on Multimedia 2022 (ACM Multimedia 2022). After two years of pure virtual MMSports workshops due to COVID-19, MMSports'22 is held on-site again. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding, and visualizing multimedia/multimodal data in sports, sports broadcasts, sports games and sports medicine. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation and understanding, for statistical analysis and evaluation, and for sensor fusion during workouts as well as competitions. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Related Workshop Proceedings are available in the ACM DL at: https://dl.acm.org/doi/proceedings/10.1145/3552437. Hideo Saito 0001, Thomas B. Moeslund, Rainer Lienhart |
ACM Multimedia | 2 |
| 2022 | Multi-Task Classification of Sewer Pipe Defects and Properties using a Cross-Task Graph Neural Network DecoderabstractThe sewerage infrastructure is one of the most important and expensive infrastructures in modern society. In order to efficiently manage the sewerage infrastructure, automated sewer inspection has to be utilized. However, while sewer defect classification has been investigated for decades, little attention has been given to classifying sewer pipe properties such as water level, pipe material, and pipe shape, which are needed to evaluate the level of sewer pipe deterioration.In this work we classify sewer pipe defects and properties concurrently and present a novel decoder-focused multi-task classification architecture Cross-Task Graph Neural Network (CT-GNN), which refines the disjointed per-task predictions using cross-task information. The CT-GNN architecture extends the traditional disjointed task-heads decoder, by utilizing a cross-task graph and unique class node embeddings. The cross-task graph can either be determined a priori based on the conditional probability between the task classes or determined dynamically using self-attention. CT-GNN can be added to any backbone and trained end-to-end at a small increase in the parameter count. We achieve state-of-the-art performance on all four classification tasks in the Sewer-ML dataset, improving defect classification and water level classification by 5.3 and 8.0 percentage points, respectively. We also outperform the single task methods as well as other multi-task classification approaches while introducing 50 times fewer parameters than previous model-focused approaches. The code and models are available at the project page http://vap.aau.dk/ctgnn. Joakim Bruslund Haurum, Meysam Madadi, Sergio Escalera, Thomas B. Moeslund |
WACV | 4 |
| 2022 | Real-world super-resolution of face-images from surveillance camerasabstractAbstract Most existing face image Super‐Resolution (SR) methods assume that the Low‐Resolution (LR) images were artificially downsampled from High‐Resolution (HR) images with bicubic interpolation. This operation changes the natural image characteristics and reduces noise. Hence, SR methods trained on such data most often fail to produce good results when applied to real LR images. To solve this problem, a novel framework for the generation of realistic LR/HR training pairs is proposed. The framework estimates realistic blur kernels, noise distributions, and JPEG compression artifacts to generate LR images with similar image characteristics as the ones in the source domain. This allows to train an SR model using high‐quality face images as Ground‐Truth (GT). For better perceptual quality, a Generative Adversarial Network (GAN) based SR model is used, where the commonly used VGG‐loss [1] is exchanged with LPIPS‐loss [2]. Experimental results on both real and artificially corrupted face images show that our method results in more detailed reconstructions with less noise compared to the existing State‐of‐the‐Art (SoTA) methods. In addition, it is shown that the traditional non‐reference Image Quality Assessment (IQA) methods fail to capture this improvement and demonstrate that the more recent NIMA metric [3] correlates better with human perception via Mean Opinion Rank (MOR). Andreas Aakerberg, Kamal Nasrollahi, Thomas B. Moeslund |
IET Image Process. | 3 |
| 2022 | Deep transfer learning in human-robot interaction for cognitive and physical rehabilitation purposes
Chaudhary Muhammad Aqdus Ilyas, Matthias Rehm, Kamal Nasrollahi, Yeganeh Madadi, Thomas B. Moeslund, Vahid Seydi |
Pattern Anal. Appl. | 5 |
| 2022 | Automatic estimation of clothing insulation rate and metabolic rate for dynamic thermal comfort assessment
Isak Worre Foged, Thomas B. Moeslund |
Pattern Anal. Appl. | 3 |
| 2022 | Visual explanation of black-box model: Similarity Difference and Uniqueness (SIDU) methodabstractExplainable Artificial Intelligence (XAI) has in recent years become a well-suited framework to generate human understandable explanations of ‘black- box’ models. In this paper, a novel XAI visual explanation algorithm known as the Similarity Difference and Uniqueness (SIDU) method that can effectively localize entire object regions responsible for prediction is presented in full detail. The SIDU algorithm robustness and effectiveness is analyzed through various computational and human subject experiments. In particular, the SIDU algorithm is assessed using three different types of evaluations (Application, Human and Functionally-Grounded) to demonstrate its superior performance. The robustness of SIDU is further studied in the presence of adversarial attack on ’black-box’ models to better understand its performance. Our code is available at: https://github.com/satyamahesh84/SIDU_XAI_CODE. Satya M. Muddamsetty, Mohammad Naser Sabet Jahromi, Andreea E. Ciontos, Laura M. Fenoy, Thomas B. Moeslund |
Pattern Recognit. | 5 |
| 2022 | Deep Pain: Exploiting Long Short-Term Memory Networks for Facial Expression ClassificationabstractPain is an unpleasant feeling that has been shown to be an important factor for the recovery of patients. Since this is costly in human resources and difficult to do objectively, there is the need for automatic systems to measure it. In this paper, contrary to current state-of-the-art techniques in pain assessment, which are based on facial features only, we suggest that the performance can be enhanced by feeding the raw frames to deep learning models, outperforming the latest state-of-the-art results while also directly facing the problem of imbalanced data. As a baseline, our approach first uses convolutional neural networks (CNNs) to learn facial features from VGG_Faces, which are then linked to a long short-term memory to exploit the temporal relation between video frames. We further compare the performances of using the so popular schema based on the canonically normalized appearance versus taking into account the whole image. As a result, we outperform current state-of-the-art area under the curve performance in the UNBC-McMaster Shoulder Pain Expression Archive Database. In addition, to evaluate the generalization properties of our proposed methodology on facial motion recognition, we also report competitive results in the Cohn Kanade+ facial expression database. Pau Rodríguez, Guillem Cucurull, Jordi Gonzàlez 0001, Josep M. Gonfaus, Kamal Nasrollahi, Thomas B. Moeslund, F. Xavier Roca |
IEEE Trans. Cybern. | 6 |
| 2022 | Effective fusion of deep multitasking representations for robust visual tracking
Seyed Mojtaba Marvasti-Zadeh, Hossein Ghanei-Yakhdan, Shohreh Kasaei, Kamal Nasrollahi, Thomas B. Moeslund |
Vis. Comput. | 5 |
| 2021 | Single-Loss Multi-task Learning For Improving Semantic Segmentation Using Super-Resolution
Andreas Aakerberg, Anders Skaarup Johansen, Kamal Nasrollahi, Thomas B. Moeslund |
CAIP (2) | 4 |
| 2021 | Object-Centric Anomaly Detection Using Memory Augmentation
Jacob V. Dueholm, Kamal Nasrollahi, Thomas B. Moeslund |
CAIP (1) | 3 |
| 2021 | Sewer-ML: A Multi-Label Sewer Defect Classification Dataset and BenchmarkabstractPerhaps surprisingly sewerage infrastructure is one of the most costly infrastructures in modern society. Sewer pipes are manually inspected to determine whether the pipes are defective. However, this process is limited by the number of qualified inspectors and the time it takes to inspect a pipe. Automatization of this process is therefore of high interest. So far, the success of computer vision approaches for sewer defect classification has been limited when compared to the success in other fields mainly due to the lack of public datasets. To this end, in this work we present a large novel and publicly available multi-label classification dataset for image-based sewer defect classification called Sewer-ML.The Sewer-ML dataset consists of 1.3 million images annotated by professional sewer inspectors from three different utility companies across nine years. Together with the dataset, we also present a benchmark algorithm and a novel metric for assessing performance. The benchmark algorithm is a result of evaluating 12 state-of-the-art algorithms, six from the sewer defect classification domain and six from the multi-label classification domain, and combining the best performing algorithms. The novel metric is a class-importance weighted F2 score, F2CIW, reflecting the economic impact of each class, used together with the normal pipe F1 score, F1Normal. The benchmark algorithm achieves an F2CIWscore of 55.11% and F1Normalscore of 90.94%, leaving ample room for improvement on the Sewer-ML dataset. The code, models, and dataset are available at the project page http://vap.aau.dk/sewer-ml Joakim Bruslund Haurum, Thomas B. Moeslund |
CVPR | 2 |
| 2021 | Generalizing Floor Plans Using Graph Neural NetworksabstractThe proliferation of indoor maps is limited by the manual process of generalizing floor plans. Previous attempts at automating similar processes use rasterization for structure. With Graph Neural Networks (GNN) it is now possible to skip rasterization and rely on the inherent structures in CAD drawings. A core component in floor plan generalization is localization of doors. We show how floor plan graphs can be extracted directly from CAD primitives and how state-of the-art GNNs can be used to classify graph nodes as door or non-door. Generalization is represented by the creation of placeholder bounding boxes using the labelled graph nodes. Our graph-based approach completely outperforms the Faster R-CNN baseline, which fail to locate any doors with the desired localization accuracy. To support further development of graph-based methods and comparison with raster-based methods, we publish a new dataset that consists of both image and graph-based floor plan representations. Code and dataset is available at https://github.com/Chrps/MapGeneralization. Christoffer P. Simonsen, Frederik M. Thiesson, Mark P. Philipsen, Thomas B. Moeslund |
ICIP | 4 |
| 2021 | Pose Estimation from RGB Images of Highly Symmetric Objects using a Novel Multi-Pose Loss and Differential RenderingabstractWe propose a novel multi-pose loss function to train a neural network for 6D pose estimation, using synthetic data and evaluating it on real images. Our loss is inspired by the VSD (Visible Surface Discrepancy) metric and relies on a differentiable renderer and CAD models. This novel multi-pose approach produces multiple weighted pose estimates to avoid getting stuck in local minima. Our method resolves pose ambiguities without using predefined symmetries. It is trained only on synthetic data. We test on real-world RGB images from the T-LESS dataset, containing highly symmetric objects common in industrial settings. We show that our solution can be used to replace the codebook in a state-of-the-art approach. So far, the codebook approach has had the shortest inference time in the field. Our approach reduces inference time further while a) avoiding discretization, b) requiring a much smaller memory footprint and c) improving pose recall.3 Stefan Hein Bengtson, Hampus Åström, Thomas B. Moeslund, Elin Anna Topp, Volker Krüger |
IROS | 3 |
| 2021 | MMSports'21: 4th International Workshop on Multimedia Content Analysis in Sports
Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 2 |
| 2021 | Human-Robot Trust Assessment Using Top-Down Visual Tracking After Robot Task Execution MistakesabstractWith increased interest in close-proximity human-robot collaboration in production settings it is important that we understand how robot behaviors and mistakes affect human-robot trust, as a lack of trust can cause loss in productivity and over-trust can lead to hazardous misuse. We designed a system for real-time human-robot trust assessment using a top-down depth camera tracking setup with the goal of using signs of physical apprehension to infer decreases in trust toward the robot. In an experiment with 20 participants we evaluated the tracking system in a repetitive collaborative pick-and-place task where the participant and the robot had to move a set of cones across a table. Midway through the tasks we disrupted the participants expectations by having the robot perform a trust-dampening action. Throughout the tasks we measured the participant’s preferred proximity and their trust toward the robot. Comparing irregular robot movements versus task execution mistakes as well simultaneous versus turn-taking collaboration, we found reported trust was significantly decreased when the robot performed an execution mistake going counter to the shared objective. This decrease was higher for participant working simultaneously as the robot. The effect of the trust-dampening actions on preferred proximity was inconclusive due to unexplained movement trends between tasks throughout the experiment. Despite being given the option to stop the robot in case of abnormal behavior, the trust-dampening actions did not increase the number of participant disruptions for the actions we tested. Kasper Hald, Matthias Rehm, Thomas B. Moeslund |
RO-MAN | 3 |
| 2021 | Memory- and time-efficient dense network for single-image super-resolutionabstractAbstract Dense connections in convolutional neural networks (CNNs), which connect each layer to every other layer, can compensate for mid/high‐frequency information loss and further enhance high‐frequency signals. However, dense CNNs suffer from high memory usage due to the accumulation of concatenating feature‐maps stored in memory. To overcome this problem, a two‐step approach is proposed that learns the representative concatenating feature‐maps. Specifically, a convolutional layer with many more filters is used before concatenating layers to learn richer feature‐maps. Therefore, the irrelevant and redundant feature‐maps are discarded in the concatenating layers. The proposed method results in 24% and 6% less memory usage and test time, respectively, in comparison to single‐image super‐resolution (SISR) with the basic dense block. It also improves the peak signal‐to‐noise ratio by 0.24 dB. Moreover, the proposed method, while producing competitive results, decreases the number of filters in concatenating layers by at least a factor of 2 and reduces the memory consumption and test time by 40% and 12%, respectively. These results suggest that the proposed approach is a more practical method for SISR. Nasrin Imanpour, Ahmad Reza Naghsh-Nilchi, S. Amirhassan Monadjemi, Hossein Karshenas, Kamal Nasrollahi, Thomas B. Moeslund |
IET Signal Process. | 6 |
| 2020 | A Context-Aware Loss Function for Action Spotting in Soccer VideosabstractIn video understanding, action spotting consists in temporally localizing human-induced events annotated with single timestamps. In this paper, we propose a novel loss function that specifically considers the temporal context naturally present around each action, rather than focusing on the single annotated frame to spot. We benchmark our loss on a large dataset of soccer videos, SoccerNet, and achieve an improvement of 12.8% over the baseline. We show the generalization capability of our loss for generic activity proposals and detection on ActivityNet, by spotting the beginning and the end of each activity. Furthermore, we provide an extended ablation study and display challenging cases for action spotting in soccer videos. Finally, we qualitatively illustrate how our loss induces a precise temporal understanding of actions and show how such semantic knowledge can be used for automatic highlights generation. Anthony Cioppa, Adrien Deliège, Silvio Giancola, Bernard Ghanem, Marc Van Droogenbroeck, Rikke Gade, Thomas B. Moeslund |
CVPR | 7 |
| 2020 | 3D-ZeF: A 3D Zebrafish Tracking Benchmark DatasetabstractIn this work we present a novel publicly available stereo based 3D RGB dataset for multi-object zebrafish tracking, called 3D-ZeF. Zebrafish is an increasingly popular model organism used for studying neurological disorders, drug addiction, and more. Behavioral analysis is often a critical part of such research. However, visual similarity, occlusion, and erratic movement of the zebrafish makes robust 3D tracking a challenging and unsolved problem. The proposed dataset consists of eight sequences with a duration between 15-120 seconds and 1-10 free moving zebrafish. The videos have been annotated with a total of 86,400 points and bounding boxes. Furthermore, we present a complexity score and a novel open-source modular baseline system for 3D tracking of zebrafish. The performance of the system is measured with respect to two detectors: a naive approach and a Faster R-CNN based fish head detector. The system reaches a MOTA of up to 77.6%. Links to the code and dataset is available at the project page http://vap.aau.dk/3d-zef. Malte Pedersen, Joakim Bruslund Haurum, Stefan Hein Bengtson, Thomas B. Moeslund |
CVPR | 4 |
| 2020 | Vision-based Individual Factors Acquisition for Thermal Comfort Assessment in a Built EnvironmentabstractTo maintain satisfactory chamber thermal environments for occupants, heating, ventilation and air conditioning (HVAC) systems have to work frequently. However, the room conditions especially the temperatures are usually set empirically which fail to consider occupants' real needs, not to mention personalized thermal comfort, therefore, the HVAC systems are underutilized and unavoidably induce energy waste. To solve this problem, a vision-based method to acquire multiple individual factors that are critical for assessing personalized thermal sensation is proposed. Specifically, with the indoor videos captured by a thermal camera as inputs, a convolutional neural network (CNN) is implemented to recognize an occupant's clothes and action type simultaneously. With a dataset of 20 persons, the experimental results show an average classification rate of 95.14% on 4 dataset partitions for a 15-category scenario, which prove the effectiveness of the proposed method. Isak Worre Foged, Thomas B. Moeslund |
FG | 3 |
| 2020 | Prediction of the Methane Production in Biogas Plants Using a Combined Gompertz and Machine Learning Model
Bolette Dybkjær Hansen, Jamshid Tamouk, Christian A. Tidmarsh, Rasmus Johansen, Thomas B. Moeslund, David G. Jensen |
ICCSA (1) | 5 |
| 2020 | One-To-One Person Re-Identification For Queue Time EstimationabstractQueue time measurements in airport check-points are essential in order to regulate staff allocation and maintain a low queue time. In this paper, we propose a method to measure queue times based on person re-identification. Specifically, we capture passenger features from entrance and exit points of an airport check-point and match features to determine queue times of passengers. However, a passenger seen at the exit, potentially, can be matched to multiple passengers seen at the entrance. Therefore, to increase the precision of re-identification, we propose a simple, yet effective, algorithm that assigns each passenger seen at the exit to only a single passenger seen at the entrance. Through experiments on a dataset collected from an airport immigration check-point, we show that our proposed assignment increases precision and recall by 22 % and 16 %, respectively, to naively assigning the best match. Aske R. Lejbølle, Benjamin Krogh, Kamal Nasrollahi, Thomas B. Moeslund |
ICIP | 4 |
| 2020 | SIDU: Similarity Difference And Uniqueness Method for Explainable AIabstractA new brand of technical artificial intelligence (Explainable AI) research has focused on trying to open up the `black box' and provide some explainability. This paper presents a novel visual explanation method for deep learning networks in the form of a saliency map that can effectively localize entire object regions. In contrast to the current state-of-the art methods, the proposed method shows quite promising visual explanations that can gain greater trust of human expert. Both quantitative and qualitative evaluations are carried out on both general and clinical data sets to confirm the effectiveness of the proposed method. Satya M. Muddamsetty, Mohammad Naser Sabet Jahromi, Thomas B. Moeslund |
ICIP | 3 |
| 2020 | Human-Robot Trust Assessment Using Motion Tracking & Galvanic Skin ResponseabstractIn this study we set out to design a computer vision-based system to assess human-robot trust in real time during close-proximity human-robot collaboration. This paper presents the setup and hardware for an augmented reality-enabled human-robot collaboration cell as well as a method of measuring operator proximity using an infrared camera. We tested this setup as a tool for assessing trust through physical apprehension signals in a collaborative drawing task, where participants hold a piece of paper on a table while the robot draws between their hands. Midway through the test we attempt to induce a decrease in trust with an unexpected change in robot speed and evaluate subject motions along with self-reported trust and emotional arousal through galvanic skin response. After performing the experiment with forty participants, we found that reported trust was significantly affected when robot movement speed was increased. The galvanic skin response measurement were not significantly different between the test conditions. The motion tracking method used in this study did not suggest that subjects' motions were significantly affected by the decrease in trust. Kasper Hald, Matthias Rehm, Thomas B. Moeslund |
IROS | 3 |
| 2020 | Sewer Deterioration Modeling: The Effect of Training a Random Forest Model on Logically Selected Data-groupsabstractBreakdown of sewers can induce significantly damage to roads and buildings placed upon it. For this reason, timely maintenance of the sewer system is essential. However, due to the under-ground position of the sewers they are very expensive to monitor, as this is done by CCTV inspection. Therefore, it is important to choose the right sewers for inspection and several decision-support tools have been developed to help the operators to select which sewers to inspect. These decision support tools all contain a model which predicts the condition of the sewers, and recently several models have been proposed in order to increase the performance. The scope of this paper is to investigate the effect of training a Random Forest model on logically selected groups of data, as opposed to training of a joined model on the full data set. The selected data groups were based on expert knowledge: The first data groups were based on the sewer material (concrete, plastic, clay, reinforced with lining and other material). The concrete data set was then further sub-divided into wastewater types (sewage, rain and combined) whereas the plastic data set was sub-divided into road classes. The results showed that the model trained on the full data set performed better than the models trained on logically selected data-groups as it encounters the heterogeneity of the data set. Furthermore, this answers an important question raised by end users of the deterioration models. Bolette Dybkjær Hansen, Søren Højmark Rasmussen, Thomas B. Moeslund, Mads Uggerby, David G. Jensen |
KES | 3 |
| 2020 | MMSports'20: 3rd International Workshop on Multimedia Content Analysis in SportsabstractThe third ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'20) is part of the ACM International Conference on Multimedia 2020 (ACM Multimedia 2020). Exceptionally, due to the corona pandemic, the workshop is held virtually. The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding and visualizing the multimedia/multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation, understanding, statistical analysis and evaluation, and sensor fusion. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 2 |
| 2020 | Deep visual unsupervised domain adaptation for classification tasks: a surveyabstractLearning methods are challenged when there is not enough labelled data. It gets worse when the existing learning data have different distributions in different domains. To deal with such situations, deep unsupervised domain adaptation techniques have newly been widely used. This study surveys such domain adaptation methods that have been used for classification tasks in computer vision. The survey includes the very recent papers on this topic that have not been included in the previous surveys and introduces a taxonomy by grouping methods published on unsupervised domain adaptation into five groups of discrepancy‐, adversarial‐, reconstruction‐, representation‐, and attention‐based methods. Yeganeh Madadi, Vahid Seydi, Kamal Nasrollahi, Reshad Hosseini, Thomas B. Moeslund |
IET Image Process. | 5 |
| 2020 | Cycle-consistent generative adversarial neural networks based low quality fingerprint enhancement
Dogus Karabulut, Pavlo Tertychnyi, Hasan Sait Arslan, Cagri Ozcinar, Kamal Nasrollahi, Joan Valls, Joan Vilaseca, Thomas B. Moeslund, Gholamreza Anbarjafari |
Multim. Tools Appl. | 8 |
| 2020 | Special issue on Advanced Machine Vision
Steven Puttemans, Toon Goedemé, Ajmal Mian, Thomas B. Moeslund, Rikke Gade |
Mach. Vis. Appl. | 4 |
| 2020 | A Double-Deep Spatio-Angular Learning Framework for Light Field-Based Face RecognitionabstractFace recognition has attracted increasing attention due to its wide range of applications, but it is still challenging when facing large variations in the biometric data characteristics. Lenslet light field cameras have recently come into prominence to capture rich spatio-angular information, thus offering new possibilities for advanced biometric recognition systems. This paper proposes a double-deep spatio-angular learning framework for light field-based face recognition, which is able to model both the intra-view/spatial and inter-view/angular information using two deep networks in sequence. This is a novel recognition framework that has never been proposed in the literature for face recognition or any other visual recognition task. The proposed double-deep learning framework includes a long short-term memory (LSTM) recurrent network, whose inputs are VGG-Face descriptions, computed using a VGG-16 convolutional neural network (CNN). The VGG-Face spatial descriptions are extracted from a selected set of 2D sub-aperture (SA) images rendered from the light field image, corresponding to different observation angles. A sequence of the VGG-Face spatial descriptions is then analyzed by the LSTM network. A comprehensive set of experiments has been conducted using the IST-EURECOM light field face database, addressing varied and challenging recognition tasks. The results show that the proposed framework achieves superior face recognition performance when compared to the state of the art. Alireza Sepas-Moghaddam, Mohammad A. Haque, Paulo Lobato Correia, Kamal Nasrollahi, Thomas B. Moeslund, Fernando Pereira 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Person Re-Identification Using Spatial and Layer-Wise AttentionabstractPerson re-identification requires extraction of discriminative features to ensure a correct match; this must be done independent of challenges, such as occlusion, view, or lighting changes. While occlusion can be eliminated by changing the camera setup from a horizontal to a vertical (overhead) viewpoint, other challenges arise as the total visible surface area of persons is decreased. As a result, methods that focus on the most discriminative regions of persons must be applied, while different domains should also be considered to extract different semantics. To further increase feature discriminability, complementary features extracted at different abstraction levels should be fused. To emphasize features at certain abstraction levels depending on the input, fusion should be done intelligently. This work considers multiple domains and feature discrimination, where a multimodal convolution neural network is applied to fuse RGB and depth information. To extract multi-local discriminative features, two different attention modules are proposed: (1) a spatial attention module, which is able to capture local information at different abstraction levels, and (2) a layer-wise attention module, which works as a dynamic weighting scheme to assign weights and fuse local abstraction-level features intelligently, depending on the input image. By fusing local and global features in a multimodal context, we show state-of-the-art accuracies on two publicly available datasets, DPI-T and TVPR, while increasing the state-of-the-art accuracy on a third dataset, OPR. Finally, through both visual and quantitative analysis, we show the ability of the proposed system to leverage multiple frames, by adapting feature weighting depending on the input. Aske R. Lejbølle, Kamal Nasrollahi, Benjamin Krogh, Thomas B. Moeslund |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | Augmented Reality Technology for Displaying Close-Proximity Sub-Surface Positions
Kasper Hald, Matthias Rehm, Thomas B. Moeslund |
INTERACT (2) | 3 |
| 2019 | MMSports'19: 2nd ACM International Workshop on Multimedia Content Analysis in SportsabstractThe second ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'19) is held in Nice, France on October 25th, 2019 co-located with the ACM International Conference on Multimedia 2019 (ACM Multimedia 2019). The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining, analyzing, understanding and visualizing the multimedia/multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation, understanding, statistical analysis and evaluation, and sensor fusion. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this workshop series on multimedia content analysis in sports. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 2 |
| 2019 | Proposing Human-Robot Trust Assessment Through Tracking Physical Apprehension Signals in Close-Proximity Human-Robot CollaborationabstractWe propose a method of human-robot trust assessment in close-proximity human-robot collaboration involving body tracking for recognition of physical signs of apprehension. We tested this by performing skeleton tracking on 30 participant while they repeated a shared task with a Sawyer robot while reporting trust between tasks. We tested different robot velocity and environment conditions with an unannounced increase in velocity midway through to provoke a dip trust. Initial analysis show significant effect for the test conditions on participant movements and reported trust as well as linear correlations between tracked signs of apprehension and reported trust. Kasper Hald, Matthias Rehm, Thomas B. Moeslund |
RO-MAN | 3 |
| 2019 | Teaching Pepper Robot to Recognize Emotions of Traumatic Brain Injured Patients Using Deep Neural NetworksabstractSocial signal extraction from the facial analysis is a popular research area in human-robot interaction. However, recognition of emotional signals from Traumatic Brain Injured (TBI) patients with the help of robots and non-intrusive sensors is yet to be explored. Existing robots have limited abilities to automatically identify human emotions and respond accordingly. Their interaction with TBI patients could be even more challenging and complex due to unique, unusual and diverse ways of expressing their emotions. To tackle the disparity in a TBI patient's Facial Expressions (FEs), a specialized deep-trained model for automatic detection of TBI patients' emotions and FE (TBI-FER model) is designed, for robot-assisted rehabilitation activities. In addition, the Pepper robot's built-in model for FE is investigated on TBI patients as well as on healthy people. Variance in their emotional expressions is determined by comparative studies. It is observed that the customized trained system is highly essential for the deployment of Pepper robot as a Socially Assistive Robot (SAR). Chaudhary Muhammad Aqdus Ilyas, Viktor Schmuck, Mohammad A. Haque, Kamal Nasrollahi, Matthias Rehm, Thomas B. Moeslund |
RO-MAN | 6 |
| 2019 | A novel deep network architecture for reconstructing RGB facial images from thermal for face recognition
Andre Litvin, Kamal Nasrollahi, Sergio Escalera, Cagri Ozcinar, Thomas B. Moeslund, Gholamreza Anbarjafari |
Multim. Tools Appl. | 5 |
| 2019 | Guest editorial: special issue on human abnormal behavioural analysis
Gholamreza Anbarjafari, Sergio Escalera, Kamal Nasrollahi, Hugo Jair Escalante, Xavier Baró, Jun Wan 0001, Thomas B. Moeslund |
Mach. Vis. Appl. | 7 |
| 2019 | Rain Removal in Traffic Surveillance: Does it Matter?abstractVarying weather conditions, including rainfall and snowfall, are generally regarded as a challenge for computer vision algorithms. One proposed solution to the challenges induced by rain and snowfall is to artificially remove the rain from images or video using rain removal algorithms. It is the promise of these algorithms that the rain-removed image frames will improve the performance of subsequent segmentation and tracking algorithms. However, rain removal algorithms are typically evaluated on their ability to remove synthetic rain on a small subset of images. Currently, their behavior is unknown on real-world videos when integrated with a typical computer vision pipeline. In this paper, we review the existing rain removal algorithms and propose a new dataset that consists of 22 traffic surveillance sequences under a broad variety of weather conditions that all include either rain or snowfall. We propose a new evaluation protocol that evaluates the rain removal algorithms on their ability to improve the performance of subsequent segmentation, instance segmentation, and feature tracking algorithms under rain and snow. If successful, the de-rained frames of a rain removal algorithm should improve segmentation performance and increase the number of accurately tracked features. The results show that a recent single-frame-based rain removal algorithm increases the segmentation performance by 19.7% on our proposed dataset, but it eventually decreases the feature tracking performance and showed mixed results with recent instance segmentation methods. However, the best video-based rain removal algorithm improves the feature tracking accuracy by 7.72%. Chris Bahnsen, Thomas B. Moeslund |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Image-based Sea/Land Map Generation from Radar Dataabstract2D radars are efficient sensors used for e.g. coastal or shipborne surveillance. However, the recorded data contains echoes from all its surroundings, without any discrimination of land, sea or occluded terrain, which degrades the performance of target detectors and trackers. We assume that a complete 360° radar scan can be used as an image and thereby exploit its spatial information with a multi-scale feature-connected convolutional autoencoder to perform image-based radar segmentation. Our method is compared against the reimplementation of a temporal-based classifier when using unfiltered radar data. The conducted experiments display that our framework can overcome the noise problems inherit in 2D radar data and discriminate the different surfaces by outperforming the temporal-based implementation with a 20% increase in mean pixel-wise accuracy, with a mAP of 67%, and a mean IoU of 58.67%. This is a promising approach towards the application of deep learning for segmentation of radar-based images. Francesc Joan Riera, Rasmus Engholm, Lars W. Jochumsen, Thomas B. Moeslund |
AVSS | 4 |
| 2018 | Changes in Facial Expression as Biometric: A Database and Benchmarks of IdentificationabstractFacial dynamics can be considered as unique signatures for discrimination between people. These have started to become important topic since many devices have the possibility of unlocking using face recognition or verification. In this work, we evaluate the efficacy of the transition frames of video in emotion as compared to the peak emotion frames for identification. For experiments with transition frames we extract features from each frame of the video from a fine-tuned VGG-Face Convolutional Neural Network (CNN) and geometric features from facial landmark points. To model the temporal context of the transition frames we train a Long-Short Term Memory (LSTM) on the geometric and the CNN features. Furthermore, we employ two fusion strategies: first, an early fusion, in which the geometric and the CNN features are stacked and fed to the LSTM. Second, a late fusion, in which the prediction of the LSTMs, trained independently on the two features, are stacked and used with a Support Vector Machine (SVM). Experimental results show that the late fusion strategy gives the best results and the transition frames give better identification results as compared to the peak emotion frames. Rain Eric Haamer, Kaustubh Kulkarni, Nasrin Imanpour, Mohammad A. Haque, Egils Avots, Michelle Breisch, Kamal Nasrollahi, Sergio Escalera, Cagri Ozcinar, Xavier Baró, Ahmad Reza Naghsh-Nilchi, Thomas B. Moeslund, Gholamreza Anbarjafari |
FG | 12 |
| 2018 | Deep Multimodal Pain Recognition: A Database and Comparison of Spatio-Temporal Visual ModalitiesabstractPain is a symptom of many disorders associated with actual or potential tissue damage in human body. Managing pain is not only a duty but also highly cost prone. The most primitive state of pain management is the assessment of pain. Traditionally it was accomplished by self-report or visual inspection by experts. However, automatic pain assessment systems from facial videos are also rapidly evolving due to the need of managing pain in a robust and cost effective way. Among different challenges of automatic pain assessment from facial video data two issues are increasingly prevalent: first, exploiting both spatial and temporal information of the face to assess pain level, and second, incorporating multiple visual modalities to capture complementary face information related to pain. Most works in the literature focus on merely exploiting spatial information on chromatic (RGB) video data on shallow learning scenarios. However, employing deep learning techniques for spatio-temporal analysis considering Depth (D) and Thermal (T) along with RGB has high potential in this area. In this paper, we present the first state-of-the-art publicly available database, 'Multimodal Intensity Pain (MIntPAIN)' database, for RGBDT pain level recognition in sequences. We provide a first baseline results including 5 pain levels recognition by analyzing independent visual modalities and their fusion with CNN and LSTM models. From the experimental evaluation we observe that fusion of modalities helps to enhance recognition performance of pain levels in comparison to isolated ones. In particular, the combination of RGB, D, and T in an early fusion fashion achieved the best recognition rate. Mohammad A. Haque, Rubén Ballester, Fatemeh Noroozi, Kaustubh Kulkarni, Christian B. Laursen, Ramin Irani, Marco Bellantonio, Sergio Escalera, Gholamreza Anbarjafari, Kamal Nasrollahi, Ole Kæseler Andersen, Erika G. Spaich, Thomas B. Moeslund |
FG | 13 |
| 2018 | Rehabilitation of Traumatic Brain Injured Patients: Patient Mood Analysis from Multimodal VideoabstractRehabilitation after traumatic brain injury (TBI) is very critical as it is largely unpredictable depending upon the nature of the injury. Rehabilitation process and recovery time also varies, as it takes months and years, depending upon the assessment of treatment, mental and physical conditions and strategies. Due to non-cooperative behaviour of patients, and increase in negative emotional expressions it is very beneficial to evaluate these expressions in a contactless way, and perform a rehabilitation physiotherapy, cognitive or other behavioral activities when the patient is in a positive mood. In this paper we have analyzed the methods for facial features extraction for TBI patients to determine optimal time to have aforementioned rehabilitation process on the basis of positive and negative facial expressions. We have employed a deep learning architecture based on convolutional neural network and long short term memory on RGB and thermal data that were collected in challenging scenarios from real patients. It automatically identifies the patient's facial expressions, and inform experts or trainers that “it is the time” to start rehabilitation session. Chaudhary Muhammad Aqdus Ilyas, Kamal Nasrollahi, Matthias Rehm, Thomas B. Moeslund |
ICIP | 4 |
| 2018 | Occlusion-aware pedestrian detectionabstractFailure in pedestrian detection systems can be extremely crucial, specifically in driverless driving. In this paper, failures in pedestrian detectors are refined by re-evaluating the results of state of the art pedestrian detection systems, via a fully convolutional neural network. The network is trained on a number of datasets which include a custom designed occluded pedestrian dataset to address the problem of occlusion. Results show that when applying the proposed network, detectors can not only maintain their state of the art performance, but they even decrease average false positives rate per image, especially in the case where pedestrians are occluded. Christos Apostolopoulos, Kamal Nasrollahi, M. Hsuan Yang, Mohammad Naser Sabet Jahromi, Thomas B. Moeslund |
ICMV | 5 |
| 2018 | Multimodal heartbeat rate estimation from the fusion of facial RGB and thermal videosabstractMeasuring Heartbeat Rate (HR) is an important tool for monitoring the health of a person. When the heart beats the influx of blood to the head causes slight involuntary movement and subtle skin color changes, which cannot be seen by the naked eye but can be tracked from facial videos using computer vision techniques and can be analyzed to estimate the HR. However, the current state of the art solutions encounter an increasing amount of complications when the subject has voluntary motion on the face or when the lighting conditions change in the video. Thus the accuracy of the HR estimation using computer vision is still inferior to that of a physical Electrocardiography (ECG) based system. The aim of this work is to improve the current non-invasive HR measurement by fusing the motion-based and color-based HR estimation methods and using them on multiple input modalities, e.g., RGB and thermal imaging. Our experiments indicate that late-fusion of the results of these methods (motion and color-based) applied to these different modalities, produces more accurate results compared to the existing solutions Anders Skaarup Johansen, Jesper W. Henriksen, Mohammad A. Haque, Mohammad Naser Sabet Jahromi, Kamal Nasrollahi, Thomas B. Moeslund |
ICMV | 6 |
| 2018 | 1st ACM International Workshop on Multimedia Content Analysis in SportsabstractThe first ACM International Workshop on Multimedia Content Analysis in Sports (ACM MMSports'18) is held in Seoul, South Korea on October 26th, 2018 and is co-located with the ACM International Conference on Multimedia 2018 (ACM Multimedia 2018). The goal of this workshop is to bring together researchers and practitioners from academia and industry to address challenges and report progress in mining and content analysis of multimedia/multimodal data in sports. The combination of sports and modern technology offers a novel and intriguing field of research with promising approaches for visual broadcast augmentation, understanding, statistical analysis and evaluation, and sensor fusion. There is a lack of research communities focusing on the fusion of multiple modalities. We are helping to close this research gap with this first workshop of a serious workshops on multimedia content analysis in sports. Rainer Lienhart, Thomas B. Moeslund, Hideo Saito 0001 |
ACM Multimedia | 2 |
| 2018 | Getting Crevices, Cracks, and Grooves in Line: Anomaly Categorization for AQC Judgment ModelsabstractThe following topics are dealt with: quality of experience; virtual reality; video streaming; computer games; video signal processing; Internet; rendering (computer graphics); learning (artificial intelligence); visual databases; and video coding. Anne Juhler Hansen, Hendrik Knoche, Thomas B. Moeslund |
QoMEX | 3 |
| 2018 | Back-dropout transfer learning for action recognitionabstractTransfer learning aims at adapting a model learned from source dataset to target dataset. It is a beneficial approach especially when annotating on the target dataset is expensive or infeasible. Transfer learning has demonstrated its powerful learning capabilities in various vision tasks. Despite transfer learning being a promising approach, it is still an open question how to adapt the model learned from the source dataset to the target dataset. One big challenge is to prevent the impact of category bias on classification performance. Dataset bias exists when two images from the same category, but from different datasets, are not classified as the same. To address this problem, a transfer learning algorithm has been proposed, called negative back‐dropout transfer learning (NB‐TL), which utilizes images that have been misclassified and further performs back‐dropout strategy on them to penalize errors. Experimental results demonstrate the effectiveness of the proposed algorithm. In particular, the authors evaluate the performance of the proposed NB‐TL algorithm on UCF 101 action recognition dataset, achieving 88.9% recognition rate. Huamin Ren, Nattiya Kanhabua, Andreas Møgelmose, Weifeng Liu 0002, Kaustubh Kulkarni, Sergio Escalera, Xavier Baró, Thomas B. Moeslund |
IET Comput. Vis. | 8 |
| 2017 | Depth Value Pre-Processing for Accurate Transfer Learning based RGB-D Object RecognitionabstractObject recognition is one of the important tasks in computer vision which has found enormous applications.Depth modality is proven to provide supplementary information to the common RGB modality for objectrecognition. In this paper, we propose methods to improve the recognition performance of an existing deeplearning based RGB-D object recognition model, namely the FusionNet proposed by Eitel et al. First, we showthat encoding the depth values as colorized surface normals is beneficial, when the model is initialized withweights learned from training on ImageNet data. Additionally, we show that the RGB stream of the FusionNetmodel can benefit from using deeper network architectures, namely the 16-layered VGGNet, in exchange forthe 8-layered CaffeNet. In combination, these changes improves the recognition performance with 2.2% incomparison to the original FusionNet, when evaluating on the Washington RGB-D Object Dataset. Andreas Aakerberg, Kamal Nasrollahi, Christoffer B. Rasmussen, Thomas B. Moeslund |
IJCCI | 4 |
| 2017 | Real-Time Barcode Detection and Classification using Deep LearningabstractBarcodes, in their different forms, can be found on almost any packages available in the market. Detecting and then decoding of barcodes have therefore great applications. We describe how to adapt the state-of-the- art deep learning-based detector of You Only Look Once (YOLO) for the purpose of detecting barcodes in a fast and reliable way. The detector is capable of detecting both 1D and QR barcodes. The detector achieves state-of-the-art results on the benchmark dataset of Muenster BarcodeDB with a detection rate of 0.991. The developed system can also find the rotation of both the 1D and QR barcodes, which gives the opportunity of rotating the detection accordingly which is shown to benefit the decoding process in a positive way. Both the detection and the rotation prediction shows real-time performance. Daniel Kold Hansen, Kamal Nasrollahi, Christoffer B. Rasmussen, Thomas B. Moeslund |
IJCCI | 4 |
| 2017 | R-FCN Object Detection Ensemble based on Object Resolution and Image QualityabstractObject detection can be difficult due to challenges such as variations in objects both inter- and intra-class. Additionally, variations can also be present between images. Based on this, research was conducted into creating an ensemble of Region-based Fully Convolutional Networks (R-FCN) object detectors. Ensemble strategies explored were firstly data sampling and selection and secondly combination strategies. Data sampling and selection aimed to create different subsets of data with respect to object size and image quality such that expert R-FCN ensemble members could be trained. Two combination strategies were explored for combining the individual member detections into an ensemble result, namely average and a weighted average. R-FCNs were trained and tested on the PASCAL VOC benchmark object detection dataset. Results proved positive with an increase in Average Precision (AP), compared to state-of-the-art similar systems, when ensemble members were combined appropriately. Christoffer B. Rasmussen, Kamal Nasrollahi, Thomas B. Moeslund |
IJCCI | 3 |
| 2017 | Locality regularized group sparse coding for action recognition
Mohammad Ali Bagheri, Qigang Gao, Sergio Escalera, Thomas B. Moeslund, Huamin Ren, Elham Etemad |
Comput. Vis. Image Underst. | 4 |
| 2017 | Computer Vision in Sports
Thomas B. Moeslund, Graham A. Thomas, Adrian Hilton 0001, Peter Carr 0001, Irfan A. Essa |
Comput. Vis. Image Underst. | 1 |
| 2017 | Computer vision for sports: Current applications and research topics
Graham A. Thomas, Rikke Gade, Thomas B. Moeslund, Peter Carr 0001, Adrian Hilton 0001 |
Comput. Vis. Image Underst. | 3 |
| 2017 | A new low-complexity patch-based image super-resolutionabstractIn this study, a novel single image super‐resolution (SR) method, which uses a generated dictionary from pairs of high‐resolution (HR) images and their corresponding low‐resolution (LR) representations, is proposed. First, HR and LR dictionaries are created by dividing HR and LR images into patches Afterwards, when performing SR, the distance between every patch of the input LR image and those of available LR patches in the LR dictionary are calculated. The minimum distance between the input LR patch and those in the LR dictionary is taken, and its counterpart from the HR dictionary will be passed through an illumination enhancement process resulting in consistency of illumination between neighbour patches. This process is applied to all patches of the LR image. Finally, in order to remove the blocking effect caused by merging the patches, an average of the obtained HR image and the interpolated image is calculated. Furthermore, it is shown that the stabe of dictionaries is reducible to a great degree. The speed of the system is improved by 62.5%. The quantitative and qualitative analyses of the experimental results show the superiority of the proposed technique over the conventional and state‐of‐the‐art methods. Pejman Rasti, Kamal Nasrollahi, Olga Orlova, Gert Tamberg, Cagri Ozcinar, Thomas B. Moeslund, Gholamreza Anbarjafari |
IET Comput. Vis. | 6 |
| 2016 | An in-depth study of sparse codes on abnormality detectionabstractSparse representation has been applied successfully in abnormal event detection, in which the baseline is to learn a dictionary accompanied by sparse codes. While much emphasis is put on discriminative dictionary construction, there are no comparative studies of sparse codes regarding abnormality detection. We present an in-depth study of two types of sparse codes solutions - greedy algorithms and convex L1-norm solutions - and their impact on abnormality detection performance. We also propose our framework of combining sparse codes with different detection methods. Our comparative experiments are carried out from various angles to better understand the applicability of sparse codes, including computation time, reconstruction error, sparsity, detection accuracy, and their performance combining various detection methods. The experiment results show that combining OMP codes with maximum coordinate detection could achieve state-of-the-art performance on the UCSD dataset [14]. Huamin Ren, Hong Pan 0001, Søren I. Olsen, Morten B. Jensen, Thomas B. Moeslund |
AVSS | 5 |
| 2016 | Task space HRI for cooperative mobile robots in fit-out operations inside ship superstructuresabstractIn the rising area of close human-robot collaboration in industrial scenarios, the human operator must be able to easily understand the intent of and data from the robot. Shipbuilding environments exhibit unique features, which make deployment of mobile robots both challenging, relevant, and interesting. One task that is still solely carried out manually today due to its complexity and high need for mobility is the fit-out operation stud welding. This paper presents the latest state-of-the-art developments in human-robot interaction (HRI) for robotic stud welding in large semi-structured manufacturing spaces. The welding itself is carried out autonomously by an autonomous industrial mobile manipulator (AIMM). A novel HRI is proposed, which employs projection mapping and an IMU device to enable intuitive and natural interaction with the robot. Task specific information is projected directly into task space as augmented reality using a projector mounted on the robot end-effector. The IMU device enables non-expert operators to program, verify, and reprogram the robot's task on-site in a ship superstructure. The usability of the system is tested in an extensive user test. It is concluded that non-experts after a short introduction are able to both modify a previous task and instruct and a new task using on average 1:01 and 1:16 minutes. Finally, the HRI has been implemented on a prototype robot and tested in an actual shipyard facility. The precision of the system, including operator inaccuracy, was evaluated to have a standard deviation of 3.6mm. Rasmus S. Andersen, Simon Bøgh, Thomas B. Moeslund, Ole Madsen |
RO-MAN | 3 |
| 2016 | Projecting robot intentions into human environmentsabstractTrained human co-workers can often easily predict each other's intentions based on prior experience. When collaborating with a robot coworker, however, intentions are hard or impossible to infer. This difficulty of mental introspection makes human-robot collaboration challenging and can lead to dangerous misunderstandings. In this paper, we present a novel, object-aware projection technique that allows robots to visualize task information and intentions on physical objects in the environment. The approach uses modern object tracking methods in order to display information at specific spatial locations taking into account the pose and shape of surrounding objects. As a result, a human co-worker can be informed in a timely manner about the safety of the workspace, the site of next robot manipulation tasks, and next subtasks to perform. A preliminary usability study compares the approach to collaboration approaches based on monitors and printed text. The study indicates that, on average, the user effectiveness and satisfaction is higher with the projection based approach. Rasmus S. Andersen, Ole Madsen, Thomas B. Moeslund, Heni Ben Amor |
RO-MAN | 3 |
| 2016 | Facial video-based detection of physical fatigue for maximal muscle activityabstractPhysical fatigue reveals the health condition of a person at, for example, health checkup, fitness assessment, or rehabilitation training. This study presents an efficient non‐contact system for detecting non‐localised physical fatigue from maximal muscle activity using facial videos acquired in a realistic environment with natural lighting where subjects were allowed to voluntarily move their head, change their facial expression, and vary their pose. The proposed method utilises a facial feature point tracking method by combining a ‘good feature to track’ and a ‘supervised descent method’ to address the challenges that originate from realistic scenario. A face quality assessment system was also incorporated in the proposed system to reduce erroneous results by discarding low quality faces that occurred in a video sequence due to problems in realistic lighting, head motion, and pose variation. Experimental results show that the proposed system outperforms video‐based existing system for physical fatigue detection. Mohammad A. Haque, Ramin Irani, Kamal Nasrollahi, Thomas B. Moeslund |
IET Comput. Vis. | 4 |
| 2016 | Multi-modal RGB-Depth-Thermal Human Body Segmentation
Cristina Palmero, Albert Clapés, Chris Bahnsen, Andreas Møgelmose, Thomas B. Moeslund, Sergio Escalera |
Int. J. Comput. Vis. | 5 |
| 2016 | Guest Editorial: Analysis and Retrieval of Events/Actions and Workflows in Video Streams
Anastasios Doulamis, Nikolaos D. Doulamis, Marco Bertini 0001, Jordi Gonzàlez 0001, Thomas B. Moeslund |
Multim. Tools Appl. | 5 |
| 2016 | Segmentation of RGB-D indoor scenes by stacking random forests and conditional random fields
Mikkel Thøgersen, Sergio Escalera, Jordi Gonzàlez 0001, Thomas B. Moeslund |
Pattern Recognit. Lett. | 4 |
| 2016 | Vision for Looking at Traffic Lights: Issues, Survey, and PerspectivesabstractThis paper presents the challenges that researchers must overcome in traffic light recognition (TLR) research and provides an overview of ongoing work. The aim is to elucidate which areas have been thoroughly researched and which have not, thereby uncovering opportunities for further improvement. An overview of the applied methods and noteworthy contributions from a wide range of recent papers is presented, along with the corresponding evaluation results. The evaluation of TLR systems is studied and discussed in depth, and we propose a common evaluation procedure, which will strengthen evaluation and ease comparison. To provide a shared basis for comparing TLR systems, we publish an extensive public data set based on footage from U.S. roads. The data set contains annotated video sequences, captured under varying light and weather conditions using a stereo camera. The data set, with its variety, size, and continuous sequences, should challenge current and future TLR systems. Morten B. Jensen, Mark P. Philipsen, Andreas Møgelmose, Thomas B. Moeslund, Mohan M. Trivedi |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2015 | Detecting road user actions in traffic intersections using RGB and thermal videoabstractThis paper investigates the development of a watch-dog system that detects a subset of road user actions in traffic intersections. Footage of the intersections is captured with RGB and thermal cameras to ensure that the road is visible round-the-clock even in difficult weather conditions. The watch-dog system consists of several, cascaded detectors which are capable of detecting specific road user actions, such as Right Turning Vehicles, Left Turning Vehicles, and Straight Going Cyclists. Experimental results on 4 hours of video from 3 different intersections show good performance and a precision above 0.93 when detecting turning vehicles. The use of both RGB and thermal video generally results in better performance, providing overall stability when observing the road. Chris Bahnsen, Thomas B. Moeslund |
AVSS | 2 |
| 2015 | Ongoing work on traffic lights: Detection and evaluationabstractResearch in traffic light recognition (TLR) has stagnated compared to related computer vision areas, such as pedestrian detection and and traffic sign recognition. We focus on the detection sub-problem, since this is the most challenging problem and solving this is the key to a successful TLR system. This is done by looking at four detectors from different author groups and their reported results. From surveying existing work it is clear that currently evaluation is limited primarily to small local datasets. In order to provide a common basis for future comparison of TLR research an extensive public database is collected based on footage from US roads. The database consists of continuous test and training video sequences, totaling 46,418 frames and 112,971 annotated traffic lights. The sequences are captured by a stereo camera mounted on the roof of a vehicle driving under both night and day conditions with varying light and weather. Mark P. Philipsen, Morten B. Jensen, Mohan M. Trivedi, Andreas Møgelmose, Thomas B. Moeslund |
AVSS | 5 |
| 2015 | Unsupervised Behavior-Specific Dictionary Learning for Abnormal Event DetectionabstractAbnormal event detection has been a challenge due to the lack of complete normal information in the training data and the volatility of the definitions of both normality and abnormality. Recent research applying sparse representation has shown its effectiveness in the expression of normal patterns. Despite progress in this area, the relationship of atoms within the dictionary is commonly neglected, thereafter anomalies which are detected based on reconstruction error could brings high false alarm - noise or infrequent normal visual features could be wrongly detected as anomalies, especially when the training data is only a small proportion of the surveillance data. Therefore, we propose behavior-specific dictionaries (BSD) through unsupervised learning, pursuing atoms from the same type of behavior to represent one behavior dictionary. To further improve the dictionary by introducing information from potential infrequent normal patterns, we refine the dictionary by searching ‘missed atoms’ that have compact coefficients. Experimental results show that our BSD algorithm outperforms state-of-the-art dictionaries in abnormal event detection on the public UCSD dataset. Moreover, BSD has less false alarms compared to state-of-the-art dictionaries especially when the training set is small, which is demonstrated on Anomaly Stairs dataset. Huamin Ren, Weifeng Liu 0002, Søren I. Olsen, Sergio Escalera, Thomas B. Moeslund |
BMVC | 5 |
| 2015 | EREL: Extremal regions of extremum levelsabstractExtremal Regions of Extremum Levels (EREL) are regions detected from a set of all extremal regions of an image. Maximally Stable Extremal Regions (MSER) which is a novel affine covariant region detector, detects regions from a same set of extremal regions as well. Although MSER results in regions with almost high repeatability, it is heavily dependent on the union-find approach which is a fairly complicated algorithm, and should be completed sequentially. Furthermore, it detects regions with low repeatability under the blur transformations. The reason for the latter shortcoming is the absence of boundaries information in stability criterion. To tackle these problems we propose to employ prior information about boundaries of regions, which results in a novel region detector algorithm that not only outperforms MSER, but avoids the MSER's rather complicated steps of union-finding. To achieve that, we introduce Maxima of Gradient Magnitudes (MGMs) and use them to find handful of Extremum Levels (ELs). The chosen ELs are then scanned to detect their Extremal Regions (ER). The proposed algorithm which is called Extremal Regions of Extremum Levels (EREL) has been tested on the public benchmark dataset of Mikolajczyk [1]. Our experimental evaluations illustrate that, in many cases EREL achieves higher repeatability scores than MSER even for very low overlap errors. Mehdi Faraji, Jamshid Shanbehzadeh, Kamal Nasrollahi, Thomas B. Moeslund |
ICIP | 4 |
| 2015 | Trajectory analysis and prediction for improved pedestrian safety: Integrated framework and evaluationsabstractThis paper presents a monocular and purely vision based pedestrian trajectory tracking and prediction framework with integrated map-based hazard inference. In Advanced Driver Assistance systems research, a lot of effort has been put into pedestrian detection over the last decade, and several pedestrian detection systems are indeed showing impressive results. Considerably less effort has been put into processing the detections further. We present a tracking system for pedestrians, which based on detection bounding boxes tracks pedestrians and is able to predict their positions in the near future. The tracking system is combined with a module which, based on the car's GPS position acquires a map and uses the road information in the map to know where the car can drive. Then the system warns the driver about pedestrians at risk, by combining the information about hazardous areas for pedestrians with a probabilistic position prediction for all observed pedestrians. Andreas Møgelmose, Mohan M. Trivedi, Thomas B. Moeslund |
Intelligent Vehicles Symposium | 3 |
| 2015 | Day and night-time drive analysis using stereo vision for naturalistic driving studiesabstractIn order to understand dangerous situations in the driving environment, naturalistic driving studies (NDS) are conducted by collecting and analyzing data from sensors looking inside and outside of the car. Manually processing the overwhelming amounts of data that are generated in such studies is very comprehensive. We propose a method for automatic data reduction for NDS based on stereo vision vehicle detection and tracking during day- and nighttime. The developed system can automatically register five NDS events, mainly related to intersections, from an existing NDS dictionary. We propose a new drive event which takes advantage of the extra dimension provided by stereo vision. In total, six drive events are selected on the basis of them being problematic to detect automatically using conventional monocular computer vision approaches. The proposed system is evaluated on day-and nighttime data, resulting in drive analysis report. The proposed system reach an overall precision of 0.78 and an overall recall of 0.72. Mark P. Philipsen, Morten B. Jensen, Ravi Kumar Satzoda, Mohan M. Trivedi, Andreas Møgelmose, Thomas B. Moeslund |
Intelligent Vehicles Symposium | 6 |
| 2015 | Circular Hough Transform and Local Circularity Measure for Weight Estimation of a Graph-Cut Based Wood Stack MeasurementabstractOne of the time consuming tasks in the timber industry is the manually measurement of features of wood stacks. Such features include, but are not limited to, the number of the logs in a stack, their diameters distribution, and their volumes. Computer vision techniques have recently been used for solving this real-world industrial application. Such techniques are facing many challenges as the task is usually performed in outdoor, uncontrolled, environments. Furthermore, the logs can vary in texture and they can be occluded by different obstacles. These all make the segmentation of the wood logs a difficult task. Graph-cut has shown to be good enough for such a segmentation. However, it is hard to find proper graph weights. This is exactly the contribution of this paper to propose a method for setting the weights of the graph. To do so, we use Circular Hough Transform (CHT) for obtaining information about the fore and background regions of a stack image, and then use this together with a Local Circularity Measure (LCM) to modify the weights of the graph to segment the wood logs from the rest of the image. We further improve the segmentation by separating overlapping logs. These segmented wood logs are finally scaled and used to acquire the necessary wood stack measurements in real-world scale (in cm). The proposed system, which works automatically, has been tested on two different datasets, containing real outdoor images of logs which vary in shapes and sizes. The experimental results show that the proposed approach not only achieves the same results as the state-of-the-art systems, it produces more stable results. Bo Galsgaard, Dennis H. Lundtoft, Ivan A. Nikolov, Kamal Nasrollahi, Thomas B. Moeslund |
WACV | 5 |
| 2015 | Quality-Aware Estimation of Facial Landmarks in Video SequencesabstractFace alignment in video is a primitive step for facial image analysis. The accuracy of the alignment greatly depends on the quality of the face image in the video frames and low quality faces are proven to cause erroneous alignment. Thus, this paper proposes a system for quality aware face alignment by using a Supervised Decent Method (SDM) along with a motion based forward extrapolation method. The proposed system first extracts faces from video frames. Then, it employs a face quality assessment technique to measure the face quality. If the face quality is high, the proposed system uses SDM for facial landmark detection. If the face quality is low the proposed system corrects the facial landmarks that are detected by SDM. Depending upon the face velocity in consecutive video frames and face quality measure, two algorithms are proposed for correction of landmarks in low quality faces by using an extrapolation polynomial. Experimental results illustrate the competency of the proposed method while comparing with the state-of-the art methods including an SDM-based method (from CVPR-2013) and a very recent method (from CVPR-2014) that uses parallel cascade of linear regression (Par-CLR). Mohammad A. Haque, Kamal Nasrollahi, Thomas B. Moeslund |
WACV | 3 |
| 2015 | Chromatic shadow detection and tracking for moving foreground segmentation
Ivan Huerta Casado, Michael B. Holte, Thomas B. Moeslund, Jordi Gonzàlez 0001 |
Image Vis. Comput. | 3 |
| 2015 | Special issue on "Soft Biometrics"
Paulo Lobato Correia, Abdenour Hadid, Thomas B. Moeslund |
Pattern Recognit. Lett. | 3 |
| 2015 | On soft biometrics
Mark S. Nixon, Paulo Lobato Correia, Kamal Nasrollahi, Thomas B. Moeslund, Abdenour Hadid, Massimo Tistarelli |
Pattern Recognit. Lett. | 4 |
| 2015 | Extremal Regions Detection Guided by Maxima of Gradient MagnitudeabstractA problem of computer vision applications is to detect regions of interest under different imaging conditions. The state-of-the-art maximally stable extremal regions (MSERs) detects affine covariant regions by applying all possible thresholds on the input image, and through three main steps including: (1) making a component tree of extremal regions' evolution; (2) obtaining region stability criterion; and (3) cleaning up. The MSER performs very well, but, it does not consider any information about the boundaries of the regions, which are important for detecting repeatable extremal regions. We have shown in this paper that employing prior information about boundaries of regions results in a novel region detector algorithm that not only outperforms MSER, but avoids the MSER's rather complicated steps of enumeration and the cleaning up. To employ the information about the region boundaries, we introduce maxima of gradient magnitudes (MGMs) which are shown to be points that are mostly around the boundaries of the regions. Having found the MGMs, the method obtains a global criterion for each level of the input image which is used to find extremum levels (ELs). The found ELs are then used to detect extremal regions. The proposed algorithm which is called extremal regions of extremum levels (EREL) has been tested on the public benchmark data set of Mikolajczyk. The obtained experimental results show that the inclusion of region boundaries through MGMs, results in a detector that detects regions with high repeatability scores and is more robust against noise compared with MSER. Mehdi Faraji, Jamshid Shanbehzadeh, Kamal Nasrollahi, Thomas B. Moeslund |
IEEE Trans. Image Process. | 4 |
| 2014 | Abnormal event detection using local sparse representationabstractWe propose to detect abnormal events via a sparse subspace clustering algorithm. Unlike most existing approaches, which search for optimized normal bases and detect abnormality based on least square error or reconstruction error from the learned normal patterns, we propose an abnormality measurement based on the difference between the normal space and local space. Specifically, we provide a reasonable normal bases through repeated K spectral clustering. Then for each testing feature we first use temporal neighbors to form a local space. An abnormal event is found if any abnormal feature is found that satisfies: the distance between its local space and the normal space is large. We evaluate our method on two public benchmark datasets: UCSD and Subway Entrance datasets. The comparison to the state-of-the-art methods validate our method's effectiveness. Huamin Ren, Thomas B. Moeslund |
AVSS | 2 |
| 2014 | Contactless measurement of muscles fatigue by tracking facial feature points in a videoabstractPhysical exercise may result in muscle tiredness which is known as muscle fatigue. This occurs when the muscles cannot exert normal force, or when more than normal effort is required. Fatigue is a vital sign, for example, for therapists to assess their patient's progress or to change their exercises when the level of the fatigue might be dangerous for the patients. The current technology for measuring tiredness, like Electromyography (EMG), requires installing some sensors on the body. In some applications, like remote patient monitoring, this however might not be possible. To deal with such cases, in this paper we present a contactless method based on computer vision techniques to measure tiredness by detecting, tracking, and analyzing some facial feature points during the exercise. Experimental results on several test subjects and comparing them against ground truth data show that the proposed system can properly find the temporal point of tiredness of the muscles when the test subjects are doing physical exercises. Ramin Irani, Kamal Nasrollahi, Thomas B. Moeslund |
ICIP | 3 |
| 2014 | Stereoscopic roadside curb height measurement using V-disparityabstractManaging road assets, such as roadside curbs, is one of the interests of municipalities. As an interesting application of computer vision, this paper proposes a system for automated measurement of the height of the roadside curbs. The developed system uses the spatial information available in the disparity image obtained from a stereo setup. Data about the geometry of the scene is extracted in the form of a row-wise histogram of the disparity map. From parameterizing the two strongest lines, each pixel can be labeled as belonging to one plane, either ground, sidewalk or curb candidates. Experimental results show that the system can measure the height of the roadside curb with good accuracy and precision. Florin Octavian Matu, Iskren Vlaykov, Mikkel Thøgersen, Kamal Nasrollahi, Thomas B. Moeslund |
ICMV | 5 |
| 2014 | RGB-D-T Based Face RecognitionabstractFacial images are of critical importance in many real-world applications from gaming to surveillance. The current literature on facial image analysis, from face detection to face and facial expression recognition, are mainly performed in either RGB, Depth (D), or both of these modalities. But, such analyzes have rarely included Thermal (T) modality. This paper paves the way for performing such facial analyzes using synchronized RGB-D-T facial images by introducing a database of 51 persons including facial images of different rotations, illuminations, and expressions. Furthermore, a face recognition algorithm has been developed to use these images. The experimental results show that face recognition using such three modalities provides better results compared to face recognition in any of such modalities in most of the cases. Olegs Nikisins, Kamal Nasrollahi, Modris Greitans, Thomas B. Moeslund |
ICPR | 4 |
| 2014 | Adaptive Non-local Means for Cost Aggregation in a Local Disparity Estimation AlgorithmabstractThe overall method used for determining disparity in a stereo setup is a widely recognized framework consisting of four steps of cost space computation, cost aggregation, disparity selection, and post-processing. In this paper a cost aggregation approach for a typical local disparity estimation method is introduced. The method introduced is built on top of an existing method called Adaptive Support-Weight using this known framework. The introduced method improves Adaptive Support-Weight method by utilizing a larger amount of data inspired by the method of Non-Local Means. The extra data is handled in a way that tries to preserve the location of depth discontinuities in the final disparity map. Experimental results on Middlebury benchmark database show that the proposed method suffers from less artifacts compared to state-of-the-art disparity estimation methods. Casper Pedersen, Kamal Nasrollahi, Thomas B. Moeslund |
ICPR | 3 |
| 2014 | Thermal cameras and applications: a survey
Rikke Gade, Thomas B. Moeslund |
Mach. Vis. Appl. | 2 |
| 2014 | Special issue on Multimedia Event Detection
Thomas B. Moeslund, Omar Javed, Yu-Gang Jiang 0001, R. Manmatha |
Mach. Vis. Appl. | 1 |
| 2014 | Super-resolution: a comprehensive survey
Kamal Nasrollahi, Thomas B. Moeslund |
Mach. Vis. Appl. | 2 |
| 2013 | Real-time acquisition of high quality face sequences from an active pan-tilt-zoom cameraabstractTraditional still camera-based facial image acquisition systems in surveillance applications produce low quality face images. This is mainly due to the distance between the camera and subjects of interest. Furthermore, people in such videos usually move around, change their head poses, and facial expressions. Moreover, the imaging conditions like illumination, occlusion, and noise may change. These all aggregate the quality of most of the detected face images in terms of measures like resolution, pose, brightness, and sharpness. To deal with these problems this paper presents an active camera-based realtime high-quality face image acquisition system, which utilizes pan-tilt-zoom parameters of a camera to focus on a human face in a scene and employs a face quality assessment method to log the best quality faces from the captured frames. The system consists of four modules: face detection, camera control, face tracking, and face quality assessment before logging. Experimental results show that the proposed system can effectively log the high quality faces from the active camera in real-time (an average of 61.74ms was spent per frame) with an accuracy of 85.27% compared to human annotated data. Mohammad A. Haque, Kamal Nasrollahi, Thomas B. Moeslund |
AVSS | 3 |
| 2013 | Tamper detection for active surveillance systemsabstractIf surveillance data are corrupted they are of no use to neither manually post-investigation nor automatic video analysis. It is therefore critical to automatically be able to detect tampering events such as defocusing, occlusion and displacement. In this work we for the first time address this important problem for an active camera. We detect these events by first comparing the incoming frames to a background model using features relevant for the three different tampering types. Individual detectors are then developed capable of monitoring long video sequences and indicating the occurrence of different tampering events. In order to assess the developed methods we have collected a large data set, which contains sequences from different active cameras at different scenarios. We evaluate our system on these data and the results are encouraging with a very high detecting rate and relatively few false positives. Theodore Tsesmelis, Lars Christensen, Preben Fihl, Thomas B. Moeslund |
AVSS | 4 |
| 2013 | Are Haar-Like Rectangular Features for Biometric Recognition Reducible?
Kamal Nasrollahi, Thomas B. Moeslund |
CIARP (2) | 2 |
| 2013 | Long-Term Occupancy Analysis Using Graph-Based Optimisation in Thermal ImageryabstractThis paper presents a robust occupancy analysis system for thermal imaging. Reliable detection of people is very hard in crowded scenes, due to occlusions and segmentation problems. We therefore propose a framework that optimises the occupancy analysis over long periods by including information on the transition in occupancy, when people enter or leave the monitored area. In stable periods, with no activity close to the borders, people are detected and counted which contributes to a weighted histogram. When activity close to the border is detected, local tracking is applied in order to identify a crossing. After a full sequence, the number of people during all periods are estimated using a probabilistic graph search optimisation. The system is tested on a total of 51,000 frames, captured in sports arenas. The mean error for a 30-minute period containing 3-13 people is 4.44 %, which is a half of the error percentage optained by detection only, and better than the results of comparable work. The framework is also tested on a public available dataset from an outdoor scene, which proves the generality of the method. Rikke Gade, Anders Jørgensen, Thomas B. Moeslund |
CVPR | 3 |
| 2013 | Haar-like features for robust real-time face recognitionabstractFace recognition is still a very challenging task when the input face image is noisy, occluded by some obstacles, of very low-resolution, not facing the camera, and not properly illuminated. These problems make the feature extraction and consequently the face recognition system unstable. The proposed system in this paper introduces the novel idea of using Haar-like features, which have commonly been used for object detection, along with a probabilistic classifier for face recognition. The proposed system is simple, real-time, effective and robust against most of the mentioned problems. Experimental results on public databases show that the proposed system indeed outperforms the state-of-the-art face recognition systems. Kamal Nasrollahi, Thomas B. Moeslund |
ICIP | 2 |
| 2013 | Action recognition using salient neighboring histogramsabstractCombining spatio-temporal interest points with Bag-of-Words models achieves state-of-the-art performance in action recognition. However, existing methods based on “bag-of-words” models either are too local to capture the variance in space/time or fail to solve the ambiguity problem in spatial and temporal dimensions. Instead, we propose a salient vocabulary construction algorithm to select visual words from a global point of view, and form compact descriptors to represent discriminative histograms in the neighborhoods. Those salient neighboring histograms are then trained to model different actions. Our approach yields a competitive result on the KTH dataset compare to state-of-the-art methods. On the more challenging UCF Sports dataset, we obtain 95.21%, which is approximately 4% better than the current best published results. Huamin Ren, Thomas B. Moeslund |
ICIP | 2 |
| 2013 | 4th ACM/IEEE ARTEMIS 2013 international workshop on analysis and retrieval of tracked events and motion in imagery streamsabstractIn this paper, we give a short summary of the papers proposed in ACMARTEMIS 2013 which is held in Barcelona Spain in conjunction with ACM Multimedia. The workshop handles the areas of features analysis both at low and high level for efficient events detection, retrieval of multimedia events and objects and video synchronization issues and also events and behavior recognition from visual data. All papers were classified into three session of a single track workshop. The first session named "Video Features and Scene Analysis" includes articles that handle low level and high level visual analysis appropriate for event detection. The second session entitled "Retrieval of Multimedia Objects/Events" applies schemes for media data retrieval and video synchronization. Finally the third session "Analysis of Visual Events" describes algorithms for detecting actions, behaviors and events in complex visual scenes. Anastasios Doulamis, Nikolaos D. Doulamis, Marco Bertini 0001, Jordi Gonzàlez 0001, Thomas B. Moeslund |
ACM Multimedia | 5 |
| 2013 | Part-Based Pedestrian Detection and Feature-Based Tracking for Driver Assistance: Real-Time, Robust Algorithms, and EvaluationabstractDetecting pedestrians is still a challenging task for automotive vision systems due to the extreme variability of targets, lighting conditions, occlusion, and high-speed vehicle motion. Much research has been focused on this problem in the last ten years and detectors based on classifiers have gained a special place among the different approaches presented. This paper presents a state-of-the-art pedestrian detection system based on a two-stage classifier. Candidates are extracted with a Haar cascade classifier trained with the Daimler Detection Benchmark data set and then validated through a part-based histogram-of-oriented-gradient (HOG) classifier with the aim of lowering the number of false positives. The surviving candidates are then filtered with feature-based tracking to enhance the recognition robustness and improve the results' stability. The system has been implemented on a prototype vehicle and offers high performance in terms of several metrics, such as detection rate, false positives per hour, and frame rate. The novelty of this system relies on the combination of a HOG part-based approach, tracking based on a specific optimized feature, and porting on a real prototype. Antonio Prioletti, Andreas Møgelmose, Paolo Grisleri, Mohan M. Trivedi, Alberto Broggi, Thomas B. Moeslund |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2012 | 3D human pose estimation using 2D body part detectors
Adela Barbulescu, Wenjuan Gong, Jordi Gonzàlez 0001, Thomas B. Moeslund, F. Xavier Roca |
ICPR | 4 |
| 2012 | Learning to detect traffic signs: Comparative evaluation of synthetic and real-world datasets
Andreas Møgelmose, Mohan M. Trivedi, Thomas B. Moeslund |
ICPR | 3 |
| 2012 | Controlling urban lighting by human motion patterns results from a full scale experimentabstractThis paper presents a full-scale experiment investigating the use of human motion intensities as input for interactive illumination of a town square in the city of Aalborg in Denmark. As illuminators sixteen 3.5 meter high RGB LED lamps were used. The activity on the square was monitored by three thermal cameras and analysed by computer vision software from which motion intensity maps and peoples trajectories were estimated and used as input to control the interactive illumination. The paper introduces a 2-layered interactive light strategy addressing ambient and effect illumination criteria totally four light scenarios were designed and tested. The result shows that in general people immersed in the street lighting did not notice that the light changed according to their presence or actions, but people watching from the edge of the square noticed the interaction between the illumination and the immersed persons. The experiment also demonstrated that interactive can give significant power savings. In the current experiment there was a difference of 92% between the most and less energy consuming light scenario. Esben Skouboe Poulsen, Hans Jørgen Andersen, Ole B. Jensen, Rikke Gade, Tobias Thyrrestrup, Thomas B. Moeslund |
ACM Multimedia | 6 |
| 2012 | Selective spatio-temporal interest points
Bhaskar Chakraborty, Michael B. Holte, Thomas B. Moeslund, Jordi Gonzàlez 0001 |
Comput. Vis. Image Underst. | 3 |
| 2012 | Semantic Understanding of Human Behaviors in Image Sequences: From video-surveillance to video-hermeneutics
Jordi Gonzàlez 0001, Thomas B. Moeslund, Liang Wang 0001 |
Comput. Vis. Image Underst. | 2 |
| 2012 | Vision-Based Traffic Sign Detection and Analysis for Intelligent Driver Assistance Systems: Perspectives and SurveyabstractIn this paper, we provide a survey of the traffic sign detection literature, detailing detection systems for traffic sign recognition (TSR) for driver assistance. We separately describe the contributions of recent works to the various stages inherent in traffic sign detection: segmentation, feature extraction, and final sign detection. While TSR is a well-established research area, we highlight open research issues in the literature, including a dearth of use of publicly available image databases and the over-representation of European traffic signs. Furthermore, we discuss future directions of TSR research, including the integration of context and localization. We also introduce a new public database containing U.S. traffic signs. Andreas Møgelmose, Mohan M. Trivedi, Thomas B. Moeslund |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2011 | A selective spatio-temporal interest point detector for human action recognition in complex scenesabstractRecent progress in the field of human action recognition points towards the use of Spatio-Temporal Interest Points (STIPs) for local descriptor-based recognition strategies. In this paper we present a new approach for STIP detection by applying surround suppression combined with local and temporal constraints. Our method is significantly different from existing STIP detectors and improves the performance by detecting more repeatable, stable and distinctive STIPs for human actors, while suppressing unwanted background STIPs. For action representation we use a bag-of-visual words (BoV) model of local N-jet features to build a vocabulary of visual-words. To this end, we introduce a novel vocabulary building strategy by combining spatial pyramid and vocabulary compression techniques, resulting in improved performance and efficiency. Action class specific Support Vector Machine (SVM) classifiers are trained for categorization of human actions. A comprehensive set of experiments on existing benchmark datasets, and more challenging datasets of complex scenes, validate our approach and show state-of-the-art performance. Bhaskar Chakraborty, Michael B. Holte, Thomas B. Moeslund, Jordi Gonzàlez 0001, F. Xavier Roca |
ICCV | 3 |
| 2011 | Extracting a Good Quality Frontal Face Image From a Low-Resolution Video SequenceabstractFeeding low-resolution and low-quality images, from inexpensive surveillance cameras, to systems like, e.g., face recognition, produces erroneous and unstable results. Therefore, there is a need for a mechanism to bridge the gap between on one hand low-resolution and low-quality images and on the other hand facial analysis systems. The proposed system in this paper deals with exactly this problem. Our approach is to apply a reconstruction-based super-resolution algorithm. Such an algorithm, however, has two main problems: first, it requires relatively similar images with not too much noise and second is that its improvement factor is limited by a factor close to two. To deal with the first problem we introduce a three-step approach, which produces a face-log containing images of similar frontal faces of the highest possible quality. To deal with the second problem, limited improvement factor, we use a learning-based super-resolution algorithm applied to the result of the reconstruction-based part to improve the quality by another factor of two. This results in an improvement factor of four for the entire system. The proposed system has been tested on 122 low-resolution sequences from two different databases. The experimental results show that the proposed system can indeed produce a high-resolution and good quality frontal face image from low-resolution video sequences. Kamal Nasrollahi, Thomas B. Moeslund |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Pose Estimation of Interacting People using Pictorial StructuresabstractPose estimation of people have had great progress in recent years but so far research has dealt with single persons.In this paper we address some of the challenges that arise when doing pose estimation of interacting people. We build on the pictorial structures framework and make important contributions by combining color-based appearance and edge information using a measure of the local quality of the appearance feature. In this way we not only combine the two types of features but dynamically find the optimal weighting of them. We further enable the method to handle occlusions by searching a foreground mask for possible occluded body parts and then applying extra strong kinematic constraints to find the true occluded body parts. The effect of applying our two contributions are show through both qualitative and quantitative tests and show a clear improvement on the ability to correctly localize body parts. Preben Fihl, Thomas B. Moeslund |
AVSS | 2 |
| 2010 | Improving stereo camera depth measurements and benefiting from intermediate resultsabstractThis paper presents a method for improving the disparity values obtained with a stereo camera by applying an iconic Kalman filter and known ego-motion. The improvements are demonstrated in an application of determining the free space in a scene as viewed by a vehicle-mounted camera. Using disparity maps from a stereo camera and known camera motion, the disparity maps are first filtered by an iconic Kalman filter, operating on each pixel individually, thereby reducing variance and increasing the density of the filtered disparity map. Then, a stochastic occupancy grid is calculated from the filtered disparity map, providing a top-down view of the scene where the uncertainty of disparity measurements are taken into account. These occupancy grids are segmented to indicate a maximum depth free of obstacles, enabling the marking of free space in the accompanying intensity image. Even without motion of the camera, the quality of the disparity map is increased significantly. Applications of the intermediate results are discussed, enabling features such as motion detection and quantifying the certainty of the measurements. The evaluation shows significant improvement in disparity variance and disparity map density, and consequently an improvement in the application of marking free space. Carsten Høilund, Thomas B. Moeslund, Claus B. Madsen, Mohan M. Trivedi |
Intelligent Vehicles Symposium | 2 |
| 2010 | View-invariant gesture recognition using 3D optical flow and harmonic motion context
Michael B. Holte, Thomas B. Moeslund, Preben Fihl |
Comput. Vis. Image Underst. | 2 |
| 2009 | Detection and removal of chromatic moving shadows in surveillance scenariosabstractSegmentation in the surveillance domain has to deal with shadows to avoid distortions when detecting moving objects. Most segmentation approaches dealing with shadow detection are typically restricted to penumbra shadows. Therefore, such techniques cannot cope well with umbra shadows. Consequently, umbra shadows are usually detected as part of moving objects. In this paper we present a novel technique based on gradient and colour models for separating chromatic moving cast shadows from detected moving objects. Firstly, both a chromatic invariant colour cone model and an invariant gradient model are built to perform automatic segmentation while detecting potential shadows. In a second step, regions corresponding to potential shadows are grouped by considering “a bluish effect” and an edge partitioning. Lastly, (i) temporal similarities between textures and (ii) spatial similarities between chrominance angle and brightness distortions are analysed for all potential shadow regions in order to finally identify umbra shadows. Unlike other approaches, our method does not make any a-priori assumptions about camera location, surface geometries, surface textures, shapes and types of shadows, objects, and background. Experimental results show the performance and accuracy of our approach in different shadowed materials and illumination conditions. Ivan Huerta Casado, Michael B. Holte, Thomas B. Moeslund, Jordi Gonzàlez 0001 |
ICCV | 3 |
| 2008 | AVSS 2008 Commentary Paper for: "Tracking People in Crowds by a Part Matching Approach"abstractOne of the major problems remaining in tracking is occlusion handling. This paper presents a system for exactly this. A human model is defined and each body part is represented by a number of features. For each new image in a sequence a foreground mask is obtained and compared to each body part of a predicted model of the human. Head detection is utilized to stabilize the tracking. The system is tested on two sequences form two public databases. Thomas B. Moeslund |
AVSS | 1 |
| 2008 | View invariant gesture recognition using 3D motion primitivesabstractThis paper presents a method for automatic recognition of human gestures. The method works with 3D image data from a range camera to achieve invariance to viewpoint. The recognition is based solely on motion from characteristic instances of the gestures. These instances are denoted 3D motion primitives. The method extracts 3D motion from range images and represent the motion from each input frame in a view invariant manner using harmonic shape context. The harmonic shape context is classified as a 3D motion primitive. A sequence of input frames results in a set of primitives that are classified as a gesture using a probabilistic edit distance method. The system has been trained on frontal images (0deg camera rotation) and tested on 240 video sequences from 0deg and 45deg. An overall recognition rate of 82.9% is achieved. The recognition rate is independent of the viewpoint which shows that the method is indeed view invariant. Michael B. Holte, Thomas B. Moeslund |
ICASSP | 2 |
| 2007 | Classification of gait types based on the duty-factorabstractThis paper deals with classification of human gait types based on the notion that different gait types are in fact different types of locomotion, i.e., running is not simply walking done faster. We present the duty-factor, which is a descriptor based on this notion. The duty-factor is independent on the speed of the human, the cameras setup etc. and hence a robust descriptor for gait classification. The duty-factor is basically a matter of measuring the ground support of the feet with respect to the stride. We estimate this by comparing the incoming silhouettes to a database of silhouettes with known ground support. Silhouettes are extracted using the codebook method and represented using shape contexts. The matching with database silhouettes is done using the Hungarian method. While manually estimated duty-factors show a clear classification the presented system contains misclassifications due to silhouette noise and ambiguities in the database silhouettes. Preben Fihl, Thomas B. Moeslund |
AVSS | 2 |
| 2006 | A survey of advances in vision-based human motion capture and analysis
Thomas B. Moeslund, Adrian Hilton 0001, Volker Krüger |
Comput. Vis. Image Underst. | 1 |
| 2003 | Computer Vision Based Head Tracking from Re-configurable 2D Markers for ARabstractThis paper presents a computer vision based head tracking system for augmented reality. A camera is attached to a head mounted display and used to track markers in the user's field of view. These markers are movable on a table and used as interface with virtual objects. Furthermore, they are used as landmarks to track the user's viewpoint and viewing direction by a homography based camera pose estimation algorithm. By integrating this computer vision tracker with a commercially available InterSense tracker, we take the advantages of the small jitter of the former one without losing the tracking speed of the later one. For static and slow head motion the system has less than 0.3mm RMS of position jitter and 0.165 degrees RMS of orientation jitter. Moritz Störring, Thomas B. Moeslund, Claus B. Madsen, Erik Granum |
ISMAR | 3 |
| 2003 | Modelling and estimating the pose of a human arm
Thomas B. Moeslund, Erik Granum |
Mach. Vis. Appl. | 1 |
| 2001 | A Survey of Computer Vision-Based Human Motion Capture
Thomas B. Moeslund, Erik Granum |
Comput. Vis. Image Underst. | 1 |
| 2000 | Multiple Cues used in Model-Based Human Motion CaptureabstractHuman motion capture has lately been the object of much attention due to commercial interests. A "touch-free" computer vision solution to the problem is desirable to avoid the intrusiveness of standard capture devices. The object to be monitored is known a priori which suggests the inclusion of a human model in the capture process. We use a model-based approach known as the analysis-by-synthesis approach. This approach is powerful but has a problem with its potential huge search space. Using multiple cues we reduce the search space by introducing constraints through the 3D locations of salient points and a silhouette of the subject. Both data types are relatively easy to derive and only require limited computational effort so the approach remains suitable for real-time applications. The approach is tested on 3D movements of a human arm and the results show that we successfully can estimate the pose of the arm using the reduced search space. Thomas B. Moeslund, Erik Granum |
FG | 1 |
| 1998 | Enhancing a WIMP based interface with speech, gaze tracking and agentsabstractThis paper describes an attempt to enhance a windows based (WIMP - Windows Icon Menu Pointer) environment. The goal is to establish whether user interaction on the common desktop PC can be augmented by adding new modalities to the WIMP interface, thus bridging the gap between todays interaction patterns and future interfaces comprising e.g. advanced conversational capabilities, VR technology, etc. Lau Bakman, Mads Blidegn, Martin Wittrup, Lars Bo Larsen, Thomas B. Moeslund |
ICSLP | 5 |
| 1998 | The intellimedia workbench - a generic environment for multimodal systemsabstractThe present paper presents a generic environment for intelligent multi media applications, denoted "The Intellimedia WorkBench ". The aim of the workbench is to facilitate development and research within the field of multi modal user interaction. Physically it is a table with various devices mounted above and around. These include: A camera and a laser projector mounted above the workbench, a microphone array mounted on the walls of the room, a speech recogniser and a speech synthesiser. The camera is attached to a vision system capable of locating various objects placed on the workbench. The paper presents two applications utilising the workbench. One is a campus information system, allowing the user to ask for directions within a part of the university campus. The second application is a pool trainer, intended to provide guidance to novice players. Introduction 1. INTRODUCTION Background. Driven by the current move towards multi modal interaction an activity was initiated at Aalb... Tom Brøndsted, Lars Bo Larsen, Michael Manthey, Paul Mc Kevitt, Thomas B. Moeslund, Kristian G. Olesen |
ICSLP | 5 |