VLDB 2026 Research / reviewers in the wild / expert
Andrew D. Bagdanov
dblp:64/3935 · also Andrew David Bagdanov, Andy Bagdanov
· DBLP profile ↗
83ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0001-6408-7043ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 57 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 5 first-author · 14 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-authorHuman-computer interaction and ubiquitous computing · 5 · 5 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | No MoCap Needed: Post-Training Motion Diffusion Models with Reinforcement Learning using Only Textual PromptsabstractDiffusion models have recently advanced human motion generation, producing realistic and diverse animations from textual prompts. However, adapting these models to unseen actions or styles typically requires additional motion capture data and full retraining, which is costly and difficult to scale. We propose a post-training framework based on Reinforcement Learning that fine-tunes pretrained motion diffusion models using only textual prompts, without requiring any motion ground truth. Our approach employs a pretrained text–motion retrieval network as a reward signal and optimizes the diffusion policy with Denoising Diffusion Policy Optimization, effectively shifting the model’s generative distribution toward the target domain without relying on paired motion data. We evaluate our method on cross-dataset adaptation and leave-one-out motion experiments using the HumanML3D and KIT-ML datasets across both latent- and joint-space diffusion architectures. Results from quantitative metrics and user studies show that our approach consistently improves the quality and diversity of generated motions, while preserving performance on the original distribution. Our approach is a flexible, data-efficient, and privacy-preserving solution for motion adaptation. Code is available on GitHub at https://github.com/miccunifi/NoMoCapNeeded. Girolamo Macaluso, Lorenzo Mandelli, Mirko Bicchierai, Stefano Berretti, Andrew D. Bagdanov |
WACV | 5 |
| 2025 | NTRL: Encounter Generation via Reinforcement Learning for Dynamic Difficulty Adjustment in Dungeons and DragonsabstractBalancing combat encounters in Dungeons & Dragons (D&D) is a complex task that requires Dungeon Masters (DM) to manually assess party strength, enemy composition, and dynamic player interactions while avoiding interruption of the narrative flow. In this paper we propose Encounter Generation via Reinforcement Learning (NTRL), a novel approach that automates Dynamic Difficulty Adjustment (DDA) in D&D via combat encounter design. By framing the problem as a contextual bandit, NTRL generates encounters based on real-time party members attributes. In comparison with classic DM heuristics, NTRL iteratively optimizes encounters to extend combat longevity ($+\mathbf{2 0 0 \%}$), increases damage dealt to party members, reducing post-combat hit points ($\mathbf{- 1 6. 6 7 \%}$), and raises the number of player deaths while maintaining low total party kills (TPK). The intensification of combat forces players to act wisely and engage in tactical maneuvers, even though the generated encounters guarantee high win rates (70 %). Even in comparison with encounters designed by human Dungeon Masters, NTRL demonstrates superior performance by enhancing the strategic depth of combat while increasing difficulty in a manner that preserves overall game fairness. Source code is available at github.com/CarloRomeo427/NTRL. Carlo Romeo, Andrew D. Bagdanov |
CoG | 2 |
| 2025 | Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality InversionabstractPre-trained multi-modal Vision-Language Models like CLIP are widely used off-the-shelf for a variety of applications. In this paper, we show that the common practice of individually exploiting the text or image encoders of these powerful multi-modal models is highly suboptimal for intra-modal tasks like image-to-image retrieval. We argue that this is inherently due to the CLIP-style inter-modal contrastive loss that does not enforce any intra-modal constraints, leading to what we call intra-modal misalignment. To demonstrate this, we leverage two optimization-based modality inversion techniques that map representations from their input modality to the complementary one without any need for auxiliary data or additional trained adapters. We empirically show that, in the intra-modal tasks of image-to-image and text-to-text retrieval, approaching these tasks inter-modally significantly improves performance with respect to intra-modal baselines on more than fifteen datasets. Additionally, we demonstrate that approaching a native inter-modal task (e.g. zero-shot image classification) intra-modally decreases performance, further validating our findings. Finally, we show that incorporating an intra-modal term in the pre-training objective or narrowing the modality gap between the text and image feature embedding spaces helps reduce the intra-modal misalignment. The code is publicly available at: https://github.com/miccunifi/Cross-the-Gap. Marco Mistretta, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini 0001, Andrew D. Bagdanov |
ICLR | 5 |
| 2025 | No Task Left Behind: Isotropic Model Merging with Common and Task-Specific SubspacesabstractModel merging integrates the weights of multiple task-specific models into a single multi-task model. Despite recent interest in the problem, a significant performance gap between the combined and single-task models remains. In this paper, we investigate the key characteristics of task matrices -- weight update matrices applied to a pre-trained model -- that enable effective merging. We show that alignment between singular components of task-specific and merged matrices strongly correlates with performance improvement over the pre-trained model. Based on this, we propose an isotropic merging framework that flattens the singular value spectrum of task matrices, enhances alignment, and reduces the performance gap. Additionally, we incorporate both common and task-specific subspaces to further improve alignment and performance. Our proposed approach achieves state-of-the-art performance on vision and language tasks across various sets of tasks and model scales. This work advances the understanding of model merging dynamics, offering an effective methodology to merge models without requiring additional training. Daniel Marczak, Simone Magistri, Sebastian Cygert, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de Weijer 0001 |
ICML | 5 |
| 2025 | ViewpointDepth: A New Dataset for Monocular Depth Estimation Under Viewpoint ShiftsabstractMonocular depth estimation is a critical task for autonomous driving and many other computer vision applications. While significant progress has been made in this field, the effects of viewpoint shifts on depth estimation models remain largely underexplored. This paper introduces a novel dataset and evaluation methodology to quantify the impact of different camera positions and orientations on monocular depth estimation performance. We propose a ground truth strategy based on homography estimation and object detection, eliminating the need for expensive LIDAR sensors. We collect a diverse dataset of road scenes from multiple viewpoints and use it to assess the robustness of a modern depth estimation model to geometric shifts. After assessing the validity of our strategy on a public dataset, we provide valuable insights into the limitations of current models and highlight the importance of considering viewpoint variations in real-world applications. Aurel Pjetri, Stefano Caprasecca, Leonardo Taccari, Matteo Simoncini, Henrique Piñeiro Monteagudo, Wallace Walter, Douglas Coimbra de Andrade, Francesco Sambo, Andrew D. Bagdanov |
IV | 9 |
| 2025 | Covariances for Free: Exploiting Mean Distributions for Training-free Federated LearningabstractUsing pre-trained models has been found to reduce the effect of data heterogeneity and speed up federated learning algorithms. Recent works have explored training-free methods using first- and second-order statistics to aggregate local client data distributions at the server and achieve high performance without any training. In this work, we propose a training-free method based on an unbiased estimator of class covariance matrices which only uses first-order statistics in the form of class means communicated by clients to the server. We show how these estimated class covariances can be used to initialize the global classifier, thus exploiting the covariances without actually sharing them. We also show that using only within-class covariances results in a better classifier initialization. Our approach improves performance in the range of 4-26% with exactly the same communication cost when compared to methods sharing only class means and achieves performance competitive or superior to methods sharing second-order statistics with dramatically less communication overhead. The proposed method is much more communication-efficient than federated prompt-tuning methods and still outperforms them. Finally, using our method to initialize classifiers and then performing federated fine-tuning or linear probing again yields better performance. Code is available at https://github.com/dipamgoswami/FedCOF. Dipam Goswami, Simone Magistri, Kai Wang 0060, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de Weijer 0001 |
NeurIPS | 5 |
| 2025 | Accurate and Efficient Low-Rank Model Merging in Core SpaceabstractIn this paper, we address the challenges associated with merging low-rank adaptations of large neural networks. With the rise of parameter-efficient adaptation techniques, such as Low-Rank Adaptation (LoRA), model fine-tuning has become more accessible. While fine-tuning models with LoRA is highly efficient, existing merging methods often sacrifice this efficiency by merging fully-sized weight matrices. We propose the Core Space merging framework, which enables the merging of LoRA-adapted models within a common alignment basis, thereby preserving the efficiency of low-rank adaptation while substantially improving accuracy across tasks. We further provide a formal proof that projection into Core Space ensures no loss of information and provide a complexity analysis showing the efficiency gains. Extensive empirical results demonstrate that Core Space significantly improves existing merging techniques and achieves state-of-the-art results on both vision and language tasks while utilizing a fraction of the computational resources. Codebase is available at https://github.com/apanariello4/core-space-merging. Aniello Panariello, Daniel Marczak, Simone Magistri, Angelo Porrello, Bartlomiej Twardowski, Andrew D. Bagdanov, Simone Calderara, Joost van de Weijer 0001 |
NeurIPS | 6 |
| 2025 | Exemplar-Free Continual Learning of Vision Transformers via Gated Class-Attention and Cascaded Feature Drift CompensationabstractAbstract Vision transformers (ViTs) have achieved remarkable successes across a broad range of computer vision applications. As a consequence, there has been increasing interest in extending continual learning theory and techniques to ViT architectures. We propose a new method for exemplar-free class incremental training of ViTs. The main challenge of exemplar-free continual learning is maintaining plasticity of the learner without causing catastrophic forgetting of previously learned tasks. This is often achieved via exemplar replay which can help recalibrate previous task classifiers to the feature drift which occurs when learning new tasks. Exemplar replay, however, comes at the cost of retaining samples from previous tasks which for many applications may not be possible. To address the problem of continual ViT training, we first propose gated class-attention to minimize the drift in the final ViT transformer block. This mask-based gating is applied to class-attention mechanism of the last transformer block and strongly regulates the weights crucial for previous tasks. Importantly, gated class-attention does not require the task-ID during inference, which distinguishes it from other parameter isolation methods. Secondly, we propose a new method of feature drift compensation that accommodates feature drift in the backbone when learning new tasks. The combination of gated class-attention and cascaded feature drift compensation allows for plasticity towards new tasks while limiting forgetting of previous ones. Extensive experiments performed on CIFAR-100, Tiny-ImageNet and ImageNet100 demonstrate that our exemplar-free method obtains competitive results when compared to rehearsal based ViT methods.(Code: https://github.com/OcraM17/GCAB-CFDC ) Marco Cotogni, Fei Yang 0004, Claudio Cusano, Andrew D. Bagdanov, Joost van de Weijer 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | EUFCC-340K: A faceted hierarchical dataset for metadata annotation in GLAM collectionsabstractAbstract In this paper, we address the challenges of automatic metadata annotation in the domain of Galleries, Libraries, Archives, and Museums (GLAMs) by introducing a novel dataset, EUFCC-340K, collected from the Europeana portal. Comprising over 340,000 images, the EUFCC-340K dataset is organized across multiple facets – Materials, Object Types, Disciplines, and Subjects – following a hierarchical structure based on the Art & Architecture Thesaurus (AAT). We developed several baseline models, incorporating multiple heads on a ConvNeXT backbone for multi-label image tagging on these facets, and fine-tuning a CLIP model with our image-text pairs. Our experiments to evaluate model robustness and generalization capabilities in two different test scenarios demonstrate the dataset’s utility in improving multi-label classification tools that have the potential to alleviate cataloging tasks in the cultural heritage sector. The EUFCC-340K dataset is publicly available at https://github.com/cesc47/EUFCC-340K . Francesc Net, Marc Folia, Pep Casals, Andrew D. Bagdanov, Lluís Gómez i Bigorda |
Multim. Tools Appl. | 4 |
| 2024 | A Benchmark Environment for Offline Reinforcement Learning in Racing GamesabstractOffline Reinforcement Learning (ORL) is a promising approach to reduce the high sample complexity of traditional Reinforcement Learning (RL) by eliminating the need for continuous environmental interactions. ORL exploits a dataset of precollected transitions and thus expands the range of application of RL to tasks in which the excessive environment queries increase training time and decrease efficiency, such as in modern AAA games. This paper introduces OfflineMania a novel environment for ORL research. It is inspired by the iconic TrackMania series and developed using the Unity 3D game engine. The environment simulates a single-agent racing game in which the objective is to complete the track through optimal navigation. We provide a variety of datasets to assess ORL performance. These datasets, created from policies of varying ability and in different sizes, aim to offer a challenging testbed for algorithm development and evaluation. We further establish a set of baselines for a range of Online RL, ORL, and hybrid Offline to Online RL approaches using our environment. Girolamo Macaluso, Alessandro Sestini, Andrew D. Bagdanov |
CoG | 3 |
| 2024 | The intrinsic convenience of federated learning in malware IoT detectionabstractThe Internet of Things is emerging as a key concept, defining a network of interconnected devices capable of seamless data collection, exchange, and analysis. However, due to their emphasis on simplicity, these devices are often vulnerable to malware attacks. This study examines the potential of machine learning methods, specifically in the context of Federated Learning, to enhance privacy protection and to benefit from IoT’s decentralized nature, such as the low overhead traffic. The proposed approach is a federated machine learning algorithm based on a central aggregator and several clients. The study aims to conduct a comprehensive analysis using the IOT-23 dataset, which contains real and labeled instances of malware infections. The test outcomes demonstrate that the proposed approach outperforms centralized approaches regarding the global area under the precision-recall curve (AUPRC) and variance, with a significance level of 0.05. Chiara Camerota, Tommaso Pecorella, Andrew D. Bagdanov |
CNSM | 3 |
| 2024 | Task-Adaptive Saliency Guidance for Exemplar-Free Class Incremental LearningabstractExemplar-free Class Incremental Learning (EFCIL) aims to sequentially learn tasks with access only to data from the current one. EFCIL is of interest because it mit-igates concerns about privacy and long-term storage of data, while at the same time alleviating the problem of catastrophic forgetting in incremental learning. In this work, we introduce task-adaptive saliency for EFCIL and propose a new framework, which we call Task-Adaptive Saliency Supervision (TASS), for mitigating the negative effects of saliency drift between different tasks. We first apply boundary-guided saliency to maintain task adaptiv-ity and plasticity on model attention. Besides, we introduce task-agnostic low-level signals as auxiliary supervision to increase the stability of model attention. Finally, we introduce a module for injecting and recovering saliency noise to increase the robustness of saliency preservation. Our experiments demonstrate that our method can better preserve saliency maps across tasks and achieve state-of-the-art results on the CIFAR-100, Tiny-ImageNet, and ImageNet-Subset EFCIL benchmarks. Code is available at https://github.com/scok30/tass. Xialei Liu, Jiang-Tian Zhai, Andrew D. Bagdanov, Ke Li 0015, Ming-Ming Cheng |
CVPR | 3 |
| 2024 | Exemplar-Free Continual Representation Learning via Learnable Drift Compensation
Alexandra Gomez-Villa, Dipam Goswami, Kai Wang 0060, Andrew D. Bagdanov, Bartlomiej Twardowski, Joost van de Weijer 0001 |
ECCV (7) | 4 |
| 2024 | Improving Zero-Shot Generalization of Learned Prompts via Unsupervised Knowledge Distillation
Marco Mistretta, Alberto Baldrati, Marco Bertini 0001, Andrew D. Bagdanov |
ECCV (84) | 4 |
| 2024 | Elastic Feature Consolidation For Cold Start Exemplar-Free Incremental LearningabstractExemplar-Free Class Incremental Learning (EFCIL) aims to learn from a sequence of tasks without having access to previous task data. In this paper, we consider the challenging Cold Start scenario in which insufficient data is available in the first task to learn a high-quality backbone. This is especially challenging for EFCIL since it requires high plasticity, which results in feature drift which is difficult to compensate for in the exemplar-free setting. To address this problem, we propose a simple and effective approach that consolidates feature representations by regularizing drift in directions highly relevant to previous tasks and employs prototypes to reduce task-recency bias. Our method, called Elastic Feature Consolidation (EFC), exploits a tractable second-order approximation of feature drift based on an Empirical Feature Matrix (EFM). The EFM induces a pseudo-metric in feature space which we use to regularize feature drift in important directions and to update Gaussian prototypes used in a novel asymmetric cross entropy loss which effectively balances prototype rehearsal with data from new tasks. Experimental results on CIFAR-100, Tiny-ImageNet, ImageNet-Subset and ImageNet-1K demonstrate that Elastic Feature Consolidation is better able to learn new tasks by maintaining model plasticity and significantly outperform the state-of-the-art. Simone Magistri, Tomaso Trinci, Albin Soutif-Cormerais, Joost van de Weijer 0001, Andrew D. Bagdanov |
ICLR | 5 |
| 2024 | Continual learning for adaptive social network identificationabstractThe popularity of social networks as primary mediums for sharing visual content has made it crucial for forensic experts to identify the original platform of multimedia content. Various methods address this challenge, but the constant emergence of new platforms and updates to existing ones often render forensic tools ineffective shortly after release. This necessitates the regular updating of methods and models, which can be particularly cumbersome for techniques based on neural networks which cannot quickly adapt to new classes without sacrificing performance on previously learned ones – a phenomenon known as catastrophic forgetting. Recently, researchers aimed at mitigating this problem via a family of techniques known as continual learning. In this paper we study the applicability of continual learning techniques to the social network identification task by evaluating two relevant forensic scenarios: Incremental Social Platform Classification, for handling newly introduced social media platforms, and Incremental Social Version Classification, for addressing updated versions of a set of existing social networks. We perform an extensive experimental evaluation of a variety of continual learning approaches applied to these two scenarios. Experimental results demonstrate that, although Continual Social Network Identification remains a difficult problem, catastrophic forgetting can be significantly mitigated in both scenarios by retaining only a fraction of the image patches from past task training samples or by employing previous tasks prototypes. Simone Magistri, Daniele Baracchi, Dasara Shullani, Andrew D. Bagdanov, Alessandro Piva |
Pattern Recognit. Lett. | 4 |
| 2024 | Automated Gameplay Testing and Validation With Curiosity-Conditioned Proximal TrajectoriesabstractThis article proposes a novel deep reinforcement learning algorithm to perform automated analysis and detection of gameplay issues in complex 3-D navigation environments. The curiosity-conditioned proximal trajectories (CCPT) method combines curiosity and imitation learning to train agents that methodically explore in the proximity of known trajectories derived from expert demonstrations. We show how our new algorithm can explore complex environments, discovering gameplay issues, and design oversights in the process, and recognize and highlight them directly to game designers. We also propose a visual analytics interface to aid interpretation of results from the method. This interface transforms information from complex models into interpretable and interactive visual forms. We further demonstrate the effectiveness of the algorithm in a novel 3-D navigation environment, which reflects the complexity of modern video games. Our results show a higher level of coverage and bug discovery than baseline methods, demonstrating that our method can be a useful tool for game designers to automatically identify design issues. Moreover, our experiments show that the visual explanations provided by the analytics interface result in a significant increase in user trust and acceptance of automated playtesting and increased confidence in the use of machine learning techniques for video game development. Alessandro Sestini, Linus Gisslén, Joakim Bergdahl, Konrad Tollmar, Andrew D. Bagdanov |
IEEE Trans. Games | 5 |
| 2023 | Towards Informed Design and Validation Assistance in Computer Games Using Imitation LearningabstractIn games, as in many other domains, design validation and testing is a significant challenge as systems are growing in size and manual testing is becoming infeasible. In this position paper we outline an approach to automated game validation based on an imitation learning technique, and provide an analysis of the potential benefits to automated game testing. The method leverages a data-driven technique, which requires little effort and time and no knowledge of machine learning or programming, that designers can use to efficiently train game testing agents. We evaluate the validity of our claim by conducting a user study with industry experts. The survey results presented in this paper demonstrate the potential of a data-driven approach to reduce effort and enhance the quality of game testing. Moreover, the survey reveals several open challenges. To this end, we analyze the identified challenges and provide a basis for further research and discussion, as well as to help guide the development of imitation learning for game testing. Alessandro Sestini, Joakim Bergdahl, Konrad Tollmar, Andrew D. Bagdanov, Linus Gisslén |
CoG | 4 |
| 2023 | Masked Autoencoders are Efficient Class Incremental LearnersabstractClass Incremental Learning (CIL) aims to sequentially learn new classes while avoiding catastrophic forgetting of previous knowledge. We propose to use Masked Autoencoders (MAEs) as efficient learners for CIL. MAEs were originally designed to learn useful representations through reconstructive unsupervised learning, and they can be easily integrated with a supervised loss for classification. Moreover, MAEs can reliably reconstruct original input images from randomly selected patches, which we use to store exemplars from past tasks more efficiently for CIL. We also propose a bilateral MAE framework to learn from image-level and embedding-level fusion, which produces better-quality reconstructed images and more stable representations. Our experiments confirm that our approach performs better than the state-of-the-art on CIFAR-100, ImageNet-Subset, and ImageNet-Full. The code is available at https://github.com/scok30/MAE-CIL. Jiang-Tian Zhai, Xialei Liu, Andrew D. Bagdanov, Ke Li 0015, Ming-Ming Cheng |
ICCV | 3 |
| 2023 | Planckian Jitter: countering the color-crippling effects of color jitter on self-supervised training
Simone Zini, Alexandra Gomez-Villa, Marco Buzzelli, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de Weijer 0001 |
ICLR | 5 |
| 2023 | Class-Incremental Learning: Survey and Performance Evaluation on Image ClassificationabstractFor future learning systems, incremental learning is desirable because it allows for: efficient resource usage by eliminating the need to retrain from scratch at the arrival of new data; reduced memory usage by preventing or limiting the amount of data required to be stored - also important when privacy limitations are imposed; and learning that more closely resembles human learning. The main challenge for incremental learning is catastrophic forgetting, which refers to the precipitous drop in performance on previously learned tasks after learning a new one. Incremental learning of deep neural networks has seen explosive growth in recent years. Initial work focused on task-incremental learning, where a task-ID is provided at inference time. Recently, we have seen a shift towards class-incremental learning where the learner must discriminate at inference time between all classes seen in previous tasks without recourse to a task-ID. In this paper, we provide a complete survey of existing class-incremental learning methods for image classification, and in particular, we perform an extensive experimental evaluation on thirteen class-incremental methods. We consider several new experimental scenarios, including a comparison of class-incremental methods on multiple large-scale image classification datasets, an investigation into small and large domain shifts, and a comparison of various network architectures. Marc Masana, Xialei Liu, Bartlomiej Twardowski, Mikel Menta, Andrew D. Bagdanov, Joost van de Weijer 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Positive Pair Distillation Considered Harmful: Continual Meta Metric Learning for Lifelong Object Re-Identification
Kai Wang 0060, Chenshen Wu, Andrew D. Bagdanov, Xialei Liu, Shiqi Yang 0002, Shangling Jui, Joost van de Weijer 0001 |
BMVC | 3 |
| 2022 | Long-Tailed Class Incremental Learning
Xialei Liu, Yusong Hu, Andrew D. Bagdanov, Ke Li 0015, Ming-Ming Cheng |
ECCV (33) | 4 |
| 2021 | Demonstration-Efficient Inverse Reinforcement Learning in Procedurally Generated EnvironmentsabstractDeep Reinforcement Learning achieves very good results in domains where reward functions can be manually engineered. At the same time, there is growing interest within the community in using games based on Procedurally Content Generation (PCG) as benchmark environments since this type of environment is perfect for studying overfitting and generalization of agents under domain shift. Inverse Reinforcement Learning (IRL) can instead extrapolate reward functions from expert demonstrations, with good results even on high-dimensional problems, however there are no examples of applying these techniques to procedurally-generated environments. This is mostly due to the number of demonstrations needed to find a good reward model. We propose a technique based on Adversarial Inverse Reinforcement Learning which can significantly decrease the need for expert demonstrations in PCG games. Through the use of an environment with a limited set of initial seed levels, plus some modifications to stabilize training, we show that our approach, DE-AIRL, is demonstration-efficient and still able to extrapolate reward functions which generalize to the fully procedural domain. We demonstrate the effectiveness of our technique on two procedural environments, MiniGrid and DeepCrawl, for a variety of tasks. Alessandro Sestini, Andrew D. Bagdanov |
CoG | 3 |
| 2021 | Policy Fusion for Adaptive and Customizable Reinforcement Learning AgentsabstractIn this article we study the problem of training intelligent agents using Reinforcement Learning for the purpose of game development. Unlike systems built to replace human players and to achieve super-human performance, our agents aim to produce meaningful interactions with the player, and at the same time demonstrate behavioral traits as desired by game designers. We show how to combine distinct behavioral policies to obtain a meaningful “fusion” policy which comprises all these behaviors. To this end, we propose four different policy fusion methods for combining pre-trained policies. We further demonstrate how these methods can be used in combination with Inverse Reinforcement Learning in order to create intelligent agents with specific behavioral styles as chosen by game designers, without having to define many and possibly poorly-designed reward functions. Experiments on two different environments indicate that entropy-weighted policy fusion significantly outperforms all others. We provide several practical examples and use-cases for how these methods are indeed useful for video game production and designers. Alessandro Sestini, Andrew D. Bagdanov |
CoG | 3 |
| 2021 | Bottom-up and Layerwise Domain Adaptation for Pedestrian Detection in Thermal ImagesabstractPedestrian detection is a canonical problem for safety and security applications, and it remains a challenging problem due to the highly variable lighting conditions in which pedestrians must be detected. This article investigates several domain adaptation approaches to adapt RGB-trained detectors to the thermal domain. Building on our earlier work on domain adaptation for privacy-preserving pedestrian detection, we conducted an extensive experimental evaluation comparing top-down and bottom-up domain adaptation and also propose two new bottom-up domain adaptation strategies. For top-down domain adaptation, we leverage a detector pre-trained on RGB imagery and efficiently adapt it to perform pedestrian detection in the thermal domain. Our bottom-up domain adaptation approaches include two steps: first, training an adapter segment corresponding to initial layers of the RGB-trained detector adapts to the new input distribution; then, we reconnect the adapter segment to the original RGB-trained detector for final adaptation with a top-down loss. To the best of our knowledge, our bottom-up domain adaptation approaches outperform the best-performing single-modality pedestrian detection results on KAIST and outperform the state of the art on FLIR. My Kieu, Andrew D. Bagdanov, Marco Bertini 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Task-Conditioned Domain Adaptation for Pedestrian Detection in Thermal Imagery
My Kieu, Andrew D. Bagdanov, Marco Bertini 0001, Alberto Del Bimbo |
ECCV (22) | 2 |
| 2020 | Robust pedestrian detection in thermal imagery using synthesized imagesabstractIn this paper we propose a method for improving pedestrian detection in the thermal domain using two stages: first, a generative data augmentation approach is used, then a domain adaptation method using generated data adapts an RGB pedestrian detector. Our model, based on the Least-Squares Generative Adversarial Network, is trained to synthesize realistic thermal versions of input RGB images which are then used to augment the limited amount of labeled thermal pedestrian images available for training. We apply our generative data augmentation strategy in order to adapt a pretrained YOLOv3 pedestrian detector to detection in the thermal-only domain. Experimental results demonstrate the effectiveness of our approach: using less than 50% of available real thermal training data, and relying on synthesized data generated by our model in the domain adaptation phase, our detector achieves state-of-the-art results on the KAIST Multispectral Pedestrian Detection Benchmark; even if more real thermal data is available adding GAN generated images to the training data results in improved performance, thus showing that these images act as an effective form of data augmentation. To the best of our knowledge, our detector achieves the best single-modality detection results on KAIST with respect to the state-of-the-art. My Kieu, Lorenzo Berlincioni, Leonardo Galteri, Marco Bertini 0001, Andrew D. Bagdanov, Alberto Del Bimbo |
ICPR | 5 |
| 2020 | RATT: Recurrent Attention to Transient Tasks for Continual Image CaptioningabstractResearch on continual learning has led to a variety of approaches to mitigating catastrophic forgetting in feed-forward classification networks. Until now surprisingly little attention has been focused on continual learning of recurrent models applied to problems like image captioning. In this paper we take a systematic look at continual learning of LSTM-based models for image captioning. We propose an attention-based approach that explicitly accommodates the transient nature of vocabularies in continual image captioning tasks -- i.e. that task vocabularies are not disjoint. We call our method Recurrent Attention to Transient Tasks (RATT), and also show how to adapt continual learning approaches based on weight regularization and knowledge distillation to recurrent continual learning problems. We apply our approaches to incremental image captioning problem on two new continual learning benchmarks we define using the MS-COCO and Flickr30 datasets. Our results demonstrate that RATT is able to sequentially learn five captioning tasks while incurring no forgetting of previously learned ones. Riccardo Del Chiaro, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de Weijer 0001 |
NeurIPS | 3 |
| 2019 | DeepPhysio: Monitored Physiotherapeutic Exercise in the Comfort of your Own HomeabstractThis paper describes an action classification pipeline for detecting and evaluating correct execution of actions in video recorded by smartphone cameras; the use case is that of simplifying monitoring of how physiotherapeutic exercises are performed by patients in the comfort of their own home, reducing the need of physical presence of therapists. Our approach is based on applying DensePose to every frame of acquired video and subsequent sequence analysis by an LSTM network. We validate our proposed recognition approach on a subset of the NTU RGB+D dataset in order to determine the best classification pipeline for this application. We also describe a mobile, cross-platform application called DeepPhysio that is designed to allow at physiotherapy patients to obtain immediate feedback about the correctness of the physical exercises. Preliminary usability analysis shows that this type of application can be effective at monitoring physiotherapy exercises. Gianmarco Sanesi, Andrew D. Bagdanov, Marco Bertini 0001, Alberto Del Bimbo |
ACM Multimedia | 2 |
| 2019 | Exploiting Unlabeled Data in CNNs by Self-Supervised Learning to RankabstractFor many applications the collection of labeled data is expensive laborious. Exploitation of unlabeled data during training is thus a long pursued objective of machine learning. Self-supervised learning addresses this by positing an auxiliary task (different, but related to the supervised task) for which data is abundantly available. In this paper, we show how ranking can be used as a proxy task for some regression problems. As another contribution, we propose an efficient backpropagation technique for Siamese networks which prevents the redundant computation introduced by the multi-branch network architecture. We apply our framework to two regression problems: Image Quality Assessment (IQA) and Crowd Counting. For both we show how to automatically generate ranked image sets from unlabeled data. Our results show that networks trained to regress to the ground truth targets for labeled data and to simultaneously learn to rank unlabeled data obtain significantly better, state-of-the-art results for both IQA and crowd counting. In addition, we show that measuring network uncertainty on the self-supervised proxy task is a good measure of informativeness of unlabeled data. This can be used to drive an algorithm for active learning and we show that this reduces labeling effort by up to 50 percent. Xialei Liu, Joost van de Weijer 0001, Andrew D. Bagdanov |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | FAST: Facilitated and Accurate Scene Text Proposals through FCN Guided Pruning
Dena Bazazian, Raul Gomez, Anguelos Nicolaou, Lluís Gómez i Bigorda, Dimosthenis Karatzas, Andrew D. Bagdanov |
Pattern Recognit. Lett. | 6 |
| 2019 | Webly-supervised zero-shot learning for artwork instance recognition
Riccardo Del Chiaro, Andrew D. Bagdanov, Alberto Del Bimbo |
Pattern Recognit. Lett. | 2 |
| 2018 | Leveraging Unlabeled Data for Crowd Counting by Learning to RankabstractWe propose a novel crowd counting approach that leverages abundantly available unlabeled crowd imagery in a learning-to-rank framework. To induce a ranking of cropped images, we use the observation that any sub-image of a crowded scene image is guaranteed to contain the same number or fewer persons than the super-image. This allows us to address the problem of limited size of existing datasets for crowd counting. We collect two crowd scene datasets from Google using keyword searches and query-by-example image retrieval, respectively. We demonstrate how to efficiently learn from these unlabeled datasets by incorporating learning-to-rank in a multi-task network which simultaneously ranks images and estimates crowd density maps. Experiments on two of the most challenging crowd counting datasets show that our approach obtains state-of-the-art results. Xialei Liu, Joost van de Weijer 0001, Andrew D. Bagdanov |
CVPR | 3 |
| 2018 | Rotate your Networks: Better Weight Consolidation and Less Catastrophic ForgettingabstractIn this paper we propose an approach to avoiding catastrophic forgetting in sequential task learning scenarios. Our technique is based on a network reparameterization that approximately diagonalizes the Fisher Information Matrix of the network parameters. This reparameterization takes the form of a factorized rotation of parameter space which, when used in conjunction with Elastic Weight Consolidation (which assumes a diagonal Fisher Information Matrix), leads to significantly better performance on lifelong learning of sequential tasks. Experimental results on the MNIST, CIFAR-100, CUB-200 and Stanford-40 datasets demonstrate that we significantly improve the results of standard elastic weight consolidation, and that we obtain competitive results when compared to the state-of-the-art in lifelong learning without forgetting. Xialei Liu, Marc Masana, Luis Herranz, Joost van de Weijer 0001, Antonio M. López 0001, Andrew D. Bagdanov |
ICPR | 6 |
| 2018 | Review on computer vision techniques in emergency situations
Laura Lopez-Fuentes, Joost van de Weijer 0001, Manuel González Hidalgo, Harald Skinnemoen, Andrew D. Bagdanov |
Multim. Tools Appl. | 5 |
| 2018 | Scale coding bag of deep features for human attribute and action recognitionabstractMost approaches to human attribute and action recognition in still images are based on image representation in which multi-scale local features are pooled across scale into a single, scale-invariant encoding. Both in bag-of-words and the recently popular representations based on convolutional neural networks, local features are computed at multiple scales. However, these multi-scale convolutional features are pooled into a single scale-invariant representation. We argue that entirely scale-invariant image representations are sub-optimal and investigate approaches to scale coding within a bag of deep features framework. Our approach encodes multi-scale information explicitly during the image encoding stage. We propose two strategies to encode multi-scale information explicitly in the final image representation. We validate our two scale coding techniques on five datasets: Willow, PASCAL VOC 2010, PASCAL VOC 2012, Stanford-40 and Human Attributes (HAT-27). On all datasets, the proposed scale coding approaches outperform both the scale-invariant method and the standard deep features of the same network. Further, combining our scale coding approaches with standard deep features leads to consistent improvement over the state of the art. Fahad Shahbaz Khan, Joost van de Weijer 0001, Rao Muhammad Anwer, Andrew D. Bagdanov, Michael Felsberg, Jorma Laaksonen |
Mach. Vis. Appl. | 4 |
| 2017 | RankIQA: Learning from Rankings for No-Reference Image Quality AssessmentabstractWe propose a no-reference image quality assessment (NR-IQA) approach that learns from rankings (RankIQA). To address the problem of limited IQA dataset size, we train a Siamese Network to rank images in terms of image quality by using synthetically generated distortions for which relative image quality is known. These ranked image sets can be automatically generated without laborious human labeling. We then use fine-tuning to transfer the knowledge represented in the trained Siamese Network to a traditional CNN that estimates absolute image quality from single images. We demonstrate how our approach can be made significantly more efficient than traditional Siamese Networks by forward propagating a batch of images through a single network and backpropagating gradients derived from all pairs of images in the batch. Experiments on the TID2013 benchmark show that we improve the state-of-theart by over 5%. Furthermore, on the LIVE benchmark we show that our approach is superior to existing NR-IQA techniques and that we even outperform the state-of-the-art in full-reference IQA (FR-IQA) methods without having to resort to high-quality reference images to infer IQA. Xialei Liu, Joost van de Weijer 0001, Andrew D. Bagdanov |
ICCV | 3 |
| 2017 | Domain-Adaptive Deep Network Compression
Marc Masana, Joost van de Weijer 0001, Luis Herranz, Andrew D. Bagdanov, José M. Álvarez 0004 |
ICCV | 4 |
| 2017 | Visual Attention Models for Scene Text RecognitionabstractIn this paper we propose an approach to lexicon-free recognition of text in scene images. Our approach relies on a LSTM-based soft visual attention model learned from convolutional features. A set of feature vectors are derived from an intermediate convolutional layer corresponding to different areas of the image. This permits encoding of spatial information into the image representation. In this way, the framework is able to learn how to selectively focus on different parts of the image. At every time step the recognizer emits one character using a weighted combination of the convolutional feature vectors according to the learned attention model. Training can be done end-to-end using only word level annotations. In addition, we show that modifying the beam search algorithm by integrating an explicit language model leads to significantly better recognition results. We validate the performance of our approach on standard SVT and ICDAR'03 scene text datasets, showing state-of-the-art performance in unconstrained text recognition. Suman K. Ghosh, Ernest Valveny, Andrew D. Bagdanov |
ICDAR | 3 |
| 2017 | Bandwidth Limited Object Recognition in High Resolution ImageryabstractThis paper proposes a novel method to optimize bandwidth usage for object detection in critical communication scenarios. We develop two operating models of active information seeking. The first model identifies promising regions in low resolution imagery and progressively requests higher resolution regions on which to perform recognition of higher semantic quality. The second model identifies promising regions in low resolution imagery while simultaneously predicting the approximate location of the object of higher semantic quality. From this general framework, we develop a car recognition system via identification of its license plate and evaluate the performance of both models on a car dataset that we introduce. Results are compared with traditional JPEG compression and demonstrate that our system saves up to one order of magnitude of bandwidth while sacrificing little in terms of recognition performance. Laura Lopez-Fuentes, Andrew D. Bagdanov, Joost van de Weijer 0001, Harald Skinnemoen |
WACV | 2 |
| 2016 | Visual Script and Language IdentificationabstractIn this paper we introduce a script identification method based on hand-crafted texture features and an artificial neural network. The proposed pipeline achieves near state-of-the-art performance for script identification of video-text and state-of-the-art performance on visual language identification of handwritten text. More than using the deep network as a classifier, the use of its intermediary activations as a learned metric demonstrates remarkable results and allows the use of discriminative models on unknown classes. Comparative experiments in video-text and text in the wild datasets provide insights on the internals of the proposed deep network. Anguelos Nicolaou, Andrew D. Bagdanov, Lluís Gómez i Bigorda, Dimosthenis Karatzas |
DAS | 2 |
| 2016 | Hierarchical part detection with deep neural networksabstractPart detection is an important aspect of object recognition. Most approaches apply object proposals to generate hundreds of possible part bounding box candidates which are then evaluated by part classifiers. Recently several methods have investigated directly regressing to a limited set of bounding boxes from deep neural network representation. However, for object parts such methods may be unfeasible due to their relatively small size with respect to the image. We propose a hierarchical method for object and part detection. In a single network we first detect the object and then regress to part location proposals based only on the feature representation inside the object. Experiments show that our hierarchical approach outperforms a network which directly regresses the part locations. We also show that our approach obtains part detection accuracy comparable or better than state-of-the-art on the CUB-200 bird and Fashionista clothing item datasets with only a fraction of the number of part proposals. Esteve Cervantes, Andrew D. Bagdanov, Marc Masana, Joost van de Weijer 0001 |
ICIP | 3 |
| 2016 | Personalized multimedia content delivery on an interactive table by passive observation of museum visitors
Svebor Karaman, Andrew D. Bagdanov, Lea Landucci, Gianpaolo D'Amico, Andrea Ferracani, Daniele Pezzatini, Alberto Del Bimbo |
Multim. Tools Appl. | 2 |
| 2015 | MORF: Multi-Objective Random Forests for face characteristic estimationabstractIn this paper we describe a technique for joint estimation of head pose and multiple soft biometrics from faces (Age, Gender and Ethnicity). Our proposed Multi-Objective Random Forests (MORF) framework is a unified model for the joint estimation of multiple characteristics that automatically adapts the measure of information gain used for evaluating the quality of weak learners. Since facial characteristics are related in the feature space, estimating all of them jointly can be beneficial as trees can learn to condition the estimation of some characteristics on others. We reformulate the splitting criterion of random trees in our multi-objective formulation and evaluate it on publicly available face characteristic estimation imagery. These preliminary experiments show promising results. Dario Di Fina, Svebor Karaman, Andrew D. Bagdanov, Alberto Del Bimbo |
AVSS | 3 |
| 2015 | ICDAR 2015 competition on Robust ReadingabstractResults of the ICDAR 2015 Robust Reading Competition are presented. A new Challenge 4 on Incidental Scene Text has been added to the Challenges on Born-Digital Images, Focused Scene Images and Video Text. Challenge 4 is run on a newly acquired dataset of 1,670 images evaluating Text Localisation, Word Recognition and End-to-End pipelines. In addition, the dataset for Challenge 3 on Video Text has been substantially updated with more video sequences and more accurate ground truth data. Finally, tasks assessing End-to-End system performance have been introduced to all Challenges. The competition took place in the first quarter of 2015, and received a total of 44 submissions. Only the tasks newly introduced in 2015 are reported on. The datasets, the ground truth specification and the evaluation protocols are presented together with the results and a brief summary of the participating methods. Dimosthenis Karatzas, Lluís Gómez i Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, Masakazu Iwamura, Jiri Matas, Lukás Neumann, Vijay Chandrasekhar 0001, Shijian Lu, Faisal Shafait, Seiichi Uchida, Ernest Valveny |
ICDAR | 5 |
| 2015 | Sparse radial sampling LBP for writer identificationabstractSampling Local Binary Patterns, a variant of Local Binary Patterns (LBP) for text-as-texture classification. By adapting and extending the standard LBP operator to the particularities of text we get a generic text-as-texture classification scheme and apply it to writer identification. In experiments on CVL and ICDAR 2013 datasets, the proposed feature-set and a simple end-to-end pipeline demonstrate State-Of-the-Art (SOA) performance. Among the SOA, the proposed method is the only one that is based on dense extraction of a single local feature descriptor. This makes it fast and applicable at the earliest stages in a DIA pipeline without the need for segmentation, binarization, or extraction of multiple features. Anguelos Nicolaou, Andrew D. Bagdanov, Marcus Liwicki, Dimosthenis Karatzas |
ICDAR | 2 |
| 2015 | Person Re-Identification by Iterative Re-Weighted Sparse RankingabstractIn this paper we introduce a method for person re-identification based on discriminative, sparse basis expansions of targets in terms of a labeled gallery of known individuals. We propose an iterative extension to sparse discriminative classifiers capable of ranking many candidate targets. The approach makes use of soft- and hard- re-weighting to redistribute energy among the most relevant contributing elements and to ensure that the best candidates are ranked at each iteration. Our approach also leverages a novel visual descriptor which we show to be discriminative while remaining robust to pose and illumination variations. An extensive comparative evaluation is given demonstrating that our approach achieves state-of-the-art performance on single- and multi-shot person re-identification scenarios on the VIPeR, i-LIDS, ETHZ, and CAVIAR4REID datasets. The combination of our descriptor and iterative sparse basis expansion improves state-of-the-art rank-1 performance by six percentage points on VIPeR and by 20 on CAVIAR4REID compared to other methods with a single gallery image per person. With multiple gallery and probe images per person our approach improves by 17 percentage points the state-of-the-art on i-LIDS and by 72 on CAVIAR4REID at rank-1. The approach is also quite efficient, capable of single-shot person re-identification over galleries containing hundreds of individuals at about 30 re-identifications per second. Giuseppe Lisanti, Iacopo Masi, Andrew D. Bagdanov, Alberto Del Bimbo |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Recognizing Actions Through Action-Specific Person DetectionabstractAction recognition in still images is a challenging problem in computer vision. To facilitate comparative evaluation independently of person detection, the standard evaluation protocol for action recognition uses an oracle person detector to obtain perfect bounding box information at both training and test time. The assumption is that, in practice, a general person detector will provide candidate bounding boxes for action recognition. In this paper, we argue that this paradigm is suboptimal and that action class labels should already be considered during the detection stage. Motivated by the observation that body pose is strongly conditioned on action class, we show that: 1) the existing state-of-the-art generic person detectors are not adequate for proposing candidate bounding boxes for action classification; 2) due to limited training examples, the direct training of action-specific person detectors is also inadequate; and 3) using only a small number of labeled action examples, the transfer learning is able to adapt an existing detector to propose higher quality bounding boxes for subsequent action classification. To the best of our knowledge, we are the first to investigate transfer learning for the task of action-specific person detection in still images. We perform extensive experiments on two benchmark data sets: 1) Stanford-40 and 2) PASCAL VOC 2012. For the action detection task (i.e., both person localization and classification of the action performed), our approach outperforms methods based on general person detection by 5.7% mean average precision (MAP) on Stanford-40 and 2.1% MAP on PASCAL VOC 2012. Our approach also significantly outperforms the state of the art with a MAP of 45.4% on Stanford-40 and 31.4% on PASCAL VOC 2012. We also evaluate our action detection approach for the task of action classification (i.e., recognizing actions without localizing them). For this task, our approach, without using any ground-truth person localization at test time, outperforms on both data sets state-of-the-art methods, which do use person locations. Fahad Shahbaz Khan, Jiaolong Xu, Joost van de Weijer 0001, Andrew D. Bagdanov, Rao Muhammad Anwer, Antonio M. López 0001 |
IEEE Trans. Image Process. | 4 |
| 2014 | Real-time people counting from depth imagery of crowded environmentsabstractIn this paper we describe a system for automatic people counting in crowded environments. The approach we propose is a counting-by-detection method based on depth imagery. It is designed to be deployed as an autonomous appliance for crowd analysis in video surveillance application scenarios. Our system performs foreground/background segmentation on depth image streams in order to coarsely segment persons, then depth information is used to localize head candidates which are then tracked in time on an automatically estimated ground plane. The system runs in real-time, at a frame-rate of about 20 fps. We collected a dataset of RGB-D sequences representing three typical and challenging surveillance scenarios, including crowds, queuing and groups. An extensive comparative evaluation is given between our system and more complex, Latent SVM-based head localization for person counting applications. Enrico Bondi, Lorenzo Seidenari, Andrew D. Bagdanov, Alberto Del Bimbo |
AVSS | 3 |
| 2014 | Fisher Vectors over Random Density Forests for Object RecognitionabstractIn this paper we describe a Fisher vector encoding of images over Random Density Forests. Random Density Forests (RDFs) are an unsupervised variation of Random Decision Forests for density estimation. In this work we train RDFs by splitting at each node in order to minimize the Gaussian differential entropy of each split. We use this as generative model of image patch features and derive the Fisher vector representation using the RDF as the underlying model. Our approach is computationally efficient, reducing the amount of Gaussian derivatives to compute, and allows more flexibility in the feature density modelling. We evaluate our approach on the PASCAL VOC 2007 dataset showing that our approach, that only uses linear classifiers, improves over bag of visual words and is comparable to the traditional Fisher vector encoding over Gaussian Mixture Models for density estimation. Claudio Baecchi, Francesco Turchini, Lorenzo Seidenari, Andrew D. Bagdanov, Alberto Del Bimbo |
ICPR | 4 |
| 2014 | Unsupervised Scene Adaptation for Faster Multi-scale Pedestrian DetectionabstractIn this paper we describe an approach to automatically improving the efficiency of soft cascade-based person detectors. Our technique addresses the two fundamental bottlenecks in cascade detectors: the number of weak classifiers that need to be evaluated in each cascade, and the total number of detection windows to be evaluated. By simply observing a soft cascade operating on a scene, we learn scale specific linear approximations of cascade traces that allows us to eliminate a large fraction of the classifier evaluation. Independently, this time by observing regions of support in the soft cascade on a training set, we learn a coarse geometric model of the scene that allows our detector to propose candidate detection windows and significantly reduce the number of windows run through the cascade. Our approaches are unsupervised and require no additional labeled person images for learning. Our linear cascade approximation results in about 28% savings in detection, while our geometric model gives a saving of over 95%, without appreciable loss of accuracy. Federico Bartoli, Giuseppe Lisanti, Svebor Karaman, Andrew D. Bagdanov, Alberto Del Bimbo |
ICPR | 4 |
| 2014 | Scale Coding Bag-of-Words for Action RecognitionabstractRecognizing human actions in still images is a challenging problem in computer vision due to significant amount of scale, illumination and pose variation. Given the bounding box of a person both at training and test time, the task is to classify the action associated with each bounding box in an image. Most state-of-the-art methods use the bag-of-words paradigm for action recognition. The bag-of-words framework employing a dense multi-scale grid sampling strategy is the de facto standard for feature detection. This results in a scale invariant image representation where all the features at multiple-scales are binned in a single histogram. We argue that such a scale invariant strategy is sub-optimal since it ignores the multi-scale information available with each bounding box of a person. This paper investigates alternative approaches to scale coding for action recognition in still images. We encode multi-scale information explicitly in three different histograms for small, medium and large scale visual-words. Our first approach exploits multi-scale information with respect to the image size. In our second approach, we encode multi-scale information relative to the size of the bounding box of a person instance. In each approach, the multi-scale histograms are then concatenated into a single representation for action classification. We validate our approaches on the Willow dataset which contains seven action categories: interacting with computer, photography, playing music, riding bike, riding horse, running and walking. Our results clearly suggest that the proposed scale coding approaches outperform the conventional scale invariant technique. Moreover, we show that our approach obtains promising results compared to more complex state-of-the-art methods. Fahad Shahbaz Khan, Joost van de Weijer 0001, Andrew D. Bagdanov, Michael Felsberg |
ICPR | 3 |
| 2014 | Multimodal page classification in administrative document image streams
Marçal Rusiñol, Volkmar Frinken, Dimosthenis Karatzas, Andrew D. Bagdanov, Josep Lladós 0001 |
Int. J. Document Anal. Recognit. | 4 |
| 2014 | Local Pyramidal Descriptors for Image RecognitionabstractIn this paper, we present a novel method to improve the flexibility of descriptor matching for image recognition by using local multiresolution pyramids in feature space. We propose that image patches be represented at multiple levels of descriptor detail and that these levels be defined in terms of local spatial pooling resolution. Preserving multiple levels of detail in local descriptors is a way of hedging one's bets on which levels will most relevant for matching during learning and recognition. We introduce the Pyramid SIFT (P-SIFT) descriptor and show that its use in four state-of-the-art image recognition pipelines improves accuracy and yields state-of-the-art results. Our technique is applicable independently of spatial pyramid matching and we show that spatial pyramids can be combined with local pyramids to obtain further improvement. We achieve state-of-the-art results on Caltech-101 (80.1%) and Caltech-256 (52.6%) when compared to other approaches based on SIFT features over intensity images. Our technique is efficient and is extremely easy to integrate into image recognition pipelines. Lorenzo Seidenari, Giuseppe Serra 0001, Andrew D. Bagdanov, Alberto Del Bimbo |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Leveraging local neighborhood topology for large scale person re-identification
Svebor Karaman, Giuseppe Lisanti, Andrew D. Bagdanov, Alberto Del Bimbo |
Pattern Recognit. | 3 |
| 2013 | Document Classification and Page Stream Segmentation for Digital Mailroom ApplicationsabstractIn this paper we present a method for the segmentation of continuous page streams into multipage documents and the simultaneous classification of the resulting documents. We first present an approach to combine the multiple pages of a document into a single feature vector that represents the whole document. Despite its simplicity and low computational cost, the proposed representation yields results comparable to more complex methods in multipage document classification tasks. We then exploit this representation in the context of page stream segmentation. The most plausible segmentation of a page stream into a sequence of multipage documents is obtained by optimizing a statistical model that represents the probability of each segmented multipage document belonging to a particular class. Experimental results are reported on a large sample of real administrative multipage documents. Albert Gordo, Marçal Rusiñol, Dimosthenis Karatzas, Andrew D. Bagdanov |
ICDAR | 4 |
| 2013 | Human action recognition using an ensemble of body-part detectorsabstractAbstract This paper describes an approach to human action recognition based on a probabilistic optimization model of body parts using hidden Markov model (HMM). Our method is able to distinguish between similar actions by only considering the body parts having major contribution to the actions, for example, legs for walking, jogging and running; arms for boxing, waving and clapping. We apply HMMs to model the stochastic movement of the body parts for action recognition. The HMM construction uses an ensemble of body‐part detectors, followed by grouping of part detections, to perform human identification. Three example‐based body‐part detectors are trained to detect three components of the human body: the head, legs and arms. These detectors cope with viewpoint changes and self‐occlusions through the use of ten sub‐classifiers that detect body parts over a specific range of viewpoints. Each sub‐classifier is a support vector machine trained on features selected for the discriminative power for each particular part/viewpoint combination. Grouping of these detections is performed using a simple geometric constraint model that yields a viewpoint‐invariant human detector. We test our approach on three publicly available action datasets: the KTH dataset, Weizmann dataset and HumanEva dataset. Our results illustrate that with a simple and compact representation we can achieve robust recognition of human actions comparable to the most complex, state‐of‐the‐art methods. Bhaskar Chakraborty, Andrew D. Bagdanov, Jordi Gonzàlez 0001, F. Xavier Roca |
Expert Syst. J. Knowl. Eng. | 2 |
| 2013 | Coloring Action Recognition in Still Images
Fahad Shahbaz Khan, Rao Muhammad Anwer, Joost van de Weijer 0001, Andrew D. Bagdanov, Antonio M. López 0001, Michael Felsberg |
Int. J. Comput. Vis. | 4 |
| 2012 | Color attributes for object detectionabstractState-of-the-art object detectors typically use shape information as a low level feature representation to capture the local structure of an object. This paper shows that early fusion of shape and color, as is popular in image classification, leads to a significant drop in performance for object detection. Moreover, such approaches also yields suboptimal results for object categories with varying importance of color and shape. In this paper we propose the use of color attributes as an explicit color representation for object detection. Color attributes are compact, computationally efficient, and when combined with traditional shape features provide state-of-the-art results for object detection. Our method is tested on the PASCAL VOC 2007 and 2009 datasets and results clearly show that our method improves over state-of-the-art techniques despite its simplicity. We also introduce a new dataset consisting of cartoon character images in which color plays a pivotal role. On this dataset, our approach yields a significant gain of 14% in mean AP over conventional state-of-the-art methods. Fahad Shahbaz Khan, Rao Muhammad Anwer, Joost van de Weijer 0001, Andrew D. Bagdanov, María Vanrell 0001, Antonio M. López 0001 |
CVPR | 4 |
| 2012 | Multi-pose face detection for accurate face logging
Andrew D. Bagdanov, Alberto Del Bimbo, Giuseppe Lisanti, Iacopo Masi |
ICPR | 1 |
| 2012 | Real-time hand status recognition from RGB-D imagery
Andrew D. Bagdanov, Alberto Del Bimbo, Lorenzo Seidenari, Lorenzo Usai |
ICPR | 1 |
| 2012 | Multipage document retrieval by textual and visual representations
Marçal Rusiñol, Dimosthenis Karatzas, Andrew D. Bagdanov, Josep Lladós 0001 |
ICPR | 3 |
| 2012 | Harmony Potentials - Fusing Global and Local Scale for Semantic Image Segmentation
Xavier Boix, Josep M. Gonfaus, Joost van de Weijer 0001, Andrew D. Bagdanov, Joan Serrat 0002, Jordi Gonzàlez 0001 |
Int. J. Comput. Vis. | 4 |
| 2011 | Adaptive Video Compression for Video Surveillance ApplicationsabstractThis article describes an approach to adaptive video coding for video surveillance applications. Using a combination of low-level features with low computational cost, we show how it is possible to control the quality of video compression so that semantically meaningful elements of the scene are encoded with higher fidelity, while background elements are allocated fewer bits in the transmitted representation. Our approach is based on adaptive smoothing of individual video frames so that image features highly correlated to semantically interesting objects are preserved. Using only low-level image features on individual frames, this adaptive smoothing can be seamlessly inserted into a video coding pipeline as a pre-processing state. Experiments show that our technique is efficient, outperforms standard H.264 encoding at comparable bit rates, and preserves features critical for downstream detection and recognition. Andrew D. Bagdanov, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari |
ISM | 1 |
| 2011 | Portmanteau Vocabularies for Multi-Cue Image RepresentationabstractWe describe a novel technique for feature combination in the bag-of-words model of image classification. Our approach builds discriminative compound words from primitive cues learned independently from training images. Our main observation is that modeling joint-cue distributions independently is more statistically robust for typical classification problems than attempting to empirically estimate the dependent, joint-cue distribution directly. We use Information theoretic vocabulary compression to find discriminative combinations of cues and the resulting vocabulary of portmanteau words is compact, has the cue binding property, and supports individual weighting of cues in the final image representation. State-of-the-art results on both the Oxford Flower-102 and Caltech-UCSD Bird-200 datasets demonstrate the effectiveness of our technique compared to other, significantly more complex approaches to multi-cue image representation Fahad Shahbaz Khan, Joost van de Weijer 0001, Andrew D. Bagdanov, María Vanrell 0001 |
NIPS | 3 |
| 2011 | Efficient discriminative multiresolution cascade for real-time human detection applications
Marco Pedersoli, Jordi Gonzàlez 0001, Andrew D. Bagdanov, F. Xavier Roca |
Pattern Recognit. Lett. | 3 |
| 2011 | Accurate Moving Cast Shadow Suppression Based on Local Color Constancy DetectionabstractThis paper describes a novel framework for detection and suppression of properly shadowed regions for most possible scenarios occurring in real video sequences. Our approach requires no prior knowledge about the scene, nor is it restricted to specific scene structures. Furthermore, the technique can detect both achromatic and chromatic shadows even in the presence of camouflage that occurs when foreground regions are very similar in color to shadowed regions. The method exploits local color constancy properties due to reflectance suppression over shadowed regions. To detect shadowed regions in a scene, the values of the background image are divided by values of the current frame in the RGB color space. We show how this luminance ratio can be used to identify segments with low gradient constancy, which in turn distinguish shadows from foreground. Experimental results on a collection of publicly available datasets illustrate the superior performance of our method compared with the most sophisticated, state-of-the-art shadow detection algorithms. These results show that our approach is robust and accurate over a broad range of shadow types and challenging video conditions. Ariel Amato, Mikhail G. Mozerov, Andrew D. Bagdanov, Jordi Gonzàlez 0001 |
IEEE Trans. Image Process. | 3 |
| 2010 | Harmony potentials for joint classification and segmentationabstractHierarchical conditional random fields have been successfully applied to object segmentation. One reason is their ability to incorporate contextual information at different scales. However, these models do not allow multiple labels to be assigned to a single node. At higher scales in the image, this yields an oversimplified model, since multiple classes can be reasonable expected to appear within one region. This simplified model especially limits the impact that observations at larger scales may have on the CRF model. Neglecting the information at larger scales is undesirable since class-label estimates based on these scales are more reliable than at smaller, noisier scales. To address this problem, we propose a new potential, called harmony potential, which can encode any possible combination of class labels. We propose an effective sampling strategy that renders tractable the underlying optimization problem. Results show that our approach obtains state-of-the-art results on two challenging datasets: Pascal VOC 2009 and MSRC-21. Josep M. Gonfaus, Xavier Boix, Joost van de Weijer 0001, Andrew D. Bagdanov, Joan Serrat 0002, Jordi Gonzàlez 0001 |
CVPR | 4 |
| 2010 | Recursive Coarse-to-Fine Localization for Fast Object Detection
Marco Pedersoli, Jordi Gonzàlez 0001, Andrew D. Bagdanov, Juan José Villanueva |
ECCV (6) | 3 |
| 2010 | Reactive Object Tracking with a Single PTZ CameraabstractIn this paper we describe a novel approach to reactive tracking of moving targets with a pan-tilt-zoom camera. The approach uses an extended Kalman filter to jointly track the object position in the real world, its velocity in 3D and the camera intrinsics, in addition to the rate of change of these parameters. The filter outputs are used as inputs to PID controllers which continuously adjust the camera motion in order to reactively track the object at a constant image velocity while simultaneously maintaining a desirable target scale in the image plane. We provide experimental results on simulated and real tracking sequences to show how our tracker is able to accurately estimate both 3D object position and camera intrinsics with very high precision over a wide range of focal lengths. Murad Al Haj, Andrew D. Bagdanov, Jordi Gonzàlez 0001, F. Xavier Roca |
ICPR | 2 |
| 2007 | Improving the robustness of particle filter-based visual trackers using online parameter adaptationabstractIn particle filter-based visual trackers, dynamic velocity components are typically incorporated into the state update equations. In these cases, there is a risk that the uncertainty in the model update stage can become amplified in unexpected and undesirable ways, leading to erroneous behavior of the tracker. Moreover, the use of a weak appearance model can make the estimates provided by the particle filter inaccurate. To deal with this problem, we propose a continuously adaptive approach to estimating uncertainty in the particle filter, one that balances the uncertainty in its static and dynamic elements. We provide quantitative performance evaluation of the resulting particle filter tracker on a set of ten video sequences. Results are reported in terms of a metric that can be used to objectively evaluate the performance of visual trackers. This metric is used to compare our modified particle filter tracker and the continuously adaptive mean shift tracker. Results show that the performance of the particle filter is significantly improved through adaptive parameter estimation, particularly in cases of occlusion and erratic, nonlinear target motion. Andrew D. Bagdanov, Alberto Del Bimbo, Fabrizio Dini, Walter Nunziati |
AVSS | 1 |
| 2006 | Learning Foveal Sensing Strategies in Unconstrained Surveillance EnvironmentsabstractIn this paper we report on techniques for automatically learning foveal sensing strategies for an active pan-tiltzoom camera. The approach uses reinforcement learning to discover foveal actions maximizing the performance of visual detectors, that are in turn assumed to be highly correlated with the task at hand. In our case, the main goal is to recognize people, hence a frontal face detection module is employed. The system uses reinforcement learning to learn if, when and how to foveate on a subject, based on its previous experience in terms or successful actions in similar situations. An action is successful if it leads to a correct face detection in the high resolution images obtained when the subject is zoomed in. In contrast with existing methods, the proposed approach obviates the need for camera calibration and camera performance modeling. Also, the method does not rely on active tracking of targets. Experimental results show how the system is capable of learning foveation strategies without requiring extensive a priori information or environmental models. Results also illustrate how the system effectively learns a strategy that allows the camera to foveate only in situations where successful detection is highly likely. Andrew D. Bagdanov, Alberto Del Bimbo, Walter Nunziati, Federico Pernici |
AVSS | 1 |
| 2006 | Boosting Color Saliency in Image Feature DetectionabstractThe aim of salient feature detection is to find distinctive local events in images. Salient features are generally determined from the local differential structure of images. They focus on the shape-saliency of the local neighborhood. The majority of these detectors are luminance-based, which has the disadvantage that the distinctiveness of the local color information is completely ignored in determining salient image features. To fully exploit the possibilities of salient point detection in color images, color distinctiveness should be taken into account in addition to shape distinctiveness. In this paper, color distinctiveness is explicitly incorporated into the design of saliency detection. The algorithm, called color saliency boosting, is based on an analysis of the statistics of color image derivatives. Color saliency boosting is designed as a generic method easily adaptable to existing feature detectors. Results show that substantial improvements in information content are acquired by targeting color salient features. Joost van de Weijer 0001, Theo Gevers, Andrew D. Bagdanov |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Towards User Transparent Data and Task Parallel Image and Video Processing: An Overview of the Parallel-Horus Project
Frank J. Seinstra, Dennis C. Koelma, Andrew D. Bagdanov |
Euro-Par | 3 |
| 2004 | Finite State Machine-Based Optimization of Data Parallel Regular Domain Problems Applied in Low-Level Image ProcessingabstractA popular approach to providing nonexperts in parallel computing with an easy-to-use programming model is to design a software library consisting of a set of preparallelized routines, and hide the intricacies of parallelization behind the library's API. However, for regular domain problems (such as simple matrix manipulations or low-level image processing applications-in which all elements in a regular subset of a dense data field are accessed in turn) speedup obtained with many such library-based parallelization tools is often suboptimal. This is because interoperation optimization (or: time-optimization of communication steps across library calls) is generally not incorporated in the library implementations. We present a simple, efficient, finite state machine-based approach for communication minimization of library-based data parallel regular domain problems. In the approach, referred to as lazy parallelization, a sequential program is parallelized automatically at runtime by inserting communication primitives and memory management operations whenever necessary. Apart from being simple and cheap, lazy parallelization guarantees to generate legal, correct, and efficient parallel programs at all times. The effectiveness of the approach is demonstrated by analyzing the performance characteristics of two typical regular domain problems obtained from the field of low-level image processing. Experimental results show significant performance improvements over nonoptimized parallel applications. Moreover, obtained communication behavior is found to be optimal with respect to the abstraction level of message passing programs. Frank J. Seinstra, Dennis C. Koelma, Andrew D. Bagdanov |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2003 | Understanding Document Analysis and Understanding (through Modeling)abstractWe turn to the viewpoint of users of a DAU system. Out of the view of users we sketch a picture of “Document Analysis and Understanding” (DAU), only a simple division of DAU into six sub-tasks, and consider this a model of DAU. We provide one elaborate example of a use of the model: the module design of the commercially successful DAU system smartFIX. We argue that such modeling of DAU can be of benefit to the whole field of DA Bertin Klein, Stefan Agne, Andrew D. Bagdanov |
ICDAR | 3 |
| 2003 | First order Gaussian graphs for efficient structure classification
Andrew D. Bagdanov, Marcel Worring |
Pattern Recognit. | 1 |
| 2002 | Multi-scale Document Description Using Rectangular Granulometries
Andrew D. Bagdanov, Marcel Worring |
Document Analysis Systems | 1 |
| 2002 | Interactive Indexing and Retrieval of Multimedia Content
Marcel Worring, Andrew D. Bagdanov, Jan C. van Gemert, Jan-Mark Geusebroek, Hoang Minh, Guus Schreiber, Cees Snoek, Jeroen Vendrig, Jan Wielemaker, Arnold W. M. Smeulders |
SOFSEM | 2 |
| 2001 | Fine-Grained Document Genre Classification Using First Order Random GraphsabstractWe approach the general problem of classifying machine-printed documents into genres. Layout is a critical factor in recognizing fine-grained genres, as document content features are similar. Document genre is determined from the layout structure detected from scanned binary images of the document pages, using no OCR results and minimal a priori knowledge of document logical structures. Our method uses the attributed relational graphs (ARGs) to represent the layout structure of document instances, and the first order random graphs (FORGs) to represent document genres. In this paper we develop our FORG-based genre classification method and present a comparative evaluation between our technique and a variety of statistical pattern classifiers. FORGs are capable of modeling common layout structure within a document genre and are shown to significantly outperform traditional pattern classification techniques when fine-grained genre distinctions must be drawn. Andrew D. Bagdanov, Marcel Worring |
ICDAR | 1 |
| 1998 | Projection profile based skew estimation algorithm for JBIG compressed images
Junichi Kanai, Andrew D. Bagdanov |
Int. J. Document Anal. Recognit. | 2 |
| 1997 | Projection profile based skew estimation algorithm for JBIG compressed imagesabstractA new projection profile based algorithm that extracts fiducial points needed to estimate a skew angle by decoding a JBIG compressed image is presented. This algorithm and three other projection profile based algorithms were tested using 460 page images and 1246 single column test zones extracted from the page images. Linear regression analyses of the experimental results showed that the new algorithm performed competitively with the other three algorithms. Andrew D. Bagdanov, Junichi Kanai |
ICDAR | 1 |