VLDB 2026 Research / reviewers in the wild / expert
Vittorio Murino
dblp:62/6790
· DBLP profile ↗
252ranked-venue papers
23as first author
33since 2021 · last 2026
0000-0002-8645-2328ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 160 · 10 first-author · 22 since 2021Artificial intelligence and machine learning · 147 · 13 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-authorSystems, architecture and hardware · 3Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SplatFill: 3D Scene Inpainting via Depth-Guided Gaussian Splatting
Mahtab Dahaghin, Milind Gajanan Padalkar, Matteo Toso, Alessio Del Bue, Vittorio Murino |
ICPR (16) | 5 |
| 2026 | HAC: Parameter-Efficient Hyperbolic Adaptation of CLIP for Zero-Shot VQA
Francesco Dibitonto, Cigdem Beyan, Vittorio Murino |
ICPR (3) | 3 |
| 2026 | Freq2Clean: Enhancing Calcium Imaging Denoising via Frequency-Domain Fusion
Valerio Morelli, Daniele Berardini, Giorgio Letti, Sebastiano Curreli, Adriano Mancini, Tommaso Fellin, Vittorio Murino |
ICPR (12) | 7 |
| 2026 | Discriminator-Guided Adaptive Diffusion for Source-Free Test-Time Adaptation Under Image Corruptions
Francesco Olivato, Cigdem Beyan, Vittorio Murino |
ICPR (8) | 3 |
| 2026 | Lose Your Self (LoYS): an adversarial entropy-based unsupervised approach for model debiasingabstractWhen spurious correlations between targets and data are present in training samples, deep neural networks may struggle to generalize, typically learning shortcuts corresponding to such undesired correlations (i.e., the bias), rather than fundamental target attributes. If bias attributes are assumed to be known, several strategies can be applied to mitigate a model’s dependency on bias, including upsampling or upweighting samples with no bias. However, this is hardly the case in real-world scenarios, and alternative unsupervised debiasing approaches have been proposed in recent years, assuming no bias information is available. In this work, we propose Lose Your Self (LoYS), a novel bias-unsupervised approach for model debiasing. The main design strategy in LoYS aims to force a model learning semantic features, discouraging it from learning bias-related descriptors, specifically focusing on the target classifier confidence. Exploiting an adversarial scheme based on entropy loss, against a shallow auxiliary classifier trained to match the predictions of a pre-trained biased model, and entropy regularizations, our experiments show how LoYS is competitive or outperforms state-of-the-art methods, on several common benchmarks. Vito Paolo Pastore, Massimiliano Ciranni, Vittorio Murino |
WACV | 3 |
| 2025 | MadCLIP: Few-Shot Medical Anomaly Detection with CLIP
Mahshid Shiri, Cigdem Beyan, Vittorio Murino |
MICCAI (6) | 3 |
| 2025 | Direction-Aware Room Impulse Response Estimation for Immersive Audio Rendering in Real EnvironmentsabstractEvolving multimedia systems are increasingly being adopted in virtual reality and gaming applications. Such systems emphasize immersion to engage users by bridging the gap between real and virtual content. In this context, visual and acoustic stimuli are the two key media that dictate such immersion. While visual 3D rendering is advancing rapidly, the same is not true for audio, where most research is limited to the reconstruction of the room impulse response (RIR) using omnidirectional audio or, at best, binaural. Such methods do not adequately account for the directions and orientations of the acoustic signals with respect to either the source or the listener, thereby compromising immersion quality. In this work, we explore the effect of adding such "directionality" to the training data to improve the estimation of the room’s acoustic parameters. A more accurate set of such parameters implies in fact a more realistic predicted RIR, leading to a more immersive experience of the acoustic scene. Specifically, we propose a novel framework driven by a suitable loss function to account for directionality in ambisonic microphones, and novel variants of loss functions for both omnidirectional and ambisonic cases. We also propose to account for microphone characteristics and their contribution to the predicted RIRs. Experiments were performed using two datasets of real recordings and the results established the efficacy of the proposed methods Giovanni Zanin, Ritujoy Biswas, Pietro Morerio, Sylvio Barbon Junior, Alberto Carini, Alessio Del Bue, Vittorio Murino |
ACM Multimedia | 7 |
| 2025 | Diffusing DeBias: Synthetic Bias Amplification for Model DebiasingabstractThe effectiveness of deep learning models in classification tasks is often challenged by the quality and quantity of training data whenever they are affected by strong spurious correlations between specific attributes and target labels. This results in a form of bias affecting training data, which typically leads to unrecoverable weak generalization in prediction. This paper addresses this problem by leveraging bias amplification with generated synthetic data only: we introduce Diffusing DeBias (DDB), a novel approach acting as a plug-in for common methods of unsupervised model debiasing, exploiting the inherent bias-learning tendency of diffusion models in data generation. Specifically, our approach adopts conditional diffusion models to generate synthetic bias-aligned images, which fully replace the original training set for learning an effective bias amplifier model to be subsequently incorporated into an end-to-end and a two-step unsupervised debiasing approach. By tackling the fundamental issue of bias-conflicting training samples’ memorization in learning auxiliary models, typical of this type of technique, our proposed method outperforms the current state-of-the-art in multiple benchmark datasets, demonstrating its potential as a versatile and effective tool for tackling bias in deep learning models. Code is available at https://github.com/Malga-Vision/DiffusingDeBias Massimiliano Ciranni, Vito Paolo Pastore, Roberto Di Via, Enzo Tartaglione, Francesca Odone, Vittorio Murino |
NeurIPS | 6 |
| 2025 | Looking at Model Debiasing through the Lens of Anomaly DetectionabstractDeep neural networks are likely to learn unintended spurious correlations between training data and labels when dealing with biased data, potentially limiting the generalization to unseen samples not presenting the same bias. In this context, model debiasing approaches can be de-vised aiming at reducing the model's dependency on such unwanted correlations, either leveraging the knowledge of bias information or not. In this work, we focus on the latter and more realistic scenario, showing the importance of accurately predicting the bias-conflicting and bias-aligned samples to obtain compelling performance in bias mitigation. On this ground, we propose to conceive the problem of model bias from an out-of-distribution perspective, intro-ducing a new bias identification method based on anomaly detection. We claim that when data is mostly biased, bias-conflicting samples can be regarded as outliers with respect to the bias-aligned distribution in the feature space of a bi-ased model, thus allowing for precisely detecting them with an anomaly detection method. Coupling the proposed bias identification approach with bias-conflicting data upsampling and augmentation in a two-step strategy, we reach state-of-the-art performance on synthetic and real benchmark datasets. Ultimately, our proposed approach shows that the data bias issue does not necessarily require complex debiasing methods, given that an accurate bias identification procedure is defined. Source code is available at https://github.com/Malga-Vision/MoDAD Vito Paolo Pastore, Massimiliano Ciranni, Davide Marinelli, Francesca Odone, Vittorio Murino |
WACV | 5 |
| 2025 | Pre-trained Multiple Latent Variable Generative Models are Good Defenders Against Adversarial AttacksabstractAttackers can deliberately perturb classifiers' input with subtle noise, altering final predictions. Among proposed countermeasures, adversarial purification employs generative networks to preprocess input images, filtering out adversarial noise. In this study, we propose specific generators, defined Multiple Latent Variable Generative Models (MLVGMs), for adversarial purification. These models possess multiple latent variables that naturally disentangle coarse from fine features. Taking advantage of these properties, we autoencode images to maintain class-relevant information, while discarding and re-sampling any detail, including adversarial noise. The procedure is completely training-free, exploring the generalization abilities of pretrained MLVGMs on the adversarial purification down-stream task. Despite the lack of large models, trained on billions of samples, we show that smaller MLVGMs are already competitive with traditional methods, and can be used as foundation models. Official code released at https://github.com/SerezD/gen_adversarial. Dario Serez, Marco Cristani, Alessio Del Bue, Vittorio Murino, Pietro Morerio |
WACV | 4 |
| 2025 | Guest Editorial: Special Issue on Multimodal Learning
Michael Ying Yang, Paolo Rota, Massimiliano Mancini, Pietro Morerio, Bodo Rosenhahn, Vittorio Murino |
Int. J. Comput. Vis. | 6 |
| 2024 | Leveraging Next-Active Objects for Context-Aware Anticipation in Egocentric VideosabstractObjects are crucial for understanding human-object interactions. By identifying the relevant objects, one can also predict potential future interactions or actions that may occur with these objects. In this paper, we study the problem of Short-Term Object interaction anticipation (STA) and propose NAOGAT (Next-Active-Object Guided Anticipation Transformer), a multi-modal end-to-end transformer network, that attends to objects in observed frames in order to anticipate the next-active-object (NAO) and, eventually, to guide the model to predict context-aware future actions. The task is challenging since it requires anticipating future action along with the object with which the action occurs and the time after which the interaction will begin, a.k.a. the time to contact (TTC). Compared to existing video modeling architectures for action anticipation, NAOGAT captures the relationship between objects and the global scene context in order to predict detections for the next active object and anticipate relevant future actions given these detections, leveraging the objects’ dynamics to improve accuracy. One of the key strengths of our approach, in fact, is its ability to exploit the motion dynamics of objects within a given clip , which is often ignored by other models, and separately decoding the object-centric and motion-centric information. Through our experiments, we show that our model outperforms existing methods on two separate datasets, Ego4D and EpicKitchens-100 ("Unseen Set"), as measured by several additional metrics, such as time to contact, and next-active-object localization. The code can be found on project page : sanketsans.github.io/wacv24 Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino, Alessio Del Bue |
WACV | 4 |
| 2024 | Computer vision and deep learning meet plankton: Milestones and future directionsabstractPlanktonic organisms play a pivotal role within aquatic ecosystems, serving as the foundation of the aquatic food chain while also playing a critical role in climate regulation and the production of oxygen. In recent years, the advent of automated systems for capturing in-situ images has led to a huge influx of plankton images, making manual classification impractical. This, at the same time, has opened up opportunities for the application of machine learning and deep learning solutions. This paper undertakes an extensive analysis of the broad range of computer vision techniques and methodologies that have emerged to facilitate the automatic analysis of small- to large-scale datasets containing plankton images. By focusing on different computer vision tasks, we present findings and limitations in order to offer a comprehensive overview of the current state-of-the-art, while also pinpointing the open challenges that demand further research and attention. Massimiliano Ciranni, Vittorio Murino, Francesca Odone, Vito Paolo Pastore |
Image Vis. Comput. | 2 |
| 2023 | Learnable Data Augmentation for One-Shot Unsupervised Domain Adaptation
Julio Ivan Davila Carrazco, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
BMVC | 4 |
| 2023 | Audio-Visual Inpainting: Reconstructing Missing Visual Information with SoundabstractWe tackle audio-visual inpainting, the problem of completing an image in such a way to be consistent with the sound associated to the scene. To this end, we propose a multimodal, audio-visual inpainting method (AVIN), and show how to leverage sound to reconstruct semantically consistent images. AVIN is a 2-stage algorithm, which first learns the scene semantics and reconstructs low resolution images based on a conditional probability distribution of pixels in the space conditioned to audio, and then refines such result with a GAN-based network to increase the resolution of the reconstructed image. We show that AVIN is able to recover the original content, especially in the hard cases where the missing area heavily degrades the scene semantics: it can perform cross-modal generation whenever no visual context is observed at all, reconstructing visual data from sound only. Code will be made available upon acceptance. Valentina Sanguineti, Sanket Kumar Thakur, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
ICASSP | 5 |
| 2023 | Enhancing Next Active Object-Based Egocentric Action Anticipation with Guided AttentionabstractShort-term action anticipation (STA) in first-person videos is a challenging task that involves understanding the next active object interactions and predicting future actions. Existing action anticipation methods have primarily focused on utilizing features extracted from video clips, but often overlooked the importance of objects and their interactions. To this end, we propose a novel approach that applies a guided attention mechanism between the objects, and the spatiotemporal features extracted from video clips, enhancing the motion and contextual information, and further decoding the object-centric and motion-centric information to address the problem of STA in egocentric videos. Our method, GANO (Guided Attention for Next active Objects) is a multi-modal, end-to-end, single transformer-based network. The experimental results performed on the largest egocentric dataset demonstrate that GANO outperforms the existing state-of-the-art methods for the prediction of the next active object label, its bounding box location, the corresponding future action, and the time to contact the object. The ablation study shows the positive contribution of the guided attention mechanism compared to other fusion methods. Moreover, it is possible to improve the next active object location and class label prediction results of GANO by just appending the learnable object tokens with the region of interest embeddings. Related implementations are available at: sanketsans.github.io/guided-attention-egocentric.html Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino, Alessio Del Bue |
ICIP | 4 |
| 2023 | Guest Editorial : Learning with Manifolds in Computer Vision
Mohamed Daoudi, Mehrtash Harandi, Vittorio Murino |
Image Vis. Comput. | 3 |
| 2023 | Efficient unsupervised learning of biological images with compressed deep featuresabstractMachine learning has significantly impacted the analysis of biological images and is now an important part of many biological data analysis pipelines. A variety of biological and biomedical domain-related tasks is gaining benefit from image analysis and pattern recognition tools developed currently. Applications include diagnostic histopathology, environmental monitoring, synthetic biology, genomics, and proteomics. Particularly in the last decade, several deep learning and advanced computer vision methods such as convolutional neural networks (CNNs), typically trained in a supervised fashion, have started to be largely employed in biological image classification. Moreover, the advancement of automatic acquisition systems has been generating a massive amount of biological data, which requires to be analyzed by domain experts. However, the cost of manual annotation of such data has become a bottleneck, impairing the application of supervised machine learning algorithms. Biological images generally have an intrinsic high variability, whose identity is sometimes hard to assign and strongly dependent on the annotator’s expertise. In this context, a limited number of annotation-free (i.e., unsupervised) learning solutions have been proposed, typically based on hand-crafted features, specifically tailored for a certain biological domain. Nonetheless, a successful unsupervised learning approach must be accurate, and sufficiently robust to deal with different biological domains. This paper aims at providing a viable solution to these issues, proposing an unsupervised learning algorithm based on compressed deep features for image classification. We exploit features extracted from ImageNet pre-trained transformers and CNNs, further compressed with a customized β-Variational AutoEncoder (β-VAE), that we call reconstruction VAE (R-VAE). We test our algorithm on biological images coming from diverse domains characterized by high variability in shape and texture information and acquired with widely differing imaging platforms. Considered image datasets range from multi-cellular organisms (plankton, coral) to sub-cellular organelles (budding yeast vacuoles, human cells’ nuclei, etc.). Our results show that the compressed deep features extracted from different pre-trained vision models establish new unsupervised learning state-of-the-art performances for the investigated datasets. Vito Paolo Pastore, Massimiliano Ciranni, Simone Bianco 0002, Jennifer Carol Fung, Vittorio Murino, Francesca Odone |
Image Vis. Comput. | 5 |
| 2023 | No Adversaries to Zero-Shot Learning: Distilling an Ensemble of Gaussian Feature GeneratorsabstractIn zero-shot learning (ZSL), the task of recognizing unseen categories when no data for training is available, state-of-the-art methods generate visual features from semantic auxiliary information (e.g., attributes). In this work, we propose a valid alternative (simpler, yet better scoring) to fulfill the very same task. We observe that, if first- and second-order statistics of the classes to be recognized were known, sampling from Gaussian distributions would synthesize visual features that are almost identical to the real ones as per classification purposes. We propose a novel mathematical framework to estimate first- and second-order statistics, even for unseen classes: our framework builds upon prior compatibility functions for ZSL and does not require additional training. Endowed with such statistics, we take advantage of a pool of class-specific Gaussian distributions to solve the feature generation stage through sampling. We exploit an ensemble mechanism to aggregate a pool of softmax classifiers, each trained in a one-seen-class-out fashion to better balance the performance over seen and unseen classes. Neural distillation is finally applied to fuse the ensemble into a single architecture which can perform inference through one forward pass only. Our method, termed Distilled Ensemble of Gaussian Generators, scores favorably with respect to state-of-the-art works. Jacopo Cavazza, Vittorio Murino, Alessio Del Bue |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Unsupervised Domain Adaptation for Video Transformers in Action RecognitionabstractOver the last few years, Unsupervised Domain Adaptation (UDA) techniques have acquired remarkable importance and popularity in computer vision. However, when compared to the extensive literature available for images, the field of videos is still relatively unexplored. On the other hand, the performance of a model in action recognition is heavily affected by domain shift. In this paper, we propose a simple and novel UDA approach for video action recognition. Our approach leverages recent advances on spatio-temporal transformers to build a robust source model that better generalises to the target domain. Furthermore, our architecture learns domain invariant features thanks to the introduction of a novel alignment loss term derived from the Information Bottleneck principle. We report results on two video action recognition benchmarks for UDA, showing state-of-the-art performance on HMDB ↔ UCF, as well as on Kinetics→NEC-Drone, which is more challenging. This demonstrates the effectiveness of our method in handling different levels of domain shift. The source code is available at https://github.com/vturrisi/UDAVT. Victor G. T. da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, Elisa Ricci 0001 |
ICPR | 6 |
| 2022 | Cleaning Noisy Labels by Negative Ensemble Learning for Source-Free Unsupervised Domain AdaptationabstractConventional Unsupervised Domain Adaptation (UDA) methods presume source and target domain data to be simultaneously available during training. Such an assumption may not hold in practice, as source data is often inaccessible (e.g., due to privacy reasons). On the contrary, a pre-trained source model is usually available, which performs poorly on target due to the well-known domain shift problem. This translates into a significant amount of misclassifications, which can be interpreted as structured noise affecting the inferred target pseudo-labels. In this work, we cast UDA as a pseudo-label refinery problem in the challenging source-free scenario. We propose Negative Ensemble Learning (NEL) technique, a unified method for adaptive noise filtering and progressive pseudo-label refinement. NEL is devised to tackle noisy pseudo-labels by enhancing diversity in ensemble members with different stochastic (i) input augmentation and (ii) feedback. The latter is achieved by leveraging the novel concept of Disjoint Residual Labels, which allow propagating diverse information to the different members. Eventually, a single model is trained with the refined pseudo-labels, which leads to a robust performance on the target domain. Extensive experiments show that the proposed method achieves state-of-the-art performance on major UDA benchmarks, such as Digit5, PACS, Visda-C, and DomainNet, without using source data samples at all. Pietro Morerio, Vittorio Murino |
WACV | 3 |
| 2022 | Dual-Head Contrastive Domain Adaptation for Video Action RecognitionabstractUnsupervised domain adaptation (UDA) methods have become very popular in computer vision. However, while several techniques have been proposed for images, much less attention has been devoted to videos. This paper introduces a novel UDA approach for action recognition from videos, inspired by recent literature on contrastive learning. In particular, we propose a novel two-headed deep architecture that simultaneously adopts cross-entropy and contrastive losses from different network branches to robustly learn a target classifier. Moreover, this work introduces a novel large-scale UDA dataset, Mixamo→Kinetics, which, to the best of our knowledge, is the first dataset that considers the domain shift arising when transferring knowledge from synthetic to real video sequences. Our extensive experimental evaluation conducted on three publicly available benchmarks and on our new Mixamo→Kinetics dataset demonstrate the effectiveness of our approach, which outperforms the current state-of-the-art methods. Code is available at https://github.com/vturrisi/CO2A. Victor G. T. da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, Elisa Ricci 0001 |
WACV | 6 |
| 2022 | Guest Editorial Introduction to the Special Issue on "Biometrics Based Methods for Healthcare Applications"
Michele Nappi, Hugo Proença 0001, Sambit Bakshi, Vittorio Murino |
Comput. Vis. Image Underst. | 4 |
| 2022 | Unsupervised Synthetic Acoustic Image Generation for Audio-Visual Scene UnderstandingabstractAcoustic images are an emergent data modality for multimodal scene understanding. Such images have the peculiarity of distinguishing the spectral signature of the sound coming from different directions in space, thus providing a richer information as compared to that derived from single or binaural microphones. However, acoustic images are typically generated by cumbersome and costly microphone arrays which are not as widespread as ordinary microphones. This paper shows that it is still possible to generate acoustic images from off-the-shelf cameras equipped with only a single microphone and how they can be exploited for audio-visual scene understanding. We propose three architectures inspired by Variational Autoencoder, U-Net and adversarial models, and we assess their advantages and drawbacks. Such models are trained to generate spatialized audio by conditioning them to the associated video sequence and its corresponding monaural audio track. Our models are trained using the data collected by a microphone array as ground truth. Thus they learn to mimic the output of an array of microphones in the very same conditions. We assess the quality of the generated acoustic images considering standard generation metrics and different downstream tasks (classification, cross-modal retrieval and sound localization). We also evaluate our proposed models by considering multimodal datasets containing acoustic images, as well as datasets containing just monaural audio signals and RGB video frames. In all of the addressed downstream tasks we obtain notable performances using the generated acoustic data, when compared to the state of the art and to the results obtained using real acoustic images as input. Valentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
IEEE Trans. Image Process. | 4 |
| 2021 | Audio-Visual Localization by Synthetic Acoustic Image GenerationabstractAcoustic images constitute an emergent data modality for multimodal scene understanding. Such images have the peculiarity to distinguish the spectral signature of sounds coming from different directions in space, thus providing richer information than the one derived from mono and binaural microphones. However, acoustic images are typically generated by cumbersome microphone arrays, which are not as widespread as ordinary microphones mounted on optical cameras. To exploit this empowered modality while using standard microphones and cameras we propose to leverage the generation of synthetic acoustic images from common audio-video data for the task of audio-visual localization. The generation of synthetic acoustic images is obtained by a novel deep architecture, based on Variational Autoencoder and U-Net models, which is trained to reconstruct the ground truth spatialized audio data collected by a microphone array, from the associated video and its corresponding monaural audio signal. Namely, the model learns how to mimic what an array of microphones can produce in the same conditions. We assess the quality of the generated synthetic acoustic images on the task of unsupervised sound source localization in a qualitative and quantitative manner, while also considering standard generation metrics. Our model is evaluated by considering both multimodal datasets containing acoustic images, used for the training, and unseen datasets containing just monaural audio signals and RGB frames, showing to reach more accurate localization results as compared to the state of the art. Valentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
AAAI | 4 |
| 2021 | Distillation Multiple Choice Learning for Multimodal Action RecognitionabstractIn this work, we address the problem of learning an ensemble of specialist networks using multimodal data, while considering the realistic and challenging scenario of possible missing modalities at test time. Our goal is to leverage the complementary information of multiple modalities to the benefit of the ensemble and each individual network. We introduce a novel Distillation Multiple Choice Learning framework for multimodal data, where different modality networks learn in a cooperative setting from scratch, strengthening one another. The modality networks learned using our method achieve significantly higher accuracy than if trained separately, due to the guidance of other modalities. We evaluate this approach on three video action recognition benchmark datasets. We obtain state-of-the-art results in comparison to other approaches that work with missing modalities at test time. Nuno C. Garcia, Sarah Adel Bargal, Vitaly Ablavsky, Pietro Morerio, Vittorio Murino, Stan Sclaroff |
WACV | 5 |
| 2021 | Transductive Zero-Shot Learning by Decoupled Feature GenerationabstractIn this paper, we address zero-shot learning (ZSL), the problem of recognizing categories for which no labeled visual data are available during training. We focus on the transductive setting, in which unlabelled visual data from unseen classes is available. State-of-the-art paradigms in ZSL typically exploit generative adversarial networks to synthesize visual features from semantic attributes. We posit that the main limitation of these approaches is to adopt a single model to face two problems: 1) generating realistic visual features, and 2) translating semantic attributes into visual cues. Differently, we propose to decouple such tasks, solving them separately. In particular, we train an unconditional generator to solely capture the complexity of the distribution of visual data and we subsequently pair it with a conditional generator devoted to enrich the prior knowledge of the data distribution with the semantic content of the class embeddings. We present a detailed ablation study to dissect the effect of our proposed decoupling approach, while demonstrating its superiority over the related state-of-the-art. Federico Marmoreo, Jacopo Cavazza, Vittorio Murino |
WACV | 3 |
| 2021 | S-VVAD: Visual Voice Activity Detection by Motion SegmentationabstractWe address the challenging Voice Activity Detection (VAD) problem, which determines "Who is Speaking and When?" in audiovisual recordings. The typical audio-based VAD systems can be ineffective in the presence of ambient noise or noise variations. Moreover, due to technical or privacy reasons, audio might not be always available. In such cases, the use of video modality to perform VAD is desirable. Almost all existing visual VAD methods rely on body part detection, e.g., face, lips, or hands. In contrast, we propose a novel visual VAD method operating directly on the entire video frame, without the explicit need of detecting a person or his/her body parts. Our method, named S-VVAD, learns body motion cues associated with speech activity within a weakly supervised segmentation framework. Therefore, it not only detects the speakers/not-speakers but simultaneously localizes the image positions of them. It is an end-to-end pipeline, person-independent and it does not require any prior knowledge nor pre-processing. S-VVAD performs well in various challenging conditions and demonstrates the state-of-the-art results on multiple datasets. Moreover, the better generalization capability of S-VVAD is confirmed for cross-dataset and person-independent scenarios. Muhammad Shahid 0002, Cigdem Beyan, Vittorio Murino |
WACV | 3 |
| 2021 | Intra-Camera Supervised Person Re-IdentificationabstractAbstract Existing person re-identification (re-id) methods mostly exploit a large set of cross-camera identity labelled training data. This requires a tedious data collection and annotation process, leading to poor scalability in practical re-id applications. On the other hand unsupervised re-id methods do not need identity label information, but they usually suffer from much inferior and insufficient model performance. To overcome these fundamental limitations, we propose a novel person re-identification paradigm based on an idea ofindependentper-camera identity annotation. This eliminates the most time-consuming and tedious inter-camera identity labelling process, significantly reducing the amount of human annotation efforts. Consequently, it gives rise to a more scalable and more feasible setting, which we callIntra-Camera Supervised (ICS)person re-id, for which we formulate a Multi-tAsk mulTi-labEl (MATE) deep learning method. Specifically, MATE is designed for self-discovering the cross-camera identity correspondence in a per-camera multi-task inference framework. Extensive experiments demonstrate the cost-effectiveness superiority of our method over the alternative approaches on three large person re-id datasets. For example, MATE yields 88.7% rank-1 score on Market-1501 in the proposed ICS person re-id setting, significantly outperforming unsupervised learning models and closely approaching conventional fully supervised learning competitors. Xiangping Zhu, Xiatian Zhu, Minxian Li, Pietro Morerio, Vittorio Murino, Shaogang Gong |
Int. J. Comput. Vis. | 5 |
| 2021 | Excitation Dropout: Encouraging Plasticity in Deep Neural Networks
Andrea Zunino, Sarah Adel Bargal, Pietro Morerio, Jianming Zhang 0001, Stan Sclaroff, Vittorio Murino |
Int. J. Comput. Vis. | 6 |
| 2021 | Guided Zoom: Zooming into Network Evidence to Refine Fine-Grained Model DecisionsabstractIn state-of-the-art deep single-label classification models, the top- k (k=2,3,4, ...) accuracy is usually significantly higher than the top-1 accuracy. This is more evident in fine-grained datasets, where differences between classes are quite subtle. Exploiting the information provided in the top k predicted classes boosts the final prediction of a model. We propose Guided Zoom, a novel way in which explainability could be used to improve model performance. We do so by making sure the model has "the right reasons" for a prediction. The reason/evidence upon which a deep neural network makes a prediction is defined to be the grounding, in the pixel space, for a specific class conditional probability in the model output. Guided Zoom examines how reasonable the evidence used to make each of the top- k predictions is. Test time evidence is deemed reasonable if it is coherent with evidence used to make similar correct decisions at training time. This leads to better informed predictions. We explore a variety of grounding techniques and study their complementarity for computing evidence. We show that Guided Zoom results in an improvement of a model's classification accuracy and achieves state-of-the-art classification performance on four fine-grained classification datasets. Our code is available at https://github.com/andreazuna89/Guided-Zoom. Sarah Adel Bargal, Andrea Zunino, Vitali Petsiuk, Jianming Zhang 0001, Kate Saenko, Vittorio Murino, Stan Sclaroff |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Personality Traits Classification Using Deep Visual Activity-Based Nonverbal Features of Key-Dynamic ImagesabstractThis paper addresses nonverbal behavior analysis for the classification of perceived personality traits using novel deep visual activity (VA)-based features extracted only from key-dynamic images. Dynamic images represent short-term VA. Key-dynamic images carry more discriminative information i.e., nonverbal features (NFs) extracted from them contribute to the classification more than NFs extracted from other dynamic images. Dynamic image construction, learning long-term VA with CNN+LSTM, and detecting spatio-temporal saliency are applied to determine key-dynamic images. Once VA-based NFs are extracted, they are encoded using covariance, and resulting representation is used for classification. This method was evaluated on two datasets: small group meetings and vlogs. For the first dataset, proposed method outperforms not only the state-of-the-art VA-based methods but also multi-modal approaches for all personality traits. For extraversion classification, it performs better than i) the most popular key-frames selection algorithm, ii) random and uniform dynamic image selection, and iii) NFs extracted from all dynamic images. Furthermore, the ablation study proves the superiority of proposed method. For the further dataset, it performs as well as the state-of-the-art visual-NFs on average, while showing improved performance for agreeableness classification. Proposed method can be adapted to any application based on nonverbal behavior analysis, thanks to being data-driven. Cigdem Beyan, Andrea Zunino, Muhammad Shahid 0002, Vittorio Murino |
IEEE Trans. Affect. Comput. | 4 |
| 2021 | RealVAD: A Real-World Dataset and A Method for Voice Activity Detection by Body Motion AnalysisabstractWe present an automatic voice activity detection (VAD) method that is solely based on visual cues. Unlike traditional approaches processing audio, we show that upper body motion analysis is desirable for the VAD task. The proposed method consists of components for body motion representation, feature extraction from a Convolutional Neural Network (CNN) architecture and unsupervised domain adaptation. The body motion representations as images are used by the feature extraction component, which is generic and person-invariant, thus, can be applied to a subject who has never been seen. The endmost component handles the domain-shift problem, which appears due to the fact that the way people move/ gesticulate while speaking might vary from subject to subject, which results in disparate body motion features and consequently poorer VAD performance. The experimental analyses applied on a publicly available real-world VAD dataset show that the proposed method performs better than the state-of-the-art video-only and multimodal VAD approaches. Moreover, the proposed method has a better generalization ability as VAD results are more consistent across different subjects. As another major contribution, we present a new multimodal dataset (called RealVAD), created from a real-world (no role-plays) panel discussion. This dataset contains many actual situations/ challenges that are missing in the previous VAD datasets. We benchmarked the RealVAD dataset by applying the proposed method as well as cross-dataset analyses. Particularly, the results of cross-dataset experiments highlight the remarkable positive contribution of the unsupervised domain adaptation applied. Cigdem Beyan, Muhammad Shahid 0002, Vittorio Murino |
IEEE Trans. Multim. | 3 |
| 2020 | Leveraging Acoustic Images for Effective Self-supervised Audio Representation Learning
Valentina Sanguineti, Pietro Morerio, Niccolò Pozzetti, Danilo Greco, Marco Cristani, Vittorio Murino |
ECCV (22) | 6 |
| 2020 | Complex-Object Visual Inspection: Empirical Studies on A Multiple Lighting SolutionabstractThe design of an automatic visual inspection system is usually performed in two stages. While the first stage consists in selecting the most suitable hardware setup for highlighting most effectively the defects on the surface to be inspected, the second stage concerns the development of algorithmic solutions to exploit the potentials offered by the collected data. In this paper, first, we present a novel illumination setup embedding four illumination configurations to resemble diffused, dark-field, and front lighting techniques. Second, we analyze the contributions brought by deploying the proposed setup in the training phase only, mimicking the scenario in which an already developed visual inspection system cannot be modified on the customer site. Along with an exhaustive set of experiments, in this paper, we demonstrate the suitability of the proposed setup for effective illumination of complex-objects, defined as manufactured items with variable surface characteristics that cannot be determined a priori. Eventually, we provide insights into the importance of multiple light configurations availability during training and their natural boosting effect which, without the need to modify the system design in the evaluation phase, lead to improvements in the overall system performance. Maya Aghaei, Matteo Bustreo, Pietro Morerio, Nicolò Carissimi, Alessio Del Bue, Vittorio Murino |
ICPR | 6 |
| 2020 | Compact CNN Structure Learning by Knowledge DistillationabstractThe concept of compressing deep Convolutional Neural Networks (CNNs) is essential to use limited computation, power, and memory resources on embedded devices. However, existing methods achieve this objective at the cost of a drop in inference accuracy in computer vision tasks. To address such a drawback, we propose a framework that leverages knowledge distillation along with customizable block-wise optimization to learn a lightweight CNN structure while preserving better control over the compression-performance tradeoff. Considering specific resource constraints, e.g., floating-point operations per inference (FLOPs) or model-parameters, our method results in a state of the art network compression while being capable of achieving better inference accuracy. In a comprehensive evaluation, we demonstrate that our method is effective, robust, and consistent with results over a variety of network architectures and datasets, at negligible training overhead. In particular, for the already compact network MobileNet_v2, our method offers up to 2× and 5.2× better model compression in terms of FLOPs and model-parameters, respectively, while getting 1.05% better model performance than the baseline network. Andrea Zunino, Pietro Morerio, Vittorio Murino |
ICPR | 4 |
| 2020 | Weakly Supervised Geodesic Segmentation of Egyptian Mummy CT ScansabstractIn this paper, we tackle the task of automatically analyzing 3D volumetric scans obtained from computed tomography (CT) devices. In particular, we address a particular task for which data is very limited: the segmentation of ancient Egyptian mummies CT scans. We aim at digitally unwrapping the mummy and identify different segments such as body, bandages and jewelry. The problem is complex because of the lack of annotated data for the different semantic regions to segment, thus discouraging the use of strongly supervised approaches. We, therefore, propose a weakly supervised and efficient interactive segmentation method to solve this challenging problem. After segmenting the wrapped mummy from its exterior region using histogram analysis and template matching, we first design a voxel distance measure to find an approximate solution for the body and bandage segments. Here, we use geodesic distances since voxel features as well as spatial relationship among voxels is incorporated in this measure. Next, we refine the solution using a GrabCut based segmentation together with a tracking method on the slices of the scan that assigns labels to different regions in the volume, using limited supervision in the form of scribbles drawn by the user. The efficiency of the proposed method is demonstrated using visualizations and validated through quantitative measures and qualitative unwrapping of the mummy. Avik Hati, Matteo Bustreo, Diego Sona, Vittorio Murino, Alessio Del Bue |
ICPR | 4 |
| 2020 | A Versatile Crack Inspection Portable System based on Classifier Ensemble and Controlled IlluminationabstractThis paper presents a novel setup for automatic visual inspection of cracks in ceramic tile as well as studies the effect of various classifiers and height-varying illumination conditions for this task. The intuition behind this setup is that cracks can be better visualized under specific lighting conditions than others. Our setup, which is designed for field work with constraints in its maximum dimensions, can acquire images for crack detection with multiple lighting conditions using the illumination sources placed at multiple heights. Crack detection is then performed by classifying patches extracted from the acquired images in a sliding window fashion. We study the effect of lights placed at various heights by training classifiers both on customized as well as state-of-the-art architectures and evaluate their performance both at patch-level and image-level, demonstrating the effectiveness of our setup. More importantly, ours is the first study that demonstrates how height-varying illumination conditions can affect crack detection with the use of existing state-of-the-art classifiers. We provide an insight about the illumination conditions that can help in improving crack detection in a challenging real-world industrial environment. Milind Gajanan Padalkar, Carlos Beltrán 0002, Matteo Bustreo, Alessio Del Bue, Vittorio Murino |
ICPR | 5 |
| 2020 | Encoding Brain Networks Through Geodesic Clustering of Functional Connectivity for Multiple Sclerosis ClassificationabstractAn important task in brain connectivity research is the classification of patients from healthy subjects. In this work, we present a two-step mathematical framework allowing to discriminate between two groups of people with an application to multiple sclerosis. The proposed approach exploits the properties of the connectivity matrices determined using the covariances between signals of a fixed set of brain areas. These positive semidefinite matrices lay on a Riemannian manifold, allowing to use a geodesic distance defined on this space. In order to generate a vector representation useful for classification purposes, but still preserving the network structure, we encoded the data exploiting the network attractors determined by a geodesic clustering of connectivity matrices. Then clustering centroids were used as a dictionary allowing to encode subject's connectivity matrices as a vector of geodesic distances. A Linear Support Vector Machine was then used to perform classification between subjects. To demonstrate the advantage of using geodesic metrics in this framework, we conducted the same analysis using Euclidean metric. Experimental results validate the fact that employing geodesic metric in this framework leads to a higher classification performance, whereas performance with a Euclidean metric was sub-optimal. Muhammad Abubakar Yamin, Paola Valsasina, Michael Dayan, Sebastiano Vascon, Jacopo Tessadori, Massimo Filippi, Vittorio Murino, Maria Assunta Rocca, Diego Sona |
ICPR | 7 |
| 2020 | Geodesic Clustering of Positive Definite Matrices For Classification of Mental Disorder Using Brain Functional ConnectivityabstractFunctional Magnetic Resonance Imaging (fMRI) is a commonly used technique to evaluate brain activity, and can be used to distinguish patients from healthy controls in a variety of diseases. In this work, we present a two-step approach to discriminate healthy subjects against those affected by either Autism Spectrum Disorder or Schizophrenia on the basis of their connectivity patterns. We exploited the property that connectivity patterns described by positive definite matrices define a Riemannian manifold. In this framework, to generate a vector representation used in the classification task, we performed a geodesic clustering of the connectivity matrices. Cluster centroids were then used as a dictionary allowing to encode all subjects graphs as vectors of geodesic distances. A linear Support Vector Machine was then used to classify subjects. To show the advantage of using geodesic distances for this problem, the same analysis was conducted using a Euclidean metric. Experiments show that employing Euclidean distances leads to a lower classification performance and possibly to the definition of the wrong number of clusters, whereas geodesic clustering results in a significantly improved accuracy. Muhammad Abubakar Yamin, Jacopo Tessadori, Muhammad Usman Akbar, Michael Dayan, Vittorio Murino, Diego Sona |
IJCNN | 5 |
| 2020 | Generative Pseudo-label Refinement for Unsupervised Domain AdaptationabstractWe investigate and characterize the inherent resilience of conditional Generative Adversarial Networks (cGANs) against noise in their conditioning labels, and exploit this fact in the context of Unsupervised Domain Adaptation (UDA). In UDA, a classifier trained on the labelled source set can be used to infer pseudo-labels on the unlabelled target set. However, this will result in a significant amount of misclassified examples (due to the well-known domain shift issue), which can be interpreted as noise injection in the ground-truth labels for the target set. We show that cGANs are, to some extent, robust against such "shift noise". Indeed, cGANs trained with noisy pseudo-labels, are able to filter such noise and generate cleaner target samples. We exploit this finding in an iterative procedure where a generative model and a classifier are jointly trained: in turn, the generator allows to sample cleaner data from the target distribution, and the classifier allows to associate better labels to target samples, progressively refining target pseudo-labels. Results on common benchmarks show that our method performs better or comparably with the unsupervised domain adaptation state of the art. Pietro Morerio, Riccardo Volpi, Ruggero Ragonesi, Vittorio Murino |
WACV | 4 |
| 2020 | Audio-Visual Model Distillation Using Acoustic ImagesabstractIn this paper, we investigate how to learn rich and robust feature representations for audio classification from visual data and acoustic images, a novel audio data modality. Former models learn audio representations from raw signals or spectral data acquired by a single microphone, with remarkable results in classification and retrieval. However, such representations are not so robust towards variable environmental sound conditions. We tackle this drawback by exploiting a new multimodal labeled action recognition dataset acquired by a hybrid audio-visual sensor that provides RGB video, raw audio signals, and spatialized acoustic data, also known as acoustic images, where the visual and acoustic images are aligned in space and synchronized in time. Using this richer information, we train audio deep learning models in a teacher-student fashion. In particular, we distill knowledge into audio networks from both visual and acoustic image teachers. Our experiments suggest that the learned representations are more powerful and have better generalization capabilities than the features learned from models trained using just single-microphone audio data. Andrés F. Pérez, Valentina Sanguineti, Pietro Morerio, Vittorio Murino |
WACV | 4 |
| 2020 | Predicting Intentions from Motion: The Subject-Adversarial Adaptation ApproachabstractAbstract This paper aims at investigating the action prediction problem from a pure kinematic perspective. Specifically, we address the problem of recognizing future actions, indeed human intentions, underlying a same initial (and apparently unrelated) motor act. This study is inspired by neuroscientific findings asserting that motor acts at the very onset are embedding information about the intention with which are performed, even when different intentions originate from a same class of movements. To demonstrate this claim in computational and empirical terms, we designed an ad hoc experiment and built a new 3D and 2D dataset where, in both training and testing, we analyze a same class of grasping movements underlying different intentions. We investigate how much the intention discriminants generalize across subjects, discovering that each subject tends to affect the prediction by his/her own bias. Inspired by the domain adaptation problem, we propose to interpret each subject as a domain, leading to a novel subject adversarial paradigm. The proposed approach favorably copes with our new problem, boosting the considered baseline features encoding 2D and 3D information and which do not exploit the subject information. Andrea Zunino, Jacopo Cavazza, Riccardo Volpi, Pietro Morerio, Andrea Cavallo, Cristina Becchio, Vittorio Murino |
Int. J. Comput. Vis. | 7 |
| 2020 | Learning with Privileged Information via Adversarial Discriminative Modality DistillationabstractHeterogeneous data modalities can provide complementary cues for several tasks, usually leading to more robust algorithms and better performance. However, while training data can be accurately collected to include a variety of sensory modalities, it is often the case that not all of them are available in real life (testing) scenarios, where a model has to be deployed. This raises the challenge of how to extract information from multimodal data in the training stage, in a form that can be exploited at test time, considering limitations such as noisy or missing modalities. This paper presents a new approach in this direction for RGB-D vision tasks, developed within the adversarial learning and privileged information frameworks. We consider the practical case of learning representations from depth and RGB videos, while relying only on RGB data at test time. We propose a new approach to train a hallucination network that learns to distill depth information via adversarial learning, resulting in a clean approach without several losses to balance or hyperparameters. We report state-of-the-art results for object classification on the NYUD dataset, and video action recognition on the largest multimodal dataset available for this task, the NTU RGB+D, as well as on the Northwestern-UCLA. Nuno C. Garcia, Pietro Morerio, Vittorio Murino |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Aggregation Signature for Small Object TrackingabstractSmall object tracking becomes an increasingly important task, which however has been largely unexplored in computer vision. The great challenges stem from the facts that: 1) small objects show extreme vague and variable appearances, and 2) they tend to be lost easier as compared to normal-sized ones due to the shaking of lens. In this paper, we propose a novel aggregation signature suitable for small object tracking, especially aiming for the challenge of sudden and large drift. We make three-fold contributions in this work. First, technically, we propose a new descriptor, named aggregation signature, based on saliency, able to represent highly distinctive features for small objects. Second, theoretically, we prove that the proposed signature matches the foreground object more accurately with a high probability. Third, experimentally, the aggregation signature achieves a high performance on multiple datasets, outperforming the state-of-the-art methods by large margins. Moreover, we contribute with two newly collected benchmark datasets, i.e., small90 and small112, for visually small object tracking. The datasets will be available in https://github.com/bczhangbczhang/. Chunlei Liu 0001, Wenrui Ding, Vittorio Murino, Baochang Zhang 0001, Jungong Han, Guodong Guo |
IEEE Trans. Image Process. | 4 |
| 2019 | Guided Zoom: Questioning Network Evidence for Fine-grained Classification
Sarah Adel Bargal, Andrea Zunino, Vitali Petsiuk, Jianming Zhang 0001, Kate Saenko, Vittorio Murino, Stan Sclaroff |
BMVC | 6 |
| 2019 | Addressing Model Vulnerability to Distributional Shifts Over Image Transformation SetsabstractWe are concerned with the vulnerability of computer vision models to distributional shifts. We formulate a combinatorial optimization problem that allows evaluating the regions in the image space where a given model is more vulnerable, in terms of image transformations applied to the input, and face it with standard search algorithms. We further embed this idea in a training procedure, where we define new data augmentation rules according to the image transformations that the current model is most vulnerable to, over iterations. An empirical evaluation on classification and semantic segmentation problems suggests that the devised algorithm allows to train models that are more robust against content-preserving image manipulations and, in general, against distributional shifts. Riccardo Volpi, Vittorio Murino |
ICCV | 2 |
| 2019 | Unsupervised Domain-Adaptive Person Re-Identification Based on AttributesabstractPedestrian attributes, e. g., hair length, clothes type and color, locally describe the semantic appearance of a person. Training person re-identification (ReID) algorithms under the supervision of such attributes have proven to be effective in extracting local features which are important for ReID. Unlike person identity, attributes are consistent across different domains (or datasets). However, most of ReID datasets lack attribute annotations. On the other hand, there are several datasets labeled with sufficient attributes for the case of pedestrian attribute recognition. Exploiting such data for ReID purpose can be a way to alleviate the shortage of attribute annotations in ReID case. In this work, an unsupervised domain adaptive ReID feature learning framework is proposed to make full use of attribute annotations. We propose to transfer attribute-related features from their original domain to the ReID one: to this end, we introduce an adversarial discriminative domain adaptation method in order to learn domain invariant features for encoding semantic attributes. Experiments on three large-scale datasets validate the effectiveness of the proposed ReID framework. Xiangping Zhu, Pietro Morerio, Vittorio Murino |
ICIP | 3 |
| 2019 | Scalable and compact 3D action recognition with approximated RBF kernel machines
Jacopo Cavazza, Pietro Morerio, Vittorio Murino |
Pattern Recognit. | 3 |
| 2019 | Adaptation of person re-identification models for on-boarding new camera(s)
Rameswar Panda, Amran Bhuiyan, Vittorio Murino, Amit K. Roy-Chowdhury |
Pattern Recognit. | 3 |
| 2019 | A Sequential Data Analysis Approach to Detect Emergent Leaders in Small GroupsabstractThis paper addresses the problem of predicting emergent leaders (ELs) in small groups, that is, meetings. This is a long-lasting research problem for social and organizational psychology and a relevant problem that recently gained momentum in social computing. Toward this goal, we propose a novel method, which analyzes the temporal dependencies of the audio-visual data by applying unsupervised deep learning generative models (feature learning). To the best of our knowledge, this is the first attempt that sequential data processing is performed for EL detection. Feature learning results in a single feature vector per a given time interval and all feature vectors representing a participant are aggregated using novel fusion techniques. Finally, the EL detection is performed using the state-of-the-art single and multiple kernel learning algorithms. The proposed method shows (significantly) improved results compared to the state-of-the-art methods and it can be adapted to analyze various small group interactions given that it is a general approach. Cigdem Beyan, Vasiliki-Maria Katsageorgiou, Vittorio Murino |
IEEE Trans. Multim. | 3 |
| 2018 | Dropout as a Low-Rank Regularizer for Matrix FactorizationabstractRegularization for matrix factorization (MF) and approximation problems has been carried out in many different ways. Due to its popularity in deep learning, dropout has been applied also for this class of problems. Despite its solid empirical performance, the theoretical properties of dropout as a regularizer remain quite elusive for this class of problems. In this paper, we present a theoretical analysis of dropout for MF, where Bernoulli random variables are used to drop columns of the factors. We demonstrate the equivalence between dropout and a fully deterministic model for MF in which the factors are regularized by the sum of the product of squared Euclidean norms of the columns. Additionally, we inspect the case of a variable sized factorization and we prove that dropout achieves the global minimum of a convex approximation problem with (squared) nuclear norm regularization. As a result, we conclude that dropout can be used as a low-rank regularizer with data dependent singular-value thresholding. Jacopo Cavazza, Pietro Morerio, Benjamin D. Haeffele, Connor Lane, Vittorio Murino, René Vidal |
AISTATS | 5 |
| 2018 | Visually-Driven Semantic Augmentation for Zero-Shot Learning
Abhinaba Roy, Jacopo Cavazza, Vittorio Murino |
BMVC | 3 |
| 2018 | Excitation Backprop for RNNs
Sarah Adel Bargal, Andrea Zunino, Donghyun Kim 0006, Jianming Zhang 0001, Vittorio Murino, Stan Sclaroff |
CVPR | 5 |
| 2018 | Adversarial Feature Augmentation for Unsupervised Domain AdaptationabstractRecent works showed that Generative Adversarial Networks (GANs) can be successfully applied in unsupervised domain adaptation, where, given a labeled source dataset and an unlabeled target dataset, the goal is to train powerful classifiers for the target samples. In particular, it was shown that a GAN objective function can be used to learn target features indistinguishable from the source ones. In this work, we extend this framework by (i) forcing the learned feature extractor to be domain-invariant, and (ii) training it through data augmentation in the feature space, namely performing feature augmentation. While data augmentation in the image space is a well established technique in deep learning, feature augmentation has not yet received the same level of attention. We accomplish it by means of a feature generator trained by playing the GAN minimax game against source features. Results show that both enforcing domain-invariance and performing feature augmentation lead to superior or comparable performance to state-of-the-art results in several unsupervised domain adaptation benchmarks. Riccardo Volpi, Pietro Morerio, Silvio Savarese, Vittorio Murino |
CVPR | 4 |
| 2018 | Modality Distillation with Multiple Stream Networks for Action Recognition
Nuno C. Garcia, Pietro Morerio, Vittorio Murino |
ECCV (8) | 3 |
| 2018 | A Multi-View Learning Approach to Deception DetectionabstractRecently, automatic deception detection has gained momentum thanks to advances in computer vision, computational linguistics and machine learning research fields. The majority of the work in this area focused on written deception and analysis of verbal features. However, according to psychology, people display various nonverbal behavioral cues, in addition to verbal ones, while lying. Therefore, it is important to utilize additional modalities such as video and audio to detect deception accurately. When multi-modal data was used for deception detection, previous studies concatenated all verbal and nonverbal features into a single vector. This concatenation might not be meaningful, because different feature groups can have different statistical properties, leading to lower classification accuracy. Following this intuition, we apply, for the first time in deception detection, a multi-view learning (MVL) approach, where each view corresponds to a feature group. This results in improved classification results over the state of the art methods. Additionally, we show that the optimized parameters of the MVL algorithm can give insights into the contribution of each feature group to the final results, thus revealing the importance of each feature and eliminating the need of performing feature selection as well. Finally, we focus on analyzing face-based low level, not hand crafted features, which are extracted using various pre-trained Deep Neural Networks (DNNs), showing that face is the most important nonverbal cue for the detection of deception. Nicolò Carissimi, Cigdem Beyan, Vittorio Murino |
FG | 3 |
| 2018 | Unsupervised Detection of White Matter Fiber Bundles with Stochastic Neural NetworksabstractExploring the human structural connectome often involves dealing with millions of white matter tracts reconstructed from diffusion MRI. Reducing the dimensionality of such data by grouping tracts into bundles can prove essential for subsequent analyses. Many unsupervised clustering algorithms aim at providing such bundles but often require the choice of a distance metric and suffer from memory storage issues relating to the size of the datasets. We propose for the first time a neural network approach for the unsupervised clustering of white matter tracts. It has the main properties of learning automatically the tract features and scaling well with the data size. As a proof of concept, we compare both quantitatively and qualitatively the computed tract clusters with a commonly used clustering method. The proposed approach shows results similar to the reference approach while not using any distance matrix or similarity metric. Michael Dayan, Vasiliki-Maria Katsageorgiou, Luca Dodero, Vittorio Murino, Diego Sona |
ICIP | 4 |
| 2018 | Minimal-Entropy Correlation Alignment for Unsupervised Deep Domain Adaptation
Pietro Morerio, Jacopo Cavazza, Vittorio Murino |
ICLR (Poster) | 3 |
| 2018 | Discriminative Latent Visual Space For Zero-Shot Object ClassificationabstractIn this paper We deal with the problem of zero-shot visual recognition. The standard zero-shot learning (ZSL) pipeline is based on the idea of learning a functional mapping from a visual embedding space to an auxiliary semantic space for a set of seen categories. In the testing phase, the task is to recognize a set of novel categories which are semantically linked to the already known ones. Although such a pipeline is inherently supervised, there exists very few endeavours in the context of ZSL that enforce discrimination in learning this mapping. In this work, we propose a novel encoder-decoder network to explore the possibility of learning an intermediate latent space for the visual features, which is deemed to be simultaneously reconstructive and discriminative. By reaching a trade-off between the joint (re)construction of the visual and the semantic embedding spaces, while ensuring separability among the known classes, the proposed model better generalizes to the unknown categories. Experimental results obtained on challenging datasets, such as AwA, CUB, and ImageNet-2, establish the efficacy of such a discriminative latent space for the standard ZSL setup. Abhinaba Roy, Biplab Banerjee, Vittorio Murino |
ICPR | 3 |
| 2018 | Multiple Mice Tracking: Occlusions Disentanglement using a Gaussian Mixture ModelabstractMouse models play an important role in preclinical research and drug discovery for human diseases. The fact that mice are a social species partaking in social interactions of high degree facilitates the study of diseases characterized by social alterations. Hence, robust animal tracking is of great importance in order to build tools capable of automatically analyzing social behavioral interactions of multiple mice. However, the presence of occlusions is a major problem in multiple mice tracking. To deal with this problem, we present here a tracking algorithm based on Kalman filter and Gaussian Mixture Modeling. Specifically, Kalman tracking is used to track the mice and when occlusions happen, we fit 2D Gaussian distributions to separate mouse blobs. This helps us to prevent mice identity swaps as it is an important feature for accurate behavior analysis. As the results of our experiments show, the proposed algorithm results in much fewer identity swaps than other state of the art algorithms. Ario Sadafi, Vasiliki-Maria Katsageorgiou, Francesco Papaleo, Vittorio Murino, Diego Sona |
ICPR | 5 |
| 2018 | Video Gesture Analysis for Autism Spectrum Disorder DetectionabstractAutism is a behavioral neurological disorder affecting a significant percentage of worldwide population. It especially starts manifesting at very low ages, but it is difficult to early diagnose it since there is not a specific exam or trial that is able to spot it safely. Its detection is in fact mainly dependent from the medical expertise used to assess the patient behavior during direct interviews. This work aims at providing an automatic objective support to the doctor for the assessment of (early) diagnosis of possible autistic subjects by only using video sequences. The underlying idea and rationale come from the psychological and neuroscience studies claiming that the executions of simple motor acts are different between pathological and healthy subjects, and this can be sufficient to discriminate between them. To this end, we devised an experiment in which we recorded, using a standard video camera, patient and healthy children performing the same simple gesture of grasping a bottle. By only processing the video clips depicting the grasping action using a recurrent deep neural network, we are able to discriminate with a good accuracy between the 2 classes of subjects. The designed deep model is also able to provide a sort of attention map in which the zones in the video of major interest are identified in space and time: this “explains” in a certain way which areas the model deems more relevant to the classification purpose, which could also be used by the doctor to make the diagnosis. In the end, this work constitutes a first step towards the development of an automatic computational system devoted to the early diagnosis of autistic subjects, providing the medical expert of a supportive objective method, potentially simple to use in clinical and also more open settings. Andrea Zunino, Pietro Morerio, Andrea Cavallo, Caterina Ansuini, Jessica Podda, Francesca Battaglia, Edvige Veneselli, Cristina Becchio, Vittorio Murino |
ICPR | 9 |
| 2018 | Investigation of Small Group Social Interactions Using Deep Visual Activity-Based Nonverbal FeaturesabstractUnderstanding small group face-to-face interactions is a prominent research problem for social psychology while the automatic realization of it recently became popular in social computing. This is mainly investigated in terms of nonverbal behaviors, as they are one of the main facet of communication. Among several multi-modal nonverbal cues, visual activity is an important one and its sufficiently good performance can be crucial for instance, when the audio sensors are missing. The existing visual activity-based nonverbal features, which are all hand-crafted, were able to perform well enough for some applications while did not perform well for some other problems. Given these observations, we claim that there is a need of more robust feature representations, which can be learned from data itself. To realize this, we propose a novel method, which is composed of optical flow computation, deep neural network based feature learning, feature encoding and classification. Additionally, a comprehensive analysis between different feature encoding techniques is also presented. The proposed method is tested on three research topics, which can be perceived during small group interactions i.e. meetings: i) emergent leader detection, ii) emergent leadership style prediction, and iii) high/low extraversion classification. The proposed method shows (significantly) better results not only as compared to the state of the art visual activity based-nonverbal features but also when the state of the art visual activity based-nonverbal features are combined with other audio-based and video-based nonverbal features. Cigdem Beyan, Muhammad Shahid 0002, Vittorio Murino |
ACM Multimedia | 3 |
| 2018 | Generalizing to Unseen Domains via Adversarial Data AugmentationabstractWe are concerned with learning models that generalize well to different unseen domains. We consider a worst-case formulation over data distributions that are near the source domain in the feature space. Only using training data from a single source distribution, we propose an iterative procedure that augments the dataset with examples from a fictitious target domain that is "hard" under the current model. We show that our iterative scheme is an adaptive data augmentation method where we append adversarial examples at each iteration. For softmax losses, we show that our method is a data-dependent regularization scheme that behaves differently from classical regularizers that regularize towards zero (e.g., ridge or lasso). On digit recognition and semantic segmentation tasks, our method learns models improve performance across a range of a priori unknown target domains. Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John C. Duchi, Vittorio Murino, Silvio Savarese |
NeurIPS | 5 |
| 2018 | Person re-identification by order-induced metric fusion
Behzad Mirmahboub, Mohamed Lamine Mekhalfi, Vittorio Murino |
Neurocomputing | 3 |
| 2018 | Discriminative body part interaction mining for mid-level action representation and classification
Abhinaba Roy, Biplab Banerjee, Vittorio Murino |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Manifold constraint transfer for visual structure-driven optimization
Baochang Zhang 0001, Alessandro Perina, Ce Li 0002, Qixiang Ye, Vittorio Murino, Alessio Del Bue |
Pattern Recognit. | 5 |
| 2018 | Audio Tracking in Noisy Environments by Acoustic Map and Spectral SignatureabstractA novel method is proposed for generic target tracking by audio measurements from a microphone array. To cope with noisy environments characterized by persistent and high energy interfering sources, a classification map (CM) based on spectral signatures is calculated by means of a machine learning algorithm. Next, the CM is combined with the acoustic map, describing the spatial distribution of sound energy, in order to obtain a cleaned joint map in which contributions from the disturbing sources are removed. A likelihood function is derived from this map and fed to a particle filter yielding the target location estimation on the acoustic image. The method is tested on two real environments, addressing both speaker and vehicle tracking. The comparison with a couple of trackers, relying on the acoustic map only, shows a sharp improvement in performance, paving the way to the application of audio tracking in real challenging environments. Marco Crocco, Samuele Martelli, Andrea Trucco, Andrea Zunino, Vittorio Murino |
IEEE Trans. Cybern. | 5 |
| 2018 | Deep Endoscope: Intelligent Duct Inspection for the Avionic IndustryabstractWe present the first autonomous endoscope for the visual inspection of very small ducts and cavities, up to a 6-mm diameter. The system has been designed, implemented, and tested in a challenging industrial scenario and in strict collaboration with an avionic industry partner. The inspected objects are metallic gearboxes eventually presenting different residuals (e.g., sand, machining swarfs, and metallic dust) inside the oil ducts. The automatic system is actuated by a robotic arm that moves the endoscope with a microcamera inside the gearbox duct, while a deep-learning-based spatio-temporal image analysis module detects, classifies, and localizes defects in real time. Feedback is given to the robotic arm in order to move or extract the endoscope given the detected anomalies. Evaluation provides a detection rate of nearly 98% given different tests with different types of residuals and duct structures. Samuele Martelli, Luca Mazzei, Carlo Canali, Paolo Guardiani, Salvatore Giunta, Alberto Ghiazza, Ivan Mondino, Ferdinando Cannella, Vittorio Murino, Alessio Del Bue |
IEEE Trans. Ind. Informatics | 9 |
| 2018 | Prediction of the Leadership Style of an Emergent Leader Using Audio and Visual Nonverbal FeaturesabstractThe coordination of a leader with group members is very important for an effective leadership given that this figure is the person who actually manages the team members to achieve a desired goal. Investigating the leadership and especially the leadership style is a prominent research topic in social and organizational psychology. However this is a new problem in social signal processing that can actually make valuable contributions by analyzing multimodal data in a more effective and efficient way. In this work we identify the leadership style of an emergent leader (i.e. the leader who naturally arises from a group not designated) as autocratic or democratic. The proposed method is applied to a dataset in-the-wild; in other words there is no role-playing which is novel for this problem. Multiple kernel learning (MKL) using multimodal nonverbal features is utilized to predict leadership styles that proved to achieve better predictions as compared to traditional learning methods. Thanks to MKL and a simple heuristic proposed the best performing features are also identified showing that better predictions can be reached only by using those features. Additionally correlation analysis between the extracted nonverbal features and the results of social psychology questionnaire is also performed. This shows that significantly high correlations exist for speaking activity based and prosodic nonverbal features. Cigdem Beyan, Francesca Capozzi, Cristina Becchio, Vittorio Murino |
IEEE Trans. Multim. | 4 |
| 2017 | Exploiting Gaussian mixture importance for person re-identificationabstractPerson re-identification (ReID) stands for the task of determining the co-occurrence of individuals across a network of cameras with disjoint viewfields. The relevant literature documents a plausible number of contributions so far. KISS metric learning is an effective ReID method. However, as reported in the existing works, KISS metric learning is sensitive to the feature dimensionality and can not capture the multi modes in the dataset. To this end, we propose in this paper a Gaussian Mixture Importance Estimation (GMIE) approach for ReID, which exploits the Gaussian Mixture Models (GMMs) to estimate the observed commonalities of similar and dissimilar person pairs in the feature space. Experiments on three benchmark datasets reveal that our method offers the property of maintaining its efficiency on high-dimensional features. Moreover, the proposed GMIE scores plausible ReID rates as compared to other works. Xiangping Zhu, Amran Bhuiyan, Mohamed Lamine Mekhalfi, Vittorio Murino |
AVSS | 4 |
| 2017 | Unsupervised Adaptive Re-identification in Open World Dynamic Camera NetworksabstractPerson re-identification is an open and challenging problem in computer vision. Existing approaches have concentrated on either designing the best feature representation or learning optimal matching metrics in a static setting where the number of cameras are fixed in a network. Most approaches have neglected the dynamic and open world nature of the re-identification problem, where a new camera may be temporarily inserted into an existing system to get additional information. To address such a novel and very practical problem, we propose an unsupervised adaptation scheme for re-identification models in a dynamic camera network. First, we formulate a domain perceptive re-identification method based on geodesic flow kernel that can effectively find the best source camera (already installed) to adapt with a newly introduced target camera, without requiring a very expensive training phase. Second, we introduce a transitive inference algorithm for re-identification that can exploit the information from best source camera to improve the accuracy across other camera pairs in a network of multiple cameras. Extensive experiments on four benchmark datasets demonstrate that the proposed approach significantly outperforms the state-of-the-art unsupervised learning based alternatives whilst being extremely efficient to compute. Rameswar Panda, Amran Bhuiyan, Vittorio Murino, Amit K. Roy-Chowdhury |
CVPR | 3 |
| 2017 | Efficient pooling of image based CNN features for action recognition in videosabstractIn this paper, we propose a new video representation incorporating image based deep features and an efficient pooling strategy for the purpose of action recognition. The Convolutional Neural Network (CNN) based features have very recently emerged as the new state of the art for image classification. Several attempts have been made to extend such CNN models for videos by explicitly focusing on the temporal evolution of the frames. Feature pooling is one such approach which represents video sequences in terms of some statistical properties of the feature dimensions over the frames. However, traditional pooling strategies including max or average pooling explicitly fail to capture the temporal progression of the frame-level contents. In contrast to previous pooling techniques, we propose a two-level video representation which separately focuses on the entire video as well as a number of video sub-volumes. In both levels, we introduce a generic time series pooling on the frame-level deep CNN features efficiently. Further, a self-tuning spectral clustering is considered to highlight video snippets which are highly probable to contain significant sub-action sequences. We validate the proposed feature encoding on the challenging KTH-actions and UCF-50 datasets and find that the proposed encoding outperforms traditional pooling based feature representations by substantial margin. Biplab Banerjee, Vittorio Murino |
ICASSP | 2 |
| 2017 | Curriculum DropoutabstractDropout is a very effective way of regularizing neural networks. Stochastically “dropping out” units with a certain probability discourages over-specific co-adaptations of feature detectors, preventing overfitting and improving network generalization. Besides, Dropout can be interpreted as an approximate model aggregation technique, where an exponential number of smaller networks are averaged in order to get a more powerful ensemble. In this paper, we show that using a fixed dropout probability during training is a suboptimal choice. We thus propose a time scheduling for the probability of retaining neurons in the network. This induces an adaptive regularization scheme that smoothly increases the difficulty of the optimization problem. This idea of “starting easy” and adaptively increasing the difficulty of the learning problem has its roots in curriculum learning and allows one to train better models. Indeed, we prove that our optimization strategy implements a very general curriculum scheme, by gradually adding noise to both the input and intermediate feature representations within the network architecture. Experiments on seven image classification datasets and different network architectures show that our method, named Curriculum Dropout, frequently yields to better generalization and, at worst, performs just as well as the standard Dropout method. Pietro Morerio, Jacopo Cavazza, Riccardo Volpi, René Vidal, Vittorio Murino |
ICCV | 5 |
| 2017 | Summarization and Classification of Wearable Camera Streams by Learning the Distributions over Deep Features of Out-of-Sample Image SequencesabstractA popular approach to training classifiers of new image classes is to use lower levels of a pre-trained feed-forward neural network and retrain only the top. Thus, most layers simply serve as highly nonlinear feature extractors. While these features were found useful for classifying a variety of scenes and objects, previous work also demonstrated unusual levels of sensitivity to the input especially for images which are veering too far away from the training distribution. This can lead to surprising results as an imperceptible change in an image can be enough to completely change the predicted class. This occurs in particular in applications involving personal data, typically acquired with wearable cameras (e.g., visual lifelogs), where the problem is also made more complex by the dearth of new labeled training data that make supervised learning with deep models difficult. To alleviate these problems, in this paper we propose a new generative model that captures the feature distribution in new data. Its latent space then becomes more representative of the new data, while still retaining the generalization properties. In particular, we use constrained Markov walks over a counting grid for modeling image sequences, which not only yield good latent representations, but allow for excellent classification with only a handful of labeled training examples of the new scenes or objects, a scenario typical in lifelogging applications. Alessandro Penna, Sadegh Mohammadi 0001, Nebojsa Jojic, Vittorio Murino |
ICCV | 4 |
| 2017 | Multi-task learning of social psychology assessments and nonverbal features for automatic leadership identificationabstractIn social psychology, the leadership investigation is performed using questionnaires which are either i) self-administered or ii) applied to group participants to evaluate other members or iii) filled by external observers. While each of these sources is informative, using them individually might not be as effective as using them jointly. This paper is the first attempt which addresses the automatic identification of leaders in small-group meetings, by learning effective models using nonverbal audio-visual features and the results of social psychology questionnaires that reflect assessments regarding leadership. Learning is based on Multi-Task Learning which is performed without using ground-truth data (GT), but using the results of questionnaires (having substantial agreement with GT), administered to external observers and the participants of the meetings, as tasks. The results show that joint learning results in better performance as compared to single task learning and other baselines. Cigdem Beyan, Francesca Capozzi, Cristina Becchio, Vittorio Murino |
ICMI | 4 |
| 2017 | Re-identification: State of the Art and Current Trends
Vittorio Murino |
ICPRAM | 1 |
| 2017 | A Novel Dictionary Learning based Multiple Instance Learning Approach to Action Recognition from VideosabstractIn this paper we deal with the problem of action recognition from unconstrained videos under the notion of multiple instance learning (MIL).The traditional MIL paradigm considers the data items as bags of instances with the constraint that the positive bags contain some class-specific instances whereas the negative bags consist of instances only from negative classes.A classifier is then further constructed using the bag level annotations and a distance metric between the bags.However, such an approach is not robust to outliers and is time consuming for a moderately large dataset.In contrast, we propose a dictionary learning based strategy to MIL which first identifies class-specific discriminative codewords, and then projects the bag-level instances into a probabilistic embedding space with respect to the selected codewords.This essentially generates a fixedlength vector representation of the bags which is specifically dominated by the properties of the class-specific instances.We introduce a novel exhaustive search strategy using a support vector machine classifier in order to highlight the class-specific codewords.The standard multiclass classification pipeline is followed henceforth in the new embedded feature space for the sake of action recognition.We validate the proposed framework on the challenging KTH and Weizmann datasets, and the results obtained are promising and comparable to representative techniques from the literature. Abhinaba Roy, Biplab Banerjee, Vittorio Murino |
ICPRAM | 3 |
| 2017 | Data-driven study of mouse sleep-stages using Restricted Boltzmann MachinesabstractRodents play an important role in sleep studies since they are the most easily available low-cost animal species sharing most genes and gene functions with humans. The scoring of sleep stages in these studies is usually based on a manual analysis of long-lasting recordings of the brain electrical activity and the activity of skeletal muscles. Hence, there is a great need for tools automating the investigation over large experiments. In this paper we present a way of analyzing this huge amount of electrophysiological data using unsupervised learning models. We show how employing latent variable models, like Restricted Boltzmann Machines, we can define different sleep-wakefulness sub-stages reflected by regularities in the data, without any prior knowledge. Our analysis shows that we can effectively discover meaningful feature representations which characterize sleep stages at a finer level than those commonly used by the experts. These feature representations can be also used to further characterize different mouse genotypes, without incurring biases of experts like in a classical analysis setup where predefined rules limit the discovery of novel insights. Vasiliki-Maria Katsageorgiou, Matteo Zanotto, Valter Tucci, Vittorio Murino, Diego Sona |
IJCNN | 4 |
| 2017 | Moving as a Leader: Detecting Emergent Leadership in Small Groups using Body PoseabstractDetecting leadership while understanding the underlying behavior is an important research topic particularly for social and organizational psychology, and has started to get attention from social signal processing research community as well. It is known that, visual activity is a useful cue to investigate the social interactions, even though previously applied nonverbal features based on head/body actions were not performing well enough for identification of emergent leaders (ELs) in small group meetings. Starting from these premises, in this study, we propose an effective method that uses 2D body pose based nonverbal features to represent the visual activity of a person. Our results suggest that, i) overall, the proposed nonverbal features derived from body pose perform better than existing visual activity based features, ii) it is possible to improve classification results by applying unsupervised feature learning as a preprocessing step, and iii) the proposed nonverbal features are able to advance the EL identification performances of other types of nonverbal features when they are used together. Cigdem Beyan, Vasiliki-Maria Katsageorgiou, Vittorio Murino |
ACM Multimedia | 3 |
| 2017 | Predicting Human Intentions from Motion Cues Only: A 2D+3D Fusion ApproachabstractIn this paper, we address the new problem of the prediction of human intentions. There is neuro-psychological evidence that actions performed by humans are anticipated by peculiar motor acts which are discriminant of the type of action going to be performed afterwards. In other words, an actual intention can be forecast by looking at the kinematics of the immediately preceding movement. To prove it in a computational and quantitative manner, we devise a new experimental setup where, without using contextual information, we predict human intentions all originating from the same motor act. We posit the problem as a classification task and we introduce a new multi-modal dataset consisting of a set of motion capture marker 3D data and 2D video sequences, where, by only analysing very similar movements in both training and test phases, we are able to predict the underlying intention, i.e., the future, never observed action. We also present an extensive experimental evaluation as a baseline, customizing state-of-the-art techniques for either 3D and 2D data analysis. Realizing that video processing methods lead to inferior performance but show complementary information with respect to 3D data sequences, we developed a 2D+3D fusion analysis where we achieve better classification accuracies, attesting the superiority of the multimodal approach for the context-free prediction of human intentions. Andrea Zunino, Jacopo Cavazza, Atesh Koul, Andrea Cavallo, Cristina Becchio, Vittorio Murino |
ACM Multimedia | 6 |
| 2017 | Distance Penalization and Fusion for Person Re-identificationabstractThis paper presents a novel person re-identification framework based on data fusion. The pipeline of the proposed method is composed of two stages. First, a metric learning paradigm is applied on a bunch of distinct feature extractors to produce an ensemble of estimated distance measures, which are subsequently penalized according to their confidence in evidencing the correct matches from the false ones, and averaged as to draw a final decision. Second, the close persons from the gallery are selected based on the previously fused distance estimates, and utilized to build a dictionary as to reconstruct a given probe pattern. Evaluated on benchmark datasets, the proposed framework advances the state-of-the-art by interesting margins. In particular, Rank1 gains amounting to about 12%, 1%, 6%, and 12%, were scored on VIPeR, CAVIAR4REID, iLIDS, and 3DPeS, respectively. Behzad Mirmahboub, Mohamed Lamine Mekhalfi, Vittorio Murino |
WACV | 3 |
| 2017 | Image and Video Understanding in Big DataabstractAn active object recognition system has the advantage of acting in the environment to capture images that are more suited for training and lead to better performance at test time. In this paper, we utilize deep convolutional neural networks for active object recognition by simultaneously predicting the object label and the next action to be performed on the object with the aim of improving recognition performance. We treat active object recognition as a reinforcement learning problem and derive the cost function to train the network for joint prediction of the object label and the action. A generative model of object similarities based on the Dirichlet distribution is proposed and embedded in the network for encoding the state of the system. The training is carried out by simultaneously minimizing the label and action prediction errors using gradient descent. We empirically show that the proposed network is able to predict both the object label and the actions on GERMS, a dataset for active object recognition. We compare the test label prediction accuracy of the proposed model with Dirichlet and Naive Bayes state encoding. The results of experiments suggest that the proposed model equipped with Dirichlet state encoding is superior in performance, and selects images that lead to better training and higher accuracy of label prediction at test time. Vittorio Murino, Shaogang Gong, Chen Change Loy, Loris Bazzani |
Comput. Vis. Image Underst. | 1 |
| 2017 | Guest Editorial: Language in Vision
Yan Yan 0002, Jiwen Lu, Ajmal Mian, Arun Ross, Vittorio Murino, Radu Horaud |
Comput. Vis. Image Underst. | 5 |
| 2017 | Automatic inspection of aeronautic components
Marco San-Biagio, Carlos Beltrán 0002, Salvatore Giunta, Alessio Del Bue, Vittorio Murino |
Mach. Vis. Appl. | 5 |
| 2017 | Adaptive Local Movement Modeling for Robust Object TrackingabstractIn this paper, we present a new strategy for modeling the motion of local patches for single-object tracking that can be seamlessly applied to most part-based trackers in the literature. The proposed adaptive local movement modeling method is able to model the local movement distribution of the image patches defining the object to track and the reliability of each image patch. Given the output of a base tracking algorithm, a Gaussian mixture model (GMM) is first used to model the distribution of the movement of local patches relative to the center of gravity of the tracked object. Then, the GMM is combined with the chosen base tracker in a boosting framework, which gives an efficient integrated scheme for the tracking task. This provides a robust procedure to detect outliers in the local motion of the patches. The algorithm is highly configurable with the possibility to change the number of local patches used for tracking and to adapt to the variations of the tracked object. The extensive tracking results on standard data sets show that equipping state-of-the-art trackers with our technique remarkably improves their performance. Baochang Zhang 0001, Alessandro Perina, Alessio Del Bue, Vittorio Murino, Jianzhuang Liu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | Exploiting Feature Correlations by Brownian Statistics for People Detection and RecognitionabstractCharacterizing an image region by its feature intercorrelations is a modern trend in computer vision. In this paper, we introduce a new image descriptor that can be seen as a natural extension of a standard covariance descriptor with the advantage of capturing nonlinear and nonmonotone dependencies. Inspired from the recent advances in mathematical statistics of Brownian motion, we can express highly complex structural information in a compact and computationally efficient manner. We show that our Brownian covariance descriptor can capture richer image characteristics than the covariance descriptor. Additionally, a detailed analysis of the Brownian manifold reveals that opposite to the classical covariance descriptor, the proposed descriptor lies in a relatively flat manifold, which can be treated as a Euclidean. This brings significant boost in the efficiency of the descriptor. The effectiveness and the generality of our approach is validated on two challenging vision tasks, pedestrian classification, and person reidentification. The experiments are carried out on multiple datasets achieving promising results. Slawomir Bak, Marco San-Biagio, Ratnesh Kumar 0003, Vittorio Murino, François Brémond |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2016 | Novel dataset for fine-grained abnormal behavior understanding in crowdabstractDespite the huge research on crowd on behavior understanding in visual surveillance community, lack of publicly available realistic datasets for evaluating crowd behavioral interaction led not to have a fair common test bed for researchers to compare the strength of their methods in the real scenarios. This work presents a novel crowd dataset contains around 45,000 video clips which annotated by one of the five different fine-grained abnormal behavior categories. We also evaluated two state-of-the-art methods on our dataset, showing that our dataset can be effectively used as a benchmark for fine-grained abnormality detection. The details of the dataset and the results of the baseline methods are presented in the paper. Javad Haddadnia, Hossein Mousavi, Maziyar Kalantarzadeh, Moin Nabi, Vittorio Murino |
AVSS | 6 |
| 2016 | Approximate Log-Hilbert-Schmidt Distances between Covariance Operators for Image ClassificationabstractThis paper presents a novel framework for visual object recognition using infinite-dimensional covariance operators of input features, in the paradigm of kernel methods on infinite-dimensional Riemannian manifolds. Our formulation provides a rich representation of image features by exploiting their non-linear correlations, using the power of kernel methods and Riemannian geometry. Theoretically, we provide an approximate formulation for the Log-Hilbert-Schmidt distance between covariance operators that is efficient to compute and scalable to large datasets. Empirically, we apply our framework to the task of image classification on eight different, challenging datasets. In almost all cases, the results obtained outperform other state of the art methods, demonstrating the competitiveness and potential of our framework. Hà Quang Minh, Marco San-Biagio, Loris Bazzani, Vittorio Murino |
CVPR | 4 |
| 2016 | Angry Crowds: Detecting Violent Events in Videos
Sadegh Mohammadi 0001, Alessandro Perina, Hamed Kiani Galoogahi, Vittorio Murino |
ECCV (7) | 4 |
| 2016 | Person re-identification using sparse representation with manifold constraintsabstractHuman re-identification is still a challenging task due to the human pose and illumination variations. Nowadays, surveillance cameras with high frame rate are capable of capturing several consecutive frames from each person. Multi-shot images provide richer information of the target person compared to a single-shot image. They, however, produce a high cost of information redundancy which may degrade the performance of re-identification systems. In this paper, we propose a novel framework that combines sparse coding and manifold constraints to extract discriminative information from multi-shot images of one pedestrian for person re-identification across a set of non-overlapped surveillance cameras. The evaluation over two standard multi-shot datasets shows very competitive accuracy of our framework against the state-of-the-art. Behzad Mirmahboub, Hamed Kiani Galoogahi, Amran Bhuiyan, Alessandro Perina, Baochang Zhang 0001, Alessio Del Bue, Vittorio Murino |
ICIP | 7 |
| 2016 | Detecting emergent leader in a meeting environment using nonverbal visual features onlyabstractIn this paper, we propose an effective method for emergent leader detection in meeting environments which is based on nonverbal visual features. Identifying emergent leader is an important issue for organizations. It is also a well-investigated topic in social psychology while a relatively new problem in social signal processing (SSP). The effectiveness of nonverbal features have been shown by many previous SSP studies. In general, the nonverbal video-based features were not more effective compared to audio-based features although, their fusion generally improved the overall performance. However, in absence of audio sensors, the accurate detection of social interactions is still crucial. Motivating from that, we propose novel, automatically extracted, nonverbal features to identify the emergent leadership. The extracted nonverbal features were based on automatically estimated visual focus of attention which is based on head pose. The evaluation of the proposed method and the defined features were realized using a new dataset which is firstly introduced in this paper including its design, collection and annotation. The effectiveness of the features and the method were also compared with many state of the art features and methods. Cigdem Beyan, Nicolò Carissimi, Francesca Capozzi, Sebastiano Vascon, Matteo Bustreo, Antonio Pierro, Cristina Becchio, Vittorio Murino |
ICMI | 8 |
| 2016 | Kernelized covariance for action recognitionabstractIn this paper we aim at increasing the descriptive power of the covariance matrix, limited in capturing linear mutual dependencies between variables only. We present a rigorous and principled mathematical pipeline to recover the kernel trick for computing the covariance matrix, enhancing it to model more complex, non-linear relationships conveyed by the raw data. To this end, we propose Kernelized-COV, which generalizes the original covariance representation without compromising the efficiency of the computation. In the experiments, we validate the proposed framework against many previous approaches in the literature, scoring on par or superior with respect to the state of the art on benchmark datasets for 3D action recognition. Jacopo Cavazza, Andrea Zunino, Marco San-Biagio, Vittorio Murino |
ICPR | 4 |
| 2016 | Unsupervised mouse behavior analysis: A data-driven study of mice interactionsabstractAutomatic analysis of rodent behavior has been receiving growing attention in recent years since rodents have been the reference species for many neuroscientific studies, with the social interaction being among the subjects of the most important ones. Systems that are employed in these studies are mainly based on tracking of mice and activity classification through supervised learning methods, trained on datasets manually annotated by experts. In this paper, we introduce a completely unsupervised way of analysing tracking data for the automatic identification of social and non-social behaviors using models capable of spotting regularities in the data. In particular, a mean-covariance Restricted Boltzmann Machine is employed to abstract higher-level behavioral configurations of mice interacting in an arena for a long time. Vasiliki-Maria Katsageorgiou, Matteo Zanotto, Valentina Ferretti 0002, Francesco Papaleo, Diego Sona, Vittorio Murino |
ICPR | 7 |
| 2016 | Fast 6D pose estimation for texture-less objects from a single RGB imageabstractA fundamental step to solve bin-picking and grasping problems is the accurate estimation of an object 3D pose. Such visual task usually rely on profusely textured objects: standard procedures such as detection of interest points or computation of appearance-based descriptors are favoured by using a highly informative surface. However, texture-less objects or their parts (i.e., those whose surface texture is poorly conditioned) are common in any environment but still challenging to deal with. This is due the fact that the distribution of surface brightness makes difficult to compute interest points or appearance-based descriptors. In this paper, we propose a method to estimate the 3D pose for texture-less objects given a coarse initialization: the pose is estimated using using edge correspondences, where the similarity measure is encoded using a pre-computed linear regression matrix. Furthermore, we also propose a method to increase the robustness of the estimated pose against background and object clutter. We validate both methods by using synthetic and real image sequences with objects with known ground truth. Enrique Muñoz, Yoshinori Konishi, Vittorio Murino, Alessio Del Bue |
ICRA | 3 |
| 2016 | Fast 6D pose from a single RGB image using Cascaded Forests TemplatesabstractThis paper presents a method for 6D pose estimation from a single RGB image for complex texture-less objects. This class of objects are common in any environment but still challenging to deal with. This is due to the fact that the distribution of surface brightness makes difficult to compute interest points or appearance-based descriptors. Here we propose a novel part-based method using an efficient template matching approach where each template independently encodes the similarity function using a Forest trained over the templates. Moreover, accuracy is even more incremented by using a cascade of the learned forest. These templates forests together with the simplicity of the computed image features allow a quick estimate of the pose achieving real-time performance. Performance are demonstrated both on synthetic and real images with known ground truth. Enrique Muñoz, Yoshinori Konishi, Carlos Beltrán 0002, Vittorio Murino, Alessio Del Bue |
IROS | 4 |
| 2016 | Effective Brain Connectivity Through a Constrained Autoregressive Model
Alessandro Crimi, Luca Dodero, Vittorio Murino, Diego Sona |
MICCAI (1) | 3 |
| 2016 | Traveling on discrete embeddings of gene expression
Pietro Lovato, Manuele Bicego, Maria Kesa, Nebojsa Jojic, Vittorio Murino, Alessandro Perina |
Artif. Intell. Medicine | 5 |
| 2016 | Detecting conversational groups in images and sequences: A robust game-theoretic approach
Sebastiano Vascon, Eyasu Zemene Mequanint, Marco Cristani, Hayley Hung, Marcello Pelillo, Vittorio Murino |
Comput. Vis. Image Underst. | 6 |
| 2016 | Bounding Multiple Gaussians Uncertainty with Application to Object Tracking
Baochang Zhang 0001, Alessandro Perina, Vittorio Murino, Jianzhuang Liu, Rongrong Ji |
Int. J. Comput. Vis. | 4 |
| 2016 | A Unifying Framework in Vector-valued Reproducing Kernel Hilbert Spaces for Manifold Regularization and Co-Regularized Multi-view LearningabstractThis paper presents a general vector-valued reproducing kernel Hilbert spaces (RKHS) framework for the problem of learning an unknown functional dependency between a structured input space and a structured output space. Our formulation encompasses both Vector-valued Manifold Regularization and Co-regularized Multi- view Learning, providing in particular a unifying framework linking these two important learning approaches. In the case of the least square loss function, we provide a closed form solution, which is obtained by solving a system of linear equations. In the case of Support Vector Machine (SVM) classification, our formulation generalizes in particular both the binary Laplacian SVM to the multi-class, multi-view settings and the multi-class Simplex Cone SVM to the semi-supervised, multi-view settings. The solution is obtained by solving a single quadratic optimization problem, as in standard SVM, via the Sequential Minimal Optimization (SMO) approach. Empirical results obtained on the task of object recognition, using several challenging data sets, demonstrate the competitiveness of our algorithms compared with other state-of-the-art methods. Hà Quang Minh, Loris Bazzani, Vittorio Murino |
J. Mach. Learn. Res. | 3 |
| 2016 | Guest Editorial Special Section on Learning in Non-(geo)metric SpacesabstractTraditional machine learning and pattern recognition techniques are intimately linked to the notion of feature spaces. Adopting this view, each object is described in terms of a vector of numerical attributes and is, therefore, mapped to a point in a Euclidean (geometric) vector space, so that the distances between the points reflect the observed (dis)similarities between the respective objects. This kind of representation is attractive because geometric spaces offer powerful analytical as well as computational tools that are simply not available in other representations. Indeed, classical machine learning methods are tightly related to geometrical concepts, and numerous powerful tools have been developed during the last few decades, starting from the maximal likelihood method in the 1920s to perceptrons in the 1960s and, more recently, to kernel machines and deep learning architectures. Marcello Pelillo, Edwin R. Hancock, Xuelong Li 0001, Vittorio Murino |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2015 | Latent subcategory models for pedestrian detection with partial occlusion handlingabstractPedestrian detection is one of the most important tasks in Computer Vision, especially in automotive and security applications. One of the most common problems in real scenarios is related to the detection of occluded pedestrians. In this paper, we propose a novel multi-cue pedestrian detection approach able to deal with non homogeneous object samples by learning latent subcategory models trained on both visual and depth-based features. We also propose a novel self-similarity based feature, namely SSTD, to encode the homogeneity in appearance of pedestrians characterized by similar occlusion patterns. Experiments are performed on the Daimler Pedestrian Detection Benchmark Dataset showing the robustness of our approach in actual scenarios. Samuele Martelli, Marco San-Biagio, Vittorio Murino |
AVSS | 3 |
| 2015 | Violence detection in crowded scenes using substantial derivativeabstractThis paper presents a novel video descriptor based on substantial derivative, an important concept in fluid mechanics, that captures the rate of change of a fluid property as it travels through a velocity field. Unlike standard approaches that only use temporal motion information, our descriptor exploits the spatio-temporal characteristic of substantial derivative. In particular, the spatial and temporal motion patterns are captured by respectively the convective and local accelerations. After estimating the convective and local field from the optic flow, we followed the standard bag-of-word procedure for each motion pattern separately, and we concatenated the two resulting histograms to form the final descriptor. We extensively evaluated the effectiveness of the proposed method on five benchmarks, including three standard datasets (Violence in Movies, Violence In Crowd, and BEHAVE), and two new video-survelliance sequences downloaded from Youtube. Our experiments show how the proposed approach sets the new state-of-the-art on all benchmarks and how the structural information captured by convective acceleration is essential to detect violent episodes in crowded scenarios. Sadegh Mohammadi 0001, Hamed Kiani Galoogahi, Alessandro Perina, Vittorio Murino |
AVSS | 4 |
| 2015 | Learning with dataset bias in latent subcategory modelsabstractLatent subcategory models (LSMs) offer significant improvements over training linear support vector machines (SVMs). Training LSMs is a challenging task due to the potentially large number of local optima in the objective function and the increased model complexity which requires large training set sizes. Often, larger datasets are available as a collection of heterogeneous datasets. However, previous work has highlighted the possible danger of simply training a model from the combined datasets, due to the presence of bias. In this paper, we present a model which jointly learns an LSM for each dataset as well as a compound LSM. The method provides a means to borrow statistical strength from the datasets while reducing their inherent bias. In experiments we demonstrate that the compound LSM, when tested on PASCAL, LabelMe, Caltech101 and SUN09 in a leave-one-dataset-out fashion, achieves an average improvement of over 6.5% over a previous SVM-based undoing bias approach and an average improvement of over 8.5% over a standard LSM trained on the concatenation of the datasets. Dimitris Stamos, Samuele Martelli, Moin Nabi, Andrew M. McDonald, Vittorio Murino, Massimiliano Pontil |
CVPR | 5 |
| 2015 | Sparse representation classification with manifold constraints transferabstractThe fact that image data samples lie on a manifold has been successfully exploited in many learning and inference problems. In this paper we leverage the specific structure of data in order to improve recognition accuracies in general recognition tasks. In particular we propose a novel framework that allows to embed manifold priors into sparse representation-based classification (SRC) approaches. We also show that manifold constraints can be transferred from the data to the optimized variables if these are linearly correlated. Using this new insight, we define an efficient alternating direction method of multipliers (ADMM) that can consistently integrate the manifold constraints during the optimization process. This is based on the property that we can recast the problem as the projection over the manifold via a linear embedding method based on the Geodesic distance. The proposed approach is successfully applied on face, digit, action and objects recognition showing a consistently increase on performance when compared to the state of the art. Baochang Zhang 0001, Alessandro Perina, Vittorio Murino, Alessio Del Bue |
CVPR | 3 |
| 2015 | Exploiting multiple detections to learn robust brightness transfer functions in re-identification systemsabstractRe-identification systems aim at recognizing the same individuals in multiple cameras and one of the most relevant problems is that the appearance of same individual varies across cameras due to illumination and viewpoint changes. This paper proposes the use of Cumulative Weighted Brightness Transfer Functions to model this appearance variations. It is multiple frame-based learning approach which leverages consecutive detections of each individual to transfer the appearance, rather than learning brightness transfer function from pairs of images. We tested our approach on standard multi-camera surveillance datasets showing consistent and significant improvements over existing methods on three different datasets without any other additional cost. Our approach is general and can be applied to any appearance-based method. Amran Bhuiyan, Alessandro Perina, Vittorio Murino |
ICIP | 3 |
| 2015 | Crowd motion monitoring using tracklet-based commotion measureabstractAbnormal detection in crowd is a challenging vision task due to the scarcity of real-world training examples and the lack of a clear definition of abnormality. To tackle these challenges, we propose a novel measure to capture the commotion of a crowd motion for the task of abnormality detection in crowd. The unsupervised nature of the proposed measure allows to detect abnormality adaptively (i.e. context dependent) with no training cost. The extensive experiments on three different levels (e.g. pixel, frame and video) show the superiority of the proposed approach compared to the state of the arts. Hossein Mousavi, Moin Nabi, Hamed Kiani Galoogahi, Alessandro Perina, Vittorio Murino |
ICIP | 5 |
| 2015 | Improving FREAK Descriptor for Image Classification
Cristina Hilario Gomez, N. V. Kartheek Medathati, Pierre Kornprobst, Vittorio Murino, Diego Sona |
ICVS | 4 |
| 2015 | Kernel-Based Analysis of Functional Brain Connectivity on Grassmann Manifold
Luca Dodero, Fabio Sambataro, Vittorio Murino, Diego Sona |
MICCAI (3) | 3 |
| 2015 | Analyzing Tracklets for the Detection of Abnormal Crowd BehaviorabstractThis paper presents a novel video descriptor, referred to as Histogram of Oriented Track lets, for recognizing abnormal situation in crowded scenes. Unlike standard approaches that use optical flow, which estimates motion vectors only from two successive frames, we built our descriptor over long-range motion trajectories which is called track lets in the literature. Following the standard procedure, we divided video sequences in spatio-temporal cuboids within which we collected statistics on the track lets passing through them. In particular, we quantized orientation and magnitude in a 2-dimensional histogram which encodes the motion patterns expected in each cuboid. We classify frames as normal and abnormal by using Latent Dirichlet Allocation and Support Vector Machines. We evaluated the effectiveness of the proposed descriptors on three datasets: UCSD, Violence in Crowds and UMN. The experiments demonstrated (i) very promising results in abnormality detection, (ii) setting new state-of-the-art on two of them, and (iii) outperforming former descriptors based on the optical flow, dense trajectories and the social force model. Hossein Mousavi, Sadegh Mohammadi 0001, Alessandro Perina, Ryad Chellali, Vittorio Murino |
WACV | 5 |
| 2015 | Semantic Multi-body Motion SegmentationabstractThis paper presents a method to deal with the multi-body segmentation problem using a set of 2D points matches between two views. The key feature of our approach is the explicit inclusion of a higher semantic information as given by general purpose object detectors that boost the segmentation of the moving objects. In the classical formulation of the problem, only 2D matched points between views are used to identify independently moving objects based on the principle that a set of points belonging to a moving object would satisfy some given multi-view relations (e.g. multi-body epipolar constraints). We improve and speedup such process by including the information that a set of 2D matches may belong to the same object given the output of a detector. As such, instead of sampling points uniformly with a RANSAC based strategy, the selection of the matches is driven by the position and score confidence of the object detectors. Evaluation on challenging synthetic and real datasets shows a remarkable improvement in respect to previous approaches, regarding both the number of iterations required to segment a scene and the effectiveness of the segmentation itself, often making the difference between satisfying segmentation and almost complete failure. Cosimo Rubino, Marco Crocco, Vittorio Murino, Alessio Del Bue |
WACV | 3 |
| 2015 | Adaptive Local Movement Modelling for Object TrackingabstractIn this paper we present a novel strategy for modelling the motion of local patches for single object tracking that can be seamlessly applied to most part-based trackers in the literature. The proposed Adaptive Local Movement Modelling (ALMM) method is able to model the local spatial distribution of the image patches defining the object to track and the reliability of each image patch. Given the output of a base tracking algorithm, a Gaussian Mixture Model (GMM) is first used to model the distribution of the movement of local patches relative to the gravity center of the tracked object. Then, the GMM is combined with the base tracker in a boosting framework, which gives a novel integrated boosting classifier for the tracking task. This provides a robust procedure to detect outliers in the local motion of the patches. The algorithm is highly configurable with the possibility to change the number of local patches used for tracking and to adapt to the variations of the tracked object. Tracking results on standard datasets show that equipping state-of-the-art trackers with our tehcnique remarkably improves their performance. Baochang Zhang 0001, Alessandro Perina, Alessio Del Bue, Vittorio Murino |
WACV | 5 |
| 2015 | Non-myopic information theoretic sensor management of a single pan-tilt-zoom camera for multiple object detection and tracking
Pietro Salvagnini, Federico Pernici, Marco Cristani, Giuseppe Lisanti, Alberto Del Bimbo, Vittorio Murino |
Comput. Vis. Image Underst. | 6 |
| 2015 | Bridging the gap in connectomic studies: A particle filtering framework for estimating structural connectivity at network scale
Simona Ullo, Vittorio Murino, Alessandro Maccione, Luca Berdondini, Diego Sona |
Medical Image Anal. | 2 |
| 2015 | Joint Individual-Group Modeling for TrackingabstractWe present a novel probabilistic framework that jointly models individuals and groups for tracking. Managing groups is challenging, primarily because of their nonlinear dynamics and complex layout which lead to repeated splitting and merging events. The proposed approach assumes a tight relation of mutual support between the modeling of individuals and groups, promoting the idea that groups are better modeled if individuals are considered and vice versa. This concept is translated in a mathematical model using a decentralized particle filtering framework which deals with a joint individual-group state space. The model factorizes the joint space into two dependent subspaces, where individuals and groups share the knowledge of the joint individual-group distribution. The assignment of people to the different groups (and thus group initialization, split and merge) is implemented by two alternative strategies: using classifiers trained beforehand on statistics of group configurations, and through online learning of a Dirichlet process mixture model, assuming that no training data is available before tracking. These strategies lead to two different methods that can be used on top of any person detector (simulated using the ground truth in our experiments). We provide convincing results on two recent challenging tracking benchmarks. Loris Bazzani, Matteo Zanotto, Marco Cristani, Vittorio Murino |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | A Game-Theoretic Probabilistic Approach for Detecting Conversational Groups
Sebastiano Vascon, Eyasu Zemene Mequanint, Marco Cristani, Hayley Hung, Marcello Pelillo, Vittorio Murino |
ACCV (5) | 6 |
| 2014 | Location recognition on lifelog images via a discriminative combination of generative models
Alessandro Perina, Matteo Zanotto, Baochang Zhang 0001, Vittorio Murino |
BMVC | 4 |
| 2014 | Weighted bag of visual words for object recognitionabstractBag of Visual words (BoV) is one of the most successful strategy for object recognition, used to represent an image as a vector of counts using a learned vocabulary. This strategy assumes that the representation is built using patches that are either densely extracted or sampled from the images using feature detectors. However, the dense strategy captures also the noisy background information, whereas the feature detection strategy can lose important parts of the objects. In this paper we propose a solution in-between these two strategies, by densely extracting patches from the image, and weighting them accordingly to their salience. Intuitively, highly salient patches have an important role in describing an object, while those with low saliency are still taken with low emphasis, instead of discarding them. We embed this idea in the word encoding mechanism adopted in the BoV approaches. The technique is successfully applied to vector quantization and Fisher vector, on Caltech-101 and Caltech-256. Marco San-Biagio, Loris Bazzani, Marco Cristani, Vittorio Murino |
ICIP | 4 |
| 2014 | A directional visual descriptor for large-scale coverage problemsabstractVisual coverage of large scale environments is a challenging problem that has many practical applications such as large scale 3D reconstruction, search and rescue and active video surveillance. In this paper, we consider a setting where mobile robots must acquire visual information using standard cameras, while minimizing associated movement costs. The main source of complexity for such scenario is the lack of a priori knowledge of 3D structures for the surrounding environment. To address this problem, we propose a novel descriptor for visual coverage that aims at measuring the orientation dependent visual information of an area, based on a regular discretization of the 3D environment in voxels. Next, we use the proposed visual descriptor to define an autonomous cooperative exploration approach, which controls the robot movements so to maximize information accuracy and minimizing movement costs. We empirically evaluate our approach in a simulation scenario based on real data for large scale 3D environments, and on widely used robotic tools (such as ROS and Stage). Experimental results show that the proposed method significantly outperforms a baseline random approach and an uncoordinated one, thus being a valid proposal for visual coverage in large scale outdoor scenarios. Marco Tamassia, Alessandro Farinelli, Vittorio Murino, Alessio Del Bue |
IROS | 3 |
| 2014 | Group-Wise Functional Community Detection through Joint Laplacian Diagonalization
Luca Dodero, Alessandro Gozzi, Adam Liska, Vittorio Murino, Diego Sona |
MICCAI (2) | 4 |
| 2014 | Mapping Brains on Grids of Features for Schizophrenia Analysis
Alessandro Perina, Denis Peruzzo, Maria Kesa, Nebojsa Jojic, Vittorio Murino, Mellani Bellani, Paolo Brambilla, Umberto Castellani |
MICCAI (2) | 5 |
| 2014 | Log-Hilbert-Schmidt metric between positive definite operators on Hilbert spaces
Hà Quang Minh, Marco San-Biagio, Vittorio Murino |
NIPS | 3 |
| 2014 | Information theoretic sensor management for multi-target tracking with a single pan-tilt-zoom cameraabstractAutomatic multiple target tracking with pan-tilt-zoom (PTZ) cameras is a hard task, with few approaches in the literature, most of them proposing simplistic scenarios. In this paper, we present a PTZ camera management framework which lies on information theoretic principles: at each time step, the next camera pose (pan, tilt, focal length) is chosen, according to a policy which ensures maximum information gain. The formulation takes into account occlusions, physical extension of targets, realistic pedestrian detectors and the mechanical constraints of the camera. Convincing comparative results on synthetic data, realistic simulations and the implementation on a real video surveillance camera validate the effectiveness of the proposed method. Pietro Salvagnini, Federico Pernici, Marco Cristani, Giuseppe Lisanti, Iacopo Masi, Alberto Del Bimbo, Vittorio Murino |
WACV | 7 |
| 2014 | Generative embeddings based on Rician mixtures for kernel-based classification of magnetic resonance images
Anna C. Carli, Mário A. T. Figueiredo, Manuele Bicego, Vittorio Murino |
Neurocomputing | 4 |
| 2014 | Encoding Structural Similarity by Cross-covariance Tensors for Image ClassificationabstractIn computer vision, an object can be modeled in two main ways: by explicitly measuring its characteristics in terms of feature vectors, and by capturing the relations which link an object with some exemplars, that is, in terms of similarities. In this paper, we propose a new similarity-based descriptor, dubbed structural similarity cross-covariance tensor (SS-CCT), where self-similarities come into play: Here the entity to be measured and the exemplar are regions of the same object, and their similarities are encoded in terms of cross-covariance matrices. These matrices are computed from a set of low-level feature vectors extracted from pairs of regions that cover the entire image. SS-CCT shares some similarities with the widely used covariance matrix descriptor, but extends its power focusing on structural similarities across multiple parts of an image, instead of capturing local similarities in a single region. The effectiveness of SS-CCT is tested on many diverse classification scenarios, considering objects and scenes on widely known benchmarks (Caltech-101, Caltech-256, PASCAL VOC 2007 and SenseCam). In all the cases, the results obtained demonstrate the superiority of our new descriptor against diverse competitors. Furthermore, we also reported an analysis on the reduced computational burden achieved by using and efficient implementation that takes advantage from the integral image representation. Marco San-Biagio, Samuele Martelli, Marco Crocco, Marco Cristani, Vittorio Murino |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2013 | Semi-supervised multi-feature learning for person re-identificationabstractPerson re-identification is probably the open challenge for low-level video surveillance in the presence of a camera network with non-overlapped fields of view. A large number of direct approaches has emerged in the last five years, often proposing novel visual features specifically designed to highlight the most discriminant aspects of people, which are invariant to pose, scale and illumination. On the other hand, learning-based methods are usually based on simpler features, and are trained on pairs of cameras to discriminate between individuals. In this paper, we present a method that joins these two ideas: given an arbitrary state-of-the-art set of features, no matter their number, dimensionality or descriptor, the proposed multi-class learning approach learns how to fuse them, ensuring that the features agree on the classification result. The approach consists of a semi-supervised multi-feature learning strategy, that requires at least a single image per person as training data. To validate our approach, we present results on different datasets, using several heterogeneous features, that set a new level of performance in the person re-identification problem. Dario Figueira, Loris Bazzani, Hà Quang Minh, Marco Cristani, Alexandre Bernardino, Vittorio Murino |
AVSS | 6 |
| 2013 | Reading between the turns: Statistical modeling for identity recognition and verification in chatsabstractIdentity safekeeping has recently become an important problem for the social web: as a case study, we focus here on instant messaging platforms, proposing novel soft-biometric cues for user recognition and verification. Specifically, we design a set of features encoding effectively how a person converses: since chats are crossbreeds of written text and face-to-face verbal communication, the features inherit equally from textual authorship attribution and conversational analysis of speech. Importantly, our cues ignore completely the semantics of the chat, relying solely on non-verbal aspects, taking care of possible privacy and ethical issues. We apply our approach on a novel dataset of 94 different individuals, whose chat conversations have been recorded for an average period of five months; recognition rate, intended as normalized AUC on CMC curve, is 95.73%, while verification rate amounts to 95.66%, as normalized AUC on ROC curve. Giorgio Roffo, Cristina Segalin, Alessandro Vinciarelli, Vittorio Murino, Marco Cristani |
AVSS | 4 |
| 2013 | SDALF+C: Augmenting the SDALF Descriptor by Relation-Based Information for Multi-shot Re-identification
Sylvie Jasmine Poletti, Vittorio Murino, Marco Cristani |
CIARP (2) | 2 |
| 2013 | Statistical Analysis of Visual Attentional Patterns for Video Surveillance
Giorgio Roffo, Marco Cristani, Frank E. Pollick, Cristina Segalin, Vittorio Murino |
CIARP (2) | 5 |
| 2013 | Encoding Classes of Unaligned Objects Using Structural Similarity Cross-Covariance Tensors
Marco San-Biagio, Samuele Martelli, Marco Crocco, Marco Cristani, Vittorio Murino |
CIARP (1) | 5 |
| 2013 | Heterogeneous Auto-similarities of Characteristics (HASC): Exploiting Relational Information for ClassificationabstractCapturing the essential characteristics of visual objects by considering how their features are inter-related is a recent philosophy of object classification. In this paper, we embed this principle in a novel image descriptor, dubbed Heterogeneous Auto-Similarities of Characteristics (HASC). HASC is applied to heterogeneous dense features maps, encoding linear relations by co variances and nonlinear associations through information-theoretic measures such as mutual information and entropy. In this way, highly complex structural information can be expressed in a compact, scale invariant and robust manner. The effectiveness of HASC is tested on many diverse detection and classification scenarios, considering objects, textures and pedestrians, on widely known benchmarks (Caltech-101, Brodatz, Daimler Multi-Cue). In all the cases, the results obtained with standard classifiers demonstrate the superiority of HASC with respect to the most adopted local feature descriptors nowadays, such as SIFT, HOG, LBP and feature co variances. In addition, HASC sets the state-of-the-art on the Brodatz texture dataset and the Daimler Multi-Cue pedestrian dataset, without exploiting ad-hoc sophisticated classifiers. Marco San-Biagio, Marco Crocco, Marco Cristani, Samuele Martelli, Vittorio Murino |
ICCV | 5 |
| 2013 | Person re-identification with a PTZ camera: An introductory studyabstractWe present an introductory study that paves the way for a new kind of person re-identification, by exploiting a single Pan-Tilt-Zoom (PTZ) camera. PTZ devices allow to zoom on body regions, acquiring discriminative visual patterns that enrich the appearance description of an individual. This intuition has been translated into a statistical direct reidentification scheme, which collects two images for each probe subject: the first image captures the probe individual, focusing on the whole body; the second can be a zoomed body part (head, torso or legs) or another whole body image, and is the outcome of an action-selection mechanism, driven by feature selection principles. The validation of this technique is also explored: in order to allow repeatability, two novel multi-resolution benchmarks have been created. On these data, we demonstrate that our approach selects effective actions, by focusing on body portions which discriminate each subject. Moreover, we show that the proposed compound of two images overwhelms standard multi-shot descriptions, composed by many more pictures. Pietro Salvagnini, Loris Bazzani, Marco Cristani, Vittorio Murino |
ICIP | 4 |
| 2013 | Multi-scale f-formation discovery for group detectionabstractWe present an unsupervised approach for the automatic detection of static interactive groups. The approach builds upon a novel multi-scale Hough voting policy, which incorporates in a flexible way the sociological notion of group as F-formation; the goal is to model at the same time small arrangements of close friends and aggregations of many individuals spread over a large area. Our technique is based on a competition of different voting sessions, each one specialized for a particular group cardinality; all the votes are then evaluated using information theoretic criteria, producing the final set of groups. The proposed technique has been applied on public benchmark sequences and a novel cocktail party dataset, evaluating new group detection metrics and obtaining state-of-the-art performances. Francesco Setti, Oswald Lanz, Roberta Ferrario, Vittorio Murino, Marco Cristani |
ICIP | 4 |
| 2013 | A unifying framework for vector-valued manifold regularization and multi-view learningabstractThis paper presents a general vector-valued reproducing kernel Hilbert spaces (RKHS) formulation for the problem of learning an unknown functional dependency between a structured input space and a structured output space, in the Semi-Supervised Learning setting. Our formulation includes as special cases Vector-valued Manifold Regularization and Multi-view Learning, thus provides in particular a unifying framework linking these two important learning approaches. In the case of least square loss function, we provide a closed form solution with an efficient implementation. Numerical experiments on challenging multi-class categorization problems show that our multi-view learning formulation achieves results which are comparable with state of the art and are significantly better than single-view learning. Hà Quang Minh, Loris Bazzani, Vittorio Murino |
ICML (2) | 3 |
| 2013 | Symmetry-driven accumulation of local features for human characterization and re-identification
Loris Bazzani, Marco Cristani, Vittorio Murino |
Comput. Vis. Image Underst. | 3 |
| 2013 | Social interactions by visual focus of attention in a three-dimensional environmentabstractAbstract In human behaviour analysis, the visual focus of attention (VFOA) of a person is a very important cue. VFOA detection is difficult, though, especially in a unconstrained and crowded environment, typical of video surveillance scenarios. In this paper, we estimate the VFOA by defining the Subjective View Frustum, which approximates the visual field of a person in a three‐dimensional representation of the scene. This opens up to several intriguing behavioural investigations. In particular, we propose the Inter‐Relation Pattern Matrix, which suggests possible social interactions between the people present in a scene. Theoretical justifications and experimental results substantiate the validity and the goodness of the analysis performed. Loris Bazzani, Marco Cristani, Diego Tosato, Michela Farenzena, Giulia Paggetti, Gloria Menegaz, Vittorio Murino |
Expert Syst. J. Knowl. Eng. | 7 |
| 2013 | Combining information theoretic kernels with generative embeddings for classification
Manuele Bicego, Aydin Ulas, Umberto Castellani, Alessandro Perina, Vittorio Murino, André F. T. Martins, Pedro M. Q. Aguiar, Mário A. T. Figueiredo |
Neurocomputing | 5 |
| 2013 | Human behavior analysis in video surveillance: A Social Signal Processing perspective
Marco Cristani, Ramachandra Raghavendra, Alessio Del Bue, Vittorio Murino |
Neurocomputing | 4 |
| 2013 | Characterizing Humans on Riemannian ManifoldsabstractIn surveillance applications, head and body orientation of people is of primary importance for assessing many behavioral traits. Unfortunately, in this context people are often encoded by a few, noisy pixels so that their characterization is difficult. We face this issue, proposing a computational framework which is based on an expressive descriptor, the covariance of features. Covariances have been employed for pedestrian detection purposes, actually a binary classification problem on Riemannian manifolds. In this paper, we show how to extend to the multiclassification case, presenting a novel descriptor, named weighted array of covariances, especially suited for dealing with tiny image representations. The extension requires a novel differential geometry approach in which covariances are projected on a unique tangent space where standard machine learning techniques can be applied. In particular, we adopt the Campbell-Baker-Hausdorff expansion as a means to approximate on the tangent space the genuine (geodesic) distances on the manifold in a very efficient way. We test our methodology on multiple benchmark datasets, and also propose new testing sets, getting convincing results in all the cases. Diego Tosato, Mauro Spera, Marco Cristani, Vittorio Murino |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Stereo-Based Framework for Pedestrian Detection with Partial Occlusion HandlingabstractThe pedestrian detection literature has been recently extended by the availability of large-scale multisensory datasets, able to capture complementary aspects of the objects of interest, namely, appearance, motion, and depth. In this paper, we exploit this multimodal scenario to propose a new set of composite descriptors dubbed CO2, CO-variances of visual features and CO-occurrences of depth fields. Covariances of visual features allow us to integrate at low-level heterogeneous visual cues related to intensity and texture. Co-occurrences of depth fields are brand new descriptors, which use range information for characterizing the global shape of a pedestrian while being also able to identify its occluded parts. This paper illustrates how these descriptors can be instantiated and combined together, improving detection capabilities taking also benefit from the proper handling of occlusions. Experimental results show that CO2, fed into a standard discriminative classification system, set state-of-the-art performances on recent multi-modal intensity- and stereo-based pedestrian datasets. Samuele Martelli, Marco Cristani, Vittorio Murino |
AVSS | 3 |
| 2012 | A Closed Form Solution for the Self-Calibration of Heterogeneous SensorsabstractWe present a novel closed-form solution for the joint self-calibration of video and range sensors. The approach single assumption is the availability of synchronous time of flight (i.e., range distances) measurements and visual position of the target on images acquired by a set of cameras. In such case, we make explicit a rank constraint that is valid for both image and range data. This rank property is used to find an initial and affine solution via bilinear factorization, which is then corrected by enforcing the metric constraints characteristic for both sensor modalities (i.e., camera and anchors constraints). The output of the algorithm is the identification of the target/range sensor position and the calibration of the cameras. The application extent of our approach is broad and versatile. In fact, with the same framework, we can deal with, but not restricted to, two very different applications. The first is aimed at calibrating cameras and microphones deployed in an unknown environment. The second uses a RGB-D device to reconstruct the 3D position of a set of keypoints using the camera and depth map images. Synthetic and real tests show the algorithm performance under different levels of noise and configurations of target locations, number of sensors and cameras. Marco Crocco, Alessio Del Bue, Igor Barros Barbosa, Vittorio Murino |
BMVC | 4 |
| 2012 | Online Bayesian Non-parametrics for Social Group Detection
Matteo Zanotto, Loris Bazzani, Marco Cristani, Vittorio Murino |
BMVC | 4 |
| 2012 | Decentralized particle filter for joint individual-group trackingabstractIn this paper, we address the task of tracking groups of people in surveillance scenarios. This is a major challenge in computer vision, since groups are structured entities, subjected to repeated split and merge events. Our solution is a joint individual-group tracking framework, inspired by a recent technique dubbed decentralized particle filtering. The proposed strategy factorizes the joint individual-group state space in two dependent subspaces where individuals and groups share the knowledge of the joint individual-group distribution. In practice, we establish a tight relation of mutual support between the modeling of individuals and that of groups, promoting the idea that groups are better tracked if individuals are considered, and viceversa. Extensive experiments on a published and novel dataset validate our intuition, opening up to many future developments. Loris Bazzani, Marco Cristani, Vittorio Murino |
CVPR | 3 |
| 2012 | A regularized spectral algorithm for Hidden Markov Models with applications in computer visionabstractHidden Markov Models (HMMs) are among the most important and widely used techniques to deal with sequential or temporal data. Their application in computer vision ranges from action/gesture recognition to videosurveillance through shape analysis. Although HMMs are often embedded in complex frameworks, this paper focuses on theoretical aspects of HMM learning. We propose a regularized algorithm for learning HMMs in the spectral framework, whose computations have no local minima. Compared with recently proposed spectral algorithms for HMMs, our method is guaranteed to produce probability values which are always physically meaningful and which, on synthetic mathematical models, give very good approximations to true probability values. Furthermore, we place no restriction on the number of symbols and the number of states. On various pattern recognition data sets, our algorithm consistently outperforms classical HMMs, both in accuracy and computational speed. This and the fact that HMMs are used in vision as building blocks for more powerful classification approaches, such as generative embedding approaches or more complex generative models, strongly support spectral HMMs (SHMMs) as a new basic tool for pattern recognition. Hà Quang Minh, Marco Cristani, Alessandro Perina, Vittorio Murino |
CVPR | 4 |
| 2012 | Learning Discriminative Spatial Relations for Detector Dictionaries: An Application to Pedestrian Detection
Enver Sangineto, Marco Cristani, Alessio Del Bue, Vittorio Murino |
ECCV (2) | 4 |
| 2012 | Low-level multimodal integration on Riemannian manifolds for automatic pedestrian detection
Marco San-Biagio, Marco Crocco, Marco Cristani, Samuele Martelli, Vittorio Murino |
FUSION | 5 |
| 2012 | A closed form solution to the microphone position self-calibration problemabstractThis paper presents a novel algorithm for the automatic 3D localization of a set of microphones in an unknown environment. Given the times of arrival at each microphone of a set of sound events, the approach simultaneously estimates the 3D positions of the sensors and the sources that have generated the events. The only assumption made is that the emission time of the sound events must be known in order to measure the time of flight for each event. A closed form solution is also proposed whenever a sound event coincides with a microphone position. Simulated and real experiments show the validity of the approach for different setups of sensors and number of events. Marco Crocco, Alessio Del Bue, Matteo Bustreo, Vittorio Murino |
ICASSP | 4 |
| 2012 | Piecewise single view Photometric Stereo with multi-view constraintsabstractThis paper presents a novel Photometric Stereo approach for static views that recasts the problem into a piecewise formulation. The proposed algorithm, called Piecewise Photometric Stereo (PPS), entails several advantages in respect to previous global approaches. It is intrinsically more efficient, since reconstructing the surface in patches is computationally faster than reconstructing the global surface. Each patch has been associated an individual photometric model rather than a single global model as used in classical approaches. In this way, the piecewise formulation may grasp more complex lighting effects. Finally, the global metric properties of the shape is preserved using the multi-view constraints. In this pipeline, structure from motion is exploited to define such set of constraints and to compose a 3D mesh representing the metric structure of the object. Real results with ground truth show the positive performance of our algorithm compared with a classical global approach for Photometric Stereo. Reza Sabzevari, Alessio Del Bue, Vittorio Murino |
ICIP | 3 |
| 2012 | A joint structural and functional analysis of in-vitro neuronal networksabstractThe acquisition, analysis and representation of experimental data describing both anatomical and functional information at cellular level is an innovative opportunity to investigate neuronal network processing and organization. In this paper we propose an image processing pipeline to study in-vitro neuronal networks with a joint analysis of anatomy and electrophysiology. Neuronal nuclei are detected by segmenting fluorescence images of neuronal cultures. The high resolution Multi Electrode Arrays (MEAs) technology is used to collect functional information on cellular electrophysiological activity. Finally, detailed maps, representing both structural and functional information, are obtained which provide statistics on neuron distribution and spiking activity. Simona Ullo, Alessio Del Bue, Alessandro Maccione, Luca Berdondini, Vittorio Murino |
ICIP | 5 |
| 2012 | Segmentation and tracking of multiple interacting mice by temperature and shape information
Luca Giancardo, Diego Sona, Diego Scheggia, Francesco Papaleo, Vittorio Murino |
ICPR | 5 |
| 2012 | Joining feature-based and similarity-based pattern description paradigms for object detection
Samuele Martelli, Marco Cristani, Loris Bazzani, Diego Tosato, Vittorio Murino |
ICPR | 5 |
| 2012 | Inner product tree for improved Orthogonal Matching Pursuit
Paolo Piro, Diego Sona, Vittorio Murino |
ICPR | 3 |
| 2012 | A multiple kernel learning approach to multi-modal pedestrian classification
Marco San-Biagio, Aydin Ulas, Marco Crocco, Marco Cristani, Umberto Castellani, Vittorio Murino |
ICPR | 6 |
| 2012 | Generative Embeddings based on Rician Mixtures - Application to Kernel-based Discriminative Classification of Magnetic Resonance Images
Anna C. Carli, Mário A. T. Figueiredo, Manuele Bicego, Vittorio Murino |
ICPRAM (1) | 4 |
| 2012 | Conversationally-inspired stylometric features for authorship attribution in instant messagingabstractAuthorship attribution (AA) aims at recognizing automatically the author of a given text sample. Traditionally applied to literary texts, AA faces now the new challenge of recognizing the identity of people involved in chat conversations. These share many aspects with spoken conversations, but AA approaches did not take it into account so far. Hence, this paper tries to fill the gap and proposes two novelties that improve the effectiveness of traditional AA approaches for this type of data: the first is to adopt features inspired by Conversation Analysis (in particular for turn-taking), the second is to extract the features from individual turns rather than from entire conversations. The experiments have been performed over a corpus of dyadic chat conversations (77 individuals in total). The performance in identifying the persons involved in each exchange, measured in terms of area under the Cumulative Match Characteristic curve, is 89.5%. Marco Cristani, Giorgio Roffo, Cristina Segalin, Loris Bazzani, Alessandro Vinciarelli, Vittorio Murino |
ACM Multimedia | 6 |
| 2012 | Stel Component Analysis: Joint Segmentation, Modeling and Recognition of Objects ClassesabstractModels that captures the common structure of an object class have appeared few years ago in the literature (Jojic and Caspi in Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), pp. 212–219, 2004 ; Winn and Jojic in Proceedings of International Conference on Computer Vision (ICCV), pp. 756–763, 2005 ); they are often referred as “stel models.” Their main characteristic is to segment objects in clear, often semantic, parts as a consequence of the modeling constraint which forces the regions belonging to a single segment to have a tight distribution over local measurements, such as color or texture. This self-similarity within a region in a single image is typical of many meaningful image parts, even when across different images of similar objects, the corresponding parts may not have similar local measurements. Moreover, the segmentation itself is expected to be consistent within a class, although still flexible. These models have been applied mostly to segmentation scenarios. In this paper, we extent those ideas (1) proposing to capture correlations that exist in structural elements of an image class due to global effects, (2) exploiting the segmentations to capture feature co-occurrences and (3) allowing the use of multiple, eventually sparse, observation of different nature. In this way we obtain richer models more suitable to recognition tasks. We accomplish these requirements using a novel approach we dubbed stel component analysis . Experimental results show the flexibility of the model as it can deal successfully with image/video segmentation and object recognition where, in particular, it can be used as an alternative of, or in conjunction with, bag-of-features and related classifiers, where stel inference provides a meaningful spatial partition of features. Alessandro Perina, Nebojsa Jojic, Marco Cristani, Vittorio Murino |
Int. J. Comput. Vis. | 4 |
| 2012 | Free Energy Score Spaces: Using Generative Information in Discriminative ClassifiersabstractA score function induced by a generative model of the data can provide a feature vector of a fixed dimension for each data sample. Data samples themselves may be of differing lengths (e.g., speech segments or other sequential data), but as a score function is based on the properties of the data generation process, it produces a fixed-length vector in a highly informative space, typically referred to as "score space." Discriminative classifiers have been shown to achieve higher performances in appropriately chosen score spaces with respect to what is achievable by either the corresponding generative likelihood-based classifiers or the discriminative classifiers using standard feature extractors. In this paper, we present a novel score space that exploits the free energy associated with a generative model. The resulting free energy score space (FESS) takes into account the latent structure of the data at various levels and can be shown to lead to classification performance that at least matches the performance of the free energy classifier based on the same generative model and the same factorization of the posterior. We also show that in several typical computer vision and computational biology applications the classifiers optimized in FESS outperform the corresponding pure generative approaches, as well as a number of previous approaches combining discriminating and generative models. Alessandro Perina, Marco Cristani, Umberto Castellani, Vittorio Murino, Nebojsa Jojic |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Multiple-shot person re-identification by chromatic and epitomic analyses
Loris Bazzani, Marco Cristani, Alessandro Perina, Vittorio Murino |
Pattern Recognit. Lett. | 4 |
| 2012 | Investigating Topic Models' Capabilities in Expression Microarray Data ClassificationabstractIn recent years a particular class of probabilistic graphical models-called topic models-has proven to represent an useful and interpretable tool for understanding and mining microarray data. In this context, such models have been almost only applied in the clustering scenario, whereas the classification task has been disregarded by researchers. In this paper, we thoroughly investigate the use of topic models for classification of microarray data, starting from ideas proposed in other fields (e.g., computer vision). A classification scheme is proposed, based on highly interpretable features extracted from topic models, resulting in a hybrid generative-discriminative approach; an extensive experimental evaluation, involving 10 different literature benchmarks, confirms the suitability of the topic models for classifying expression microarray data. Manuele Bicego, Pietro Lovato, Alessandro Perina, Marianna Fasoli, Massimo Delledonne, Mario Pezzotti, Annalisa Polverari, Vittorio Murino |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2011 | Custom Pictorial Structures for Re-identificationabstractWe propose a novel methodology for re-identification, based on Pictorial Structures (PS). Whenever face or other biometric information is missing, humans recognize an individual by selectively focusing on the body parts, looking for part-to-part correspondences. We want to take inspiration from this strategy in a re-identification context, using PS to achieve this objective. For single image re-identification, we adopt PS to localize the parts, extract and match their descriptors. When multiple images of a single individual are available, we propose a new algorithm to customize the fit of PS on that specific person, leading to what we call a Custom Pictorial Structure (CPS). CPS learns the appearance of an individual, improving the localization of its parts, thus obtaining more reliable visual characteristics for re-identification. It is based on the statistical learning of pixel attributes collected through spatio-temporal reasoning. The use of PS and CPS leads to state-of-the-art results on all the available public benchmarks, and opens a fresh new direction for research on re-identification. Dong Seon Cheng, Marco Cristani, Michele Stoppa, Loris Bazzani, Vittorio Murino |
BMVC | 5 |
| 2011 | Social interaction discovery by statistical analysis of F-formationsabstractWe present a novel approach for detecting social interactions in a crowded scene by employing solely visual cues. The detection of social interactions in unconstrained scenarios is a valuable and important task, especially for surveillance purposes. Our proposal is inspired by the social signaling literature, and in particular it considers the sociological notion of F-formation. An F-formation is a set of possible configurations in space that people may assume while participating in a social interaction. Our system takes as input the positions of the people in a scene and their (head) orientations; then, employing a voting strategy based on the Hough transform, it recognizes F-formations and the individuals associated with them. Experiments on simulations and real data promote our idea. Marco Cristani, Loris Bazzani, Giulia Paggetti, Andrea Fossati, Diego Tosato, Alessio Del Bue, Gloria Menegaz, Vittorio Murino |
BMVC | 8 |
| 2011 | Multimodal Schizophrenia Detection by Multiclassification Analysis
Aydin Ulas, Umberto Castellani, Pasquale Mirtuono, Manuele Bicego, Vittorio Murino, Stefania Cerruti, Marcella Bellani, Manfredo Atzori, Gianluca Rambaldelli, Michele Tansella, Paolo Brambilla |
CIARP | 5 |
| 2011 | Fast FPGA-based architecture for pedestrian detection based on covariance matricesabstractPedestrian detection is a crucial task in several video surveillance and automotive scenarios, but only a few detection systems are designed to be realized on an embedded architecture, allowing to increase the processing speed which is one of the key requirements in real applications. In this paper, we propose a novel SoC (System on Chip) architecture for fast pedestrian detection in video. Our implementation is based on a linear SVM (Support Vector Machine) classification frame- work, learned on a set of overlapped image patches. Each patch is described by a covariance matrix of a set of image features. Exploiting the inner parallelism of the FPGA (Field Programmable Gate Array) boards, we dramatically accelerate the covariance matrices computation that plays a crucial role in the framework. In the experiments, we show the effectiveness and the efficiency of our pedestrian detection system, reaching a detection speed of 132 fps at VGA resolution. Samuele Martelli, Diego Tosato, Marco Cristani, Vittorio Murino |
ICIP | 4 |
| 2011 | Learning attentional policies for tracking and recognition in video with deep networks
Loris Bazzani, Nando de Freitas, Hugo Larochelle, Vittorio Murino, Jo-Anne Ting |
ICML | 4 |
| 2011 | A Method for Asteroids 3D Surface Reconstruction from Close Approach Distances
Luca Baglivo, Alessio Del Bue, Massimo Lunardelli, Francesco Setti, Vittorio Murino, Mariolino De Cecco |
ICVS | 5 |
| 2011 | An Experimental Framework for Evaluating PTZ Tracking Algorithms
Pietro Salvagnini, Marco Cristani, Alessio Del Bue, Vittorio Murino |
ICVS | 4 |
| 2011 | A New Shape Diffusion Descriptor for Brain Classification
Umberto Castellani, Pasquale Mirtuono, Vittorio Murino, Marcella Bellani, Gianluca Rambaldelli, Michele Tansella, Paolo Brambilla |
MICCAI (2) | 3 |
| 2011 | Statistical 3D Shape Analysis by Local Generative DescriptorsabstractIn this paper, we propose a new approach for surface representation. Generative models are exploited for encoding the variations of local geometric properties of 3D shapes. Surfaces are locally modeled as a stochastic process which spans a neighborhood area through a set of circular geodesic pathways, captured by a modified version of a Hidden Markov Model (HMM) named multicircular HMM (MC-HMM). The approach proposed consists of two main phases: 1) local geometric feature collection and 2) MC-HMM parameter estimation. The effectiveness of our proposal is demonstrated by several applicative scenarios, all using well-known benchmark data sets, such as multiple view registration, matching of deformable shapes, and object recognition on cluttered scenes. The results achieved are very promising and open up the use of generative models as geometric descriptors in an extensive range of applications. Umberto Castellani, Marco Cristani, Vittorio Murino |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Generative modeling and classification of dialogs by a low-level turn-taking feature
Marco Cristani, Anna Pesarin, Carlo Drioli, Alessandro Tavano, Alessandro Perina, Vittorio Murino |
Pattern Recognit. | 6 |
| 2010 | Person re-identification by symmetry-driven accumulation of local featuresabstractIn this paper, we present an appearance-based method for person re-identification. It consists in the extraction of features that model three complementary aspects of the human appearance: the overall chromatic content, the spatial arrangement of colors into stable regions, and the presence of recurrent local motifs with high entropy. All this information is derived from different body parts, and weighted opportunely by exploiting symmetry and asymmetry perceptual principles. In this way, robustness against very low resolution, occlusions and pose, viewpoint and illumination changes is achieved. The approach applies to situations where the number of candidates varies continuously, considering single images or bunch of frames for each individual. It has been tested on several public benchmark datasets (ViPER, iLIDS, ETHZ), gaining new state-of-the-art performances. Michela Farenzena, Loris Bazzani, Alessandro Perina, Vittorio Murino, Marco Cristani |
CVPR | 4 |
| 2010 | Object Recognition with Hierarchical Stel Models
Alessandro Perina, Nebojsa Jojic, Umberto Castellani, Marco Cristani, Vittorio Murino |
ECCV (6) | 5 |
| 2010 | Multi-class Classification on Riemannian Manifolds for Video Surveillance
Diego Tosato, Michela Farenzena, Mauro Spera, Vittorio Murino, Marco Cristani |
ECCV (2) | 4 |
| 2010 | Collaborative particle filters for group trackingabstractTracking groups of people is a highly informative task in surveillance, and it represents a still open and little explored issue. In this paper, we propose a brand new framework for group tracking, that consists in two separate particle filters, one focusing on groups as atomic entities (the multi-group tracker), and the other modeling each individual separately (the multi-object tracker). The latter helps the multi-group tracker in better defining the nature of a group, evaluating the membership of each individual with respect to different groups, and allowing a robust management of the occlusions. The coupling of the two processes is theoretically founded due to the revision of the posterior distribution of the multi-group tracker with the statistics accumulated by the multi-object tracker. Experimental comparative results certify the goodness of the proposed technique. Loris Bazzani, Marco Cristani, Vittorio Murino |
ICIP | 3 |
| 2010 | Combining free energy score spaces with information theoretic kernels: Application to scene classificationabstractMost approaches to learn classifiers for structured objects (e.g., images) use generative models in a classical Bayesian framework. However, state-of-the-art classifiers for vectorial data (e.g., support vector machines) are learned discriminatively. A generative embedding is a mapping from the object space into a fixed dimensional score space, induced by a generative model, usually learned from data. The fixed dimensionality of these generative score spaces makes them adequate for discriminative learning of classifiers, thus bringing together the best of the discriminative and generative paradigms. In particular, it was recently shown that this hybrid approach outperforms a classifier obtained directly for the generative model upon which the score space was built. Using a generative embedding involves two steps: (i) defining and learning the generative model and using it to build the embedding; (ii) discriminatively learning a (maybe kernel) classifier on the adopted score space. The literature on generative embeddings is essentially focused on step (i), usually using some standard off-the-shelf tool for step (ii). In this paper, we adopt a different approach, by focusing also on the discriminative learning step. In particular, we combine two very recent and top performing tools in each of the steps: (i) the free energy score space; (ii) non-extensive information theoretic kernels. In this paper, we apply this methodology in scene recognition. Experimental results on two benchmark datasets shows that our approach yields state-of-the-art performance. Manuele Bicego, Alessandro Perina, Vittorio Murino, André F. T. Martins, Pedro M. Q. Aguiar, Mário A. T. Figueiredo |
ICIP | 3 |
| 2010 | Part-based human detection on Riemannian manifoldsabstractIn this paper we propose a novel part-based framework for pedestrian detection. We model a human as a hierarchy of fixed overlapped parts, each of which described by covariances of features. Each part is modeled by a boosted classifier, learnt using Logitboost on Riemannian manifolds. All the classifiers are then linked to form a high-level classifier, through weighted summation, whose weights are estimated during the learning. The final classifier is simple, light and robust. The experimental results show that we outperform the state-of-the-art human detection performances on the INRIA person dataset. Diego Tosato, Michela Farenzena, Marco Cristani, Vittorio Murino |
ICIP | 4 |
| 2010 | Multiple-Shot Person Re-identification by HPE SignatureabstractIn this paper, we propose a novel appearance-based method for person re-identification, that condenses a set of frames of the same individual into a highly informative signature, called Histogram Plus Epitome, HPE. It incorporates complementary global and local statistical descriptions of the human appearance, focusing on the overall chromatic content, via histograms representation, and on the presence of recurrent local patches, via epitome estimation. The matching of HPEs provides optimal performances against low resolution, occlusions, pose and illumination variations, defining novel state-of-the-art results on all the datasets considered. Loris Bazzani, Marco Cristani, Alessandro Perina, Michela Farenzena, Vittorio Murino |
ICPR | 5 |
| 2010 | 2D Shape Recognition Using Information Theoretic KernelsabstractIn this paper, a novel approach for contour based 2D shape recognition is proposed, using a class of information theoretic kernels recently introduced. This kind of kernels, based on a non-extensive generalization of the classical Shannon information theory, are defined on probability measures. In the proposed approach, chain code representations are first extracted from the contours; then n-gram statistics are computed and used as input to the information theoretic kernels. We tested different versions of such kernels, using support vector machine and nearest neighbor classifiers. An experimental evaluation on the Chicken pieces dataset shows that the proposed approach significantly outperforms the current state-of-the-art methods. Manuele Bicego, André F. T. Martins, Vittorio Murino, Pedro M. Q. Aguiar, Mário A. T. Figueiredo |
ICPR | 3 |
| 2010 | Nonlinear Mappings for Generative Kernels on Latent Variable ModelsabstractGenerative kernels have emerged in the last years as an effective method for mixing discriminative and generative approaches. In particular, in this paper, we focus on kernels defined on generative models with latent variables (e.g. the states in a Hidden Markov Model). The basic idea underlying these kernels is to compare objects, via a inner product, in a feature space where the dimensions are related to the latent variables of the model. Here we propose to enhance these kernels via a nonlinear normalization of the space, namely a nonlinear mapping of space dimensions able to exploit their discriminative characteristics. In this paper we investigate three possible nonlinear mappings, for two HMM-based generative kernels, testing them in different sequence classification problems, with really promising results. Anna C. Carli, Manuele Bicego, Sisto Baldo, Vittorio Murino |
ICPR | 4 |
| 2010 | 2LDA: Segmentation for RecognitionabstractFollowing the trend of “segmentation for recognition”, we present 2LDA, a novel generative model to automatically segment an image in 2 segments, background and foreground, while inferring a latent Dirichlet allocation (LDA) topic distribution on both segments. The idea is to merge two separate modules, LDA and the segmentation module, explicitly considering (and exchanging) the uncertainty between them. The resulting model adds spatial relationships to LDA, which in turn helps in using the topics to segment an image. The experimental results show that, unlike LDA, our model can be used to recognize objects, and also outperforms the state of the art algorithms. Alessandro Perina, Marco Cristani, Vittorio Murino |
ICPR | 3 |
| 2010 | A Re-evaluation of Pedestrian Detection on Riemannian ManifoldsabstractBoosting covariance data on Riemannian manifolds has proven to be a convenient strategy in a pedestrian detection context. In this paper we show that the detection performances of the state-of-the-art approach of Tuzel et al. can be greatly improved, from both a computational and a qualitative point of view, by considering practical and theoretical issues, and allowing also the estimation of occlusions in a fine way. The resulting detection system reaches the best performance on the INRIA dataset, setting novel state-of-the art results. Diego Tosato, Michela Farenzena, Marco Cristani, Vittorio Murino |
ICPR | 4 |
| 2010 | Brain Morphometry by Probabilistic Latent Semantic Analysis
Umberto Castellani, Alessandro Perina, Vittorio Murino, Marcella Bellani, Gianluca Rambaldelli, Michele Tansella, Paolo Brambilla |
MICCAI (2) | 3 |
| 2010 | Pervasive video analysis: workshop overviewabstractThis workshop aims at tackling the novel challenging scenarios in pervasive video analysis which require not only to address specific problems (e.g., tracking, recognition) on a single view, but to deal with a set of distributed observations, eventually integrated with subjective mobile video streams. Accepted papers cover a wide range of subjects going from the joint analysis of video sequences, taken from fixed location and mobile cameras, to situation awareness and understanding. Hamid K. Aghajan, Marco Cristani, Vittorio Murino, Nicu Sebe |
ACM Multimedia | 3 |
| 2010 | Toward an automatically generated soundtrack from low-level cross-modal correlations for automotive scenariosabstractIn this paper, we propose a novel recommendation policy for driving scenarios. While driving a car, listening to an audio track may enrich the atmosphere, conveying emotions that let the driver sense a more arousing experience. Here, we are introducing a recommendation policy that, given a video sequence taken by a camera mounted onboard a car, chooses the most suitable audio piece from a predetermined set of melodies. The mixing mechanism takes inspiration from a set of generic qualitative aesthetical rules for cross-modal linking, realized by associating audio and video features. The contribution of this paper is to translate such qualitative rules into quantitative terms, learning from an extensive training dataset cross-modal statistical correlations, and validating them in a thoroughly way. In this way, we are able to define what are the audio and video features that correlate at best (i.e., promoting or rejecting some aesthetical rules), and what are their correlation intensities. This knowledge is then employed for the realization of the recommendation policy. A set of user studies illustrate and validate the policy, thus encouraging further developments toward a real implementation in an automotive application. Marco Cristani, Anna Pesarin, Carlo Drioli, Vittorio Murino, Antonio Rodà, Michele Grapulin, Nicu Sebe |
ACM Multimedia | 4 |
| 2010 | Structural epitome: a way to summarize one's visual experienceabstractIn order to study the properties of total visual input in humans, a single subject wore a camera for two weeks capturing, on average, an image every 20 seconds (www.research.microsoft.com/~jojic/aihs). The resulting new dataset contains a mix of indoor and outdoor scenes as well as numerous foreground objects. Our first analysis goal is to create a visual summary of the subject’s two weeks of life using unsupervised algorithms that would automatically discover recurrent scenes, familiar faces or common actions. Direct application of existing algorithms, such as panoramic stitching (e.g. Photosynth) or appearance-based clustering models (e.g. the epitome), is impractical due to either the large dataset size or the dramatic variation in the lighting conditions. As a remedy to these problems, we introduce a novel image representation, the “stel epitome,” and an associated efficient learning algorithm. In our model, each image or image patch is characterized by a hidden mapping T, which, as in previous epitome models, defines a mapping between the image-coordinates and the coordinates in the large all-I-have-seen" epitome matrix. The limited epitome real-estate forces the mappings of different images to overlap, with this overlap indicating image similarity. However, in our model the image similarity does not depend on direct pixel-to-pixel intensity/color/feature comparisons as in previous epitome models, but on spatial configuration of scene or object parts, as the model is based on the palette-invariant stel models. As a result, stel epitomes capture structure that is invariant to non-structural changes, such as illumination, that tend to uniformly affect pixels belonging to a single scene or object part." Nebojsa Jojic, Alessandro Perina, Vittorio Murino |
NIPS | 3 |
| 2010 | A real-time versatile roadway path extraction and tracking on an FPGA platform
Roberto Marzotto, Paul Zoratti, Daniele Bagni, Andrea Colombari, Vittorio Murino |
Comput. Vis. Image Underst. | 5 |
| 2010 | Learning natural scene categories by selective multi-scale feature extraction
Alessandro Perina, Marco Cristani, Vittorio Murino |
Image Vis. Comput. | 3 |
| 2009 | Learning Approach to Analyze Tumour Heterogeneity in DCE-MRI Data During Anti-cancer Treatment
Alessandro Daducci, Umberto Castellani, Marco Cristani, Paolo Farace, Pasquina Marzola, Andrea Sbarbati, Vittorio Murino |
AIME | 7 |
| 2009 | Stel component analysis: Modeling spatial correlations in image class structureabstractAs a useful concept in the study of the low level image class structure, we introduce the notion of a structure element - `stel.' The notion is related to the notions of a pixel, superpixel, segment or a part, but instead of referring to an element or a region of a single image, stel is a probabilistic element of an entire image class. Stels often define clear object or scene parts as a consequence of the modeling constraint which forces the regions belonging to a single stel to have a tight distribution over local measurements, such as color or texture. This self-similarity within a region in a single image is typical of most meaningful image parts, even when in different images of similar objects the corresponding parts may not have similar local measurements. The stel itself is expected to be consistent within a class, yet flexible, which we accomplish using a novel approach we dubbed stel component analysis. Experimental results show how stel component analysis can assist in image/video segmentation and object recognition where, in particular, it can be used as an alternative of, or in conjunction with, bag-of-features and related classifiers, where stel inference provides a meaningful spatial partition of features. Nebojsa Jojic, Alessandro Perina, Marco Cristani, Vittorio Murino, Brendan J. Frey |
CVPR | 4 |
| 2009 | A hybrid generative/discriminative classification framework based on free-energy termsabstractHybrid, generative-discriminative, techniques have proven to be valuable approaches in tackling difficult object or scene recognition problems. In general, a generative model over the available data for each image class is first learned providing a relatively comprehensive statistical multi-level representation. In this way, new meaningful image features become available, which encode the degree of fitness of the data with respect to the model at different representation levels. Such features are then fed into a discriminative classifier which can exploit the intrinsic data separability. In this paper, we propose the use of variational free energy terms as feature vectors, so that the degree of fitness of the data and the uncertainty over the generative process are explicitly included in the data description. The proposed method is automatically superior to a pure generative classification, and we also experimentally validate it on a wide selection of generative models applied to challenging benchmarks in hard computer vision tasks such as scene, object, and shape recognition. In several instances, the proposed approach outperforms the current state-of-the-art techniques as for classification results, while also showing to be computationally inexpensive. Alessandro Perina, Marco Cristani, Umberto Castellani, Vittorio Murino, Nebojsa Jojic |
ICCV | 4 |
| 2009 | Online subjective feature selection for occlusion management in tracking applicationsabstractMost of the state-of-the-art tracking algorithms are prone to error when dealing with occlusions, especially when the involved moving objects are hardly discernible in appearance. In this paper, we propose a multi-object particle filtering tracking framework particularly suited to manage the occlusion problem. The presented solution consists in the introduction of a online subjective feature selection mechanism, which highlights and employs the most discriminant features characterizing a single object with respect to the neighbouring objects. The policy adopted fits formally in the observation step of the particle filtering process, it is effective and not computationally costly. Trials carried out on illustrative synthetic data and on recent challenging benchmark sequences report compelling performances and encourage further development of the technique. Loris Bazzani, Marco Cristani, Manuele Bicego, Vittorio Murino |
ICIP | 4 |
| 2009 | Free energy score spaceabstractScore functions induced by generative models extract fixed-dimension feature vectors from different-length data observations by subsuming the process of data generation, projecting them in highly informative spaces called score spaces. In this way, standard discriminative classifiers are proved to achieve higher performances than a solely generative or discriminative approach. In this paper, we present a novel score space that exploits the free energy associated to a generative model through a score function. This function aims at capturing both the uncertainty of the model learning and ``local compliance of data observations with respect to the generative process. Theoretical justifications and convincing comparative classification results on various generative models prove the goodness of the proposed strategy. Alessandro Perina, Marco Cristani, Umberto Castellani, Vittorio Murino, Nebojsa Jojic |
NIPS | 4 |
| 2009 | Fully non-homogeneous hidden Markov model double net: A generative model for haplotype reconstruction and block discovery
Alessandro Perina, Marco Cristani, Luciano Xumerle, Vittorio Murino, Pier Franco Pignatti, Giovanni Malerba |
Artif. Intell. Medicine | 4 |
| 2008 | Unsupervised Learning of Saliency Concepts for Natural Image Classification and Retrieval
Alessandro Perina, Marco Cristani, Vittorio Murino |
CIARP | 3 |
| 2008 | Geo-located image analysis using latent representationsabstractImage categorization is undoubtedly one of the most challenging open problems faced in computer vision, far from being solved by employing pure visual cues. Recently, additional textual ldquotagsrdquo can be associated to images, enriching their semantic interpretation beyond the pure visual aspect, and helping to bridge the so-called semantic gap. One of the latest class of tags consists in geo-location data, containing information about the geographical site where an image has been captured. Such data motivate, if not require, novel strategies to categorize images, and pose new problems to focus on. In this paper, we present a statistical method for geo-located image categorization, in which categories are formed by clustering geographically proximal images with similar visual appearance. The proposed strategy permits also to deal with the geo-recognition problem, i.e., to infer the geographical area depicted by images with no available location information. The method lies in the wide literature on statistical latent representations, in particular, the probabilistic latent semantic analysis (pLSA) paradigm has been extended, introducing a latent aspect which characterizes peculiar visual features of different geographical zones. Experiments on categorization and georecognition have been carried out employing a well-known geographical image repository: results are actually very promising, opening new interesting challenges and applications in this research field. Marco Cristani, Alessandro Perina, Umberto Castellani, Vittorio Murino |
CVPR | 4 |
| 2008 | A statistical signature for automatic dialogue classificationabstractIn the last few years, there has been a certain attention to the problem of human-human communication, trying to devise artificial systems able to mediate a conversational setting between two or more people. In this paper, we designed an automatic system based on a generative structure able to classify hard dialog acts. The generative model is composed by integrating a hierarchical Gaussian mixture model and the Influence Model, originating a brand new method able to deal with such difficult scenarios. The method has been tested on a set of conversational settings involving dialogues between adults and children and adults, in flat and arguing discussions, proving very accurate classification results. Anna Pesarin, Marco Cristani, Vittorio Murino, Carlo Drioli, Alessandro Perina, Alessandro Tavano |
ICPR | 3 |
| 2008 | Geo-located Image Grouping Using Latent Descriptions
Marco Cristani, Alessandro Perina, Vittorio Murino |
ICVS | 3 |
| 2008 | Visual MRI: Merging information visualization and non-parametric clustering techniques for MRI dataset analysis
Umberto Castellani, Marco Cristani, Carlo Combi, Vittorio Murino, Andrea Sbarbati, Pasquina Marzola |
Artif. Intell. Medicine | 4 |
| 2008 | Sparse points matching by combining 3D mesh saliency with statistical descriptorsabstractAbstract This paper proposes new methodology for the detection and matching of salient points over several views of an object. The process is composed by three main phases. In the first step, detection is carried out by adopting a new perceptually‐inspired 3D saliency measure. Such measure allows the detection of few sparse salient points that characterize distinctive portions of the surface. In the second step, a statistical learning approach is considered to describe salient points across different views. Each salient point is modelled by a Hidden Markov Model (HMM), which is trained in an unsupervised way by using contextual 3D neighborhood information, thus providing a robust and invariant point signature. Finally, in the third step, matching among points of different views is performed by evaluating a pairwise similarity measure among HMMs. An extensive and comparative experimental session has been carried out, considering real objects acquired by a 3D scanner from different points of view, where objects come from standard 3D databases. Results are promising, as the detection of salient points is reliable, and the matching is robust and accurate. Umberto Castellani, Marco Cristani, Simone Fantoni, Vittorio Murino |
Comput. Graph. Forum | 4 |
| 2007 | Automatic selection of MRF control parameters by reactive tabu search
Umberto Castellani, Andrea Fusiello, Riccardo Gherardi, Vittorio Murino |
Image Vis. Comput. | 4 |
| 2007 | Segmentation and tracking of multiple video objects
Andrea Colombari, Andrea Fusiello, Vittorio Murino |
Pattern Recognit. | 3 |
| 2007 | Audio-Visual Event Recognition in Surveillance Video SequencesabstractIn the context of the automated surveillance field, automatic scene analysis and understanding systems typically consider only visual information, whereas other modalities, such as audio, are typically disregarded. This paper presents a new method able to integrate audio and visual information for scene analysis in a typical surveillance scenario, using only one camera and one monaural microphone. Visual information is analyzed by a standard visual background/foreground (BG/FG) modelling module, enhanced with a novelty detection stage and coupled with an audio BG/FG modelling scheme. These processes permit one to detect separate audio and visual patterns representing unusual unimodal events in a scene. The integration of audio and visual data is subsequently performed by exploiting the concept of synchrony between such events. The audio-visual (AV) association is carried out online and without need for training sequences, and is actually based on the computation of a characteristic feature called audio-video concurrence matrix, allowing one to detect and segment AV events, as well as to discriminate between them. Experimental tests involving classification and clustering of events show all the potentialities of the proposed approach, also in comparison with the results obtained by employing the single modalities and without considering the synchrony issue Marco Cristani, Manuele Bicego, Vittorio Murino |
IEEE Trans. Multim. | 3 |
| 2006 | Acoustic Range Image Segmentation by Effective Mean ShiftabstractImage perception in underwater environment is a difficult task for a human operator, and data segmentation becomes a crucial step toward an higher level interpretation and recognition of the observing scenarios. This paper contributes to the related state of the art, by fitting the mean shift clustering paradigm to the segmentation of acoustical range images, providing a segmentation approach in which whatever parameter tuning is absent. Moreover, the method exploits actively the connectivity information provided by the range map, by using reverse projection as acceleration technique. Therefore, the method is able to produce, starting from raw range data, meaningful segmented clouds of points in a fully automatic and efficient fashion. Umberto Castellani, Marco Cristani, Vittorio Murino |
ICIP | 3 |
| 2006 | Clustering Under Prior Knowledge with Application to Image SegmentationabstractThis paper proposes a new approach to model-based clustering under prior knowl- edge. The proposed formulation can be interpreted from two different angles: as penalized logistic regression, where the class labels are only indirectly observed (via the probability density of each class); as finite mixture learning under a group- ing prior. To estimate the parameters of the proposed model, we derive a (gener- alized) EM algorithm with a closed-form E-step, in contrast with other recent approaches to semi-supervised probabilistic clustering which require Gibbs sam- pling or suboptimal shortcuts. We show that our approach is ideally suited for image segmentation: it avoids the combinatorial nature Markov random field pri- ors, and opens the door to more sophisticated spatial priors (e.g., wavelet-based) in a simple and computationally efficient way. Finally, we extend our formulation to work in unsupervised, semi-supervised, or discriminative modes. Mário A. T. Figueiredo, Dong Seon Cheng, Vittorio Murino |
NIPS | 3 |
| 2006 | Unsupervised scene analysis: A hidden Markov model approach
Manuele Bicego, Marco Cristani, Vittorio Murino |
Comput. Vis. Image Underst. | 3 |
| 2006 | Similarity-based pattern recognition
Manuele Bicego, Vittorio Murino, Marcello Pelillo, Andrea Torsello |
Pattern Recognit. | 2 |
| 2005 | Towards Information Visualization and Clustering Techniques for MRI Data Sets
Umberto Castellani, Carlo Combi, Pasquina Marzola, Vittorio Murino, Andrea Sbarbati, Marco Zampieri |
AIME | 4 |
| 2005 | Uncalibrated interpolation of rigid displacements for view synthesisabstractIn this paper we present a method for novel view synthesis from two uncalibrated reference views. Snapshots of a scene are created as if they were taken from a different "virtual" viewpoint. The relative affine structure is used to describe the geometry of the scene and then to extrapolate and interpolate novel views. The contribution of this paper is an automatic method for specifying the virtual viewpoint in an uncalibrated setting, based on the interpolation and extrapolation of the epipolar geometry linking the reference views. Experimental results using synthetic and real images are shown. Andrea Colombari, Andrea Fusiello, Vittorio Murino |
ICIP (1) | 3 |
| 2005 | A supervised data-driven approach for microarray spot quality classification
Manuele Bicego, Maria Rosario Martinez, Vittorio Murino |
Pattern Anal. Appl. | 3 |
| 2005 | A Hidden Markov Model approach for appearance-based 3D object recognition
Manuele Bicego, Umberto Castellani, Vittorio Murino |
Pattern Recognit. Lett. | 3 |
| 2005 | A complete system for on-line 3D modelling from acoustic images
Umberto Castellani, Andrea Fusiello, Vittorio Murino, Laura Papaleo, Enrico Puppo, Massimiliano Pittore |
Signal Process. Image Commun. | 3 |
| 2004 | High Resolution Video Mosaicing with Global Alignment
Roberto Marzotto, Andrea Fusiello, Vittorio Murino |
CVPR (1) | 3 |
| 2004 | Audio-Video Integration for Background Modelling
Marco Cristani, Manuele Bicego, Vittorio Murino |
ECCV (2) | 3 |
| 2004 | Investigating Hidden Markov Models' Capabilities in 2D Shape ClassificationabstractIn this paper, Hidden Markov Models (HMMs) are investigated for the purpose of classifying planar shapes represented by their curvature coefficients. In the training phase, special attention is devoted to the initialization and model selection issues, which make the learning phase particularly effective. The results of tests on different data sets show that the proposed system is able to accurately classify objects that were translated, rotated, occluded, or deformed by shearing, also in the presence of noise. Manuele Bicego, Vittorio Murino |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Similarity-based classification of sequences using hidden Markov models
Manuele Bicego, Vittorio Murino, Mário A. T. Figueiredo |
Pattern Recognit. | 2 |
| 2004 | Augmented Scene Modeling and Visualization by Optical and Acoustic Sensor IntegrationabstractIn this paper, underwater scene modeling from multisensor data is addressed. Acoustic and optical devices aboard an underwater vehicle are used to sense the environment in order to produce an output that is readily understandable even by an inexperienced operator. The main idea is to integrate multiple-sensor data by geometrically registering such data to a model. The geometrical structure of this model is a priori known but not ad hoc designed for this purpose. As a result, the vehicle pose is derived and model objects can be superimposed upon actual images, thus generating an augmented-reality representation. Results on a real underwater scene are reported, showing the effectiveness of the proposed approach. Andrea Fusiello, Vittorio Murino |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2003 | Automatic road extraction from aerial images by probabilistic contour trackingabstractIn this paper a new automatic approach to road extraction from aerial images is proposed. This method improves a recently introduced promising approach to probabilistic contour tracking, originally semi-automatic, by adding a fully automatic initialization strategy and a merging methodology, able to combine the different obtained results. The initialization strategy is based on the Hough transform and on some topological considerations, and the merging step is based on a new introduced quality measure, based on color and gradient information. Experimental results on real highly complex images show that the proposed approach is a promising and fully automatic method for extracting roads from images, even in presence of highly urbanized areas, occlusions or shadows. Manuele Bicego, Silvio Dalfini, Gianni Vernazza, Vittorio Murino |
ICIP (3) | 4 |
| 2003 | Mosaic of a video shot with multiple moving objectsabstractIn this paper we describe an application which takes a video shot as input and produces a compact representation composed by a background layer and segmented moving objects. We deal with the problems of global registration, super-resolution mosaicing, objects segmentation and tracking. Global registration is achieved with a graph-based technique that exploits situations when the camera returns to a previously seen area. Objects segmentation is based on motion analysis using a robust statistical model of the background. Tracking is based on blob matching using singular value decomposition. Andrea Fusiello, Michele Aprile, Roberto Marzotto, Vittorio Murino |
ICIP (2) | 4 |
| 2003 | A sequential pruning strategy for the selection of the number of states in hidden Markov models
Manuele Bicego, Vittorio Murino, Mário A. T. Figueiredo |
Pattern Recognit. Lett. | 2 |
| 2003 | Special issue on 3-D image analysis and modeling
Hongbin Zha, Hideo Saito 0001, Vittorio Murino, Andrea Fusiello |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2002 | Model Acquisition by Registration of Multiple Acoustic Range Views
Andrea Fusiello, Umberto Castellani, Lucca Ronchetti, Vittorio Murino |
ECCV (2) | 4 |
| 2002 | Registration of very time-distant aerial imagesabstractWe address the alignment of historical and present-day aerial photographs. Historical images refer to regions bombed during the Second World War. In these regions, the risk of unexploded bombs is still high, especially where the bombing was more frequent. Alignment is required to fill in an unexploded bombs risk map. The task is challenging because many features in the historical images have changed or are missing (and vice versa). Moreover, in the historical images, bomb craters introduce large gray level variations so that it is difficult to extract features automatically. This work propose a semi-automatic application for image alignment in order to improve accuracy and to speed up the alignment process. Vittorio Murino, Umberto Castellani, Alberto Etrari, Andrea Fusiello |
ICIP (3) | 1 |
| 2002 | A Multimodal Electronic Travel Aid DeviceabstractThis paper describes an electronic travel aid device, that may enable blind individuals to "see the world with their ears". A wearable prototype will be assembled using low-cost hardware: earphones, sunglasses fitted with two micro cameras, and a palmtop computer. The system, which currently runs on a desktop computer, is able to detect the light spot produced by a laser pointer, compute its angular position and depth, and generate a corresponding sound providing auditory cues for perception of the position and distance of the pointed surface patch. It permits different sonification modes that can be chosen by drawing, with the laser pointer, a predefined stroke which will be recognized by a hidden Markov model. In this way a blind person can use a common pointer as a replacement for the cane and will interact with the device using a flexible and natural sketch based interface. Andrea Fusiello, Antonello Panuccio, Vittorio Murino, Federico Fontana, Davide Rocchesso |
ICMI | 3 |
| 2002 | A Cross-Modal Electronic Travel Aid Device
Federico Fontana, Andrea Fusiello, Michele Gobbi, Vittorio Murino, Davide Rocchesso, Luca Sartor, Antonello Panuccio |
Mobile HCI | 4 |
| 2002 | Registration of Multiple Acoustic Range Views for Underwater Scene Reconstruction
Umberto Castellani, Andrea Fusiello, Vittorio Murino |
Comput. Vis. Image Underst. | 3 |
| 2001 | Disparity map restoration by integration of confidence in Markov random fields modelsabstractThis paper proposes some Markov random field (MRF) models for the restoration of stereo disparity maps. The main aspect is the use of confidence maps provided by the symmetric multiple windows (SMW) stereo algorithm to guide the restoration process. The SMW algorithm is an adaptive, multiple-window scheme using left-right consistency to compute disparity and its associated confidence in the presence of occlusions. The MRF approach allows the combining in a single functional of all the available information: observed data with its confidence, noise, and a-priori hypotheses. Optimal estimates of the disparity are obtained by minimizing an energy functional using simulated annealing. Results with a real stereo pair show the improvement obtained by restoration using the MRF approach integrating confidence data. Andrea Fusiello, Umberto Castellani, Vittorio Murino |
ICIP (2) | 3 |
| 2001 | Artificial Neural Networks for Image Analysis and Computer Vision
Vittorio Murino, Gianni Vernazza |
Image Vis. Comput. | 1 |
| 2001 | Reconstruction and segmentation of underwater acoustic images combining confidence information in MRF models
Vittorio Murino |
Pattern Recognit. | 1 |
| 2000 | 3D Mosaicing for Environment ReconstructionabstractThis paper proposes a technique for the 3D reconstruction of an underwater environment from multiple range views. The final target of the work lies in improving the understanding of a human operator guiding an underwater remotely operated vehicle (ROV) equipped with an acoustic camera, which provides a sequence of 3D images in real time. Since the field of view is narrow we devise a technique for the reconstruction of relevant information of the image sequence up to building a mosaic of the surrounding scene. Due to the very noisy nature of the data and the low range resolution, smoothing, segmentation, registration, and fusion problems have been tackled. Examples on real images are presented to show the promising performances of the algorithm. Vittorio Murino, Andrea Fusiello, Nicola Iuretigh, Enrico Puppo |
ICPR | 1 |
| 2000 | Underwater Computer Vision and Pattern Recognition
Vittorio Murino, Andrea Trucco |
Comput. Vis. Image Underst. | 1 |
| 2000 | Three-dimensional image generation and processing in underwater acoustic visionabstractUnderwater exploration is becoming more and more important for many applications involving physical, biological, geological, archaeological, and industrial issues. This paper aims at surveying the up-to-date advances in acoustic acquisition systems and data processing techniques, especially focusing on three-dimensional (3-D) short-range imaging for scene reconstruction and understanding. In fact, the advent of smarter and more efficient imaging systems has allowed the generation of good quality high-resolution images and the related design of proper techniques for underwater scene understanding. The term acoustic vision is introduced to generally describe all data processing (especially image processing) methods devoted to the interpretation of a scene. Since acoustics is also used for medical applications, a short overview of the related systems for biomedical acoustic image for motion is provided. The final goal of the paper is to establish the state of-the art of the techniques and algorithms for acoustic image generation and processing, providing technical details and results for the most promising techniques, and pointing out the potential capabilities of this technology for underwater scene understanding. Vittorio Murino, Andrea Trucco |
Proc. IEEE | 1 |
| 1999 | Three-Dimensional Skeleton Extraction by Point Set ContractionabstractIn this paper, a skeleton extraction method for unstructured three-dimensional data is presented. The algorithm is based on a point set contraction procedure and is proved to be robust to noise and coarse resolution of original data. Preliminary experiments performed on both synthetic and real data are provided, showing the goodness of the proposed method. Riccardo Giannitrapani, Vittorio Murino |
ICIP (1) | 2 |
| 1999 | A confidence-based approach to enhancing underwater acoustic image formationabstractThis paper describes a flexible technique to enhance the formation of short-range acoustic images so as to improve image quality and facilitate the tasks of subsequent postprocessing methods. The proposed methodology operates as an ideal interface between the signals formed by a focused beamforming technique (i.e., the beam signals) and the related image, whether a two-dimensional (2-D) or three-dimensional (3-D) one. To this end, a reliability measure has been introduced, called confidence, which allows one to perform a rapid examination of the beam signals and is aimed at accurately detecting echoes backscattered from a scene. The confidence-based approach exploits the physics of the process of image formation and generic a priori knowledge of a scene to synthesize model-based signals to be compared with actual backscattered echoes, giving, at the same time, a measure of the reliability of their similarity. The objectives that can be attained by this method can be summarized in a reduction in artifacts due to the lowering of the side-lobe level, a better lateral resolution, a greater accuracy in range determination, a direct estimation of the reliability of the information acquired, thus leading to a higher image quality and hence a better scene understanding. Tests on both simulated and actual data (concerning both 2-D and 3-D images) show the higher efficiency of the proposed confidence-based approach, as compared with more traditional techniques. Vittorio Murino, Andrea Trucco |
IEEE Trans. Image Process. | 1 |
| 1998 | Edge/Region-Based Segmentation and Reconstruction of Underwater Acoustic Images by Markov Random FieldsabstractThis paper describes a technique for the reconstruction and segmentation of three-dimensional acoustical images using a coupled Random Fields able to actively integrate confidence information associated with acquired data. Beamforming, a method widely applied in acoustic imaging, is used to build a three-dimensional image, associated point by point with another kind of information representing the reliability (i.e. "confidence") of such an image. Unfortunately, this kind of images is plagued by several problems due to the nature of the signal and to the related sensing system, thus heavily affecting data quality. Specifically, speckle noise and the broad directivity characteristic of the sensor lead to very degraded images. In the proposed algorithm, range and confidence images are modelled as Markov Random Fields whose associated probability distributions are specified by a single energy functional. A three-fold process has been applied able to reconstruct, segment, and restore the involved acoustic images exploiting both types of data. Our approach showed better performances with respect to other MRF-based methods as well as classical methods disregarding reliability information. Optimal (in the Maximum A-Posteriori probability sense) estimates of the 3D and confidence images are obtained by minimizing the energy functional by using simulated annealing. Vittorio Murino, Andrea Trucco |
CVPR | 1 |
| 1998 | A Probabilistic Approach to the Coupled Reconstruction and Restoration of Underwater Acoustic ImagesabstractDescribes a probabilistic technique for the coupled reconstruction and restoration of underwater acoustic images. The technique is founded on the physics of the image-formation process. Beamforming, a method widely applied in acoustic imaging, is used to build a range image from backscattered echoes, associated point by point with another type of information representing the reliability (or confidence) of such an image. Unfortunately, this kind of images is plagued by problems due to the nature of the signal and to the related sensing system. In the proposed algorithm, the range and confidence images are modeled as Markov random fields whose associated probability distributions are specified by a single energy function. This function has been designed to fully embed the physics of the acoustic image-formation process by modeling a priori knowledge of the acoustic system, the considered scene, and the noise-affecting measures and also by integrating reliability information to allow the coupled and simultaneous reconstruction and restoration of both images. Optimal (in the maximum a posteriori probability sense) estimates of the reconstructed range image map and the restored confidence image are obtained by minimizing the energy function using simulated annealing. Experimental results show the improvement of the processed images over those obtained by other methods performing separate reconstruction and restoration processes that disregard reliability information. Vittorio Murino, Andrea Trucco, Carlo S. Regazzoni |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | Noisy texture classification: A higher-order statistics approach
Vittorio Murino, Cinthya Ottonello, Sergio Pagnan |
Pattern Recognit. | 1 |
| 1998 | Structured neural networks for pattern recognitionabstractThis paper proposes a novel approach for the design of structures of neural networks for pattern recognition. The basic idea lies in subdividing the whole classification problem in smaller and simpler problems at different levels, each managed by appropriate components of a complex neural architecture. Three neural structures are presented and applied in a surveillance system aimed at monitoring a railway waiting room classifying potential dangerous situations. Each architecture is composed by nodes, which are actual multilayer perceptrons trained to discriminate between subsets of classes until a complete separation among the classes is achieved. This approach showed better performances with respect to a classical statistical classification procedures and to a single neural network. Vittorio Murino |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1997 | Object Pose Estimation in Underwater Acoustic ImagesabstractWe address the problem of the recognition of man-made objects and the estimation of the related orientation in 2D acoustic images acquired with a forward looking sonar or an acoustic camera. A voting-based approach is described that is able to recognize objects and to estimate their two-dimensional pose by using information coming from boundary segments and their angular relations. The method is directly applied to the edge discontinuities of underwater acoustic images, whose quality is usually affected by some undesired effects such as object blurring, speckle noise, and geometrical distortions degrading the edge detection. The voting approach is robust with respect to these effects, so that good results are obtained even with images of poor quality. The sequences of simulated and real acoustic images are presented in order to test the validity of the proposed method in terms of the average estimation error and computational load. Vittorio Murino, Gian Luca Foresti, Andrea Trucco |
ICIP (1) | 1 |
| 1997 | A Belief-Based Approach for Adaptive Image ProcessingabstractThis paper proposes a new approach to the problem of intelligently regulating image-processing parameters of a distributed network. The proposed approach is based on two-step probabilistic process: (a) belief updating, which consists in computing a functional cost at each node of the network and, (b) belief maximization, which depends on maximizing this functional cost by using a stochastic optimization algorithm. The architecture of an image processing system, consisting of three modules connected in a chain-like structure, is presented as an example showing the capabilities of the proposed approach. Each module is provided with a priori information about the set of parameters that manage a particular data transformation, and with evaluation criteria to judge data quality and to decide on the parameters to be adjusted. Experimental results obtained by using a digitally controlled camera and lens objective, are presented to show the validity of the proposed approach. Vittorio Murino, Gian Luca Foresti, Carlo S. Regazzoni |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1997 | 2D into 3D Hough-space mapping for planar object pose estimation
Vittorio Murino, Gian Luca Foresti |
Image Vis. Comput. | 1 |
| 1997 | Beam pattern formulation and analysis for wide-band beamforming systems using sparse arrays
Vittorio Murino, Andrea Trucco, Alessandra Tesei |
Signal Process. | 1 |
| 1996 | Spurious effects reduction by the reconstruction of acoustic images from bispectrumabstractA method to improve the quality of acoustic images obtained by systems based on beamforming is described. The specific goal is to eliminate two additive spurious terms present in the beam signals. The first effect is due to side lobes and the second one is due to the nonlinear effects introduced by the modulus extractor present at the output of the beamformer. The relative importance of these two terms is assessed, showing that the term due to modulus extraction has a strong negative incidence on the final image quality. The Gaussian distribution of these spurious terms is discussed and verified in the case of punctiform scenes. Owing to the Gaussianity of the spurious terms, a method to eliminate them from the beam signals, based on higher order spectra analysis, is proposed, resulting in an evident image quality and reliability improvement. Andrea Trucco, Cinthya Ottonello, Vittorio Murino |
ICIP (3) | 3 |
| 1996 | Grouping as a Searching Process for Minimum-Energy Configurations of Labelled Random Fields
Vittorio Murino, Carlo S. Regazzoni, Gian Luca Foresti |
Comput. Vis. Image Underst. | 1 |
| 1996 | A distributed probabilistic system for adaptive regulation of image processing parametersabstractA distributed optimization framework and its application to the regulation of the behavior of a network of interacting image processing algorithms are presented. The algorithm parameters used to regulate information extraction are explicitly represented as state variables associated with all network nodes. Nodes are also provided with message-passing procedures to represent dependences between parameter settings at adjacent levels. The regulation problem is defined as a joint-probability maximization of a conditional probabilistic measure evaluated over the space of possible configurations of the whole set of state variables (i.e., parameters). The global optimization problem is partitioned and solved in a distributed way, by considering local probabilistic measures for selecting and estimating the parameters related to specific algorithms used within the network. The problem representation allows a spatially varying tuning of parameters, depending on the different informative contents of the subareas of an image. An application of the proposed approach to an image processing problem is described. The processing chain chosen as an example consists of four modules. The first three algorithms correspond to network nodes. The topmost node is devoted to integrating information derived from applying different parameter settings to the algorithms of the chain. The nodes associated with data-transformation processes to be regulated are represented by an optical sensor and two filtering units (for edge-preserving and edge-extracting filterings), and a straight-segment detection module is used as an integration site. Vittorio Murino, Gian Luca Foresti, Carlo S. Regazzoni |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1995 | Simulated annealing approach for the design of unequally spaced arraysabstractA synthesis method aimed at designing an array antenna is proposed. Simulated annealing (SA), which is a probabilistic methodology used to solve combinatorial optimization problems, has been utilized to optimize the position and the weighting coefficients of the array elements in order to improve the antenna performance. The sensor position and related weighting coefficients are considered as parameters to be tuned in order to constrain the directivity function (i.e., the beam power pattern) of an antenna to satisfy specific requirements. Conventional beamforming is utilized to compute the beam power pattern having the desired properties, such as a narrow width of the main lobe, and side lobe amplitudes under a certain threshold, etc., taking into account the need to reduce to number of sensors with a small spatial aperture. Several results are presented showing a notable improvement in the antenna performance utilizing the SA approach with respect to those considered in the literature. Vittorio Murino |
ICASSP | 1 |
| 1995 | A Multilevel Fusion Approach to Object Identification in Outdoor Road ScenesabstractThe task of object identification is fundamental to the operations of an autonomous vehicle. It can be accomplished by using techniques based on a Multisensor Fusion framework, which allows the integration of data coming from different sensors. In this paper, an approach to the synergic interpretation of data provided by thermal and visual sensors is proposed. Such integration is justified by the necessity for solving the ambiguities that may arise from separate data interpretations. The architecture of a distributed Knowledge-Based system is described. It performs an Intelligent Data Fusion process by integrating, in an opportunistic way, data acquired with a thermal and a video (b/w) camera. Data integration is performed at various architecture levels in order to increase the robustness of the whole recognition process. A priori models allow the system to obtain interesting data from both sensors; to transform such data into intermediate symbolic objects; and, finally, to recognize environmental situations on which to perform further processing. Some results are reported for different environmental conditions (i.e. a road scene by day and by night, with and without the presence of obstacles). Vittorio Murino, Carlo S. Regazzoni, Gian Luca Foresti, Gianni Vernazza |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1994 | Underwater 3D Imaging by FFT Dynamic Focusing BeamformingabstractDescribes an application of the dynamic focusing technique for underwater 3D acoustic imaging. In particular, the integration of dynamic focalization with an FFT implementation of phase-shift beamforming has been developed in order to avoid an increase in computational complexity. The resulting system, which uses narrow-band signals, permits the formation of 3D images at a high frame rate and without the limitations imposed by the depth of field on fixed-focused beamforming. Some simulations made it possible to test the proposed system, and the related results are reported.> Vittorio Murino, Andrea Trucco |
ICIP (1) | 1 |
| 1993 | Distributed spatial reasoning for multisensory image interpretation
Gian Luca Foresti, Vittorio Murino, Carlo S. Regazzoni, Gianni Vernazza |
Signal Process. | 2 |
| 1992 | Distributed Belief Revision for Adaptive Image Processing Regulation
Vittorio Murino, Massimiliano F. Peri, Carlo S. Regazzoni |
ECCV | 1 |
| 1992 | Multilevel GMRF-based segmentation of image sequencesabstractA probabilistic method for obtaining a complete image representation on the basis of spatial-temporal knowledge is presented. The main goal of the algorithm is to obtain a consistent segmentation of a noisy image sequence. Consistent means that the same region must maintain the same label in all consequent images of the sequence where it appears. To this end, a processing scheme is presented which extends Bayesian networks of Gibbs-Markov random fields (GMRF) to segmentation of dynamic scenes.> Carlo S. Regazzoni, Vittorio Murino |
ICPR (2) | 2 |
| 1991 | A numerical and symbolic fusion method for interpretation of image sequenceabstractThe problem of analyzing a sequence of images by taking into account the symbolic and numerical content of the signal is considered. An algorithm for segmentation and tracking of regions among images of a sequence is presented. The method is based on a Gibbs Markov random field (GMRF) model which couples the image process to a spatial-temporal region process. Optical flow field is used to adaptively decide the temporal clique to be used during annealing of energy. Displacement vectors of pixels belonging to recognized regions are predicted by using prior knowledge about object behavior.> Fabio Arduini, R. Cabri, Gian Luca Foresti, Vittorio Murino, Carlo S. Regazzoni |
ICASSP | 4 |
| 1990 | Three-dimensional reconstruction and lateral views in optical microscopyabstractThe three-dimensional properties of an optical microscope are analyzed and a defocusing technique is proposed to recover the spatial distribution of the specimens under investigation. Limitations and real capabilities of the 3D reconstruction are pointed out. As a result, lateral views are obtained by means of operations in the spatial frequency domain. In such a way, it is possible to represent side views of an object within the angular aperture range of the microscope. A theory concerning image formation is discussed and simulations of side view reconstructions are reported. Tullio Tommasi, Bruno Bianco, Vittorio Murino, Alessandra Oneto, Alberto Diaspro |
VCIP | 3 |