EDBT 2026 Demo / reviewers in the wild / expert
Andrea Cavallaro
dblp:36/368
· DBLP profile ↗
211ranked-venue papers
14as first author
35since 2021 · last 2026
0000-0001-5086-7858ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 158 · 13 first-author · 24 since 2021Artificial intelligence and machine learning · 59 · 1 first-author · 15 since 2021Systems, architecture and hardware · 11 · 3 since 2021Security and privacy · 5 · 2 since 2021Databases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Contextual Scalarisation Thompson Sampling for Multi-objective Decisions in Public Media
Théo Maëtz, Luc Guillet, Andrea Cavallaro |
ICPR (13) | 3 |
| 2026 | High-Resolution Open-Vocabulary Object 6D Pose EstimationabstractThe generalisation to unseen objects in the 6D pose estimation task is very challenging. While Vision-Language Models (VLMs) enable using natural language descriptions to support 6D pose estimation of unseen objects, these solutions underperform compared to model-based methods. In this work we present Horyon, an open-vocabulary VLM-based architecture that addresses relative pose estimation between two scenes of an unseen object, described by a textual prompt only. We use the textual prompt to identify the unseen object in the scenes and then obtain high-resolution multi-scale features. These features are used to extract cross-scene matches for registration. We evaluate our model on a benchmark with a large variety of unseen objects across four datasets, namely REAL275, Toyota-Light, Linemod, and YCB-Video. Our method achieves state-of-the-art performance on all datasets, outperforming by 12.6 in Average Recall the previous best-performing approach. Jaime Corsetti, Davide Boscaini, Francesco Giuliari, Changjae Oh, Andrea Cavallaro, Fabio Poiesi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | On the Consistency Between Subjective and Objective Evaluation for Speech Enhancement Under Low-SNR Drone NoiseabstractSpeech enhancement in the presence of strong drone noise at extremely low signal-to-noise ratios (SNRs) remains challenging, and conventional objective metrics may fail to accurately reflect perceived speech quality. In this paper, we evaluate several deep learning-based speech enhancement algorithms under severe drone-noise conditions using a diverse set of objective metrics, including signal fidelity, intelligibility, and perceptual quality measures. The results reveal substantial inconsistencies among metrics, leading to conflicting conclusions regarding algorithm performance. To address this issue, we conduct a controlled subjective listening experiment under extremely low-SNR conditions to assess the perceptual relevance of these metrics. The consistency analysis shows that learning-based perceptual metrics (e.g., NISQA) and intelligibility-oriented measures (e.g., ESTOI) achieve the highest agreement with subjective judgments, followed by SI-SDR, while traditional quality metrics such as PESQ exhibit weaker agreement. These findings highlight the limitations of conventional metrics in extremely low-SNR conditions and provide guidance for selecting perceptually meaningful evaluation criteria in drone audition. Dmitrii Mukhutdinov, Ashish Alex, Andrea Cavallaro, Lin Wang 0009 |
IEEE Signal Process. Lett. | 3 |
| 2025 | 3D Face Reconstruction Error Decomposed: A Modular Benchmark for Fair and Fast Method EvaluationabstractComputing the standard benchmark metric for 3D face reconstruction, namely geometric error, requires a number of steps, such as mesh cropping, rigid alignment, or point correspondence. Current benchmark tools are monolithic (they implement a specific combination of these steps), even though there is no consensus on the best way to measure error. We present a toolkit for a Modularized 3D Face reconstruction Benchmark (M3DFB), where the fundamental components of error computation are segregated and interchangeable, allowing one to quantify the effect of each. Furthermore, we propose a new component, namely correction, and present a computationally efficient approach that penalizes for mesh topology inconsistency. Using this toolkit, we test 16 error estimators with 10 reconstruction methods on two real and two synthetic datasets. Critically, the widely used ICP-based estimator provides the worst benchmarking performance, as it significantly alters the true ranking of the top- 5 reconstruction methods. Notably, the correlation of ICP with the true error can be as low as 0.41. Moreover, non-rigid alignment leads to significant improvement (correlation larger than 0.90), highlighting the importance of annotating 3D landmarks on datasets. Finally, the proposed correction scheme, together with non-rigid warping, leads to an accuracy on a par with the best non-rigid ICP-based estimators, but runs an order of magnitude faster. Our open-source codebase is designed for researchers to easily compare alternatives for each component, thus helping accelerating progress in benchmarking for 3D face reconstruction and, furthermore, supporting the improvement of learned reconstruction methods, which depend on accurate error estimation for effective training. Evangelos Sariyanidi, Claudio Ferrari, Federico Nocentini, Stefano Berretti, Andrea Cavallaro, Birkan Tunç |
FG | 5 |
| 2025 | Improving Generalization of Language-Conditioned Robot ManipulationabstractThe control of robots for manipulation tasks generally relies on visual input. Recent advances in vision-language models (VLMs) enable the use of natural language instructions to condition visual input and control robots in a wider range of environments. However, existing methods require a large amount of data to fine-tune VLMs for operating in unseen environments. In this paper, we present a framework that learns object-arrangement tasks from just a few demonstrations. We propose a two-stage framework that divides object-arrangement tasks into a target localization stage, for picking the object, and a region determination stage for placing the object. We present an instance-level semantic fusion module that aligns the instance-level image crops with the text embedding, enabling the model to identify the target objects defined by the natural language instructions. We validate our method on both simulation and real-world robotic environments. Our method, fine-tuned with a few demonstrations, improves generalization capability and demonstrates zero-shot ability in real-robot manipulation scenarios. Chenglin Cui, Chaoran Zhu, Changjae Oh, Andrea Cavallaro |
IROS | 4 |
| 2025 | NaviFormer: A Deep Reinforcement Learning Transformer-like Model to Holistically Solve the Navigation ProblemabstractPath planning is usually solved by addressing either the (high-level) route planning problem (waypoint sequencing to achieve the final goal) or the (low-level) path planning problem (trajectory prediction between two waypoints avoiding collisions). However, real-world problems usually require simultaneous solutions to the route and path planning subproblems with a holistic and efficient approach. In this paper, we introduce NaviFormer, a deep reinforcement learning model based on a Transformer architecture that solves the global navigation problem by predicting both high-level routes and low-level trajectories. To evaluate NaviFormer, several experiments have been conducted, including comparisons with other algorithms. Results show competitive accuracy from NaviFormer since it can understand the constraints and difficulties of each subproblem and act consequently to improve performance. Moreover, its superior computation speed proves its suitability for real-time missions. Daniel Fuertes, Andrea Cavallaro, Carlos R. del-Blanco, Fernando Jaureguizar, Narciso García |
IROS | 2 |
| 2025 | MM-HSD: Multi-Modal Hate Speech Detection in VideosabstractWhile hate speech detection (HSD) has been extensively studied in text, existing multi-modal approaches remain limited, particularly in videos. As modalities are not always individually informative, simple fusion methods fail to fully capture inter-modal dependencies. Moreover, previous work often omits relevant modalities such as on-screen text and audio, which may contain subtle hateful content and thus provide essential cues, both individually and in combination with others. In this paper, we present MM-HSD, a multi-modal model for HSD in videos that integrates video frames, audio, and text derived from speech transcripts and from frames (i.e.on-screen text) together with features extracted by Cross-Modal Attention (CMA). We are the first to use CMA as an early feature extractor for HSD in videos, to systematically compare query/key configurations, and to evaluate the interactions between different modalities in the CMA block. Our approach leads to improved performance when on-screen text is used as a query and the rest of the modalities serve as a key. Experiments on the HateMM dataset show that MM-HSD outperforms state-of-the-art methods on M-F1 score (0.874), using concatenation of transcript, audio, video, on-screen text, and CMA for feature extraction on raw embeddings of the modalities. The code is available at https://github.com/idiap/mm-hsd. Berta Céspedes-Sarrias, Carlos Collado-Capell, Pablo Rodenas-Ruiz, Olena Hrynenko, Andrea Cavallaro |
ACM Multimedia | 5 |
| 2025 | Learning human-to-robot handovers through 3D scene reconstructionabstractLearning robot manipulation policies from raw, real-world image data requires a large number of robot-action trials in the physical environment. Although training using simulations offers a cost-effective alternative, the visual domain gap between simulation and robot workspace remains a major limitation. Gaussian Splatting visual reconstruction methods have recently provided new directions for robot manipulation by generating realistic environments. In this paper, we propose the first method for learning supervised-based robot handovers solely from RGB images without the need of real-robot training or real-robot data collection. The proposed policy learner, Human-to-Robot Handover using Sparse-View Gaussian Splatting (H2RH-SGS), leverages sparse-view Gaussian Splatting reconstruction of human-to-robot handover scenes to generate robot demonstrations containing image-action pairs captured with a camera mounted on the robot gripper. As a result, the simulated camera pose changes in the reconstructed scene can be directly translated into gripper pose changes. We train a robot policy on demonstrations collected with 16 household objects and directly deploy this policy in the real environment. Experiments in both Gaussian Splatting reconstructed scene and real-world human-to-robot handover experiments demonstrate that H2RH-SGS serves as a new and effective representation for the human-to-robot handover task. Yuekun Wu, Yik Lung Pang, Andrea Cavallaro, Changjae Oh |
RO-MAN | 3 |
| 2025 | Identifying Privacy PersonasabstractPrivacy personas capture the differences in user segments with respect to one’s knowledge, behavioural patterns, level of self-efficacy, and perception of the importance of privacy protection. Modelling these differences is essential for appropriately choosing personalised communication about privacy (e.g. to increase literacy) and for defining suitable choices for privacy enhancing technologies (PETs). While various privacy personas have been derived in the literature, they group together people who differ from each other in terms of important attributes such as perceived or desired level of control, and motivation to use PET. To address this lack of granularity and comprehensiveness in describing personas, we propose eight personas that we derive by combining qualitative and quantitative analysis of the responses to an interactive educational questionnaire. We design an analysis pipeline that uses divisive hierarchical clustering and Boschloo’s statistical test of homogeneity of proportions to ensure that the elicited clusters differ from each other based on a statistical measure. Additionally, we propose a new measure for calculating distances between questionnaire responses, that accounts for the type of the question (closed- vs open-ended) used to derive traits. We show that the proposed privacy personas statistically differ from each other. We statistically validate the proposed personas and also compare them with personas in the literature, showing that they provide a more granular and comprehensive understanding of user segments, which will allow to better assist users with their privacy needs. Olena Hrynenko, Andrea Cavallaro |
Proc. Priv. Enhancing Technol. | 2 |
| 2025 | Learning Privacy from Visual EntitiesabstractSubjective interpretation and content diversity make predicting whether an image is private or public a challenging task. Graph neural networks combined with convolutional neural networks (CNNs), which consist of 14,000 to 500 millions parameters, generate features for visual entities (e.g., scene and object types) and identify the entities that contribute to the decision. In this paper, we show that using a simpler combination of transfer learning and a CNN to relate privacy with scene types optimises only 732 parameters while achieving comparable performance to that of graph-based methods. On the contrary, end-to-end training of graph-based methods can mask the contribution of individual components to the classification performance. Furthermore, we show that a high-dimensional feature vector, extracted with CNNs for each visual entity, is unnecessary and complexifies the model. The graph component has also negligible impact on performance, which is driven by fine-tuning the CNN to optimize image features for privacy nodes. Alessio Xompero, Andrea Cavallaro |
Proc. Priv. Enhancing Technol. | 2 |
| 2024 | Open-vocabulary object 6D pose estimationabstractWe introduce the new setting of open-vocabulary object 6D pose estimation, in which a textual prompt is used to specify the object of interest. In contrast to existing approaches, in our setting (i) the object of interest is speci-fied solely through the textual prompt, (ii) no object model (e.g., CAD or video sequence) is required at inference, and (iii) the object is imaged from two RGBD viewpoints of dif-ferent scenes. To operate in this setting, we introduce a novel approach that leverages a Vision-Language Model to segment the object of interest from the scenes and to esti-mate its relative 6D pose. The key of our approach is a carefully devised strategy to fuse object-level information provided by the prompt with local image features, resulting in a feature space that can generalize to novel concepts. We validate our approach on a new benchmark based on two popular datasets, REAL275 and Toyota-Light, which collectively encompass 34 object instances appearing in four thousand image pairs. The results demonstrate that our approach outperforms both a well-established hand-crafted method and a recent deep learning-based base-line in estimating the relative 6D pose of objects in dif-ferent scenes. Code and dataset are available at https://jcorsetti.github.io/oryon. Jaime Corsetti, Davide Boscaini, Changjae Oh, Andrea Cavallaro, Fabio Poiesi |
CVPR | 4 |
| 2024 | Test-time adaptation for 6D pose trackingabstractWe propose a test-time adaptation for 6D object pose tracking that learns to adapt a pre-trained model to track the 6D pose of novel objects. We consider the problem of 6D object pose tracking as a 3D keypoint detection and matching task and present a model that extracts 3D keypoints. Given an RGB-D image and the mask of a target object for each frame, the proposed model consists of the self- and cross-attention modules to produce the features that aggregate the information within and across frames, respectively. By using the keypoints detected from the features for each frame, we estimate the pose changes between two frames, which enables 6D pose tracking when the 6D pose of a target object in the initial frame is given. Our model is first trained in a source domain, a category-level tracking dataset where the ground truth 6D pose of the object is available. To deploy this pre-trained model to track novel objects, we present a test-time adaptation strategy that trains the model to adapt to the target novel object by self-supervised learning. Given an RGB-D video sequence of the novel object, the proposed self-supervised losses encourage the model to estimate the 6D pose changes that can keep the photometric and geometric consistency of the object. We validate our method on the NOCS-REAL275 dataset and our collected dataset, and the results show the advantages of tracking novel objects. The collected dataset and visualisation of tracking results are available: https://bartektian.github.io/TA-6DT.html Changjae Oh, Andrea Cavallaro |
Pattern Recognit. | 3 |
| 2023 | Human-interpretable and deep features for image privacy classificationabstractPrivacy is a complex, subjective and contextual concept that is difficult to define. Therefore, the annotation of images to train privacy classifiers is a challenging task. In this paper, we analyse privacy classification datasets and the properties of controversial images that are annotated with contrasting privacy labels by different assessors. We discuss suitable features for image privacy classification and propose eight privacy-specific and human-interpretable features. These features increase the performance of deep learning models and, on their own, improve the image representation for privacy classification compared with much higher dimensional deep features. Darya Baranouskaya, Andrea Cavallaro |
ICIP | 2 |
| 2023 | Digital assets rights management through smart legal contracts and smart contractsabstractIntellectual property rights (IPR) management needs to evolve in a digital world where not only companies but also many independent content creators contribute to our culture with their art, music, and videos. In this respect, blockchain has recently emerged as a promising infrastructure providing a trustworthy and immutable environment that, thanks to smart contracts, may enable more agile management of digital rights and streamline royalty payments. However, no widespread consensus has been reached on the ability of this technology to adequately manage and transfer IPR. This paper presents an innovative approach to digital rights management developed within the scope of an international research endeavour co-financed by the European Commission named MediaVerse. The approach proposes the combined usage of smart legal contracts and blockchain smart contracts to take care of the legally-binding contractual aspects of IPR and, at the same time, the need for notarization, rights transfer, and royalty payments. The work conducted represents a contribution to advancing the current literature on IPR management that may lead to an improved and fairer monetization process for content creators as a means of individual empowerment. Enrico Ferro, Marco Saltarella, Domenico Rotondi, Marco Giovanelli, Giacomo Corrias, Roberto H. Moncada, Andrea Cavallaro, Alfredo Favenza |
Blockchain Res. Appl. | 7 |
| 2023 | Data augmentation for speech separationabstractDeep learning models have advanced the state of the art of monaural speech separation. However, the performance of a separation model considerably decreases when tested on unseen speakers and noisy conditions. Separation models trained with data augmentation generalize better to unseen conditions. In this paper, we conduct a comprehensive survey of data augmentation techniques and apply them to improve the generalization of time-domain speech separation models. The augmentation techniques include seven source-preserving approaches (Gaussian noise, Gain, Time masking, frequency masking, Short noise, Time stretch, and Pitch shift) and three non-source preserving approaches (Dynamix mixing, Mixup, and Cutmix). After hyperparameter search for each augmentation method, we test the generalization of the augmented model by cross-corpus testing on three datasets (LibriMix, TIMIT, and VCTK), and identify the best augmentation combination that enhances generalization. Experimental results indicate that a combination of several non-source preserving strategies (CutMix, Mixup, and Dynamic mixing) resulted in the best generalization performance. Finally, the augmentation combinations also improved the performance of the speech separation model even when fewer training data are available. Ashish Alex, Lin Wang 0009, Paolo Gastaldo, Andrea Cavallaro |
Speech Commun. | 4 |
| 2022 | Implicit texture mapping for multi-view video synthesis
Mohamed Ilyes Lakhal, Oswald Lanz, Andrea Cavallaro |
BMVC | 3 |
| 2022 | Selective Colour Restoration of Underwater Surfaces
Chau Yi Li, Andrea Cavallaro |
BMVC | 2 |
| 2022 | Training Privacy-Preserving Video Analytics Pipelines by Suppressing Features That Reveal Information About Private AttributesabstractDeep neural networks are increasingly deployed for scene analytics, including to evaluate the attention and reaction of people exposed to out-of-home advertisements. However, the features extracted by a deep neural network that was trained to predict a specific, consensual attribute (e.g. emotion) may also encode and thus reveal information about private, protected attributes (e.g. age or gender). In this work, we focus on such leakage of private information at inference time. We consider an adversary with access to the features extracted by the layers of a deployed neural network and use these features to predict private attributes. To prevent the success of such an attack, we modify the training of the network using a confusion loss that encourages the extraction of features that make it difficult for the adversary to accurately predict private attributes. We validate this training approach on image-based tasks using a publicly available dataset. Results show that, compared to the original network, the proposed PrivateNet can reduce the leakage of private information of a state-of-the-art emotion recognition classifier by 2.88% for gender and by 13.06% for age group, with a minimal effect on task accuracy. Chau Yi Li, Andrea Cavallaro |
ICASSP | 2 |
| 2022 | Is Cross-Attention Preferable to Self-Attention for Multi-Modal Emotion Recognition?abstractHumans express their emotions via facial expressions, voice intonation and word choices. To infer the nature of the underlying emotion, recognition models may use a single modality, such as vision, audio, and text, or a combination of modalities. Generally, models that fuse complementary information from multiple modalities outperform their uni-modal counterparts. However, a successful model that fuses modalities requires components that can effectively aggregate task-relevant information from each modality. As cross-modal attention is seen as an effective mechanism for multi-modal fusion, in this paper we quantify the gain that such a mechanism brings compared to the corresponding self-attention mechanism. To this end, we implement and compare a cross-attention and a self-attention model. In addition to attention, each model uses convolutional layers for local feature extraction and recurrent layers for global sequential modelling. We compare the models using different modality combinations for a 7-class emotion classification task using the IEMOCAP dataset. Experimental results indicate that albeit both models improve upon the state-of-the-art in terms of weighted and unweighted accuracy for tri- and bi-modal configurations, their performance is generally statistically comparable. The code to replicate the experiments is available at https://github.com/smartcameras/SelfCrossAttn Vandana Rajan, Alessio Brutti, Andrea Cavallaro |
ICASSP | 3 |
| 2022 | Audio-Visual Object Classification for Human-Robot CollaborationabstractHuman-robot collaboration requires the contactless estimation of the physical properties of containers manipulated by a person, for example while pouring content in a cup or moving a food box. Acoustic and visual signals can be used to estimate the physical properties of such objects, which may vary substantially in shape, material and size, and also be occluded by the hands of the person. To facilitate comparisons and stimulate progress in solving this problem, we present the CORSMAL challenge and a dataset to assess the performance of the algorithms through a set of well-defined performance scores. The tasks of the challenge are the estimation of the mass, capacity, and dimensions of the object (container), and the classification of the type and amount of its content. A novel feature of the challenge is our real-to-simulation framework for visualising and assessing the impact of estimation errors in human-to-robot handovers. Alessio Xompero, Yik Lung Pang, T. Patten, A. Prabhakar, Berk Çalli, Andrea Cavallaro |
ICASSP | 6 |
| 2022 | On The Limits of Perceptual Quality Measures for Enhanced Underwater ImagesabstractThe appearance of objects in underwater images is degraded by the selective attenuation of light, which reduces contrast and causes a colour cast. This degradation depends on the water environment, and increases with depth and with the distance of the object from the camera. Despite an increasing volume of works in underwater image enhancement and restoration, the lack of a commonly accepted evaluation measure is hindering the progress as it is difficult to compare methods. In this paper, we review commonly used colour accuracy measures, such as colour reproduction error and CIEDE2000, and no-reference image quality measures, such as UIQM, UCIQE and CCF, which have not yet been systematically validated. We show that none of the no-reference quality measures satisfactorily rates the quality of enhanced underwater images and discuss their main shortcomings. Images and results are available at https://puiqe.eecs.qmul.ac.uk. Chau Yi Li, Andrea Cavallaro |
ICIP | 2 |
| 2022 | Cluster-Based 3D Keypoint Detection for Category-Agnostic 6D Pose TrackingabstractWe present a model for category-agnostic 6D pose tracking. We tackle object pose tracking as a 3D keypoint detection and matching task that does not require ground-truth annotation of the keypoints. Using RGB-D data and the target object mask as inputs, we spatially segment the point cloud of the object into clusters. Each 3D point in the cluster is characterised by features encoding appearance and geometric information. We use these features to detect a keypoint for each cluster and, with the detected keypoint sets from two time instants, we recover the pose change through least-squares optimisation. The loss functions are designed to ensure that the detected keypoints are consistent over time and suitable for pose tracking. Andrea Cavallaro, Changjae Oh |
ICIP | 2 |
| 2022 | Generating gender-ambiguous voices for privacy-preserving speech recognitionabstractOur voice encodes a uniquely identifiable pattern which can be used to infer private attributes, such as gender or identity, that an individual might wish not to reveal when using a speech recognition service.To prevent attribute inference attacks alongside speech recognition tasks, we present a generative adversarial network, GenGAN, that synthesises voices that conceal the gender or identity of a speaker.The proposed network includes a generator with a U-Net architecture that learns to fool a discriminator.We condition the generator only on gender information and use an adversarial loss between signal distortion and privacy preservation.We show that GenGAN improves the tradeoff between privacy and utility compared to privacy-preserving representation learning methods that consider gender information as a sensitive attribute to protect. Dimitrios Stoidis, Andrea Cavallaro |
INTERSPEECH | 2 |
| 2022 | Learning Generalisable Omni-Scale Representations for Person Re-IdentificationabstractAn effective person re-identification (re-ID) model should learn feature representations that are both discriminative, for distinguishing similar-looking people, and generalisable, for deployment across datasets without any adaptation. In this paper, we develop novel CNN architectures to address both challenges. First, we present a re-ID CNN termed omni-scale network (OSNet) to learn features that not only capture different spatial scales but also encapsulate a synergistic combination of multiple scales, namely omni-scale features. The basic building block consists of multiple convolutional streams, each detecting features at a certain scale. For omni-scale feature learning, a unified aggregation gate is introduced to dynamically fuse multi-scale features with channel-wise weights. OSNet is lightweight as its building blocks comprise factorised convolutions. Second, to improve generalisable feature learning, we introduce instance normalisation (IN) layers into OSNet to cope with cross-dataset discrepancies. Further, to determine the optimal placements of these IN layers in the architecture, we formulate an efficient differentiable architecture search algorithm. Extensive experiments show that, in the conventional same-dataset setting, OSNet achieves state-of-the-art performance, despite being much smaller than existing re-ID models. In the more challenging yet practical cross-dataset setting, OSNet beats most recent unsupervised domain adaptation methods without using any target data. Kaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Audio-Visual Tracking of Concurrent SpeakersabstractAudio-visual tracking of an unknown number of concurrent speakers in 3D is a challenging task, especially when sound and video are collected with a compact sensing platform. In this paper, we propose a tracker that builds on generative and discriminative audio-visual likelihood models formulated in a particle filtering framework. We localize multiple concurrent speakers with a de-emphasized acoustic map assisted by the image detection-derived 3D video observations. The 3D multi-modal observations are either assigned to existing tracks for discriminative likelihood computation or used to initialize new tracks. The generative likelihoods rely on color distribution of the target and the de-emphasized acoustic map value. Experiments on AV16.3 and CAV3D datasets show that the proposed tracker outperforms the uni-modal trackers and the state-of-the-art approaches both in 3D and on the image plane. Xinyuan Qian 0001, Alessio Brutti, Oswald Lanz, Maurizio Omologo, Andrea Cavallaro |
IEEE Trans. Multim. | 5 |
| 2021 | Robust Latent Representations Via Cross-Modal Translation and AlignmentabstractMulti-modal learning relates information across observation modalities of the same physical phenomenon to leverage complementary information. Most multi-modal machine learning methods require that all the modalities used for training are also available for testing. This is a limitation when signals from some modalities are unavailable or severely degraded. To address this limitation, we aim to improve the testing performance of uni-modal systems using multiple modalities during training only. The proposed multi-modal training framework uses cross-modal translation and correlation-based latent space alignment to improve the representations of a worse performing (or weaker) modality. The translation from the weaker to the better performing (or stronger) modality generates a multi-modal intermediate encoding that is representative of both modalities. This encoding is then correlated with the stronger modality representation in a shared latent space. We validate the proposed framework on the AVEC 2016 dataset (RECOLA) for continuous emotion recognition and show the effectiveness of the framework that achieves state-of- the-art (uni-modal) performance for weaker modalities. Vandana Rajan, Alessio Brutti, Andrea Cavallaro |
ICASSP | 3 |
| 2021 | FoolHD: Fooling Speaker Identification by Highly Imperceptible Adversarial DisturbancesabstractSpeaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial attack that generates imperceptible adversarial perturbations against a speaker identification model. Our approach, FoolHD, uses a Gated Convolutional Autoencoder that operates in the DCT domain and is trained with a multi-objective loss function, to generate and conceal the adversarial perturbation within the original audio files. In addition to hindering speaker identification performance, this multi-objective loss accounts for human perception through a frame-wise cosine similarity between MFCC feature vectors extracted from the original and adversarial audio files. We validate the effectiveness of FoolHD with a 250-speaker identification x-vector network, trained using VoxCeleb, in terms of accuracy, success rate, and imperceptibility. Our results show that FoolHD generates highly imperceptible adversarial audio files (average PESQ scores above 4.30), while achieving a success rate of 99.6% and 99.2% in misleading the speaker identification model, for untargeted and targeted settings, respectively. Ali Shahin Shamsabadi, Francisco Teixeira, Alberto Abad, Bhiksha Raj, Andrea Cavallaro, Isabel Trancoso |
ICASSP | 5 |
| 2021 | On the Reversibility of Adversarial AttacksabstractAdversarial attacks modify images with perturbations that change the prediction of classifiers. These modified images, known as adversarial examples, expose the vulnerabilities of deep neural network classifiers. In this paper, we investigate the predictability of the mapping between the classes predicted for original images and for their corresponding adversarial examples. This predictability relates to the possibility of retrieving the original predictions and hence reversing the induced misclassification. We refer to this property as the reversibility of an adversarial attack, and quantify reversibility as the accuracy in retrieving the original class or the true class of an adversarial example. We present an approach that reverses the effect of an adversarial attack on a classifier using a prior set of classification results. We analyse the reversibility of state-of-the-art adversarial attacks on benchmark classifiers and discuss the factors that affect the reversibility. Chau Yi Li, Ricardo Sanchez-Matilla, Ali Shahin Shamsabadi, Riccardo Mazzon, Andrea Cavallaro |
ICIP | 5 |
| 2021 | Improving Filling Level Classification with Adversarial TrainingabstractWe investigate the problem of classifying–from a single image–the level of content in a cup or a drinking glass. This problem is made challenging by several ambiguities caused by transparencies, shape variations and partial occlusions, and by the availability of only small training datasets. In this paper, we tackle this problem with an appropriate strategy for transfer learning. Specifically, we use adversarial training in a generic source dataset and then refine the training with a task-specific dataset. We also discuss and experimentally evaluate several training strategies and their combination on a range of container types of the CORSMAL Containers Manipulation dataset. We show that transfer learning with adversarial training in the source domain consistently improves the classification accuracy on the test set and limits the overfitting of the classifier to specific features of the training data. Apostolos Modas, Alessio Xompero, Ricardo Sanchez-Matilla, Pascal Frossard, Andrea Cavallaro |
ICIP | 5 |
| 2021 | Protecting Gender and Identity with Disentangled Speech RepresentationsabstractBesides its linguistic content, our speech is rich in biometric information that can be inferred by classifiers. Learning privacy-preserving representations for speech signals enables downstream tasks without sharing unnecessary, private information about an individual. In this paper, we show that protecting gender information in speech is more effective than modelling speaker-identity information only when generating a non-sensitive representation of speech. Our method relies on reconstructing speech by decoding linguistic content along with gender information using a variational autoencoder. Specifically, we exploit disentangled representation learning to encode information about different attributes into separate subspaces that can be factorised independently. We present a novel way to encode gender information and disentangle two sensitive biometric identifiers, namely gender and identity, in a privacy-protecting setting. Experiments on the LibriSpeech dataset show that gender recognition and speaker verification can be reduced to a random guess, protecting against classification-based attacks. Dimitrios Stoidis, Andrea Cavallaro |
Interspeech | 2 |
| 2021 | OHPL: One-shot Hand-eye Policy LearnerabstractThe control of a robot for manipulation tasks generally relies on object detection and pose estimation. An attractive alternative is to learn control policies directly from raw input data. However, this approach is time-consuming and expensive since learning the policy requires many trials with robot actions in the physical environment. To reduce the training cost, the policy can be learned in simulation with a large set of synthetic images. The limit of this approach is the domain gap between the simulation and the robot workspace. In this paper, we propose to learn a policy for robot reaching movements from a single image captured directly in the robot workspace from a camera placed on the end-effector (a hand-eye camera). The idea behind the proposed policy learner is that view changes seen from the hand-eye camera produced by actions in the robot workspace are analogous to locating a region-of-interest in a single image by performing sequential object localisation. This similar view change enables training of object reaching policies using reinforcement-learning-based sequential object localisation. To facilitate the adaptation of the policy to view changes in the robot workspace, we further present a dynamic filter that learns to bias an input state to remove irrelevant information for an action decision. The proposed policy learner can be used as a powerful representation for robotic tasks, and we validate it on static and moving object reaching tasks. Changjae Oh, Yik Lung Pang, Andrea Cavallaro |
IROS | 3 |
| 2021 | Mixup Augmentation for Generalizable Speech SeparationabstractDeep learning has advanced the state of the art of single-channel speech separation. However, separation models may overfit the training data and generalization across datasets is still an open problem in real-world conditions with noise. In this paper we address the generalization problem with Mixup as data augmentation approach. Mixup creates new training examples from linear combinations of samples during mini-batch training. We propose four variations of Mixup and assess the improved generalization of a speech separation model, DPRNN, with cross-corpus evaluation on LibriMix, TIMIT and VCTK datasets. DPRNN allows efficient modelling of longer input sequences by splitting the learnt representation from input mixture segment into small chunks and performing intra and inter chunk operations iteratively. We show that training DPRNN with the proposed Data-only Mixup augmentation variation improves performance on an unseen dataset in noisy conditions when compared to the baseline SpecAugment augmented models, while having comparable performance on the source dataset. Ashish Alex, Lin Wang 0009, Paolo Gastaldo, Andrea Cavallaro |
MMSP | 4 |
| 2021 | Towards safe human-to-robot handovers of unknown containersabstractSafe human-to-robot handovers of unknown objects require accurate estimation of hand poses and object properties, such as shape, trajectory, and weight. Accurately estimating these properties requires the use of scanned 3D object models or expensive equipment, such as motion capture systems and markers, or both. However, testing handover algorithms with robots may be dangerous for the human and, when the object is an open container with liquids, for the robot. In this paper, we propose a real-to-simulation framework to develop safe human-to-robot handovers with estimations of the physical properties of unknown cups or drinking glasses and estimations of the human hands from videos of a human manipulating the container. We complete the handover in simulation, and we estimate a region that is not occluded by the hand of the human holding the container. We also quantify the safeness of the human and object in simulation. We validate the framework using public recordings of containers manipulated before a handover and show the safeness of the handover when using noisy estimates from a range of perceptual algorithms. Yik Lung Pang, Alessio Xompero, Changjae Oh, Andrea Cavallaro |
RO-MAN | 4 |
| 2021 | View-Action Representation Learning for Active First-Person VisionabstractIn visual navigation, a moving agent equipped with a camera is traditionally controlled by an input action and the estimation of the features from a sensory state (i.e. the camera view) is treated as a pre-processing step to perform high-level vision tasks. In this paper, we present a representation learning approach that, instead, considers both state and action as inputs. We condition the encoded feature from the state transition network on the action that changes the view of the camera, thus describing the scene more effectively. Specifically, we introduce an action representation module that generates decoded higher dimensional representations from an input action to increase the representational power. We then fuse the output from the action representation module with the intermediate response of the state transition network that predicts the future state. To enhance the discrimination capability among predictions from different input actions, we further introduce triplet ranking loss and $N$ -tuplet loss functions, which in turn can be integrated with the regression loss. We demonstrate the proposed representation learning approach in reinforcement and imitation learning-based mapless navigation tasks, where the camera agent learns to navigate only through the view of the camera and the performed action, without external information. Changjae Oh, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Semantically Adversarial Learnable FiltersabstractWe present an adversarial framework to craft perturbations that mislead classifiers by accounting for the image content and the semantics of the labels. The proposed framework combines a structure loss and a semantic adversarial loss in a multi-task objective function to train a fully convolutional neural network. The structure loss helps generate perturbations whose type and magnitude are defined by a target image processing filter. The semantic adversarial loss considers groups of (semantic) labels to craft perturbations that prevent the filtered image from being classified with a label in the same group. We validate our framework with three different target filters, namely detail enhancement, log transformation and gamma correction filters; and evaluate the adversarially filtered images against three classifiers, ResNet50, ResNet18 and AlexNet, pre-trained on ImageNet. We show that the proposed framework generates filtered images with a high success rate, robustness, and transferability to unseen classifiers. We also discuss objective and subjective evaluations of the adversarial perturbations. Ali Shahin Shamsabadi, Changjae Oh, Andrea Cavallaro |
IEEE Trans. Image Process. | 3 |
| 2020 | Novel-View Human Action Synthesis
Mohamed Ilyes Lakhal, Davide Boscaini, Fabio Poiesi, Oswald Lanz, Andrea Cavallaro |
ACCV (4) | 5 |
| 2020 | ColorFool: Semantic Adversarial ColorizationabstractAdversarial attacks that generate small Lp norm perturbations to mislead classifiers have limited success in black-box settings and with unseen classifiers. These attacks are also not robust to defenses that use denoising filters and to adversarial training procedures. Instead, adversarial attacks that generate unrestricted perturbations are more robust to defenses, are generally more successful in black-box settings and are more transferable to unseen classifiers. However, unrestricted perturbations may be noticeable to humans. In this paper, we propose a content-based black-box adversarial attack that generates unrestricted perturbations by exploiting image semantics to selectively modify colors within chosen ranges that are perceived as natural by humans. We show that the proposed approach, ColorFool, outperforms in terms of success rate, robustness to defense frameworks and transferability, five state-of-the-art adversarial attacks on two different tasks, scene and object classification, when attacking three state-of-the-art deep neural networks using three standard datasets. The source code is available at https://github.com/smartcameras/ColorFool. Ali Shahin Shamsabadi, Ricardo Sanchez-Matilla, Andrea Cavallaro |
CVPR | 3 |
| 2020 | Edgefool: an Adversarial Image Enhancement FilterabstractAdversarial examples are intentionally perturbed images that mislead classifiers. These images can, however, be easily detected using denoising algorithms, when high-frequency spatial perturbations are used, or can be noticed by humans, when perturbations are large. In this paper, we propose EdgeFool, an adversarial image enhancement filter that learns structure-aware adversarial perturbations. Edge-Fool generates adversarial images with perturbations that enhance image details via training a fully convolutional neural network end-to-end with a multi-task loss function. This loss function accounts for both image detail enhancement and class misleading objectives. We evaluate EdgeFool on three classifiers (ResNet-50, ResNet-18 and AlexNet) using two datasets (ImageNet and Private-Places365) and compare it with six adversarial methods (DeepFool, SparseFool, Carlini-Wagner, SemanticAdv, Non-targeted and Private Fast Gradient Sign Methods). Code is available at https://github.com/smartcameras/EdgeFool.git. Ali Shahin Shamsabadi, Changjae Oh, Andrea Cavallaro |
ICASSP | 3 |
| 2020 | Multi-View Shape Estimation of Transparent ContainersabstractThe 3D localisation of an object and the estimation of its properties, such as shape and dimensions, are challenging under varying degrees of transparency and lighting conditions. In this paper, we propose a method for jointly localising container-like objects and estimating their dimensions using two wide-baseline, calibrated RGB cameras. Under the assumption of circular symmetry along the vertical axis, we estimate the dimensions of an object with a generative 3D sampling model of sparse circumferences, iterative shape fitting and image re-projection to verify the sampling hypotheses in each camera using semantic segmentation masks. We evaluate the proposed method on a novel dataset of objects with different degrees of transparency and captured under different backgrounds and illumination conditions. Our method, which is based on RGB images only, outperforms in terms of localisation success and dimension estimation accuracy a deep-learning based approach that uses depth maps. Alessio Xompero, Ricardo Sanchez-Matilla, Apostolos Modas, Pascal Frossard, Andrea Cavallaro |
ICASSP | 5 |
| 2020 | Cast-Gan: Learning To Remove Colour Cast From Underwater ImagesabstractUnderwater images are degraded by blur and colour cast caused by the attenuation of light in water. To remove the colour cast with neural networks, images of the scene taken under white illumination are needed as reference for training, but are generally unavailable. As an alternative, one can use surrogate reference images taken close to the water surface or degraded images synthesised from reference datasets. However, the former still suffer from colour cast and the latter generally have limited colour diversity. To address these problems, we exploit open data and typical colour distributions of objects to create a synthetic image dataset that reflects degradations naturally occurring in underwater photography. We use this dataset to train Cast-GAN, a Generative Adversarial Network whose loss function includes terms that eliminate artefacts that are typical of underwater images enhanced with neural networks. We compare the enhancement results of Cast-GAN with four state-of-the-art methods and validate the cast removal with a subjective evaluation. Chau Yi Li, Andrea Cavallaro |
ICIP | 2 |
| 2020 | Privacy-preserving Machine Learning for Multimedia Data
Andrea Cavallaro |
ICPRAM | 1 |
| 2020 | Deep Learning for Privacy in MultimediaabstractWe discuss the design and evaluation of machine learning algorithms that provide users with more control on the multimedia information they share. We introduce privacy threats for multimedia data and key features of privacy protection. We cover privacy threats and mitigating actions for images, videos, and motion-sensor data from mobile and wearable devices, and their protection from unwanted, automatic inferences. The tutorial offers theoretical explanations followed by examples with software developed by the presenters and distributed as open source. Andrea Cavallaro, Mohammad Malekzadeh, Ali Shahin Shamsabadi |
ACM Multimedia | 1 |
| 2020 | DarkneTZ: towards model privacy at the edge using trusted execution environmentsabstractWe present DarkneTZ, a framework that uses an edge device's Trusted Execution Environment (TEE) in conjunction with model partitioning to limit the attack surface against Deep Neural Networks (DNNs). Increasingly, edge devices (smartphones and consumer IoT devices) are equipped with pre-trained DNNs for a variety of applications. This trend comes with privacy risks as models can leak information about their training data through effective membership inference attacks (MIAs). Fan Mo 0004, Ali Shahin Shamsabadi, Kleomenis Katevas, Soteris Demetriou, Ilias Leontiadis, Andrea Cavallaro, Hamed Haddadi 0001 |
MobiSys | 6 |
| 2020 | Privacy and utility preserving sensor-data transformations
Mohammad Malekzadeh, Richard G. Clegg, Andrea Cavallaro, Hamed Haddadi 0001 |
Pervasive Mob. Comput. | 3 |
| 2020 | A Blind Source Separation Framework for Ego-Noise Reduction on Multi-Rotor DronesabstractAcoustic sensing from a multi-rotor drone is heavily degraded by the strong ego-noise produced by the rotating motors and propellers. To address this problem, we propose a blind source separation (BSS) framework that extracts a target sound from noisy multi-channel signals captured by a microphone array mounted on a drone. The proposed method addresses the challenging problem of permutation alignment, in extremely low signal-to-noise-ratio scenarios (e.g. SNR <; -15 dB), by performing clustering on the time activities of the separated signals across frequencies. Since initialization plays an important role to the success of clustering, we propose a pre-processing algorithm which uses time-frequency spatial filtering (TFS) to generate a reference to pre-align the permutation. The pre-alignment not only improves the performance of clustering and permutation alignment, but also solves the target-channel selection problem for BSS. The proposed method integrates the advantages of both TFS and BSS. Experimental results with real-recorded data show that the proposed method is capable of processing the audio stream continuously in a blockwise manner and also remarkably outperforms the state-of-the-art. Lin Wang 0009, Andrea Cavallaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | PrivEdge: From Local to Distributed Private Training and PredictionabstractMachine Learning as a Service (MLaaS) operators provide model training and prediction on the cloud. MLaaS applications often rely on centralised collection and aggregation of user data, which could lead to significant privacy concerns when dealing with sensitive personal data. To address this problem, we propose PrivEdge, a technique for privacy-preserving MLaaS that safeguards the privacy of users who provide their data for training, as well as users who use the prediction service. With PrivEdge, each user independently uses their private data to locally train a one-class reconstructive adversarial network that succinctly represents their training data. As sending the model parameters to the service provider in the clear would reveal private information, PrivEdge secret-shares the parameters among two non-colluding MLaaS providers, to then provide cryptographically private prediction services through secure multi-party computation techniques. We quantify the benefits of PrivEdge and compare its performance with state-of-the-art centralised architectures on three privacy-sensitive image-based tasks: individual identification, writer identification, and handwritten letter recognition. Experimental results show that PrivEdge has high precision and recall in preserving privacy, as well as in distinguishing between private and non-private images. Moreover, we show the robustness of PrivEdge to image compression and biased training data. The source code is available at https://github.com/smartcameras/PrivEdge. Ali Shahin Shamsabadi, Adrià Gascón, Hamed Haddadi 0001, Andrea Cavallaro |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | A Spatio-Temporal Multi-Scale Binary DescriptorabstractBinary descriptors are widely used for multi-view matching and robotic navigation. However, their matching performance decreases considerably under severe scale and viewpoint changes in non-planar scenes. To overcome this problem, we propose to encode the varying appearance of selected 3D scene points tracked by a moving camera with compact spatio-temporal descriptors. To this end, we first track interest points and capture their temporal variations at multiple scales. Then, we validate feature tracks through 3D reconstruction and compress the temporal sequence of descriptors by encoding the most frequent and stable binary values. Finally, we determine multiscale correspondences across views with a matching strategy that handles severe scale differences. The proposed spatio-temporal multi-scale approach is generic and can be used with a variety of binary descriptors. We show the effectiveness of the joint multiscale extraction and temporal reduction through comparisons of different temporal reduction strategies and the application to several binary descriptors. Alessio Xompero, Oswald Lanz, Andrea Cavallaro |
IEEE Trans. Image Process. | 3 |
| 2020 | Exploiting Vulnerabilities of Deep Neural Networks for Privacy ProtectionabstractAdversarial perturbations can be added to images to protect their content from unwanted inferences. These perturbations may, however, be ineffective against classifiers that were not seen during the generation of the perturbation, or against defenses based on re-quantization, median filtering or JPEG compression. To address these limitations, we present an adversarial attack that is specifically designed to protect visual content against unseen classifiers and known defenses. We craft perturbations using an iterative process that is based on the Fast Gradient Signed Method and that randomly selects a classifier and a defense, at each iteration. This randomization prevents an undesirable overfitting to a specific classifier or defense. We validate the proposed attack in both targeted and untargeted settings on the private classes of the Places365-Standard dataset. Using ResNet18, ResNet50, AlexNet and DenseNet161 as classifiers, the performance of the proposed attack exceeds that of eleven state-of-the-art attacks. Ricardo Sanchez-Matilla, Chau Yi Li, Ali Shahin Shamsabadi, Riccardo Mazzon, Andrea Cavallaro |
IEEE Trans. Multim. | 5 |
| 2019 | Poster: Towards Characterizing and Limiting Information Exposure in DNN LayersabstractPre-trained Deep Neural Network (DNN) models are increasingly used in smartphones and other user devices to enable prediction services, leading to potential disclosures of (sensitive) information from training data captured inside these models. Based on the concept of generalization error, we propose a framework to measure the amount of sensitive information memorized in each layer of a DNN. Our results show that, when considered individually, the last layers encode a larger amount of information from the training data compared to the first layers. We find that the same DNN architecture trained with different datasets has similar exposure per layer. We evaluate an architecture to protect the most sensitive layers within an on-device Trusted Execution Environment (TEE) against potential white-box membership inference attacks without the significant computational overhead. Fan Mo 0004, Ali Shahin Shamsabadi, Kleomenis Katevas, Andrea Cavallaro, Hamed Haddadi 0001 |
CCS | 4 |
| 2019 | Accurate Target Annotation in 3D from Multimodal StreamsabstractAccurate annotation is fundamental to quantify the performance of multi-sensor and multi-modal object detectors and trackers. However, invasive or expensive instrumentation is needed to automatically generate these annotations. To mitigate this problem, we present a multi-modal approach that leverages annotations from reference streams (e.g. individual camera views) and measurements from unannotated additional streams (e.g. audio) to infer 3D trajectories through an optimization. The core of our approach is a multi-modal extension of Bundle Adjustment with a cross-modal correspondence detection that selectively uses measurements in the optimization. We apply the proposed approach to fully annotate a new multi-modal and multi-view dataset for multi-speaker 3D tracking. Oswald Lanz, Alessio Brutti, Alessio Xompero, Xinyuan Qian 0001, Maurizio Omologo, Andrea Cavallaro |
ICASSP | 6 |
| 2019 | Scene Privacy ProtectionabstractImages shared on social media are routinely analysed by classifiers for content annotation and user profiling. These automatic inferences reveal to the service provider sensitive information that a naive user might want to keep private. To address this problem, we present a method designed to distort the image data so as to hinder the inference of a classifier without affecting the utility for social media users. The proposed approach is based on the Fast Gradient Sign Method (FGSM) and limits the likelihood that automatic inference can expose the true class of a distorted image. Experimental results on a scene classification task show that the proposed method, private FGSM, achieves a desirable trade-off between the drop in classification accuracy and the distortion on the private classes of the Places365-Standard dataset using ResNet50. The classifier is misled 94.40% of the times in the top-5 classes with only a small average reduction of three image quality measures (SSIM, PSNR, BRISQUE). Chau Yi Li, Ali Shahin Shamsabadi, Ricardo Sanchez-Matilla, Riccardo Mazzon, Andrea Cavallaro |
ICASSP | 5 |
| 2019 | Robust Compressive Sensing of Multiband Spectrum with Partial and Incorrect PriorsabstractCompressive sensing has been applied in wideband spectrum sensing to achieve sub-Nyquist sampling. Prior information of the multiband spectrum occupancy, e.g. from geo-location database, can be utilized by compressive spectrum sensing (CSS) to enhance the sensing performance. However, these priors are prone to be partially missing and may also contain incorrect information. We hereby propose a CSS scheme aided by priors and robust to priors imperfections, and moreover, a novel and practical algorithm to provide robust channel sparsity estimation needed by the CSS scheme. Simulations show prominent enhancement of detection performance and lower iteration counts by employing priors in the proposed CSS scheme. Jiadong Yu, Andrea Cavallaro, Yue Gao 0001 |
ICC | 3 |
| 2019 | View-LSTM: Novel-View Video Synthesis Through View DecompositionabstractWe tackle the problem of synthesizing a video of multiple moving people as seen from a novel view, given only an input video and depth information or human poses of the novel view as prior. This problem requires a model that learns to transform input features into target features while maintaining temporal consistency. To this end, we learn an invariant feature from the input video that is shared across all viewpoints of the same scene and a view-dependent feature obtained using the target priors. The proposed approach, View-LSTM, is a recurrent neural network structure that accounts for the temporal consistency and target feature approximation constraints. We validate View-LSTM by designing an end-to-end generator for novel-view video synthesis. Experiments on a large multi-view action recognition dataset validate the proposed model. Mohamed Ilyes Lakhal, Oswald Lanz, Andrea Cavallaro |
ICCV | 3 |
| 2019 | Omni-Scale Feature Learning for Person Re-IdentificationabstractAs an instance-level recognition problem, person re-identification (ReID) relies on discriminative features, which not only capture different spatial scales but also encapsulate an arbitrary combination of multiple scales. We callse features of both homogeneous and heterogeneous scales omni-scale features. In this paper, a novel deep ReID CNN is designed, termed Omni-Scale Network (OSNet), for omni-scale feature learning. This is achieved by designing a residual block composed of multiple convolutional feature streams, each detecting features at a certain scale. Importantly, a novel unified aggregation gate is introduced to dynamically fuse multi-scale features with input-dependent channel-wise weights. To efficiently learn spatial-channel correlations and avoid overfitting, the building block uses both pointwise and depthwise convolutions. By stacking such blocks layer-by-layer, our OSNet is extremely lightweight and can be trained from scratch on existing ReID benchmarks. Despite its small model size, our OSNet achieves state-of-the-art performance on six person-ReID datasets. Code and models are available at: https://github.com/KaiyangZhou/deep-person-reid. Kaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao Xiang 0002 |
ICCV | 3 |
| 2019 | Learnable Masks for Pose-Guided View SynthesisabstractPose-guided human view synthesis uses a target pose to generate the appearance of a new view of a person. The input view and the target pose can be processed separately with UNet architectures that combine the results in a late fusion stage. UNet architectures link their encoder and decoder with skip connections that preserve the location of spatial features by injecting input information in the decoding process. However, direct skip connections may transfer irrelevant information to the decoder. We overcome this limitation with learnable masks for skip connections that encourage the decoder to use only relevant information from the encoder. We show that adding the proposed mask to UNet architectures improves the performance of view synthesis with only a slight increase in inference time. Mohamed Ilyes Lakhal, Oswald Lanz, Andrea Cavallaro |
ICIP | 3 |
| 2019 | A Predictor of Moving Objects for First-Person VisionabstractPredicting the motion of objects captured by a moving camera is important for first-person vision tasks. In this paper, we present an accurate model to forecast the position of moving objects by disentangling global and object motion without the need of camera calibration or planarity assumptions. Our predictor uses past observations to model online the motion of objects by selectively tracking a spatially balanced set of keypoints and estimating scene transformations between pairs of frames. We show that we can forecast up to 60% more accurately than state-of-the-art predictors while being resilient to noisy observations. Moreover, the proposed predictor is robust to frame-rate reduction and outperforms alternative approaches while processing only 33% of the frames with moving cameras. We also show the benefit of integrating the proposed predictor in a multi-object tracker. Ricardo Sanchez-Matilla, Andrea Cavallaro |
ICIP | 2 |
| 2019 | Learning Action Representations for Self-supervised Visual ExplorationabstractLearning to efficiently navigate an environment using only an on-board camera is a difficult task for an agent when the final goal is far from the initial state and extrinsic rewards are sparse. To address this problem, we present a self-supervised prediction network to train the agent with intrinsic rewards that relate to achieving the desired final goal. The network learns to predict its future camera view (the future state) from a current state-action pair through an Action Representation Module that decodes input actions as higher dimensional representations. To increase the representational power of the network during exploration we fuse the responses from the Action Representation Module in the transition network, which predicts the future state. Moreover, to enhance the discrimination capability between predictions from different input actions we introduce joint regression and triplet ranking loss functions. We show that, despite the sparse extrinsic rewards, by learning action representations we achieve a faster training convergence than state-of-the-art methods with only a small increase in the number of the model parameters. Changjae Oh, Andrea Cavallaro |
ICRA | 2 |
| 2019 | Audio-visual sensing from a quadcopter: dataset and baselines for source localization and sound enhancementabstractWe present an audio-visual dataset recorded outdoors from a quadcopter and discuss baseline results for multiple applications. The dataset includes a scenario for source localization and sound enhancement with up to two static sources, and a scenario for source localization and tracking with a moving sound source. These sensing tasks are made challenging by the strong and time-varying ego-noise generated by the rotating motors and propellers. The dataset was collected using a small circular array with 8 microphones and a camera mounted on the quadcopter. The camera view was used to facilitate the annotation of the sound-source positions and can also be used for multi-modal sensing tasks. We discuss the audio-visual calibration procedure that is needed to generate the annotation for the dataset, which we make available to the research community. Lin Wang 0009, Ricardo Sanchez-Matilla, Andrea Cavallaro |
IROS | 3 |
| 2019 | ConflictNET: End-to-End Learning for Speech-Based Conflict Intensity EstimationabstractComputational paralinguistics aims to infer human emotions, personality traits and behavioural patterns from speech signals. In particular, verbal conflict is an important example of human-interaction behaviour, whose detection would enable monitoring and feedback in a variety of applications. The majority of methods for detection and intensity estimation of verbal conflict apply off-the-shelf classifiers/regressors to generic hand-crafted acoustic features. Generating conflict-specific features requires refinement steps and the availability of metadata, such as the number of speakers and their speech overlap duration. Moreover, most techniques treat feature extraction and regression as independent modules, which require separate training and parameter tuning. To address these limitations, we propose the first end-to-end convolutional-recurrent neural network architecture that learns conflict-specific features directly from raw speech waveforms, without using explicit domain knowledge or metadata. Additionally, to selectively focus the model on portions of speech containing verbal conflict instances, we include a global attention interface that learns the alignment between layers of the recurrent network. Experimental results on the SSPNet Conflict Corpus show that our end-to-end architecture achieves state-of-the-art performance in terms of Pearson Correlation Coefficient. Vandana Rajan, Alessio Brutti, Andrea Cavallaro |
IEEE Signal Process. Lett. | 3 |
| 2019 | Multi-Speaker Tracking From an Audio-Visual Sensing DeviceabstractCompact multi-sensor platforms are portable and thus desirable for robotics and personal-assistance tasks. However, compared to physically distributed sensors, the size of these platforms makes person tracking more difficult. To address this challenge, we propose a novel 3-D audio-visual people tracker that exploits visual observations (object detections) to guide the acoustic processing by constraining the acoustic likelihood on the horizontal plane defined by the predicted height of a speaker. This solution allows the tracker to estimate, with a small microphone array, the distance of a sound. Moreover, we apply a color-based visual likelihood on the image plane to compensate for misdetections. Finally, we use a 3-D particle filter and greedy data association to combine visual observations, color-based, and acoustic likelihoods to track the position of multiple simultaneous speakers. We compare the proposed multimodal 3-D tracker against two state-of-the-art methods on the AV16.3 dataset and on a newly collected dataset with co-located sensors, which we make available to the research community. Experimental results show that our multimodal approach outperforms the other methods both in 3-D and on the image plane. Xinyuan Qian 0001, Alessio Brutti, Oswald Lanz, Maurizio Omologo, Andrea Cavallaro |
IEEE Trans. Multim. | 5 |
| 2019 | Guest Editorial Trustworthiness in Social Multimedia Analytics and DeliveryabstractThe papers in this special issue focus on trustworthiness in multimedia communications. Recently, social multimedia content is being delivered to users with a high quality of experience (QoE) with the advance of multimedia technologies and social networks. However, as a huge amount of social users have various demands to exchange and share multimedia content with each other, it becomes a new challenge for the current social multimedia analytics and delivery to deal with the various attacks perpetrated by malicious users or through spam contents. Therefore, the trust and risk management for social multimedia content based on the social tie of users become of prime importance to face the unpredicted threats and subsequent damage. This Special Section aims to provide a premier forum for researchers working on the trust-based social multimedia analytics and delivery. It also provides the opportunity for both academic and industrial researchers to discuss recent results and provide solutions to the above-mentioned challenges. Zhou Su 0001, Qing Fang, Sanjeev Mehrotra, Ali C. Begen, Qiang Ye 0001, Andrea Cavallaro |
IEEE Trans. Multim. | 7 |
| 2018 | Video Summarisation by Classification with Deep Reinforcement Learning
Kaiyang Zhou, Tao Xiang 0002, Andrea Cavallaro |
BMVC | 3 |
| 2018 | Learning Switching Models for Abnormality Detection for Autonomous DrivingabstractWe present an approach to learn a model to estimate the dynamical states at continuous and discrete inference levels when trajectory information is available. We learn from sparse data a probabilistic switching model that generates trajectories associated with a stationary plan of an agent. The learned generative model is used within a Markov Jump Linear System (MJLSs) to switch among set of space dependent linear filters that analyze new trajectories and detect deviations from the learned model based on internal innovation measurements. We show examples of application of the proposed approach to learn filters for evaluating deviations from a reference human driving task execution that includes static and dynamic obstacle avoidance. Mohamad Baydoun, Damian Campo, Valentina Sanguineti, Lucio Marcenaro, Andrea Cavallaro, Carlo S. Regazzoni |
FUSION | 5 |
| 2018 | Multi-Camera Matching of Spatio-Temporal Binary FeaturesabstractLocal image features are generally robust to different geometric and photometric transformations on planar surfaces or under narrow baseline views. However, the matching performance decreases considerably across cameras with unknown poses separated by a wide baseline. To address this problem, we accumulate temporal information within each view by tracking local binary features, which encode intensity comparisons of pixel pairs in an image patch. We then encode the spatio-temporal features into fixed-length binary descriptors by selecting temporally dominant binary values. We complement the descriptor with a binary vector that identifies intensity comparisons that are temporally unstable. Finally, we use this additional vector to ignore the corresponding binary values in the fixed-length binary descriptor when matching the features across cameras. We analyse the performance of the proposed approach and compare it with baselines. Alessio Xompero, Oswald Lanz, Andrea Cavallaro |
FUSION | 3 |
| 2018 | a Multi-Perspective Approach to Anomaly Detection for Self -Aware Embodied AgentsabstractThis paper focuses on multi-sensor anomaly detection for moving cognitive agents using both external and private first-person visual observations. Both observation types are used to characterize agents motion in a given environment. The proposed method generates locally uniform motion models by dividing a Gaussian process that approximates agents displacements on the scene and provides a Shared Level (SL) self-awareness based on Environment Centered (EC) models. Such models are then used to train in a semi-unsupervised way a set of Generative Adversarial Networks (GANs) that produce an estimation of external and internal parameters of moving agents. Obtained results exemplify the feasibility of using multi-perspective data for predicting and analyzing trajectory information. Mohamad Baydoun, Mahdyar Ravanbakhsh, Damian Campo, Pablo Marín-Plaza, David Martín 0001, Lucio Marcenaro, Andrea Cavallaro, Carlo S. Regazzoni |
ICASSP | 7 |
| 2018 | 3D Mouth Tracking from a Compact Microphone Array Co-Located with a cameraabstractWe address the 3D audio-visual mouth tracking problem when using a compact platform with co-located audio-visual sensors, without a depth camera. In particular, we propose a multi-modal particle filter that combines a face detector and 3D hypothesis mapping to the image plane. The audio likelihood computation is assisted by video, which relies on a GCC-PHAT based acoustic map. By combining audio and video inputs, the proposed approach can cope with a reverberant and noisy environment, and can deal with situations when the person is occluded, outside the Field of View (FoV), or not facing the sensors. Experimental results show that the proposed tracker is accurate both in 3D and on the image plane. Xinyuan Qian 0001, Alessio Xompero, Andrea Cavallaro, Alessio Brutti, Oswald Lanz, Maurizio Omologo |
ICASSP | 3 |
| 2018 | Concurrent Target Following with Active Directional SensorsabstractWe propose a collision-avoidance tracker for agents with a directional sensor that aim to maintain a moving target in their field of view. The proposed tracker addresses the view maintenance issue within an Optimal Reciprocal Collision Avoidance (ORCA) framework. Our tracking agents adaptively share the responsibility of avoiding each other and minimise with a smooth actuation the deviation angle from their heading direction to their target. Experimental results with real people trajectories from public datasets show that the proposed method improves view maintenance. Yiming Wang 0002, Andrea Cavallaro |
ICASSP | 2 |
| 2018 | Unsupervised Trajectory Modeling Based on Discrete Descriptors for Classifying Moving Objects in Video SequencesabstractThis paper focuses on modeling and classifying trajectories from video sequences. Location, velocity and time of appearance are considered as features for recognizing and modeling motions of objects. In a training phase, a discretization of the proposed features is performed by using a self-organizing map approach such that a set of clusters (feature vocabulary) is created for describing trajectories. A cluster dissimilarity measure based on a weighted fusion of features facilitates the recognition of trajectory classes in an incremental way. As a result, an unsupervised method for encoding observed motion information and identifying trajectory patterns is proposed in this article. The method is evaluated with real and simulated data. Additionally' comparisons with previous works show the benefits of our method when encoding and identifying motion patterns in video sequences. Damian Campo, Mohamad Baydoun, Lucio Marcenaro, Andrea Cavallaro, Carlo S. Regazzoni |
ICIP | 4 |
| 2018 | Background Light Estimation for Depth-Dependent Underwater Image RestorationabstractLight undergoes a wavelength-dependent attenuation and loses energy along its propagation path in water. In particular, the absorption of red wavelengths is greater than that of green and blue wavelengths in open ocean waters. This reduces the red intensity of the scene radiance reaching the camera and results in non-uniform light, known as background light, due to the scene depth. Restoration methods that compensate for this colour loss often assume constant background light and distort the colour of the water region(s). To address this problem, we propose a restoration method that compensates for the colour loss due to the scene-to-camera distance of non-water regions without altering the colour of pixels representing water. This restoration is achieved by ensuring background light candidates are selected from pixels representing water and then estimating the non-uniform background light without prior knowledge of the scene depth. Experimental results shows that the proposed approach outperforms existing methods in preserving the colour of water regions. Chau Yi Li, Andrea Cavallaro |
ICIP | 2 |
| 2018 | Confidence Intervals for Tracking Performance ScoresabstractThe objective evaluation of trackers quantifies the discrepancy between tracking results and a manually annotated ground truth. As generating ground truth for a video dataset is tedious and time-consuming, often only keyframes are manually annotated. The annotation between these keyframes is then obtained semi-automatically, for example with linear interpolation. This approximation has two main undesirable consequences: first, interpolated annotations may drift from the actual object, especially with moving cameras; second, trackers that use linear prediction or regularize trajectories with linear interpolation unfairly gain a higher tracking evaluation score. This problem may become even more important when semi-automatically annotated datasets are used to train machine learning modules. To account for these annotation inaccuracies for a given dataset, we identify objects whose annotations are interpolated and propose a simple method that analyzes existing annotations and produces a confidence interval to complement tracking scores. These confidence intervals quantify the uncertainty in the annotation and allow us to appropriately interpret the ranking of trackers with respect to the chosen tracking performance score. Ricardo Sanchez-Matilla, Andrea Cavallaro |
ICIP | 2 |
| 2018 | Distributed One-Class LearningabstractWe propose a cloud-based filter trained to block third parties from uploading privacy-sensitive images of others to online social media. The proposed filter uses Distributed One-Class Learning, which decomposes the cloud-based filter into multiple one-class classifiers. Each one-class classifier captures the properties of a class of privacy-sensitive images with an autoencoder. The multi-class filter is then reconstructed by combining the parameters of the one-class autoen-coders. The training takes place on edge devices (e.g. smartphones) and therefore users do not need to upload their private and/or sensitive images to the cloud. A major advantage of the proposed filter over existing distributed learning approaches is that users cannot access, even indirectly, the parameters of other users. Moreover, the filter can cope with the imbalanced and complex distribution of the image content and the independent probability of addition of new users. We evaluate the performance of the proposed distributed filter using the exemplar task of blocking a user from sharing privacy-sensitive images of other users. In particular, we validate the behavior of the proposed multi-class filter with non-privacy-sensitive images, the accuracy when the number of classes increases, and the robustness to attacks when an adversary user has access to privacy-sensitive images of other users. Ali Shahin Shamsabadi, Hamed Haddadi 0001, Andrea Cavallaro |
ICIP | 3 |
| 2018 | MORB: A Multi-Scale Binary DescriptorabstractLocal image features play an important role in matching images under different geometric and photometric transformations. However, as the scale difference across views increases, the matching performance may considerably decrease. To address this problem we propose MORB, a multi-scale binary descriptor that is based on ORB and that improves the accuracy of feature matching under scale changes. MORB describes an image patch at different scales using an oriented sampling pattern of intensity comparisons in a predefined set of pixel pairs. We also propose a matching strategy that estimates the cross-scale match between MORB descriptors across views. Experiments show that MORB outperforms state-of-the-art binary descriptors under several transformations. Alessio Xompero, Oswald Lanz, Andrea Cavallaro |
ICIP | 3 |
| 2018 | Tracking a moving sound source from a multi-rotor droneabstractWe propose a method to track from a multi-rotor drone a moving source, such as a human speaker or an emergency whistle, whose sound is mixed with the strong ego-noise generated by rotating motors and propellers. The proposed method is independent of the specific drone and does not need pre-training nor reference signals. We first employ a time-frequency spatial filter to estimate, on short audio segments, the direction of arrival of the moving source and then we track these noisy estimations with a particle filter. We quantitatively evaluate the results using a ground-truth trajectory of the sound source obtained with an on-board camera and compare the performance of the proposed method with baseline solutions. Lin Wang 0009, Ricardo Sanchez-Matilla, Andrea Cavallaro |
IROS | 3 |
| 2018 | A distributed vision-based consensus model for aerial-robotic teamsabstractWe present a distributed model for a team of autonomous aerial robots to collaboratively track a target without external control. The model uses distributed consensus to coordinate actions and to maintain formation via geometric constraints. Each robot uses its ego-centric view of a target and the relative distance from its two closest neighbors to infer its steering commands. To account for noisy and missing target detections, the robots exchange their estimated target position and formation configuration through shared PID-controlled steering responses. We show that the proposed model enables the team to maintain the view of a maneuvering target with varying acceleration under noisy detections and failures up to situations when all robots but one lose the target from their field of view. Fabio Poiesi, Andrea Cavallaro |
IROS | 2 |
| 2018 | Temporally Smooth Privacy-Protected Airborne VideosabstractRecreational videography from small drones can capture bystanders who may be uncomfortable about appearing in those videos. Existing privacy filters, such as scrambling and hopping blur, address this issue through de-identification but generate temporal distortions that manifest themselves as flicker. To address this problem, we present a robust spatiotemporal hopping blur filter that protects privacy through de-identification of face regions. The proposed filter is meant for on-board installation and produces temporally smooth and pleasant videos. We apply hopping blur to protect each frame against identification attacks, and minimise artefacts and flicker introduced by the hopping blur. We evaluate the proposed filter against different identification attacks and by assessing the quality of the resulting videos using a subjective test and objective measures. Omair Sarwar, Andrea Cavallaro, Bernhard Rinner |
IROS | 2 |
| 2018 | Underwater image and video dehazing with pure haze region segmentationabstractUnderwater scenes captured by cameras are plagued with poor contrast and a spectral distortion, which are the result of the scattering and absorptive properties of water. In this paper we present a novel dehazing method that improves visibility in images and videos by detecting and segmenting image regions that contain only water. The colour of these regions, which we refer to as pure haze regions, is similar to the haze that is removed during the dehazing process. Moreover, we propose a semantic white balancing approach for illuminant estimation that uses the dominant colour of the water to address the spectral distortion present in underwater scenes. To validate the results of our method and compare them to those obtained with state-of-the-art approaches, we perform extensive subjective evaluation tests using images captured in a variety of water types and underwater videos captured onboard an underwater vehicle. Simon Emberton, Lars Chittka, Andrea Cavallaro |
Comput. Vis. Image Underst. | 3 |
| 2018 | Pseudo-Determined Blind Source Separation for Ad-hoc Microphone NetworksabstractWe propose a pseudo-determined blind source separation framework that exploits the information from a large number of microphones in an ad-hoc network to extract and enhance sound sources in a reverberant scenario. After compensating for the time offsets and sampling rate mismatch between (asynchronous) signals, we interpret as a determined M × M mixture the over-determined M × N mixture, where M ) N is the number of microphones and N is the number of sources. Next, we propose a pseudodetermined mixture model that can apply an M × M independent component analysis (ICA) directly to the M-channel recordings. Moreover, we propose a reference-based permutation alignment scheme that aligns the permutation of the ICA outputs and classifies them into target channels, which contain the N sources, and nontarget channels, which contain reverberation residuals. Finally, using the signals from nontarget channels, we estimate in each target channel the power spectral density of the noise component that we suppress with a spectral postfilter. Interestingly, we also obtain late-reverberation suppression as byproduct. Experiments show that each processing block improves incrementally source separation and that the performance of the proposed pseudodetermined separation improves as the number of microphones increases. Lin Wang 0009, Andrea Cavallaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Cooperative Robots to Observe Moving Targets: ReviewabstractThe deployment of multiple robots for achieving a common goal helps to improve the performance, efficiency, and/or robustness in a variety of tasks. In particular, the observation of moving targets is an important multirobot application that still exhibits numerous open challenges, including the effective coordination of the robots. This paper reviews control techniques for cooperative mobile robots monitoring multiple targets. The simultaneous movement of robots and targets makes this problem particularly interesting, and our review systematically addresses this cooperative multirobot problem for the first time. We classify and critically discuss the control techniques: cooperative multirobot observation of multiple moving targets, cooperative search, acquisition, and track, cooperative tracking, and multirobot pursuit evasion. We also identify the five major elements that characterize this problem, namely, the coordination method, the environment, the target, the robot and its sensor(s). These elements are used to systematically analyze the control techniques. The majority of the studied work is based on simulation and laboratory studies, which may not accurately reflect real-world operational conditions. Importantly, while our systematic analysis is focused on multitarget observation, our proposed classification is useful also for related multirobot applications. Asif Khan 0003, Bernhard Rinner, Andrea Cavallaro |
IEEE Trans. Cybern. | 3 |
| 2018 | Constrained Optimization for Plane-Based StereoabstractDepth and surface normal estimation are crucial components in understanding 3D scene geometry from calibrated stereo images. In this paper, we propose visibility and disparity magnitude constraints for slanted patches in the scene. These constraints can be used to associate geometrically feasible planes with each point in the disparity space. The new constraints are validated in the PatchMatch Stereo framework. We use these new constraints not only for initialization, but also in the local plane refinement step of this iterative algorithm. The proposed constraints increase the probability of estimating correct plane parameters, and lead to an improved 3D reconstruction of the scene. Furthermore, the proposed constrained initialization reduces the number of iterations before convergence to the optimal plane parameters. In addition, as most stereo image pairs are not perfectly rectified, we modify the view propagation process by assigning the plane parameters to the neighbors of the candidate pixel. To update the plane parameters in the plane refinement step, we use a gradient free non-linear optimizer. The benefits of the new initialization, propagation, and refinement schemes are demonstrated. Shahnawaz Ahmed, Miles E. Hansard, Andrea Cavallaro |
IEEE Trans. Image Process. | 3 |
| 2018 | Layered Scene Models from Single Hazy ImagesabstractThis paper describes the construction of a layered scene model, based on a single hazy image that has sufficient depth variation. A depth map and radiance image are estimated by standard dehazing methods. The radiance image is then segmented into a small number of clusters, and a corresponding scene plane is estimated for each. This provides the basic structure of a layered scene model, without the need for multiple views, or image correspondences. We show that problems of gap filling and depth blending can be addressed systematically, with respect to the layered depth structure. The final models, which resemble cardboard 'pop-ups', are visually convincing. An implementation is described, and subjective depth preferences are tested in a psychophysical experiment. Lingyun Zhao 0001, Miles E. Hansard, Andrea Cavallaro |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2017 | Modeling and classification of trajectories based on a Gaussian process decomposition into discrete componentsabstractWe present a method to model and classify trajectory data that come from surveillance videos. Observations of the locations of moving entities are used to estimate their expected velocity in the scene. Such estimation is performed by a Gaussian process regression that enables to approximate probabilistically the expected velocity of entities given some observed evidence in the scene. Subsequently, regions where estimations have high certainty are decomposed into zones by superpixel segmentation. Each zone represents a region where motions of entities can be explained by a quasilinear dynamical model. We evaluated the proposed method with two datasets and confirmed its reliability for characterizing and classifying trajectories. Damian Campo, Mohamad Baydoun, Lucio Marcenaro, Andrea Cavallaro, Carlo S. Regazzoni |
AVSS | 4 |
| 2017 | A batch asynchronous tracker for wireless smart-camera networksabstractDistributed tracking in wireless smart-camera networks is affected by varying local processing delays that generally depend on the current scene complexity. As a consequence, each camera makes target information available to the network at different time instants. These unknown delays compound the drifts caused by local clocks and may induce tracking failures when target information is fused. To address this problem, we propose a distributed batch asynchronous tracker for fully connected wireless smart-camera networks. The cameras use the information filter to estimate the target state information and to predict corresponding information of other cameras based on the asynchronous information received from them. Finally, the temporally aligned information is fused. We show that the proposed approach achieves higher tracking accuracy than the state of the art under varying degrees of asynchronism. Sandeep Katragadda, Andrea Cavallaro |
AVSS | 2 |
| 2017 | Active visual tracking in multi-agent scenariosabstractWe propose an active visual tracker with collision avoidance for camera-equipped robots in dense multi-agent scenarios. The objective of each tracking agent (robot) is to maintain visual fixation on its moving target while updating its velocity to avoid other agents. However, when multiple robots are present or targets intensively intersect each other, robots may have no accessible collision-avoiding paths. We address this problem with an adaptive mechanism that sets the pair-wise responsibilities to increase the total accessible collision-avoiding controls. The final collision-avoiding control accounts for motion smoothness and view performance, i.e. maintaining the target centered in the field of view and at a certain size. We validate the proposed approach under different target-intersecting scenarios and compare it with the Optimal Reciprocal Collision Avoidance and the Reciprocal Velocity Obstacle methods. Yiming Wang 0002, Andrea Cavallaro |
AVSS | 2 |
| 2017 | Average consensus-based asynchronous trackingabstractTarget tracking in a network of wireless cameras may fail if data are captured or exchanged asynchronously. Unlike traditional sensor networks, video processing may generate significant delays that also vary from camera to camera. Moreover, the continuous and rapid change of the dynamics of the consensus variable (the target state) makes tracking even more challenging under these conditions. To address this problem, we propose a consensus approach that enables each camera to predict information of other cameras with respect to its own capturing time-stamp based on the received information. This prediction is key to compensate for asynchronous data exchanges. Simulations show the performance improvement with the proposed approach compared to the state of the art in the presence of asynchronous frame captures and random processing delays. Sandeep Katragadda, Carlo S. Regazzoni, Andrea Cavallaro |
ICASSP | 3 |
| 2017 | 3D audio-visual speaker tracking with an adaptive particle filterabstractWe propose an audio-visual fusion algorithm for 3D speaker tracking from a localised multi-modal sensor platform composed of a camera and a small microphone array. After extracting audio-visual cues from individual modalities we fuse them adaptively using their reliability in a particle filter framework. The reliability of the audio signal is measured based on the maximum Global Coherence Field (GCF) peak value at each frame. The visual reliability is based on colour-histogram matching with detection results compared with a reference image in the RGB space. Experiments on the AV16.3 dataset show that the proposed adaptive audio-visual tracker outperforms both the individual modalities and a classical approach with fixed parameters in terms of tracking accuracy. Xinyuan Qian 0001, Alessio Brutti, Maurizio Omologo, Andrea Cavallaro |
ICASSP | 4 |
| 2017 | Time-frequency processing for sound source localization from a micro aerial vehicleabstractWe address the problem of sound source localization with a microphone array mounted on a micro aerial vehicle (MAV). Due to the noise generated by motors and propellers, this scenario is characterized by extremely low signal-to-noise ratios (SNR). Based on the observation that the energy of MAV sound recordings is usually concentrated at isolated time-frequency bins, we propose a time-frequency processing framework to address this problem. We first estimate the direction of arrival of the sound at individual time-frequency bins. Then we formulate a set of spatially informed filters pointing at candidate directions in the search space. The output of the filtering tends to present high non-Gaussianity when the spatial filter is steered towards the target sound source. Finally, by measuring the non-Gaussianity of the spatial filtering outputs we build a spatial likelihood function from which we estimate the direction of the target sound. Experimental results with real-recorded MAV ego-noise show the superiority of the proposed method over the state of the art in performing source localization robustly. Lin Wang 0009, Andrea Cavallaro |
ICASSP | 2 |
| 2017 | Efficient estimation of target detection qualityabstractThe capability of determining the quality of target detections is important for applications using smart cameras, such as autonomous robotics and surveillance. We propose to estimate the quality of target detections by integrating the target location uncertainty over polygonal domains, which represent the fields of view of the cameras. We define a framework based on numerical integration that easily accommodates multiple models for uncertainty and fields of view. We perform quadrature-based integration combined with importance sampling to provide accurate quality estimations while reducing the computational cost. The proposed method outperforms alternative approaches in terms of estimation accuracy and execution time. We validate the proposed approach with a recent distributed multi-camera multi-target tracker and improved it by considering realistic fields of view. Results demonstrate the effectiveness of the proposed method in decreasing the state estimation error. Juan C. SanMiguel, Andrea Cavallaro |
ICIP | 2 |
| 2017 | Privacy Protection in Online MultimediaabstractOnline multimedia has been growing rapidly due to ubiquitous mobile phones, widely deployed surveillance cameras, dashcams and mini-drones. When one takes photographs or videos at a public location, it is highly likely that some other people ("bystanders") also appear in the visual data. The data may be available online, such as shared by social media, and questions about privacy arise. This panel discusses the issues about privacy in online multimedia from legal, technological, and social aspects. Yung-Hsiang Lu, Andrea Cavallaro, Catherine Crump, Gerald Friedland, Keith Winstein |
ACM Multimedia | 2 |
| 2017 | Multi-Modal Localization and Enhancement of Multiple Sound Sources from a Micro Aerial VehicleabstractThe ego-noise generated by the motors and propellers of a micro aerial vehicle (MAV) masks the environmental sounds and considerably degrades the quality of the on-board sound recording. Sound enhancement approaches generally require knowledge of the direction of arrival of the target sound sources, which are difficult to estimate due to the low signal-to-noise-ratio (SNR) caused by the ego-noise and the interferences between multiple sources. To address this problem, we propose a multi-modal analysis approach that jointly exploits audio and video to enhance the sounds of multiple targets captured from an MAV equipped with a microphone array and a video camera. We first address audio-visual calibration via camera resectioning, audio-visual temporal alignment and geometrical alignment to jointly use the features in the audio and video streams, which are independently generated. The spatial information from the video is used to assist sound enhancement by tracking multiple potential sound sources with a particle filter. Then we infer the directions of arrival of the target sources from the video tracking results and extract the sound from the desired direction with a time-frequency spatial filter, which suppresses the ego-noise by exploiting its time-frequency sparsity. Experimental demonstration results with real outdoor data verify the robustness of the proposed multi-modal approach for multiple speakers in extremely low-SNR scenarios. Ricardo Sanchez-Matilla, Lin Wang 0009, Andrea Cavallaro |
ACM Multimedia | 3 |
| 2017 | Hierarchical modeling for first-person vision activity recognition
Girmaw Abebe, Andrea Cavallaro |
Neurocomputing | 2 |
| 2017 | Multi-Tracker Partition FusionabstractWe propose a decision-level approach to fuse the output of multiple trackers based on their estimated individual performance. The proposed approach is composed of three main steps. First, we group trackers into clusters based on the spatiotemporal pair-wise correlation of their short-term trajectories. Then, we evaluate performance based on reverse-time analysis with an adaptive reference frame and define the cluster with trackers that appear to be successfully following the target as the on-target cluster. Finally, the state estimations produced by trackers in the on-target cluster are fused to obtain the target state. The proposed fusion approach uses standard tracker outputs and can therefore combine various types of trackers. We tested the proposed approach with several combinations of state-of-the-art trackers and also compared it with individual trackers and other fusion approaches. The results show that the proposed approach improves the state estimation accuracy under multiple tracking challenges. ObaidUllah Khalid, Juan C. SanMiguel, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Support Vector Motion ClusteringabstractWe present a closed-loop unsupervised clustering method for motion vectors extracted from highly dynamic video scenes. Motion vectors are assigned to nonconvex homogeneous clusters characterizing direction, size and shape of regions with multiple independent activities. The proposed method is based on support vector clustering. Cluster labels are propagated over time via incremental learning. The proposed method uses a kernel function that maps the input motion vectors into a high-dimensional space to produce nonconvex clusters. We improve the mapping effectiveness by quantifying feature similarities via a blend of position and orientation affinities. We use the Quasiconformal Kernel Transformation to boost the discrimination of outliers. The temporal propagation of the clusters’ identities is achieved via incremental learning based on the concept of feature obsolescence to deal with appearing and disappearing features. Moreover, we design an online clustering performance prediction algorithm used as a feedback that refines the cluster model at each frame in an unsupervised manner. We evaluate the proposed method on synthetic data sets and real-world crowded videos and show that our solution outperforms state-of-the-art approaches. Isah Abdullahi Lawal, Fabio Poiesi, Davide Anguita, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Energy Consumption Models for Smart Camera NetworksabstractCamera networks require heavy visual data processing and high-bandwidth communication. In this paper, we identify key factors underpinning the development of resource-aware algorithms and we propose a comprehensive energy consumption model for the resources employed by smart camera networks, which are composed of cameras that process data locally and collaborate with their neighbors. We account for the main parameters that influence consumption when sensing (frame size and frame rate), processing (dynamic frequency scaling and task load), and communication (output power and bandwidth) are considered. Next, we define an abstraction based on clock frequency and duty cycle that accounts for active, idle, and sleep operational states. We demonstrate the importance of the proposed model for a multicamera tracking task and show how one may significantly reduce consumption with only minor performance degradation when choosing to operate with an appropriately reduced hardware capacity. Moreover, we quantify the dependency on local computation resources and bandwidth availability. The proposed consumption model can be easily adjusted to account for new platforms, thus providing a valuable tool for the design of resource-aware algorithms and further research in resource-aware camera networks. Juan C. SanMiguel, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Guest Editorial Introduction to the Special Issue on Group and Crowd Behavior Analysis for Intelligent Multicamera Video SurveillanceabstractDespite significant progress in human behavior analysis over the past few years, most of today’s state-of-the-art algorithms focus on analyzing individual behavior in a simple environment monitored by a single camera. Recently, the widespread availability of cameras and a growing need for public safety have shifted the attention of researchers in video surveillance from individual behavior analysis to group and crowd behavior analysis in multicamera networks. Group behavior analysis provides a novel level for describing events, which are semantically more meaningful, highlighting barely visible relational connections among people. Crowd behavior analysis can also be used for anomaly detection such as panic scenarios, dangerous situations, and illegal behaviors in public spaces. Hongxun Yao, Andrea Cavallaro, Thierry Bouwmans, Zhengyou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Online Cross-Modal Adaptation for Audio-Visual Person Identification With Wearable CamerasabstractWe propose an audio-visual target identification approach for egocentric data with cross-modal model adaptation. The proposed approach blindly and iteratively adapts the time-dependent models of each modality to varying target appearance and environmental conditions using the posterior of the other modality. The adaptation is unsupervised and performed online; thus, models can be improved as new unlabeled data become available. In particular, accurate models do not deteriorate when a modality is underperforming thanks to an appropriate selection of the parameters in the adaptation. Importantly, unlike traditional audio-visual integration methods, the proposed approach is also useful for temporal intervals during which only one modality is available or when different modalities are used for different tasks. We evaluate the proposed method in an end-to-end multimodal person identification application with two challenging real-world datasets and show that the proposed approach successfully adapts models in presence of mild mismatch. We also show that the proposed approach is beneficial to other multimodal score fusion algorithms. Alessio Brutti, Andrea Cavallaro |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2017 | Robust Registration of Dynamic Facial SequencesabstractAccurate face registration is a key step for several image analysis applications. However, existing registration methods are prone to temporal drift errors or jitter among consecutive frames. In this paper, we propose an iterative rigid registration framework that estimates the misalignment with trained regressors. The input of the regressors is a robust motion representation that encodes the motion between a misaligned frame and the reference frame(s), and enables reliable performance under non-uniform illumination variations. Drift errors are reduced when the motion representation is computed from multiple reference frames. Furthermore, we use the L2norm of the representation as a cue for performing coarse-to-fine registration efficiently. Importantly, the framework can identify registration failures and correct them. Experiments show that the proposed approach achieves significantly higher registration accuracy than the state-of-the-art techniques in challenging sequences. Evangelos Sariyanidi, Hatice Gunes, Andrea Cavallaro |
IEEE Trans. Image Process. | 3 |
| 2017 | Biologically Inspired Motion Encoding for Robust Global Motion EstimationabstractThe growing use of cameras embedded in autonomous robotic platforms and worn by people is increasing the importance of accurate global motion estimation (GME). However, existing GME methods may degrade considerably under illumination variations. In this paper, we address this problem by proposing a biologically-inspired GME method that achieves high estimation accuracy in the presence of illumination variations. We mimic the early layers of the human visual cortex with the spatio-temporal Gabor motion energy by adopting the pioneering model of Adelson and Bergen and we provide the closed-form expressions that enable the study and adaptation of this model to different application needs. Moreover, we propose a normalisation scheme for motion energy to tackle temporal illumination variations. Finally, we provide an overall GME scheme which, to the best of our knowledge, achieves the highest accuracy on the Pose, Illumination, and Expression (PIE) database. Evangelos Sariyanidi, Hatice Gunes, Andrea Cavallaro |
IEEE Trans. Image Process. | 3 |
| 2017 | Learning Bases of Activity for Facial Expression RecognitionabstractThe extraction of descriptive features from sequences of faces is a fundamental problem in facial expression analysis. Facial expressions are represented by psychologists as a combination of elementary movements known as action units: each movement is localised and its intensity is specified with a score that is small when the movement is subtle and large when the movement is pronounced. Inspired by this approach, we propose a novel data-driven feature extraction framework that represents facial expression variations as a linear combination of localised basis functions, whose coefficients are proportional to movement intensity. We show that the linear basis functions required by this framework can be obtained by training a sparse linear model with Gabor phase shifts computed from facial videos. The proposed framework addresses generalisation issues that are not addressed by existing learnt representations, and achieves, with the same learning parameters, state-of-the-art results in recognising both posed expressions and spontaneous micro-expressions. This performance is confirmed even when the data used to train the model differ from test data in terms of the intensity of facial movements and frame rate. Evangelos Sariyanidi, Hatice Gunes, Andrea Cavallaro |
IEEE Trans. Image Process. | 3 |
| 2016 | Design space exploration for adaptive privacy protection in airborne imagesabstractAirborne cameras on low-flying unmanned vehicles introduce new privacy challenges due to their mobility and viewing angles. In this paper, we focus on face recognition from airborne cameras and explore the design space to determine when a face in an airborne image is inherently protected, that is when an individual is not recognizable. Moreover, when individuals are recognizable by facial recognition algorithms, we propose an adaptive filtering mechanism to lower the face resolution in order to preserve privacy while ensuring a minimum reduction of the fidelity of the image. In particular, we estimate the resolution of faces captured at different altitudes and tilt angles using the data from navigation sensors and ascertain when the captured face is inherently protected. When the face is unprotected, we define a mechanism that automatically configures the strength of a privacy protection filter to improve the trade-off between privacy protection and fidelity of an aerial image or video. Omair Sarwar, Bernhard Rinner, Andrea Cavallaro |
AVSS | 3 |
| 2016 | Prioritized target tracking with active collaborative camerasabstractMobile cameras on robotic platforms can support fixed multi-camera installations to improve coverage and target localization accuracy. We propose a novel collaborative framework for prioritized target tracking that complement static cameras with mobile cameras, which track targets on demand. Upon receiving a request from static cameras, a mobile camera selects (or switches to) a target to track using a local selection criterion that accounts for target priority, view quality and energy consumption. Mobile cameras use a receding horizon scheme to minimize tracking uncertainty as well as energy consumption when planning their path. We validate the proposed framework in simulated realistic scenarios and show that it improves tracking accuracy and target observation time with reduced energy consumption compared to a framework with only static cameras and compared to a state-of-the-art motion strategy. Yiming Wang 0002, Andrea Cavallaro |
AVSS | 2 |
| 2016 | Ear in the sky: Ego-noise reduction for auditory micro aerial vehiclesabstractWe investigate the spectral and spatial characteristics of the ego-noise of a multirotor micro aerial vehicle (MAV) using audio signals captured with multiple onboard microphones and derive a noise model that grounds the feasibility of microphone-array techniques for noise reduction. The spectral analysis suggests that the ego-noise consists of narrowband harmonic noise and broadband noise, whose spectra vary dynamically with the motor rotation speed. The spatial analysis suggests that the ego-noise of a P-rotor MAV can be modeled as P directional noises plus one diffuse noise. Moreover, because of the fixed positions of the microphones and motors, we can assume that the acoustic mixing network of the ego-noise is stationary. We validate the proposed noise model and the stationary mixing assumption by applying blind source separation to multi-channel recordings from both a static and a moving MAV and quantify the signal-to-noise ratio improvement. Moreover, we make all the audio recordings publicly available. Lin Wang 0009, Andrea Cavallaro |
AVSS | 2 |
| 2016 | Detecting tracking errors via forecasting
ObaidUllah Khalid, Andrea Cavallaro, Bernhard Rinner |
BMVC | 2 |
| 2016 | Detection of fast incoming objects with a moving camera
Fabio Poiesi, Andrea Cavallaro |
BMVC | 2 |
| 2016 | Rate-adaptive multicast video streaming from teams of micro aerial vehiclesabstractVideo multicasting from cameras mounted on micro aerial vehicles (MAVs) is desirable for applications such as search and rescue, surveillance and disaster management. Because of the mobility of the video sources and the high data-rate of videos, the transmission rate should be adapted to the task at hand. Rate-adaptive video multicast streaming in 802.11 requires wireless link estimation as well as frequent feedback from multiple receivers. We propose an application layer rate-adaptive video multicast streaming framework using 802.11 adhoc network that is applicable when both the sender and the receiver nodes are mobile. The receiver nodes of a multicast group are dynamically elected based on their changing link conditions to gain feedback. An Application Layer Video Multicast Gateway (ALVM-GW) adapts the transmission rate and the video encoding rate based on the received feedback. Emulation results show that the proposed approach has balanced performance in terms of goodput, delay and packet loss. Raheeb Muzaffar, Vladimir Vukadinovic, Andrea Cavallaro |
ICRA | 3 |
| 2016 | Application-Layer Rate-Adaptive Multicast Video Streaming over 802.11 for Mobile DevicesabstractMulticast video streaming over IEEE 802.11 is unreliable due to the lack of feedback from receivers. High data rates and variable link conditions require feedback from the receivers for link estimation to improve reliability and rate adaptation accordingly. In this paper, we validate on a test platform an application-layer rate-adaptive video multicast streaming framework using an 802.11 ad-hoc network applicable for mobile senders and receivers. Experimental results serve as a proof of concept and show the performance in terms of goodput, delay, packet loss, and received video quality. Raheeb Muzaffar, Evsen Yanmaz, Christian Bettstetter, Andrea Cavallaro |
ACM Multimedia | 4 |
| 2016 | Robust multi-dimensional motion features for first-person vision activity recognitionabstractWe propose robust multi-dimensional motion features for human activity recognition from first-person videos. The proposed features encode information about motion magnitude, direction and variation, and combine them with virtual inertial data generated from the video itself. The use of grid flow representation, per-frame normalization and temporal feature accumulation enhances the robustness of our new representation. Results on multiple datasets demonstrate that the proposed feature representation outperforms existing motion features, and importantly it does so independently of the classifier. Moreover, the proposed multi-dimensional motion features are general enough to make them suitable for vision tasks beyond those related to wearable cameras. Girmaw Abebe, Andrea Cavallaro, Xavier Parra Llanas |
Comput. Vis. Image Underst. | 2 |
| 2016 | ViComp: composition of user-generated videosabstractWe propose ViComp, an automatic audio-visual camera selection framework for composing uninterrupted recordings from multiple user-generated videos (UGVs) of the same event. We design an automatic audio-based cut-point selection method to segment the UGV. ViComp combines segments of UGVs using a rank-based camera selection strategy by considering audio-visual quality and camera selection history. We analyze the audio to maintain audio continuity. To filter video segments which contain visual degradations, we perform spatial and spatio-temporal quality assessment. We validate the proposed framework with subjective tests and compare it with state-of-the-art methods. Sophia Bano, Andrea Cavallaro |
Multim. Tools Appl. | 2 |
| 2016 | Special issue on "Video analytics for audience measurement in retail and digital signage"
Sebastiano Battiato, Andrea Cavallaro, Cosimo Distante |
Pattern Recognit. Lett. | 2 |
| 2016 | An Iterative Approach to Source Counting and Localization Using Two Distant MicrophonesabstractWe propose a time difference of arrival (TDOA) estimation framework based on time-frequency inter-channel phase difference (IPD) to count and localize multiple acoustic sources in a reverberant environment using two distant microphones. The time-frequency (T-F) processing enables exploitation of the nonstationarity and sparsity of audio signals, increasing robustness to multiple sources and ambient noise. For inter-channel phase difference estimation, we use a cost function, which is equivalent to the generalized cross correlation with phase transform (GCC) algorithm and which is robust to spatial aliasing caused by large inter-microphone distances. To estimate the number of sources, we further propose an iterative contribution removal (ICR) algorithm to count and locate the sources using the peaks of the GCC function. In each iteration, we first use IPD to calculate the GCC function, whose highest peak is detected as the location of a sound source; then we detect the T-F bins that are associated with this source and remove them from the IPD set. The proposed ICR algorithm successfully solves the GCC peak ambiguities between multiple sources and multiple reverberant paths. Lin Wang 0009, Tsz-Kin Hon, Joshua D. Reiss, Andrea Cavallaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | Over-Determined Source Separation and Localization Using Distributed MicrophonesabstractWe propose an overdetermined source separation and localization method for a set of M microphones distributed around an unknown number, N <; M, of sources. We reformulate the overdetermined acoustic mixing procedure with a new determined mixing model and apply a determined M × M independent component analysis ('CA) in each frequency bin directly. The reformulated 'CA operates without knowing N and also leads to better separation in reverberant scenarios. To solve the challenging permutation ambiguity problem, we first employ a time activity-based clustering approach to cluster the separated frequency components into M channels. We then propose a remixing procedure to detect and merge channels from the same source. The detection is done by analyzing time and frequency activities, spectral likeliness, and spatial location. To estimate the spatial location, we propose a time-frequency masking-based steered response power algorithm. Simulated and real-data experiments in a very challenging reverberant scenario confirm the effectiveness of the proposed method in obtaining the number of sources, the separated signals, and the location and spatial likelihood of each source. Lin Wang 0009, Joshua D. Reiss, Andrea Cavallaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Dynamic Bayesian Network modeling for self- and cross-correcting trackingabstractWe present a generic formulation of self- and cross-correcting Bayesian trackers using a Dynamic Bayesian Network. Correction operations in a tracker such as parameter tuning, model updates and re-initialization are represented using hidden variables together with the target state and measurement variables in the Dynamic Bayesian network model. The representation allows one to model different self- and cross-correcting tracking frameworks under the same formulation and facilitates comparison and the design of new trackers. The proposed model is demonstrated with three state-of-the-art trackers that are based on different principles to implement online correction of target tracking. Tewodros Atanaw Biresaw, Andrea Cavallaro, Carlo S. Regazzoni |
AVSS | 2 |
| 2015 | Coalition formation for distributed tracking in wireless camera networksabstractWe present a fully distributed framework for multi-target tracking with bandwidth-limited (wireless) camera networks. Cameras self-organize into coalitions to perform the task of distributed target tracking via local interactions. Each camera joins the coalitions based on considerations of marginal utility, which takes into account tracking confidence and communication performance in the neighborhood of the camera. The proposed framework achieves higher tracking accuracy and quicker convergence than decentralized tracking or distributed tracking without coalition formation. Moreover, the communication cost of the proposed framework is considerably reduced compared to distributed tracking without coalition formation and comparable to decentralized tracking as the number of targets increases. Yiming Wang 0002, Andrea Cavallaro |
AVSS | 2 |
| 2015 | Hierarchical rank-based veiling light estimation for underwater dehazingabstractCurrent dehazing approaches are often hindered when scenes contain bright objects which can cause veiling light and transmission estimation methods to fail. This paper introduces a single image dehazing approach for underwater images with novel veiling light and transmission estimation steps which deal with issues arising from bright objects. We use features to hierarchically rank regions of an image and to select the most likely veiling light candidate. A region-based approach is used to find optimal transmission values for areas that suffer from oversaturation. We also locate background regions through superpixel segmentation and clustering, and adapt the transmission values in these regions so to avoid artefacts. We validate the performance of our approach in comparison to the state of the art in underwater dehazing through subjective evaluation and with commonly used quantitative measures. Simon Emberton, Lars Chittka, Andrea Cavallaro |
BMVC | 3 |
| 2015 | Refining graph matching using inherent structure informationabstractWe present a graph matching refinement framework that improves the performance of a given graph matching algorithm. Our method synergistically uses the inherent structure information embedded globally in the active association graph, and locally on each individual graph. The combination of such information reveals how consistent each candidate match is with its global and local contexts. In doing so, the proposed method removes most false matches and improves precision. The validation on standard benchmark datasets demonstrates the effectiveness of our method. Wenzhao Li, Yi-Zhe Song, Andrea Cavallaro |
ICME | 3 |
| 2015 | Multiscale observation of multiple moving targets using Micro Aerial VehiclesabstractThis paper presents a centralized algorithm for multi-scale observation of multiple moving targets using a team of Micro Aerial Vehicles (MAVs). The proposed algorithm is appropriate when MAVs can observe targets at different elevations with the objective of jointly maximizing duration and resolution of observation for each target. The MAVs share the workload using a greedy assignment of locations and targets to MAVs. The proposed algorithm uses a quad-tree data structure to model the movement decisions of MAVs as well as the variable qualities (resolutions) of observations. We consider cases where there is uncertainty in the target observations (i.e., measurement noise), the number of targets is larger than that of the MAVs and the combined field of views (FOVs) of the sensors cannot cover the whole search region. Simulation results confirm the effectiveness of the proposed algorithm. Asif Khan 0003, Bernhard Rinner, Andrea Cavallaro |
IROS | 3 |
| 2015 | Distributed vision-based flying cameras to film a moving targetabstractFormations of camera-equipped quadrotors (flying cameras) have the actuation agility to track moving targets from multiple viewing angles. In this paper we propose an infrastructure-free distributed control method for multiple flying cameras tracking a moving object. The proposed vision-based servoing can deal with noisy and missing target observations, accounts for quadrotor oscillations and does not require an external positioning system. The flight direction of each camera is inferred via geometric derivation, and the formation is maintained by employing a distributed algorithm that uses the target position information on the camera plane and the position of neighboring flying cameras. Simulations show that the proposed solution enables the tracking of a moving target by the cameras flying in formation despite noisy target detections and when the target is outside some of the fields of view. Fabio Poiesi, Andrea Cavallaro |
IROS | 2 |
| 2015 | Gyro-based Camera-motion Detection in User-generated VideosabstractWe propose a gyro-based camera-motion detection method for videos captured with smartphones. First, the delay between the acquisition of video and gyroscope data is estimated using similarities induced by camera motion in the two sensor modalities. Pan, tilt and shake are then detected using the dominant motions and high frequencies in the gyroscope data. Morphological operations are applied to remove outliers and to identify segments with continuous camera-motion. We compare the proposed method with existing methods that use visual or inertial sensor data. Sophia Bano, Andrea Cavallaro, Xavier Parra Llanas |
ACM Multimedia | 2 |
| 2015 | Temporal validation of Particle Filters for video tracking
Juan C. SanMiguel, Andrea Cavallaro |
Comput. Vis. Image Underst. | 2 |
| 2015 | Correlation-based self-correcting tracking
Tewodros Atanaw Biresaw, Andrea Cavallaro, Carlo S. Regazzoni |
Neurocomputing | 2 |
| 2015 | Discovery and organization of multi-camera user-generated videos of the same event
Sophia Bano, Andrea Cavallaro |
Inf. Sci. | 2 |
| 2015 | Audio-visual events for multi-camera synchronization
Anna Llagostera Casanovas, Andrea Cavallaro |
Multim. Tools Appl. | 2 |
| 2015 | Automatic Analysis of Facial Affect: A Survey of Registration, Representation, and RecognitionabstractAutomatic affect analysis has attracted great interest in various contexts including the recognition of action units and basic or non-basic emotions. In spite of major efforts, there are several open questions on what the important cues to interpret facial expressions are and how to encode them. In this paper, we review the progress across a range of affect recognition applications to shed light on these fundamental questions. We analyse the state-of-the-art solutions by decomposing their pipelines into fundamental components, namely face registration, representation, dimensionality reduction and recognition. We discuss the role of these components and highlight the models and new trends that are followed in their design. Moreover, we provide a comprehensive analysis of facial representations by uncovering their advantages and limitations; we elaborate on the type of information they encode and discuss how they deal with the key challenges of illumination variations, registration errors, head-pose variations, occlusions, and identity bias. This survey allows us to identify open issues and to define future directions for designing real-world affect recognition systems. Evangelos Sariyanidi, Hatice Gunes, Andrea Cavallaro |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Audio Fingerprinting for Multi-Device Self-LocalizationabstractWe investigate the self-localization problem of an ad-hoc network of randomly distributed and independent devices in an open-space environment with low reverberation but heavy noise (e.g. smartphones recording videos of an outdoor event). Assuming a sufficient number of sound sources, we estimate the distance between a pair of devices from the extreme (minimum and maximum) time difference of arrivals (TDOAs) from the sources to the pair of devices without knowing the time offset. The obtained inter-device distances are then exploited to derive the geometrical configuration of the network. In particular, we propose a robust audio fingerprinting algorithm for noisy recordings and perform landmark matching to construct a histogram of the TDOAs of multiple sources. The extreme TDOAs can be estimated from this histogram. By using audio fingerprinting features, the proposed algorithm works robustly in very noisy environments. Experiments with free-field simulation and open-space recordings prove the effectiveness of the proposed algorithm. Tsz-Kin Hon, Lin Wang 0009, Joshua D. Reiss, Andrea Cavallaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Tracker-Level Fusion for Robust Bayesian Visual TrackingabstractWe propose a tracker-level fusion framework for robust visual tracking. The framework combines trackers addressing different tracking challenges to improve the overall performance. A novelty of the proposed framework is the inclusion of an online performance measure to identify the track quality level of each tracker so as to guide the fusion. The fusion is then based on appropriately mixing the prior state of the trackers. Moreover, the track-quality level is used to update the target appearance model. We demonstrate the framework with two Bayesian trackers on video sequences with various challenges and show its robustness compared with the independent use of the two individual trackers, and also compared with state-of-the-art trackers that use tracker-level fusion. Tewodros Atanaw Biresaw, Andrea Cavallaro, Carlo S. Regazzoni |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Tracking Multiple High-Density Homogeneous TargetsabstractWe present a framework for multitarget detection and tracking that infers candidate target locations in videos containing a high density of homogeneous targets. We propose a gradient-climbing technique and an isocontor slicing approach for intensity maps to localize targets. The former uses Markov chain Monte Carlo to iteratively fit a shape model onto the target locations, whereas the latter uses the intensity values at different levels to find consistent object shapes. We generate trajectories by recursively associating detections with a hierarchical graph-based tracker on temporal windows. The solution to the graph is obtained with a greedy algorithm that accounts for false-positive associations. The edges of the graph are weighted with a likelihood function based on location information. We evaluate the performance of the proposed framework on challenging datasets containing videos with high density of targets and compare it with six alternative trackers. Fabio Poiesi, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Probabilistic Subpixel Temporal Registration for Facial Expression Analysis
Evangelos Sariyanidi, Hatice Gunes, Andrea Cavallaro |
ACCV (4) | 3 |
| 2014 | Consensus protocols for distributed tracking in wireless camera networks
Sandeep Katragadda, Juan C. SanMiguel, Andrea Cavallaro |
FUSION | 3 |
| 2014 | Parallel particle-PHD filterabstractThe complexity of multi-target tracking grows faster than linearly with the increase of the numbers of objects, thus making the design of real-time trackers a challenging task for scenarios with a large number of targets. The Probability Hypothesis Density (PHD) filter is known to help reducing this complexity. However, this reduction may not suffice in critical situations when the number of targets, dimension of the state vector, clutter conditions and sample rate are high. To address this problem, we propose a parallelization scheme for the particle PHD filter. The proposed scheme exploits the knowledge of mutual interacting targets in the scene to help fragmentation and to reduce the workload of individual processors. We compare the proposed approach with alternative parallelization schemes and discuss its advantages and limitations using the results obtained on two multi-target tracking datasets. Marco Del Coco, Andrea Cavallaro |
ICASSP | 2 |
| 2014 | Low-cost multi-camera object matchingabstractWe propose an object matching approach aimed at smartphone cameras that exploits the well-known concept of local sets of features for object representation. We also enable the temporal alignment of cameras by exploiting the frames of detected objects to group objects appeared in the same time interval for the assignment within each camera. The proposed approach does not need training thus making it suitable for matching during short temporal intervals. We use both outdoor and indoor datasets for the evaluation, and show that the proposed method reduces up to 95% the amount of information to be stored and communicated. Syed Fahad Tahir, Andrea Cavallaro |
ICASSP | 2 |
| 2014 | Trajectory clustering for motion pattern extraction in aerial videosabstractWe present an end-to-end approach for trajectory clustering from aerial videos that enables the extraction of motion patterns in urban scenes. Camera motion is first compensated by mapping object trajectories on a reference plane. Then clustering is performed based on statistics from the Discrete Wavelet Transform coefficients extracted from the trajectories. Finally, motion patterns are identified by distance minimization from the centroids of the trajectory clusters. The experimental validation on four datasets shows the effectiveness of the proposed approach in extracting trajectory clusters. We also make available two new real-world aerial video datasets together with the estimated object trajectories and ground-truth cluster labeling. Tahir Nawaz 0001, Andrea Cavallaro, Bernhard Rinner |
ICIP | 2 |
| 2014 | Assessing tracking assessment measuresabstractWe propose a methodology to quantitatively compare the relative performance of tracking evaluation measures. The proposed methodology is based on determining the probabilistic agreement between tracking result decisions made by measures and those made by humans. We use tracking results on publicly available datasets with different target types and varying challenges, and collect the judgments of 90 skilled, semi-skilled and unskilled human subjects using a web-based performance assessment test. The analysis of the agreements allows us to highlight the variation in performance of the different measures and the most appropriate ones for the various stages of tracking performance evaluation. Tahir Nawaz 0001, Fabio Poiesi, Andrea Cavallaro |
ICIP | 3 |
| 2014 | Camera Localization UsingTrajectories and MapsabstractWe propose a new Bayesian framework for automatically determining the position (location and orientation) of an uncalibrated camera using the observations of moving objects and a schematic map of the passable areas of the environment. Our approach takes advantage of static and dynamic information on the scene structures through prior probability distributions for object dynamics. The proposed approach restricts plausible positions where the sensor can be located while taking into account the inherent ambiguity of the given setting. The proposed framework samples from the posterior probability distribution for the camera position via data driven MCMC, guided by an initial geometric analysis that restricts the search space. A Kullback-Leibler divergence analysis is then used that yields the final camera position estimate, while explicitly isolating ambiguous settings. The proposed approach is evaluated in synthetic and real environments, showing its satisfactory performance in both ambiguous and unambiguous settings. Raúl Mohedano, Andrea Cavallaro, Narciso García |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Cost-Effective Features for Reidentification in Camera NetworksabstractNetworks of smart cameras share large amounts of data to accomplish tasks such as reidentification. We propose a feature-selection method that minimizes the data needed to represent the appearance of objects by learning the most appropriate feature set for the task at hand (person reidentification). The computational cost for feature extraction and the cost for storing the feature descriptor are considered jointly with feature performance to select cost-effective good features. This selection allows us to improve intercamera reidentification while reducing the bandwidth that is necessary to share data across the camera network. We also rank the selected features in the order of effectiveness for the task to enable a further reduction of the feature set by dropping the least effective features when application constraints require this adaptation. We compare the proposed approach with state-of-the-art methods on the iLIDS and VIPeR datasets and show that the proposed approach considerably reduces network traffic due to intercamera feature sharing while keeping the reidentification performance at an equivalent or better level compared with the state of the art. Syed Fahad Tahir, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Multiview Matching of Articulated ObjectsabstractWe address the problem of multiview association of articulated objects observed using possibly moving and hand-held cameras. Starting from trajectory data, we encode the temporal evolution of the objects and perform matching without making assumptions on scene geometry and with only weak assumptions on the field-of-view overlaps. After generating a viewpoint invariant representation using self-similarity matrices, we put in correspondence the spatio-temporal object descriptions using spectral methods on the resulting matching graph. We validate the proposed method on three publicly available real-world datasets and compare it with alternative approaches. Moreover, we present an extensive analysis of the accuracy of the proposed method in different contexts, with varying noise levels on the input data, varying amount of overlap between the fields of view, and varying duration of the available observations. Luca Zini, Francesca Odone, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Measures of Effective Video TrackingabstractTo evaluate multitarget video tracking results, one needs to quantify the accuracy of the estimated target-size and the cardinality error as well as measure the frequency of occurrence of ID changes. In this paper, we survey existing multitarget tracking performance scores and, after discussing their limitations, we propose three parameter-independent measures for evaluating multitarget video tracking. The measures consider target-size variations, combine accuracy and cardinality errors, quantify long-term tracking accuracy at different accuracy levels, and evaluate ID changes relative to the duration of the track in which they occur. We conduct an extensive experimental validation of the proposed measures by comparing them with existing ones and by evaluating four state-of-the-art trackers on challenging real-world publicly-available data sets. The software implementing the proposed measures is made available online to facilitate their use by the research community. Tahir Nawaz 0001, Fabio Poiesi, Andrea Cavallaro |
IEEE Trans. Image Process. | 3 |
| 2014 | Resource Allocation for Personalized Video SummarizationabstractWe propose a hybrid personalized summarization framework that combines adaptive fast-forwarding and content truncation to generate comfortable and compact video summaries. We formulate video summarization as a discrete optimization problem, where the optimal summary is determined by adopting Lagrangian relaxation and convex-hull approximation to solve a resource allocation problem. To trade-off playback speed and perceptual comfort we consider information associated to the still content of the scene, which is essential to evaluate the relevance of a video, and information associated to the scene activity, which is more relevant for visual comfort. We perform clip-level fast-forwarding by selecting the playback speeds from discrete options, which naturally include content truncation as special case with infinite playback speed. We demonstrate the proposed summarization framework in two use cases, namely summarization of broadcasted soccer videos and surveillance videos. Objective and subjective experiments are performed to demonstrate the relevance and efficiency of the proposed method. Fan Chen 0002, Christophe De Vleeschouwer, Andrea Cavallaro |
IEEE Trans. Multim. | 3 |
| 2013 | Detection and tracking of groups in crowdabstractWe propose a method to detect and track interacting people by employing a framework based on a Social Force Model (SFM). The method embeds plausible human behaviors to predict interactions in a crowd by iteratively minimizing the error between predictions and measurements. We model people approaching a group and restrict the group formation based on the relative velocity of candidate group members. The detected groups are then tracked by linking their interaction centers over time using a buffered graph-based tracker. We show how the proposed framework outperforms existing group localization techniques on three publicly available datasets, with improvements of up to 13% on group detection. Riccardo Mazzon, Fabio Poiesi, Andrea Cavallaro |
AVSS | 3 |
| 2013 | Local Zernike Moment Representation for Facial Affect RecognitionabstractLocal representations became popular for facial affect recognition as they efficiently capture the image discontinuities, which play an important role for interpreting facial actions. We propose to use Local Zernike Moments (ZMs) [4] due to their useful and compact description of the image discontinuities and texture. Their main advantage in comparison to well-established alternatives such as Local Binary Patterns (LBPs) [5], is their flexibility in terms of the size and level of detail of the local description. We introduce a local ZM-based representation which involves a non-linear encoding layer (quantisation). The functionality of this layer is mapping similar facial configurations together and increasing compactness. We demonstrate the use of the local ZM-based representation for posed and naturalistic affect recognition on standard datasets, and show its superiority to alternative approaches for both tasks. Contemporary representations are often designed as frameworks consisting of three layers [2]: (Local) feature extraction, non-linear encoding and pooling. Non-linear encoding aims at enhancing the relevance of local features by increasing their robustness against image noise. Pooling describes small spatial neighbourhoods as single entities, ignoring the precise location of the encoded features, and increasing the tolerance against small geometric inconsistencies. In what follows, we describe the proposed local ZM-based representation scheme in terms of this threelayered framework. Feature Extraction – Local Zernike Moments: The computation of (complex) ZMs can be considered equivalent to representing an image in an alternative space. As shown in Figure 1-a, an image is decomposed onto a set of basis matrices (ZM bases), which are useful for describing the variation at different directions and scales. ZM bases are orthogonal, therefore there is no overlap in the information conveyed by each feature (ZM coefficient). ZMs are usually computed for the entire image, however in this case, ZMs cannot capture the local variation due to ZM bases lacking localisation [3]. In contrary, when computed around local neighbourhoods across the image, they become an efficient tool for describing the image discontinuities which are essential to interpreting facial activity. Non-linear Encoding – Quantisation: We perform quantisation via converting local features into binary values. Such coarse quantisation increases compactness and allows us to code each local block only with a single integer. Figure 1-b illustrates the process of obtaining the Quantised Local ZM (QLZM) image. Firstly, local ZM coefficients are computed across the input image (LZM layer) — each image in the LZM layer (LZM image) contains the features that are extracted through a particular ZM basis. Next, each LZM image is converted into a binary image by quantising each pixel via the signum(·) function. Finally, the QLZM image is obtained by combining all of the binary images. Specifically, each pixel in a particular location of the QLZM image is an integer (QLZM integer), computed by concatenating all of the binary values in the corresponding location of all binary images. The QLZM image is similar to an LBP-transformed image, in the sense that it contains integers of a limited range. Yet, the physical meaning of the information encoded by each integer is quite different. LBP integers describe a circular block by considering only the values along the border, neglecting the pixels that remain inside the block. Therefore, the efficient operation scale of LBPs is usually limited to 3-5 pixels [1, 5]. QLZM integers, on the other hand, describe blocks as a whole, and provide flexibility in terms of operation scale without major loss of information. Pooling – Histograms: Our representation scheme pools encoded features over local histograms. Figure 1-c illustrates the overall pipeline of the proposed representation scheme. Firstly, the QLZM image is computed through the process that is illustrated in detail in Figure 1-b. Next, . . . . . . ... = ZM Coefficients (local features) ZM Bases Evangelos Sariyanidi, Hatice Gunes, Muhittin Gökmen, Andrea Cavallaro |
BMVC | 4 |
| 2013 | Distributed measurement selection for energy-efficient radio tracking
Valerio Targon, Andrea Cavallaro |
FUSION | 2 |
| 2013 | Detecting group interactions by online association of trajectory dataabstractWe propose a method for detecting group interactions for groups of varying number of objects. We model each object as a moving agent with a direction-aware interest map and group interactions as mutual interests between objects. After grouping objects into unit interactions individually in each frame, we solve the temporal association problem by tracking group interaction over consecutive frames. Optimal grouping is obtained by finding the maximum weight spanning tree of a directed graph formed by objects and their potential interactions. Experimental results show that our method obtained around 80% recalling rates on two publicly available datasets. Fan Chen 0002, Andrea Cavallaro |
ICASSP | 2 |
| 2013 | Multi-target tracking on confidence maps: An application to people tracking
Fabio Poiesi, Riccardo Mazzon, Andrea Cavallaro |
Comput. Vis. Image Underst. | 3 |
| 2013 | Multi-camera tracking using a Multi-Goal Social Force Model
Riccardo Mazzon, Andrea Cavallaro |
Neurocomputing | 2 |
| 2013 | Video-Based Human Behavior Understanding: A SurveyabstractUnderstanding human behaviors is a challenging problem in computer vision that has recently seen important advances. Human behavior understanding combines image and signal processing, feature extraction, machine learning, and 3-D geometry. Application scenarios range from surveillance to indexing and retrieval, from patient care to industrial safety and sports analysis. Given the broad set of techniques used in video-based behavior understanding and the fast progress in this area, in this paper we organize and survey the corresponding literature, define unambiguous key terms, and discuss links among fundamental building blocks ranging from human detection to action and interaction recognition. The advantages and the drawbacks of the methods are critically discussed, providing a comprehensive coverage of key aspects of video-based human behavior understanding, available datasets for experimentation and comparisons, and important open research issues. Paulo Vinicius Koerich Borges, Nicola Conci, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Guest Editorial: Special issue on intelligent video surveillance for public security and personal privacyabstractThis Special Issue offers an overview of ongoing research on intelligent video surveillance (IVS) techniques, and brings together cutting-edge research work on security and privacy problems with respect to technological, behavioral, legal, and cultural aspects. We received 34 submissions and each submission was rigorously reviewed by at least two experts in the related fields based on the criteria of originality, significance, quality, and clarity. Eventually, 12 papers were accepted for the Special Issue, spanning a variety of topics including privacy protection, background modeling, tracking, action/activity analysis, and crowd behavior perception. The papers constituting this issue are then briefly summarized. Noboru Babaguchi, Andrea Cavallaro, Rama Chellappa, Frédéric Dufaux, Liang Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | A Protocol for Evaluating Video Trackers Under Real-World ConditionsabstractThe absence of a commonly adopted performance evaluation framework is hampering advances in the design of effective video trackers. In this paper, we present a single-score evaluation measure and a protocol to objectively compare trackers. The proposed measure evaluates tracking accuracy and failure, and combines them for both summative and formative performance assessment. The proposed protocol is composed of a set of trials that evaluate the robustness of trackers on a range of test scenarios representing several real-world conditions. The protocol is validated on a set of sequences with a diversity of targets (head, vehicle and person) and challenges (occlusions, background clutter, pose changes and scale changes) using six state-of-the-art trackers, highlighting their strengths and weaknesses on more than 187000 frames. The software implementing the protocol and the evaluation results are made available online and new results can be included, thus facilitating the comparison of trackers. Tahir Nawaz 0001, Andrea Cavallaro |
IEEE Trans. Image Process. | 2 |
| 2012 | People-background segmentation with unequal error costabstractWe address the problem of segmenting a video in two classes of different semantic value, namely background and people, with the goal of guaranteeing that no people (or body parts) are classified as background. Body parts classified as background are given a higher classification error cost (segmentation with bias on background), as opposed to traditional approaches focused on people detection. To generate the people-background segmentation mask, the proposed approach first combines detection confidence maps of body parts and then extends them in order to derive a background mask, which is finally post-processed using morphological operators. Experiments validate the performance of our algorithm in different complex indoor and outdoor scenes with both static and moving cameras. Álvaro García-Martín, Andrea Cavallaro, José María Martínez Sanchez |
ICIP | 2 |
| 2012 | Standalone evaluation of deterministic video trackingabstractWe present an approach for performance evaluation of deterministic video trackers without ground-truth data. The proposed approach detects if a tracker is correctly operating over time using two main steps. First, it transforms the output of the localization step into a distribution of the target state, which emulates a multi-hypothesis tracker. Then, the uncertainty of such distribution is estimated to determine the time instants when the tracker is stable. A time-reversed analysis is used to identify tracker recovery after unsuccessful operation. The proposed approach is demonstrated on the well-known MeanShift tracker. The results over a heterogeneous dataset show that the proposed approach outperforms the related state-of-the-art methods in presence of tracking challenges such as occlusions, illumination and scale changes, and clutter. Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez |
ICIP | 2 |
| 2012 | Interaction recognition in wide areas using audiovisual sensorsabstractWe present an event recognition framework to detect interactions among objects, for example people, using a network of cameras and associated microphone pairs. The complementarity of the video and audio modalities is exploited to cover wide areas. In particular, object movements in portions of the scene that are not covered by the cameras' fields of view are estimated using the input from microphones. After estimating trajectories using audio-visual features, we recognize interactions based on a Coupled Hidden Markov Model Maximum a Posteriori (CHMM-MAP) approach. The states of the CHMM are initialized via Gaussian Mixture Model (GMM) clustering on a multi-dimensional feature space. Evaluation and comparison with three alternative methods demonstrate the effectiveness of the proposed CHMM-MAP trained on multiple features on both synthetic and real data. Murtaza Taj, Andrea Cavallaro |
ICIP | 2 |
| 2012 | Person re-identification in crowd
Riccardo Mazzon, Syed Fahad Tahir, Andrea Cavallaro |
Pattern Recognit. Lett. | 3 |
| 2012 | Adaptive Appearance Modeling for Video Tracking: Survey and EvaluationabstractLong-term video tracking is of great importance for many applications in real-world scenarios. A key component for achieving long-term tracking is the tracker's capability of updating its internal representation of targets (the appearance model) to changing conditions. Given the rapid but fragmented development of this research area, we propose a unified conceptual framework for appearance model adaptation that enables a principled comparison of different approaches. Moreover, we introduce a novel evaluation methodology that enables simultaneous analysis of tracking accuracy and tracking success, without the need of setting application-dependent thresholds. Based on the proposed framework and this novel evaluation methodology, we conduct an extensive experimental comparison of trackers that perform appearance model adaptation. Theoretical and experimental analyses allow us to identify the most effective approaches as well as to highlight design choices that favor resilience to errors during the update process. We conclude the paper with a list of key open research challenges that have been singled out by means of our experimental comparison. Samuele Salti, Andrea Cavallaro, Luigi Di Stefano |
IEEE Trans. Image Process. | 2 |
| 2012 | Adaptive Online Performance Evaluation of Video TrackersabstractWe propose an adaptive framework to estimate the quality of video tracking algorithms without ground-truth data. The framework is divided into two main stages, namely, the estimation of the tracker condition to identify temporal segments during which a target is lost and the measurement of the quality of the estimated track when the tracker is successful. A key novelty of the proposed framework is the capability of evaluating video trackers with multiple failures and recoveries over long sequences. Successful tracking is identified by analyzing the uncertainty of the tracker, whereas track recovery from errors is determined based on the time-reversibility constraint. The proposed approach is demonstrated on a particle filter tracker over a heterogeneous data set. Experimental results show the effectiveness and robustness of the proposed framework that improves state-of-the-art approaches in the presence of tracking challenges such as occlusions, illumination changes, and clutter and on sequences containing multiple tracking errors and recoveries. Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez |
IEEE Trans. Image Process. | 2 |
| 2011 | Abnormal motion detection in crowded scenes using local spatio-temporal analysisabstractWe present a motion classification approach to detect movements of interest (abnormal motion) based on local feature modeling within spatio-temporal detectors. The modeling is performed using motion vectors and local detectors. The detectors are trained independently for learning abnormal motion based on labeled samples. Each detector is assigned an abnormality score, both in space and time, which is the basis of the final classification. The spatial relationship across detectors is used to discriminate simultaneous occurrences of abnormal motion. The performance of the proposed method is evaluated on 52 hours of the multi-camera surveillance dataset of the TRECVID 2010 challenge. Fahad Daniyal, Andrea Cavallaro |
ICASSP | 2 |
| 2011 | PFT: A protocol for evaluating video trackersabstractThe growing interest in developing video tracking algorithms has not been accompanied by the development of commonly used evaluation criteria to assess and to compare their performance. Researchers often present trackers' results on different datasets and evaluate them with different performance measures thus hindering both formative and summative quality assessment. In this paper, we present a protocol to evaluate the performance of tracking algorithms that tests video trackers using a set of trials and a pre-defined set of sequences and that enables objective and reproducible performance evaluation of trackers using ground truth information. Each trial highlights strengths and weaknesses of a tracker on simulated test scenarios on real sequences that represent real-world scenarios. Moreover a new evaluation measure is introduced that allows us to summarize the performance of a tracker based on the lost-track-ratio curve. The validation and the effectiveness of the proposed protocol is demonstrated experimentally on three trackers and its implementation is made available online to the research community. Tahir Nawaz 0001, Andrea Cavallaro |
ICIP | 2 |
| 2011 | Efficient depth blurring with occlusion handlingabstractWe present a novel algorithm for enhancing an image or video frame with depth of field. The algorithm deals with occlusive blurring effects as occur in real cameras and can handle a variation in blur which is continuous up to blur level quantization granularity, with asymptotic complexity O(N log2N) in terms of time and memory for an N-pixel image, irrespective of the complexity of the variation in blur. The proposed algorithm is a postfiltering approach which, unlike prior algorithms, does not suffer from intensity leakage. Experimental results show the algorithm to be 3 to 4 times faster than an existing algorithm, which had the same asymptotic complexity. Tim Popkin, Andrea Cavallaro, David Hands |
ICIP | 2 |
| 2011 | Special Issue on Video Analysis on Resource-Limited SystemsabstractThe 17 papers in this special issue focus on resource-limited systems. Rama Chellappa, Andrea Cavallaro, Ying Wu 0001, Caifeng Shan, Yun Fu 0001, Kari Pulli |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Image Coding Using Depth Blurring for Aesthetically Acceptable DistortionabstractWe introduce the concept of depth-based blurring to achieve an aesthetically acceptable distortion when reducing the bitrate in image coding. The proposed depth-based blurring is a prefiltering that reduces high-frequency components by mimicking the limited depth of field effect that occurs in cameras. To cope with the challenge of avoiding intensity leakage at the boundaries of objects when blurring at different depth levels, we introduce a selective blurring algorithm that simulates occlusion effects as occur in natural blurring. The proposed algorithm can handle any number of blurring and occlusion levels. Subjective experiments show that the proposed algorithm outperforms foveation filtering, which is the dominant approach for bitrate reduction by space-variant prefiltering. Tim Popkin, Andrea Cavallaro, David Hands |
IEEE Trans. Image Process. | 2 |
| 2010 | Local Abnormality Detection in Video Using Subspace LearningabstractOn-line abnormality detection in video without the use of object detection and tracking is a desirable task in surveillance.We address this problem for the case when labeled information about normal events is limited and information about abnormal events is not available. We formulate this problem as a one-class classification, where multiple local novelty classifiers (detectors) are used to first learn normal actions based on motion information and then to detect abnormal instances. Each detector is associated to a small region of interest and is trained over labeled samples projected on an appropriate subspace. We discover this subspace by using both labeled and unlabeled segments.We investigate the use of subspace learning and compare two methodologies based on linear (Principal Components Analysis) and on non-linear subspace learning (Locality Preserving Projections), respectively. Experimental results on a real underground station dataset shows that the linear approach is better suited for cases where the subspace learning is restricted to the labeled samples, whereas the non-linear approach is preferable in the presence of additional unlabeled data. Ioannis Tziakos, Andrea Cavallaro, Li-Qun Xu |
AVSS | 2 |
| 2010 | Detector-less ball localization using context and motion flow analysisabstractWe present a technique for estimating the location of the ball during a basketball game without using a detector. The technique is based on the analysis of the dynamics in the scene and allows us to overcome the challenges due to frequent occlusions of the ball and its similarity in appearance with the background. Based on the assumption that the ball is the point of focus of the game and that the motion flow of the players is dependent on its position during attack actions, the most probable candidates for the ball location are extracted from each frame. These candidates are then validated over time using a Kalman filter. Experimental results on a real basketball dataset show that the location of the ball can be estimated with an average accuracy of 82%. Fabio Poiesi, Fahad Daniyal, Andrea Cavallaro |
ICIP | 3 |
| 2010 | Evaluation of on-line quality estimators for object trackingabstractFailure of tracking algorithms is inevitable in real and on-line tracking systems. The online estimation of the track quality is therefore desirable for detecting tracking failures while the algorithm is operating. In this paper, we propose a taxonomy and present a comparative evaluation of online quality estimators for video object tracking. The measures are compared over a heterogeneous video dataset with standard sequences. Among other results, the experiments show, that the Observation Likelihood (OL) measure is an appropriate quality measure for overall tracking performance evaluation, while the Template Inverse Matching (TIM) measure is appropriate to detect the start and the end instants of tracking failures. Juan C. SanMiguel, Andrea Cavallaro, José María Martínez Sanchez |
ICIP | 2 |
| 2010 | Special issue on multi-camera and multi-modal sensor fusion
Andrea Cavallaro, Hamid K. Aghajan |
Comput. Vis. Image Underst. | 1 |
| 2010 | Event monitoring via local motion abnormality detection in non-linear subspace
Ioannis Tziakos, Andrea Cavallaro, Li-Qun Xu |
Neurocomputing | 2 |
| 2010 | Content and task-based view selection from multiple video streams
Fahad Daniyal, Murtaza Taj, Andrea Cavallaro |
Multim. Tools Appl. | 3 |
| 2010 | Accurate and Efficient Method for Smoothly Space-Variant Gaussian BlurringabstractThis paper presents a computationally efficient algorithm for smoothly space-variant Gaussian blurring of images. The proposed algorithm uses a specialized filter bank with optimal filters computed through principal component analysis. This filter bank approximates perfect space-variant Gaussian blurring to arbitrarily high accuracy and at greatly reduced computational cost compared to the brute force approach of employing a separate low-pass filter at each image location. This is particularly important for spatially variant image processing such as foveated coding. Experimental results show that the proposed algorithm provides typically 10 to 15 dB better approximation of perfect Gaussian blurring than the blended Gaussian pyramid blurring approach when using a bank of just eight filters. Tim Popkin, Andrea Cavallaro, David Hands |
IEEE Trans. Image Process. | 2 |
| 2009 | Trajectory Association and Fusion across Partially Overlapping CamerasabstractWe present a novel unsupervised inter-camera trajectory correspondence algorithm that does not require prior knowledge of the camera placement. The approach consists of three steps, namely association, fusion and linkage. For association, local trajectory pairs corresponding to the same physical object are estimated using multiple spatio-temporal features on a common ground-plane. To disambiguate spurious associations, we employ a hybrid approach that utilizes the matching results on the image- and ground-plane. The trajectory segments after association are fused by adaptive averaging. Finally, linkage integrates segments and generates a single trajectory of an object across the entire observed area. We evaluated the performance of the proposed approach on a simulated and two real scenarios with simultaneous moving objects observed by multiple cameras and compared it with state-of-the-art algorithms. Convincing results are observed in favor of the proposed approach. Nadeem Anjum, Andrea Cavallaro |
AVSS | 2 |
| 2009 | Compact Signatures for 3D Face Recognition under Varying ExpressionsabstractWe present a novel approach to 3D face recognition using compact face signatures based on automatically detected 3D landmarks. We represent the face geometry with inter-landmark distances within selected regions of interest to achieve robustness to expression variations. The inter-landmark distances are compressed through principal component analysis and linear discriminant analysis is then applied on the reduced features to maximize the separation between face classes. The classification of a probe face is based on a nearest mean classifier after transforming the probe onto the subspace. We analyze the performance of different landmark combinations (signatures) to determine a signature that is robust to expressions. The selected signature is then used to train a point distribution model for the automatic localization of the landmarks, without any prior knowledge of scale, pose, orientation or texture. We evaluate the proposed approach on a challenging publicly available facial expression database (BU-3DFE) and achieve 96.5% recognition rate using the automatically localized signature. Moreover, because of its compactness the face signature can be stored on 2D barcodes and used for radio-frequency identification. Fahad Daniyal, Prathap M. Nair, Andrea Cavallaro |
AVSS | 3 |
| 2009 | Grouping motion trajectoriesabstractWe present a method to group trajectories of moving objects extracted from real-world surveillance videos. The trajectories are first mapped into a low dimensionality feature space generated through linear regression. Next the regression coefficients are clustered by a Gaussian mixture model initialized by K-means for improved efficiency. The model selection problem is solved with Bayesian information criterion that penalizes models with high complexity. We demonstrate the proposed approach on both synthetic and real-world scenes. Experimental results show that the proposed clustering method outperforms K-means and mixture of regression models, while also reducing the computational complexity compared to the latter. Samuel Pachoud, Emilio Maggio, Andrea Cavallaro |
ICASSP | 3 |
| 2009 | Distance blurring for space-variant image codingabstractWe present a novel selective blurring algorithm that mimics the optical distance blur effects that occur naturally in cameras and eyes. The proposed algorithm provides a realistic simulation of distance blurring, with the desirable properties of aiming to mimic occlusion effects as occur in natural blurring, and of being able to handle any number of blurring and occlusion levels with the same order of computational complexity. We have performed subjective experiments to compare the perceived quality of distance blurred images with that of foveation-filtered images under equivalent conditions, when both are used as space-variant prefiltering stages prior to a JPEG encoder. The results show that the distance-based blurring was significantly preferable to the foveation blurring for four out of nine test images, whereas a significant converse preference was found for only one test image. Tim Popkin, Andrea Cavallaro, David Hands |
ICASSP | 2 |
| 2009 | Multi-foveation filteringabstractWe present a method for computing a function of average multi-viewer eye sensitivity based on the Geisler & Perry contrast threshold formula, and, from this, the cut-off frequency map (as used in foveation filtering) that is optimal in the sense of discarding frequencies in least-noticeable-first order. Existing approaches usually solve the multi-viewer foveation problem as a number of single-viewer foveations, effectively taking collective sensitivity to be the maximum of the individual viewer eye sensitivities. This has inherent problems such as over-sensitivity to outliers which are not problems with the proposed approach. Furthermore, the proposed approach can be employed in the infinite-viewer (probability-based) scenario without additional cost. Tim Popkin, Andrea Cavallaro, David Hands |
ICASSP | 2 |
| 2009 | Audio-assisted trajectory estimation in non-overlapping multi-camera networksabstractWe present an algorithm to improve trajectory estimation in networks of non-overlapping cameras using audio measurements. The algorithm fuses audiovisual cues in each camera's field of view and recovers trajectories in unobserved regions using microphones only. Audio source localization is performed using stereo audio and cycloptic vision (STAC) sensor by estimating the time difference of arrival (TDOA) between microphone pair and then by computing the cross correlation. Audio estimates are then smoothed using Kalman filtering. The audio-visual fusion is performed using a dynamic weighting strategy. We show that using a multi-modal sensor with combined visual (narrow) and audio (wider) field of view can enable extended target tracking in non-overlapping camera settings. In particular, the weighting scheme improves performance in the overlapping regions. The algorithm is evaluated in several multi-sensor configurations using synthetic data and compared with state of the art algorithm. Murtaza Taj, Andrea Cavallaro |
ICASSP | 2 |
| 2009 | Accurate appearance-based Bayesian tracking for maneuvering targets
Emilio Maggio, Andrea Cavallaro |
Comput. Vis. Image Underst. | 2 |
| 2009 | Video event segmentation and visualisation in non-linear subspace
Ioannis Tziakos, Andrea Cavallaro, Li-Qun Xu |
Pattern Recognit. Lett. | 2 |
| 2009 | Learning Scene Context for Multiple Object TrackingabstractWe propose a framework for multitarget tracking with feedback that accounts for scene contextual information. We demonstrate the framework on two types of context-dependent events, namely target births (i.e., objects entering the scene or reappearing after occlusion) and spatially persistent clutter. The spatial distributions of birth and clutter events are incrementally learned based on mixtures of Gaussians. The corresponding models are used by a probability hypothesis density (PHD) filter that spatially modulates its strength based on the learned contextual information. Experimental results on a large video surveillance dataset using a standard evaluation protocol show that the feedback improves the tracking accuracy from 9% to 14% by reducing the number of false detections and false trajectories. This performance improvement is achieved without increasing the computational complexity of the tracker. Emilio Maggio, Andrea Cavallaro |
IEEE Trans. Image Process. | 2 |
| 2009 | 3-D Face Detection, Landmark Localization, and Registration Using a Point Distribution ModelabstractWe present an accurate and robust framework for detecting and segmenting faces, localizing landmarks, and achieving fine registration of face meshes based on the fitting of a facial model. This model is based on a 3-D Point Distribution Model (PDM) that is fitted without relying on texture, pose, or orientation information. Fitting is initialized using candidate locations on the mesh, which are extracted from low-level curvature-based feature maps. Face detection is performed by classifying the transformations between model points and candidate vertices based on the upper-bound of the deviation of the parameters from the mean model. Landmark localization is performed on the segmented face by finding the transformation that minimizes the deviation of the model from the mean shape. Face registration is obtained using prior anthropometric knowledge and the localized landmarks. The performance of face detection is evaluated on a database of faces and non-face objects where we achieve an accuracy of 99.6%. We also demonstrate face detection and segmentation on objects with different scale and pose. The robustness of landmark localization is evaluated with noisy data and by varying the number of shapes and model points used in the model learning phase. Finally, face registration is compared with the traditional Iterative Closest Point (ICP) method and evaluated through a face retrieval and recognition framework on the GavabDB dataset, where we achieve a recognition rate of 87.4% and a retrieval rate of 83.9%. Prathap M. Nair, Andrea Cavallaro |
IEEE Trans. Multim. | 2 |
| 2008 | Object and Scene-Centric Activity Detection Using State Occupancy Duration ModelingabstractWe propose a video event analysis framework based on object segmentation and tracking, combined with a Hidden Semi-Markov Model (HSMM) that uses state occupancy duration modeling. The observations generated by a multi-object detector and tracker are used as emitting symbols and the corresponding probabilities are computed using multivariate Gaussians. Next, we recognize events by estimating the most likely object state sequence using a HSMM decoding strategy, based on the Viterbi algorithm. Moreover,the duration distribution enforces the state transition after certain time and hence better models the events constrained on time intervals. We demonstrate and evaluate the proposed framework on a dataset of approximately 20 K frames, and show that the duration modeling improves the event detection results by 7% to 11%, compared to state-of-the-art HMMs. Murtaza Taj, Andrea Cavallaro |
AVSS | 2 |
| 2008 | Matching 3D Faces with Partial DataabstractWe present a novel approach to matching 3D faces with expressions, deformations and outliers. The matching is performed through an accurate and robust algorithm for registering face meshes. The registration algorithm incorporates prior anthropometric knowledge through the use of suitable landmarks and regions of the face to be used with the Iterative Closest Point (ICP) registration algorithm. The localization of landmarks and regions is achieved through the fitting of a 3D Point Distribution Model (PDM) and is independent of texture, pose and orientation information. We show that the use of expression-invariant facial regions for registration and similarity estimation outperforms the use of the entire face region. Evaluation is performed on the challenging GavabDB database and we achieve 93.7 % rank-1 recognition with an overall retrieval accuracy of 91.1%. 1 Prathap M. Nair, Andrea Cavallaro |
BMVC | 2 |
| 2008 | Video Augmentation for Improving Audio Speech Recognition under NoiseabstractFor the recognition of speech, in particular spoken digits, captured in video with poor sound due to noise, we develop a novel audio-visual fusion technique that performs significantly better than utilising either audio or video signal alone. Specifically, we present an audio-visual intermediate fusion strategy to locate speaker dependant pronounced digits in continuous video recorded with sound. A model template for each digit is represented in a single audio-visual feature space using a set of spatio-temporal visual features at multiple scales together with a set of thirteen Mel Frequency Cepstral Coefficients as audio features. Using a unified structure for both visual and audio feature selection and extraction, we solve the problem of one-to-one correspondence between the audio and visual spaces caused by differences in data sampling rates. To combine the two modalities, we adopt an intermediate fusion strategy by combining the two modalities in a probabilistic sequence matching function, permitting automatic segmentation of a continuous probe video sequence and matching with available model templates. For experiments, the CUAVE [17] database was used to compare our scheme with two alternative methods. The evaluation shows that the proposed approach outperforms the others both in recognition accuracy and robustness in coping with variations in probe sequences. 1 Samuel Pachoud, Shaogang Gong, Andrea Cavallaro |
BMVC | 3 |
| 2008 | Macro-cuboïd based probabilistic matching for lip-reading digitsabstractIn this paper, we present a spatio-temporal feature representation and a probabilistic matching function to recognise lip movements from pronounced digits. Our model (1) automatically selects spatio-temporal features extracted from 10 digit model templates and (2) matches them with probe video sequences. Spatio-temporal features embed lip movements from pronouncing digits and contain more discriminative information than spatial features alone. A model template for each digit is represented by a set of spatio-temporal features at multiple scales. A probabilistic sequence matching function automatically segments a probe video sequence and matches the most likely sequence of digits recognised in the probe sequence. We demonstrate the proposed approach using the CUAVE database and compare our representational scheme with three alternative methods, based on optical flow, intensity gradient and block matching, respectively. The evaluation shows that the proposed approach outperforms the others in recognition accuracy and is robust in coping with variations in probe sequences. Samuel Pachoud, Shaogang Gong, Andrea Cavallaro |
CVPR | 3 |
| 2008 | SHREC'08 entry: Registration and retrieval of 3D faces using a Point Distribution ModelabstractWe present a 3D face retrieval approach based on the fitting of a 3D Point Distribution Model (PDM). The model is fitted on a face without relying on texture, pose or orientation information. The retrieval is based on fine registration obtained using anthropometric knowledge and facial landmarks that are localized using the model. Prathap M. Nair, Andrea Cavallaro |
Shape Modeling International | 2 |
| 2008 | Multifeature Object Trajectory Clustering for Video AnalysisabstractWe present a novel multifeature video object trajectory clustering algorithm that estimates common patterns of behaviors and isolates outliers. The proposed algorithm is based on four main steps, namely the extraction of a set of representative trajectory features, non-parametric clustering, cluster merging and information fusion for the identification of normal and rare object motion patterns. First we transform the trajectories into a set of feature spaces on which mean-shift identifies the modes and the corresponding clusters. Furthermore, a merging procedure is devised to refine these results by combining similar adjacent clusters. The final common patterns are estimated by fusing the clustering results across all feature spaces. Clusters corresponding to reoccurring trajectories are considered as normal, whereas sparse trajectories are associated to abnormal and rare events. The performance of the proposed algorithm is evaluated on standard data-sets and compared with state-of-the-art techniques. Experimental results show that the proposed approach outperforms state-of-the-art algorithms both in terms of accuracy and robustness in discovering common patterns in video as well as in recognizing outliers. Nadeem Anjum, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Efficient Multitarget Visual Tracking Using Random Finite SetsabstractWe propose a filtering framework for multitarget tracking that is based on the probability hypothesis density (PHD) filter and data association using graph matching. This framework can be combined with any object detectors that generate positional and dimensional information of objects of interest. The PHD filter compensates for missing detections and removes noise and clutter. Moreover, this filter reduces the growth in complexity with the number of targets from exponential to linear by propagating the first-order moment of the multitarget posterior, instead of the full posterior. In order to account for the nature of the PHD propagation, we propose a novel particle resampling strategy and we adapt dynamic and observation models to cope with varying object scales. The proposed resampling strategy allows us to use the PHD filter when a priori knowledge of the scene is not available. Moreover, the dynamic and observation models are not limited to the PHD filter and can be applied to any Bayesian tracker that can handle state-dependent variances. Extensive experimental results on a large video surveillance dataset using a standard evaluation protocol show that the proposed filtering framework improves the accuracy of the tracker, especially in cluttered scenes. Emilio Maggio, Murtaza Taj, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Single camera calibration for trajectory-based behavior analysisabstractPerspective deformations on the image plane make the analysis of object behaviors difficult in surveillance video. In this paper, we improve the results of trajectory-based scene analysis by using single camera calibration for perspective rectification. First, the ground-plane view is estimated from perspective images captured from a single camera. Next, unsupervised fuzzy clustering is applied on the transformed trajectories to group similar behaviors and to isolate outliers. We evaluate the proposed approach on real outdoor surveillance scenarios with standard datasets and show that perspective rectification improves the accuracy of the trajectory clustering results. Nadeem Anjum, Andrea Cavallaro |
AVSS | 2 |
| 2007 | Relative Position Estimation of Non-Overlapping CamerasabstractWe present an algorithm for the estimation of the relative camera position in a network of cameras with non-overlapping fields of view. The algorithm estimates the missing trajectory information in the unobserved areas of the multi-sensor configuration using both parametric and non-parametric algorithms. First, Kalman filtering is used to estimate the trajectories in the unobserved regions. Next, linear regression estimates the position of the target based upon the motion model generated from the measured positions in the field of view of each sensor. Finally, the relative orientation of the sensors is calculated using the observed and estimated target position from adjacent cameras. We demonstrate the algorithm on both synthetic and real data. Nadeem Anjum, Murtaza Taj, Andrea Cavallaro |
ICASSP (2) | 3 |
| 2007 | Hands-On Experience in Image Processing: The Automated Lecture CameramanabstractWe present a system for live recorded lecture-based distance learning delivery that was designed and implemented based on the framework defined in A. Cavallaro et al, 2005. The system is built by students based on previous students' projects and is deployed in a real distance learning scenario. The video capturing process of a lecture is automated using a robotic camera that tracks the movements of a lecturer during the delivery of a traditional class. The robotic camera is guided by the results of an image processing module based on face detection. The video of the lecturer is synchronized with the presentation slides and with the audio of the lecture. The system was evaluated based on students' feedback. Andrea Cavallaro, Ruchira Chandrasekera, Murtaza Taj |
ICASSP (3) | 1 |
| 2007 | Particle PHD Filtering for Multi-Target Visual TrackingabstractWe propose a multi-target tracking algorithm based on the probability hypothesis density (PHD) filter and data association using graph matching. The PHD filter is used to compensate for miss-detections and to remove noise and clutter. This filter propagates the first order moment of the multi-target posterior (instead of the full posterior) to reduce the growth in complexity with the number of targets from exponential to linear. Next the filtered states are associated using graph matching. Experimental results on face, people and vehicle tracking show that the proposed multi-target tracking algorithm improves the accuracy of the tracker, especially in cluttered scenes. Emilio Maggio, Elisa Piccardo, Carlo S. Regazzoni, Andrea Cavallaro |
ICASSP (1) | 4 |
| 2007 | Tracking Atoms with Particles for Audio-Visual Source LocalizationabstractWe present a general framework and an efficient algorithm for tracking relevant video structures. The structures to be tracked are implicitly defined by a matching pursuit procedure that extracts and ranks the most important image contours. Based on the ranking, the contours are automatically selected to initialize a particle filtering tracker. The proposed algorithm deals with salient video entities whose behavior has an intuitive meaning, related to the physics of the signal. Moreover, as the interactions between such structures are easily defined, the inference of higher level signal configurations can be made intuitive. The proposed algorithm improves the performance of existing video structures trackers, while reducing the computational complexity. The algorithm is demonstrated on audio-visual source localization. Gianluca Monaci, Pierre Vandergheynst, Emilio Maggio, Andrea Cavallaro |
ICASSP (2) | 4 |
| 2007 | Unsupervised Fuzzy Clustering for Trajectory AnalysisabstractWe propose an unsupervised fuzzy approach for motion trajectory clustering. The proposed approach is divided into three main steps: first Mean-shift is used for local mode seeking by analyzing trajectory data over multiple feature spaces. This step generates a set of tentative clusters. Next, adjacent clusters are combined by analysing the cluster attributes across all feature spaces. Sparse clusters are finally considered as generated by outlier object behaviors and then removed. The performance of the proposed algorithm is evaluated on real outdoor video surveillance scenarios with standard data-sets and it is compared with state-of-the-art techniques. Nadeem Anjum, Andrea Cavallaro |
ICIP (3) | 2 |
| 2007 | Multi-Modal Particle Filtering Tracking using Appearance, Motion and Audio LikelihoodsabstractWe propose a multi-modal object tracking algorithm that combines appearance, motion and audio information in a particle filter. The proposed tracker fuses at the likelihood level the audio-visual observations captured with a video camera coupled with two microphones. Two video likelihoods are computed that are based on a 3D color histogram appearance model and on a color change detection, whereas an audio likelihood provides information about the direction of arrival of a target. The direction of arrival is computed based on a multi-band generalized cross-correlation function enhanced with a noise suppression and reverberation filtering that uses the precedence effect. We evaluate the tracker on single and multi-modality tracking and quantify the performance improvement introduced by integrating audio and visual information in the tracking process. Matteo Bregonzio, Murtaza Taj, Andrea Cavallaro |
ICIP (5) | 3 |
| 2007 | Region Segmentation and Feature Point Extraction on 3D Faces using a Point Distribution ModelabstractWe present a novel approach to accurately detect landmarks and segment regions on face meshes without the use of texture, pose or orientation information. The proposed approach is based on a 3D point distribution model (PDM) that is fitted to the region of interest using candidate vertices extracted from low-level feature maps. The robustness of the algorithm is evaluated in the presence of noise and at the variation of the number of scans and model points used in the learning phase. Experimental results demonstrate the accuracy of the proposed method in detecting landmarks, with an improvement of 55% over a state-of-the-art non-statistical approach. Prathap M. Nair, Andrea Cavallaro |
ICIP (3) | 2 |
| 2007 | Multi-Camera Scene Analysis using an Object-Centric Continuous Distribution Hidden Markov ModelabstractWe propose a multi-camera event detection framework that can operate on a common ground plane as well as on the image plane. The proposed event detector is based on an object-centric state modeling that uses a continuous distribution hidden Markov model (CDHMM). Video objects are first detected using statistical change detection and then tracked using graph matching. Next, the algorithm recognizes events by estimating the most likely object state sequence using a HMM decoding strategy, based on the Viterbi algorithm. We demonstrate and evaluate the proposed framework on standard event detection datasets with single and multiple cameras, with both overlapping and non-overlapping fields of view. Murtaza Taj, Andrea Cavallaro |
ICIP (4) | 2 |
| 2007 | Adaptive Multifeature Tracking in a Particle Filtering FrameworkabstractIn this paper, we propose a tracking algorithm based on an adaptive multifeature statistical target model. The features are combined in a single particle filter by weighting their contributions using a novel reliability measure derived from the particle distribution in the state space. This measure estimates the reliability of the information by measuring the spatial uncertainty of features. A modified resampling strategy is also devised to account for the needs of the feature reliability estimation. We demonstrate the algorithm using color and orientation features. Color is described with partwise normalized histograms. Orientation is described with histograms of the gradient directions that represent the shape and the internal edges of a target. A feedback from the state estimation is used to align the orientation histograms as well as to adapt the scales of the filters to compute the gradient. Experimental results over a set of real-world sequences show that the proposed feature weighting procedure outperforms state-of-the-art solutions and that the proposed adaptive multifeature tracker improves the reliability of the target estimate while eliminating the need of manually selecting each feature's relevance. Emilio Maggio, F. Smerladi, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2005 | Surveillance video for mobile devicesabstractIn this paper, we present a video encoding scheme that uses object-based adaptation to deliver surveillance video to mobile devices. The method relies on a set of complementary video adaptation strategies and generates content that matches various appliance and network resources. Prior to encoding, some of the adaptation strategies exploit video object segmentation and selective filtering in order to improve the perceived quality. Moreover, object segmentation enables the generation of automatic summaries and of simplified versions of the monitored scene. The performance of individual adaptation strategies is assessed using an objective video quality metric, which is also used to select the strategy that provides maximum value for the user under a given set of constraints. We demonstrate the effectiveness of the scheme on standard surveillance test sequences and realistic mobile client resource profiles. Olivier Steiger, Touradj Ebrahimi, Andrea Cavallaro |
AVSS | 3 |
| 2005 | Performance evaluation of event detection solutions: the CREDS experienceabstractIn video surveillance projects, automatic and real-time event detection solutions are required to guarantee an efficient and cost-effective use of the infrastructure. Many solutions have been proposed to automatically detect a variety of events of interest. However, not all solutions and technologies may satisfy all the requirements of the surveillance scenario. For this reason, performance evaluation of existing event detection solutions becomes an important step in the deployment of video surveillance projects. In this paper, we propose a practical approach that aims at minimizing the ground truth generation problem and the expertise required to evaluate and compare the results by introducing specific requirements of specific event detection scenarios. This approach is believed to be applicable for an initial evaluation of candidate solutions to a specific surveillance scenario before more exhaustive tests in an integrated environment. The proposed method is under evaluation in the framework of the challenge of real-time event detection solutions (CREDS). Francesco Ziliani, Sergio A. Velastin, Fatih Porikli, Lucio Marcenaro, Timothy P. Kelliher, Andrea Cavallaro, Philippe Bruneaut |
AVSS | 6 |
| 2005 | Combining Colour and Orientation for Adaptive Particle Filter-based TrackingabstractWe propose an accurate tracking algorithm based on a multi-feature statistical model. The model combines in a single particle filter colour and gradient-based orientation information. A reliability measure derived from the particle distribution is used to adaptively weigh the contribution of the two features. Furthermore, information from the tracker is used to set the dimension of the filters for the computation of the gradient, effectively solving the scale selection problem. Experiments over a set of real-world sequences show that the adaptive use of colour and orientation information improves over either feature taken separately, both in terms of tracking accuracy and of reduction of lost tracks. Also, the automatic scale selection for the derivative filters results in increased robustness. 2 Emilio Maggio, Fabrizio Smeraldi, Andrea Cavallaro |
BMVC | 3 |
| 2005 | Image analysis and computer vision for undergraduatesabstractReal hands-on experience can help students gain a better understanding of theoretical problems in image analysis and computer vision and allows them to put into practice and improve their knowledge of digital signal processing, mathematics, statistics, perception and psychophysics. However, important efforts are necessary to enable students to develop a computer vision application because of the lack of extensively tested and well documented software platforms. In this paper, we describe our experience with an open source library addressed at researchers and developers in computer vision, the OpenCV library, its limits when used by students, and how we adapted it for teaching purposes by producing a set of appropriate tutorials. These tutorials help the students reduce the average time for installation and setup from 1 week to 4 hours and help them design an end-to-end image analysis and computer vision project. Finally, we discuss our experience of using this framework for undergraduate as well postgraduate student projects. Andrea Cavallaro |
ICASSP (5) | 1 |
| 2005 | Hybrid Particle Filter and Mean Shift tracker with adaptive transition modelabstractWe propose a tracking algorithm based on a combination of particle filter and mean shift, and enhanced with a new adaptive state transition model. The particle filter is robust to partial and total occlusions, can deal with multi-modal pdf and can recover lost tracks. However, its complexity dramatically increases with the dimensionality of the sampled pdf. Mean shift has a low complexity, but is unable to deal with multi-modal pdf. To overcome these problems, the proposed tracker first produces a smaller number of samples than the particle filter and then shifts the samples toward a close local maximum using mean shift. The transition model predicts the state based on adaptive variances. Experimental results show that the combined tracker outperforms the particle filter and mean shift in terms of accuracy in estimating the target size and position while generating 80% less samples than the particle filter. Emilio Maggio, Andrea Cavallaro |
ICASSP (2) | 2 |
| 2005 | Multi-part target representation for color trackingabstractThis paper presents an effective target representation based on multiple colour histograms computed on semi-overlapping image areas. This solution introduces spatial information in the representation, without compromising the benefits of the histograms. In particular, target rotation and scaling can be accounted for, thus improving the tracker robustness to false targets. We demonstrate that the proposed target representation outperforms the standard single histogram model and non-overlapping multi-part representations, using state-of-the-art tracking algorithms. Experimental results show that the proposed representation achieves an improvement of tracking accuracy and a reduction of track losses, without increasing significantly the computational complexity. Emilio Maggio, Andrea Cavallaro |
ICIP (1) | 2 |
| 2005 | Evaluating Perceptually Prefiltered VideoabstractPerceptual prefiltering is the process of enhancing relevant portions of an image or of a video, and of simplifying contextual information in order to improve the perceived quality or the compression ratio. In this paper, we discuss the results of subjective quality evaluation experiments performed to assess the impact of perceptual prefiltering on video coding and we propose an objective quality metric that mimics the behavior of human observers. The predicted performance of the proposed metric is consistent with the subjective evaluation scores. Experimental results demonstrate that perceptual prefiltering leads to quality improvements by up to 10% at low bit rates Olivier Steiger, Touradj Ebrahimi, Andrea Cavallaro |
ICME | 3 |
| 2005 | Tracking Video Objects in Cluttered BackgroundabstractWe present an algorithm for tracking video objects which is based on a hybrid strategy. This strategy uses both object and region information to solve the correspondence problem. Low-level descriptors are exploited to track object's regions and to cope with track management issues. Appearance and disappearance of objects, splitting and partial occlusions are resolved through interactions between regions and objects. Experimental results demonstrate that this approach has the ability to deal with multiple deformable objects, whose shape varies over time. Furthermore, it is very simple, because the tracking is based on the descriptors, which represent a very compact piece of information about regions, and they are easy to define and track automatically. Finally, this procedure implicitly provides one with a description of the objects and their track, thus enabling indexing and manipulation of the video content. Andrea Cavallaro, Olivier Steiger, Touradj Ebrahimi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | Semantic video analysis for adaptive content delivery and automatic descriptionabstractWe present an encoding framework which exploits semantics for video content delivery. The video content is organized based on the idea of main content message. In the work reported in this paper, the main content message is extracted from the video data through semantic video analysis, an application-dependent process that separates relevant information from non relevant information. We use here semantic analysis and the corresponding content annotation under a new perspective: the results of the analysis are exploited for object-based encoders, such as MPEG-4, as well as for frame-based encoders, such as MPEG-1. Moreover, the use of MPEG-7 content descriptors in conjunction with the video is used for improving content visualization for narrow channels and devices with limited capabilities. Finally, we analyze and evaluate the impact of semantic video analysis in video encoding and show that the use of semantic video analysis prior to encoding sensibly reduces the bandwidth requirements compared to traditional encoders not only for an object-based encoder but also for a frame-based encoder. Andrea Cavallaro, Olivier Steiger, Touradj Ebrahimi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2004 | Color image scalable coding with matching pursuitabstractThe paper presents a new scalable and highly flexible color image coder based on a matching pursuit expansion. The matching pursuit algorithm provides an intrinsically progressive stream and the proposed coder allows us to reconstruct color information from the first bit received. In order to capture edges in natural images efficiently, the dictionary of atoms is built by translation, rotation and anisotropic refinement of a wavelet-like mother function. This dictionary is moreover invariant under shifts and isotropic scaling, thus leading to very simple spatial resizing operations. This flexibility and adaptivity of the MP coder makes it appropriate for asymmetric applications with heterogeneous end user terminals. Rosa M. Figueras i Ventura, Pierre Vandergheynst, Pascal Frossard, Andrea Cavallaro |
ICASSP (3) | 4 |
| 2004 | Segmentation-driven perceptual quality metricsabstractWe present a full-reference and a no-reference perceptual video quality metric that incorporate both low-level and high-level aspects of vision. Low-level aspects include color perception, contrast sensitivity, masking as well as artifact analysis. High-level aspects take into account the cognitive behavior of an observer when watching a video by means of semantic segmentation. Using the special case of semantic face segmentation, we evaluate the proposed segmentation-driven perceptual quality metrics using a range of test sequences and demonstrate an improvement of their prediction performance. Andrea Cavallaro, Stefan Winkler 0001 |
ICIP | 1 |
| 2004 | Cast shadow segmentation using invariant color features
Elena Salvador, Andrea Cavallaro, Touradj Ebrahimi |
Comput. Vis. Image Underst. | 2 |
| 2003 | Semantic segmentation and description for video transcodingabstractWe present an automatic content-based video transcoding algorithm, which is based on how humans perceive visual information. The transcoder support multiple video objects and their description. First the video is decomposed into meaningful objects through semantic segmentation. Then the transcoder adapts its behavior to code relevant (foreground) and non-relevant objects differently. Both objects-based and frame-based encoders are combined with semantic segmentation. Experimental results show that the use of semantics and description prior to transcoding reduces the bandwidth requirements and makes it possible to adapt the video representation to limited network and terminal device capabilities still retaining the essential information. Andrea Cavallaro, Olivier Steiger, Touradj Ebrahimi |
ICME | 1 |
| 2003 | Object-based video: extraction tools, evaluation metrics, and applications
Andrea Cavallaro, Touradj Ebrahimi |
VCIP | 1 |
| 2002 | Objective evaluation of segmentation quality using spatio-temporal contextabstractIn this paper, we propose an automatic method for the objective evaluation of segmentation results. The method is based on computing the deviation of the segmentation results from a reference segmentation. The discrepancy between two results is weighted based on spatial and temporal contextual information, by taking into account the way humans perceive visual information. The metric is useful for applications where the final judge of the quality is a human observer or the results of segmentation are otherwise processed in a human-like fashion. The proposed evaluation has been applied both to automatically provide a ranking among different segmentation algorithms and to optimally set the parameters of a given algorithm. Andrea Cavallaro, Elisa Drelie Gelasca, Touradj Ebrahimi |
ICIP (3) | 1 |
| 2002 | Accurate video object segmentation through change detectionabstractWe propose an algorithm for the accurate extraction of video objects from color sequences. The semantics defining the video objects is motion, and the extraction algorithm is based on change detection. The color difference between frames is modeled so as to separate the contributions caused by sensor noise and illumination variations from those caused by meaningful objects. Sensor noise is eliminated by using a probability-based classification, and local illumination variations are removed using a knowledge-based approach that is formulated as a hypothesize-and-test scheme. Experimental results show that the proposed method provides accurate contours of multiple deformable objects, thus providing a reliable input to object-based applications such as those supported by the MPEG-4 and MPEG-7 standards. Andrea Cavallaro, Touradj Ebrahimi |
ICME (1) | 1 |
| 2002 | Multiple video object tracking in complex scenesabstractWe present an automatic video object tracking algorithm capable of dealing with multiple simultaneous objects. The tracking is based on interactions between high-level and low-level image analysis results. The high-level result is a partition defining video objects, and the low-level result is a partition formed by homogeneous regions. For each region, a set of characteristic descriptors is produced. These region descriptors, and not regions themselves, are used to track the regions (and thus the objects) along time. Track management issues such as appearance and disappearance of objects, splitting and partial occlusions are resolved through interactions between regions and objects. Defining the tracking based on the parts of objects, identified by region segmentation, has led to a flexible technique that exploits the nature of the video object tracking problem. Experimental results show that the proposed method is able to track multiple rigid and deformable objects in indoor and outdoor scenes. Andrea Cavallaro, Olivier Steiger, Touradj Ebrahimi |
ACM Multimedia | 1 |
| 2002 | MPEG-7 description of generic video objects for scene reconstruction
Olivier Steiger, Andrea Cavallaro, Touradj Ebrahimi |
VCIP | 2 |
| 2001 | Shadow identification and classification using invariant color modelsabstractA novel approach to shadow detection is presented. The method is based on the use of invariant color models to identify and to classify shadows in digital images. The procedure is divided into two levels: first, shadow candidate regions are extracted; then, by using the invariant color features, shadow candidate pixels are classified as self shadow points or as cast shadow points. The use of invariant color features allows a low complexity of the classification stage. Experimental results show that the method succeeds in detecting and classifying shadows within the environmental constrains assumed as hypotheses, which are less restrictive than state-of-the-art methods with respect to illumination conditions and the scene's layout. Elena Salvador, Andrea Cavallaro, Touradj Ebrahimi |
ICASSP | 2 |
| 2001 | Video object extraction based on adaptive background and statistical change detection
Andrea Cavallaro, Touradj Ebrahimi |
VCIP | 1 |