EDBT 2026 Demo / reviewers in the wild / expert
Andrea Prati 0001
dblp:p/AndreaPrati
· DBLP profile ↗
70ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-1211-529XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 32 · 3 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mcgm-styler: free-form styler for mask conditional text-to-image generative modelabstractAbstract Generative models for text-to-image synthesis have made significant advancements in recent years, enabling the creation of highly detailed and stylistically diverse images. In this work, we introduce MCGM-Styler, an extension of our previous MCGM as reported (MCGM: Mask conditional text-to-image generative model, 2024) model, which generates images based on masked conditions to specify the action pose of subjects in a source image. Our key contribution is the addition of a new training step that enables the model to also perform style transfer, allowing it to generate images that not only force the pose by mask condition, but also adhere to both single and multiple target artistic styles. Unlike traditional approaches that require large datasets, MCGM-Styler is trained on a single image, making it highly efficient and adaptable. The model can handle scenes with one or multiple subjects, generating coherent and stylistically consistent outputs so the user can generate any subject(s) in any pose and style or generate an image with a mix of different styles. We evaluate our approach against existing works, particularly the DreamStyler model (Ahn et al. in Proc. AAAI Conf. Artif. Intell. 38:674–681), a state-of-the-art method for style transfer. Our results demonstrate that the MCGM-Styler achieves superior performance in preserving not only the pose of the concept, but also the style fidelity, highlighting its effectiveness in controllable image generation. Rami Skaik, Leonardo Rossi, Tomaso Fontanini, Andrea Prati 0001 |
Vis. Comput. | 4 |
| 2025 | Mamba-ST: State Space Model for Efficient Style TransferabstractThe goal of style transfer is, given a content image and a style source, generating a new image preserving the content but with the artistic representation of the style source. Most of the state-of-the-art architectures use transformers or diffusion-based models to perform this task, despite the heavy computational burden that they require. In particular, transformers use self- and cross-attention layers which have large memory footprint, while diffusion models require high inference time. To overcome the above, this paper explores a novel design of Mamba, an emergent State-Space Model (SSM), called Mamba-ST, to perform style transfer. To do so, we adapt Mamba linear equation to simulate the behavior of cross-attention layers, which are able to combine two separate embeddings into a single output, but drastically reducing memory usage and time complexity. We modified the Mamba's inner equations so to accept inputs from, and combine, two separate data streams. To the best of our knowledge, this is the first attempt to adapt the equations of SSMs to a vision task like style transfer without requiring any other module like cross-attention or custom normalization layers. An extensive set of experiments demonstrates the superiority and efficiency of our method in performing style transfer compared to transformers and diffusion models. Results show improved quality in terms of both ArtFID and FID metrics. Code is available at https://github.com/FilippoBotti/MambaST. Filippo Botti, Alex Ergasti, Leonardo Rossi, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati 0001 |
WACV | 7 |
| 2025 | Swin2-MoSE: A new single image supersolution model for remote sensingabstractAbstract Due to the limitations of current optical and sensor technologies and the high cost of updating them, the spectral and spatial resolution of satellites may not always meet desired requirements. For these reasons, Remote‐Sensing Single‐Image Super‐Resolution (RS‐SISR) techniques have gained significant interest. In this paper, Swin2‐MoSE model is proposed, an enhanced version of Swin2SR. The model introduces MoE‐SM, an enhanced Mixture‐of‐Experts (MoE) to replace the Feed‐Forward inside all Transformer block. MoE‐SM is designed with Smart‐Merger, and new layer for merging the output of individual experts, and with a new way to split the work between experts, defining a new per‐example strategy instead of the commonly used per‐token one. Furthermore, it is analyzed how positional encodings interact with each other, demonstrating that per‐channel bias and per‐head bias can positively cooperate. Finally, the authors propose to use a combination of Normalized‐Cross‐Correlation (NCC) and Structural Similarity Index Measure (SSIM) losses, to avoid typical MSE loss limitations. Experimental results demonstrate that Swin2‐MoSE outperforms any Swin derived models by up to 0.377–0.958 dB (PSNR) on task of , and resolution‐upscaling ( and OLI2MSI datasets). It also outperforms SOTA models by a good margin, proving to be competitive and with excellent potential, especially for complex tasks. Additionally, an analysis of computational costs is also performed. Finally, the efficacy of Swin2‐MoSE is shown, applying it to a semantic segmentation task (SeasoNet dataset). Code and pretrained are available on https://github.com/IMPLabUniPr/swin2‐mose/tree/official_code Leonardo Rossi, Vittorio Bernuzzi, Tomaso Fontanini, Massimo Bertozzi, Andrea Prati 0001 |
IET Image Process. | 5 |
| 2025 | MARS: Paying More Attention to Visual Attributes for Text-Based Person SearchabstractText-Based Person Search (TBPS) is a problem that gained significant interest within the research community. The task is that of retrieving one or more images of a specific individual based on a textual description. The multi-modal nature of the task requires learning representations that bridge text and image data within a shared latent space. Existing TBPS systems face two major challenges. One is defined as inter-identity noise that is due to the inherent vagueness and imprecision of text descriptions, and it indicates how descriptions of visual attributes can be generally associated to different people; the other is the intra-identity variations, which are all those nuisances, e.g., pose, illumination, that can alter the visual appearance of the same textual attributes for a given subject. To address these issues, this article presents a novel TBPS architecture named Mae-Attribute-Relation-Sensitive (MARS), which enhances current state-of-the-art models by introducing two key components: a Visual Reconstruction Loss and an Attribute Loss. The former employs a Masked AutoEncoder trained to reconstruct randomly masked image patches with the aid of the textual description. In doing so the model is encouraged to learn more expressive representations and textual–visual relations in the latent space. The attribute loss, instead, balances the contribution of different types of attributes, defined as adjective–noun chunks of text. This loss ensures that every attribute is taken into consideration in the person retrieval process. Extensive experiments on three commonly used datasets, namely CUHK-PEDES, ICFG-PEDES, and RSTPReid, report performance improvements, with significant gains in the Mean Average Precision (mAP) metric w.r.t. the current state of the art. Code will be available at https://github.com/ErgastiAlex/MARS . Alex Ergasti, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Memory-Augmented Online Video Anomaly DetectionabstractThe ability to understand the surrounding scene is of paramount importance for Autonomous Vehicles (AVs). This paper presents a system capable to work in an online fashion, giving an immediate response to the arise of anomalies surrounding the AV, exploiting only the videos captured by a dash-mounted camera. Our architecture, called MOVAD, relies on two main modules: a Short-Term Memory Module to extract information related to the ongoing action, implemented by a Video Swin Transformer (VST), and a Long-Term Memory Module injected inside the classifier that considers also remote past information and action context thanks to the use of a Long-Short Term Memory (LSTM) network. The strengths of MOVAD are not only linked to its excellent performance, but also to its straightforward and modular architecture, trained in a end-to-end fashion with only RGB frames with as less assumptions as possible, which makes it easy to implement and play with. We evaluated the performance of our method on Detection of Traffic Anomaly (DoTA) dataset, a challenging collection of dash-mounted camera videos of accidents. After an extensive ablation study, MOVAD is able to reach an AUC score of 82.17%, surpassing the current state-of-the-art by +2.87 AUC. Our code and pretrained are available online on https://github.com/IMPLabUniPr/movad/tree/movad_vad Leonardo Rossi, Vittorio Bernuzzi, Tomaso Fontanini, Massimo Bertozzi, Andrea Prati 0001 |
ICASSP | 5 |
| 2023 | FrankenMask: Manipulating semantic masks with transformers for face parts editingabstractIn this paper, we propose FrankenMask, a novel framework that allows swapping and rearranging face parts in semantic masks for automatic editing of shape-related facial attributes. This is a novel yet challenging task as substituting face parts in a semantic mask requires to account for possible spatial misalignment and the adaptation of surrounding regions. We obtain such a feature by combining a Transformer encoder to learn the spatial relationships of facial parts, with an encoder–decoder architecture, which reconstructs a complete mask from the composition of local parts. Reconstruction and attribute classification results demonstrate the effective synthesis of facial images, while showing the generation of accurate and plausible facial attributes. Code is available at https://github.com/TFonta/FrankenMask_semantic. Tomaso Fontanini, Claudio Ferrari, Giuseppe Lisanti, Leonardo Galteri, Stefano Berretti, Massimo Bertozzi, Andrea Prati 0001 |
Pattern Recognit. Lett. | 7 |
| 2023 | Unsupervised Discovery and Manipulation of Continuous Disentangled Factors of VariationabstractLearning a disentangled representation of a distribution in a completely unsupervised way is a challenging task that has drawn attention recently. In particular, much focus has been put in separating factors of variation (i.e., attributes) within the latent code of a Generative Adversarial Network (GAN). Achieving that permits control of the presence or absence of those factors in the generated samples by simply editing a small portion of the latent code. Nevertheless, existing methods that perform very well in a noise-to-image setting often fail when dealing with a real data distribution, i.e., when the discovered attributes need to be applied to real images. However, some methods are able to extract and apply a style to a sample but struggle to maintain its content and identity, while others are not able to locally apply attributes and end up achieving only a global manipulation of the original image. In this article, we propose a completely (i.e., truly ) unsupervised method that is able to extract a disentangled set of attributes from a data distribution and apply them to new samples from the same distribution by preserving their content. This is achieved by using an image-to-image GAN that maps an image and a random set of continuous attributes to a new image that includes those attributes. Indeed, these attributes are initially unknown and they are discovered during training by maximizing the mutual information between the generated samples and the attributes’ vector. Finally, the obtained disentangled set of continuous attributes can be used to freely manipulate the input samples. We prove the effectiveness of our method over a series of datasets and show its application on various tasks, such as attribute editing, data augmentation, and style transfer. Tomaso Fontanini, Luca Donati, Massimo Bertozzi, Andrea Prati 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Aspect-Based Emotion Analysis and Multimodal Coreference: A Case Study of Customer Comments on Adidas Instagram PostsabstractWhile aspect-based sentiment analysis of user-generated content has received a lot of attention in the past years, emotion detection at the aspect level has been relatively unexplored. Moreover, given the rise of more visual content on social media platforms, we want to meet the ever-growing share of multimodal content. In this paper, we present a multimodal dataset for Aspect-Based Emotion Analysis (ABEA). Additionally, we take the first steps in investigating the utility of multimodal coreference resolution in an ABEA framework. The presented dataset consists of 4,900 comments on 175 images and is annotated with aspect and emotion categories and the emotional dimensions of valence and arousal. Our preliminary experiments suggest that ABEA does not benefit from multimodal coreference resolution, and that aspect and emotion classification only requires textual information. However, when more specific information about the aspects is desired, image recognition could be essential. Luna De Bruyne, Akbar Karimi 0001, Orphée De Clercq, Andrea Prati 0001, Véronique Hoste |
LREC | 4 |
| 2022 | Self-Balanced R-CNN for instance segmentation
Leonardo Rossi, Akbar Karimi 0001, Andrea Prati 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Recursively Refined R-CNN: Instance Segmentation with Self-RoI Rebalancing
Leonardo Rossi, Akbar Karimi 0001, Andrea Prati 0001 |
CAIP (1) | 3 |
| 2021 | A Real-Time Approach for Automatic Food Quality Assessment Based on Shape AnalysisabstractProducts sorting is a task of paramount importance for many countries’ agricultural industry. An accurate quality check assures that good products are not wasted, and rotten, broken and bent food are properly discarded, which is extremely important for food production chains. Such products sorting and quality controls are often performed with consolidated instruments, since simple systems are easier to maintain, validate, and they speed up the processing in terms of production line speed and products per second. Moreover, industries often lack advanced formation, required for more sophisticated solutions. As a result, the sorting task for many food products is mainly done by color information only. Sorting machines typically detect the color response of products to specific LEDs with various light wavelengths. Unfortunately, a color check is often not enough to detect some very common defects. The shape of a product, instead, reveals many important defects and is highly reliable in detecting external objects mixed with food. Also, shape can be used to take detailed measurements of a product, such as its area, length, width, anisotropy, etc. This paper proposes a complete treatment of the problem of sorting food by its shape. It treats real-world problems such as accuracy, execution time, latency and it provides an overview of a full system used on state-of-the-art measurement machines. Luca Donati, Eleonora Iotti, Andrea Prati 0001 |
Int. J. Comput. Intell. Appl. | 3 |
| 2021 | LSH kNN graph for diffusion on image retrieval
Federico Magliani, Andrea Prati 0001 |
Inf. Retr. J. | 2 |
| 2021 | Bag of indexes: a multi-index scheme for efficient approximate nearest neighbor search
Federico Magliani, Tomaso Fontanini, Andrea Prati 0001 |
Multim. Tools Appl. | 3 |
| 2020 | Adversarial Training for Aspect-Based Sentiment Analysis with BERTabstractAspect-Based Sentiment Analysis (ABSA) studies the extraction of sentiments and their targets. Collecting labeled data for this task in order to help neural networks generalize better can be laborious and time-consuming. As an alternative, similar data to the real-world examples can be produced artificially through an adversarial process which is carried out in the embedding space. Although these examples are not real sentences, they have been shown to act as a regularization method which can make neural networks more robust. In this work, we fine- tune the general purpose BERT and domain specific post-trained BERT (BERT-PT) using adversarial training. After improving the results of post-trained BERT with different hyperparameters, we propose a novel architecture called BERT Adversarial Training (BAT) to utilize adversarial training for the two major tasks of Aspect Extraction and Aspect Sentiment Classification in sentiment analysis. The proposed model outperforms the general BERT as well as the in-domain post-trained BERT in both tasks. To the best of our knowledge, this is the first study on the application of adversarial training in ABSA. The code is publicly available on a GitHub repository at https://github.com/IMPLabUniPr/Adversarial-Training-for-ABSA. Akbar Karimi 0001, Leonardo Rossi, Andrea Prati 0001 |
ICPR | 3 |
| 2020 | A Novel Region of Interest Extraction Layer for Instance SegmentationabstractGiven the wide diffusion of deep neural network architectures for computer vision tasks, several new applications are nowadays more and more feasible. Among them, a particular attention has been recently given to instance segmentation, by exploiting the results achievable by two-stage networks (such as Mask R-CNN or Faster R-CNN), derived from R-CNN. In these complex architectures, a crucial role is played by the Region of Interest (RoI) extraction layer, devoted to extracting a coherent subset of features from a single Feature Pyramid Network (FPN) layer attached on top of a backbone. This paper is motivated by the need to overcome the limitations of existing RoI extractors which select only one (the best) layer from FPN. Our intuition is that all the layers of FPN retain useful information. Therefore, the proposed layer (called Generic RoI Extractor - GRoIE) introduces non-local building blocks and attention mechanisms to boost the performance. A comprehensive ablation study at component level is conducted to find the best set of algorithms and parameters for the GRoIE layer. Moreover, GRoIE can be integrated seamlessly with every two-stage architecture for both object detection and instance segmentation tasks. Therefore, the improvements brought about by the use of GRoIE in different state-of-the-art architectures are also evaluated. The proposed layer leads up to gain a 1.1% AP improvement on bounding box detection and 1.7% AP improvement on instance segmentation. The code is publicly available on GitHub repository at https://github.com/IMPLabUniPr/mmdetection/tree/groie_dev. Leonardo Rossi, Akbar Karimi 0001, Andrea Prati 0001 |
ICPR | 3 |
| 2020 | MetalGAN: Multi-domain label-less image synthesis using cGANs and meta-learning
Tomaso Fontanini, Eleonora Iotti, Luca Donati, Andrea Prati 0001 |
Neural Networks | 4 |
| 2019 | Multi-target Tracking in Multiple Non-overlapping Cameras Using Fast-Constrained Dominant Sets
Yonatan Tariku, Eyasu Zemene Mequanint, Andrea Prati 0001, Marcello Pelillo, Mubarak Shah |
Int. J. Comput. Vis. | 3 |
| 2019 | A complete hand-drawn sketch vectorization framework
Luca Donati, Simone Cesano, Andrea Prati 0001 |
Multim. Tools Appl. | 3 |
| 2019 | Large-Scale Image Geo-Localization Using Dominant SetsabstractThis paper presents a new approach for the challenging problem of geo-localization using image matching in a structured database of city-wide reference images with known GPS coordinates. We cast the geo-localization as a clustering problem of local image features. Akin to existing approaches to the problem, our framework builds on low-level features which allow local matching between images. For each local feature in the query image, we find its approximate nearest neighbors in the reference set. Next, we cluster the features from reference images using Dominant Set clustering, which affords several advantages over existing approaches. First, it permits variable number of nodes in the cluster, which we use to dynamically select the number of nearest neighbors for each query feature based on its discrimination value. Second, this approach is several orders of magnitude faster than existing approaches. Thus, we obtain multiple clusters (different local maximizers) and obtain a robust final solution to the problem using multiple weak solutions through constrained Dominant Set clustering on global image features, where we enforce the constraint that the query image must be included in the cluster. This second level of clustering also bypasses heuristic approaches to voting and selecting the reference image that matches to the query. We evaluate the proposed framework on an existing dataset of 102k street view images as well as a new larger dataset of 300k images, and show that it outperforms the state-of-the-art by 20 and 7 percent, respectively, on the two datasets. Eyasu Zemene Mequanint, Yonatan Tariku, Haroon Idrees, Andrea Prati 0001, Marcello Pelillo, Mubarak Shah |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | A technology platform for automatic high-level tennis game analysis
Vito Renò, Nicola Mosca, Massimiliano Nitti, Tiziana D'Orazio, Cataldo Guaragnella, Donato Campagnoli, Andrea Prati 0001, Ettore Stella |
Comput. Vis. Image Underst. | 7 |
| 2016 | Simultaneous clustering and outlier detection using dominant setsabstractWe present a unified approach for simultaneous clustering and outlier detection in data. We utilize some properties of a family of quadratic optimization problems related to dominant sets, a well-known graph-theoretic notion of a cluster which generalizes the concept of a maximal clique to edge-weighted graphs. Unlike most (all) of the previous techniques, in our framework the number of clusters arises intuitively and outliers are obliterated automatically. The resulting algorithm discovers both parameters from the data. Experiments on real and on large scale synthetic dataset demonstrate the effectiveness of our approach and the utility of carrying out both clustering and outlier detection in a concurrent manner. Eyasu Zemene Mequanint, Yonatan Tariku, Andrea Prati 0001, Marcello Pelillo |
ICPR | 3 |
| 2016 | Multi-object tracking using dominant setsabstractMulti‐object tracking is an interesting but challenging task in the field of computer vision. Most previous works based on data association techniques merely take into account the relationship between detection responses in a locally limited temporal domain, which makes them inherently prone to identity switches and difficulties in handling long‐term occlusions. In this study, a dominant set clustering based tracker is proposed, which formulates the tracking task as a problem of finding dominant sets in an auxiliary edge weighted graph. Unlike most techniques which are limited in temporal locality (i.e. few frames are considered), the authors utilised a pairwise relationships (in appearance and position) between different detections across the whole temporal span of the video for data association in a global manner. Meanwhile, temporal sliding window technique is utilised to find tracklets and perform further merging on them. The authors’ robust tracklet merging step renders the tracker to long term occlusions with more robustness. The authors present results on three different challenging datasets (i.e. PETS2009‐S2L1, TUD‐standemitte and ETH dataset (‘sunny day’ sequence)), and show significant improvements compared with several state‐of‐art methods. Yonatan Tariku, Eyasu Zemene Mequanint, Marcello Pelillo, Andrea Prati 0001 |
IET Comput. Vis. | 4 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 19 |
| 2015 | Editorial introduction to the special issue on "Image Understanding for Real-World Distributed Video Networks" - Computer Vision and Image Understanding Journal
Bir Bhanu, Andrea Prati 0001, Faisal Z. Qureshi |
Comput. Vis. Image Underst. | 2 |
| 2015 | A deep analysis on age estimation
Ivan Huerta Casado, Carles Fernández, Carlos Segura, Javier Hernando, Andrea Prati 0001 |
Pattern Recognit. Lett. | 5 |
| 2014 | A fast and effective ellipse detector for embedded vision applications
Michele Fornaciari, Andrea Prati 0001, Rita Cucchiara |
Pattern Recognit. | 2 |
| 2013 | A people counting system for business analyticsabstractThis paper deals with people counting in stores for business analytics using stereo vision. Among the several problems in this type of applications, two are the most relevant for our purposes: the management of occlusions and the distinction between adult people (potential customers) and other objects (children, trolleys, strollers, animals, etc.). The proposed solution uses a novel approach for object detection (based on background suppression on a so-called “depth bird-eye view” and the clustering on the 3D point cloud by means of mean shift with a cylindrical kernel) followed by an adult people classifier which exploits a fitness measure with respect to a cylindrical human body model. The fitness is computed using Montecarlo sampling to estimate the volume occupation. Experiments are conducted on two real setups (including a store in a normal day of activity) and compared with a previous work. The results demonstrate the accuracy of the proposed solution. Carlo Pane, Marco Gasparini, Andrea Prati 0001, Giovanni Gualdi, Rita Cucchiara |
AVSS | 3 |
| 2013 | Editorial to the 'pattern recognition and artificial intelligence for human behaviour analysis' special sectionabstractThe Pattern Recognition (PR) and Artificial Intelligence (AI) scientific communities have shared knowledge and effort in order to obtain more effective solutions for many different research areas. However, although the techniques and approaches are somewhat similar, the two communities often tackle problems from rather different perspectives. In the first paper ‘Social Interactions by Visual Focus of Attention in a Three-Dimensional Environment', by Bazzani, Tosato, Cristani, Farenzena, Paggetti, Menegaz and Murino, a novel approach to social interaction discovery is presented; instead of using global or local appearance features, the authors exploit the Subjective View Frustum, which approximates the visual field of a person in a three-dimensional representation of the scene. The main contribution of the second paper 'Human action recognition using an ensemble of body-part detectors', by Chakraborty, Bagdanov, Gonzalez and Roca, is to transform the problem of action recognition into that of recognising the distinctive motion of specific body parts, for instance, the legs for walking, the hands for boxing, etc. The intuition behind the approach is that several human actions can be described more compactly and effectively by considering only the relevant motions of the body parts actually performing the actions. We hope you enjoy the special section. Luca Iocchi is Associate Professor at Sapienza University of Rome, Italy. His main research interests are in the areas of cognitive robotics, action planning, multi-robot coordination, robot perception, robot learning, sensor data fusion. He is being involved in several projects aiming at developing intelligent robotic systems and intelligent surveillance systems. He is active in many conferences and journals related to artificial intelligence and robotics, as well as in the organisation of scientific competitions, such as RoboCup@Home. Andrea Prati is Associate Professor at the University IUAV of Venice. He collaborated in several research projects at regional, national and international level. His research interests belong to different themes, from embedded devices for sensor networks in computer vision applications, to robotic vision, to multimedia, to performance analysis for multimedia computers. However, his main research activity is on video-surveillance topics: object tracking in distributed, multi-camera environments; analysis and removal of the shadows; behaviour analysis through trajectory classification. Andrea Prati is author of more than 130 papers in international journals and conference proceedings; he has been invited speaker and reviewer for many international journals. He is also a member of the Editorial Board of Journal of Optical Engineering (SPIE) and Journal on Ambient Intelligence and Smart Environments (IOS Press). He has also been the Program Chair of ICIAP 2007. He has been the PC of ACM/IEEE Intl Conf on Distributed Smart Cameras (ICDSC) in 2011 and 2012, and will be for 2013 edition in Palm Springs, CA (USA). He is also organising as General Chair the 2014 ICDSC edition in Venice. He is a senior member of IEEE, and a member of ACM and GIRPR. Roberto Vezzani is an Assistant Professor at University of Modena and Reggio Emilia and he works in the Engineering Department 'Enzo Ferrari'. His research interests mainly belong to video surveillance systems, with particular focus on behaviour analysis, people tracking and re-identification. He is the author of the ViSOR web repository, an online platform for sharing research videos and annotations developed within the European project VidiVideo. He was the technical coordinator of the European Project THIS, for transport hub intelligent video surveillance. He is author of more than 50 papers on international journals and conferences. Luca Iocchi, Andrea Prati 0001, Roberto Vezzani |
Expert Syst. J. Knowl. Eng. | 2 |
| 2012 | Real-time object detection and localization with SIFT-based clustering
Paolo Piccinini, Andrea Prati 0001, Rita Cucchiara |
Image Vis. Comput. | 2 |
| 2012 | Multistage Particle Windows for Fast and Accurate Object DetectionabstractThe common paradigm employed for object detection is the sliding window (SW) search. This approach generates grid-distributed patches, at all possible positions and sizes, which are evaluated by a binary classifier: The tradeoff between computational burden and detection accuracy is the real critical point of sliding windows; several methods have been proposed to speed up the search such as adding complementary features. We propose a paradigm that differs from any previous approach since it casts object detection into a statistical-based search using a Monte Carlo sampling for estimating the likelihood density function with Gaussian kernels. The estimation relies on a multistage strategy where the proposal distribution is progressively refined by taking into account the feedback of the classifiers. The method can be easily plugged into a Bayesian-recursive framework to exploit the temporal coherency of the target objects in videos. Several tests on pedestrian and face detection, both on images and videos, with different types of classifiers (cascade of boosted classifiers, soft cascades, and SVM) and features (covariance matrices, Haar-like features, integral channel features, and histogram of oriented gradients) demonstrate that the proposed method provides higher detection rates and accuracy as well as a lower computational burden w.r.t. sliding window detection. Giovanni Gualdi, Andrea Prati 0001, Rita Cucchiara |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | A multi-stage pedestrian detection using monolithic classifiersabstractDespite the many efforts in finding effective feature sets or accurate classifiers for people detection, few works have addressed ways for reducing the computational burden introduced by the sliding window paradigm. This paper proposes a multi-stage procedure for refining the search for pedestrians using the HOG features and the monolithic SVM classifier. The multi-stage procedure is based on particle-based estimation of pdfs and exploits the margin provided by the classifier to draw more particles on the areas where the classifier's response is higher. This iterative algorithm achieves the same accuracy than sliding window using less particles (and thus being more efficient) and, conversely, is more accurate when configured to work at the same computational load. Experimental results on publicly available datasets demonstrate that this method, previously proposed for boosted classifiers only, can be successfully applied to monolithic classifiers. Giovanni Gualdi, Andrea Prati 0001, Rita Cucchiara |
AVSS | 2 |
| 2011 | Detecting anomalies in people's trajectories using spectral graph analysis
Simone Calderara, Uri Heinemann, Andrea Prati 0001, Rita Cucchiara, Naftali Tishby |
Comput. Vis. Image Underst. | 3 |
| 2011 | Mixtures of von Mises Distributions for People Trajectory Shape AnalysisabstractPeople trajectory analysis is a recurrent task in many pattern recognition applications, such as surveillance, behavior analysis, video annotation, and many others. In this paper, we propose a new framework for analyzing trajectory shape, invariant to spatial shifts of the people motion in the scene. In order to cope with the noise and the uncertainty of the trajectory samples, we propose to describe the trajectories as a sequence of angles modeled by distributions of circular statistics, i.e., a mixture of von Mises (MovM) distributions. To deal with MovM, we define a new specific expectation-maximization (EM) algorithm for estimating the parameters and derive a closed form of the Bhattacharyya distance between single von Mises pdfs. Trajectories are then modeled with a sequence of symbols, corresponding to the most suitable distribution in the mixture, and compared each other after a global alignment procedure to cope with trajectories of different lengths. The trajectories in the training set are clustered according to their shape similarity in an off-line phase, and testing trajectories are then classified with a specific on-line EM, based on sufficient statistics. The approach is particularly suitable for classifying people trajectories in video surveillance, searching for abnormal (i.e., infrequent) paths. Tests on synthetic and real data are provided with also a complete comparison with other circular statistical and alignment methods. Simone Calderara, Andrea Prati 0001, Rita Cucchiara |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Multi-stage Sampling with Boosting Cascades for Pedestrian Detection in Images and Videos
Giovanni Gualdi, Andrea Prati 0001, Rita Cucchiara |
ECCV (6) | 2 |
| 2010 | Alignment-Based Similarity of People Trajectories Using Semi-directional StatisticsabstractThis paper presents a method for comparing people trajectories for video surveillance applications, based on semi-directional statistics. In fact, the modelling of a trajectory as a sequence of angles, speeds and time lags, requires the use of a statistical tool capable to jointly consider periodic and linear variables. Our statistical method is compared with two state-of-the-art methods. Simone Calderara, Andrea Prati 0001, Rita Cucchiara |
ICPR | 2 |
| 2009 | Learning People Trajectories Using Semi-directional StatisticsabstractThis paper proposes a system for people trajectory shape analysis by exploiting a statistical approach which accounts for sequences of both directional (the directions of the trajectory) and linear (the speeds) data. A semi-directional distribution (AWLG - Approximated Wrapped and Linear Gaussian) is used with a mixture to find main directions and speeds. A variational version of the mutual information criterion is proposed to prove the statistical dependency of the data. Then, in order to compare data sequences, we define an inexact method with a Kullback-Leibler-based distance measure and employ a global alignment technique is to handle sequences of different lengths and with local shifts or deformations. A comprehensive analysis of variable dependency and parameter estimation techniques are reported and evaluated on both synthetic and real data sets. Simone Calderara, Andrea Prati 0001, Rita Cucchiara |
AVSS | 2 |
| 2009 | A fast multi-model approach for object duplicate extractionabstractThis paper presents an innovative approach for localizing and segmenting duplicate objects for industrial applications. The working conditions are challenging, with complex heavily-occluded objects, arranged at random in the scene. To account for high flexibility and processing speed, this approach exploits SIFT keypoint extraction and mean shift clustering to efficiently partition the correspondences between the object model and the duplicates onto the different object instances. The re-projection (by means of an Euclidean transform) of some delimiting points onto the current image is used to segment the object shapes. This procedure is compared in terms of accuracy with existing homography-based solutions which make use of RANSAC to eliminate outliers in the homography estimation. Moreover, in order to improve the extraction in the case of reflective or transparent objects, multiple object models are used and fused together. Experimental results on different and challenging kinds of objects are reported. Paolo Piccinini, Andrea Prati 0001, Rita Cucchiara |
WACV | 2 |
| 2008 | Action Signature: A Novel Holistic Representation for Action RecognitionabstractRecognizing different actions with a unique approach can be a difficult task. This paper proposes a novel holistic representation of actions that we called "action signature". This 1D trajectory is obtained by parsing the 2D image containing the orientations of the gradient calculated on the motion feature map called motion-history image. In this way, the trajectory is a sketch representation of how the object motion varies in time. A robust statistical framework based on mixtures of von Mises distributions and dynamic programming for sequence alignment are used to compare and classify actions/trajectories. The experimental results show a rather high accuracy in distinguishing quite complicated actions, such as drinking, jumping, or abandoning an object. Simone Calderara, Rita Cucchiara, Andrea Prati 0001 |
AVSS | 3 |
| 2008 | Commentary Paper 2 on "A Localized Approach to Abandoned Luggage Detection with Foreground-Mask Sampling"abstractThis paper proposes an approach to abandoned luggage detection that mimics human behavior in monitoring a scene: first the abandoned luggage is detected through a foreground-mask sampling; then people nearby the detected object are tracked to eventually associate the left object with its possessor; ultimately, MAP principle is applied to perform reasoning. Andrea Prati 0001 |
AVSS | 1 |
| 2008 | Commentary Paper 1 on "Automatic Detection of Adverse Weather Conditions in Traffic Scenes"abstractThis paper describes two solutions for detecting snow and fog, respectively, in traffic scenes. The former weather condition is detected by modeling the video luminance with a mixture of Gaussians (MoG), while the latter is detected analyzing the Fourier harmonic frequencies. Andrea Prati 0001 |
AVSS | 1 |
| 2008 | Using circular statistics for trajectory shape analysisabstractThe analysis of patterns of movement is a crucial task for several surveillance applications, for instance to classify normal or abnormal people trajectories on the basis of their occurrence. This paper proposes to model the shape of a single trajectory as a sequence of angles described using a mixture of Von Mises (MoVM) distribution. A complete EM (expectation maximization) algorithm is derived for MoVM parameters estimation and an on-line version proposed to meet real time requirement. Maximum-A-Posteriori is used to encode the trajectory as a sequence of symbols corresponding to the MoVM components. Iterative k-medoids clustering groups trajectories in a variable number of similarity classes. The similarity is computed aligning (with dynamic programming) two sequences and considering as symbol-to-symbol distance the Bhattacharyya distance between von Mises distributions. Extensive experiments have been performed on both synthetic and real data. Andrea Prati 0001, Simone Calderara, Rita Cucchiara |
CVPR | 1 |
| 2008 | ACM multimedia 2008: 1st workshop on vision networks for behavior analysis (VNBA 2008)abstractThe VNBA workshop marks a new era of the successful series of the Video Surveillance and Sensor Networks (VSSN) workshops, held until 2006 in conjunction with the ACM Multimedia conference. This new version of the workshop inherits from VSSN the experienced Technical Program Committee as well as the interests of its community, but shifts the focus to cover higher level topics and applications under the common framework of "behaviour analysis", hence aiming to adapt to the evolved directions of interest in the field, and reaching out to other research communities with overlapping interests. Hamid K. Aghajan, Andrea Prati 0001 |
ACM Multimedia | 2 |
| 2008 | HECOL: Homography and epipolar-based consistent labeling for outdoor park surveillance
Simone Calderara, Andrea Prati 0001, Rita Cucchiara |
Comput. Vis. Image Underst. | 2 |
| 2008 | Bayesian-Competitive Consistent Labeling for People SurveillanceabstractThis paper presents a novel and robust approach to consistent labeling for people surveillance in multi-camera systems. A general framework scalable to any number of cameras with overlapped views is devised. An off-line training process automatically computes ground-plane homography and recovers epipolar geometry. When a new object is detected in any one camera, hypotheses for potential matching objects in the other cameras are established. Each of the hypotheses is evaluated using a prior and likelihood value. The prior accounts for the positions of the potential matching objects, while the likelihood is computed by warping the vertical axis of the new object on the field of view of the other cameras and measuring the amount of match. In the likelihood, two contributions (forward and backward) are considered so as to correctly handle the case of groups of people merged into single objects. Eventually, a maximum-a-posteriori approach estimates the best label assignment for the new object. Comparisons with other methods based on homography and extensive outdoor experiments demonstrate that the proposed approach is accurate and robust in coping with segmentation errors and in disambiguating groups. Simone Calderara, Rita Cucchiara, Andrea Prati 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Video Streaming for Mobile Video SurveillanceabstractMobile video surveillance represents a new paradigm that encompasses, on the one side, ubiquitous video acquisition and, on the other side, ubiquitous video processing and viewing, addressing both computer-based and human-based surveillance. To this aim, systems must provide efficient video streaming with low latency and low frame skipping, even over limited bandwidth networks. This work presents MoSES (MObile Streaming for vidEo Surveillance), an effective system for mobile video surveillance for both PC and PDA clients; it relies over H.264/AVC video coding and GPRS/EDGE-GPRS network. Adaptive control algorithms are employed to achieve the best tradeoff between low latency and good video fluidity. MoSES provides a good-quality video streaming that is used as input to computer-based video surveillance applications for people segmentation and tracking. In this paper new and general-purpose methodologies for streaming performance evaluation are also proposed and used to compare MoSES with existing solutions in terms of different parameters (latency, image quality, video fluidity, and frame losses), as well as in terms of performance in people segmentation and tracking. Giovanni Gualdi, Andrea Prati 0001, Rita Cucchiara |
IEEE Trans. Multim. | 2 |
| 2007 | Detection of abnormal behaviors using a mixture of Von Mises distributionsabstractThis paper proposes the use of a mixture of Von Mises distributions to detect abnormal behaviors of moving people. The mixture is created from an unsupervised training set by exploiting k-medoids clustering algorithm based on Bhattacharyya distance between distributions. The extracted medoids are used as modes in the multi-modal mixture whose weights are the priors of the specific medoid. Given the mixture model a new trajectory is verified on the model by considering each direction composing it as independent. Experiments over a real scenario composed of multiple, partially-overlapped cameras are reported. Simone Calderara, Rita Cucchiara, Andrea Prati 0001 |
AVSS | 3 |
| 2007 | An Open Source Architecture for Low-Latency Video Streaming on PDAsabstractThis paper presents a open-source system for low- latency video streaming on PDAs, specifically addressing mobile video surveillance requirements. The system is based on H.264 and suitably modified to obtain the best trade-off between image quality and video fluidity, working also at very limited bandwidths. Moreover, the used controls allow to keep the number of lost frames very low. A large set of experiments and comparisons have been carried out and the achieved results demonstrate the efficacy and efficiency of our system. Giovanni Gualdi, Andrea Prati 0001, Rita Cucchiara |
ISM | 2 |
| 2007 | A multi-camera vision system for fall detection and alarm generationabstractAbstract: In‐house video surveillance can represent an excellent support for people with some difficulties (e.g. elderly or disabled people) living alone and with a limited autonomy. New hardware technologies and in particular digital cameras are now affordable and they have recently gained credit as tools for (semi‐)automatically assuring people's safety. In this paper a multi‐camera vision system for detecting and tracking people and recognizing dangerous behaviours and events such as a fall is presented. In such a situation a suitable alarm can be sent, e.g. by means of an SMS. A novel technique of warping people's silhouette is proposed to exchange visual information between partially overlapped cameras whenever a camera handover occurs. Finally, a multi‐client and multi‐threaded transcoding video server delivers live video streams to operators/remote users in order to check the validity of a received alarm. Semantic and event‐based transcoding algorithms are used to optimize the bandwidth usage. A two‐room setup has been created in our laboratory to test the performance of the overall system and some of the results obtained are reported. Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani |
Expert Syst. J. Knowl. Eng. | 2 |
| 2007 | Expert environments: machine intelligence methods for ambient intelligenceabstractThe ideas put forward by Donald A. Norman (1999) in his monograph entitled The Invisible Computer can be considered the main source of inspiration of a new research area, called ambient intelligence. Ambient intelligence, commonly abbreviated as AmI, is primarily concerned with human–environment interactions. An environment is seen anthropomorphically, as an intelligent agent able to interact with users, creating for them processes to interpret, inform, communicate and dialogue (Abowd & Mynatt, 2000; Remagnino & Foresti, 2005). The history of ambient intelligence starts in Europe in 2001 with the Fifth European Framework Program. At that time, the IST Program Advisory Group (ISTAG) of the European Commission (Directorate General on Information Society and the Media) introduced the concept of ambient intelligence by publishing the report Scenarios for Ambient Intelligence in 2010 (Ducatel et al., 2001). Since then, ambient intelligence has been recognized in Europe as one of the key concepts related to the information and communication technology society. An updated version of the report was published in 2003 under the title Ambient Intelligence: from Vision to Reality (Ducatel et al., 2003). Ambient intelligence's emphasis is on support to human interactions with the environment, user-friendliness, ubiquitous accessibility etc. and requires competences from many research areas, ranging from computer vision, machine learning, distributed computing and middleware, context awareness systems, sensor networks etc. Enticing illustrative scenarios have been published, in which the user wears technology that communicates with systems and devices present in the environment in order to provide information and receive services. Since its inception, ambient intelligence has inspired the design and implementation of system prototypes for intelligent spaces, developed and tested in controlled environments. The scope of this special issue is to publish innovative ideas on a selection of topics related to ambient intelligence. Cameras and computer vision algorithms, for instance, can be used to unobtrusively acquire awareness of the environment. In the paper entitled ‘A multi-camera vision system for fall detection and alarm generation’, Cucchiara et al. propose a system for using cameras to detect people's falls. The use of cameras is preferred in ambient intelligence to accelerometers or other devices because they are not intrusive, people tend to get used to their presence and people's actions are less affected. In this paper, a system of multiple cameras is used to monitor people's movements in house environments and detect changes in people's posture with the specific goal to identify falls. In the paper entitled ‘Understanding intention of movement from electroencephalograms’, Lakany and Conway analyse people's intentions, using electroencephalogram waves. Their method is non-intrusive and it uses a brain–computer interface. Support vector machines are used to select features and perform classification of intentions, and tests are performed to detect the direction of users' movements. While the previous papers address ambient intelligence from the point of view of sensing technologies and algorithms, the following two papers are mainly focused on knowledge representation for context awareness. The paper entitled ‘Knowledge representation for ambient security’ by Snidaro and Foresti proposes an ontology-based methodology for representing the interaction between the user and the environment with specific reference to security scenarios. Similarly, in the paper entitled ‘Context-aware environments: from specification to implementation’ Reigner et al. are concerned with the problem of implementing a context model for a smart environment. Their paper proposes interesting approaches based on ‘networks of situations’, introducing a comparison of the use of Petri nets and hidden Markov models. Finally, in the paper entitled ‘Collection, storage and application of human knowledge in expert system development’ Balch et al. propose an analysis of the knowledge engineering flow, encompassing knowledge acquisition, representation and inference. This flow analysis is presented and applied to the petroleum industry application domain, and specific software tools that use fuzzy logic are utilized. Paolo Remagnino, Andrea Prati 0001, Gian Luca Foresti, Rita Cucchiara |
Expert Syst. J. Knowl. Eng. | 2 |
| 2006 | Group Detection at Camera Handoff for Collecting People Appearance in Multi-camera SystemsabstractLogging information on moving objects is crucial in video surveillance systems. Distributed multi-camera systems can provide the appearance of objects/people from different viewpoints and at different resolutions, allowing a more complete and precise logging of the information. This is achieved through consistent labeling to correlate collected information of the same person. This paper proposes a novel approach to consistent labeling also capable to fully characterize groups of people and to manage miss segmentations. The ground-plane homography and the epipolar geometry are automatically learned and exploited to warp objects' principal axes between overlapped cameras. A MAP estimator that exploits two contributions (forward and backward) is used to choose the most probable label configuration to be assigned at the handoff of a new object. Extensive experiments demonstrate the accuracy of the proposed method in detecting single and simultaneous handoffs, miss segmentations, and groups. Simone Calderara, Rita Cucchiara, Andrea Prati 0001 |
AVSS | 3 |
| 2006 | Low-Latency Live Video Streaming over Low-Capacity NetworksabstractThis paper presents an effective system for streaming over low-capacity networks (such as GPRS and EGPRS) of live videos with low latency. Existing solutions are either too complex or not suitable to our scope. For this reason, we developed a complete, ready-to-use streaming system based on H.264/AVC codec and UDP/IP stack. The system employs adaptive controls to achieve the best tradeoff between low latency and good video fluency, by keeping the UDP buffer occupancy at the decoder side between two given levels. Our experiments demonstrate that this system is able to transmit live videos at CIF format and 10 fps over GPRS/EGPRS with very low latency (1.73 sec on average, basically due to the network delay), good fluency and average quality, measured with PSNR, of 31 dB on GPRS at 23 kbps at 10 fps Giovanni Gualdi, Rita Cucchiara, Andrea Prati 0001 |
ISM | 3 |
| 2006 | A semi-automatic system for segmentation of cardiac M-mode images
Luca Bertelli, Rita Cucchiara, Giovanni Paternostro, Andrea Prati 0001 |
Pattern Anal. Appl. | 4 |
| 2006 | A system for automatic face obscuration for privacy purposes
Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani |
Pattern Recognit. Lett. | 2 |
| 2006 | Semantic adaptation of sport videos with user-centred performance analysisabstractIn semantic video adaptation measures of performance must consider the impact of the errors in the automatic annotation over the adaptation in relationship with the preferences and expectations of the user. In this paper, we define two new performance measures Viewing Quality Loss and Bit-rate Cost Increase,that are obtained from classical peak signal-to-noise ration (PSNR) and bitrate, and relate the results of semantic adaptation to the errors in the annotation of events and objects and the user's preferences and expectations. We present and discuss results obtained with a system that performs automatic annotation of soccer sport video highlights and applies different coding strategies to different parts of the video according to their relative importance for the end user. With reference to this framework, we analyze how highlights' statistics and the errors of the annotation engine influence the performance of semantic adaptation and reflect into the quality of the video displayed at the user's client and the increase of transmission costs. Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001 |
IEEE Trans. Multim. | 4 |
| 2005 | Entry edge of field of view for multi-camera tracking in distributed video surveillanceabstractEfficient solution to people tracking in distributed video surveillance is requested to monitor crowded and large environments. This paper proposes a novel use of the entry edges of field of view (E/sup 2/oFoV) to solve the consistent labeling problem between partially overlapped views. An automatic and reliable procedure allows obtaining the homographic transformation between two overlapped views, without any manual calibration of the cameras. Through the homography, the consistent labeling is established each time a new track is detected in one of the cameras. A camera transition graph (CTG) is defined to speed up the establishment process by reducing the search space. Experimental results prove the effectiveness of the proposed solution also in challenging conditions. Simone Calderara, Roberto Vezzani, Andrea Prati 0001, Rita Cucchiara |
AVSS | 3 |
| 2005 | Posture classification in a multi-camera indoor environmentabstractPosture classification is a key process for analyzing the people's behaviour. Computer vision techniques can be helpful in automating this process, but cluttered environments and consequent occlusions make this task often difficult. Different views provided by multiple cameras can be exploited to solve occlusions by warping known object appearance into the occluded view. To this aim, this paper describes an approach to posture classification based on projection histograms, reinforced by HMM for assuring temporal coherence of the posture. The single camera posture classification is then exploited in the multi-camera system to solve the cases in which the occlusions make the classification impossible. Experimental results of the classification from both the single camera and the multi-camera system are provided. Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani |
ICIP (1) | 2 |
| 2005 | On the usefulness of object shape coding with MPEG-4abstractThis paper reports the results of an in-depth analysis of the degree of usefulness of object shape coding in video compression. In particular, MPEG-4 is used as reference standard. The influence of different coding parameters on the performance is deeply examined and discussions on the results are provided. Object shape coding is compared with classical (MPEG-2) frame-based coding both at an objective level (by comparing PSNR/quality and bitrate/filesize) and at a subjective level (asking to a set of users to express their opinion on overall quality, cognitive effectiveness, and willingness to pay). In conclusion, this paper aims at answering to the question whether it is convenient to use object shape coding instead of frame-based coding or not. Andrea Prati 0001, Rita Cucchiara |
ISM | 1 |
| 2005 | An Integrated Framework for Semantic Annotation and Adaptation
Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001 |
Multim. Tools Appl. | 4 |
| 2005 | Probabilistic posture classification for Human-behavior analysisabstractComputer vision and ubiquitous multimedia access nowadays make feasible the development of a mostly automated system for human-behavior analysis. In this context, our proposal is to analyze human behaviors by classifying the posture of the monitored person and, consequently, detecting corresponding events and alarm situations, like a fall. To this aim, our approach can be divided in two phases: for each frame, the projection histograms (Haritaoglu et al., 1998) of each person are computed and compared with the probabilistic projection maps stored for each posture during the training phase; then, the obtained posture is further validated exploiting the information extracted by a tracking module in order to take into account the reliability of the classification of the first phase. Moreover, the tracking algorithm is used to handle occlusions, making the system particularly robust even in indoors environments. Extensive experimental results demonstrate a promising average accuracy of more than 95% in correctly classifying human postures, even in the case of challenging conditions. Rita Cucchiara, Costantino Grana, Andrea Prati 0001, Roberto Vezzani |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2004 | Content-based video adaptation with user's preferencesabstractWe present an integrated system that has been designed to support automatic semantic extraction of highlights in sports video and automatic video adaptation according to user's preferences. To analyze the user's satisfaction, we propose a new performance measure that explicitly takes into account the user's preferences and considers the number and type of errors produced by the annotation engine and the way in which these errors affect the compressed video quality and bandwidth allocation. We provide experimental results with application to soccer and swimming. Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001 |
ICME | 4 |
| 2004 | Neighbor cache prefetching for multimedia image and video processingabstractCache performance is strongly influenced by the type of locality embodied in programs. In particular, multimedia programs handling images and videos are characterized by a bidimensional spatial locality, which is not adequately exploited by standard caches. In this paper we propose novel cache prefetching techniques for image data, called neighbor prefetching, able to improve exploitation of bidimensional spatial locality. A performance comparison is provided against other assessed prefetching techniques on a multimedia workload (with MPEG-2 and MPEG-4 decoding, image processing, and visual object segmentation), including a detailed evaluation of both the miss rate and the memory access time. Results prove that neighbor prefetching achieves a significant reduction in the time due to delayed memory cycles (more than 97% on MPEG-4 with respect to 75% of the second performing technique). This reduction leads to a substantial speedup on the overall memory access time (up to 140% for MPEG-4). Performance has been measured with the PRIMA trace-driven simulator, specifically devised to support cache prefetching. Rita Cucchiara, Massimo Piccardi, Andrea Prati 0001 |
IEEE Trans. Multim. | 3 |
| 2003 | Object Segmentation in Videos from Moving Camera with MRFs on Color and Motion FeaturesabstractIn this paper we address the problem of fast segmenting moving objects in video acquired by moving camera or more generally with a moving background. We present an approach based on a color segmentation followed by a region-merging on motion through Markov random fields (MRFs). The technique we propose is inspired by the work of Gelgon and Bouthemy (2000), that has been modified to reduce computational cost in order to achieve a fast segmentation (about ten frame per second). To this aim a modified region matching algorithm (namely partitioned region matching) and an innovative arc-based MRF optimization algorithm with a suitable definition of the motion reliability are proposed. Results on both synthetic and real sequences are reported to confirm validity of our solution. Rita Cucchiara, Andrea Prati 0001, Roberto Vezzani |
CVPR (1) | 2 |
| 2003 | Object and event detection for semantic annotation and transcodingabstractVideo annotation provides a suitable way to describe, organize, and index stored videos. On the other hand, transcoding aims at adapting content to the user/client capabilities and requirements. Both cues are now mandatory, given the tremendous demand of multimedia access from remote clients, in particular nowadays that new terminals with limited resources (PDAs, HCCs, Smart phones) have access to the network. In this paper we propose a unified framework to define event-based and object-based semantic extraction from video to provide both semantic video annotation for video stored and semantic on-line transcoding from live cameras. Two case studies (highlights' extraction from soccer videos for the annotation and people behavior detection in domotic application for transcoding) and corresponding experimental results are reported. Marco Bertini 0001, Rita Cucchiara, Alberto Del Bimbo, Andrea Prati 0001 |
ICME | 4 |
| 2003 | Improving Data Prefetching Efficacy in Multimedia Applications
Rita Cucchiara, Andrea Prati 0001, Massimo Piccardi |
Multim. Tools Appl. | 2 |
| 2003 | Detecting Moving Objects, Ghosts, and Shadows in Video StreamsabstractBackground subtraction methods are widely exploited for moving object detection in videos in many applications, such as traffic monitoring, human motion capture, and video surveillance. How to correctly and efficiently model and update the background model and how to deal with shadows are two of the most distinguishing and challenging aspects of such approaches. The article proposes a general-purpose method that combines statistical assumptions with the object-level knowledge of moving objects, apparent objects (ghosts), and shadows acquired in the processing of the previous frames. Pixels belonging to moving objects, ghosts, and shadows are processed differently in order to supply an object-based selective update. The proposed approach exploits color information for both background subtraction and shadow detection to improve object segmentation and background update. The approach proves fast, flexible, and precise in terms of both pixel accuracy and reactivity to background changes. Rita Cucchiara, Costantino Grana, Massimo Piccardi, Andrea Prati 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2003 | Detecting Moving Shadows: Algorithms and EvaluationabstractMoving shadows need careful consideration in the development of robust dynamic scene analysis systems. Moving shadow detection is critical for accurate object detection in video streams since shadow points are often misclassified as object points, causing errors in segmentation and tracking. Many algorithms have been proposed in the literature that deal with shadows. However, a comparative evaluation of the existing approaches is still lacking. In this paper, we present a comprehensive survey of moving shadow detection approaches. We organize contributions reported in the literature in four classes two of them are statistical and two are deterministic. We also present a comparative empirical evaluation of representative algorithms selected from these four classes. Novel quantitative (detection and discrimination rate) and qualitative metrics (scene and object independence, flexibility to shadow situations, and robustness to noise) are proposed to evaluate these classes of algorithms on a benchmark suite of indoor and outdoor video sequences. These video sequences and associated "ground-truth" data are made available at http://cvrr.ucsd.edu/aton/shadow to allow for others in the community to experiment with new algorithms and metrics. Andrea Prati 0001, Ivana Mikic, Mohan M. Trivedi, Rita Cucchiara |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Semantic transcoding for live video serverabstractIn this paper we present transcoding techniques for a video server architecture that enables the user to access live video streams by using different devices with different capabilities. For live videos, annotation methods cannot be exploited. Instead we propose methods of on-the-fly transcoding that adapt the video content with respect to the user resources and the video semantic. Thus we propose an object-based transcoding with classes of relevance (for instance People, Face and Background). To compare the different strategies we propose a metric based on the Weighted Mean Square Error that allows the analysis of different application scenarios by means of a class-wise distortion measure. The obtained results show that the use of semantic can improve the bandwidth to distortion ratio significantly. Rita Cucchiara, Costantino Grana, Andrea Prati 0001 |
ACM Multimedia | 3 |
| 2001 | Analysis and Detection of Shadows in Video Streams: A Comparative EvaluationabstractRobustness to changes in illumination conditions as well as viewing perspectives is an important requirement for many computer vision applications. One of the key factors in enhancing the robustness of dynamic scene analysis is that of accurate and reliable means for shadow detection. Shadow detection is critical for correct object detection in image sequences. Many algorithms have been proposed in the literature that deal with shadows. However, a comparative evaluation of the existing approaches is still lacking. In this paper, the full range of problems underlying the shadow detection is identified and discussed. We classify the proposed solutions to this problem using a taxonomy of four main classes, deterministic model and non-model based, and statistical parametric and nonparametric. Novel quantitative (detection and discrimination accuracy) and qualitative metrics (scene and object independence, flexibility to shadow situations and robustness to noise) are proposed to evaluate these classes of algorithms on a benchmark suite of indoor and outdoor video sequences. Andrea Prati 0001, Rita Cucchiara, Ivana Mikic, Mohan M. Trivedi |
CVPR (2) | 1 |
| 2000 | Focus based Feature Extraction for Pallets RecognitionabstractVisual recognition for object grasping is a well-known challenge for robot automation in industrial applications. A typical example is pallet recognition in industrial environment for pick-and-place automated process. The aim of vision and reasoning algorithms is to help robots in choosing the best pallets holes location. This work proposes an application-based approach, which full all requirements, dealing with every kind of occlusions and light situ-ations possible. Even some meaning noise (or meaning misunderstand-ing) is considered. A pallet model, with limited degrees of freedom, is de-scribed and, starting from it, a complete approach to pallet recognition is out-lined. In the model we dene both virtual and real corners, that are geomet-rical object proprieties computed by different image analysis operators. Real corners are perceived by processing brightness information directly from the image, while virtual corners are inferred at a higher level of abstraction. A nal reasoning stage selects the best solution tting the model. Experimental results and performance are reported in order to demonstrate the suitability of the proposed approach. 1 Rita Cucchiara, Massimo Piccardi, Andrea Prati 0001 |
BMVC | 3 |
| 2000 | Exploring multimedia applications locality to improve cache performance
Andrea Prati 0001 |
ACM Multimedia | 1 |