VLDB 2026 Research / reviewers in the wild / expert
Nicola Conci
dblp:32/3652
· DBLP profile ↗
64ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-7858-0928ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 16 · 11 since 2021Computer networks · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NFlowAD: A normalizing flow model for anomaly detection in human motion animationsabstractAnomaly detection has been extensively investigated in numerous application areas. Hand-crafted rules have gradually given way to supervised classification techniques, which frequently rely on a small number of anomaly labels and related architectures. When it comes to human motion, abnormalities emerge at a fine-grained temporal or joint level rather than over a whole video sequence. This study introduces NFlowAD, a self-supervised system that analyzes body joints to detect irregularities in human motion. It blends normalizing flows with masked motion modeling to describe normal motion data without the need for anomaly labels. Inference uses both reconstruction mistakes and flow-based likelihoods to detect anomalies. The validation pipeline on various state-of-the-art datasets demonstrates NFlowAD’s efficiency in recognizing, locating, and analyzing anomalous motion sequences, while maintaining robust detection and interpretability. Mahamat Issa Choueb, Praveen Kumar Sekharamantry, Giulia Martinelli, Francesco G. B. De Natale, Nicola Conci |
Signal Process. Image Commun. | 5 |
| 2025 | WCT-Enhanced Instance Normalization for Unsupervised Domain Adaptation in Object DetectionabstractSafety-critical applications like video surveillance, traffic monitoring, and autonomous driving often face a lack of labeled data and performance degradation due to domain shifts between training and testing sets. Unsupervised Domain Adaptation (UDA) offers a viable solution when annotating the target domain is costly and time consuming. We propose a novel WCT-IN stylization module that injects Gaussian noise into the style feature representation, combining Adaptive Instance Normalization (AdaIN) for style transfer with Whitening and Coloring Transform (WCT) for color correction and feature transformation. This generates intermediate images that bridge the source–target domain gap while preserving source labels. Integrated with a consistency learning framework, our method improves adaptation performance. Extensive experiments on benchmark datasets (Cityscapes, Foggy Cityscapes, Sim10K, and KITTI) demonstrate the effectiveness of our approach in enhancing object detection under domain shifts. Andualem Welabo Tulu, Nicola Conci |
AVSS | 2 |
| 2025 | MEET: The Music Event Emotion Tracking MetaverseabstractThis paper presents the Music Event Emotion Tracking (MEET) Metaverse, which is one of the demo results produced within the FUN-Media project. The MEET Metaverse is a virtual disco where multiple users can join together through their avatars to enjoy musical events. The peculiar characteristic of the MEET Metaverse is that the emotions of participants are inferred from their facial expressions and speech, and are used to select the next song to be played in the disco based on the average emotional state of participants. Moreover, avatars’ facial poses are updated based on participants’ emotions, and realistic avatar animations are reproduced using a combination of motion retargeting and high-fidelity appearance modeling. Luigi Atzori, Gülnaziye Bingöl, Concetta Cantone, Nicola Conci, Matteo Fasa, Alessandro Floris, Giulia Martinelli, Marina Samarotto, Salvatore Serrano |
QoMEX | 4 |
| 2025 | Signal processing for haptic surface modeling: A reviewabstractHaptic feedback has been integrated into Virtual and Augmented Reality, complementing acoustic and visual information and contributing to an all-round immersive experience in multiple fields, spanning from the medical domain to entertainment and gaming. Haptic technologies involve complex cross-disciplinary research that encompasses sensing, data representation, interactive rendering, perception, and quality of experience. The standard processing pipeline, consists of (I) sensing physical features in the real world using a transducer, (II) modeling and storing the collected information in some digital format, (III) communicating the information, and finally, (IV) rendering the haptic information through appropriate devices, thus producing a user experience (V) perceptually close to the original physical world. Among these areas, sensing, rendering and perception have been deeply investigated and are the subject of different comprehensive surveys available in the literature. Differently, research dealing with haptic surface modeling and data representation still lacks a comprehensive dissection. In this work, we aim at providing an overview on modeling and representation of haptic surfaces from a signal processing perspective, covering the aspects that lie in between haptic information acquisition on one side and rendering and perception on the other side. We analyze, categorize, and compare research papers that address the haptic surface modeling and data representation, pointing out existing gaps and possible research directions. Antonio Luigi Stefani, Niccolò Bisagno, Andrea Rosani, Nicola Conci, Francesco G. B. De Natale |
Signal Process. Image Commun. | 4 |
| 2024 | Lagrangian Hashing for Compressed Neural Field Representations
Shrisudhan Govindarajan, Zeno Sambugaro, Akhmedkhan Shabanov, Towaki Takikawa, Daniel Rebain, Nicola Conci, Kwang Moo Yi, Andrea Tagliasacchi |
ECCV (27) | 7 |
| 2024 | MapFlow: Multi-Agent Pedestrian Trajectory Prediction Using Normalizing FlowabstractIn the task of pedestrian trajectory prediction, multi-modal prediction has recently emerged, demonstrating how a good model should predict multiple socially acceptable futures. With this respect, Normalizing Flows (NFs) have shown remarkable generative capabilities that make them particularly suitable for multi-modal trajectory prediction. By sampling from the learned distribution, NFs can produce multiple socially acceptable trajectories, each one paired with its corresponding likelihood score. Taking advantage of the multi-modal prediction coupled with the likelihood score, with MapFlow we introduce a solution based on NFs that improves the accuracy in prediction by incorporating in the model the social influence of neighboring pedestrians.1 Antonio Luigi Stefani, Niccolò Bisagno, Nicola Conci |
ICASSP | 3 |
| 2024 | Unicrowd Simulator: Visual and Behavioral Fidelity For The Generation of Crowd DatasetsabstractWe introduce UniCrowd1, a human crowd simulator for the modeling of human-related dynamics. The simulator is accompanied by a meticulously collected dataset within its synthetic environment, along with a comprehensive validation pipeline. Leveraging simulation as a powerful tool for generating annotated data, UniCrowd addresses the increasing demand for large training datasets, mimicking both the behavioral and visual aspect of crowds. Recent advancements in rendering and virtualization engines have enhanced the simulators capabilities to represent complex scenes, encompassing environmental factors such as weather conditions, surface reflectance, and human-related events like actions and behaviors. The adaptability and the non-deterministic nature of the human behavioral module of UniCrowd, coupled with its 3D rendering represents an improvement over available crowd simulators. We demonstrate the suitability of our simulator and its associated dataset for various computer vision tasks. We highlight applications such as detection and segmentation, as well as specialized tasks including crowd counting, human pose estimation, trajectory analysis and prediction.1The simulator and the dataset can be accessed at github.com/mmlabcvUniCrowd Niccolò Bisagno, Antonio Luigi Stefani, Nicola Garau, Francesco G. B. De Natale, Nicola Conci |
ICIP | 5 |
| 2024 | All Skeletons are Created Equal! A Domain Adaptation Transformer to Handle Multiple TopologiesabstractDesigning an effective Human Pose Estimation (HPE) pipeline necessitates handling diverse datasets, a task often deemed necessary but burdensome by researchers. Existing datasets are annotated using different conventions, with the human body represented as a parametric 3D model or as a collection of 2D/3D joints and bones, known as skeleton-based annotation. Despite its widespread use in training both 2D and 3D HPE networks, the lack of standardization in the topologies of joint-based pose annotations requires considerable difficulties when evaluating the algorithms performances across different datasets. To solve this issue, we introduce a novel self-supervised human pose domain adaptation approach to map a given skeletal model into a target model of choice. We design a transformer-based architecture trained to reconstruct missing joints within a given topology, aligning them with the target model. During testing, the network seamlessly reconstructs the missing joints, treating them as if masked, based on the common joints shared by the original topology and the target one. Unlike previous approaches, our method works with arbitrary body poses, joints number, and skeleton topologies, across multiple datasets.11The code is available at https://github.com/mmlab-cv/DAT.git Giulia Martinelli, Nicola Garau, Niccolò Bisagno, Nicola Conci |
ICIP | 4 |
| 2024 | Dynamic Crowd Routing: RL-Driven Crowd DynamicsabstractThe simulation of crowds is complex and challenging. Every individual in a crowd exhibits a different behaviour, targets a different goal, and undergoes different types of interactions. Within crowds, groups can be identified in both static and dynamic configurations, with varying levels of responsiveness, leading to the emergence of complex avoidance mechanisms. In the past, rule-based models have been proposed to simulate crowds, unlocking the potential for large-scale simulations. Over the years, learning-based solutions have been presented, achieving acceptable results despite the lack of high-quality ground truth data for training. While both rule-based and learning-based methods have recently been integrated into 3D simulation engines, they usually rely on navigation meshes or B-spline functions, hindering their generalization to open-world scenarios. In this work, we propose a reinforcement learning-based solution to learn meaningful crowd dynamics inside the Unreal Engine 3D engine, enabling massive and highly dynamic crowd simulations. We show how our proposed method makes it possible to simulate crowd setups that require complex dynamic routing mechanisms, which are otherwise hard to achieve using rule-based approaches or even deep learning-based methods. Our approach also al-lows us to easily collect large synthetic datasets that are both photorealistic and provide accurate ground truth data without the need for any manual annotation. Some demonstration videos are available at mmlab-cv.github.io/DynamicCrowdRouting; Code, complete experiments and analysis will be made available upon acceptance. Daniele Della Pietra, Nicola Garau, Nicola Conci, Fabrizio Granelli |
MMSP | 3 |
| 2024 | MoMa: Skinned motion retargeting using masked pose modeling
Giulia Martinelli, Nicola Garau, Niccolò Bisagno, Nicola Conci |
Comput. Vis. Image Underst. | 4 |
| 2024 | Agglomerator++: Interpretable part-whole hierarchies and latent space representations in neural networksabstractDeep neural networks achieve outstanding results in a large variety of tasks, often outperforming human experts. However, a known limitation of current neural architectures is the poor accessibility in understanding and interpreting the network’s response to a given input. This is directly related to the huge number of variables and the associated non-linearities of neural models, which are often used as black boxes. This lack of transparency, particularly in crucial areas like autonomous driving, security, and healthcare, can trigger skepticism and limit trust, despite the networks’ high performance. In this work, we want to advance the interpretability in neural networks. We present Agglomerator++, a framework capable of providing a representation of part-whole hierarchies from visual cues and organizing the input distribution to match the conceptual-semantic hierarchical structure between classes. We evaluate our method on common datasets, such as SmallNORB, MNIST, FashionMNIST, CIFAR-10, and CIFAR-100, showing that our solution delivers a more interpretable model compared to other state-of-the-art approaches. Our code is available at https://mmlab-cv.github.io/Agglomeratorplusplus/ . • We introduce a novel model, called Agglomerator++, mimicking the functioning of the cortical columns in the human brain. • Our solution provides interpretability of relationships in data, namely the hierarchical organization of the feature space. • We introduce positional encoding and input masking during pre-training for self- supervised reconstruction. • This neural representation is more efficient and closely resembles human lexical similarities. Zeno Sambugaro, Nicola Garau, Niccolò Bisagno, Nicola Conci |
Comput. Vis. Image Underst. | 4 |
| 2023 | CapsulePose: A variational CapsNet for real-time end-to-end 3D human pose estimation
Nicola Garau, Nicola Conci |
Neurocomputing | 2 |
| 2022 | Interpretable part-whole hierarchies and conceptual-semantic relationships in neural networksabstractDeep neural networks achieve outstanding results in a large variety of tasks, often outperforming human experts. However, a known limitation of current neural architectures is the poor accessibility to understand and interpret the network response to a given input. This is directly related to the huge number of variables and the associated non-linearities of neural models, which are often used as black boxes. When it comes to critical applications as autonomous driving, security and safety, medicine and health, the lack of interpretability of the network behavior tends to induce skepticism and limited trustworthiness, despite the accurate performance of such systems in the given task. Furthermore, a single metric, such as the classification accuracy, provides a non-exhaustive evaluation of most realworld scenarios. In this paper, we want to make a step forward towards interpretability in neural networks, providing new tools to interpret their behavior. We present Agglomerator, a framework capable of providing a representation of part-whole hierarchies from visual cues and organizing the input distribution matching the conceptual-semantic hierarchical structure between classes. We evaluate our method on common datasets, such as SmallNORB, MNIST, FashionMNIST, CIFAR-10, and CIFAR-100, providing a more interpretable model than other state-of-the-art approaches. Nicola Garau, Niccolò Bisagno, Zeno Sambugaro, Nicola Conci |
CVPR | 4 |
| 2022 | Human trajectory forecasting using a flow-based generative model
Bo Zhang 0045, Tao Wang 0110, Changdong Zhou, Nicola Conci, Hongbo Liu 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | A multimodal framework for the evaluation of patients' weaknesses, supporting the design of customised AAL solutions
Nicola Garau, Damiano Fruet, Alessandro Luchetti, Francesco G. B. De Natale, Nicola Conci |
Expert Syst. Appl. | 5 |
| 2022 | Editorial for the special issue on deep learning for precise and efficient object detection
Yanwei Pang, Jungong Han, Nicola Conci |
Pattern Recognit. Lett. | 4 |
| 2021 | Out-of-Distribution Detection Using Union of 1-Dimensional SubspacesabstractThe goal of out-of-distribution (OOD) detection is to handle the situations where the test samples are drawn from a different distribution than the training data. In this paper, we argue that OOD samples can be detected more easily if the training data is embedded into a low-dimensional space, such that the embedded training samples lie on a union of 1-dimensional subspaces. We show that such embedding of the in-distribution (ID) samples provides us with two main advantages. First, due to compact representation in the feature space, OOD samples are less likely to occupy the same region as the known classes. Second, the first singular vector of ID samples belonging to a 1-dimensional subspace can be used as their robust representative. Motivated by these observations, we train a deep neural network such that the ID samples are embedded onto a union of 1-dimensional subspaces. At the test time, employing sampling techniques used for approximate Bayesian inference in deep learning, input samples are detected as OOD if they occupy the region corresponding to the ID samples with probability 0. Spectral components of the ID samples are used as robust representative of this region. Our method does not have any hyperparameter to be tuned using extra information and it can be applied on different modalities with minimal change. The effectiveness of the proposed method is demonstrated on different benchmark datasets, both in the image and video classification domains. Alireza Zaeemzadeh, Niccolò Bisagno, Zeno Sambugaro, Nicola Conci, Nazanin Rahnavard, Mubarak Shah |
CVPR | 4 |
| 2021 | DECA: Deep viewpoint-Equivariant human pose estimation using Capsule AutoencodersabstractHuman Pose Estimation (HPE) aims at retrieving the 3D position of human joints from images or videos. We show that current 3D HPE methods suffer a lack of viewpoint equivariance, namely they tend to fail or perform poorly when dealing with viewpoints unseen at training time. Deep learning methods often rely on either scale-invariant, translation-invariant, or rotation-invariant operations, such as max-pooling. However, the adoption of such procedures does not necessarily improve viewpoint generalization, rather leading to more data-dependent methods. To tackle this issue, we propose a novel capsule autoencoder network with fast Variational Bayes capsule routing, named DECA. By modeling each joint as a capsule entity, combined with the routing algorithm, our approach can preserve the joints’ hierarchical and geometrical structure in the feature space, independently from the viewpoint. By achieving viewpoint equivariance, we drastically reduce the network data dependency at training time, resulting in an improved ability to generalize for unseen viewpoints. In the experimental validation, we outperform other methods on depth images from both seen and unseen viewpoints, both top-view, and front-view. In the RGB domain, the same network gives state-of-the-art results on the challenging viewpoint transfer task, also establishing a new framework for top-view HPE. The code can be found at https://github.com/mmlab-cv/DECA. Nicola Garau, Niccolò Bisagno, Piotr Bródka, Nicola Conci |
ICCV | 4 |
| 2021 | Embedding group and obstacle information in LSTM networks for human trajectory prediction in crowded scenes
Niccolò Bisagno, Cristiano Saltori, Bo Zhang 0045, Francesco G. B. De Natale, Nicola Conci |
Comput. Vis. Image Underst. | 5 |
| 2021 | Where Are They Going? Predicting Human Behaviors in Crowded ScenesabstractIn this article, we propose a framework for crowd behavior prediction in complicated scenarios. The fundamental framework is designed using the standard encoder-decoder scheme, which is built upon the long short-term memory module to capture the temporal evolution of crowd behaviors. To model interactions among humans and environments, we embed both the social and the physical attention mechanisms into the long short-term memory. The social attention component can model the interactions among different pedestrians, whereas the physical attention component helps to understand the spatial configurations of the scene. Since pedestrians’ behaviors demonstrate multi-modal properties, we use the generative model to produce multiple acceptable future paths. The proposed framework not only predicts an individual’s trajectory accurately but also forecasts the ongoing group behaviors by leveraging on the coherent filtering approach. Experiments are carried out on the standard crowd benchmarks (namely, the ETH, the UCY, the CUHK crowd, and the CrowdFlow datasets), which demonstrate that the proposed framework is effective in forecasting crowd behaviors in complex scenarios. Bo Zhang 0045, Niccolò Bisagno, Nicola Conci, Francesco G. B. De Natale, Hongbo Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Weight Estimation from an RGB-D camera in top-view configurationabstractThe development of so-called soft-biometrics aims at providing information related to the physical and behavioural characteristics of a person. This paper focuses on body weight estimation based on the observation from a top-view RGB-D camera. In fact, the capability to estimate the weight of a person can be of help in many different applications, from health-related scenarios, to business intelligence and retail analytics. To deal with this issue, a TVWE (Top-View Weight Estimation) framework is proposed with the aim of predicting the weight. The approach relies on the adoption of Deep Neural Networks (DNNs) that have been trained on depth data. Each network has also been modified in their top section to replace classification with prediction inference. The performance of five state-of-art DNNs have been compared, namely VGG16, ResNet, Inception, DenseNet and Efficient-Net. In addition, a convolutional auto-encoder has also been included for completeness. Considering the limited literature in this domain, the TVWE framework has been evaluated on a new publicly available dataset: “VRAI Weight estimation Dataset”, which also collects, for each subject, labels related to weight, gender, and height. The experimental results have demonstrated that the proposed methods are suitable for this task, bringing different and significant insights for the application of the solution in different domains. Marco Mameli, Marina Paolanti, Nicola Conci, Filippo Tessaro, Emanuele Frontoni, Primo Zingaretti |
ICPR | 3 |
| 2020 | AG-GAN: An Attentive Group-Aware GAN for pedestrian trajectory predictionabstractUnderstanding human behaviors in crowded scenarios requires analyzing not only the position of the subjects in space, but also the scene context. Existing approaches mostly rely on the motion history of each pedestrian and model the interactions among people by considering the entire surrounding neighborhood. In our approach, we address the problem of motion prediction by applying coherent group clustering and a global attention mechanism on the LSTM-based Generative Adversarial Networks (GANs). The proposed model consists of an attentive group-aware GAN that observes the agents' past motion and predicts future paths, using (i) a group pooling module to model neighborhood interaction, and (ii) an attention module to specifically focus on hidden states. The experimental results demonstrate that our proposal outperforms state-of-the-art models on common benchmark datasets, and is able to generate socially-acceptable trajectories. Niccolò Bisagno, Syed Zohaib Hassan, Nicola Conci |
ICPR | 4 |
| 2019 | Person Head Detection Based Deep Model for People Counting in Sports VideosabstractPeople counting in sports venues is emerging as a new domain in the field of video surveillance. People counting in these venues faces many key challenges, such as severe occlusions, few pixels per head, and significant variations in person's head sizes due to wide sport areas. We propose a deep model based method, which works as a head detector and takes into consideration the scale variations of heads in videos. Our method is based on the notion that head is the most visible part in the sports venues where large number of people are gathered. To cope with the problem of different scales, we generate scale aware head proposals based on scale map. Scale aware proposals are then fed to the Convolutional Neural Network (CNN) and it provides a response matrix containing the presence probabilities of people observed across scene scales. We then use non-maximal suppression to get the accurate head positions. For the performance evaluation, we carry out extensive experiments on two standard datasets and compare the results with state-of-the-art (SoA) methods. The results in terms of Average Precision (AvP), Average Recall (AvR), and Average F1-Score (AvF-Score) show that our method is better than SoA methods. Sultan Daud Khan, Mohib Ullah, Nicola Conci, Faouzi Alaya Cheikh, Azeddine Beghdadi |
AVSS | 4 |
| 2019 | Indoor object recognition in RGBD images with complex-valued neural networks for visually-impaired people
Rim Trabelsi, Issam Jabri, Farid Melgani, Fethi Smach, Nicola Conci, Ammar Bouallègue |
Neurocomputing | 5 |
| 2019 | Social media and satellites - Disaster event detection, linking and summarization
Kashif Ahmad, Konstantin Pogorelov, Michael Riegler 0001, Nicola Conci, Pål Halvorsen |
Multim. Tools Appl. | 4 |
| 2019 | Natural disasters detection in social media and satellite imagery: a survey
Naina Said, Kashif Ahmad, Michael Riegler 0001, Konstantin Pogorelov, Laiq Hassan, Nasir Ahmad, Nicola Conci |
Multim. Tools Appl. | 7 |
| 2019 | Automatic detection of passable roads after floods in remote sensed and social media data
Kashif Ahmad, Konstantin Pogorelov, Michael Riegler 0001, Olga Ostroukhova, Pål Halvorsen, Nicola Conci, Rozenn Dahyot |
Signal Process. Image Commun. | 6 |
| 2019 | How Deep Features Have Improved Event Recognition in Multimedia: A SurveyabstractEvent recognition is one of the areas in multimedia that is attracting great attention of researchers. Being applicable in a wide range of applications, from personal to collective events, a number of interesting solutions for event recognition using multimedia information sources have been proposed. On the other hand, following their immense success in classification, object recognition, and detection, deep learning has been shown to perform well in event recognition tasks also. Thus, a large portion of the literature on event analysis relies nowadays on deep learning architectures. In this article, we provide an extensive overview of the existing literature in this field, analyzing how deep features and deep learning architectures have changed the performance of event recognition frameworks. The literature on event-based analysis of multimedia contents can be categorized into four groups, namely (i) event recognition in single images; (ii) event recognition in personal photo collections; (iii) event recognition in videos; and (iv) event recognition in audio recordings. In this article, we extensively review different deep-learning-based frameworks for event recognition in these four domains. Furthermore, we also review some benchmark datasets made available to the scientific community to validate novel event recognition pipelines. In the final part of the manuscript, we also provide a detailed discussion on basic insights gathered from the literature review, and identify future trends and challenges. Kashif Ahmad, Nicola Conci |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | A saliency-based approach to event recognition
Kashif Ahmad, Nicola Conci, Francesco G. B. De Natale |
Signal Process. Image Commun. | 2 |
| 2018 | Ensemble of Deep Models for Event RecognitionabstractIn this article, we address the problem of recognizing an event from a single related picture. Given the large number of event classes and the limited information contained in a single shot, the problem is known to be particularly hard. To achieve a reliable detection, we propose a combination of multiple classifiers, and we compare three alternative strategies to fuse the results of each classifier, namely: (i) induced order weighted averaging operators, (ii) genetic algorithms, and (iii) particle swarm optimization. Each method is aimed at determining the optimal weights to be assigned to the decision scores yielded by different deep models, according to the relevant optimization strategy. Experimental tests have been performed on three event recognition datasets, evaluating the performance of various deep models, both alone and selectively combined. Experimental results demonstrate that the proposed approach outperforms traditional multiple classifier solutions based on uniform weighting, and outperforms recent state-of-the-art approaches. Kashif Ahmad, Mohamed Lamine Mekhalfi, Nicola Conci, Farid Melgani, Francesco G. B. De Natale |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2017 | Data-Driven crowd simulationabstractIn this work we propose a framework for data-driven crowd simulation starting from a small set of trajectories. To model the macroscopic behavior, our method extracts pedestrian trajectories from real videos, clusters all trajectories of pedestrians who intend to reach the same goal, and computes the velocity field associated with each exit region in the scene to guide virtual agents toward their destinations, namely, the goal-dependent path selection. While at the microscopic level, the simulation is performed using on the one hand the Social Force Model to handle the collision-avoidance among agents and on the other hand the computed velocity fields to model the macroscopic behavior. The experimental results demonstrate that the velocity field can be exploited to effectively reproduce crowd behaviors. Niccolò Bisagno, Nicola Conci, Bo Zhang 0045 |
AVSS | 2 |
| 2017 | A pool of deep models for event recognitionabstractThis paper proposes a novel two-stage framework for event recognition in still images. First, for a generic event image, deep features, obtained via different pre-trained models, are fed into an ensemble of classifiers, whose posterior classification probabilities are thereafter fused by means of an order-induced scheme, which penalizes the yielded scores according to their confidence in classifying the image at hand, and then averages them. Second, we combine the fusion results with a reverse matching paradigm in order to draw the final output of our proposed pipeline. We evaluate our approach on three challenging datasets and we show that better results can be attained, advancing recent leading works. Kashif Ahmad, Mohamed Lamine Mekhalfi, Nicola Conci, Giulia Boato, Farid Melgani, Francesco G. B. De Natale |
ICIP | 3 |
| 2017 | The JORD System: Linking Sky and Social Multimedia Data to Natural DisastersabstractBeing able to automatically link social media information and data to remote-sensed data holds large possibilities for society and research. In this paper, we present a system called JORD that is able to autonomously collect social media data about technological and environmental disasters, and link it automatically to remote-sensed data. In addition, we demonstrate that queries in local languages that are relevant to the exact position of natural disasters retrieve more accurate information about a disaster event. To show the capabilities of the system, we present some examples of disaster events detected by the system. To evaluate the quality of the provided information and usefulness of JORD from the potential users point of view we include a crowdsourced user study. Kashif Ahmad, Michael Riegler 0001, Ans Riaz, Nicola Conci, Duc-Tien Dang-Nguyen, Pål Halvorsen |
ICMR | 4 |
| 2017 | Complex-Valued Representation for RGB-D Object Recognition
Rim Trabelsi, Issam Jabri, Farid Melgani, Fethi Smach, Nicola Conci, Ammar Bouallègue |
PSIVT | 5 |
| 2017 | The S-Hock dataset: A new benchmark for spectator crowd analysis
Francesco Setti, Davide Conigliaro, Paolo Rota, Chiara Bassetti, Nicola Conci, Nicu Sebe, Marco Cristani |
Comput. Vis. Image Underst. | 5 |
| 2017 | Automatic Synchronization of Multi-user Photo GalleriesabstractIn this paper we address the issue of photo galleries synchronization, where pictures related to the same event are collected by different users. Existing solutions to address the problem are usually based on unrealistic assumptions, like time consistency across photo galleries, and often heavily rely on heuristics, therefore limiting the applicability to real-world scenarios. We propose a solution that achieves better generalization performance for the synchronization task compared to the available literature. The method is characterized by three stages: at first, deep convolutional neural network features are used to assess the visual similarity among the photos; then, pairs of similar photos are detected across different galleries and used to construct a graph; eventually, a probabilistic graphical model is used to estimate the temporal offset of each pair of galleries, by traversing the minimum spanning tree extracted from this graph. The experimental evaluation is conducted on four publicly available datasets covering different types of events, demonstrating the strength of our proposed method. A thorough discussion of the obtained results is provided for a critical assessment of the quality in synchronization. Emanuele Sansone, Konstantinos Apostolidis, Nicola Conci, Giulia Boato, Vasileios Mezaris, Francesco G. B. De Natale |
IEEE Trans. Multim. | 3 |
| 2016 | Crowd behavior identificationabstractIn this paper we present a novel method for crowd behavior identification. In our method, the motion flow field is obtained from the video by computing the dense optical flow. Then, a thermal diffusion process (TDP) is exploited to increase the coherence of the motion flow. Approximating the moving particles to individuals, their interaction forces are computed using a modified variant of the social force model (M-SFM) to highlight potential particles of interest. Besides capturing the effect of neighboring individuals on each other, the M-SFM also takes into account the crowd disorder, usually triggered by regions of high interactions. The experimental evaluation is conducted on a set of benchmark video sequences, commonly used for crowd motion analysis, and the obtained results are compared against a state of the art technique. Mohib Ullah, Nicola Conci, Francesco G. B. De Natale |
ICIP | 3 |
| 2016 | USED: a large-scale social event detection datasetabstractEvent discovery from single pictures is a challenging problem that has raised significant interest in the last decade. During this time, a number of interesting solutions have been proposed to tackle event discovery in still images. However, a large scale benchmarking image dataset for the evaluation and comparison of event discovery algorithms from single images is still lagging behind. To this aim, in this paper we provide a large-scale properly annotated and balanced dataset of 490,000 images, covering every aspect of 14 different types of social events, selected among the most shared ones in the social network. Such a large scale collection of event-related images is intended to become a powerful support tool for the research community in multimedia analysis by providing a common benchmark for training, testing, validation and comparison of existing and novel algorithms. In this paper, we provide a detailed description of how the dataset is collected, organized and how it can be beneficial for the researchers in the multimedia analysis domain. Moreover, a deep learning based approach is introduced into event discovery from single images as one of the possible applications of this dataset with a belief that deep learning can prove to be a breakthrough also in this research area. By providing this dataset, we hope to gather research community in the multimedia and signal processing domains to advance this application. Kashif Ahmad, Nicola Conci, Giulia Boato, Francesco G. B. De Natale |
MMSys | 2 |
| 2015 | The S-HOCK dataset: Analyzing crowds at the stadiumabstractThe topic of crowd modeling in computer vision usually assumes a single generic typology of crowd, which is very simplistic. In this paper we adopt a taxonomy that is widely accepted in sociology, focusing on a particular category, the spectator crowd, which is formed by people “interested in watching something specific that they came to see” [6]. This can be found at the stadiums, amphitheaters, cinema, etc. In particular, we propose a novel dataset, the Spectators Hockey (S-HOCK), which deals with 4 hockey matches during an international tournament. In the dataset, a massive annotation has been carried out, focusing on the spectators at different levels of details: at a higher level, people have been labeled depending on the team they are supporting and the fact that they know the people close to them; going to the lower levels, standard pose information has been considered (regarding the head, the body) but also fine grained actions such as hands on hips, clapping hands etc. The labeling focused on the game field also, permitting to relate what is going on in the match with the crowd behavior. This brought to more than 100 millions of annotations, useful for standard applications as people counting and head pose estimation but also for novel tasks as spectator categorization. For all of these we provide protocols and baseline results, encouraging further research. Davide Conigliaro, Paolo Rota, Francesco Setti, Chiara Bassetti, Nicola Conci, Nicu Sebe, Marco Cristani |
CVPR | 5 |
| 2015 | Real-life violent social interaction detectionabstractThis paper proposes a method to detect and localize dyadic human interactions in real videos. The idea stems from the significant difference between an action performed by a single subject and an interaction between two persons. In the first case all the visual information is concentrated on the subject, while in the latter case the action of a person is related to the interacting person's attitude, following an action/reaction principle. This kind of behavior is significant especially in natural and real scenarios, in which people are moving freely without the awareness of being recorded. To highlight these features and provide researchers with a common ground for comparisons, we have collected and annotated a new dataset, retrieving from YouTube 30 different videos of a specific type of interaction, namely urban fight situations. The proposed dataset is one of the most challenging annotated video collection concerning dyadic interactions, due to the intrinsic intra-class variability characterizing real fights. In addition, we provide an extensive experimental analysis on this dataset and we demonstrate that the visual information extracted in the area associated to the interpersonal space plays a fundamental role in detecting fights. Paolo Rota, Nicola Conci, Nicu Sebe, James M. Rehg |
ICIP | 2 |
| 2015 | Traffic accident detection through a hydrodynamic lensabstractIn this paper we present a novel method for automatic traffic accident detection, based on Smoothed Particles Hydrodynamics (SPH). In our method, a motion flow field is obtained from the video through dense optical flow extraction. Then a thermal diffusion process (TDP) is exploited to turn the motion flow field into a coherent motion field. Approximating the moving particles to individuals, their interaction forces, represented as endothermic reactions, are computed using the enthalpy measure, thus obtaining the potential particles of interest. Furthermore, we exploit SPH that accumulates the contribution of each particle in a weighted form, based on a kernel function. The experimental evaluation is conducted on a set of video sequences collected from Youtube, and the obtained results are compared against a state of the art technique. Mohib Ullah, Hina Afridi, Nicola Conci, Francesco G. B. De Natale |
ICIP | 4 |
| 2015 | Human interaction recognition in the wild: Analyzing trajectory clustering from multiple-instance-learning perspectiveabstractIn this paper, we propose a framework to recognize complex human interactions. First, we adopt trajectories to represent human motion in a video. Then, the extracted trajectories are clustered into different groups (named as local motion patterns) using the coherent filtering algorithm. As trajectories within the same group exhibit similar motion properties (i.e., velocity, direction), we adopt the histogram of large-displacement optical flow (denoted as HO-LDOF) as the group motion feature vector. Thus, each video can be briefly represented by a collection of local motion patterns that are described by the HO-LDOF. Finally, classification is achieved using the citation-KNN, which is a typical multiple-instance-learning algorithm. Experimental results on the TV human interaction dataset and the UT human interaction dataset demonstrate the applicability of our method. Bo Zhang 0045, Paolo Rota, Nicola Conci, Francesco G. B. De Natale |
ICME | 3 |
| 2015 | Computer vision based room interior designabstractThis paper introduces a new application of computer vision. To the best of the author’s knowledge, it is the first attempt to incorporate computer vision techniques into room interior designing. The computer vision based interior designing is achieved in two steps: object identification and color assignment. The image segmentation approach is used for the identification of the objects in the room and different color schemes are used for color assignment to these objects. The proposed approach is applied to simple as well as complex images from online sources. The proposed approach not only accelerated the process of interior designing but also made it very efficient by giving multiple alternatives. Nasir Ahmad, Kashif Ahmad, Nicola Conci |
ICMV | 4 |
| 2015 | Segmentation of Discriminative Patches in Human Activity VideoabstractIn this article, we present a novel approach to segment discriminative patches in human activity videos. First, we adopt the spatio-temporal interest points (STIPs) to represent significant motion patterns in the video sequence. Then, nonnegative sparse coding is exploited to generate a sparse representation of each STIP descriptor. We construct the feature vector for each video by applying a two-stage sum-pooling and l 2 -normalization operation. After training a multi-class classifier through the error-correcting code SVM, the discriminative portion of each video is determined as the patch that has the highest confidence while also being correctly classified according to the video category. Experimental results show that the video patches extracted by our method are more separable, while preserving the perceptually relevant portion of each activity. Bo Zhang 0045, Nicola Conci, Francesco G. B. De Natale |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2014 | Camera viewpoint change detection for interaction analysis in TV showsabstractIn this paper, we propose a novel approach to detect abrupt camera viewpoint changes in edited video materials (movies, TV shows), to improve human activity recognition. The motivation for this work lies in the difficulty of correctly identifying actions in case of camera motion and viewpoint changes, because of the abrupt variations in the appearance model of the scene, which significantly deteriorate the continuity of the spatio-temporal features under investigation. To this aim, we compute the motion interchange pattern (MIP) for each pixel in a video, from which a feature descriptor is constructed for the entire frame. The change in camera viewpoint is achieved through the one-class SVM. We apply our detector on the TV human interaction dataset (TVHI). The experimental results show that our approach can distinguish the abrupt changes with a high accuracy, allowing for an improvement also in the activity recognition performance. Bo Zhang 0045, Nicola Conci, Francesco G. B. De Natale |
ICIP | 2 |
| 2014 | You Talkin' to Me?: Recognizing Complex Human Interactions in Unconstrained VideosabstractNowadays, due to the exponential growth of the user generated videos and the prevailing videos sharing communities such as YouTube and Hulu, recognizing complex human activities in the wild becomes increasingly important in the research community. These videos are hard to study due to the frequent changes of camera viewpoint, multiple people moving in the scene, fast body movements, and varied lengths of video clips. In this paper, we propose a novel framework to analyze human interactions in TV shows. Firstly, we exploit the motion interchange pattern (MIP) to detect camera viewpoint changes in a video, and extract the salient motion points in the bounding box that covers the region of interest (ROI) in each frame. Then, we compute the large displacement optical flow for the salient pixels in the bounding box, and build the histogram of oriented optical flow as the motion feature vector for each frame. Finally, the self-similarity matrix (SSM) is adopted to capture the global temporal correlation of frames in a video. After extracting the SSM descriptors, the video feature vector can be constructed through different encoding approaches. The proposed framework works well in practice for unconstrained videos. We validate our approach on the TV human interaction (TVHI) dataset, and the experimental results demonstrate the efficacy of our strategy. Bo Zhang 0045, Yan Yan 0002, Nicola Conci, Nicu Sebe |
ACM Multimedia | 3 |
| 2014 | Collaborative creativity: The Music Room
Fabio Morreale, Antonella De Angeli, Raul Masu, Paolo Rota, Nicola Conci |
Pers. Ubiquitous Comput. | 5 |
| 2013 | Structured learning for crowd motion segmentationabstractIn this paper we present a novel method for motion segmentation in crowded scenes, based on statistical modeling for structured prediction using a Conditional Random Field (CRF). As opposed to other conditional Markov models, CRF overcomes the label bias problem, making it suitable for crowd motion analysis. In our method, a grid of particles is initialized on the scene, and advected using optical flow. The particles are exploited to extract motion patterns, used as input priors for CRF training. Furthermore, we exploit min cut/max flow algorithm to remove the residual noise and highlight the main directions of crowd motion. The experimental evaluation is conducted on a set of benchmark video sequences, commonly used for crowd motion analysis, and the obtained results are compared against other state of the art techniques. Nicola Conci |
ICIP | 2 |
| 2013 | Recognition of social interactions based on feature selection from visual codebooksabstractIn this paper we propose a novel method to recognize different types of two-person interactions in video sequences. After extracting the spatio-temporal interest points (STIPs) from the visual scene through the 3D Harris detector, K-means clustering is applied to construct the visual codebook. We adopt a new feature selection procedure, called knowledge gain, based on the rough set theory to identify the most meaningful visual words in the codebook. For each video sequence, the histogram of selected visual words is used to train a multi-class SVM classifier. The algorithm is tested on two different datasets in order to demonstrate the applicability of the technique in different environmental configurations. Experimental results show that knowledge gain can improve the classification performance. Bo Zhang 0045, Francesco G. B. De Natale, Nicola Conci |
ICIP | 3 |
| 2013 | Swarm Optimization of Structuring Elements for VHR Image ClassificationabstractMathematical morphology has shown to be an effective tool to extract spatial information for remote-sensing image classification. Its application is performed by means of a structuring element (SE), whose shape and size play a fundamental role for appropriately extracting structures in complex regions such as urban areas. In this letter, we propose a novel method, which automatically tailors both the shape and the size of the SE according to the considered classification task. For this purpose, the SE design is formulated as an optimization problem within a particle swarm optimization framework. The experiments conducted on two real images suggest that better accuracies can be achieved with respect to the common procedure for finding the best regular SE, which, so far, is heuristically done. Abdelhamid Daamouche, Farid Melgani, Naif Alajlan, Nicola Conci |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2013 | Video-Based Human Behavior Understanding: A SurveyabstractUnderstanding human behaviors is a challenging problem in computer vision that has recently seen important advances. Human behavior understanding combines image and signal processing, feature extraction, machine learning, and 3-D geometry. Application scenarios range from surveillance to indexing and retrieval, from patient care to industrial safety and sports analysis. Given the broad set of techniques used in video-based behavior understanding and the fast progress in this area, in this paper we organize and survey the corresponding literature, define unambiguous key terms, and discuss links among fundamental building blocks ranging from human detection to action and interaction recognition. The advantages and the drawbacks of the methods are critically discussed, providing a comprehensive coverage of key aspects of video-based human behavior understanding, available datasets for experimentation and comparisons, and important open research issues. Paulo Vinicius Koerich Borges, Nicola Conci, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Optical Image Classification: A Ground-Truth Design FrameworkabstractIn the remote sensing field, ground-truth design for collecting training samples represents a tricky and critical problem since it has a direct impact on most of the subsequent image processing and analysis steps. In this paper, we propose a novel framework for assisting a human user in designing ground-truth by photointerpretation for optical remote sensing image classification. The proposed approach is (almost) completely automatic and comprehensive since it aims at assisting the human user from the first to the last step of the process. It is based on unsupervised methods of segmentation and clustering, in order to investigate both the spatial and the spectral information in the process of ground-truth design. The resulting ground-truth is classifier-free and can be further improved by making it classifier-driven through an active learning process. To validate the proposed framework, an experimental study was conducted on very high spatial resolution and hyperspectral images acquired by the IKONOS and the Reflective Optics System Imaging Spectrometer sensors, respectively. The obtained results show the usefulness and effectiveness of the proposed approach. Edoardo Pasolli, Farid Melgani, Naif Alajlan, Nicola Conci |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2012 | Context-aware mobile crowdsourcingabstractUbiquity of internet-connected media-and sensor-equipped portable devices has emerged a range of opportunities for direct involvement of citizens into public decision making, leading to a new participatory format of public administration functioning. Intersecting the power of the crowdsourcing problem-solving paradigm by directly relying on human intelligence, with instantaneity and situation-awareness of mobile technologies, one gets a context-aware crowdsourcing approach for problem-solving in the right circumstances with the right people. In this paper, we present a prototype implementation of a context-aware mobile crowdsourcing system that enables the deployment and execution of crowdsourcing campaigns with users carrying mobile devices. The system is designed to maximize conditions for user participation, while minimizing the usage of energy. The paper describes the system architecture, defines an optimized sampling algorithm, and outlines a preliminary experimentation study carried out. Andrei Tamilin, Iacopo Carreras, Emmanuel Ssebaggala, Alfonse Opira, Nicola Conci |
UbiComp | 5 |
| 2010 | Detection and enhancement of moving objects in surveillance centric codingabstractThe coexistence of multiple cameras, especially in wireless video surveillance systems imposes severe constraints in terms of computational resources, power supply and bandwidth. These limitations hamper coding and transmission of high-quality video streams. In this paper, we propose a new approach for video coding and transmission of surveillance video. It integrates the coding sub-system and the motion detection module to enhance the quality of moving objects in low-bit rate streams. At the sender side, videos are downsampled and compressed. At the decoder, the video is upsampled and the foreground quality is enhanced by detecting meaningful edges of moving objects via the Hough transform. The quality of the background is also enhanced through progressive update of the background model. Nicola Conci, Ebroul Izquierdo |
ICASSP | 1 |
| 2010 | Learning and matching human activities using regular expressionsabstractIn this paper we propose a novel method to analyze trajectories in surveillance scenarios relying on automatically learned Context-Free Grammars. Given a training corpus of trajectories associated to a set of actions, an initial processing is carried out to extract the syntactical structure of the activities; then, the rules characterizing different behaviors are retrieved and coded as CFG models. The classification of the new trajectories vs the learned templates is performed through a parsing engine allowing the online recognition as well as the detection of nested activities. The proposed system has been validated in the framework of assisted living applications. The obtained results demonstrate the capability of the system in recognizing activity patterns in different configurations, also in presence of noise. Mattia Daldoss, Nicola Piotto, Nicola Conci, Francesco G. B. De Natale |
ICIP | 3 |
| 2009 | Hierarchical Matching of 3D Pedestrian Trajectories for Surveillance ApplicationsabstractIn this paper we propose a string-based approach to effectively represent trajectories in the 3D space. The strategy is coupled with a syntactical matching algorithm that allows evaluating the similarity of the retrieved data with pre-stored templates. The symbolic representation of the trajectory, is the core of the proposed system, which helps discriminating among different tracks using a modified version of the edit-distance. The hierarchical application of the algorithm on the spatial and temporal components helps detecting anomalous trajectories, and has proven to be robust in automatically learning new instances or classes of paths. We present the results achieved by performing a number of tests in an indoor lab used as a testbed for assisted living applications. The algorithm can discriminate among different classes of trajectories and can recognize actions and detect anomalies within the same class. Nicola Piotto, Francesco G. B. De Natale, Nicola Conci |
AVSS | 3 |
| 2009 | Camera placement using particle swarm optimization in visual surveillance applicationsabstractIn this paper, we propose a novel approach to automatically plan video cameras positioning in indoor environments for surveillance applications. In order to ensure maximum coverage of the observed scene, we have implemented an ad-hoc tool based on the particle swarm optimization. The camera is modeled as a 2D function. A Rayleigh distribution is used to characterize the relevance of the observed object with respect to the distance from the camera and a Gaussian distribution is adopted to model the horizontal field-of-view. Knowing the environment characteristics and the location of obstacles and walls, it is possible to derive a reliable positioning of the sensors. Results are first presented for very simple scenarios; the algorithm is then run to solve the problem in a more realistic environment model. Nicola Conci, Leonardo Lizzi |
ICIP | 1 |
| 2009 | Hand tracking and trajectory analysis for physical rehabilitationabstractIn this work we present a framework for physical rehabilitation, which is based on hand tracking. One particular requirement in physical rehabilitation is the capability of the patient to correctly reproduce a specific path, following an example provided by the medical staff. Currently, these assignments are typically performed manually, and a nurse or doctor, who supervises the correctness of the movement, constantly assists the patient throughout the whole rehabilitation process. With the proposed system, our aim is to provide medical institutions and patients with a low-cost and portable instrument to automatically assess the rehabilitation improvements. To evaluate the performance of the exercise, and to determine the distance between the trial and the reference path, we adopted the dynamic time warping (DTW) and the longest common sub-sequence (LCSS) as discriminating metrics. Trajectories and numerical values are then stored to track the history of the patient and appraise the improvements of the rehabilitation process over time. Thanks to the tests conducted with real patients, it has been possible to evaluate the quality of the proposed tool, in terms of both graphical interface and functionalities. Giulia Boato, Nicola Conci, Mattia Daldoss, Francesco G. B. De Natale, Nicola Piotto |
MMSP | 2 |
| 2009 | Syntactic Matching of Trajectories for Ambient Intelligence ApplicationsabstractIn this paper we propose a novel approach for syntactic description and matching of object trajectories in digital video, suitable for classification and recognition purposes. Trajectories are first segmented by detecting the meaningful discontinuities in time and space, and are successively expressed through an ad-hoc syntax. A suitable metric is then proposed, which allows determining the similarity among trajectories, based on the so-called inexact or approximate matching. The metric mimics the algorithms used in bio-informatics to match DNA sequences, and returns a score, which allows identifying the analogies among different trajectories on both global and local basis. The tool can therefore be adopted for the analysis, classification, and learning of motion patterns, in activity detection or behavioral understanding. Nicola Piotto, Nicola Conci, Francesco G. B. De Natale |
IEEE Trans. Multim. | 2 |
| 2008 | Syntactic matching of pedestrian trajectories for behavioral analysisabstractIn the present work we propose a new approach to dynamically characterize trajectories for a syntactic spatio-temporal alignment that can be applied in the context of behavioral analysis and anomalous activity detection. The developed architecture is based on a symbolic representation of the trajectory, exploiting the framework of the so-called edit-distance. The acquired trajectory samples are filtered to identify the most significant spatio-temporal discontinuities: these key points are converted into a string-based domain where the matching of trajectory pairs can be expressed in terms of global alignment between symbols, similarly to DNA string matching algorithms. The extraction, characterization and alignment of trajectories have been tested in different environments, demonstrating the reliability of the achieved results and the viability of the solution for video surveillance and domotics applications. Nicola Piotto, Nicola Conci, Francesco G. B. De Natale |
MMSP | 2 |
| 2007 | Natural Human-Machine Interface using an Interactive Virtual BlackboardabstractInput peripherals such as mouse, tablet or touchscreen, significantly contributed to ease the attitude of humans towards computing machines. They reduce the need of a keyboard and make the interaction with the computer faster and more instinctive, in particular for unskilled users. Next step would be the complete removal of any tangible device, towards the concept of "disappearing computer". In this paper we propose an interactive virtual blackboard, based on a video processing and gesture recognition engine, which enables the user interacting almost seamlessly with the system giving commands, writing, and manipulating objects on a projected visual interface. Nicola Conci, Paolo Ceresato, Francesco G. B. De Natale |
ICIP (5) | 1 |
| 2007 | Multiple description video coding using coefficients ordering and interpolation
Nicola Conci, Francesco G. B. De Natale |
Signal Process. Image Commun. | 1 |
| 2005 | A cross-layer approach for efficient MPEG-4 video streaming using multicarrier spread-spectrum transmission and unequal error protectionabstractIn this work, a novel methodology for the efficient multiplexing and streaming of MPEG4 video over wireless networks is presented and discussed. The proposed cross-layer adaptation jointly exploits variable-bitrate (VBR) multi-carrier code-division multiplexing (MC-CDM) and MPEG4 fine-grain-scalability (FGS) in order to provide unequal error protection to the transmitted video stream. A shared bandwidth is partitioned into orthogonal sub-channels in order to multiplex different layers of MPEG4-coded signals. Lower layers are assigned a higher number of sub-channels (and hence an increased frequency diversity) as compared to FGS enhancement layers, in order to provide a differentiated protection against channel degradations. Results achieved in terms of PSNR show that the VBR MC-CDM technique can provide better results than conventional MPEG4 single-layer multicarrier spread spectrum transmission. Nicola Conci, Giovanni Berlanda Scorza, Claudio Sacchi |
ICIP (1) | 1 |
| 2005 | A wireless multimedia framework for the management of emergency situations in automotive applications: The AIDER system
Nicola Conci, Francesco G. B. De Natale, J. Bustamante, S. Zangherati |
Signal Process. Image Commun. | 1 |