VLDB 2026 Research / reviewers in the wild / expert
Lamberto Ballan
dblp:24/1297
· DBLP profile ↗
54ranked-venue papers
15as first author
22since 2021 · last 2026
0000-0003-0819-851XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 13 first-author · 13 since 2021Artificial intelligence and machine learning · 30 · 4 first-author · 17 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models
Kuinan Hou, Jing Mi, Marco Zorzi, Lamberto Ballan, Alberto Testolin |
ICPR (11) | 4 |
| 2026 | Personal-3D: A Comprehensive Benchmark for Personalized Embodied AI AgentsabstractAbstract Despite significant progress in Embodied AI, current agents largely operate under generic task specifications and struggle to reason about user-specific semantics that naturally arise in human-centered environments. In domestic settings, objects are often associated with particular individuals, requiring agents to interpret personalized instructions such as ownership and preference when navigating and acting in 3D spaces. We introduce PersONAL-3D ( PERS onalized O bject N avigation A nd L ocalization), a benchmark designed to study personalized spatial reasoning in embodied environments. PersONAL-3D focuses on domestic scenarios in which an agent must navigate to target objects associated with specific individuals, given natural-language instructions such as “find Lily’s backpack” . The benchmark includes 2,000+ curated evaluation episodes across 30+ photorealistic HM3D homes. Each episode pairs a natural-language scene description that specifies object ownership with a user-specific query, requiring models to ground personalized semantics in 3D space. PersONAL-3D supports two evaluation settings: (1) Personalized Active Navigation in previously unseen environments, and (2) Personalized Object Grounding in pre-explored scenes or directly on 3D point clouds. Experiments with state-of-the-art baselines reveal a substantial gap to human performance, with the best navigation model underperforming humans by about 45 percentage points in Success Rate and 25 percentage points in Path Efficiency, indicating that current methods struggle to perceive, act, and reason over personalized information in embodied contexts. This work highlights personalization as a critical and largely unsolved challenge for embodied AI systems operating in real-world assistive scenarios. Filippo Ziliotto, Jelin Raphael Akkara, Alessandro Daniele, Lamberto Ballan, Luciano Serafini, Tommaso Campari |
Int. J. Comput. Vis. | 4 |
| 2025 | TANGO: Training-free Embodied AI Agents for Open-world TasksabstractLarge Language Models (LLMs) have demonstrated excellent capabilities in composing various modules together to create programs that can perform complex reasoning tasks on images. In this paper, we propose TANGO, an approach that extends the program composition via LLMs already observed for images, aiming to integrate those capabilities into embodied agents capable of observing and acting in the world. Specifically, by employing a simple PointGoal Navigation model combined with a memory-based exploration policy as a foundational primitive for guiding an agent through the world, we show how a single model can address diverse tasks without additional training. We task an LLM with composing the provided primitives to solve a specific task, using only a few in-context examples in the prompt. We evaluate our approach on three key Embodied AI tasks: Open-Set ObjectGoal Navigation, Multi-Modal Lifelong Navigation, and Open Embodied Question Answering, achieving state-of-the-art results without any specific fine-tuning in challenging zero-shot scenarios. Filippo Ziliotto, Tommaso Campari, Luciano Serafini, Lamberto Ballan |
CVPR | 4 |
| 2025 | Following the Human Thread in Social NavigationabstractThe success of collaboration between humans and robots in shared environments relies on the robot's real-time adaptation to human motion. Specifically, in Social Navigation, the agent should be close enough to assist but ready to back up to let the human move freely, avoiding collisions. Human trajectories emerge as crucial cues in Social Navigation, but they are partially observable from the robot's egocentric view and computationally complex to process.
We present the first Social Dynamics Adaptation model (SDA) based on the robot's state-action history to infer the social dynamics. We propose a two-stage Reinforcement Learning framework: the first learns to encode the human trajectories into social dynamics and learns a motion policy conditioned on this encoded information, the current status, and the previous action. Here, the trajectories are fully visible, i.e., assumed as privileged information. In the second stage, the trained policy operates without direct access to trajectories. Instead, the model infers the social dynamics solely from the history of previous actions and statuses in real-time.
Tested on the novel Habitat 3.0 platform, SDA sets a novel state-of-the-art (SotA) performance in finding and following humans.
The code can be found at https://github.com/L-Scofano/SDA. Luca Scofano, Alessio Sampieri, Tommaso Campari, Valentino Sacco, Indro Spinelli, Lamberto Ballan, Fabio Galasso |
ICLR | 6 |
| 2025 | Temporally-Aware Supervised Contrastive Learning for Polyp Counting in Colonoscopy
Luca Parolari, Andrea Cherubini, Lamberto Ballan, Carlo Biffi |
MICCAI (10) | 3 |
| 2025 | KD-Mamba: Selective state space models with knowledge distillation for trajectory prediction
Shaokang Cheng, Shiru Qu, Lamberto Ballan |
Comput. Vis. Image Underst. | 4 |
| 2024 | Harlequin: Color-Driven Generation of Synthetic Data for Referring Expression Comprehension
Luca Parolari, Elena Izzo, Lamberto Ballan |
ICPR (18) | 3 |
| 2024 | Distilling Knowledge for Short-to-Long Term Trajectory PredictionabstractLong-term trajectory forecasting is an important and challenging problem in the fields of computer vision, machine learning, and robotics. One fundamental difficulty stands in the evolution of the trajectory that becomes more and more uncertain and unpredictable as the time horizon grows, subsequently increasing the complexity of the problem. To overcome this issue, in this paper, we propose Di-Long, a new method that employs the distillation of a short-term trajectory model forecaster that guides a student network for long-term trajectory prediction during the training process. Given a total sequence length that comprehends the allowed observation for the student network and the complementary target sequence, we let the student and the teacher solve two different related tasks defined over the same full trajectory: the student observes a short sequence and predicts a long trajectory, whereas the teacher observes a longer sequence and predicts the remaining short target trajectory. The teacher’s task is less uncertain, and we use its accurate predictions to guide the student through our knowledge distillation framework, reducing long-term future uncertainty. Our experiments show that our proposed Di-Long method is effective for long-term forecasting and achieves state-of-the-art performance on the Intersection Drone Dataset (inD) and the Stanford Drone Dataset (SDD). Guglielmo Camporese, Shaokang Cheng, Lamberto Ballan |
IROS | 4 |
| 2024 | Multi-modal transformer with language modality distillation for early pedestrian action anticipationabstractLanguage-vision integration has become an increasingly popular research direction within the computer vision field. In recent years, there has been a growing recognition of the importance of incorporating linguistic information into visual tasks, particularly in domains such as action anticipation. This integration allows anticipation models to leverage textual descriptions to gain deeper contextual understanding, leading to more accurate predictions. In this work, we focus on pedestrian action anticipation, where the objective is the early prediction of pedestrians’ future actions in urban environments. Our method relies on a multi-modal transformer model that encodes past observations and produces predictions at different anticipation times, employing a learned mask technique to filter out redundancy in the observed frames. Instead of relying solely on visual cues extracted from images or videos, we explore the impact of integrating textual information in enriching the input modalities of our pedestrian action anticipation model. We investigate various techniques for generating descriptive captions corresponding to input images, aiming to enhance the anticipation performance. Evaluation results on available public benchmarks demonstrate the effectiveness of our method in improving the prediction performance at different anticipation times compared to previous works. Additionally, incorporating the language modality in our anticipation model proved significant improvement, reaching a 29.5% increase in the F1 score at 1-second anticipation and a 16.66% increase at 4-second anticipation. These results underscore the potential of language-vision integration in advancing pedestrian action anticipation in complex urban environments. Nada Osman, Guglielmo Camporese, Lamberto Ballan |
Comput. Vis. Image Underst. | 3 |
| 2024 | FastSTI: A Fast Conditional Pseudo Numerical Diffusion Model for Spatio-Temporal Traffic Data ImputationabstractHigh-quality spatiotemporal traffic data is crucial for intelligent transportation systems (ITS) and their data-driven applications. Inevitably, the issue of missing data caused by various disturbances threatens the reliability of data acquisition. Recent studies of diffusion probability models have demonstrated the superiority of deep generative models in imputation tasks by precisely capturing the spatio-temporal correlation of traffic data. One drawback of diffusion models is their slow sampling/denoising process. In this work, we aim to accelerate the imputation process while retaining the performance. We propose a fast conditional diffusion model for spatiotemporal traffic data imputation (FastSTI). To speed up the process yet, obtain better performance, we propose the application of a high-order pseudo-numerical solver. Our method further revs the imputation by introducing a predefined alignment strategy of variance schedule during the sampling process. Evaluating FastSTI on two types of real-world traffic datasets (traffic speed and flow) with different missing data scenarios proves its ability to impute higher-quality samples in only six sampling steps, especially under high missing rates (60%$\sim ~90$%). The experimental results illustrate a speed-up of$\textbf {8.3} \times $faster than the current state-of-the-art model while achieving better performance. Shaokang Cheng, Nada Osman, Shiru Qu, Lamberto Ballan |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Weakly-Supervised Visual-Textual Grounding with Semantic Prior Refinement
Davide Rigoni 0001, Luca Parolari, Luciano Serafini, Alessandro Sperduti, Lamberto Ballan |
BMVC | 5 |
| 2023 | TAMformer: Multi-Modal Transformer with Learned Attention Mask for Early Intent PredictionabstractHuman intention prediction is a growing area of research where an activity in a video has to be anticipated by a vision-based system. To this end, the model creates a representation of the past, and subsequently, it produces future hypotheses about upcoming scenarios. In this work, we focus on pedestrians’ early intention prediction in which, from a current observation of an urban scene, the model predicts the future activity of pedestrians that approach the street. Our method is based on a multi-modal transformer that encodes past observations and produces multiple predictions at different anticipation times. Moreover, we propose to learn the attention masks of our transformer-based model (Temporal Adaptive Mask Transformer) in order to weigh differently present and past temporal dependencies. We investigate our method on several public benchmarks for early intention prediction, improving the prediction performances at different anticipation times compared to the previous works. Nada Osman, Guglielmo Camporese, Lamberto Ballan |
ICASSP | 3 |
| 2023 | Exploiting Proximity-Aware Tasks for Embodied Social NavigationabstractLearning how to navigate among humans in an occluded and spatially constrained indoor environment, is a key ability required to embodied agents to be integrated into our society. In this paper, we propose an end-to-end architecture that exploits Proximity-Aware Tasks (referred as to Risk and Proximity Compass) to inject into a reinforcement learning navigation policy the ability to infer common-sense social behaviours. To this end, our tasks exploit the notion of immediate and future dangers of collision. Furthermore, we propose an evaluation protocol specifically designed for the Social Navigation Task in simulated environments. This is done to capture fine-grained features and characteristics of the policy by analyzing the minimal unit of human-robot spatial interaction, called Encounter. We validate our approach on Gibson4+ and Habitat-Matterport3D datasets. Enrico Cancelli, Tommaso Campari, Luciano Serafini, Angel X. Chang, Lamberto Ballan |
ICCV | 5 |
| 2023 | On the problem of recommendation for sensitive users and influential items: Simultaneously maintaining interest and diversityabstractRecommender systems, in real-world circumstances, tend to limit user exposure to certain topics and to overexpose them to others to maximize performance. However, repeated exposure to biased content could lead to the so-called echo chamber phenomenon: especially in social network environments, people encounter only information that reflects their previous beliefs and opinions, reinforcing them. This phenomenon could have worrying consequences for society, including the spread of aggressive, unhealthy, or risky behaviors. Some persons can be more affected than others by echo-chambers. We define as sensitive the users whose behavior could be influenced by the over- or under-exposure to certain items due to the echo-chamber effect, and as influential the items that could influence the behavior of such users. In this paper, we address the problem of recommending influential items to sensitive users. We formalize the problem and propose three techniques that can be used to diversify the distributions of influential items in order to positively affect sensitive users’ behavior. Recommendations that meet this diversity criterion could potentially avoid dangerous societal consequences and simultaneously promote healthier lifestyles. We tested the proposed techniques in a real-world dataset by considering two different case studies that involved potentially aggressive and potentially depressed users. All techniques have been proven to be effective and allow high performance to be maintained while diversifying recommendations. Alvise De Biasio, Merylin Monaro, Luca Oneto, Lamberto Ballan, Nicolò Navarin |
Knowl. Based Syst. | 4 |
| 2022 | Where are my Neighbors? Exploiting Patches Relations in Self-Supervised Vision Transformer
Guglielmo Camporese, Elena Izzo, Lamberto Ballan |
BMVC | 3 |
| 2022 | Online Learning of Reusable Abstract Models for Object Goal NavigationabstractIn this paper, we present a novel approach to incrementally learn an Abstract Model of an unknown environment, and show how an agent can reuse the learned model for tackling the Object Goal Navigation task. The Abstract Model is a finite state machine in which each state is an abstraction of a state of the environment, as perceived by the agent in a certain position and orientation. The perceptions are high-dimensional sensory data (e.g., RGB-D images), and the abstraction is reached by exploiting image segmentation and the Taskonomy model bank. The learning of the Abstract Model is accomplished by executing actions, observing the reached state, and updating the Abstract Model with the acquired information. The learned models are memorized by the agent, and they are reused whenever it recognizes to be in an environment that corresponds to the stored model. We investigate the effectiveness of the proposed approach for the Object Goal Navigation task, relying on public benchmarks. Our results show that the reuse of learned Abstract Models can boost performance on Object Goal Navigation. Tommaso Campari, Leonardo Lamanna 0001, Paolo Traverso, Luciano Serafini, Lamberto Ballan |
CVPR | 5 |
| 2022 | How many Observations are Enough? Knowledge Distillation for Trajectory ForecastingabstractAccurate prediction of future human positions is an essential task for modern video-surveillance systems. Current state-of-the-art models usually rely on a “history” of past tracked locations (e.g., 3 to 5 seconds) to predict a plausible sequence of future locations (e.g., up to the next 5 seconds). We feel that this common schema neglects critical traits of realistic applications: as the collection of input trajectories involves machine perception (i.e., detection and tracking), incorrect detection and fragmentation errors may accumulate in crowded scenes, leading to tracking drifts. On this account, the model would be fed with corrupted and noisy input data, thus fatally affecting its prediction performance. In this regard, we focus on delivering accurate predictions when only few input observations are used, thus potentially lowering the risks associated with automatic perception. To this end, we conceive a novel distillation strategy that allows a knowledge transfer from a teacher network to a student one, the latter fed with fewer observations (just two ones). We show that a properly defined teacher super-vision allows a student network to perform comparably to state-of-the-art approaches that demand more observations. Besides, extensive experiments on common trajectory forecasting datasets highlight that our student network better generalizes to unseen scenarios. Alessio Monti, Angelo Porrello, Simone Calderara, Pasquale Coscia, Lamberto Ballan, Rita Cucchiara |
CVPR | 5 |
| 2022 | Early Pedestrian Intent Prediction via Features EstimationabstractAnticipating human motion is an essential requirement for autonomous vehicles and robots in order to primary guarantee people’s safety. In urban scenarios, they interact with humans, the surrounding environment, and other vehicles relying on several cues to forecast crossing or not crossing intentions. For these reasons, this challenging task is often tackled using both visual and non-visual features to anticipate future actions from 2 s to 1 s earlier the event. Our work primarily aims to revise this standard evaluation protocol to forecast crossing events as early as possible. To this end, we conceive a solution upon an extensively used model for egocentric action anticipation (RU-LSTM), proposing to envision future features, or modalities, that can better infer human intentions using a properly attention-based fusion mechanism. We validate our model against JAAD and PIE datasets and demonstrate that an intent prediction model can benefit from these additional clues for anticipating pedestrians crossing events. Nada Osman, Enrico Cancelli, Guglielmo Camporese, Pasquale Coscia, Lamberto Ballan |
ICIP | 5 |
| 2022 | Aligning and linking entity mentions in image, text, and knowledge base
Shahi Dost, Luciano Serafini, Marco Rospocher, Lamberto Ballan, Alessandro Sperduti |
Data Knowl. Eng. | 4 |
| 2021 | Conditional Variational Capsule Network for Open Set RecognitionabstractIn open set recognition, a classifier has to detect unknown classes that are not known at training time. In order to recognize new categories, the classifier has to project the input samples of known classes in very compact and separated regions of the features space for discriminating samples of unknown classes. Recently proposed Capsule Networks have shown to outperform alternatives in many fields, particularly in image recognition, however they have not been fully applied yet to open-set recognition. In capsule networks, scalar neurons are replaced by capsule vectors or matrices, whose entries represent different proper-ties of objects. In our proposal, during training, capsules features of the same known class are encouraged to match a pre-defined gaussian, one for each class. To this end, we use the variational autoencoder framework, with a set of gaussian priors as the approximation for the posterior distribution. In this way, we are able to control the compactness of the features of the same class around the center of the gaussians, thus controlling the ability of the classifier in detecting samples from unknown classes. We conducted several experiments and ablation of our model, obtaining state of the art results on different datasets in the open set recognition and unknown detection tasks. Yunrui Guo, Guglielmo Camporese, Wenjing Yang 0002, Alessandro Sperduti, Lamberto Ballan |
ICCV | 5 |
| 2021 | AC-VRNN: Attentive Conditional-VRNN for multi-future trajectory predictionabstractAnticipating human motion in crowded scenarios is essential for developing intelligent transportation systems, social-aware robots and advanced video surveillance applications. A key component of this task is represented by the inherently multi-modal nature of human paths which makes socially acceptable multiple futures when human interactions are involved. To this end, we propose a generative architecture for multi-future trajectory predictions based on Conditional Variational Recurrent Neural Networks (C-VRNNs). Conditioning mainly relies on prior belief maps, representing most likely moving directions and forcing the model to consider past observed dynamics in generating future positions. Human interactions are modelled with a graph-based attention mechanism enabling an online attentive hidden state refinement of the recurrent estimation. To corroborate our model, we perform extensive experiments on publicly-available datasets (e.g., ETH/UCY, Stanford Drone Dataset, STATS SportVU NBA, Intersection Drone Dataset and TrajNet++) and demonstrate its effectiveness in crowded scenes compared to several state-of-the-art methods. Alessia Bertugli, Simone Calderara, Pasquale Coscia, Lamberto Ballan, Rita Cucchiara |
Comput. Vis. Image Underst. | 4 |
| 2021 | Am I Done? Predicting Action Progress in VideosabstractIn this article, we deal with the problem of predicting action progress in videos. We argue that this is an extremely important task, since it can be valuable for a wide range of interaction applications. To this end, we introduce a novel approach, named ProgressNet, capable of predicting when an action takes place in a video, where it is located within the frames, and how far it has progressed during its execution. To provide a general definition of action progress, we ground our work in the linguistics literature, borrowing terms and concepts to understand which actions can be the subject of progress estimation. As a result, we define a categorization of actions and their phases. Motivated by the recent success obtained from the interaction of Convolutional and Recurrent Neural Networks, our model is based on a combination of the Faster R-CNN framework, to make framewise predictions, and LSTM networks, to estimate action progress through time. After introducing two evaluation protocols for the task at hand, we demonstrate the capability of our model to effectively predict action progress on the UCF-101 and J-HMDB datasets. Federico Becattini, Tiberio Uricchio, Lorenzo Seidenari, Lamberto Ballan, Alberto Del Bimbo |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | VT-LINKER: Visual-Textual-Knowledge Entity Linkerabstract"A picture is worth a thousand words", the adage reads. However, pictures cannot replace words in terms of their ability to efficiently convey clear (mostly) unambiguous and concise knowledge. Images and text, indeed, reveal different and complementary information that, if combined, result in more information than the sum of that contained in the single media. The combination of visual and textual information can be obtained by linking the entities mentioned in the text with those shown in the pictures. To further integrate this with agent background knowledge, an additional step is necessary. That is, either finding the entities in the agent knowledge base that correspond to those mentioned in the text or shown in the picture or, extending the knowledge base with the newly discovered entities. We call this complex task Visual-Textual-Knowledge Entity Linking (VTKEL). In this paper, we precisely define the VTKEL task and present two datasets composed of 1k and 30k pictures, annotated with visual and textual entities and linked to the YAGO ontology. Successively, we develop the first unsupervised algorithm for the solution of VTKEL task. The evaluation of the algorithm shows promising results on both 1k and 30k VTKEL datasets. Shahi Dost, Luciano Serafini, Marco Rospocher, Lamberto Ballan, Alessandro Sperduti |
ECAI | 4 |
| 2020 | Knowledge Distillation for Action Anticipation via Label SmoothingabstractHuman capability to anticipate near future from visual observations and non-verbal cues is essential for developing intelligent systems that need to interact with people. Several research areas, such as human-robot interaction (HRI), assisted living or autonomous driving need to foresee future events to avoid crashes or help people. Egocentric scenarios are classic examples where action anticipation is applied due to their numerous applications. Such challenging task demands to capture and model domain's hidden structure to reduce prediction uncertainty. Since multiple actions may equally occur in the future, we treat action anticipation as a multi-label problem with missing labels extending the concept of label smoothing. This idea resembles the knowledge distillation process since useful information is injected into the model during training. We implement a multi-modal framework based on long short-term memory (LSTM) networks to summarize past observations and make predictions at different time steps. We perform extensive experiments on EPIC-Kitchens and EGTEA Gaze+ datasets including more than 2500 and 100 action classes, respectively. The experiments show that label smoothing systematically improves performance of state-of-the-art models for action anticipation. Guglielmo Camporese, Pasquale Coscia, Antonino Furnari, Giovanni Maria Farinella, Lamberto Ballan |
ICPR | 5 |
| 2020 | A CNN-RNN Framework for Image Annotation from Visual Cues and Social Network MetadataabstractImages represent a commonly used form of visual communication among people. Nevertheless, image classification may be a challenging task when dealing with unclear or non-common images needing more context to be correctly annotated. Metadata accompanying images on social-media represent an ideal source of additional information for retrieving proper neighborhoods easing image annotation task. To this end, we blend visual features extracted from neighbors and their metadata to jointly leverage context and visual cues. Our models use multiple semantic embeddings to achieve the dual objective of being robust to vocabulary changes between train and test sets and decoupling the architecture from the low-level metadata representation. Convolutional and recurrent neural networks (CNNs-RNNs) are jointly adopted to infer similarity among neighbors and query images. We perform comprehensive experiments on the NUS-WIDE dataset showing that our models outperform state-of-the-art architectures based on images and metadata, and decrease both sensory and semantic gaps to better annotate images. Tobia Tesan, Pasquale Coscia, Lamberto Ballan |
ICPR | 3 |
| 2020 | Jointly Linking Visual and Textual Entity Mentions with Background Knowledge
Shahi Dost, Luciano Serafini, Marco Rospocher, Lamberto Ballan, Alessandro Sperduti |
NLDB | 4 |
| 2018 | Context-Aware Trajectory PredictionabstractHuman motion and behaviour in crowded spaces is influenced by several factors, such as the dynamics of other moving agents in the scene, as well as the static elements that might be perceived as points of attraction or obstacles. In this work, we present a new model for human trajectory prediction which is able to take advantage of both human-human and human-space interactions. The future trajectory of humans, are generated by observing their past positions and interactions with the surroundings. To this end, we propose a “context-aware” recurrent neural network LSTM model, which can learn and predict human motion in crowded spaces such as a sidewalk, a museum or a shopping mall. We evaluate our model on a public pedestrian datasets, and we contribute a new challenging dataset that collects videos of humans that navigate in a (real) crowded space such as a big museum. Results show that our approach can predict human trajectories better when compared to previous state-of-the-art forecasting models. Federico Bartoli, Giuseppe Lisanti, Lamberto Ballan, Alberto Del Bimbo |
ICPR | 3 |
| 2018 | Guest Editorial
Lamberto Ballan, Shih-Fu Chang, Gang Hua 0001, Thomas Mensink, Greg Mori, Rahul Sukthankar |
Comput. Vis. Image Underst. | 1 |
| 2018 | Learning without prejudice: Avoiding bias in webly-supervised action recognition
Christian Rupprecht 0001, Ansh Kapil, Lamberto Ballan, Federico Tombari |
Comput. Vis. Image Underst. | 4 |
| 2018 | Long-term path prediction in urban scenarios using circular distributions
Pasquale Coscia, Francesco Castaldo, Francesco Palmieri 0001, Alexandre Alahi, Silvio Savarese, Lamberto Ballan |
Image Vis. Comput. | 6 |
| 2017 | Effective Fisher vector aggregation for 3D object retrievalabstractWe formulate the task of 3D object retrieval as a visual search problem where a database containing videos of objects captured manually from different viewpoints is queried using a single image. We propose to aggregate visual information of similar views and use the Fisher vector (FV) framework to compactly represent a database of objects. Large-scale experiments on an existing video dataset that we complemented with image queries, shows that our aggregation schemes significantly outperform standard retrieval techniques. When representing our database with only 4 FVs per object, our approach performs with a mean average precision (mAP) of 73.0% on our dataset while the baseline (no aggregation) only reaches a mAP of 43.8%. It can also reach a 72.0% mAP level with a 10× smaller database than the baseline. Jean-Baptiste Boin, André Araújo 0001, Lamberto Ballan, Bernd Girod |
ICASSP | 3 |
| 2017 | Automatic image annotation via label transfer in the semantic space
Tiberio Uricchio, Lamberto Ballan, Lorenzo Seidenari, Alberto Del Bimbo |
Pattern Recognit. | 2 |
| 2016 | Knowledge Transfer for Scene-Specific Motion Prediction
Lamberto Ballan, Francesco Castaldo, Alexandre Alahi, Francesco Palmieri 0001, Silvio Savarese |
ECCV (1) | 1 |
| 2016 | Point-based path prediction from polar histograms
Pasquale Coscia, Francesco Castaldo, Francesco Palmieri 0001, Lamberto Ballan, Alexandre Alahi, Silvio Savarese |
FUSION | 4 |
| 2015 | Love Thy Neighbors: Image Annotation by Exploiting Image MetadataabstractSome images that are difficult to recognize on their own may become more clear in the context of a neighborhood of related images with similar social-network metadata. We build on this intuition to improve multilabel image annotation. Our model uses image metadata nonparametrically to generate neighborhoods of related images using Jaccard similarities, then uses a deep neural network to blend visual information from the image and its neighbors. Prior work typically models image metadata parametrically, in contrast, our nonparametric treatment allows our model to perform well even when the vocabulary of metadata changes between training and testing. We perform comprehensive experiments on the NUS-WIDE dataset, where we show that our model outperforms state-of-the-art methods for multilabel image annotation even when our model is forced to generalize to new types of metadata. Justin Johnson 0001, Lamberto Ballan, Li Fei-Fei 0001 |
ICCV | 2 |
| 2015 | Image Tag Assignment, Refinement and RetrievalabstractThis tutorial focuses on challenges and solutions for content-based image annotation and retrieval in the context of online image sharing and tagging. We present a unified review on three closely linked problems, i.e., tag assignment, tag refinement, and tag-based image retrieval. We introduce a taxonomy to structure the growing literature, understand the ingredients of the main works, clarify their connections and difference, and recognize their merits and limitations. Moreover, we present an open-source testbed, with training sets of varying sizes and three test datasets, to evaluate methods of varied learning complexity. A selected set of eleven representative works have been implemented and evaluated. During the tutorial we provide a practice session for hands on experience with the methods, software and datasets. For repeatable experiments all data and code are online at http://www.micc.unifi.it/tagsurvey Xirong Li 0001, Tiberio Uricchio, Lamberto Ballan, Marco Bertini 0001, Cees Snoek, Alberto Del Bimbo |
ACM Multimedia | 3 |
| 2015 | A data-driven approach for tag refinement and localization in web videos
Lamberto Ballan, Marco Bertini 0001, Giuseppe Serra 0001, Alberto Del Bimbo |
Comput. Vis. Image Underst. | 1 |
| 2015 | Data-driven approaches for social image and video tagging
Lamberto Ballan, Marco Bertini 0001, Tiberio Uricchio, Alberto Del Bimbo |
Multim. Tools Appl. | 1 |
| 2014 | A Cross-media Model for Automatic Image AnnotationabstractAutomatic image annotation is still an important open problem in multimedia and computer vision. The success of media sharing websites has led to the availability of large collections of images tagged with human-provided labels. Many approaches previously proposed in the literature do not accurately capture the intricate dependencies between image content and annotations. We propose a learning procedure based on Kernel Canonical Correlation Analysis which finds a mapping between visual and textual words by projecting them into a latent meaning space. The learned mapping is then used to annotate new images using advanced nearest-neighbor voting methods. We evaluate our approach on three popular datasets, and show clear improvements over several approaches relying on more standard representations. Lamberto Ballan, Tiberio Uricchio, Lorenzo Seidenari, Alberto Del Bimbo |
ICMR | 1 |
| 2013 | An evaluation of nearest-neighbor methods for tag refinementabstractThe success of media sharing and social networks has led to the availability of extremely large quantities of images that are tagged by users. The need of methods to manage efficiently and effectively the combination of media and metadata poses significant challenges. In particular, automatic image annotation of social images has become an important research topic for the multimedia community. In this paper we propose and thoroughly evaluate the use of nearest-neighbor methods for tag refinement. Extensive and rigorous evaluation using two standard large-scale datasets shows that the performance of these methods is comparable with that of more complex and computationally intensive approaches and that, differently from these latter approaches, nearest-neighbor methods can be applied to `web-scale' data. Tiberio Uricchio, Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo |
ICME | 2 |
| 2013 | Copy-move forgery detection and localization by means of robust clustering with J-Linkage
Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Luca Del Tongo, Giuseppe Serra 0001 |
Signal Process. Image Commun. | 2 |
| 2013 | Context-Dependent Logo Matching and RecognitionabstractWe contribute, through this paper, to the design of a novel variational framework able to match and recognize multiple instances of multiple reference logos in image archives. Reference logos and test images are seen as constellations of local features (interest points, regions, etc.) and matched by minimizing an energy function mixing: 1) a fidelity term that measures the quality of feature matching, 2) a neighborhood criterion that captures feature co-occurrence/geometry, and 3) a regularization term that controls the smoothness of the matching solution. We also introduce a detection/recognition procedure and study its theoretical consistency. Finally, we show the validity of our method through extensive experiments on the challenging MICC-Logos dataset. Our method overtakes, by 20%, baseline as well as state-of-the-art matching/recognition procedures. Hichem Sahbi, Lamberto Ballan, Giuseppe Serra 0001, Alberto Del Bimbo |
IEEE Trans. Image Process. | 2 |
| 2012 | Combining generative and discriminative models for classifying social images from 101 object categories
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Andrea M. Serain, Giuseppe Serra 0001, Benito F. Zaccone |
ICPR | 1 |
| 2012 | Effective Codebooks for Human Action Representation and Classification in Unconstrained VideosabstractRecognition and classification of human actions for annotation of unconstrained video sequences has proven to be challenging because of the variations in the environment, appearance of actors, modalities in which the same action is performed by different persons, speed and duration, and points of view from which the event is observed. This variability reflects in the difficulty of defining effective descriptors and deriving appropriate and effective codebooks for action categorization. In this paper, we propose a novel and effective solution to classify human actions in unconstrained videos. It improves on previous contributions through the definition of a novel local descriptor that uses image gradient and optic flow to respectively model the appearance and motion of human actions at interest point regions. In the formation of the codebook, we employ radius-based clustering with soft assignment in order to create a rich vocabulary that may account for the high variability of human actions. We show that our solution scores very good performance with no need of parameter tuning. We also show that a strong reduction of computation time can be obtained by applying codebook size reduction with Deep Belief Networks with little loss of accuracy. Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001 |
IEEE Trans. Multim. | 1 |
| 2011 | Enriching and localizing semantic tags in internet videosabstractTagging of multimedia content is becoming more and more widespread as web 2.0 sites, like Flickr and Facebook for images, YouTube and Vimeo for videos, have popularized tagging functionalities among their users. These user-generated tags are used to retrieve multimedia content, and to ease browsing and exploration of media collections, e.g.~using tag clouds. However, not all media are equally tagged by users: using the current browsers is easy to tag a single photo, and even tagging a part of a photo, like a face, has become common in sites like Flickr and Facebook; on the other hand tagging a video sequence is more complicated and time consuming, so that users just tag the overall content of a video. In this paper we present a system for automatic video annotation that increases the number of tags originally provided by users, and localizes them temporally, associating tags to shots. This approach exploits collective knowledge embedded in tags and Wikipedia, and visual similarity of keyframes and images uploaded to social sites like YouTube and Flickr. Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001 |
ACM Multimedia | 1 |
| 2011 | Event detection and recognition for semantic annotation of video
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001 |
Multim. Tools Appl. | 1 |
| 2011 | A SIFT-Based Forensic Method for Copy-Move Attack Detection and Transformation RecoveryabstractOne of the principal problems in image forensics is determining if a particular image is authentic or not. This can be a crucial task when images are used as basic evidence to influence judgment like, for example, in a court of law. To carry out such forensic analysis, various technological instruments have been developed in the literature. In this paper, the problem of detecting if an image has been forged is investigated; in particular, attention has been paid to the case in which an area of an image is copied and then pasted onto another zone to create a duplication or to cancel something that was awkward. Generally, to adapt the image patch to the new context a geometric transformation is needed. To detect such modifications, a novel methodology based on scale invariant features transform (SIFT) is proposed. Such a method allows us to both understand if a copy-move attack has occurred and, furthermore, to recover the geometric transformation used to perform cloning. Extensive experimental results are presented to confirm that the technique is able to precisely individuate the altered area and, in addition, to estimate the geometric transformation parameters with high reliability. The method also deals with multiple cloning. Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Giuseppe Serra 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2010 | Geometric tampering estimation by means of a SIFT-based forensic analysisabstractIn many application scenarios digital images play a basic role and often it is important to assess if their content is realistic or has been manipulated to mislead watcher's opinion. Image forensics tools provide answers to similar questions. This paper, in particular, focuses on the problem of detecting if a feigned image has been created by cloning an area of the image onto another zone to make a duplication or to cancel something awkward. The proposed method is based on SIFT features and allows both to understand which are the image points involved in the counterfeit attack and, furthermore, to recover the parameters of the geometric transformation. Experimental results are provided to witness the powerfulness of the proposed technique. Irene Amerini, Lamberto Ballan, Roberto Caldelli, Alberto Del Bimbo, Giuseppe Serra 0001 |
ICASSP | 2 |
| 2010 | Video event classification using string kernels
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001 |
Multim. Tools Appl. | 1 |
| 2010 | Semantic annotation of soccer videos by visual instance clustering and spatial/temporal reasoning in ontologies
Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001 |
Multim. Tools Appl. | 1 |
| 2009 | Recognizing human actions by fusing spatio-temporal appearance and motion descriptorsabstractIn this paper we propose a new method for human action categorization by using an effective combination of a new 3D gradient descriptor with an optic flow descriptor, to represent spatio-temporal interest points. These points are used to represent video sequences using a bag of spatio-temporal visual words, following the successful results achieved in object and scene classification. We extensively test our approach on the standard KTH and Weizmann actions datasets, showing its validity and good performance. Experimental results outperform state-of-the-art methods, without requiring fine parameter tuning. Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Lorenzo Seidenari, Giuseppe Serra 0001 |
ICIP | 1 |
| 2009 | Deep networks for audio event classification in soccer videosabstractIn this work is presented a novel approach for the classification of audio concepts in broadcast soccer videos using deep belief network (DBN), a probabilistic neural network with several hidden layers. Comparison with support vector machine (SVM) classifiers has been carried on, showing that our preliminary results are promisingly comparable to the state-of-the-art. Lamberto Ballan, Alessio Bazzica, Marco Bertini 0001, Alberto Del Bimbo, Giuseppe Serra 0001 |
ICME | 1 |
| 2008 | Automatic trademark detection and recognition in sport videosabstractIn this paper we describe a system for automatic detection and recognition of trademarks in sports videos. We propose a compact representation of trademarks based on SIFT feature points and a matching algorithm to robustly detect and retrieve trademarks in a variety of different sports video types. Trademark localization is performed through robust clustering of matched feature points in the video frame. A supervised machine learning approach is used to automatically adapt the similarity threshold used to assess the trademark matches. Experimental results are provided, along with an analysis of the precision and recall. Results show that our proposed technique is efficient and effectively detects and classifies trademarks. Lamberto Ballan, Marco Bertini 0001, Alberto Del Bimbo, Arjun Jain |
ICME | 1 |
| 2008 | A system for automatic detection and recognition of advertising trademarks in sports videosabstractIn this technical demonstration we show the current version of our trademark detection and recognition system that has been developed in collaboration with a sport marketing firm1 with the aim of evaluating the visibility of advertising trademarks in broadcast sporting events. We propose a semi-automatic system for detecting and retrieving trademark appearances in sports videos. A human annotator supervises the results of the automatic annotation through an interface that shows the time and the position of the detected trademarks; due to this fact the aim of the system is to provide a good recall figure, so that the supervisor can safely skip the parts of the video that have been marked as not containing a trademark, thus speeding up his work. Lamberto Ballan, Marco Bertini 0001, Arjun Jain |
ACM Multimedia | 1 |