EDBT 2026 Demo / reviewers in the wild / expert
Giovanni Maria Farinella
dblp:31/6643
· DBLP profile ↗
111ranked-venue papers
9as first author
44since 2021 · last 2026
0000-0002-6034-0432ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 70 · 5 first-author · 20 since 2021Artificial intelligence and machine learning · 61 · 4 first-author · 35 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos
Daniele Materia, Francesco Ragusa, Giovanni Maria Farinella |
ICPR (8) | 3 |
| 2026 | Online Episodic Memory Visual Query Localization with Egocentric Streaming Object MemoryabstractEpisodic memory retrieval enables wearable cameras to recall objects or events previously observed in video. However, existing formulations assume an "offline" setting with full video access at query time, limiting their applicability in real-world scenarios with power and storage-constrained wearable devices. Towards more application-ready episodic memory systems, we introduce Online Visual Query 2D (OVQ2D), a task where models process video streams online, observing each frame only once, and retrieve object localizations using a compact memory instead of full video history. We address OVQ2D with ESOM (Egocentric Streaming Object Memory), a novel framework integrating an object discovery module, an object tracking module, and a memory module that find, track, and store spatiotemporal object information for efficient querying. Experiments on Ego4D demonstrate ESOM’s superiority over other online approaches, though OVQ2D remains challenging, with top performance at only 4% success. ESOM’s accuracy increases markedly with perfect object tracking (31.91%), discovery (40.55%), or both (81.92%), underscoring the need of applied research on these components. Zaira Manigrasso, Matteo Dunnhofer, Antonino Furnari, Moritz Nottebaum, Antonio Finocchiaro, Davide Marana, Rosario Forte, Giovanni Maria Farinella, Christian Micheloni |
WACV | 8 |
| 2026 | ProSkill: Segment-Level Skill Assessment in Procedural VideosabstractSkill assessment in procedural videos is crucial for the objective evaluation of human performance in settings such as manufacturing and procedural daily tasks. Current research on skill assessment has predominantly focused on sports and lacks large-scale datasets for complex procedural activities. Existing studies typically involve only a limited number of actions, focus on either pairwise assessments (e.g., A is better than B) or on binary labels (e.g., good execution vs needs improvement). In response to these shortcomings, we introduce ProSkill, the first benchmark dataset for action-level skill assessment in procedural tasks. ProSkill provides absolute skill assessment annotations, along with pairwise ones. This is enabled by a novel and scalable annotation protocol that allows for the creation of an absolute skill assessment ranking starting from pairwise assessments. This protocol leverages a Swiss Tournament scheme for efficient pairwise comparisons, which are then aggregated into consistent, continuous global scores using an ELO-based rating system. We use our dataset to benchmark the main state-of-the-art skill assessment algorithms, including both ranking-based and pairwise paradigms. The suboptimal results achieved by the current state-of-the-art highlight the challenges and thus the value of ProSkill in the context of skill assessment for procedural videos. All data and code are available at https://fpv-iplab.github.io/ProSkill/. Michele Mazzamuto, Daniele Di Mauro, Gianpiero Francesca, Giovanni Maria Farinella, Antonino Furnari |
WACV | 4 |
| 2026 | Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistanceabstractWe present Ego-EXTRA, a video-language Egocentric Dataset for EXpert-TRAinee assistance. Ego-EXTRA features 50 hours of unscripted egocentric videos of subjects performing procedural activities (the trainees) while guided by real-world experts who provide guidance and answer specific questions using natural language. Following a ``Wizard of OZ'' data collection paradigm, the expert enacts a wearable intelligent assistant, looking at the activities performed by the trainee exclusively from their egocentric point of view, answering questions when asked by the trainee, or proactively interacting with suggestions during the procedures. This unique data collection protocol enables Ego-EXTRA to capture a high-quality dialogue in which expert-level feedback is provided to the trainee. Two-way dialogues between experts and trainees are recorded, transcribed, and used to create a novel benchmark comprising more than 15k high-quality Visual Question Answer sets, which we use to evaluate Multimodal Large Language Models. The results show that Ego-EXTRA is challenging and highlight the limitations of current models when used to provide expert-level assistance to the user. The Ego-EXTRA dataset is publicly available to support the benchmark of egocentric video-language assistants: https://fpv-iplab.github.io/Ego-EXTRA/. Francesco Ragusa, Michele Mazzamuto, Rosario Forte, Irene D'Ambra, James Fort, Jakob J. Engel, Antonino Furnari, Giovanni Maria Farinella |
WACV | 8 |
| 2026 | TI-PREGO: Chain of Thought and In-Context Learning for online mistake detection in PRocedural EGOcentric videos
Leonardo Plini, Luca Scofano, Edoardo De Matteis, Guido Maria D'Amely di Melendugno, Alessandro Flaborea, Andrea Sanchietti, Giovanni Maria Farinella, Fabio Galasso, Antonino Furnari |
Comput. Vis. Image Underst. | 7 |
| 2026 | Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction DetectionabstractAbstract In this work, we explore the role of synthetic data in improving the detection of Hand-Object Interactions from egocentric images. Through extensive experimentation and comparative analysis on VISOR , EgoHOS , and ENIGMA-51 datasets, our findings demonstrate the potential of synthetic data to significantly improve HOI detection, particularly when real labeled data are scarce or unavailable. By using synthetic data and only $$10\%$$ 10 % of the real labeled data, we achieve improvements in Overall AP over models trained exclusively on real data, with gains of $$+5.67\%$$ + 5.67 % on VISOR , $$+8.24\%$$ + 8.24 % on EgoHOS , and $$+11.69\%$$ + 11.69 % on ENIGMA-51 . Furthermore, we systematically study how aligning synthetic data to specific real-world benchmarks with respect to objects, grasps, and environments, showing that the effectiveness of synthetic data consistently improves with better synthetic-real alignment. As a result of this work, we release a new data generation pipeline and the new HOI-Synth benchmark, which augments existing datasets with synthetic images of hand-object interaction. These data are automatically annotated with hand-object contact states, bounding boxes, and pixel-wise segmentation masks. All data, code, and tools for synthetic data generation are available at: https://fpv-iplab.github.io/HOI-Synth/ . Rosario Leonardi, Antonino Furnari, Francesco Ragusa, Giovanni Maria Farinella |
Int. J. Comput. Vis. | 4 |
| 2026 | Exocentric-to-Egocentric Adaptation for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs
Camillo Quattrocchi, Antonino Furnari, Daniele Di Mauro, Mario Valerio Giuffrida, Giovanni Maria Farinella |
Int. J. Comput. Vis. | 5 |
| 2026 | Integrating Affordances and Attention Models for Short-Term Object Interaction AnticipationabstractShort-Term object-interaction Anticipation (STA) consists in detecting the location of the next-active objects, the noun and verb categories of the interaction, as well as the time to contact from the observation of egocentric video. This ability is fundamental for wearable assistants to understand user's goals and provide timely assistance, or to enable human-robot interaction. In this work, we present a method to improve the performance of STA predictions. Our contributions are two-fold: 1) We propose STAformer and STAformer++, two novel attention-based architectures integrating frame-guided temporal pooling, dual image-video attention, and multiscale feature fusion to support STA predictions from an image-input video pair; 2) We introduce two novel modules to ground STA predictions on human behavior by modeling affordances. First, we integrate an environment affordance model which acts as a persistent memory of interactions that can take place in a given physical scene. We explore how to integrate environment affordances via simple late fusion and with an approach which adaptively learns how to best fuse affordances with end-to-end predictions. Second, we predict interaction hotspots from the observation of hands and object trajectories, increasing confidence in STA predictions localized around the hotspot. Our results show significant improvements on Overall Top-5 mAP, with gain up to $+23\%$+23% on Ego4D and $+31\%$+31% on a novel set of curated EPIC-Kitchens STA labels. We released the https://github.com/lmur98/AFFttention code, annotations, and pre-extracted affordances on Ego4D and EPIC-Kitchens to encourage future research in this area. Lorenzo Mur-Labadia, Ruben Martinez-Cantin, Josechu J. Guerrero, Giovanni Maria Farinella, Antonino Furnari |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Task Graph Maximum Likelihood Estimation for Procedural Activity Understanding in Egocentric VideosabstractHumans engage daily in procedural activities such as cooking a recipe or fixing a bike, which can be described as goal-oriented sequences of key-steps following certain ordering constraints. Task graphs mined from videos or textual descriptions have recently gained popularity as a human-readable, holistic representation of procedural activities encoding a partial ordering over key-steps, and have shown promise in supporting downstream video understanding tasks. While previous works generally relied on hand-crafted procedures to extract task graphs from videos, this paper introduces an approach based on gradient-based maximum likelihood optimization of edge weights, which can be used to directly estimate an adjacency matrix and can also be naturally plugged into more complex neural network architectures. We validate the ability of the proposed approach to generate accurate task graphs on the CaptainCook4D and EgoPER datasets. Moreover, we extend our validation analysis to the EgoProceL dataset, which we manually annotate with task graphs as an additional contribution. The three datasets together constitute a new benchmark for task graph learning, where our approach obtains improvements of +14.5%, +10.2% and +13.6% in$F_{1}$score, respectively, over previous approaches. Thanks to the differentiability of the proposed framework, we also introduce a feature-based approach for predicting task graphs from key-step textual or video embeddings, which exhibits emerging video understanding abilities. Beyond that, task graphs learned with our approach obtain top performance in the Ego-Exo4D procedure understanding benchmark including 5 different downstream tasks, with gains of up to +4.61%, +0.10%, +5.02%, +8.62%, and +15.16% in finding Previous Keysteps, Optional Keysteps, Procedural Mistakes, Missing Keysteps, and Future Keysteps, respectively. We finally show significant enhancements to the challenging task of online mistake detection in procedural egocentric videos, achieving notable gains of +19.8% and +6.4% in the Assembly101-O and EPIC-Tent-O datasets, respectively, compared to the state of the art. The code for replicating the experiments is available athttps://github.com/fpv-iplab/Differentiable-Task-Graph-Learning. Luigi Seminara, Giovanni Maria Farinella, Antonino Furnari |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human ActivitiesabstractWe address the challenge of unsupervised mistake detection in egocentric video of skilled human activities through the analysis of gaze signals. While traditional methods rely on manually labeled mistakes, our approach does not require mistake annotations, hence overcoming the need of domain-specific labeled data. Based on the observation that eye movements closely follow object manipulation activities, we assess to what extent eye-gaze signals can support mistake detection, proposing to identify deviations in attention patterns measured through a gaze tracker with respect to those estimated by a gaze prediction model. Since predicting gaze in video is characterized by high uncertainty, we propose a novel gaze completion task, where eye fixations are predicted from visual observations and partial gaze trajectories, and contribute a novel gaze completion approach which explicitly models correlations between gaze information and local visual tokens. Inconsistencies between predicted and observed gaze trajectories act as an indicator to identify mistakes. Experiments highlight the effectiveness of the proposed approach in different settings, with relative gains up to +14%, +11%, and +5% in EPIC-Tent, HoloAssist and IndustReal respectively, remarkably matching results of supervised approaches without seeing any labels. We further show that gaze-based analysis is particularly useful in the presence of skilled actions, low action execution confidence, and actions requiring hand-eye coordination and object manipulation skills. Our method is ranked first on the HoloAssist Mistake Detection challenge. Michele Mazzamuto, Antonino Furnari, Yoichi Sato 0001, Giovanni Maria Farinella |
CVPR | 4 |
| 2025 | CrowdSim++: Unifying Crowd Navigation and Obstacle Avoidance
Marco Rosano, Danilo Leocata, Antonino Furnari, Giovanni Maria Farinella |
ICAART (2) | 4 |
| 2025 | Egocentric action anticipation from untrimmed videosabstractAbstract Egocentric action anticipation involves predicting future actions performed by the camera wearer from egocentric video. Although the task has recently gained attention in the research community, current approaches often assume that input videos are ‘trimmed’, meaning that a short video sequence is sampled a fixed time before the beginning of the action. However, trimmed action anticipation has limited applicability in real‐world scenarios, where it is crucial to deal with ‘untrimmed’ video inputs and the exact moment of action initiation cannot be assumed at test time. To address these limitations, an untrimmed action anticipation task is proposed, which, akin to temporal action detection, assumes that the input video is untrimmed at test time, while still requiring predictions to be made before actions take place. The authors introduce a benchmark evaluation procedure for methods designed to address this novel task and compare several baselines on the EPIC‐KITCHENS‐100 dataset. Through our experimental evaluation, testing a variety of models, the authors aim to better understand their performance in untrimmed action anticipation. Our results reveal that the performance of current models designed for trimmed action anticipation is limited, emphasising the need for further research in this area. Ivan Rodin, Antonino Furnari, Giovanni Maria Farinella |
IET Comput. Vis. | 3 |
| 2025 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions—including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. https://ego-exo4d-data.org/ Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, David Crandall, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
Int. J. Comput. Vis. | 89 |
| 2025 | Ego4D: Around the World in 3,600 Hours of Egocentric VideoabstractWe introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Devansh Kukreja, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik |
IEEE Trans. Pattern Anal. Mach. Intell. | 68 |
| 2024 | PREGO: Online Mistake Detection in PRocedural EGOcentric VideosabstractPromptly identifying procedural errors from egocentric videos in an online setting is highly challenging and valuable for detecting mistakes as soon as they happen. This capability has a wide range of applications across various fields, such as manufacturing and healthcare. The nature of procedural mistakes is open-set since novel types of failures might occur, which calls for one-class classifiers trained on correctly executed procedures. However, no technique can currently detect open-set procedural mistakes online. We propose PREGO, the first online one-class classification model for mistake detection in PRocedural EGOcentric videos. PREGO is based on an online action recognition component to model the current action, and a symbolic reasoning module to predict the next actions. Mistake detection is performed by comparing the recognized current action with the expected future one. We evaluate PREGO on two procedural egocentric video datasets, Assembly101 and Epic-tent, which we adapt for online benchmarking of procedural mistake detection to establish suitable benchmarks, thus defining the Assembly101-O and Epic-tent-O datasets, respectively. The code is available at https://github.com/alef/abo/PREGO. Alessandro Flaborea, Guido Maria D'Amely di Melendugno, Leonardo Plini, Luca Scofano, Edoardo De Matteis, Antonino Furnari, Giovanni Maria Farinella, Fabio Galasso |
CVPR | 7 |
| 2024 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multi-modal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured ego-centric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is un-precedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions-including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
CVPR | 88 |
| 2024 | Action Scene Graphs for Long-Form Understanding of Egocentric VideosabstractWe present Egocentric Action Scene Graphs (EASGs), a new representation for long-form understanding of egocentric videos. EASGs extend standard manually-annotated representations of egocentric videos, such as verb-noun action labels, by providing a temporally evolving graph-based description of the actions performed by the camera wearer, including interacted objects, their relationships, and how actions unfold in time. Through a novel annotation procedure, we extend the Ego4D dataset adding manually labeled Egocentric Action Scene Graphs which offer a rich set of annotations for long-from egocentric video understanding. We hence define the EASG generation task and provide a baseline approach, establishing preliminary benchmarks. Experiments on two downstream tasks, action anticipation and activity summarization, highlight the effectiveness of EASGs for long-form egocentric video understanding. We will release the dataset and code to replicate experiments and annotations11The code is available at https://github.com/fpv-iplab/EASG. Ivan Rodin, Antonino Furnari, Kyle Min 0001, Subarna Tripathi, Giovanni Maria Farinella |
CVPR | 5 |
| 2024 | Dynamic Price Prediction for Revenue Management System in Hospitality Sector
Susanna Saitta, Vito D'Amico, Giovanni Maria Farinella |
DATA | 3 |
| 2024 | Are Synthetic Data Useful for Egocentric Hand-Object Interaction Detection?
Rosario Leonardi, Antonino Furnari, Francesco Ragusa, Giovanni Maria Farinella |
ECCV (71) | 4 |
| 2024 | AFF-ttention! Affordances and Attention Models for Short-Term Object Interaction Anticipation
Lorenzo Mur-Labadia, Ruben Martinez-Cantin, Josechu J. Guerrero, Giovanni Maria Farinella, Antonino Furnari |
ECCV (23) | 4 |
| 2024 | Synchronization Is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs
Camillo Quattrocchi, Antonino Furnari, Daniele Di Mauro, Mario Valerio Giuffrida, Giovanni Maria Farinella |
ECCV (72) | 5 |
| 2024 | The Analog Layer: Simulating Imperfect Computations in Neural Networks to Improve Robustness and Generalization Ability
Giovanni Maria Manduca, Antonino Furnari, Giovanni Maria Farinella |
ICPR (26) | 3 |
| 2024 | Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosabstractProcedural activities are sequences of key-steps aimed at achieving specific goals. They are crucial to build intelligent agents able to assist users effectively. In this context, task graphs have emerged as a human-understandable representation of procedural activities, encoding a partial ordering over the key-steps. While previous works generally relied on hand-crafted procedures to extract task graphs from videos, in this paper, we propose an approach based on direct maximum likelihood optimization of edges' weights, which allows gradient-based learning of task graphs and can be naturally plugged into neural network architectures. Experiments on the CaptainCook4D dataset demonstrate the ability of our approach to predict accurate task graphs from the observation of action sequences, with an improvement of +16.7% over previous approaches. Owing to the differentiability of the proposed framework, we also introduce a feature-based approach, aiming to predict task graphs from key-step textual or video embeddings, for which we observe emerging video understanding abilities. Task graphs learned with our approach are also shown to significantly enhance online mistake detection in procedural egocentric videos, achieving notable gains of +19.8% and +7.5% on the Assembly101-O and EPIC-Tent-O datasets. Code for replicating the experiments is available at https://github.com/fpv-iplab/Differentiable-Task-Graph-Learning. Luigi Seminara, Giovanni Maria Farinella, Antonino Furnari |
NeurIPS | 2 |
| 2024 | ENIGMA-51: Towards a Fine-Grained Understanding of Human Behavior in Industrial ScenariosabstractENIGMA-51 is a new egocentric dataset acquired in an industrial scenario by 19 subjects who followed instructions to complete the repair of electrical boards using industrial tools (e.g., electric screwdriver) and equipments (e.g., oscilloscope). The 51 egocentric video sequences are densely annotated with a rich set of labels that enable the systematic study of human behavior in the industrial domain. We provide benchmarks on four tasks related to human behavior: 1) untrimmed temporal detection of human-object interactions, 2) egocentric human-object interaction detection, 3) short-term object interaction anticipation and 4) natural language understanding of intents and entities. Baseline results show that the ENIGMA-51 dataset poses a challenging benchmark to study human behavior in industrial scenarios. We publicly release the dataset at https://iplab.dmi.unict.it/ENIGMA-51. Francesco Ragusa, Rosario Leonardi, Michele Mazzamuto, Claudia Bonanno, Rosario Scavo, Antonino Furnari, Giovanni Maria Farinella |
WACV | 7 |
| 2024 | Exploiting multimodal synthetic data for egocentric human-object interaction detection in an industrial scenarioabstractIn this paper, we tackle the problem of Egocentric Human-Object Interaction (EHOI) detection in an industrial setting. To overcome the lack of public datasets in this context, we propose a pipeline and a tool for generating synthetic images of EHOIs paired with several annotations and data signals (e.g., depth maps or segmentation masks). Using the proposed pipeline, we present EgoISM-HOI a new multimodal dataset composed of synthetic EHOI images in an industrial environment with rich annotations of hands and objects. To demonstrate the utility and effectiveness of synthetic EHOI data produced by the proposed tool, we designed a new method that predicts and combines different multimodal signals to detect EHOIs in RGB images. Our study shows that exploiting synthetic data to pre-train the proposed method significantly improves performance when tested on real-world data. Moreover, to fully understand the usefulness of our method, we conducted an in-depth analysis in which we compared and highlighted the superiority of the proposed approach over different state-of-the-art class-agnostic methods. To support research in this field, we publicly release the datasets, source code, and pre-trained models at https://iplab.dmi.unict.it/egoism-hoi. Rosario Leonardi, Francesco Ragusa, Antonino Furnari, Giovanni Maria Farinella |
Comput. Vis. Image Underst. | 4 |
| 2024 | An Outlook into the Future of Egocentric VisionabstractAbstract What will the future be? We wonder! In this survey, we explore the gap between current research in egocentric vision and the ever-anticipated future, where wearable computing, with outward facing cameras and digital overlays, is expected to be integrated in our every day lives. To understand this gap, the article starts by envisaging the future through character-based stories, showcasing through examples the limitations of current technology. We then provide a mapping between this future and previously defined research tasks. For each task, we survey its seminal works, current state-of-the-art methodologies and available datasets, then reflect on shortcomings that limit its applicability to future research. Note that this survey focuses on software models for egocentric vision, independent of any specific hardware. The paper concludes with recommendations for areas of immediate explorations so as to unlock our path to the future always-on, personalised and life-enhancing egocentric vision. Chiara Plizzari, Gabriele Goletto, Antonino Furnari, Siddhant Bansal, Francesco Ragusa, Giovanni Maria Farinella, Dima Damen, Tatiana Tommasi |
Int. J. Comput. Vis. | 6 |
| 2023 | Egocentric Action Anticipation for Personal HealthabstractThe egocentric action anticipation task consists in predicting future (unobserved) actions based on input from a wearable first-person view camera. In this work, we are focusing on applying egocentric action anticipation methods to the personal health domain, i.e., utilizing them for the analysis of dietary and hygienic activities routine and for the prevention of undesirable behavior (such as tasting food before washing hands or adding sugar). We collect a dataset of egocentric videos, capturing specific activities related to food preparation; fine-tune existing action anticipation models on this dataset and analyse effects caused by domain shift. Ivan Rodin, Antonino Furnari, Dimitrios Mavroeidis, Giovanni Maria Farinella |
ICASSP | 4 |
| 2023 | Streaming egocentric action anticipation: An evaluation scheme and approachabstractEgocentric action anticipation aims to predict the future actions the camera wearer will perform from the observation of the past. While predictions about the future should be available before the predicted events take place, most approaches do not pay attention to the computational time required to make such predictions. As a result, current evaluation schemes assume that predictions are available right after the input video is observed, i.e., presuming a negligible runtime, which may lead to overly optimistic evaluations. We propose a streaming egocentric action evaluation scheme which assumes that predictions are performed online and made available only after the model has processed the current input segment, which depends on its runtime. To evaluate all models considering the same prediction horizon, we hence propose that slower models should base their predictions on temporal segments sampled ahead of time. Based on the observation that model runtime can affect performance in the considered streaming evaluation scenario, we further propose a lightweight action anticipation model based on feed-forward 3D CNNs which is optimized using knowledge distillation techniques with a novel past-to-future distillation loss. Experiments on the three popular datasets EPIC-KITCHENS-55, EPIC-KITCHENS-100 and EGTEA Gaze+ show that (i) the proposed evaluation scheme induces a different ranking on state-of-the-art methods as compared to classic evaluations, (ii) lightweight approaches tend to outmatch more computationally expensive ones, and (iii) the proposed model based on feed-forward 3D CNNs and knowledge distillation outperforms current art in the streaming egocentric action anticipation scenario. Antonino Furnari, Giovanni Maria Farinella |
Comput. Vis. Image Underst. | 2 |
| 2023 | MECCANO: A multimodal egocentric dataset for humans behavior understanding in the industrial-like domainabstractWearable cameras allow to acquire images and videos from the user’s perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still understudied in egocentric settings and in particular in industrial scenarios. To encourage research in this field, we present MECCANO, a multimodal dataset of egocentric videos to study humans behavior understanding in industrial-like settings. The multimodality is characterized by the presence of gaze signals, depth maps and RGB videos acquired simultaneously with a custom headset. The dataset has been explicitly labeled for fundamental tasks in the context of human behavior understanding from a first person view, such as recognizing and anticipating human-object interactions. With the MECCANO dataset, we explored six different tasks including (1) Action Recognition, (2) Active Objects Detection and Recognition, (3) Egocentric Human-Objects Interaction Detection, (4) Egocentric Gaze Estimation, (5) Action Anticipation and (6) Next-Active Objects Detection. We propose a benchmark aimed to study human behavior in the considered industrial-like scenario which demonstrates that the investigated tasks and the considered scenario are challenging for state-of-the-art algorithms. To support research in this field, we publicy release the dataset at https://iplab.dmi.unict.it/MECCANO/. Francesco Ragusa, Antonino Furnari, Giovanni Maria Farinella |
Comput. Vis. Image Underst. | 3 |
| 2023 | Visual Object Tracking in First Person VisionabstractThe understanding of human-object interactions is fundamental in First Person Vision (FPV). Visual tracking algorithms which follow the objects manipulated by the camera wearer can provide useful information to effectively model such interactions. In the last years, the computer vision community has significantly improved the performance of tracking algorithms for a large variety of target objects and scenarios. Despite a few previous attempts to exploit trackers in the FPV domain, a methodical analysis of the performance of state-of-the-art trackers is still missing. This research gap raises the question of whether current solutions can be used "off-the-shelf" or more domain-specific investigations should be carried out. This paper aims to provide answers to such questions. We present the first systematic investigation of single object tracking in FPV. Our study extensively analyses the performance of 42 algorithms including generic object trackers and baseline FPV-specific trackers. The analysis is carried out by focusing on different aspects of the FPV setting, introducing new performance measures, and in relation to FPV-specific tasks. The study is made possible through the introduction of TREK-150, a novel benchmark dataset composed of 150 densely annotated video sequences. Our results show that object tracking in FPV poses new challenges to current visual trackers. We highlight the factors causing such behavior and point out possible research directions. Despite their difficulties, we prove that trackers bring benefits to FPV downstream tasks requiring short-term object tracking. We expect that generic object tracking will gain popularity in FPV as new and FPV-specific methodologies are investigated. Supplementary Information: The online version contains supplementary material available at 10.1007/s11263-022-01694-6. Matteo Dunnhofer, Antonino Furnari, Giovanni Maria Farinella, Christian Micheloni |
Int. J. Comput. Vis. | 3 |
| 2023 | Editorial: Special Section on Egocentric PerceptionabstractThe papers in this special issue focus on egocentric perception. It gathers recent advances in this field that brings together multiple communities including computer vision, machine learning, and multimedia. Antonino Furnari, David Crandall, Dima Damen, Kristen Grauman, Giovanni Maria Farinella |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Ego4D: Around the World in 3, 000 Hours of Egocentric VideoabstractWe introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/ Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik |
CVPR | 67 |
| 2022 | Towards Streaming Egocentric Action AnticipationabstractEgocentric action anticipation consists in predicting the future actions a camera wearer will likely perform based on past video observations. While in a real-world system it is fundamental to produce such predictions before the action begins, past works have not generally paid attention to model runtime during evaluation. Indeed, current evaluation schemes assume that predictions can be made offline, and hence that computational resources are not limited. In this paper, we propose a "streaming" evaluation protocol which explicitly considers model runtime for performance assessment, assuming that predictions will be available only after the current video segment is processed, which depends on the processing time of a method. Following the proposed evaluation scheme, we benchmark different state-of-the-art approaches for egocentric action anticipation on two popular datasets. Our analysis shows that models with a smaller runtime tend to outperform heavier models in the considered streaming scenario, thus changing the rankings observed in standard offline evaluations. Based on this observation, we propose a lightweight action anticipation model consisting in a simple feed-forward 3D CNN, which we propose to optimize using knowledge distillation techniques and a custom loss. The results show that the proposed approach outperforms prior art in the streaming scenario, also in combination with other lightweight models. Antonino Furnari, Giovanni Maria Farinella |
ICPR | 2 |
| 2022 | MADiMa'22: 7th International Workshop on Multimedia Assisted Dietary ManagementabstractThis abstract provides a summary and overview of the 7th International Workshop on Multimedia Assisted Dietary Management. Stavroula G. Mougiakakou, Giovanni Maria Farinella, Keiji Yanai, Dario Allegra |
ACM Multimedia | 2 |
| 2022 | A multi camera unsupervised domain adaptation pipeline for object detection in cultural sites through adversarial learning and self-training
Giovanni Pasqualino, Antonino Furnari, Giovanni Maria Farinella |
Comput. Vis. Image Underst. | 3 |
| 2022 | Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100abstractAbstract This paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS. The effort culminates in EPIC-KITCHENS-100, a collection of 100 hours, 20M frames, 90K actions in 700 variable-length videos, capturing long-term unscripted activities in 45 environments, using head-mounted cameras. Compared to its previous version (Damen in Scaling egocentric vision: ECCV, 2018), EPIC-KITCHENS-100 has been annotated using a novel pipeline that allows denser (54% more actions per minute) and more complete annotations of fine-grained actions (+128% more action segments). This collection enables new challenges such as action detection and evaluating the “test of time”—i.e. whether models trained on data collected in 2018 can generalise to new footage collected two years later. The dataset is aligned with 6 challenges: action recognition (full and weak supervision), action detection, action anticipation, cross-modal retrieval (from captions), as well as unsupervised domain adaptation for action recognition. For each challenge, we define the task, provide baselines and evaluation metrics. Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, Michael Wray |
Int. J. Comput. Vis. | 3 |
| 2021 | The MECCANO Dataset: Understanding Human-Object Interactions from Egocentric Videos in an Industrial-like DomainabstractWearable cameras allow to collect images and videos of humans interacting with the world. While human-object interactions have been thoroughly investigated in third person vision, the problem has been understudied in egocentric settings and in industrial scenarios. To fill this gap, we introduce MECCANO, the first dataset of egocentric videos to study human-object interactions in industrial-like settings. MECCANO has been acquired by 20 participants who were asked to build a motorbike model, for which they had to interact with tiny objects and tools. The dataset has been explicitly labeled for the task of recognizing human-object interactions from an egocentric perspective. Specifically, each interaction has been labeled both temporally (with action segments) and spatially (with active object bounding boxes). With the proposed dataset, we investigate four different tasks including 1) action recognition, 2) active object detection, 3) active object recognition and 4) egocentric human-object interaction detection, which is a revisited version of the standard human-object interaction detection task. Baseline results show that the MECCANO dataset is a challenging benchmark to study egocentric human-object interactions in industrial-like scenarios. We publicy release the dataset at https://iplab.dmi.unict.it/MECCANO/. Francesco Ragusa, Antonino Furnari, Salvatore Livatino, Giovanni Maria Farinella |
WACV | 4 |
| 2021 | Predicting the future from first person (egocentric) vision: A survey
Ivan Rodin, Antonino Furnari, Dimitrios Mavroeidis, Giovanni Maria Farinella |
Comput. Vis. Image Underst. | 4 |
| 2021 | Editorial: Computer Vision Theory and Applications at VISAPP 2020
Petia Radeva, Giovanni Maria Farinella |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | An unsupervised domain adaptation scheme for single-stage artwork recognition in cultural sites
Giovanni Pasqualino, Antonino Furnari, Giovanni Signorello, Giovanni Maria Farinella |
Image Vis. Comput. | 4 |
| 2021 | Exploiting objective text description of images for visual sentiment analysis
Alessandro Ortis, Giovanni Maria Farinella, Giovanni Torrisi, Sebastiano Battiato |
Multim. Tools Appl. | 2 |
| 2021 | The EPIC-KITCHENS Dataset: Collection, Challenges and BaselinesabstractSince its introduction in 2018, EPIC-KITCHENS has attracted attention as the largest egocentric video benchmark, offering a unique viewpoint on people's interaction with objects, their attention, and even intention. In this paper, we detail how this large-scale dataset was captured by 32 participants in their native kitchen environments, and densely annotated with actions and object interactions. Our videos depict nonscripted daily activities, as recording is started every time a participant entered their kitchen. Recording took place in four countries by participants belonging to ten different nationalities, resulting in highly diverse kitchen habits and cooking styles. Our dataset features 55 hours of video consisting of 11.5M frames, which we densely labelled for a total of 39.6K action segments and 454.2K object bounding boxes. Our annotation is unique in that we had the participants narrate their own videos (after recording), thus reflecting true intention, and we crowd-sourced ground-truths based on these. We describe our object, action and anticipation challenges, and evaluate several baselines over two test splits, seen and unseen kitchens. We introduce new baselines that highlight the multimodal nature of the dataset and the importance of explicit temporal modelling to discriminate fine-grained actions (e.g., 'closing a tap' from 'opening' it up). Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, Michael Wray |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Rolling-Unrolling LSTMs for Action Anticipation from First-Person VideoabstractIn this paper, we tackle the problem of egocentric action anticipation, i.e., predicting what actions the camera wearer will perform in the near future and which objects they will interact with. Specifically, we contribute Rolling-Unrolling LSTM, a learning architecture to anticipate actions from egocentric videos. The method is based on three components: 1) an architecture comprised of two LSTMs to model the sub-tasks of summarizing the past and inferring the future, 2) a Sequence Completion Pre-Training technique which encourages the LSTMs to focus on the different sub-tasks, and 3) a Modality ATTention (MATT) mechanism to efficiently fuse multi-modal predictions performed by processing RGB frames, optical flow fields and object-based features. The proposed approach is validated on EPIC-Kitchens, EGTEA Gaze+ and ActivityNet. The experiments show that the proposed architecture is state-of-the-art in the domain of egocentric videos, achieving top performances in the 2019 EPIC-Kitchens egocentric action anticipation challenge. The approach also achieves competitive performance on ActivityNet with respect to methods not based on unsupervised pre-training and generalizes to the tasks of early action recognition and action recognition. To encourage research on this challenging topic, we made our code, trained models, and pre-extracted features available at our web page: http://iplab.dmi.unict.it/rulstm. Antonino Furnari, Giovanni Maria Farinella |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | EgoCart: A Benchmark Dataset for Large-Scale Indoor Image-Based Localization in Retail StoresabstractWe consider the task of localizing shopping carts in a retail store from egocentric images. Addressing this task allows to infer information on the behavior of the customers to understand how they move in the store and what they pay more attention to. To study the problem, we propose a large dataset of images collected in a real retail store. The dataset comprises 19, 531 RGB images along with depth maps, ground truth camera poses, as well as class labels specifying the areas of the store in which each image has been acquired. We release the dataset to the public to encourage research in large-scale image-based indoor localization and to address the scarcity of large datasets to tackle the problem. We hence perform a benchmark of several image-based localization techniques exploiting images and depth information on the proposed dataset. In our study, both localization performances and space/time requirements are compared. The results show that, while state-of-the-art approaches allow to achieve good results, there is space for improvement. Emiliano Spera, Antonino Furnari, Sebastiano Battiato, Giovanni Maria Farinella |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Knowledge Distillation for Action Anticipation via Label SmoothingabstractHuman capability to anticipate near future from visual observations and non-verbal cues is essential for developing intelligent systems that need to interact with people. Several research areas, such as human-robot interaction (HRI), assisted living or autonomous driving need to foresee future events to avoid crashes or help people. Egocentric scenarios are classic examples where action anticipation is applied due to their numerous applications. Such challenging task demands to capture and model domain's hidden structure to reduce prediction uncertainty. Since multiple actions may equally occur in the future, we treat action anticipation as a multi-label problem with missing labels extending the concept of label smoothing. This idea resembles the knowledge distillation process since useful information is injected into the model during training. We implement a multi-modal framework based on long short-term memory (LSTM) networks to summarize past observations and make predictions at different time steps. We perform extensive experiments on EPIC-Kitchens and EGTEA Gaze+ datasets including more than 2500 and 100 action classes, respectively. The experiments show that label smoothing systematically improves performance of state-of-the-art models for action anticipation. Guglielmo Camporese, Pasquale Coscia, Antonino Furnari, Giovanni Maria Farinella, Lamberto Ballan |
ICPR | 4 |
| 2020 | Unsupervised Domain Adaptation for Object Detection in Cultural SitesabstractThe ability to detect objects in cultural sites from the egocentric point of view of the user can enable interesting applications for both the visitors and the manager of the site. Unfortunately, current object detection algorithms have to be trained on large amounts of labeled data, the collection of which is costly and time-consuming. While synthetic data generated from the 3D model of the cultural site can be used to train object detection algorithms, a significant drop in performance is generally observed when such algorithms are deployed to work with real images. In this paper, we consider the problem of unsupervised domain adaptation for object detection in cultural sites. Specifically, we assume the availability of synthetic labeled images and real unlabeled images for training. To study the problem, we propose a dataset containing 75244 synthetic and 2190 real images with annotations for 16 different artworks. We hence investigate different domain adaptation techniques based on image-to-image translation and feature alignment. Our analysis points out that such techniques can be useful to address the domain adaptation issue, while there is still plenty of space for improvement on the proposed dataset. We release the dataset at our web page to encourage research on this challenging topic: https://iplab.dmi.unict.it/EGO-CH-OBJ-ADAPT/. Giovanni Pasqualino, Antonino Furnari, Giovanni Maria Farinella |
ICPR | 3 |
| 2020 | Semantic Object Segmentation in Cultural Sites using Real and Synthetic DataabstractWe consider the problem of object segmentation in cultural sites. Since collecting and labeling large datasets of real images is challenging, we investigate whether the use of synthetic images can be useful to achieve good segmentation performance on real data. To perform the study, we collected a new dataset comprising both real and synthetic images of 24 artworks in a cultural site. The synthetic images have been automatically generated from the 3D model of the considered cultural site using a tool developed for that purpose. Real and synthetic images have been labeled for the task of semantic segmentation of artworks. We compare three different approaches to perform object segmentation exploiting real and synthetic data. The experimental results point out that the use of synthetic data helps to improve the performances of segmentation algorithms when tested on real images. Satisfactory performance is achieved exploiting semantic segmentation together with image-to-image translation and including a small amount of real data during training. To encourage research on the topic, we publicly release the proposed dataset at the following url:https://iplab.dmi.unict.it/EGO-CH-OBJ-SEG/. Francesco Ragusa, Daniele Di Mauro, Alfio Palermo, Antonino Furnari, Giovanni Maria Farinella |
ICPR | 5 |
| 2020 | On Embodied Visual Navigation in Real Environments Through HabitatabstractVisual navigation models based on deep learning can learn effective policies when trained on large amounts of visual observations through reinforcement learning. Unfortunately, collecting the required experience in the real world requires the deployment of a robotic platform, which is expensive and time-consuming. To deal with this limitation, several simulation platforms have been proposed in order to train visual navigation policies on virtual environments efficiently. Despite the advantages they offer, simulators present a limited realism in terms of appearance and physical dynamics, leading to navigation policies that do not generalize in the real world. In this paper, we propose a tool based on the Habitat simulator which exploits real world images of the environment, together with sensor and actuator noise models, to produce more realistic navigation episodes. We perform a range of experiments to assess the ability of such policies to generalize using virtual and real-world images, as well as observations transformed with unsupervised domain adaptation approaches. We also assess the impact of sensor and actuation noise on the navigation performance and investigate whether it allows to learn more robust navigation policies. We show that our tool can effectively help to train and evaluate navigation policies on real-world observations without running navigation episodes in the real world. Marco Rosano, Antonino Furnari, Luigi Gulino, Giovanni Maria Farinella |
ICPR | 4 |
| 2020 | Anticipating Activity from Multimodal SignalsabstractImages, videos, audio signals, sensor data, can be easily collected in huge quantity by different devices and processed in order to emulate the human capability of elaborating a variety of different stimuli. Are multimodal signals useful to understand and anticipate human actions if acquired from the user viewpoint? This paper proposes to build an embedding space where inputs of different nature, but semantically correlated, are projected in a new representation space and properly exploited to anticipate the future user activity. To this purpose, we built a new multimodal dataset comprising video, audio, tri-axial acceleration, angular velocity, tri-axial magnetic field, pressure and temperature. To benchmark the proposed multimodal anticipation challenge, we consider classic classifiers on top of deep learning methods used to build the embedding space representing multimodal signals. The achieved results show that the exploitation of different modalities is useful to improve the anticipation of the future activity. Tiziana Rotondo, Giovanni Maria Farinella, Davide Giacalone, Sebastiano Mauro Strano, Valeria Tomaselli, Sebastiano Battiato |
ICPR | 2 |
| 2020 | Virtual to Real Unsupervised Domain Adaptation for Image-Based Localization in Cultural SitesabstractThe ability to localize the visitors of a cultural site from egocentric images can allow applications to understand where people go and what they pay attention to in the site. Current pipelines to tackle the problem require the collection and labeling of large amounts of images, which is challenging, especially in large-scale indoor environments. On the contrary, virtual images of a cultural site can be generated and automatically labeled using dedicated tools with minimum effort. In this paper, we investigate whether unsupervised domain adaptation techniques can be used to train localization models on labeled virtual data and unlabeled real data, and deploy them to work with real images. To perform this study, we propose a new dataset of both real and virtual images acquired in a cultural site which are labeled for room-based localization as well as for 3 DOF camera pose estimation. We hence compare two approaches to unsupervised domain adaptation: mid-level representations and image-to-image translation. Our analysis shows that both approaches can be used to reduce the domain gap arising from the different data sources and that the proposed dataset is a challenging benchmark for unsupervised domain adaptation for image-based localization. Santi A. Orlando, Antonino Furnari, Giovanni Maria Farinella |
IPAS | 3 |
| 2020 | On the Exploitation of Temporal Redundancy to Improve Polyp Detection in ColonoscopyabstractColonoscopy is currently the most effective screening method to find precancerous colon polyps and plan their removal. Computer-aided polyp detection can reduce polyp miss detection rates and help doctors find the most critical regions to pay attention to. The challenge in detecting polyps is due to the polyp's morphology and size, and these fall into false-negative. Indeed, polyps may exhibit high variability in shapes (e.g., depressed, flat, pedunculated, etc ...). Moreover, the water injected from the endoscope results in artifacts which impede the detection, and the lubricating mucus causes light artifacts due its glossiness. To address this problem, we propose a mask-based attention mechanism to ensure that the employed detector focuses on particular regions of the image in order to reduce misdetection rate. Our contribution takes advantage of information on polyp's position over time within a video sequence. We provide such information through a binary mask which points out the last-known polyp's position. The proposed approach is validated on a dataset that has been labeled by colonoscopy experts. It contains about 200 videos and more than 500 different polyps with high variability in size and textures. Experimental results show that the proposed attention mechanism recover a smaller number of false negatives and achieves an Fl-score of 80.21%. Giovanna Pappalardo, Dario Allegra, Filippo Stanco, Giovanni Maria Farinella |
IPAS | 4 |
| 2020 | Survey on visual sentiment analysisabstractVisual Sentiment Analysis aims to understand how images affect people, in terms of evoked emotions. Although this field is rather new, a broad range of techniques have been developed for various data sources and problems, resulting in a large body of research. This paper reviews pertinent publications and tries to present an exhaustive overview of the field. After a description of the task and the related applications, the subject is tackled under different main headings. The paper also describes principles of design of general Visual Sentiment Analysis systems from three main points of view: emotional models, dataset definition, feature design. A formalization of the problem is discussed, considering different levels of granularity, as well as the components that can affect the sentiment toward an image in different ways. To this aim, this paper considers a structured formalization of the problem which is usually used for the analysis of text, and discusses it's suitability in the context of Visual Sentiment Analysis. The paper also includes a description of new challenges, the evaluation from the viewpoint of progress toward more sophisticated systems and related practical applications, as well as a summary of the insights resulting from this study. Alessandro Ortis, Giovanni Maria Farinella, Sebastiano Battiato |
IET Image Process. | 2 |
| 2020 | Learning and recognition for assistive computer vision
Giovanni Maria Farinella, Marco Leo, Gérard G. Medioni, Mohan M. Trivedi |
Pattern Recognit. Lett. | 1 |
| 2020 | SceneAdapt: Scene-based domain adaptation for semantic segmentation using adversarial learning
Daniele Di Mauro, Antonino Furnari, Giuseppe Patanè 0002, Sebastiano Battiato, Giovanni Maria Farinella |
Pattern Recognit. Lett. | 5 |
| 2020 | Egocentric visitor localization and artwork detection in cultural sites using synthetic data
Santi A. Orlando, Antonino Furnari, Giovanni Maria Farinella |
Pattern Recognit. Lett. | 3 |
| 2020 | EGO-CH: Dataset and fundamental tasks for visitors behavioral understanding using egocentric vision
Francesco Ragusa, Antonino Furnari, Sebastiano Battiato, Giovanni Signorello, Giovanni Maria Farinella |
Pattern Recognit. Lett. | 5 |
| 2019 | What Would You Expect? Anticipating Egocentric Actions With Rolling-Unrolling LSTMs and Modality AttentionabstractEgocentric action anticipation consists in understanding which objects the camera wearer will interact with in the near future and which actions they will perform. We tackle the problem proposing an architecture able to anticipate actions at multiple temporal scales using two LSTMs to 1) summarize the past, and 2) formulate predictions about the future. The input video is processed considering three complimentary modalities: appearance (RGB), motion (optical flow) and objects (object-based features). Modality-specific predictions are fused using a novel Modality ATTention (MATT) mechanism which learns to weigh modalities in an adaptive fashion. Extensive evaluations on two large-scale benchmark datasets show that our method outperforms prior art by up to +7% on the challenging EPIC-Kitchens dataset including more than 2500 actions, and generalizes to EGTEA Gaze+. Our approach is also shown to generalize to the tasks of early action recognition and action recognition. Our method is ranked first in the public leaderboard of the EPIC-Kitchens egocentric action anticipation challenge 2019. Please see the project web page for code and additional details: http://iplab.dmi.unict.it/rulstm. Antonino Furnari, Giovanni Maria Farinella |
ICCV | 2 |
| 2019 | Egocentric Action Anticipation by Disentangling Encoding and InferenceabstractEgocentric action anticipation consists in predicting future actions from videos collected by means of a wearable camera. Action anticipation methods should be able to continuously 1) summarize the past and 2) predict possible future actions. We observe that action anticipation benefits from explicitly disentangling the two tasks. To this aim, we introduce a learning architecture which makes use of a "rolling" LSTM to continuously summarize the past and an "unrolling" LSTM to anticipate future actions at multiple temporal scales. The model includes a spatial and a temporal branch which process RGB images and optical flow fields independently. The predictions performed by the two branches are fused using a novel modality attention mechanism which leverages the complementary nature of the modalities. Experiments on the EPIC-KITCHENS dataset show that the proposed method surpasses the state-of-the-art by +4.02% and +6.39% when considering Top-1 and Top-5 accuracy respectively. Please see the project webpage at http://iplab.dmi.unict.it/rulstm/. Antonino Furnari, Giovanni Maria Farinella |
ICIP | 2 |
| 2019 | Siamese Ballistics Neural NetworkabstractFirearm identification is crucial in many investigative scenario. The crime scene often contains traces left by firearms in terms of bullets and cartridges. Traces analysis is a fundamental step in the Forensics Ballistics Analysis Process to identify which firearm fired a specific cartridge. In this paper we present a fully automated technique to compare cartridges represented as a set of 3D point-clouds. The overall approach is based on Siamese Neural Network learning paradigm that we use to build a suitable embedding space where the 3D point-cloud of the cartridges are compared. The proposed approach has been assessed by considering the NBTRD dataset. Obtained results support the exploitation of the proposed technique in ballistic analysis. Oliver Giudice, Luca Guarnera, Antonino Barbaro Paratore, Giovanni Maria Farinella, Sebastiano Battiato |
ICIP | 4 |
| 2019 | MADiMA'19: 5th International Workshop on Multimedia Assisted Dietary ManagementabstractThis abstract provides a summary and overview of the 5th International Workshop on Multimedia Assisted Dietary Management. Stavroula G. Mougiakakou, Giovanni Maria Farinella, Keiji Yanai, Dario Allegra |
ACM Multimedia | 2 |
| 2019 | Estimating the occupancy status of parking areas by counting cars and non-empty stalls
Daniele Di Mauro, Antonino Furnari, Giuseppe Patanè 0002, Sebastiano Battiato, Giovanni Maria Farinella |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | Egocentric visitors localization in natural sites
Filippo L. M. Milotta, Antonino Furnari, Sebastiano Battiato, Giovanni Signorello, Giovanni Maria Farinella |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | A context-driven privacy enforcement system for autonomous media capture devices
Giovanni Maria Farinella, Christian Napoli 0001, Gabriele Nicotra, Salvatore Riccobene |
Multim. Tools Appl. | 1 |
| 2018 | Scene Adaptation for Semantic Segmentation using Adversarial LearningabstractSemantic Segmentation algorithms based on the deep learning paradigm have reached outstanding performances. However, in order to achieve good results in a new domain, it is generally demanded to fine-tune a pre-trained deep architecture using new labeled data coming from the target application domain. The fine-tuning procedure is also required when the domain application settings change, e. g., when a camera is moved, or a new camera is installed. This implies the collection and pixel-wise la-beling of images to be used for training, which slows down the deployment of semantic segmentation systems in real industrial scenarios and increases the industrial costs. Taking into account the aforementioned issues, in this paper we propose an approach based on Adversarial Learning to perform scene adaptation for semantic segmentation. We frame scene adaptation as the task of predicting semantic segmentation masks for images belonging to a Target Scene Context given labeled images coming from a Source Scene Context and unlabeled images coming from the Target Scene Context. Experiments highlight that the proposed method achieves promising performances both when the two scenes contain similar content (i.e., they are related to two different points of view of the same scene) and when the observed scenes contain unrelated content (i.e., they account to completely different scenes). Daniele Di Mauro, Antonino Furnari, Giuseppe Patanè 0002, Sebastiano Battiato, Giovanni Maria Farinella |
AVSS | 5 |
| 2018 | Visual Sentiment Analysis Based on on Objective Text Description of ImagesabstractVisual Sentiment Analysis aims to estimate the polarity of the sentiment evoked by images in terms of positive or negative sentiment. To this aim, most of the state of the art works exploit the text associated to a social post provided by the user. However, such textual data is typically noisy due to the subjectivity of the user which usually includes text useful to maximize the diffusion of the social post. In this paper we extract and employ an Objective Text description of images automatically extracted from the visual content rather than the classic Subjective Text provided by the users. The proposed method defines a multimodal embedding space based on the contribute of both visual and textual features. The sentiment polarity is then inferred by a supervised Support Vector Machine trained on the representations of the obtained embedding space. Experiments performed on a representative dataset of 47235 labelled samples demonstrate that the exploitation of the proposed Objective Text helps to outperform state-of-the-art for sentiment polarity estimation. Alessandro Ortis, Giovanni Maria Farinella, Giovanni Torrisi, Sebastiano Battiato |
CBMI | 2 |
| 2018 | Scaling Egocentric Vision: The Dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, Michael Wray |
ECCV (4) | 3 |
| 2018 | Egocentric Shopping Cart LocalizationabstractThis work investigates the new problem of image-based egocentric shopping cart localization in retail stores. The contribution of our work is two-fold. First, we propose a novel large-scale dataset for image-based egocentric shopping cart localization. The dataset has been collected using cameras placed on shopping carts in a large retail store. It contains a total of 19,531 image frames, each labelled with its six Degrees Of Freedom pose. We study the localization problem by analysing how cart locations should be represented and estimated, and how to assess the localization results. Second, we benchmark different image-based techniques to address the task. Specifically, we investigate two families of algorithms: classic methods based on image retrieval and emerging methods based on regression. Experimental results show that methods based on image retrieval largely outperform regression-based approaches. We also point out that deep metric learning techniques allow to learn better visual representations w.r.t. other architectures, and are useful to improve the localization results of both retrieval-based and regression-based approaches. Our findings suggest that deep metric learning techniques can help bridge the gap between retrieval-based and regression-based methods. Emiliano Spera, Antonino Furnari, Sebastiano Battiato, Giovanni Maria Farinella |
ICPR | 4 |
| 2018 | Preface
Sebastiano Battiato, Patrizia Daniele, Giovanni Maria Farinella, Sofia Giuffrè, Laura Scrimali 0001 |
J. Glob. Optim. | 3 |
| 2018 | Personal-location-based temporal segmentation of egocentric videos for lifelogging applications
Antonino Furnari, Sebastiano Battiato, Giovanni Maria Farinella |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Market basket analysis from egocentric videos
Vito Santarcangelo, Giovanni Maria Farinella, Antonino Furnari, Sebastiano Battiato |
Pattern Recognit. Lett. | 2 |
| 2017 | Park SmartabstractThe paper presents Park Smart, a solution which aim is to solve the pain of finding a free parking space in public and private areas (e.g. cities, malls, etc.), and hence to optimize parking stalls allocation as well as to increase revenues for the companies which manage them. The proposed solution exploits cutting edge technologies such as IoT, Cloud Computing and Deep Learning. Daniele Di Mauro, Marco Moltisanti, Giuseppe Patanè 0002, Sebastiano Battiato, Giovanni Maria Farinella |
AVSS | 5 |
| 2017 | Computer vision for assistive technologies
Marco Leo, Gérard G. Medioni, Mohan M. Trivedi, Takeo Kanade, Giovanni Maria Farinella |
Comput. Vis. Image Underst. | 5 |
| 2017 | Next-active-object prediction from egocentric videos
Antonino Furnari, Sebastiano Battiato, Kristen Grauman, Giovanni Maria Farinella |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | Distortion adaptive Sobel filters for the gradient estimation of wide angle images
Antonino Furnari, Giovanni Maria Farinella, Arcangelo Bruna, Sebastiano Battiato |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Organizing egocentric videos of daily living activities
Alessandro Ortis, Giovanni Maria Farinella, Valeria D'Amico, Luca Addesso, Giovanni Torrisi, Sebastiano Battiato |
Pattern Recognit. | 2 |
| 2017 | Recognizing Personal Locations From Egocentric VideosabstractContextual awareness in wearable computing allows for construction of intelligent systems, which are able to interact with the user in a more natural way. In this paper, we study how personal locations arising from the user's daily activities can be recognized from egocentric videos. We assume that few training samples are available for learning purposes. Considering the diversity of the devices available on the market, we introduce a benchmark dataset containing egocentric videos of eight personal locations acquired by a user with four different wearable cameras. To make our analysis useful in real-world scenarios, we propose a method to reject negative locations, i.e., those not belonging to any of the categories of interest for the end-user. We assess the performances of the main state-of-the-art representations for scene and object classification on the considered task, as well as the influence of device-specific factors such as the field of view and the wearing modality. Concerning the different device-specific factors, experiments revealed that the best results are obtained using a head-mounted wide-angular device. Our analysis shows the effectiveness of using representations based on convolutional neural networks, employing basic transfer learning techniques and an entropy-based rejection algorithm. Antonino Furnari, Giovanni Maria Farinella, Sebastiano Battiato |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2017 | Affine Covariant Features for Fisheye Distortion Local ModelingabstractPerspective cameras are the most popular imaging sensors used in computer vision. However, many application fields, including automotive, surveillance, and robotics, require the use of wide angle cameras (e.g., fisheye), which allow to acquire a larger portion of the scene using a single device at the cost of the introduction of noticeable radial distortion in the images. Affine covariant feature detectors have proved successful in a variety of computer vision applications, including object recognition, image registration, and visual search. Moreover, their robustness to a series of variabilities related to both the scene and the image acquisition process has been thoroughly studied in the literature. In this paper, we investigate their effectiveness on fisheye images providing both theoretical and experimental analyses. As theoretical outcome, we show that the inherently non-linear radial distortion can be locally approximated by linear functions with a reasonably small error. The experimental analysis builds on Mikolajczyk's benchmark to assess the robustness of three popular affine region detectors (i.e., maximally stable extremal regions, and Harris and Hessian affine region detectors), with respect to different variabilities as well as to radial distortion. To support the evaluations, we rely on the Oxford data set and introduce a novel benchmark data set comprising 50 images depicting different scene categories. Experiments are carried out on rectilinear images to which radial distortion is artificially added, and on real-world images acquired using fisheye lenses. Our analysis points out that affine region detectors can be effectively employed directly on fisheye images and that the radial distortion is locally modeled as an additional affine variability. Antonino Furnari, Giovanni Maria Farinella, Arcangelo Bruna, Sebastiano Battiato |
IEEE Trans. Image Process. | 2 |
| 2017 | Guest Editorial Nutrition Informatics: From Food Monitoring to Dietary ManagementabstractThe papers in this special section address the concept of nutrition informatics from food monitoring to dietary management. Non-communicable diseases (NCD) account for a massively increasing proportion of the global health burden. A number of behavioral and physiological factors are related to the rising onset of NCD worldwide, with unhealthy eating playing a key role among them. In parallel, food allergies and associated acute and sometimes life-threatening reactions are a public health problem. Thus, balanced nutrition with a proper diet is the key to the prevention of diet related diseases. The recent advances in smartphone technologies, wearable sensors, computer vision and machine learning will bring the applications of nutrition informatics closer to the individuals and enable them to make better decisions regarding their daily lives. Stavroula G. Mougiakakou, Giovanni Maria Farinella, Keiji Yanai, Edward Sazonov |
IEEE J. Biomed. Health Informatics | 2 |
| 2016 | Learning Approaches for Parking Lots Classification
Daniele Di Mauro, Sebastiano Battiato, Giuseppe Patanè 0002, Marco Leotta, Daniele Maio, Giovanni Maria Farinella |
ACIVS | 6 |
| 2016 | The Social PictureabstractWe present The Social Picture, a framework to collect and explore huge amount of crowdsourced social images about public events, cultural heritage sites and other customized private events.The Social Picture aims to create social communities of users that contribute to the creation of image collections about common interests. The collections can be explored through a number of advanced Computer Vision and Machine Learning algorithms, able to capture the visual content of images in order to organize them in a semantic way. The interfaces of The Social Picture allow the users to create customized collections by exploiting semantic filters based on visual features, social network tags, geolocation, and other information related to the images. Sebastiano Battiato, Giovanni Maria Farinella, Filippo L. M. Milotta, Alessandro Ortis, Luca Addesso, Antonino Casella, Valeria D'Amico, Giovanni Torrisi |
ICMR | 2 |
| 2016 | Overview of the ACM MultiMedia 2016 International Workshop on Multimedia Assisted Dietary ManagementabstractThis abstract provides a summary and overview of the 2nd international workshop on multimedia assisted dietary management. Stavroula G. Mougiakakou, Giovanni Maria Farinella, Keiji Yanai |
ACM Multimedia | 2 |
| 2016 | Special issue on Assistive Computer Vision and Robotics - Part I
Giovanni Maria Farinella, Takeo Kanade, Marco Leo, Gérard G. Medioni, Mohan M. Trivedi |
Comput. Vis. Image Underst. | 1 |
| 2016 | Special Issue on Assistive Computer Vision and Robotics - "Assistive Solutions for Mobility, Communication and HMI"
Giovanni Maria Farinella, Takeo Kanade, Marco Leo, Gérard G. Medioni, Mohan M. Trivedi |
Comput. Vis. Image Underst. | 1 |
| 2016 | Aligning shapes for symbol classification and retrieval
Sebastiano Battiato, Giovanni Maria Farinella, Oliver Giudice, Giovanni Puglisi |
Multim. Tools Appl. | 2 |
| 2016 | Semantic segmentation of images exploiting DCT based features and random forest
Daniele Ravì, M. Bober, Giovanni Maria Farinella, Mirko Guarnera, Sebastiano Battiato |
Pattern Recognit. | 3 |
| 2015 | On Blind Source Camera Identification
Giovanni Maria Farinella, Mario Valerio Giuffrida, V. Digiacomo, Sebastiano Battiato |
ACIVS | 1 |
| 2015 | A Mobile Application for Braille to Black Conversion
Giovanni Maria Farinella, Paolo Leonardi, Filippo Stanco |
ACIVS | 1 |
| 2015 | Exploring Protected Nature Through Multimodal Navigation of Multimedia Contents
Giovanni Signorello, Giovanni Maria Farinella, Giovanni Gallo, Luciano Santo, Antonino Lopes, Emanuele Scuderi |
ACIVS | 2 |
| 2015 | Fast and Low Power Consumption Outliers Removal for Motion Vector Estimation
Giuseppe Spampinato, Arcangelo Bruna, Giovanni Maria Farinella, Sebastiano Battiato, Giovanni Puglisi |
ACIVS | 3 |
| 2015 | An Electronic Travel Aid to Assist Blind and Visually Impaired People to Avoid Obstacles
Filippo L. M. Milotta, Dario Allegra, Filippo Stanco, Giovanni Maria Farinella |
CAIP (2) | 4 |
| 2015 | Generalized Sobel Filters for gradient estimation of distorted imagesabstractIn this paper we tackle the problem of correctly estimating the gradient of distorted images. The proper estimation of the gradient in the presence of distortion is of great interest due to the large number of applications relying on wide angle cameras (e.g., in surveillance, automotive, robotics). To this aim we propose the Generalized Sobel Filters (GSF), a family of adaptive Sobel filters able to correctly estimate the gradient of distorted images. To assess the performances of the proposed method, we acquired a benchmark dataset of high resolution images belonging to different categories which are relevant to application domains where the gradient estimation is usually employed. We build an objective evaluation pipeline and perform experiments which show that our method outperforms the state-of-the-art. Antonino Furnari, Giovanni Maria Farinella, Arcangelo Bruna, Sebastiano Battiato |
ICIP | 2 |
| 2015 | RECfusion: Automatic Video Curation Driven by Visual Content PopularityabstractThe proliferation of mobile devices and the diffusion of social media have changed the communication paradigm of people that share multimedia data by allowing new interaction models (e.g., social networks). In social events (e.g., concerts), the automatic video understanding goal includes the interpretation of which visual contents are the most popular. The popularity of a visual content depends on how many people are looking at that scene, and therefore it could be obtained through the "visual consensus" among multiple video streams acquired by the different users devices. In this work we present RECfusion, a system able to automatically create a single video from multiple video sources by taking into account the popularity of the acquired scenes. The frames composing the final popular video are selected from the different video streams by considering those visual scenes which are pointed and recorded by the highest number of users' devices. Results on two benchmark datasets confirm the effectiveness of the proposed system. Alessandro Ortis, Giovanni Maria Farinella, Valeria D'Amico, Luca Addesso, Giovanni Torrisi, Sebastiano Battiato |
ACM Multimedia | 2 |
| 2015 | An integrated system for vehicle tracking and classification
Sebastiano Battiato, Giovanni Maria Farinella, Antonino Furnari, Giovanni Puglisi, Anique Snijders, Jelmer Spiekstra |
Expert Syst. Appl. | 2 |
| 2015 | Representing scenes for real-time context classification on mobile devices
Giovanni Maria Farinella, Daniele Ravì, Valeria Tomaselli, Mirko Guarnera, Sebastiano Battiato |
Pattern Recognit. | 1 |
| 2014 | Classifying food images represented as Bag of TextonsabstractThe classification of food images is an interesting and challenging problem since the high variability of the image content which makes the task difficult for current state-of-the-art classification methods. The image representation to be employed in the classification engine plays an important role. We believe that texture features have been not properly considered in this application domain. This paper points out, through a set of experiments, that textures are fundamental to properly recognize different food items. For this purpose the bag of visual words model (BoW) is employed. Images are processed with a bank of rotation and scale invariant filters and then a small codebook of Textons is built for each food class. The learned class-based Textons are hence collected in a single visual dictionary. The food images are represented as visual words distributions (Bag of Textons) and a Support Vector Machine is used for the classification stage. The experiments demonstrate that the image representation based on Bag of Textons is more accurate than existing (and more complex) approaches in classifying the 61 classes of the Pittsburgh Fast-Food Image Dataset. Giovanni Maria Farinella, Marco Moltisanti, Sebastiano Battiato |
ICIP | 1 |
| 2014 | Affine region detectors on the fisheye domainabstractFeature extractors play an important role in different Computer Vision application domains such as registration, recognition and visual search. Different detectors have been proposed and evaluated so far assuming images taken with classic cameras. However, many operating cameras (e.g., in surveillance and automotive) are built considering a fisheye model and a preprocessing step is performed to remove the distortion of the images before running a detector. The following question arises: are the current detectors suitable to work directly in the fisheye domain? To answer this question, in this paper a benchmark dataset and objective evaluation measures are considered to evaluate the performances of the state-of-the-art detectors in the fisheye domain. Test images are properly generated starting from benchmark rectilinear images and considering different fisheye focal lengths. The experiments evaluate the performances of the detectors against both increasing fisheye distortion and the combination of the fisheye distortion with photometric and geometric variability of the image content. The experiments demonstrate that affine covariant detectors can be employed directly in the fisheye domain. Furthermore, although the transformation between the rectilinear and the fisheye coordinates is not affine, we show that the mapping can be locally approximated by linear functions with a small error. Antonino Furnari, Giovanni Maria Farinella, Giovanni Puglisi, Arcangelo Bruna, Sebastiano Battiato |
ICIP | 2 |
| 2014 | Aligning codebooks for near duplicate image detection
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Puglisi, Daniele Ravì |
Multim. Tools Appl. | 2 |
| 2014 | Saliency-Based Selection of Gradient Vector Flow Paths for Content Aware Image ResizingabstractContent-aware image resizing techniques allow to take into account the visual content of images during the resizing process. The basic idea beyond these algorithms is the removal of vertical and/or horizontal paths of pixels (i.e., seams) containing low salient information. In this paper, we present a method which exploits the gradient vector flow (GVF) of the image to establish the paths to be considered during the resizing. The relevance of each GVF path is straightforward derived from an energy map related to the magnitude of the GVF associated to the image to be resized. To make more relevant, the visual content of the images during the content-aware resizing, we also propose to select the generated GVF paths based on their visual saliency properties. In this way, visually important image regions are better preserved in the final resized image. The proposed technique has been tested, both qualitatively and quantitatively, by considering a representative data set of 1000 images labeled with corresponding salient objects (i.e., ground-truth maps). Experimental results demonstrate that our method preserves crucial salient regions better than other state-of-the-art algorithms. Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Puglisi, Daniele Ravì |
IEEE Trans. Image Process. | 2 |
| 2012 | Content-aware image resizing with seam selection based on Gradient Vector FlowabstractContent-aware image resizing is an effective technique that allows to take into account the visual content of images during the resizing process. The basic idea beyond these algorithms is the resizing of an image by considering vertical and/or horizontal paths of pixels (i.e., seams) which contain low salient information. In this paper we exploit the Gradient Vector Flow (GVF) of the image to establish the paths to be considered during the resizing. The relevance of each path is derived from a saliency map obtained by considering the magnitude of the GVF associated to the image under consideration. The proposed technique has been tested, both qualitatively and quantitatively, by considering a representative set of images labeled with corresponding salient objects (i.e., ground-truth maps). Experimental results demonstrate that our method preserves crucial salient regions better than other state-of-the-art algorithms. Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Puglisi, Daniele Ravì |
ICIP | 2 |
| 2012 | Aligning Bags of Shape Contexts for Blurred Shape Model based symbol classification
Sebastiano Battiato, Giovanni Maria Farinella, Oliver Giudice, Giovanni Puglisi |
ICPR | 2 |
| 2012 | Robust Image Alignment for Tampering DetectionabstractThe widespread use of classic and newest technologies available on Internet (e.g., emails, social networks, digital repositories) has induced a growing interest on systems able to protect the visual content against malicious manipulations that could be performed during their transmission. One of the main problems addressed in this context is the authentication of the image received in a communication. This task is usually performed by localizing the regions of the image which have been tampered. To this aim the aligned image should be first registered with the one at the sender by exploiting the information provided by a specific component of the forensic hash associated to the image. In this paper we propose a robust alignment method which makes use of an image hash component based on the Bag of Features paradigm. The proposed signature is attached to the image before transmission and then analyzed at destination to recover the geometric transformations which have been applied to the received image. The estimator is based on a voting procedure in the parameter space of the model used to recover the geometric transformation occurred into the manipulated image. The proposed image hash encodes the spatial distribution of the image features to deal with highly textured and contrasted tampering patterns. A block-wise tampering detection which exploits an histograms of oriented gradients representation is also proposed. A non-uniform quantization of the histogram of oriented gradient space is used to build the signature of each image block for tampering purposes. Experiments show that the proposed approach obtains good margin of performances with respect to state-of-the art methods. Sebastiano Battiato, Giovanni Maria Farinella, Enrico Messina, Giovanni Puglisi |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2011 | Understanding geometric manipulations of images through bovw-based hashingabstractThe increasing use of low cost imaging devices and the innovations in terms of media distribution technologies induce a growing interest on technologies able to protect digital visual media against malicious manipulations of the visual contents. One of the main problems addressed in this research area is the blind detection of traces of forgery on an image obtained through the internet. Specifically, in this paper we consider the context of communications, where malicious image manipulations should be detected by a receiver. In the proposed method, an image hash based on the Bag of Visual Words paradigm is attached as signature to the image before trans mission. The forensic hash is then analyzed at destination to detect the geometric transformations which have been applied to the received image. This task is fundamental for further processing which usually assumes that the received image is aligned with the original one, as in the case of tampering detection systems. Experiments show that the proposed approach outperforms state-of-the art methods by obtaining a good margin in terms of performances. Sebastiano Battiato, Giovanni Maria Farinella, Enrico Messina, Giovanni Puglisi |
ICME | 2 |
| 2011 | Robust image registration and tampering localization exploiting bag of features based forensic signatureabstractThe distribution of digital images with the classic and newest technologies available on Internet (e.g., emails, social networks, digital repositories) has induced a growing interest on systems able to protect the visual content against malicious manipulations that could be performed during their transmission. One of the main problems addressed in this context is the authentication of the image received in a communication. This task is usually performed by localizing the regions of the image which have been tampered. To this aim the received image should be first registered with the one at the sender by exploiting the information provided by a specific component of the forensic hash associated with the image. In this paper we propose a robust alignment method which makes use of an image signature based on the Bag of Features paradigm. The alignment is based on a voting procedure in the parameter space of the model used to recover the geometric transformation occurred into the manipulated image. Experiments show that the proposed approach obtains good margin in terms of performances with respect to state-of-the art methods. Sebastiano Battiato, Giovanni Maria Farinella, Enrico Messina, Giovanni Puglisi |
ACM Multimedia | 2 |
| 2010 | Red-eyes removal through cluster based Linear Discriminant AnalysisabstractRed-eye artifact is a well-known problem in digital photography. Since the large diffusion of mobile devices with embedded camera and flashgun, automatic detection and correction of red-eyes have become an important task. In this paper we describe a technique that makes use of three steps to identify and correct red-eyes. First, red-eye candidates are extracted from the input image by using simple color segmentation coupled with geometrical constraints. A set of linear discriminant classifiers is then learned on the clustered patches space, and hence employed to distinguish between eyes and non-eyes patches. The proposed cluster-based Linear Discriminant Analysis is used to deal with the multi-modally nature of the input space. The third step of the pipeline is devoted to artifacts correction through de-saturation and brightness reduction. Experimental results on a large dataset of images demonstrate the effectiveness of the pro- posed pipeline that outperforms other existing solutions in terms of hit rates maximization, false positives reduction and ad-hoc quality measure. Sebastiano Battiato, Giovanni Maria Farinella, Mirko Guarnera, Giuseppe Messina, Daniele Ravì |
ICIP | 2 |
| 2010 | Boosting Gray Codes for Red Eyes RemovalabstractSince the large diffusion of digital camera and mobile devices with embedded camera and flashgun, the red-eyes artifacts have de-facto become a critical problem. The technique herein described makes use of three main steps to identify and remove red-eyes. First, red eyes candidates are extracted from the input image by using an image filtering pipeline. A set of classifiers is then learned on gray code features extracted in the clustered patches space, and hence employed to distinguish between eyes and non-eyes patches. Once red-eyes are detected, artifacts are removed through desaturation and brightness reduction. The proposed method has been tested on large dataset of images achieving effective results in terms of hit rates maximization, false positives reduction and quality measure. Sebastiano Battiato, Giovanni Maria Farinella, Mirko Guarnera, Giuseppe Messina, Daniele Ravì |
ICPR | 2 |
| 2010 | Exploiting visual and text features for direct marketing learning in time and space constrained domains
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Giuffrida, Catarina Sismeiro, Giuseppe Tribulato |
Pattern Anal. Appl. | 2 |
| 2009 | Spatial Hierarchy of Textons Distributions for Scene Classification
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Gallo, Daniele Ravì |
MMM | 2 |
| 2009 | Using visual and text features for direct marketing on multimedia messaging services domain
Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Giuffrida, Catarina Sismeiro, Giuseppe Tribulato |
Multim. Tools Appl. | 2 |
| 2008 | Scene categorization using bag of Textons on spatial hierarchyabstractThis paper proposes a method to recognize scene categories using bags of visual words obtained hierarchically partitioning into subregion the input images. Specifically, for each subregions the texton histogram and the extension of the sub-region is taken into account. The bags of visual words, obtained in this way, are weighted and used in a similarity measure during the categorization. Experimental tests using ten different scene categories show that the proposed approach achieves good performances with respect to the state of the art methods. Sebastiano Battiato, Giovanni Maria Farinella, Giovanni Gallo, Daniele Ravì |
ICIP | 2 |
| 2007 | Digital Mosaic Frameworks - An OverviewabstractAbstract Art often provides valuable hints for technological innovations especially in the field of Image Processing and Computer Graphics. In this paper we survey in a unified framework several methods to transform raster input images into good quality mosaics. For each of the major different approaches in literature the paper reports a short description and a discussion of the most relevant issues. To complete the survey comparisons among the different techniques both in terms of visual quality and computational complexity are provided. Sebastiano Battiato, Gianpiero di Blasi, Giovanni Maria Farinella, Giovanni Gallo |
Comput. Graph. Forum | 3 |
| 2006 | Objective Outcome Evaluation of Breast Surgery
Giovanni Maria Farinella, Gaetano Impoco, Giovanni Gallo, Salvatore Spoto, Giuseppe Catanuto, Maurizio B. Nava |
MICCAI (1) | 1 |