EDBT 2026 Demo / reviewers in the wild / expert
Alessio Del Bue
dblp:73/6117
· DBLP profile ↗
132ranked-venue papers
11as first author
60since 2021 · last 2026
0000-0002-2262-4872ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 89 · 9 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 88 · 7 first-author · 38 since 2021Systems, architecture and hardware · 11 · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SplatFill: 3D Scene Inpainting via Depth-Guided Gaussian Splatting
Mahtab Dahaghin, Milind Gajanan Padalkar, Matteo Toso, Alessio Del Bue, Vittorio Murino |
ICPR (16) | 4 |
| 2026 | Uncertainty-guided Open-Set Source-Free Unsupervised Domain Adaptation with Target-private Class SegregationabstractStandard Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target, requiring simultaneous access to both source and target data. Moreover, UDA approaches commonly assume that source and target domains share the same labels space. Yet, these two assumptions are hardly satisfied in real-world scenarios. This paper considers the more challenging Source-Free Open-set Domain Adaptation (SF-OSDA) setting, where both assumptions do not hold. We propose a novel approach for SF-OSDA that takes advantage of the granularity of target-private categories by segregating their samples into multiple unknown classes. Starting from an initial clustering-based pseudo-labels initialisation, our method progressively improves the segregation of target-private samples by refining their pseudo-labels with the guide of an uncertainty-based sample selection module. Additionally, we propose a novel contrastive loss, named NL-InfoNCELoss, that, integrating negative learning into self-supervised contrastive learning, enhances the model robustness to noisy pseudo-labels. Extensive experiments on benchmark datasets demonstrate the superiority of our proposed approach over competing methods, establishing new state-of-the-art performance. Notably, additional analyses show that our method is able to learn the underlying semantics of novel classes, opening the possibility to perform novel class discovery. Mattia Litrico, Davide Talon, Sebastiano Battiato, Alessio Del Bue, Mario Valerio Giuffrida, Pietro Morerio |
Int. J. Comput. Vis. | 4 |
| 2025 | Maps from Motion (MfM): Generating 2D Semantic Maps from Sparse Multi-View ImagesabstractWorld-wide detailed 2D maps require enormous collective efforts. OpenStreetMap is the result of 11 million registered users manually annotating the GPS location of over 1.75 billion entries, including distinctive landmarks and common urban objects. At the same time, manual annotations can include errors and are slow to update, limiting the map's accuracy. Mapsfrom Motion (MfM) is a step for-ward to automatize such time-consuming map making procedure by computing 2D maps of semantic objects directly from a collection of uncalibrated multi-view images. From each image, we extract a set of object detections, and estimate their spatial arrangement in a top-down local map centered in the reference frame of the camera that captured the image. Aligning these local maps is not a trivial problem, since they provide incomplete, noisy fragments of the scene, and matching detections across them is unreliable because of the presence of repeated pattern and the limited appearance variability of urban objects. We address this with a novel graph-based framework, that encodes the spatial and semantic distribution of the objects detected in each image, and learns how to combine them to predict the objects' poses in a global reference system, while taking into account all possible detection matches and preserving the topology observed in each image. Despite the complexity of the problem, our best model achieves global2D registration with an average accuracy within 4 meters (i.e. below GPS accuracy) even on sparse sequences with strong view-point change, on which COLMAP has an 80% failure rate. We provide extensive evaluation on synthetic and real-world data, showing how the method obtains a solution even in scenarios where standard optimization techniques fail. Find more information at matteot90.github.io/MapsFromMotion. Matteo Toso, Stefano Fiorini, Stuart Jamea, Alessio Del Bue |
3DV | 4 |
| 2025 | Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems ApproachabstractProgress in Embodied AI has made it possible for end-to-end-trained agents to navigate in photo-realistic environments with high-level reasoning and zero-shot or language-conditioned behavior, but benchmarks are still dominated by simulation. In this work, we focus on the fine-grained behavior of fast-moving real robots and present a large-scale experimental study involving 262 navigation episodes in a real environment with a physical robot, where we analyze the type of reasoning emerging from end-to-end training. In particular, we study the presence of realistic dynamics which the agent learned for open-loop forecasting, and their interplay with sensing. We analyze the way the agent uses latent memory to hold elements of the scene structure and information gathered during exploration. We probe the planning capabilities of the agent, and find in its memory evidence for somewhat precise plans over a limited horizon. Furthermore, we show in a post-hoc analysis that the value function learned by the agent relates to long-term planning. Put together, our experiments paint a new picture on how using tools from computer vision and sequential decision making have led to new capabilities in robotics and control. An interactive tool is available [here]. Steeven Janny, Hervé Poirier, Leonid Antsfeld, Guillaume Bono, Gianluca Monaci, Boris Chidlovskii, Francesco Giuliari, Alessio Del Bue, Christian Wolf 0001 |
CVPR | 8 |
| 2025 | Embodied Image Captioning: Self-Supervised Learning Agents for Spatially Coherent Image DescriptionsabstractWe present a self-supervised method to improve an agent's abilities in describing arbitrary objects while actively exploring a generic environment. This is a challenging problem, as current models struggle to obtain coherent image captions due to different camera viewpoints and clutter. We propose a three-phase framework to fine-tune existing captioning models that enhances caption accuracy and consistency across views via a consensus mechanism. First, an agent explores the environment, collecting noisy image-caption pairs. Then, a consistent pseudo-caption for each object instance is distilled via consensus using a large language model. Finally, these pseudo-captions are used to fine-tune an off-the-shelf captioning model, with the addition of contrastive learning. We analyse the performance of the combination of captioning models, exploration policies, pseudo-labeling methods, and fine-tuning strategies, on our manually labeled test set. Results show that a policy can be trained to mine samples with higher disagreement compared to classical baselines. Our pseudo-captioning method, in combination with all policies, has a higher semantic similarity compared to other existing methods, and fine-tuning improves caption accuracy and consistency by a significant margin. Code and test set annotations available at https://hsp-iit.github.io/embodied-captioning/ Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo Natale |
ICCV | 5 |
| 2025 | ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco ReconstructionabstractThe task of reassembly is a significant challenge across multiple domains, including archaeology, genomics, and molecular docking, requiring the precise placement and orientation of elements to reconstruct an original structure. In this work, we address key limitations in state-of-the-art Deep Learning methods for reassembly, namely i) scalability; ii) multimodality; and iii) real-world applicability: beyond square or simple geometric shapes, realistic and complex erosion, or other real-world problems. We propose ReassembleNet, a method that reduces complexity by representing each input piece as a set of contour keypoints and learning to select the most informative ones by Graph Neural Networks pooling inspired techniques. ReassembleNet effectively lowers computational complexity while enabling the integration of features from multiple modalities, including both geometric and texture data. Further enhanced through pretraining on a semi-synthetic dataset. We then apply diffusion-based pose estimation to recover the original structure. We improve on prior methods by 57% and 87% for RMSE Rotation and Translation, respectively. Adeela Islam, Stefano Fiorini, Stuart James, Pietro Morerio, Alessio Del Bue |
ICCV | 5 |
| 2025 | BillBoard Splatting (BBSplat): Learnable Textured Primitives for Novel View SynthesisabstractWe present billboard Splatting (BBSplat) - a novel approach for novel view synthesis based on textured geometric primitives. BBSplat represents the scene as a set of optimizable textured planar primitives with learnable RGB textures and alpha-maps to control their shape. BBSplat primitives can be used in any Gaussian Splatting pipeline as drop-in replacements for Gaussians. The proposed primitives close the rendering quality gap between 2D and 3D Gaussian Splatting (GS), enabling the accurate extraction of 3D mesh as in the 2DGS framework. Additionally, the explicit nature of planar primitives enables the use of the ray-tracing effects in rasterization. Our novel regularization term encourages textures to have a sparser structure, enabling an efficient compression that leads to a reduction in the storage space of the model up to x17 times compared to 3DGS. Our experiments show the efficiency of BBSplat on standard datasets of real indoor and outdoor scenes such as Tanks&Temples, DTU, and Mip-NeRF-360. Namely, we achieve a state-of-the-art PSNR of 29.72 for DTU at Full HD resolution. David Svitov, Pietro Morerio, Lourdes Agapito, Alessio Del Bue |
ICCV | 4 |
| 2025 | GRASPLAT: Enabling dexterous grasping through novel view synthesisabstractAchieving dexterous robotic grasping with multi-fingered hands remains a significant challenge. While existing methods rely on complete 3D scans to predict grasp poses, these approaches face limitations due to the difficulty of acquiring high-quality 3D data in real-world scenarios. In this paper, we introduce GRASPLAT, a novel grasping framework that leverages consistent 3D information while being trained solely on RGB images. Our key insight is that by synthesizing physically plausible images of a hand grasping an object, we can regress the corresponding hand joints for a successful grasp. To achieve this, we utilize 3D Gaussian Splatting to generate high-fidelity novel views of real hand-object interactions, enabling end-to-end training with RGB data. Unlike prior methods, our approach incorporates a photometric loss that refines grasp predictions by minimizing discrepancies between rendered and real images. We conduct extensive experiments on both synthetic and real-world grasping datasets, demonstrating that GRASPLAT improves grasp success rates up to 36.9% over existing image-based methods. Project page: https://mbortolon97.github.io/grasplat/ Matteo Bortolon, Nuno Ferreira Duarte, Plinio Moreno, Fabio Poiesi, José Santos-Victor, Alessio Del Bue |
IROS | 6 |
| 2025 | Measuring Uncertainty in Shape Completion to Improve Grasp QualityabstractShape completion networks have been used recently in real-world robotic experiments to complete the missing/hidden information in environments where objects are only observed in one or few instances where self-occlusions are bound to occur. Nowadays, most approaches rely on deep neural networks that handle rich 3D point cloud data that lead to more precise and realistic object geometries. However, these models still suffer from inaccuracies due to its nondeterministic/stochastic inferences which could lead to poor performance in grasping scenarios where these errors compound to unsuccessful grasps. We present an approach to calculate the uncertainty of a 3D shape completion model during inference of single view point clouds of an object on a table top. In addition, we propose an update to grasp pose algorithms quality score by introducing the uncertainty of the completed point cloud present in the grasp candidates. To test our full pipeline we perform real world grasping with a 7dof robotic arm with a 2 finger gripper on a large set of household objects and compare against previous approaches that do not measure uncertainty. Our approach ranks the grasp quality better, leading to higher grasp success rate for the rank 5 grasp candidates compared to state of the art. Project Page: https://nunoduarte.github.io/pages.3dsgrasp++ Nuno Ferreira Duarte, Seyed Saber Mohammadi, Plinio Moreno, Alessio Del Bue, José Santos-Victor |
IROS | 4 |
| 2025 | Direction-Aware Room Impulse Response Estimation for Immersive Audio Rendering in Real EnvironmentsabstractEvolving multimedia systems are increasingly being adopted in virtual reality and gaming applications. Such systems emphasize immersion to engage users by bridging the gap between real and virtual content. In this context, visual and acoustic stimuli are the two key media that dictate such immersion. While visual 3D rendering is advancing rapidly, the same is not true for audio, where most research is limited to the reconstruction of the room impulse response (RIR) using omnidirectional audio or, at best, binaural. Such methods do not adequately account for the directions and orientations of the acoustic signals with respect to either the source or the listener, thereby compromising immersion quality. In this work, we explore the effect of adding such "directionality" to the training data to improve the estimation of the room’s acoustic parameters. A more accurate set of such parameters implies in fact a more realistic predicted RIR, leading to a more immersive experience of the acoustic scene. Specifically, we propose a novel framework driven by a suitable loss function to account for directionality in ambisonic microphones, and novel variants of loss functions for both omnidirectional and ambisonic cases. We also propose to account for microphone characteristics and their contribution to the predicted RIRs. Experiments were performed using two datasets of real recordings and the results established the efficacy of the proposed methods Giovanni Zanin, Ritujoy Biswas, Pietro Morerio, Sylvio Barbon Junior, Alberto Carini, Alessio Del Bue, Vittorio Murino |
ACM Multimedia | 6 |
| 2025 | Pre-trained Multiple Latent Variable Generative Models are Good Defenders Against Adversarial AttacksabstractAttackers can deliberately perturb classifiers' input with subtle noise, altering final predictions. Among proposed countermeasures, adversarial purification employs generative networks to preprocess input images, filtering out adversarial noise. In this study, we propose specific generators, defined Multiple Latent Variable Generative Models (MLVGMs), for adversarial purification. These models possess multiple latent variables that naturally disentangle coarse from fine features. Taking advantage of these properties, we autoencode images to maintain class-relevant information, while discarding and re-sampling any detail, including adversarial noise. The procedure is completely training-free, exploring the generalization abilities of pretrained MLVGMs on the adversarial purification down-stream task. Despite the lack of large models, trained on billions of samples, we show that smaller MLVGMs are already competitive with traditional methods, and can be used as foundation models. Official code released at https://github.com/SerezD/gen_adversarial. Dario Serez, Marco Cristani, Alessio Del Bue, Vittorio Murino, Pietro Morerio |
WACV | 3 |
| 2025 | CDHN: Cross-domain hallucination network for 3D keypoints estimationabstractThis paper presents a novel method to estimate sparse 3D keypoints from single-view RGB images . Our network is trained in two steps using a knowledge distillation framework. In the first step, the teacher is trained to extract 3D features from point cloud data, which are used in combination with 2D features to estimate the 3D keypoints. In the second step, the teacher teaches the student module to hallucinate the 3D features from RGB images that are similar to those extracted from the point clouds. This procedure helps the network during inference to extract 2D and 3D features directly from images, without requiring point clouds as input. Moreover, the network also predicts a confidence score for every keypoint, which is used to select the valid ones from a set of N predicted keypoints. This allows the prediction of different number of keypoints depending on the object’s geometry. We use the estimated keypoints for computing the relative pose between two views of an object. The results are compared with those of KP-Net and StarMap , which are the state-of-the-art for estimating 3D keypoints from a single-view RGB image. The average angular distance error of our approach (5.94°) is 8.46° and 55.26° lower than that of KP-Net (14.40°) and StarMap (61.20°), respectively. Mohammad Zohaib, Milind Gajanan Padalkar, Pietro Morerio, Matteo Taiana, Alessio Del Bue |
Pattern Recognit. | 5 |
| 2025 | GANzzle++: Generative approaches for jigsaw puzzle solving as local to global assignment in latent spatial representationsabstractJigsaw puzzles are a popular and enjoyable pastime that humans can easily solve, even with many pieces. However, solving a jigsaw is a combinatorial problem, and the space of possible solutions is exponential in the number of pieces, intractable for pairwise solutions. In contrast to the classical pairwise local matching of pieces based on edge heuristics, we estimate an approximate solution image, i.e., a mental image , of the puzzle and exploit it to guide the placement of pieces as a piece-to-global assignment problem. Therefore, from unordered pieces, we consider conditioned generation approaches, including Generative Adversarial Networks (GAN) models, Slot Attention (SA) and Vision Transformers (ViT), to recover the solution image. Given the generated solution representation, we cast the jigsaw solving as a 1-to-1 assignment matching problem using Hungarian attention, which places pieces in corresponding positions in the global solution estimate. Results show that the newly proposed GANzzle-SA and GANzzle-VIT benefit from the early fusion strategy where pieces are jointly compressed and gathered for global structure recovery. A single deep learning model generalizes to puzzles of different sizes and improves the performances by a large margin. Evaluated on PuzzleCelebA and PuzzleWikiArts, our approaches bridge the gap of deep learning strategies with respect to optimization-based classic puzzle solvers. • We present new generative modules for estimating the jigsaw solution image. • We show the effect of estimating the target image for placement of pieces. • We evaluate on open datasets showing a large margin improvement. Davide Talon, Alessio Del Bue, Stuart James |
Pattern Recognit. Lett. | 2 |
| 2024 | PRAGO: Differentiable Multi-View Pose Optimization From Objectness DetectionsabstractRobustly estimating camera poses from a set of images is a fundamental task which remains challenging for differentiable methods, especially in the case of small and sparse camera pose graphs. To overcome this challenge, we propose Pose-refined Rotation Averaging Graph Optimization (PRAGO). From a set of objectness detections on unordered images, our method reconstructs the rotational pose, and in turn, the absolute pose, in a differentiable manner benefiting from the optimization of a sequence of geometrical tasks. We show how our objectness pose-refinement module in PRAGO is able to refine the inherent ambiguities in pairwise relative pose estimation without removing edges and avoiding making early decisions on the viability of graph edges. PRAGO then refines the absolute rotations through iterative graph construction, reweighting the graph edges to compute the final rotational pose, which can be converted into absolute poses using translation averaging. We show that PRAGO is able to outperform non-differentiable solvers on small and sparse scenes extracted from 7-Scenes achieving a relative improvement of 21% for rotations while achieving similar translation estimates. Matteo Taiana, Matteo Toso, Stuart James, Alessio Del Bue |
3DV | 4 |
| 2024 | HAHA: Highly Articulated Gaussian Human Avatars with Textured Mesh Prior
David Svitov, Pietro Morerio, Lourdes Agapito, Alessio Del Bue |
ACCV (9) | 4 |
| 2024 | DiffAssemble: A Unified Graph-Diffusion Model for 2D and 3D ReassemblyabstractReassembly tasks play a fundamental role in many fields and multiple approaches exist to solve specific reassembly problems. In this context, we posit that a general unified model can effectively address them all, irrespective of the input data type (images, 3D, etc.). We introduce DiffAssemble, a Graph Neural Network (GNN)-based architecture that learns to solve reassembly tasks using a diffusion model formulation. Our method treats the elements of a set, whether pieces of 2D patch or 3D object fragments, as nodes of a spatial graph. Training is performed by introducing noise into the position and rotation of the elements and iteratively denoising them to reconstruct the coherent initial pose. DiffAssemble achieves state-of-the-art (SOTA) results in most 2D and 3D reassembly tasks and is the first learning-based approach that solves 2D puzzles for both rotation and translation. Furthermore, we highlight its remarkable reduction in run-time, performing 11 times faster than the quickest optimization-based method for puzzle solving. Code available at https://github.com/IIT-PAVIS/DiffAssemble. Gianluca Scarpellini, Stefano Fiorini, Francesco Giuliari, Pietro Morerio, Alessio Del Bue |
CVPR | 5 |
| 2024 | 6DGS: 6D Pose Estimation from a Single Image and a 3D Gaussian Splatting Model
Matteo Bortolon, Theodore Tsesmelis, Stuart James, Fabio Poiesi, Alessio Del Bue |
ECCV (52) | 5 |
| 2024 | Look Around and Learn: Self-training Object Detection by Exploration
Gianluca Scarpellini, Stefano Rosa, Pietro Morerio, Lorenzo Natale, Alessio Del Bue |
ECCV (56) | 5 |
| 2024 | SelfGeo: Self-supervised and Geodesic-Consistent Estimation of Keypoints on Deformable Shapes
Mohammad Zohaib, Luca Cosmo, Alessio Del Bue |
ECCV (85) | 3 |
| 2024 | Contrastive Gaussian Clustering for Weakly Supervised 3D Scene SegmentationabstractAbstract 3D scene segmentation is a crucial task in Computer Vision, with applications in autonomous driving, augmented reality, and robotics. Traditional methods often struggle to provide consistent and accurate segmentation across different viewpoints. To address this, we look at the growing field of novel view synthesis. Methods like NeRF and 3DGS take a set of images and implicitly learn a multi-view consistent representation of the geometry of the scene; the same strategy can be extended to learn a 3D segmentation of the scene that is consistent with the 2D segmentation of an initial training set of input images. We introduce Contrastive Gaussian Clustering, a novel approach for novel segmentation view synthesis and 3D scene segmentation. We extend 3D Gaussian Splatting to include a learnable 3D feature field, which allows us to cluster the 3D Gaussians into objects. Using a combination of contrastive learning and spatial regularization, our model can be trained on inconsistent 2D segmentation labels, and still learn to generate multi-view consistent masks. Moreover, the resulting model is extremely accurate, improving the IoU accuracy of the predicted masks by $$+8\%$$ + 8 % over the state of the art. Code and trained models are available at https://github.com/MyrnaCCS/contrastive-gaussian-clustering . Myrna C. Silva, Mahtab Dahaghin, Matteo Toso, Alessio Del Bue |
ICPR (23) | 4 |
| 2024 | IFFNeRF: Initialisation Free and Fast 6DoF pose estimation from a single image and a NeRF modelabstractWe introduce IFFNeRF to estimate the six degrees-of-freedom (6DoF) camera pose of a given image, building on the Neural Radiance Fields (NeRF) formulation. IFFNeRF is specifically designed to operate in real-time and eliminates the need for an initial pose guess that is proximate to the sought solution. IFFNeRF utilizes the Metropolis-Hasting algorithm to sample surface points from within the NeRF model. From these sampled points, we cast rays and deduce the color for each ray through pixel-level view synthesis. The camera pose can then be estimated as the solution to a Least Squares problem by selecting correspondences between the query image and the resulting bundle. We facilitate this process through a learned attention mechanism, bridging the query image embedding with the embedding of parameterized rays, thereby matching rays pertinent to the image. Through synthetic and real evaluation settings, we show that our method can improve the angular and translation error accuracy by 80.1% and 67.3%, respectively, compared to iNeRF while performing at 34fps on consumer hardware and not requiring the initial pose guess. Project page: https://mbortolon97.github.io/frenerf/ Matteo Bortolon, Theodore Tsesmelis, Stuart James, Fabio Poiesi, Alessio Del Bue |
ICRA | 5 |
| 2024 | Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language NavigationabstractVision-and-Language Navigation in Continuous Environments (VLN-CE) is one of the most intuitive yet challenging embodied AI tasks. Agents are tasked to navigate towards a target goal by executing a set of low-level actions, following a series of natural language instructions. All VLN-CE methods in the literature assume that language instructions are exact. However, in practice, instructions given by humans can contain errors when describing a spatial environment due to inaccurate memory or confusion. Current VLN-CE benchmarks do not address this scenario, making the state-of-the-art methods in VLN-CE fragile in the presence of erroneous instructions from human users. For the first time, we propose a novel benchmark dataset that introduces various types of instruction errors considering potential human causes. This benchmark provides valuable insight into the robustness of VLN systems in continuous environments. We observe a noticeable performance drop (up to −25%) in Success Rate when evaluating the state-of-the-art VLN-CE methods on our benchmark. Moreover, we formally define the task of Instruction Error Detection and Localization, and establish an evaluation protocol on top of our benchmark dataset. We also propose an effective method, based on a cross-modal transformer architecture, that achieves the best performance in error detection and localization, compared to baselines. Surprisingly, our proposed method has revealed errors in the validation set of the two commonly used datasets for VLN-CE, i.e., R2R-CE and RxR-CE, demonstrating the utility of our technique in other tasks. Francesco Taioli, Stefano Rosa, Alberto Castellini, Lorenzo Natale, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Yiming Wang 0002 |
IROS | 5 |
| 2024 | Re-assembling the past: The RePAIR dataset and benchmark for real world 2D and 3D puzzle solvingabstractThis paper proposes the RePAIR dataset that represents a challenging benchmark to test modern computational and data driven methods for puzzle-solving and reassembly tasks. Our dataset has unique properties that are uncommon to current benchmarks for 2D and 3D puzzle solving. The fragments and fractures are realistic, caused by a collapse of a fresco during a World War II bombing at the Pompeii archaeological park. The fragments are also eroded and have missing pieces with irregular shapes and different dimensions, challenging further the reassembly algorithms. The dataset is multi-modal providing high resolution images with characteristic pictorial elements, detailed 3D scans of the fragments and meta-data annotated by the archaeologists. Ground truth has been generated through several years of unceasing fieldwork, including the excavation and cleaning of each fragment, followed by manual puzzle solving by archaeologists of a subset of approx. 1000 pieces among the 16000 available. After digitizing all the fragments in 3D, a benchmark was prepared to challenge current reassembly and puzzle-solving methods that often solve more simplistic synthetic scenarios. The tested baselines show that there clearly exists a gap to fill in solving this computationally complex problem. Theodore Tsesmelis, Luca Palmieri 0002, Marina Khoroshiltseva, Adeela Islam, Gur Elkin, Ofir Itzhak Shahar, Gianluca Scarpellini, Stefano Fiorini, Yaniv Ohayon, Nadav Alali, Sinem Aslan, Pietro Morerio, Sebastiano Vascon, Elena Gravina, Maria Cristina Napolitano, Giuseppe Scarpati, Gabriel Zuchtriegel, Alexandra Spühler, Michel E. Fuchs, Stuart James, Ohad Ben-Shahar, Marcello Pelillo, Alessio Del Bue |
NeurIPS | 23 |
| 2024 | I2EDL: Interactive Instruction Error Detection and LocalizationabstractIn the Vision-and-Language Navigation in Continuous Environments (VLN-CE) task, the human user guides an autonomous agent to reach a target goal via a series of low-level actions following a textual instruction in natural language. However, most existing methods do not address the likely case where users may make mistakes when providing such instruction (e.g., "turn left" instead of "turn right"). In this work, we address a novel task of Interactive VLN in Continuous Environments (IVLN-CE), which allows the agent to interact with the user during the VLN-CE navigation to verify any doubts regarding the instruction errors. We propose an Interactive Instruction Error Detector and Localizer (I2EDL) that triggers the user-agent interaction upon the detection of instruction errors during the navigation. We leverage a pre-trained module to detect instruction errors and pinpoint them in the instruction by cross-referencing the textual input and past observations. In such way, the agent is able to query the user for a timely correction, without demanding the user's cognitive load, as we locate the probable errors to a precise part of the instruction. We evaluate the proposed I2EDL on a dataset of instructions containing errors, and further devise a novel metric, the Success weighted by Interaction Number (SIN), to reflect both the navigation performance and the interaction effectiveness. We show how the proposed method can ask focused requests for corrections to the user, which in turn increases the navigation success, while minimizing the interactions. Francesco Taioli, Stefano Rosa, Alberto Castellini, Lorenzo Natale, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Yiming Wang 0002 |
RO-MAN | 5 |
| 2024 | Leveraging Next-Active Objects for Context-Aware Anticipation in Egocentric VideosabstractObjects are crucial for understanding human-object interactions. By identifying the relevant objects, one can also predict potential future interactions or actions that may occur with these objects. In this paper, we study the problem of Short-Term Object interaction anticipation (STA) and propose NAOGAT (Next-Active-Object Guided Anticipation Transformer), a multi-modal end-to-end transformer network, that attends to objects in observed frames in order to anticipate the next-active-object (NAO) and, eventually, to guide the model to predict context-aware future actions. The task is challenging since it requires anticipating future action along with the object with which the action occurs and the time after which the interaction will begin, a.k.a. the time to contact (TTC). Compared to existing video modeling architectures for action anticipation, NAOGAT captures the relationship between objects and the global scene context in order to predict detections for the next active object and anticipate relevant future actions given these detections, leveraging the objects’ dynamics to improve accuracy. One of the key strengths of our approach, in fact, is its ability to exploit the motion dynamics of objects within a given clip , which is often ignored by other models, and separately decoding the object-centric and motion-centric information. Through our experiments, we show that our model outperforms existing methods on two separate datasets, Ego4D and EpicKitchens-100 ("Unseen Set"), as measured by several additional metrics, such as time to contact, and next-active-object localization. The code can be found on project page : sanketsans.github.io/wacv24 Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino, Alessio Del Bue |
WACV | 5 |
| 2024 | Unsupervised Active Visual Search With Monte Carlo Planning Under Uncertain DetectionsabstractWe propose a solution for Active Visual Search of objects in an environment, whose 2D floor map is the only known information. Our solution has three key features that make it more plausible and robust to detector failures compared to state-of-the-art methods: i) it is unsupervised as it does not need any training sessions. ii) During the exploration, a probability distribution on the 2D floor map is updated according to an intuitive mechanism, while an improved belief update increases the effectiveness of the agent's exploration. iii) We incorporate the awareness that an object detector may fail into the aforementioned probability modelling by exploiting the success statistics of a specific detector. Our solution is dubbed POMP-BE-PD (Pomcp-based Online Motion Planning with Belief by Exploration and Probabilistic Detection). It uses the current pose of an agent and an RGB-D observation to learn an optimal search policy, exploiting a POMDP solved by a Monte-Carlo planning approach. On the Active Vision Dataset Benchmark, we increase the average success rate over all the environments by a significant 35 % while decreasing the average path length by 4 % with respect to competing methods. Thus, our results are state-of-the-art, even without any training procedure. Francesco Taioli, Francesco Giuliari, Yiming Wang 0002, Riccardo Berra, Alberto Castellini, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Francesco Setti |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Positional diffusion: Graph-based diffusion models for set orderingabstractPositional reasoning is the process of ordering an unsorted set of parts into a consistent structure. To address this problem, we present Positional Diffusion , a plug-and-play graph formulation with Diffusion Probabilistic Models. Using a diffusion process, we add Gaussian noise to the set elements’ position and map them to a random position in a continuous space. Positional Diffusion learns to reverse the noising process and recover the original positions through an Attention-based Graph Neural Network. To evaluate our method, we conduct extensive experiments on three different tasks and seven datasets, comparing our approach against the state-of-the-art methods for visual puzzle-solving, sentence ordering, and room arrangement, demonstrating that our method outperforms long-lasting research on puzzle solving with up to + 17 % compared to the second-best deep learning method, and performs on par against the state-of-the-art methods on sentence ordering and room rearrangement. Our work highlights the suitability of diffusion models for ordering problems and proposes a novel formulation and method for solving various ordering tasks. We release our code at https://github.com/IIT-PAVIS/Positional_Diffusion . • The article presents a novel method for Ordering Elements of a Set in 1D and 2D space. • We propose a task-agnostic method, Positional Diffusion for different ordering tasks • Our approach combines Graph Neural Networks with Diffusion Probabilistic Models. • Without any task-specific modes, our method can outperform task-specific approaches. • We test our approach on Sentence ordering, Visual Puzzles, and Furniture Arrangement. Francesco Giuliari, Gianluca Scarpellini, Stefano Fiorini, Stuart James, Pietro Morerio, Yiming Wang 0002, Alessio Del Bue |
Pattern Recognit. Lett. | 7 |
| 2023 | Learnable Data Augmentation for One-Shot Unsupervised Domain Adaptation
Julio Ivan Davila Carrazco, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
BMVC | 3 |
| 2023 | Guiding Pseudo-labels with Uncertainty Estimation for Source-free Unsupervised Domain AdaptationabstractStandard Unsupervised Domain Adaptation (UDA) methods assume the availability of both source and target data during the adaptation. In this work, we investigate Source-free Unsupervised Domain Adaptation (SF-UDA), a specific case of UDA where a model is adapted to a target domain without access to source data. We propose a novel approach for the SF-UDA setting based on a loss reweighting strategy that brings robustness against the noise that inevitably affects the pseudo-labels. The classification loss is reweighted based on the reliability of the pseudo-labels that is measured by estimating their uncertainty. Guided by such reweighting strategy, the pseudo-labels are progressively refined by aggregating knowledge from neighbouring samples. Furthermore, a self-supervised contrastive framework is leveraged as a target space regulariser to enhance such knowledge aggregation. A novel negative pairs exclusion strategy is proposed to identify and exclude negative pairs made of samples sharing the same class, even in presence of some noise in the pseudo-labels. Our method outperforms previous methods on three major benchmarks by a large margin. We set the new SF-UDA state-of-the-art on VisDA-C and DomainNet with a performance gain of + 1.8% on both benchmarks and on PACS with + 12.3% in the single-source setting and +6.6% in multi-target adaptation. Additional analyses demonstrate that the proposed approach is robust to the noise, which results in significantly more accurate pseudo-labels compared to state-of-the-art approaches. Mattia Litrico, Alessio Del Bue, Pietro Morerio |
CVPR | 2 |
| 2023 | Audio-Visual Inpainting: Reconstructing Missing Visual Information with SoundabstractWe tackle audio-visual inpainting, the problem of completing an image in such a way to be consistent with the sound associated to the scene. To this end, we propose a multimodal, audio-visual inpainting method (AVIN), and show how to leverage sound to reconstruct semantically consistent images. AVIN is a 2-stage algorithm, which first learns the scene semantics and reconstructs low resolution images based on a conditional probability distribution of pixels in the space conditioned to audio, and then refines such result with a GAN-based network to increase the resolution of the reconstructed image. We show that AVIN is able to recover the original content, especially in the hard cases where the missing area heavily degrades the scene semantics: it can perform cross-modal generation whenever no visual context is observed at all, reconstructing visual data from sound only. Code will be made available upon acceptance. Valentina Sanguineti, Sanket Kumar Thakur, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
ICASSP | 4 |
| 2023 | Person Re-Identification without Identification via Event AnonymizationabstractWide-scale use of visual surveillance in public spaces puts individual privacy at stake while increasing resource consumption (energy, bandwidth, and computation). Neuromorphic vision sensors (event-cameras) have been recently considered a valid solution to the privacy issue because they do not capture detailed RGB visual information of the subjects in the scene. However, recent deep learning architectures have been able to reconstruct images from event cameras with high fidelity, reintroducing a potential threat to privacy for event-based vision applications. In this paper, we aim to anonymize event-streams to protect the identity of human subjects against such image reconstruction attacks. To achieve this, we propose an end-to-end network architecture jointly optimized for the twofold objective of preserving privacy and performing a downstream task such as person ReId. Our network learns to scramble events, enforcing the degradation of images recovered from the privacy attacker. In this work, we also bring to the community the first ever event-based person ReId dataset gathered to evaluate the performance of our approach. We validate our approach with extensive experiments and report results on the synthetic event data simulated from the publicly available SoftBio dataset and our proposed Event-ReId dataset. The code is available at https://github.com/IIT-PAVIS/ReId_without_Id Shafiq Ahmad, Pietro Morerio, Alessio Del Bue |
ICCV | 3 |
| 2023 | SC3K: Self-supervised and Coherent 3D Keypoints Estimation from Rotated, Noisy, and Decimated Point Cloud DataabstractThis paper proposes a new method to infer keypoints from arbitrary object categories in practical scenarios where point cloud data (PCD) are noisy, down-sampled and arbitrarily rotated. Our proposed model adheres to the following principles: i) keypoints inference is fully unsupervised (no annotation given), ii) keypoints position error should be low and resilient to PCD perturbations (robustness), iii) keypoints should not change their indexes for the intra-class objects (semantic coherence), iv) keypoints should be close to or proximal to PCD surface (compactness). We achieve these desiderata by proposing a new self-supervised training strategy for keypoints estimation that does not assume any a priori knowledge of the object class, and a model architecture with coupled auxiliary losses that promotes the desired keypoints properties. We compare the keypoints estimated by the proposed approach with those of the state-of-the-art unsupervised approaches. The experiments show that our approach outperforms by estimating keypoints with improved coverage (+9.41%) while being semantically consistent (+4.66%) that best characterize the object’s 3D shape for downstream tasks. Code and data are available at: https://github.com/IIT-PAVIS/SC3K Mohammad Zohaib, Alessio Del Bue |
ICCV | 2 |
| 2023 | Inclusive Digital Storytelling: Artificial Intelligence and Augmented Reality to Re-centre Stories from the Margins
Valentina Nisi, Stuart James, Paulo Bala, Alessio Del Bue, Nuno Nunes 0001 |
ICIDS (1) | 4 |
| 2023 | Enhancing Next Active Object-Based Egocentric Action Anticipation with Guided AttentionabstractShort-term action anticipation (STA) in first-person videos is a challenging task that involves understanding the next active object interactions and predicting future actions. Existing action anticipation methods have primarily focused on utilizing features extracted from video clips, but often overlooked the importance of objects and their interactions. To this end, we propose a novel approach that applies a guided attention mechanism between the objects, and the spatiotemporal features extracted from video clips, enhancing the motion and contextual information, and further decoding the object-centric and motion-centric information to address the problem of STA in egocentric videos. Our method, GANO (Guided Attention for Next active Objects) is a multi-modal, end-to-end, single transformer-based network. The experimental results performed on the largest egocentric dataset demonstrate that GANO outperforms the existing state-of-the-art methods for the prediction of the next active object label, its bounding box location, the corresponding future action, and the time to contact the object. The ablation study shows the positive contribution of the guided attention mechanism compared to other fusion methods. Moreover, it is possible to improve the next active object location and class label prediction results of GANO by just appending the learnable object tokens with the region of interest embeddings. Related implementations are available at: sanketsans.github.io/guided-attention-egocentric.html Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Vittorio Murino, Alessio Del Bue |
ICIP | 5 |
| 2023 | 3DSGrasp: 3D Shape-Completion for Robotic GraspabstractReal-world robotic grasping can be done robustly if a complete 3D Point Cloud Data (PCD) of an object is available. However, in practice, PCDs are often incomplete when objects are viewed from few and sparse viewpoints before the grasping action, leading to the generation of wrong or inaccurate grasp poses. We propose a novel grasping strategy, named 3DSGrasp, that predicts the missing geometry from the partial PCD to produce reliable grasp poses. Our proposed PCD completion network is a Transformer-based encoder-decoder network with an Offset-Attention layer. Our network is inherently invariant to the object pose and point's permutation, which generates PCDs that are geometrically consistent and completed properly. Experiments on a wide range of partial PCD show that 3DSGrasp outperforms the best state-of-the-art method on PCD completion tasks and largely improves the grasping success rate in real-world scenarios. The code and dataset are available at: https://github.com/NunoDuarte/3DSGrasp. Seyed Saber Mohammadi, Nuno Ferreira Duarte, Dimitrios Dimou, Yiming Wang 0002, Matteo Taiana, Pietro Morerio, Atabak Dehban, Plinio Moreno, Alexandre Bernardino, Alessio Del Bue, José Santos-Victor |
ICRA | 10 |
| 2023 | Towards Equivariant Optical Flow Estimation with Deep LearningabstractMethods for Optical Flow (OF) estimation based on Deep Learning have considerably improved traditional approaches in challenging and realistic conditions. However, data-driven approaches can inherently be biased, leading to unexpected under-performance in real application scenarios. In this paper, we first observe that the OF estimation accuracy varies with motion direction, and name this phenomenon ‘OF sign imbalance’. The sign imbalance cannot be assessed by means of the endpoint-error (EPE), the typical training and evaluation metric for Deep Optical Flow estimators. This paper tackles this issue by proposing a new metric to assess the sign imbalance, which is compared to the endpoint-error. We provide an extensive evaluation of the sign imbalance for the state-of-the-art optical flow estimators. Based on the evaluation, we propose two strategies to mitigate the phenomenon, i) by constraining the model estimations during inference, and, ii) by constraining the loss function during training. Testing and training code is available at: www.github.com/stsavian/equivariant_of_estimation. Stefano Savian, Pietro Morerio, Alessio Del Bue, Andrea Janes, Tammam Tillo |
WACV | 3 |
| 2023 | "Connected to the people": Social Inclusion & Cohesion in Action through a Cultural Heritage Digital ToolabstractCurrent cultural policies are evolving from social inclusion (removing barriers and promoting equality for participation in culture) to social cohesion (fostering solid bonds between groups despite their differences). Digital interventions can create spaces that promote social inclusion and cohesion. In this paper, we report on the design and evaluation of a cultural heritage and digital storytelling application supporting a participatory approach to culture and hosting society. We evaluate our intervention in three marginalized communities with different social-cultural contexts: migrant women in Barcelona, a community living in a priority neighbourhood in Paris and second and third-generation migrants in Lisbon. Through an analysis of their application use, our findings point at their needs and desires, highlighting how the app can support social inclusion as the first step towards cohesion, but that these are heterogeneous concepts susceptible to nuanced appropriations by the different communities. Valentina Nisi, Paulo Bala, Vanessa Cesário, Stuart James, Alessio Del Bue, Nuno Nunes 0001 |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2023 | No Adversaries to Zero-Shot Learning: Distilling an Ensemble of Gaussian Feature GeneratorsabstractIn zero-shot learning (ZSL), the task of recognizing unseen categories when no data for training is available, state-of-the-art methods generate visual features from semantic auxiliary information (e.g., attributes). In this work, we propose a valid alternative (simpler, yet better scoring) to fulfill the very same task. We observe that, if first- and second-order statistics of the classes to be recognized were known, sampling from Gaussian distributions would synthesize visual features that are almost identical to the real ones as per classification purposes. We propose a novel mathematical framework to estimate first- and second-order statistics, even for unseen classes: our framework builds upon prior compatibility functions for ZSL and does not require additional training. Endowed with such statistics, we take advantage of a pool of class-specific Gaussian distributions to solve the feature generation stage through sampling. We exploit an ensemble mechanism to aggregate a pool of softmax classifiers, each trained in a one-seen-class-out fashion to better balance the performance over seen and unseen classes. Neural distillation is finally applied to fuse the ensemble into a single architecture which can perform inference through one forward pass only. Our method, termed Distilled Ensemble of Gaussian Generators, scores favorably with respect to state-of-the-art works. Jacopo Cavazza, Vittorio Murino, Alessio Del Bue |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Leveraging Commonsense for Object Localisation in Partial ScenesabstractWe propose an end-to-end solution to address the problem of object localisation in partial scenes, where we aim to estimate the position of an object in an unknown area given only a partial 3D scan of the scene. We propose a novel scene representation to facilitate the geometric reasoning, Directed Spatial Commonsense Graph (D-SCG), a spatial scene graph that is enriched with additional concept nodes from a commonsense knowledge base. Specifically, the nodes of D-SCG represent the scene objects and the edges are their relative positions. Each object node is then connected via different commonsense relationships to a set of concept nodes. With the proposed graph-based scene representation, we estimate the unknown position of the target object using a Graph Neural Network that implements a sparse attentional message passing mechanism. The network first predicts the relative positions between the target object and each visible object by learning a rich representation of the objects via aggregating both the object nodes and the concept nodes in D-SCG. These relative positions then are merged to obtain the final position. We evaluate our method using Partial ScanNet, improving the state-of-the-art by 5.9% in terms of the localisation accuracy at a 8x faster training speed. Francesco Giuliari, Geri Skenderi, Marco Cristani, Alessio Del Bue, Yiming Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Locality-aware subgraphs for inductive link prediction in knowledge graphsabstractRecent methods for inductive reasoning on Knowledge Graphs (KGs) transform the link prediction problem into a graph classification task. They first extract a subgraph around each target link based on the k-hop neighborhood of the target entities, encode the subgraphs using a Graph Neural Network (GNN), then learn a function that maps subgraph structural patterns to link existence. Although these methods have witnessed great successes, increasing k often leads to an exponential expansion of the neighborhood, thereby degrading the GNN expressivity due to oversmoothing. In this paper, we formulate the subgraph extraction as a local clustering procedure that aims at sampling tightly-related subgraphs around the target links, based on a personalized PageRank (PPR) approach. Empirically, on three real-world KGs, we show that reasoning over subgraphs extracted by PPR-based local clustering can lead to a more accurate link prediction model than relying on neighbors within fixed hop distances. Furthermore, we investigate graph properties such as average clustering coefficient and node degree, and show that there is a relation between these and the performance of subgraph-based link prediction. Hebatallah A. Mohamed Hassan 0001, Diego Pilutti, Stuart James, Alessio Del Bue, Marcello Pelillo, Sebastiano Vascon |
Pattern Recognit. Lett. | 4 |
| 2022 | Spatial Commonsense Graph for Object Localisation in Partial ScenesabstractWe solve object localisation in partial scenes, a new problem of estimating the unknown position of an object (e.g. where is the bag?) given a partial 3D scan of a scene. The proposed solution is based on a novel scene graph model, the Spatial Commonsense Graph (SCG), where objects are the nodes and edges define pairwise distances between them, enriched by concept nodes and relationships from a commonsense knowledge base. This allows SCG to better generalise its spatial inference over unknown 3D scenes. The SCG is used to estimate the unknown position of the target object in two steps: first, we feed the SCG into a novel Proximity Prediction Network, a graph neural network that uses attention to perform distance prediction between the node representing the target object and the nodes representing the observed objects in the SCG; second, we propose a Localisation Module based on circular intersection to estimate the object position using all the predicted pairwise distances in order to be independent of any reference system. We create a new dataset of partially reconstructed scenes to benchmark our method and baselines for object localisation in partial scenes, where our proposed method achieves the best localisation performance. Francesco Giuliari, Geri Skenderi, Marco Cristani, Yiming Wang 0002, Alessio Del Bue |
CVPR | 5 |
| 2022 | PoserNet: Refining Relative Camera Poses Exploiting Object Detections
Matteo Taiana, Matteo Toso, Stuart James, Alessio Del Bue |
ECCV (33) | 4 |
| 2022 | Fusion and Orthogonal Projection for Improved Face-Voice AssociationabstractWe study the problem of learning association between face and voice. Prior works adopt pairwise or triplet loss formulations to learn an embedding space amenable for associated matching and verification tasks. Albeit showing some progress, such loss formulations are restrictive due to dependency on distance-dependent margin parameter, poor run-time training complexity, and reliance on carefully crafted negative mining procedures. In this work, we hypothesize that enriched feature representation coupled with an effective yet efficient supervision is necessary in realizing a discriminative joint embedding space for improved face-voice association. To this end, we propose a light-weight, plug-and-play mechanism that exploits the complementary cues in both modalities to form enriched fused embeddings and clusters them based on their identity labels via orthogonality constraints. We coin our proposed mechanism as fusion and orthogonal projection (FOP) and instantiate in a two-stream pipeline. The overall resulting framework is evaluated on a large-scale VoxCeleb dataset with a multitude of tasks, including cross-modal verification and matching. Our method performs favourably against the current state-of-the-art methods and our proposed supervision formulation is more effective and efficient than the ones employed by the contemporary methods. Muhammad Saad Saeed, Muhammad Haris Khan, Shah Nawaz, Muhammad Haroon Yousaf, Alessio Del Bue |
ICASSP | 5 |
| 2022 | Writing with (Digital) Scissors: Designing a Text Editing Tool for Assisted Storytelling Using Crowd-Generated Content
Paulo Bala, Stuart James, Alessio Del Bue, Valentina Nisi |
ICIDS | 3 |
| 2022 | Ganzzle: Reframing Jigsaw Puzzle Solving as a Retrieval Task using a Generative Mental ImageabstractPuzzle solving is a combinatorial challenge due to the difficulty of matching adjacent pieces. Instead, we infer a mental image from all pieces, which a given piece can then be matched against avoiding the combinatorial explosion. Exploiting advancements in Generative Adversarial methods, we learn how to reconstruct the image given a set of unordered pieces, allowing the model to learn a joint embedding space to match an encoding of each piece to the cropped layer of the generator. Therefore we frame the problem as a R@1 retrieval task, and then solve the linear assignment using differentiable Hungarian attention, making the process end-to-end. In doing so our model is puzzle size agnostic, in contrast to prior deep learning methods which are single size. We evaluate on two new large-scale datasets, where our model is on par with deep learning methods, while generalizing to multiple puzzle sizes. Davide Talon, Alessio Del Bue, Stuart James |
ICIP | 2 |
| 2022 | Fast re-OBJ: real-time object re-identification in rigid scenes
Ertugrul Bayraktar, Yiming Wang 0002, Alessio Del Bue |
Mach. Vis. Appl. | 3 |
| 2022 | Unsupervised Synthetic Acoustic Image Generation for Audio-Visual Scene UnderstandingabstractAcoustic images are an emergent data modality for multimodal scene understanding. Such images have the peculiarity of distinguishing the spectral signature of the sound coming from different directions in space, thus providing a richer information as compared to that derived from single or binaural microphones. However, acoustic images are typically generated by cumbersome and costly microphone arrays which are not as widespread as ordinary microphones. This paper shows that it is still possible to generate acoustic images from off-the-shelf cameras equipped with only a single microphone and how they can be exploited for audio-visual scene understanding. We propose three architectures inspired by Variational Autoencoder, U-Net and adversarial models, and we assess their advantages and drawbacks. Such models are trained to generate spatialized audio by conditioning them to the associated video sequence and its corresponding monaural audio track. Our models are trained using the data collected by a microphone array as ground truth. Thus they learn to mimic the output of an array of microphones in the very same conditions. We assess the quality of the generated acoustic images considering standard generation metrics and different downstream tasks (classification, cross-modal retrieval and sound localization). We also evaluate our proposed models by considering multimodal datasets containing acoustic images, as well as datasets containing just monaural audio signals and RGB video frames. In all of the addressed downstream tasks we obtain notable performances using the generated acoustic data, when compared to the state of the art and to the results obtained using real acoustic images as input. Valentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
IEEE Trans. Image Process. | 3 |
| 2021 | Audio-Visual Localization by Synthetic Acoustic Image GenerationabstractAcoustic images constitute an emergent data modality for multimodal scene understanding. Such images have the peculiarity to distinguish the spectral signature of sounds coming from different directions in space, thus providing richer information than the one derived from mono and binaural microphones. However, acoustic images are typically generated by cumbersome microphone arrays, which are not as widespread as ordinary microphones mounted on optical cameras. To exploit this empowered modality while using standard microphones and cameras we propose to leverage the generation of synthetic acoustic images from common audio-video data for the task of audio-visual localization. The generation of synthetic acoustic images is obtained by a novel deep architecture, based on Variational Autoencoder and U-Net models, which is trained to reconstruct the ground truth spatialized audio data collected by a microphone array, from the associated video and its corresponding monaural audio signal. Namely, the model learns how to mimic what an array of microphones can produce in the same conditions. We assess the quality of the generated synthetic acoustic images on the task of unsupervised sound source localization in a qualitative and quantitative manner, while also considering standard generation metrics. Our model is evaluated by considering both multimodal datasets containing acoustic images, used for the training, and unseen datasets containing just monaural audio signals and RGB frames, showing to reach more accurate localization results as compared to the state of the art. Valentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
AAAI | 3 |
| 2021 | Unsupervised Human Action Recognition with Skeletal Graph Laplacian and Self-Supervised Viewpoints Invariance
Giancarlo Paoletti, Jacopo Cavazza, Cigdem Beyan, Alessio Del Bue |
BMVC | 4 |
| 2021 | (Just) A Spoonful of Refinements Helps the Registration Error Go DownabstractWe tackle data-driven 3D point cloud registration. Given point correspondences, the standard Kabsch algorithm provides an optimal rotation estimate. This allows to train registration models in an end-to-end manner by differentiating the SVD operation. However, given the initial rotation estimate supplied by Kabsch, we show we can improve point correspondence learning during model training by extending the original optimization problem. In particular, we linearize the governing constraints of the rotation matrix and solve the resulting linear system of equations. We then iteratively produce new solutions by updating the initial estimate. Our experiments show that, by plugging our differentiable layer to existing learning-based registration methods, we improve the correspondence matching quality. This yields up to a 7% decrease in rotation error for correspondence-based data-driven registration methods. Sérgio Agostinho, Aljosa Osep, Alessio Del Bue, Laura Leal-Taixé |
ICCV | 3 |
| 2021 | Lights: Light Specularity Dataset For Specular Detection In Multi-ViewabstractSpecular highlights are commonplace in images, however, methods for detecting them and removing the phenomenon are particularly challenging. A reason for this is the difficulty in creating a dataset for training or evaluation, as in the real world, we lack the necessary control over the environment. Therefore, we propose a novel physically-based rendered LIGHT Specularity (LIGHTS) Dataset for the evaluation of the specular highlight detection task. Our dataset consists of 18 high-quality architectural scenes, where each scene is rendered with multiple views. In total, the dataset contains 2, 603 views with an average of 145 views per scene. Additionally, we propose a simple aggregation based method for specular highlight detection that outperforms prior work by 3.6% in two orders of magnitude less time on our dataset. Mohamed Dahy Elkhouly, Theodore Tsesmelis, Alessio Del Bue, Stuart James |
ICIP | 3 |
| 2021 | Pointview-GCN: 3D Shape Classification With Multi-View Point CloudsabstractWe address 3D shape classification with partial point cloud inputs captured from multiple viewpoints around the object. Different from existing methods that perform classification on the complete point cloud by first registering multi-view capturing, we propose PointView-GCN with multi-level Graph Convolutional Networks (GCNs) to hierarchically aggregate the shape features of single-view point clouds, in order to encode both the geometrical cues of an object and their multi-view relations. With experiments on our novel single-view datasets, we prove that PointView-GCN produces a more descriptive global shape feature which stably improves the classification accuracy by $\sim 5$% compared to the classifiers with single-view point clouds, and outperforms the state-of-the-art methods with the complete point clouds on ModelNet40. Seyed Saber Mohammadi, Yiming Wang 0002, Alessio Del Bue |
ICIP | 3 |
| 2021 | End-To-End Pairwise Human Proxemics from Uncalibrated Single ImagesabstractIn this work, we address the ill-posed problem of estimating pairwise metric distances between people using only a single uncalibrated image. We propose an end-to-end model, DeepProx, that takes as inputs two skeletal joints as a set of 2D image coordinates and outputs the metric distance between them. We show that an increased performance is achieved by a geometrical loss over simplified camera parameters provided at training time. Further, DeepProx achieves a remarkable generalisation over novel viewpoints through domain generalisation techniques. We validate our proposed method quantitatively and qualitatively against baselines on public datasets for which we provided groundtruth on interpersonal distances. Pietro Morerio, Matteo Bustreo, Yiming Wang 0002, Alessio Del Bue |
ICIP | 4 |
| 2021 | Multi-Illumination Fusion With Crack Enhancement Using Cycle-Consistent LossesabstractThis paper addresses for the first time, the problem of multi-illumination fusion with crack enhancement. Our models are trained using cycle-consistent losses to combine crack details from several mutually registered multi-illumination images of ceramic tiles, into a single representative image. Using real-world industrial data, we show that the crack locations are enhanced in the fused images, making them easily noticeable for remote inspection, and demonstrate the effectiveness of our method compared to a multi-exposure fusion technique. Milind Gajanan Padalkar, Carlos Beltrán 0002, Alessio Del Bue |
ICIP | 3 |
| 2021 | Predicting Gaze from Egocentric Social Interaction Videos and IMU DataabstractGaze prediction in egocentric videos is a fairly new research topic, which might have several applications for assistive technology (e.g., supporting blind people in their daily interactions), security (e.g., attention tracking in risky work environments), education (e.g., augmented / mixed reality training simulators, immersive games) and so forth. Egocentric gaze is typically estimated from video while few works attempt to use inertial measurement unit (IMU) data, a sensor modality often available in wearable devices (e.g., augmented reality headsets). Instead, in this paper, we examine whether joint learning of egocentric video and corresponding IMU data can improve the first-person gaze prediction compared to using these modalities separately. In this respect, we propose a multimodal network and evaluate it on several unconstrained social interaction scenarios captured by a first-person perspective. The proposed multimodal network achieves better results compared to unimodal methods as well as several (multimodal) baselines, showing that using egocentric video together with IMU data can boost the first-person gaze estimation performance. Sanket Kumar Thakur, Cigdem Beyan, Pietro Morerio, Alessio Del Bue |
ICMI | 4 |
| 2021 | POMP++: Pomcp-based Active Visual Search in unknown indoor environmentsabstractIn this paper, we focus on the problem of learning online an optimal policy for Active Visual Search (AVS) of objects in unknown indoor environments. We propose POMP++, a planning strategy that introduces a novel formulation on top of the classic Partially Observable Monte Carlo Planning (POMCP) framework, to allow training-free online policy learning in unknown environments. We present a new belief reinvigoration strategy that enables the use of POMCP with a dynamically growing state space to address the online generation of the floor map. We evaluate our method on two public benchmark datasets, AVD that is acquired by real robotic platforms and Habitat ObjectNav that is rendered from real 3D scene scans, achieving the best success rate with an improvement of >10% over the state-of-the-art methods. Francesco Giuliari, Alberto Castellini, Riccardo Berra, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Francesco Setti, Yiming Wang 0002 |
IROS | 4 |
| 2021 | Single Image Human Proxemics Estimation for Visual Social DistancingabstractIn this work, we address the problem of estimating the so-called "Social Distancing" given a single uncalibrated image in unconstrained scenarios. Our approach proposes a semi-automatic solution to approximate the homography matrix between the scene ground and image plane. With the estimated homography, we then leverage an off-the-shelf pose detector to detect body poses on the image and to reason upon their inter-personal distances using the length of their body-parts. Inter-personal distances are further locally inspected to detect possible violations of the social distancing rules. We validate our proposed method quantitatively and qualitatively against baselines on public domain datasets for which we provided groundtruth on interpersonal distances. Besides, we demonstrate the application of our method deployed in a real testing scenario where statistics on the inter-personal distances are currently used to improve the safety in a critical environment. Maya Aghaei, Matteo Bustreo, Yiming Wang 0002, Gian Luca Bailo, Pietro Morerio, Alessio Del Bue |
WACV | 6 |
| 2021 | A Benchmark and Evaluation of Non-Rigid Structure from MotionabstractAbstract Non-rigid structure from motion (nrsfm), is a long standing and central problem in computer vision and its solution is necessary for obtaining 3D information from multiple images when the scene is dynamic. A main issue regarding the further development of this important computer vision topic, is the lack of high quality data sets. We here address this issue by presenting a data set created for this purpose, which is made publicly available, and considerably larger than the previous state of the art. To validate the applicability of this data set, and provide an investigation into the state of the art of nrsfm, including potential directions forward, we here present a benchmark and a scrupulous evaluation using this data set. This benchmark evaluates 18 different methods with available code that reasonably spans the state of the art in sparse nrsfm. This new public data set and evaluation protocol will provide benchmark tools for further development in this challenging field. Sebastian Nesgaard Jensen, Mads Brix Doest, Henrik Aanæs, Alessio Del Bue |
Int. J. Comput. Vis. | 4 |
| 2021 | AIforCOVID: Predicting the clinical outcomes in patients with COVID-19 applying AI to chest-X-rays. An Italian multicentre studyabstractRecent epidemiological data report that worldwide more than 53 million people have been infected by SARS-CoV-2, resulting in 1.3 million deaths. The disease has been spreading very rapidly and few months after the identification of the first infected, shortage of hospital resources quickly became a problem. In this work we investigate whether artificial intelligence working with chest X-ray (CXR) scans and clinical data can be used as a possible tool for the early identification of patients at risk of severe outcome, like intensive care or death. Indeed, further to induce lower radiation dose than computed tomography (CT), CXR is a simpler and faster radiological technique, being also more widespread. In this respect, we present three approaches that use features extracted from CXR images, either handcrafted or automatically learnt by convolutional neuronal networks, which are then integrated with the clinical data. As a further contribution, this work introduces a repository that collects data from 820 patients enrolled in six Italian hospitals in spring 2020 during the first COVID-19 emergency. The dataset includes CXR images, several clinical attributes and clinical outcomes. Exhaustive evaluation shows promising performance both in 10-fold and leave-one-centre-out cross-validation, suggesting that clinical data and images have the potential to provide useful information for the management of patients and hospital resources. Paolo Soda, Natascha Claudia D'Amico, Jacopo Tessadori, Giovanni Valbusa, Valerio Guarrasi, Chandra Bortolotto, Muhammad Usman Akbar, Rosa Sicilia, Ermanno Cordelli, Deborah Fazzini, Michaela Cellina, Giancarlo Oliva, Giovanni Callea, Silvia Panella, Maurizio Cariati, Diletta Cozzi, Vittorio Miele, Elvira Stellato, Gianpaolo Carrafiello, Giulia Castorani, Annalisa Simeone, Lorenzo Preda, Giulio Iannello, Alessio Del Bue, Fabio Tedoldi, Marco Alì, Diego Sona, Sergio Papa |
Medical Image Anal. | 24 |
| 2021 | Forecasting People Trajectories and Head Poses by Jointly Reasoning on Tracklets and VisletsabstractIn this article, we explore the correlation between people trajectories and their head orientations. We argue that people trajectory and head pose forecasting can be modelled as a joint problem. Recent approaches on trajectory forecasting leverage short-term trajectories (aka tracklets) of pedestrians to predict their future paths. In addition, sociological cues, such as expected destination or pedestrian interaction, are often combined with tracklets. In this article, we propose MiXing-LSTM (MX-LSTM) to capture the interplay between positions and head orientations (vislets) thanks to a joint unconstrained optimization of full covariance matrices during the LSTM backpropagation. We additionally exploit the head orientations as a proxy for the visual attention, when modeling social interactions. MX-LSTM predicts future pedestrians location and head pose, increasing the standard capabilities of the current approaches on long-term trajectory forecasting. Compared to the state-of-the-art, our approach shows better performances on an extensive set of public benchmarks. MX-LSTM is particularly effective when people move slowly, i.e., the most challenging scenario for all other models. The proposed approach also allows for accurate predictions on a longer time horizon. Irtiza Hasan, Francesco Setti, Theodore Tsesmelis, Vasileios Belagiannis, Sikandar Amin, Alessio Del Bue, Marco Cristani, Fabio Galasso |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | POMP: Pomcp-based Online Motion Planning for active visual search in indoor environments
Yiming Wang 0002, Francesco Giuliari, Riccardo Berra, Alberto Castellini, Alessio Del Bue, Alessandro Farinelli, Marco Cristani, Francesco Setti |
BMVC | 5 |
| 2020 | Where to Explore Next? ExHistCNN for History-Aware Autonomous 3D Exploration
Yiming Wang 0002, Alessio Del Bue |
ECCV (29) | 2 |
| 2020 | Analysis of Face-Touching Behavior in Large Scale Social Interaction DatasetabstractWe present the first publicly available annotations for the analysis of face-touching behavior. These annotations are for a dataset composed of audio-visual recordings of small group social interactions with a total number of 64 videos, each one lasting between 12 to 30 minutes and showing a single person while participating to four-people meetings. They were performed by in total 16 annotators with an almost perfect agreement (Cohen's Kappa=0.89) on average. In total, 74K and 2M video frames were labelled as face-touch and no-face-touch, respectively. Given the dataset and the collected annotations, we also present an extensive evaluation of several methods: rule-based, supervised learning with hand-crafted features and feature learning and inference with a Convolutional Neural Network (CNN) for Face-Touching detection. Our evaluation indicates that among all, CNN performed the best, reaching 83.76% F1-score and 0.84 Matthews Correlation Coefficient. To foster future research in this problem, code and dataset were made publicly available (github.com/IIT-PAVIS/Face-Touching-Behavior), providing all video frames, face-touch annotations, body pose estimations including face and hands key-points detection, face bounding boxes as well as the baseline methods implemented and the cross-validation splits used for training and evaluating our models. Cigdem Beyan, Matteo Bustreo, Muhammad Shahid 0002, Gian Luca Bailo, Nicolò Carissimi, Alessio Del Bue |
ICMI | 6 |
| 2020 | Complex-Object Visual Inspection: Empirical Studies on A Multiple Lighting SolutionabstractThe design of an automatic visual inspection system is usually performed in two stages. While the first stage consists in selecting the most suitable hardware setup for highlighting most effectively the defects on the surface to be inspected, the second stage concerns the development of algorithmic solutions to exploit the potentials offered by the collected data. In this paper, first, we present a novel illumination setup embedding four illumination configurations to resemble diffused, dark-field, and front lighting techniques. Second, we analyze the contributions brought by deploying the proposed setup in the training phase only, mimicking the scenario in which an already developed visual inspection system cannot be modified on the customer site. Along with an exhaustive set of experiments, in this paper, we demonstrate the suitability of the proposed setup for effective illumination of complex-objects, defined as manufactured items with variable surface characteristics that cannot be determined a priori. Eventually, we provide insights into the importance of multiple light configurations availability during training and their natural boosting effect which, without the need to modify the system design in the evaluation phase, lead to improvements in the overall system performance. Maya Aghaei, Matteo Bustreo, Pietro Morerio, Nicolò Carissimi, Alessio Del Bue, Vittorio Murino |
ICPR | 5 |
| 2020 | Are Multiple Cross-Correlation Identities better than just Two? Improving the Estimate of Time Differences-of-Arrivals from Blind Audio SignalsabstractGiven an unknown audio source, the estimation of time differences-of-arrivals (TDOAs) can be efficiently and robustly solved using blind channel identification and exploiting the cross-correlation identity (CCI). Prior “blind” works have improved the estimate of TDOAs by means of different algorithmic solutions and optimization strategies, while always sticking to the case N=2 microphones. But what if we can obtain a direct improvement in performance by just increasing N? In this paper we try to investigate this direction, showing that, despite the arguable simplicity, this is capable of (sharply) improving upon state-of-the-art blind channel identification methods based on CCI, without modifying the computational pipeline. Inspired by our results, we seek to warm up the community and the practitioners by paving the way (with two concrete, yet preliminary, examples) towards joint approaches in which advances in the optimization are combined with an increased number of microphones, in order to achieve further improvements. Danilo Greco, Jacopo Cavazza, Alessio Del Bue |
ICPR | 3 |
| 2020 | Weakly Supervised Geodesic Segmentation of Egyptian Mummy CT ScansabstractIn this paper, we tackle the task of automatically analyzing 3D volumetric scans obtained from computed tomography (CT) devices. In particular, we address a particular task for which data is very limited: the segmentation of ancient Egyptian mummies CT scans. We aim at digitally unwrapping the mummy and identify different segments such as body, bandages and jewelry. The problem is complex because of the lack of annotated data for the different semantic regions to segment, thus discouraging the use of strongly supervised approaches. We, therefore, propose a weakly supervised and efficient interactive segmentation method to solve this challenging problem. After segmenting the wrapped mummy from its exterior region using histogram analysis and template matching, we first design a voxel distance measure to find an approximate solution for the body and bandage segments. Here, we use geodesic distances since voxel features as well as spatial relationship among voxels is incorporated in this measure. Next, we refine the solution using a GrabCut based segmentation together with a tracking method on the slices of the scan that assigns labels to different regions in the volume, using limited supervision in the form of scribbles drawn by the user. The efficiency of the proposed method is demonstrated using visualizations and validated through quantitative measures and qualitative unwrapping of the mummy. Avik Hati, Matteo Bustreo, Diego Sona, Vittorio Murino, Alessio Del Bue |
ICPR | 5 |
| 2020 | A Versatile Crack Inspection Portable System based on Classifier Ensemble and Controlled IlluminationabstractThis paper presents a novel setup for automatic visual inspection of cracks in ceramic tile as well as studies the effect of various classifiers and height-varying illumination conditions for this task. The intuition behind this setup is that cracks can be better visualized under specific lighting conditions than others. Our setup, which is designed for field work with constraints in its maximum dimensions, can acquire images for crack detection with multiple lighting conditions using the illumination sources placed at multiple heights. Crack detection is then performed by classifying patches extracted from the acquired images in a sliding window fashion. We study the effect of lights placed at various heights by training classifiers both on customized as well as state-of-the-art architectures and evaluate their performance both at patch-level and image-level, demonstrating the effectiveness of our setup. More importantly, ours is the first study that demonstrates how height-varying illumination conditions can affect crack detection with the use of existing state-of-the-art classifiers. We provide an insight about the illumination conditions that can help in improving crack detection in a challenging real-world industrial environment. Milind Gajanan Padalkar, Carlos Beltrán 0002, Matteo Bustreo, Alessio Del Bue, Vittorio Murino |
ICPR | 4 |
| 2020 | Subspace Clustering for Action Recognition with Covariance Representations and Temporal PruningabstractThis paper tackles the problem of human action recognition, defined as classifying which action is displayed in a trimmed sequence, from skeletal data. Albeit state-of-the-art approaches designed for this application are all supervised, in this paper we pursue a more challenging direction: solving the problem with unsupervised learning. To this end, we propose a novel subspace clustering method, which exploits covariance matrix to enhance the action's discriminability and a times-tamp pruning approach that allow us to better handle the temporal dimension of the data. Through a broad experimental validation, we show that our computational pipeline surpasses existing unsupervised approaches but also can result in favorable performances as compared to the supervised methods. The code is available here: https://github.com/IIT-PAVIS/subspace-clustering-action-recognition Giancarlo Paoletti, Jacopo Cavazza, Cigdem Beyan, Alessio Del Bue |
ICPR | 4 |
| 2020 | Revisiting Projective Structure from Motion: A Robust and Efficient Incremental SolutionabstractThis paper presents a solution to the Projective Structure from Motion (PSfM) problem able to deal efficiently with missing data, outliers and, for the first time, large scale 3D reconstruction scenarios. By embedding the projective depths into the projective parameters of the points and views, we decrease the number of unknowns to estimate and improve computational speed by optimizing standard linear Least Squares systems instead of homogeneous ones. In order to do so, we show that an extension of the linear constraints from the Generalized Projective Reconstruction Theorem can be transferred to the projective parameters, ensuring also a valid projective reconstruction in the process. We use an incremental approach that, starting from a solvable sub-problem, incrementally adds views and points until completion with a robust, outliers free, procedure. To prevent error accumulation, a refinement based on alternation between new estimations of views and points is used. This can also be done with constrained non-linear optimization. Experiments with simulated data shows that our approach is performing well, both in term of the quality of the reconstruction and the capacity to handle missing data and outliers with a reduced computational time. Finally, results on real datasets shows the ability of the method to be used in medium and large scale 3D reconstruction scenarios with high ratios of missing data (up to 98 percent). Ludovic Magerand, Alessio Del Bue |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Machine Learning for Cultural Heritage: A SurveyabstractThe application of Machine Learning (ML) to Cultural Heritage (CH) has evolved since basic statistical approaches such as Linear Regression to complex Deep Learning models. The question remains how much of this actively improves on the underlying algorithm versus using it within a ‘black box’ setting. We survey across ML and CH literature to identify the theoretical changes which contribute to the algorithm and in turn them suitable for CH applications. Alternatively, and most commonly, when there are no changes, we review the CH applications, features and pre/post-processing which make the algorithm suitable for its use. We analyse the dominant divides within ML, Supervised, Semi-supervised and Unsupervised, and reflect on a variety of algorithms that have been extensively used. From such an analysis, we give a critical look at the use of ML in CH and consider why CH has only limited adoption of ML. Marco Fiorucci, Marina Khoroshiltseva, Massimiliano Pontil, Arianna Traviglia, Alessio Del Bue, Stuart James |
Pattern Recognit. Lett. | 5 |
| 2019 | Human-Centric Light Sensing and Estimation From RGBD Images: The Invisible Light SwitchabstractLighting design in indoor environments is of primary importance for at least two reasons: 1) people should perceive an adequate light; 2) an effective lighting design means consistent energy saving. We present the Invisible Light Switch (ILS) to address both aspects. ILS dynamically adjusts the room illumination level to save energy while maintaining constant the light level perception of the users. So the energy saving is invisible to them. Our proposed ILS leverages a radiosity model to estimate the light level which is perceived by a person within an indoor environment, taking into account the person position and her/his viewing frustum (head pose). ILS may therefore dim those luminaires, which are not seen by the user, resulting in an effective energy saving, especially in large open offices (where light may otherwise be ON everywhere for a single person). To quantify the system performance, we have collected a new dataset where people wear luxmeter devices while working in office rooms. The luxmeters measure the amount of light (in Lux) reaching the people gaze, which we consider a proxy to their illumination level perception. Our initial results are promising: in a room with 8 LED luminaires, the energy consumption in a day may be reduced from 18585 to 6206 watts with ILS (currently needing 1560 watts for operations). While doing so, the drop in perceived lighting decreases by just 200 lux, a value considered negligible when the original illumination level is above 1200 lux, as is normally the case in offices. Theodore Tsesmelis, Irtiza Hasan, Marco Cristani, Alessio Del Bue, Fabio Galasso |
WACV | 4 |
| 2019 | RGBD2lux: Dense Light Intensity Estimation With an RGBD SensorabstractLighting design and modelling or industrial applications like luminaire planning and commissioning rely heavily on time-consuming manual measurements or on physically coherent computational simulations. Regarding the latter, standard approaches are based on CAD modeling simulations and offline rendering, with long processing times and therefore inflexible workflows. Thus, in this paper we propose a computer vision based system to measure lighting with just a single RGBD camera. The proposed method uses both depth data and images from the sensor to provide a dense measure of light intensity in the field of view of the camera. We evaluate our system on novel ground truth data and compare it to state-of-the-art commercial light planning software. Our system provides improved performance, while being completely automated, given that the CAD model is extracted from the depth and the albedo estimated with the support of RGB images. To the best of our knowledge, this is the first automatic framework for the estimation of lighting in general indoor scenarios from RGBD input. Theodore Tsesmelis, Irtiza Hasan, Marco Cristani, Fabio Galasso, Alessio Del Bue |
WACV | 5 |
| 2018 | Visual Graphs from Motion (VGfM): Scene Understanding with Object Geometry Reasoning
Paul Gay, Stuart James, Alessio Del Bue |
ACCV (3) | 3 |
| 2018 | MX-LSTM: Mixing Tracklets and Vislets to Jointly Forecast Trajectories and Head PosesabstractRecent approaches on trajectory forecasting use tracklets to predict the future positions of pedestrians exploiting Long Short Term Memory (LSTM) architectures. This paper shows that adding vislets, that is, short sequences of head pose estimations, allows to increase significantly the trajectory forecasting performance. We then propose to use vislets in a novel framework called MX-LSTM, capturing the interplay between tracklets and vislets thanks to a joint unconstrained optimization of full covariance matrices during the LSTM backpropagation. At the same time, MX-LSTM predicts the future head poses, increasing the standard capabilities of the long-term trajectory forecasting approaches. With standard head pose estimators and an attentional-based social pooling, MX-LSTM scores the new trajectory forecasting state-of-the-art in all the considered datasets (Zara01, Zara02, UCY, and TownCentre) with a dramatic margin when the pedestrians slow down, a case where most of the forecasting approaches struggle to provide an accurate solution. Irtiza Hasan, Francesco Setti, Theodore Tsesmelis, Alessio Del Bue, Fabio Galasso, Marco Cristani |
CVPR | 4 |
| 2018 | ROOM REFLECTORS ESTIMATION FROM SOUND BY GREEDY ITERATIVE APPROACHabstractRoom reconstruction from sound is the problem of estimating indoor space boundaries given only a set of sound events acquired by an array of microphones. In order to find a solution in realistic scenarios, it is necessary a robust and practical method that can solve the echo labelling problem, i.e. assigning at each signal delay the correct reflector that has generated it. Although being an NP-hard problem, in this paper we demonstrate that it is possible to solve the echo labelling problem in a reasonable computational time without the need of additional hypotheses on the echoes order of arrival. Marco Crocco, Andrea Trucco, Alessio Del Bue |
ICASSP | 3 |
| 2018 | Practical Motion Segmentation for Urban Street View ScenesabstractThough a long-studied problem, motion segmentation has yet to migrate into practical applications. We argue that a vital step towards that goal lies in addressing motion segmentation for the specific setting of interest. To this end, this paper presents a new approach for image-based motion segmentation in the case of vehicles navigating inside an urban environment. We exploit two application-specific factors - the restricted camera movement and the known type of moving objects - to deal with the two major limiting factors - missing data and strong perspective effects - that affect most previous “generic” motion segmentation algorithms. By constraining the geometry and exploiting known semantic classes in the scene, we achieve much higher accuracy than previous approaches. In addition to the novel algorithm, we contribute a more realistic motion segmentation benchmark dataset for moving platforms by annotating real video sequences from the KITTI dataset. Experiments on this dataset and other synthetic data confirm the effectiveness of the proposed approach. Cosimo Rubino, Alessio Del Bue, Tat-Jun Chin |
ICRA | 2 |
| 2018 | Multi-view Aggregation for Color Naming with Shadow Detection and RemovalabstractThis paper presents a set of methods for classifying the color attribute of objects when multiple images of the same objects are available. This problem is more complex than the single image estimation since varying environmental effects, such as, shadows or specularities from light sources, can result in poor accuracy. These depend primarily on the camera positions and the material type of the objects. Single image techniques focus on improving the discrimination of between colors, whereas in multi-view systems additional information is available but should be utilized wisely. To this end, we propose three methods to aggregate image pixel information in multi-view that boost the performance of color name classification. Moreover, we study the effect of shadows by employing automatic shadow detection and correction techniques on the color naming problem. We tested our proposals on a new multi-view color names dataset (M3DCN) which contain indoor and outdoor objects. The experimental evaluation shows that one out of the three presented aggregation methods is very efficient and it achieves the highest accuracy in term of classification results. Also, we experimentally show that addressing visual outliers like shadow in multi-view images improves the performance of the color attribute decision process. Mohamed Dahy Elkhouly, Stuart James, Alessio Del Bue |
IPAS | 3 |
| 2018 | "Seeing is Believing": Pedestrian Trajectory Forecasting Using Visual Frustum of AttentionabstractIn this paper we show the importance of the head pose estimation in the task of trajectory forecasting. This cue, when produced by an oracle and injected in a novel socially-based energy minimization approach, allows to get state-of-the-art performances on four different forecasting benchmarks, without relying on additional information such as expected destination and desired speed, which are supposed to be know beforehand for most of the current forecasting techniques. Our approach uses the head pose estimation for two aims: 1) to define a view frustum of attention, highlighting the people a given subject is more interested about, in order to avoid collisions; 2) to give a shorttime estimation of what would be the desired destination point. Moreover, we show that when the head pose estimation is given by a real detector, though the performance decreases, it still remains at the level of the top score forecasting systems. Irtiza Hasan, Francesco Setti, Theodore Tsesmelis, Alessio Del Bue, Marco Cristani, Fabio Galasso |
WACV | 4 |
| 2018 | Factorization based structure from motion with object priors
Paul Gay, Cosimo Rubino, Marco Crocco, Alessio Del Bue |
Comput. Vis. Image Underst. | 4 |
| 2018 | 3D Object Localisation from Multi-View Image DetectionsabstractIn this work we present a novel approach to recover objects 3D position and occupancy in a generic scene using only 2D object detections from multiple view images. The method reformulates the problem as the estimation of a quadric (ellipsoid) in 3D given a set of 2D ellipses fitted to the object detection bounding boxes in multiple views. We show that a closed-form solution exists in the dual-space using a minimum of three views while a solution with two views is possible through the use of non-linear optimisation and object constraints on the size of the object shape. In order to make the solution robust toward inaccurate bounding boxes, a likely occurrence in object detection methods, we introduce a data preconditioning technique and a non-linear refinement of the closed form solution based on implicit subspace constraints. Results on synthetic tests and on different real datasets, involving challenging scenarios, demonstrate the applicability and potential of our method in several realistic scenarios. Cosimo Rubino, Marco Crocco, Alessio Del Bue |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Manifold constraint transfer for visual structure-driven optimization
Baochang Zhang 0001, Alessandro Perina, Ce Li 0002, Qixiang Ye, Vittorio Murino, Alessio Del Bue |
Pattern Recognit. | 6 |
| 2018 | Deep Endoscope: Intelligent Duct Inspection for the Avionic IndustryabstractWe present the first autonomous endoscope for the visual inspection of very small ducts and cavities, up to a 6-mm diameter. The system has been designed, implemented, and tested in a challenging industrial scenario and in strict collaboration with an avionic industry partner. The inspected objects are metallic gearboxes eventually presenting different residuals (e.g., sand, machining swarfs, and metallic dust) inside the oil ducts. The automatic system is actuated by a robotic arm that moves the endoscope with a microcamera inside the gearbox duct, while a deep-learning-based spatio-temporal image analysis module detects, classifies, and localizes defects in real time. Feedback is given to the robotic arm in order to move or extract the endoscope given the detected anomalies. Evaluation provides a detection rate of nearly 98% given different tests with different types of residuals and duct structures. Samuele Martelli, Luca Mazzei, Carlo Canali, Paolo Guardiani, Salvatore Giunta, Alberto Ghiazza, Ivan Mondino, Ferdinando Cannella, Vittorio Murino, Alessio Del Bue |
IEEE Trans. Ind. Informatics | 10 |
| 2017 | Probabilistic Structure from Motion with Objects (PSfMO)abstractThis paper proposes a probabilistic approach to recover affine camera calibration and objects position/occupancy from multi-view images using solely the information from image detections. We show that remarkable object localisation and volumetric occupancy can be recovered by including both geometrical constraints and prior information given by objects CAD models from the ShapeNet dataset. This can be done by recasting the problem in the context of a probabilistic framework based on PPCA that enforces both geometrical constraints and the associated semantic given by the object category extracted by the object detector We present results on synthetic data and extensive real evaluation on the ScanNet datasets on more than 1200 image sequences to show the validity of our approach in realistic scenarios. In particular, we show that 3D statistical priors are key to obtain reliable reconstruction especially when the input detections are noisy, a likely case in real scenes. Paul Gay, Vaibhav Bansal, Cosimo Rubino, Alessio Del Bue |
ICCV | 4 |
| 2017 | Practical Projective Structure from Motion (P2SfM)abstractThis paper presents a solution to the Projective Structure from Motion (PSfM) problem able to deal efficiently with missing data, outliers and, for the first time, large scale 3D reconstruction scenarios. By embedding the projective depths into the projective parameters of the points and views, we decrease the number of unknowns to estimate and improve computational speed by optimizing standard linear Least Squares systems instead of homogeneous ones. In order to do so, we show that an extension of the linear constraints from the Generalized Projective Reconstruction Theorem can be transferred to the projective parameters, ensuring also a valid projective reconstruction in the process. We use an incremental approach that, starting from a solvable sub-problem, incrementally adds views and points until completion with a robust, outliers free, procedure. Experiments with simulated data shows that our approach is performing well, both in term of the quality of the reconstruction and the capacity to handle missing data and outliers with a reduced computational time. Finally, results on real datasets shows the ability of the method to be used in medium and large scale 3D reconstruction scenarios with high ratios of missing data (up to 98%). Ludovic Magerand, Alessio Del Bue |
ICCV | 2 |
| 2017 | Tiny head pose classification by bodily cuesabstractThe head pose is an important cue for computer vision. Traditionally considered in human computer interaction applications, it becomes very hard to model in surveillance scenarios, due to the tiny head size. Additionally, no public dataset contains continuous head pose annotations in open scenery, making the challenge even harder to face. Here we present a framework based on Faster RCNN, which introduces a branch in the network architecture related to the head pose estimation. The key idea is to leverage the presence of the people body to better infer the head pose, through a joint optimization process. Additionally, we enrich the Town Center dataset with head pose labels, promoting further study on this topic. Results on this novel benchmark and ablation studies on other task-specific datasets promote our idea and confirm the importance of the body cues to contextualize the head pose estimation. Irtiza Hasan, Theodore Tsesmelis, Fabio Galasso, Alessio Del Bue, Marco Cristani |
ICIP | 4 |
| 2017 | Automatic inspection of aeronautic components
Marco San-Biagio, Carlos Beltrán 0002, Salvatore Giunta, Alessio Del Bue, Vittorio Murino |
Mach. Vis. Appl. | 4 |
| 2017 | Adaptive Local Movement Modeling for Robust Object TrackingabstractIn this paper, we present a new strategy for modeling the motion of local patches for single-object tracking that can be seamlessly applied to most part-based trackers in the literature. The proposed adaptive local movement modeling method is able to model the local movement distribution of the image patches defining the object to track and the reliability of each image patch. Given the output of a base tracking algorithm, a Gaussian mixture model (GMM) is first used to model the distribution of the movement of local patches relative to the center of gravity of the tracked object. Then, the GMM is combined with the chosen base tracker in a boosting framework, which gives an efficient integrated scheme for the tracking task. This provides a robust procedure to detect outliers in the local motion of the patches. The algorithm is highly configurable with the possibility to change the number of local patches used for tracking and to adapt to the variations of the tracked object. The extensive tracking results on standard data sets show that equipping state-of-the-art trackers with our technique remarkably improves their performance. Baochang Zhang 0001, Alessandro Perina, Alessio Del Bue, Vittorio Murino, Jianzhuang Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Structure from Motion with ObjectsabstractThis paper shows for the first time that is possible to reconstruct the position of rigid objects and to jointly recover affine camera calibration solely from a set of object detections in a video sequence. In practice, this work can be considered as the extension of Tomasi and Kanade factorization method using objects. Instead of using points to form a rank constrained measurement matrix, we can form a matrix with similar rank properties using 2D object detection proposals. In detail, we first fit an ellipse onto the image plane at each bounding box as given by the object detector. The collection of all the ellipses in the dual space is used to create a measurement matrix that gives a specific rank constraint. This matrix can be factorised and metrically upgraded in order to provide the affine camera matrices and the 3D position of the objects as an ellipsoid. Moreover, we recover the full 3D quadric thus giving additional information about object occupancy and 3D pose. Finally, we also show that 2D points measurements can be seamlessly included in the framework to reduce the number of objects required. This last aspect unifies the classical point-based Tomasi and Kanade approach with objects in a unique framework. Experiments with synthetic and real data show the feasibility of our approach for the affine camera case. Marco Crocco, Cosimo Rubino, Alessio Del Bue |
CVPR | 3 |
| 2016 | Estimation of TDOA for room reflections by iterative weighted l1 constraintabstractEstimation of Time Difference of Arrivals (TDOAs) corresponding to early room reflections can be formulated as a Blind Channel Identification (BCI) problem, exploiting signals acquired by a set of microphones, given an unknown transmitted source. To cope with the intrinsic noise sensitivity and ill-posedness of the problem, sparsity and non-negativity priors on the Acoustic Impulse Response (AIR) of the room can be exploited. Here we propose a novel iterative method, resulting in a sequence of convex problems relying on a weighted l1constraint. The proposed method allows to outperform current state of the art on speech and non-speech real signals, while lowering the number of parameters to tune and the sensitivity of solution given the few free parameters. Marco Crocco, Alessio Del Bue |
ICASSP | 2 |
| 2016 | Learning-based approach to boost detection rate and localisation accuracy in single molecule localisation microscopyabstractThis paper introduces a novel system for the analysis of superresolution microscopy images using a learning based approach boosting performance and simplicity of use. Key component of single-molecule-localisation (SML) microscopy techniques is the ability to localise single emitting molecules in a stack of noisy images with a high degree of accuracy. To this end, we propose a SVM-based detector coupled with a weighted Least Squares fitting approach that favourably compares with the state of the art on established evaluation protocols. Silvia Colabrese, Marco Castello, Giuseppe Vicidomini, Alessio Del Bue |
ICIP | 4 |
| 2016 | Person re-identification using sparse representation with manifold constraintsabstractHuman re-identification is still a challenging task due to the human pose and illumination variations. Nowadays, surveillance cameras with high frame rate are capable of capturing several consecutive frames from each person. Multi-shot images provide richer information of the target person compared to a single-shot image. They, however, produce a high cost of information redundancy which may degrade the performance of re-identification systems. In this paper, we propose a novel framework that combines sparse coding and manifold constraints to extract discriminative information from multi-shot images of one pedestrian for person re-identification across a set of non-overlapped surveillance cameras. The evaluation over two standard multi-shot datasets shows very competitive accuracy of our framework against the state-of-the-art. Behzad Mirmahboub, Hamed Kiani Galoogahi, Amran Bhuiyan, Alessandro Perina, Baochang Zhang 0001, Alessio Del Bue, Vittorio Murino |
ICIP | 6 |
| 2016 | Fast 6D pose estimation for texture-less objects from a single RGB imageabstractA fundamental step to solve bin-picking and grasping problems is the accurate estimation of an object 3D pose. Such visual task usually rely on profusely textured objects: standard procedures such as detection of interest points or computation of appearance-based descriptors are favoured by using a highly informative surface. However, texture-less objects or their parts (i.e., those whose surface texture is poorly conditioned) are common in any environment but still challenging to deal with. This is due the fact that the distribution of surface brightness makes difficult to compute interest points or appearance-based descriptors. In this paper, we propose a method to estimate the 3D pose for texture-less objects given a coarse initialization: the pose is estimated using using edge correspondences, where the similarity measure is encoded using a pre-computed linear regression matrix. Furthermore, we also propose a method to increase the robustness of the estimated pose against background and object clutter. We validate both methods by using synthetic and real image sequences with objects with known ground truth. Enrique Muñoz, Yoshinori Konishi, Vittorio Murino, Alessio Del Bue |
ICRA | 4 |
| 2016 | Fast 6D pose from a single RGB image using Cascaded Forests TemplatesabstractThis paper presents a method for 6D pose estimation from a single RGB image for complex texture-less objects. This class of objects are common in any environment but still challenging to deal with. This is due to the fact that the distribution of surface brightness makes difficult to compute interest points or appearance-based descriptors. Here we propose a novel part-based method using an efficient template matching approach where each template independently encodes the similarity function using a Forest trained over the templates. Moreover, accuracy is even more incremented by using a cascade of the learned forest. These templates forests together with the simplicity of the computed image features allow a quick estimate of the pose achieving real-time performance. Performance are demonstrated both on synthetic and real images with known ground truth. Enrique Muñoz, Yoshinori Konishi, Carlos Beltrán 0002, Vittorio Murino, Alessio Del Bue |
IROS | 5 |
| 2015 | Sparse representation classification with manifold constraints transferabstractThe fact that image data samples lie on a manifold has been successfully exploited in many learning and inference problems. In this paper we leverage the specific structure of data in order to improve recognition accuracies in general recognition tasks. In particular we propose a novel framework that allows to embed manifold priors into sparse representation-based classification (SRC) approaches. We also show that manifold constraints can be transferred from the data to the optimized variables if these are linearly correlated. Using this new insight, we define an efficient alternating direction method of multipliers (ADMM) that can consistently integrate the manifold constraints during the optimization process. This is based on the property that we can recast the problem as the projection over the manifold via a linear embedding method based on the Geodesic distance. The proposed approach is successfully applied on face, digit, action and objects recognition showing a consistently increase on performance when compared to the state of the art. Baochang Zhang 0001, Alessandro Perina, Vittorio Murino, Alessio Del Bue |
CVPR | 4 |
| 2015 | Semantic Multi-body Motion SegmentationabstractThis paper presents a method to deal with the multi-body segmentation problem using a set of 2D points matches between two views. The key feature of our approach is the explicit inclusion of a higher semantic information as given by general purpose object detectors that boost the segmentation of the moving objects. In the classical formulation of the problem, only 2D matched points between views are used to identify independently moving objects based on the principle that a set of points belonging to a moving object would satisfy some given multi-view relations (e.g. multi-body epipolar constraints). We improve and speedup such process by including the information that a set of 2D matches may belong to the same object given the output of a detector. As such, instead of sampling points uniformly with a RANSAC based strategy, the selection of the matches is driven by the position and score confidence of the object detectors. Evaluation on challenging synthetic and real datasets shows a remarkable improvement in respect to previous approaches, regarding both the number of iterations required to segment a scene and the effectiveness of the segmentation itself, often making the difference between satisfying segmentation and almost complete failure. Cosimo Rubino, Marco Crocco, Vittorio Murino, Alessio Del Bue |
WACV | 4 |
| 2015 | Adaptive Local Movement Modelling for Object TrackingabstractIn this paper we present a novel strategy for modelling the motion of local patches for single object tracking that can be seamlessly applied to most part-based trackers in the literature. The proposed Adaptive Local Movement Modelling (ALMM) method is able to model the local spatial distribution of the image patches defining the object to track and the reliability of each image patch. Given the output of a base tracking algorithm, a Gaussian Mixture Model (GMM) is first used to model the distribution of the movement of local patches relative to the gravity center of the tracked object. Then, the GMM is combined with the base tracker in a boosting framework, which gives a novel integrated boosting classifier for the tracking task. This provides a robust procedure to detect outliers in the local motion of the patches. The algorithm is highly configurable with the possibility to change the number of local patches used for tracking and to adapt to the variations of the tracked object. Tracking results on standard datasets show that equipping state-of-the-art trackers with our tehcnique remarkably improves their performance. Baochang Zhang 0001, Alessandro Perina, Alessio Del Bue, Vittorio Murino |
WACV | 4 |
| 2015 | Garment-based motion capture (GaMoCap): high-density capture of human shape in motion
Nicolò Biasi, Francesco Setti, Alessio Del Bue, Mattia Tavernini, Massimo Lunardelli, Alberto Fornaser, Mauro Da Lio, Mariolino De Cecco |
Mach. Vis. Appl. | 3 |
| 2014 | A directional visual descriptor for large-scale coverage problemsabstractVisual coverage of large scale environments is a challenging problem that has many practical applications such as large scale 3D reconstruction, search and rescue and active video surveillance. In this paper, we consider a setting where mobile robots must acquire visual information using standard cameras, while minimizing associated movement costs. The main source of complexity for such scenario is the lack of a priori knowledge of 3D structures for the surrounding environment. To address this problem, we propose a novel descriptor for visual coverage that aims at measuring the orientation dependent visual information of an area, based on a regular discretization of the 3D environment in voxels. Next, we use the proposed visual descriptor to define an autonomous cooperative exploration approach, which controls the robot movements so to maximize information accuracy and minimizing movement costs. We empirically evaluate our approach in a simulation scenario based on real data for large scale 3D environments, and on widely used robotic tools (such as ROS and Stage). Experimental results show that the proposed method significantly outperforms a baseline random approach and an uncoordinated one, thus being a valid proposal for visual coverage in large scale outdoor scenarios. Marco Tamassia, Alessandro Farinelli, Vittorio Murino, Alessio Del Bue |
IROS | 4 |
| 2013 | Joint estimation of segmentation and structure from motion
Luca Zappella, Alessio Del Bue, Xavier Lladó, Joaquim Salvi |
Comput. Vis. Image Underst. | 2 |
| 2013 | Adaptive Non-rigid Registration and Structure from Motion from Image Trajectories
Alessio Del Bue |
Int. J. Comput. Vis. | 1 |
| 2013 | Human behavior analysis in video surveillance: A Social Signal Processing perspective
Marco Cristani, Ramachandra Raghavendra, Alessio Del Bue, Vittorio Murino |
Neurocomputing | 3 |
| 2012 | A Closed Form Solution for the Self-Calibration of Heterogeneous SensorsabstractWe present a novel closed-form solution for the joint self-calibration of video and range sensors. The approach single assumption is the availability of synchronous time of flight (i.e., range distances) measurements and visual position of the target on images acquired by a set of cameras. In such case, we make explicit a rank constraint that is valid for both image and range data. This rank property is used to find an initial and affine solution via bilinear factorization, which is then corrected by enforcing the metric constraints characteristic for both sensor modalities (i.e., camera and anchors constraints). The output of the algorithm is the identification of the target/range sensor position and the calibration of the cameras. The application extent of our approach is broad and versatile. In fact, with the same framework, we can deal with, but not restricted to, two very different applications. The first is aimed at calibrating cameras and microphones deployed in an unknown environment. The second uses a RGB-D device to reconstruct the 3D position of a set of keypoints using the camera and depth map images. Synthetic and real tests show the algorithm performance under different levels of noise and configurations of target locations, number of sensors and cameras. Marco Crocco, Alessio Del Bue, Igor Barros Barbosa, Vittorio Murino |
BMVC | 2 |
| 2012 | Artistic Image Classification: An Analysis on the PRINTART Database
Gustavo Carneiro 0001, Nuno Pinho da Silva, Alessio Del Bue, João Paulo Costeira |
ECCV (4) | 3 |
| 2012 | Learning Discriminative Spatial Relations for Detector Dictionaries: An Application to Pedestrian Detection
Enver Sangineto, Marco Cristani, Alessio Del Bue, Vittorio Murino |
ECCV (2) | 3 |
| 2012 | A closed form solution to the microphone position self-calibration problemabstractThis paper presents a novel algorithm for the automatic 3D localization of a set of microphones in an unknown environment. Given the times of arrival at each microphone of a set of sound events, the approach simultaneously estimates the 3D positions of the sensors and the sources that have generated the events. The only assumption made is that the emission time of the sound events must be known in order to measure the time of flight for each event. A closed form solution is also proposed whenever a sound event coincides with a microphone position. Simulated and real experiments show the validity of the approach for different setups of sensors and number of events. Marco Crocco, Alessio Del Bue, Matteo Bustreo, Vittorio Murino |
ICASSP | 2 |
| 2012 | Piecewise single view Photometric Stereo with multi-view constraintsabstractThis paper presents a novel Photometric Stereo approach for static views that recasts the problem into a piecewise formulation. The proposed algorithm, called Piecewise Photometric Stereo (PPS), entails several advantages in respect to previous global approaches. It is intrinsically more efficient, since reconstructing the surface in patches is computationally faster than reconstructing the global surface. Each patch has been associated an individual photometric model rather than a single global model as used in classical approaches. In this way, the piecewise formulation may grasp more complex lighting effects. Finally, the global metric properties of the shape is preserved using the multi-view constraints. In this pipeline, structure from motion is exploited to define such set of constraints and to compose a 3D mesh representing the metric structure of the object. Real results with ground truth show the positive performance of our algorithm compared with a classical global approach for Photometric Stereo. Reza Sabzevari, Alessio Del Bue, Vittorio Murino |
ICIP | 2 |
| 2012 | A joint structural and functional analysis of in-vitro neuronal networksabstractThe acquisition, analysis and representation of experimental data describing both anatomical and functional information at cellular level is an innovative opportunity to investigate neuronal network processing and organization. In this paper we propose an image processing pipeline to study in-vitro neuronal networks with a joint analysis of anatomy and electrophysiology. Neuronal nuclei are detected by segmenting fluorescence images of neuronal cultures. The high resolution Multi Electrode Arrays (MEAs) technology is used to collect functional information on cellular electrophysiological activity. Finally, detailed maps, representing both structural and functional information, are obtained which provide statistics on neuron distribution and spiking activity. Simona Ullo, Alessio Del Bue, Alessandro Maccione, Luca Berdondini, Vittorio Murino |
ICIP | 2 |
| 2012 | Optimal Metric Projections for Deformable and Articulated Structure-from-Motion
Marco Paladini, Alessio Del Bue, João M. F. Xavier, Lourdes Agapito, Marko Stosic, Marija Dodig |
Int. J. Comput. Vis. | 2 |
| 2012 | Bilinear Modeling via Augmented Lagrange Multipliers (BALM)abstractThis paper presents a unified approach to solve different bilinear factorization problems in computer vision in the presence of missing data in the measurements. The problem is formulated as a constrained optimization where one of the factors must lie on a specific manifold. To achieve this, we introduce an equivalent reformulation of the bilinear factorization problem that decouples the core bilinear aspect from the manifold specificity. We then tackle the resulting constrained optimization problem via Augmented Lagrange Multipliers. The strength and the novelty of our approach is that this framework can seamlessly handle different computer vision problems. The algorithm is such that only a projector onto the manifold constraint is needed. We present experiments and results for some popular factorization problems in computer vision such as rigid, non-rigid, and articulated Structure from Motion, photometric stereo, and 2D-3D non-rigid registration. Alessio Del Bue, João M. F. Xavier, Lourdes Agapito, Marco Paladini |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Social interaction discovery by statistical analysis of F-formationsabstractWe present a novel approach for detecting social interactions in a crowded scene by employing solely visual cues. The detection of social interactions in unconstrained scenarios is a valuable and important task, especially for surveillance purposes. Our proposal is inspired by the social signaling literature, and in particular it considers the sociological notion of F-formation. An F-formation is a set of possible configurations in space that people may assume while participating in a social interaction. Our system takes as input the positions of the people in a scene and their (head) orientations; then, employing a voting strategy based on the Hough transform, it recognizes F-formations and the individuals associated with them. Experiments on simulations and real data promote our idea. Marco Cristani, Loris Bazzani, Giulia Paggetti, Andrea Fossati, Diego Tosato, Alessio Del Bue, Gloria Menegaz, Vittorio Murino |
BMVC | 6 |
| 2011 | Multiview 3D warpsabstractImage registration and 3D reconstruction are fundamental computer vision and medical imaging problems. They are particularly challenging when the input data are images of a deforming body obtained by a single moving camera. We propose a new modelling framework, the multiview 3D warps. Existing models are twofold: they estimate inter-image warps which are often inconsistent between the different images and do not model the underlying 3D structure, or reconstruct just a sparse set of points. In contrast, our multiview 3D warps combine the advantages of both; they have an explicit 3D component and a set of 3D deformations combined with projection to 2D. They thus capture the dense deforming body's time-varying shape and camera pose. The advantages over the classical solutions are numerous: thanks to our feature-based estimation method for the multiview 3D warps, one can not only augment the original images but also retarget or clone the observed body's 3D deformations by changing the pose. Experimental results on simulated and real data are reported, confirming the advantages of our framework over existing methods. Alessio Del Bue, Adrien Bartoli |
ICCV | 1 |
| 2011 | A Method for Asteroids 3D Surface Reconstruction from Close Approach Distances
Luca Baglivo, Alessio Del Bue, Massimo Lunardelli, Francesco Setti, Vittorio Murino, Mariolino De Cecco |
ICVS | 2 |
| 2011 | An Experimental Framework for Evaluating PTZ Tracking Algorithms
Pietro Salvagnini, Marco Cristani, Alessio Del Bue, Vittorio Murino |
ICVS | 3 |
| 2011 | Simultaneous motion segmentation and Structure from MotionabstractThis paper presents a novel approach to simultaneously compute the motion segmentation and the 3D reconstruction of a set of 2D points extracted from an image sequence. Starting from an initial segmentation, our method proposes an iterative procedure that corrects the misclassified points while reconstructing the 3D scene, which is composed of objects that move independently. This optimization procedure is made by considering two well-known principles: firstly, in multi-body Structure from Motion the matrix describing the 3D shape is sparse, secondly, the segmented 2D points must give a valid 3D reconstruction given the rotational metric constraints. Our formulation results in a bilinear optimization where sparsity and metric constraints are enforced at each iteration of the algorithm. The final result is the corrected segmentation, the 3D structure of the moving objects and an orthographic camera matrix for each motion and each frame. Results are shown on synthetic sequences and a preliminary application on real sequences of the Hopkins 155 database is presented. Luca Zappella, Alessio Del Bue, Xavier Lladó, Joaquim Salvi |
WACV | 2 |
| 2011 | Reconstruction of non-rigid 3D shapes from stereo-motion
Xavier Lladó, Alessio Del Bue, Arnau Oliver, Joaquim Salvi, Lourdes Agapito |
Pattern Recognit. Lett. | 2 |
| 2010 | Adaptive Metric Registration of 3D Models to Non-rigid Image Trajectories
Alessio Del Bue |
ECCV (3) | 1 |
| 2010 | Bilinear Factorization via Augmented Lagrange Multipliers
Alessio Del Bue, João M. F. Xavier, Lourdes Agapito, Marco Paladini |
ECCV (4) | 1 |
| 2010 | Piecewise Quadratic Reconstruction of Non-Rigid Surfaces from Monocular Sequences
João Fayad, Lourdes Agapito, Alessio Del Bue |
ECCV (4) | 3 |
| 2010 | Non-rigid metric reconstruction from perspective cameras
Xavier Lladó, Alessio Del Bue, Lourdes Agapito |
Image Vis. Comput. | 2 |
| 2009 | Non-rigid Structure from Motion using Quadratic Deformation ModelsabstractIn this paper we present a new approach to the modelling of non-rigid 3D surfaces from the observation of 2D motion in images captured by an orthographic camera. Our aim is to characterize strong variations of the shape due, for instance, to bending motions. Such motions are hard to describe with previously used deformation models, such as the linear basis shapes model, which would tend to overestimate the dimensionality of the deformable data. Our approach uses a quadratic deformation model which is able to represent non-linear non-rigid motions such as bending, stretching, shearing and twisting. The model is bilinear and thus fits easily into previous schemes for Non-Rigid Structure from Motion (NRSfM). We formulate the NRSfM problem using a non-linear optimization scheme to minimize image reprojection error and recover the camera parameters, the 3D shape at rest and the quadratic deformation transformations. Our experiments with synthetic and real data show examples in which methods based on the linear basis shape model perform poorly or do not converge and instead the quadratic model is able to achieve accurate 3D reconstructions. © 2009. The copyright of this document resides with its authors. João Fayad, Alessio Del Bue, Lourdes Agapito, Pedro Aguiar |
BMVC | 2 |
| 2009 | Factorization for non-rigid and articulated structure using metric projectionsabstractThis paper describes a new algorithm for recovering the 3D shape and motion of deformable and articulated objects purely from uncalibrated 2D image measurements using an iterative factorization approach. Most solutions to non-rigid and articulated structure from motion require metric constraints to be enforced on the motion matrix to solve for the transformation that upgrades the solution to metric space. While in the case of rigid structure the metric upgrade step is simple since the motion constraints are linear, deformability in the shape introduces non-linearities. In this paper we propose an alternating least-squares approach associated with a globally optimal projection step onto the manifold of metric constraints. An important advantage of this new algorithm is its ability to handle missing data which becomes crucial when dealing with real video sequences with self-occlusions. We show successful results of our algorithms on synthetic and real sequences of both deformable and articulated data. Marco Paladini, Alessio Del Bue, Marko Stosic, Marija Dodig, João M. F. Xavier, Lourdes Agapito |
CVPR | 2 |
| 2009 | 2D-3D registration of deformable shapes with manifold projection Alessio Del BueabstractWe present an algorithm able to register a known 3D deformable model to a set of 2D matched points extracted from a single image. Unlike previous approaches, the problem is solved simultaneously for both the rigid and non-rigid parameters of the model. The key advantage of our approach is the projection of an initial affine estimation of the motion parameters into the motion manifold corresponding to the exact parametrization of the problem. This projection is formulated as theminimization of the distance between the affine solution and the surface of the manifold. Such optimization results in a quadratically constrained quadratic minimization problem that can be efficiently solved with standard optimization tools. Synthetic and real tests demonstrate the effectiveness of the approach. Alessio Del Bue, Marko Stosic, Marija Dodig, João M. F. Xavier |
ICIP | 1 |
| 2008 | A Model of Brightness Variations Due to Illumination Changes and Non-rigid Motion Using Spherical HarmonicsabstractPixel brightness variations in an image sequence depend both on the objects ‘surface reflectance and on the motion of the camera and object. In the case of rigid shapes some proposed models have been very successful explaining the relation among these strongly coupled components. On the other hand, shapes which deform pose new challenges since the relation between pixel brightness variation with non-rigid motion is not yet clear. In this paper, we introduce a new model which describes brightness variations with two independent components represented as linear basis shapes. Lighting influence is represented in terms of Spherical Harmonics and non-rigid motion as a linear model which represents image coordinates displacement. We then propose an efficient procedure for the estimation of this image model in two distinct steps. First, shape normal’s and albedo are estimated using standard photometric stereo on a sequence with varying lighting and no deformable motion. Then, given the knowledge of the object’s shape normal’s and albedo, we efficiently compute the 2D coordinates bases by minimizing image pixel residuals over an image sequence with constant lighting and only non-rigid motion. Experiments on real tests show the effectiveness of our approach in a face modelling context. José Miguel Buenaposada, Alessio Del Bue, Enrique Muñoz, Luis Baumela |
BMVC | 2 |
| 2008 | A factorization approach to structure from motion with shape priorsabstractThis paper presents an approach for including 3D prior models into a factorization framework for structure from motion. The proposed method computes a closed-form affine fit which mixes the information from the data and the 3D prior on the shape structure. Moreover, it is general in regards to different classes of objects treated: rigid, articulated and deformable. The inclusion of the shape prior may aid the inference of camera motion and 3D structure components whenever the data is degenerate (i.e. nearly planar motion of the projected shape). A final non-linear optimization stage, which includes the shape priors as a quadratic cost, upgrades the affine fit to metric. Results on real and synthetic image sequences, which present predominant degenerate motion, make clear the improvements over the 3D reconstruction. Alessio Del Bue |
CVPR | 1 |
| 2008 | Recovering Euclidean deformable models from stereo-motionabstractIn this paper we present a novel Structure from Motion (SfM) approach able to infer 3D deformable models from uncalibrated stereo images. Using a stereo setup dramatically improves the 3D model estimation when the observed 3D shape is mostly deforming without undergoing strong rigid motion. Our approach first calibrates the stereo system automatically and then computes a single metric rigid structure for each frame. Afterwards, these 3D shapes are aligned to a reference view using a RANSAC method in order to compute the mean shape of the object and to select the subset of points on the object which have remained rigid throughout the sequence without deforming. The selected rigid points are then used to compute frame-wise shape registration and to extract the motion parameters robustly from frame to frame. Finally, all this information is used in a global optimization stage with bundle adjustment which allows to refine the frame-wise initial solution and also to recover the non-rigid 3D model. We show results on synthetic and real data that prove the performance of the proposed method even when there is no rigid motion in the original sequence. Xavier Lladó, Alessio Del Bue, Lourdes Agapito |
ICPR | 2 |
| 2007 | Non-rigid structure from motion using ranklet-based tracking and non-linear optimization
Alessio Del Bue, Fabrizio Smeraldi, Lourdes Agapito |
Image Vis. Comput. | 1 |
| 2006 | Non-Rigid Metric Shape and Motion Recovery from Uncalibrated Images Using PriorsabstractIn this paper we focus on the estimation of the 3D Euclidean shape and motion of a non-rigid object which is moving rigidly while deforming and is observed by a perspective camera. Our method exploits the fact that it is often a reasonable assumption that some of the points are deforming throughout the sequence while others remain rigid. First we use an automatic segmentation algorithm to identify the set of rigid points which in turn is used to estimate the internal camera calibration parameters and the overall rigid motion. Finally we formalise the problem of non-rigid shape estimation as a constrained non-linear minimization adding priors on the degree of deformability of each point. We perform experiments on synthetic and real data which show firstly that even when using a minimal set of rigid points it is possible to obtain reliable metric information and secondly that the shape priors help to disambiguate the contribution to the image motion caused by the deformation and the perspective distortion. Alessio Del Bue, Xavier Lladó, Lourdes Agapito |
CVPR (1) | 1 |
| 2006 | Non-Rigid Stereo Factorization
Alessio Del Bue, Lourdes Agapito |
Int. J. Comput. Vis. | 1 |
| 2005 | Non-rigid 3D Factorization for Projective ReconstructionabstractIn this paper we address the problem of projective reconstruction for deformable objects. Recent work in non-rigid factorization has proved that it is possible to model deformations as a linear combination of basis shapes, allowing the recovery of camera motion and 3D shape under weak perspective viewing conditions. However, the performance of these methods degrades when the object of interest is close to the camera and strong perspective distortion is present in the data. The main contribution of this work is the proposal of a practical method for the recovery of projective depths, camera motion and non-rigid 3D shape from a sequence of images under strong perspective conditions. Our approach is based on minimizing 2D reprojection errors, solving the minimization as four weighted least squares problems. Results using synthetic and real data are given to illustrate the performance of our method. 1 Xavier Lladó, Alessio Del Bue, Lourdes Agapito |
BMVC | 2 |
| 2005 | Tracking points on deformable objects with rankletsabstractWe present a robust algorithm for point tracking on deformable objects. The key elements are the use of orientation selective rank features (ranklets), local filter adaptation and dynamic model update. A multi-scale vector of ranklets is used to encode a neighbourhood of each tracked point. The shape of the filters is optimised for each neighbourhood independently. Substantial appearance variations are catered for by maintaining a stack of models for each tracked point. This enables the system to recalibrate whenever the object reverts to its original appearance. Fabrizio Smeraldi, Alessio Del Bue, Lourdes Agapito |
ICIP (3) | 2 |
| 2002 | Multivariate Saddle Point Detection for Statistical Clustering
Dorin Comaniciu, Visvanathan Ramesh, Alessio Del Bue |
ECCV (3) | 3 |
| 2002 | Smart cameras with real-time video object generationabstractThe paper presents a system for video object generation and selective encoding with applications in surveillance, mobile videophones, and the automotive industry. Object tracking and MPEG-4 compression are performed in real-time. The system belongs to a new generation of intelligent vision sensors called smart cameras, which execute autonomous vision tasks and report events and data to a remote base-station. A detection module signals the presence of an object of interest within the camera field of view, while the tracking part follows the target to generate temporal trajectories. The compression is MPEG-4 compliant and implements the simple profile of the standard, which is capable of encoding up to four video objects. At the same time, the compression is selective, maintaining a higher quality for foreground objects and a lower quality for background representation. This property contributes to bandwidth reduction while preserving the essential information of foreground objects. The system performance is demonstrated in experiments that involve objects representing faces and vehicles seen from both static and moving cameras. Alessio Del Bue, Dorin Comaniciu, Visvanathan Ramesh, Carlo S. Regazzoni |
ICIP (3) | 1 |