Loris Bazzani

dblp:76/7404 · DBLP profile ↗
← Back
32ranked-venue papers
11as first author
5since 2021 · last 2025
0009-0003-1970-1085ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 8 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Learning Visual Hierarchies in Hyperbolic Space for Image Retrieval
Sameera Ramasinghe, Chenchen Hu, Julien Monteil, Loris Bazzani, Thalaiyasingam Ajanthan
ICCV5
2025 LATTECLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
abstract
Large-scale vision-language pre-trained (VLP) models (e.g., CLIP [46]) are renowned for their versatility, as they can be applied to diverse applications in a zero-shot setup. However, when these models are used in specific domains, their performance often falls short due to domain gaps or the under-representation of these domains in the training data. While fine-tuning VLP models on custom datasets with human-annotated labels can address this issue, annotating even a small-scale dataset (e.g., 100k samples) can be an expensive endeavor, often requiring expert annotators if the task is complex. To address these challenges, we propose LATTECLIP, an unsupervised method for fine-tuning CLIP models on classification with known class names in custom domains, without relying on human annotations. Our method leverages Large Multimodal Models (LMMs) to generate expressive textual descriptions for both individual images and groups of images. These provide additional contextual information to guide the fine-tuning process in the custom domains. Since LMM-generated descriptions are prone to hallucination or missing details, we introduce a novel strategy to distill only the useful information and stabilise the training. Specifically, we learn rich per-class prototype representations from noisy generated texts and dual pseudo-labels. Our experiments on 10 domain-specific datasets show that LATTECLIP outperforms pre-trained zero-shot methods by an average improvement of +4.74 points in top-1 accuracy and other state-of-the-art unsupervised methods by +3.45 points.
Anh-Quan Cao, Maximilian Jaritz, Matthieu Guillaumin, Raoul de Charette, Loris Bazzani
WACV5
2024 ViewFusion: Towards Multi-View Consistency via Interpolated Denoising
abstract
Novel-view synthesis through diffusion models has demonstrated remarkable potential for generating diverse and high-quality images. Yet, the independent process of image generation in these prevailing methods leads to challenges in maintaining multiple view consistency. To address this, we introduce ViewFusion, a novel, training-free algorithm that can be seamlessly integrated into existing pre-trained diffusion models. Our approach adopts an auto-regressive method that implicitly leverages previously generated views as context for next view generation, ensuring robust multi-view consistency during the novel-view generation process. Through a diffusion process that fuses known-view information via interpolated denoising, our framework successfully extends single-view conditioned models to work in multiple-view conditional settings without any additional fine-tuning. Extensive experimental results demonstrate the effectiveness of View Fusion in generating consistent and detailed novel views.
Xianghui Yang, Sameera Ramasinghe, Loris Bazzani, Gil Avraham, Anton van den Hengel
CVPR4
2021 Revamping Cross-Modal Recipe Retrieval With Hierarchical Transformers and Self-Supervised Learning
abstract
Cross-modal recipe retrieval has recently gained substantial attention due to the importance of food in people’s lives, as well as the availability of vast amounts of digital cooking recipes and food images to train machine learning models. In this work, we revisit existing approaches for cross-modal recipe retrieval and propose a simplified end-to-end model based on well established and high performing encoders for text and images. We introduce a hierarchical recipe Transformer which attentively encodes individual recipe components (titles, ingredients and instructions). Further, we propose a self-supervised loss function computed on top of pairs of individual recipe components, which is able to leverage semantic relationships within recipes, and enables training using both image-recipe and recipe-only samples. We conduct a thorough analysis and ablation studies to validate our design choices. As a result, our proposed method achieves state-of-the-art performance in the cross-modal recipe retrieval task on the Recipe1M dataset. We make code and models publicly available1.
Amaia Salvador, Erhan Gundogdu, Loris Bazzani, Michael Donoser
CVPR3
2021 Learning Attribute-driven Disentangled Representations for Interactive Fashion Retrieval
abstract
Interactive retrieval for online fashion shopping provides the ability to change image retrieval results according to the user feedback. One common problem in interactive retrieval is that a specific user interaction (e.g., changing the color of a T-shirt) causes other aspects to change inadvertently (e.g., the retrieved item has a sleeve type different than the query). This is a consequence of existing methods learning visual representations that are semantically entangled in the embedding space, which limits the controllability of the retrieved results. We propose to leverage on the semantics of visual attributes to train convolutional networks that learn attribute-specific subspaces for each attribute to obtain disentangled representations. Thus operations, such as swapping out a particular attribute value for another, impact the attribute at hand and leave others untouched. We show that our model can be tailored to deal with different retrieval tasks while maintaining its disentanglement property. We obtain state-of-the-art performance on three interactive fashion retrieval tasks: attribute manipulation retrieval, conditional similarity retrieval, and outfit complementary item retrieval. Code and models are publicly available1.
Yuxin Hou, Eleonora Vig, Michael Donoser, Loris Bazzani
ICCV4
2020 Image Search With Text Feedback by Visiolinguistic Attention Learning
abstract
Image search with text feedback has promising impacts in various real-world applications, such as e-commerce and internet search. Given a reference image and text feedback from user, the goal is to retrieve images that not only resemble the input image, but also change certain aspects in accordance with the given text. This is a challenging task as it requires the synergistic understanding of both image and text. In this work, we tackle this task by a novel Visiolin-guistic Attention Learning (VAL) framework. Specifically, we propose a composite transformer that can be seamlessly plugged in a CNN to selectively preserve and transform the visual features conditioned on language semantics. By inserting multiple composite transformers at varying depths, VAL is incentive to encapsulate the multi-granular visiolinguistic information, thus yielding an expressive representation for effective image search. We conduct comprehensive evaluation on three datasets: Fashion200k, Shoes and FashionIQ. Extensive experiments show our model exceeds existing approaches on all datasets, demonstrating consistent superiority in coping with various text feedbacks, including attribute-like and natural language descriptions.
Yanbei Chen, Shaogang Gong, Loris Bazzani
CVPR3
2020 Learning Joint Visual Semantic Matching Embeddings for Language-Guided Retrieval
Yanbei Chen, Loris Bazzani
ECCV (22)2
2017 Recurrent Mixture Density Network for Spatiotemporal Visual Attention
Loris Bazzani, Hugo Larochelle, Lorenzo Torresani
ICLR (Poster)1
2017 Image and Video Understanding in Big Data
abstract
An active object recognition system has the advantage of acting in the environment to capture images that are more suited for training and lead to better performance at test time. In this paper, we utilize deep convolutional neural networks for active object recognition by simultaneously predicting the object label and the next action to be performed on the object with the aim of improving recognition performance. We treat active object recognition as a reinforcement learning problem and derive the cost function to train the network for joint prediction of the object label and the action. A generative model of object similarities based on the Dirichlet distribution is proposed and embedded in the network for encoding the state of the system. The training is carried out by simultaneously minimizing the label and action prediction errors using gradient descent. We empirically show that the proposed network is able to predict both the object label and the actions on GERMS, a dataset for active object recognition. We compare the test label prediction accuracy of the proposed model with Dirichlet and Naive Bayes state encoding. The results of experiments suggest that the proposed model equipped with Dirichlet state encoding is superior in performance, and selects images that lead to better training and higher accuracy of label prediction at test time.
Vittorio Murino, Shaogang Gong, Chen Change Loy, Loris Bazzani
Comput. Vis. Image Underst.4
2016 Approximate Log-Hilbert-Schmidt Distances between Covariance Operators for Image Classification
abstract
This paper presents a novel framework for visual object recognition using infinite-dimensional covariance operators of input features, in the paradigm of kernel methods on infinite-dimensional Riemannian manifolds. Our formulation provides a rich representation of image features by exploiting their non-linear correlations, using the power of kernel methods and Riemannian geometry. Theoretically, we provide an approximate formulation for the Log-Hilbert-Schmidt distance between covariance operators that is efficient to compute and scalable to large datasets. Empirically, we apply our framework to the task of image classification on eight different, challenging datasets. In almost all cases, the results obtained outperform other state of the art methods, demonstrating the competitiveness and potential of our framework.
Hà Quang Minh, Marco San-Biagio, Loris Bazzani, Vittorio Murino
CVPR3
2016 Self-taught object localization with deep networks
abstract
This paper introduces self-taught object localization, a novel approach that leverages deep convolutional networks trained for whole-image recognition to localize objects in images without additional human supervision, i.e., without using any ground-truth bounding boxes for training. The key idea is to analyze the change in the recognition scores when artificially masking out different regions of the image. The masking out of a region that includes the object typically causes a significant drop in recognition score. This idea is embedded into an agglomerative clustering technique that generates self-taught localization hypotheses. Our object localization scheme outperforms existing proposal methods in both precision and recall for small number of subwindow proposals (e.g., on ILSVRC-2012 it produces a relative gain of 23.4% over the state-of-the-art for top-1 hypothesis). Furthermore, our experiments show that the annotations automatically-generated by our method can be used to train object detectors yielding recognition results remarkably close to those obtained by training on manually-annotated bounding boxes.
Loris Bazzani, Alessandro Bergamo, Dragomir Anguelov, Lorenzo Torresani
WACV1
2016 A Unifying Framework in Vector-valued Reproducing Kernel Hilbert Spaces for Manifold Regularization and Co-Regularized Multi-view Learning
abstract
This paper presents a general vector-valued reproducing kernel Hilbert spaces (RKHS) framework for the problem of learning an unknown functional dependency between a structured input space and a structured output space. Our formulation encompasses both Vector-valued Manifold Regularization and Co-regularized Multi- view Learning, providing in particular a unifying framework linking these two important learning approaches. In the case of the least square loss function, we provide a closed form solution, which is obtained by solving a system of linear equations. In the case of Support Vector Machine (SVM) classification, our formulation generalizes in particular both the binary Laplacian SVM to the multi-class, multi-view settings and the multi-class Simplex Cone SVM to the semi-supervised, multi-view settings. The solution is obtained by solving a single quadratic optimization problem, as in standard SVM, via the Sequential Minimal Optimization (SMO) approach. Empirical results obtained on the task of object recognition, using several challenging data sets, demonstrate the competitiveness of our algorithms compared with other state-of-the-art methods.
Hà Quang Minh, Loris Bazzani, Vittorio Murino
J. Mach. Learn. Res.2
2015 Joint Individual-Group Modeling for Tracking
abstract
We present a novel probabilistic framework that jointly models individuals and groups for tracking. Managing groups is challenging, primarily because of their nonlinear dynamics and complex layout which lead to repeated splitting and merging events. The proposed approach assumes a tight relation of mutual support between the modeling of individuals and groups, promoting the idea that groups are better modeled if individuals are considered and vice versa. This concept is translated in a mathematical model using a decentralized particle filtering framework which deals with a joint individual-group state space. The model factorizes the joint space into two dependent subspaces, where individuals and groups share the knowledge of the joint individual-group distribution. The assignment of people to the different groups (and thus group initialization, split and merge) is implemented by two alternative strategies: using classifiers trained beforehand on statistics of group configurations, and through online learning of a Dirichlet process mixture model, assuming that no training data is available before tracking. These strategies lead to two different methods that can be used on top of any person detector (simulated using the ground truth in our experiments). We provide convincing results on two recent challenging tracking benchmarks.
Loris Bazzani, Matteo Zanotto, Marco Cristani, Vittorio Murino
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Weighted bag of visual words for object recognition
abstract
Bag of Visual words (BoV) is one of the most successful strategy for object recognition, used to represent an image as a vector of counts using a learned vocabulary. This strategy assumes that the representation is built using patches that are either densely extracted or sampled from the images using feature detectors. However, the dense strategy captures also the noisy background information, whereas the feature detection strategy can lose important parts of the objects. In this paper we propose a solution in-between these two strategies, by densely extracting patches from the image, and weighting them accordingly to their salience. Intuitively, highly salient patches have an important role in describing an object, while those with low saliency are still taken with low emphasis, instead of discarding them. We embed this idea in the word encoding mechanism adopted in the BoV approaches. The technique is successfully applied to vector quantization and Fisher vector, on Caltech-101 and Caltech-256.
Marco San-Biagio, Loris Bazzani, Marco Cristani, Vittorio Murino
ICIP2
2013 Semi-supervised multi-feature learning for person re-identification
abstract
Person re-identification is probably the open challenge for low-level video surveillance in the presence of a camera network with non-overlapped fields of view. A large number of direct approaches has emerged in the last five years, often proposing novel visual features specifically designed to highlight the most discriminant aspects of people, which are invariant to pose, scale and illumination. On the other hand, learning-based methods are usually based on simpler features, and are trained on pairs of cameras to discriminate between individuals. In this paper, we present a method that joins these two ideas: given an arbitrary state-of-the-art set of features, no matter their number, dimensionality or descriptor, the proposed multi-class learning approach learns how to fuse them, ensuring that the features agree on the classification result. The approach consists of a semi-supervised multi-feature learning strategy, that requires at least a single image per person as training data. To validate our approach, we present results on different datasets, using several heterogeneous features, that set a new level of performance in the person re-identification problem.
Dario Figueira, Loris Bazzani, Hà Quang Minh, Marco Cristani, Alexandre Bernardino, Vittorio Murino
AVSS2
2013 Person re-identification with a PTZ camera: An introductory study
abstract
We present an introductory study that paves the way for a new kind of person re-identification, by exploiting a single Pan-Tilt-Zoom (PTZ) camera. PTZ devices allow to zoom on body regions, acquiring discriminative visual patterns that enrich the appearance description of an individual. This intuition has been translated into a statistical direct reidentification scheme, which collects two images for each probe subject: the first image captures the probe individual, focusing on the whole body; the second can be a zoomed body part (head, torso or legs) or another whole body image, and is the outcome of an action-selection mechanism, driven by feature selection principles. The validation of this technique is also explored: in order to allow repeatability, two novel multi-resolution benchmarks have been created. On these data, we demonstrate that our approach selects effective actions, by focusing on body portions which discriminate each subject. Moreover, we show that the proposed compound of two images overwhelms standard multi-shot descriptions, composed by many more pictures.
Pietro Salvagnini, Loris Bazzani, Marco Cristani, Vittorio Murino
ICIP2
2013 A unifying framework for vector-valued manifold regularization and multi-view learning
abstract
This paper presents a general vector-valued reproducing kernel Hilbert spaces (RKHS) formulation for the problem of learning an unknown functional dependency between a structured input space and a structured output space, in the Semi-Supervised Learning setting. Our formulation includes as special cases Vector-valued Manifold Regularization and Multi-view Learning, thus provides in particular a unifying framework linking these two important learning approaches. In the case of least square loss function, we provide a closed form solution with an efficient implementation. Numerical experiments on challenging multi-class categorization problems show that our multi-view learning formulation achieves results which are comparable with state of the art and are significantly better than single-view learning.
Hà Quang Minh, Loris Bazzani, Vittorio Murino
ICML (2)2
2013 Symmetry-driven accumulation of local features for human characterization and re-identification
Loris Bazzani, Marco Cristani, Vittorio Murino
Comput. Vis. Image Underst.1
2013 Social interactions by visual focus of attention in a three-dimensional environment
abstract
Abstract In human behaviour analysis, the visual focus of attention (VFOA) of a person is a very important cue. VFOA detection is difficult, though, especially in a unconstrained and crowded environment, typical of video surveillance scenarios. In this paper, we estimate the VFOA by defining the Subjective View Frustum, which approximates the visual field of a person in a three‐dimensional representation of the scene. This opens up to several intriguing behavioural investigations. In particular, we propose the Inter‐Relation Pattern Matrix, which suggests possible social interactions between the people present in a scene. Theoretical justifications and experimental results substantiate the validity and the goodness of the analysis performed.
Loris Bazzani, Marco Cristani, Diego Tosato, Michela Farenzena, Giulia Paggetti, Gloria Menegaz, Vittorio Murino
Expert Syst. J. Knowl. Eng.1
2012 Online Bayesian Non-parametrics for Social Group Detection
Matteo Zanotto, Loris Bazzani, Marco Cristani, Vittorio Murino
BMVC2
2012 Decentralized particle filter for joint individual-group tracking
abstract
In this paper, we address the task of tracking groups of people in surveillance scenarios. This is a major challenge in computer vision, since groups are structured entities, subjected to repeated split and merge events. Our solution is a joint individual-group tracking framework, inspired by a recent technique dubbed decentralized particle filtering. The proposed strategy factorizes the joint individual-group state space in two dependent subspaces where individuals and groups share the knowledge of the joint individual-group distribution. In practice, we establish a tight relation of mutual support between the modeling of individuals and that of groups, promoting the idea that groups are better tracked if individuals are considered, and viceversa. Extensive experiments on a published and novel dataset validate our intuition, opening up to many future developments.
Loris Bazzani, Marco Cristani, Vittorio Murino
CVPR1
2012 Joining feature-based and similarity-based pattern description paradigms for object detection
Samuele Martelli, Marco Cristani, Loris Bazzani, Diego Tosato, Vittorio Murino
ICPR3
2012 Conversationally-inspired stylometric features for authorship attribution in instant messaging
abstract
Authorship attribution (AA) aims at recognizing automatically the author of a given text sample. Traditionally applied to literary texts, AA faces now the new challenge of recognizing the identity of people involved in chat conversations. These share many aspects with spoken conversations, but AA approaches did not take it into account so far. Hence, this paper tries to fill the gap and proposes two novelties that improve the effectiveness of traditional AA approaches for this type of data: the first is to adopt features inspired by Conversation Analysis (in particular for turn-taking), the second is to extract the features from individual turns rather than from entire conversations. The experiments have been performed over a corpus of dyadic chat conversations (77 individuals in total). The performance in identifying the persons involved in each exchange, measured in terms of area under the Cumulative Match Characteristic curve, is 89.5%.
Marco Cristani, Giorgio Roffo, Cristina Segalin, Loris Bazzani, Alessandro Vinciarelli, Vittorio Murino
ACM Multimedia4
2012 Learning Where to Attend with Deep Architectures for Image Tracking
abstract
We discuss an attentional model for simultaneous object tracking and recognition that is driven by gaze data. Motivated by theories of perception, the model consists of two interacting pathways, identity and control, intended to mirror the what and where pathways in neuroscience models. The identity pathway models object appearance and performs classification using deep (factored)-restricted Boltzmann machines. At each point in time, the observations consist of foveated images, with decaying resolution toward the periphery of the gaze. The control pathway models the location, orientation, scale, and speed of the attended object. The posterior distribution of these states is estimated with particle filtering. Deeper in the control pathway, we encounter an attentional mechanism that learns to select gazes so as to minimize tracking uncertainty. Unlike in our previous work, we introduce gaze selection strategies that operate in the presence of partial information and on a continuous action space. We show that a straightforward extension of the existing approach to the partial information setting results in poor performance, and we propose an alternative method based on modeling the reward surface as a gaussian process. This approach gives good performance in the presence of partial information and allows us to expand the action space from a small, discrete set of fixation points to a continuous domain.
Misha Denil, Loris Bazzani, Hugo Larochelle, Nando de Freitas
Neural Comput.2
2012 Multiple-shot person re-identification by chromatic and epitomic analyses
Loris Bazzani, Marco Cristani, Alessandro Perina, Vittorio Murino
Pattern Recognit. Lett.1
2011 Custom Pictorial Structures for Re-identification
abstract
We propose a novel methodology for re-identification, based on Pictorial Structures (PS). Whenever face or other biometric information is missing, humans recognize an individual by selectively focusing on the body parts, looking for part-to-part correspondences. We want to take inspiration from this strategy in a re-identification context, using PS to achieve this objective. For single image re-identification, we adopt PS to localize the parts, extract and match their descriptors. When multiple images of a single individual are available, we propose a new algorithm to customize the fit of PS on that specific person, leading to what we call a Custom Pictorial Structure (CPS). CPS learns the appearance of an individual, improving the localization of its parts, thus obtaining more reliable visual characteristics for re-identification. It is based on the statistical learning of pixel attributes collected through spatio-temporal reasoning. The use of PS and CPS leads to state-of-the-art results on all the available public benchmarks, and opens a fresh new direction for research on re-identification.
Dong Seon Cheng, Marco Cristani, Michele Stoppa, Loris Bazzani, Vittorio Murino
BMVC4
2011 Social interaction discovery by statistical analysis of F-formations
abstract
We present a novel approach for detecting social interactions in a crowded scene by employing solely visual cues. The detection of social interactions in unconstrained scenarios is a valuable and important task, especially for surveillance purposes. Our proposal is inspired by the social signaling literature, and in particular it considers the sociological notion of F-formation. An F-formation is a set of possible configurations in space that people may assume while participating in a social interaction. Our system takes as input the positions of the people in a scene and their (head) orientations; then, employing a voting strategy based on the Hough transform, it recognizes F-formations and the individuals associated with them. Experiments on simulations and real data promote our idea.
Marco Cristani, Loris Bazzani, Giulia Paggetti, Andrea Fossati, Diego Tosato, Alessio Del Bue, Gloria Menegaz, Vittorio Murino
BMVC2
2011 Learning attentional policies for tracking and recognition in video with deep networks
Loris Bazzani, Nando de Freitas, Hugo Larochelle, Vittorio Murino, Jo-Anne Ting
ICML1
2010 Person re-identification by symmetry-driven accumulation of local features
abstract
In this paper, we present an appearance-based method for person re-identification. It consists in the extraction of features that model three complementary aspects of the human appearance: the overall chromatic content, the spatial arrangement of colors into stable regions, and the presence of recurrent local motifs with high entropy. All this information is derived from different body parts, and weighted opportunely by exploiting symmetry and asymmetry perceptual principles. In this way, robustness against very low resolution, occlusions and pose, viewpoint and illumination changes is achieved. The approach applies to situations where the number of candidates varies continuously, considering single images or bunch of frames for each individual. It has been tested on several public benchmark datasets (ViPER, iLIDS, ETHZ), gaining new state-of-the-art performances.
Michela Farenzena, Loris Bazzani, Alessandro Perina, Vittorio Murino, Marco Cristani
CVPR2
2010 Collaborative particle filters for group tracking
abstract
Tracking groups of people is a highly informative task in surveillance, and it represents a still open and little explored issue. In this paper, we propose a brand new framework for group tracking, that consists in two separate particle filters, one focusing on groups as atomic entities (the multi-group tracker), and the other modeling each individual separately (the multi-object tracker). The latter helps the multi-group tracker in better defining the nature of a group, evaluating the membership of each individual with respect to different groups, and allowing a robust management of the occlusions. The coupling of the two processes is theoretically founded due to the revision of the posterior distribution of the multi-group tracker with the statistics accumulated by the multi-object tracker. Experimental comparative results certify the goodness of the proposed technique.
Loris Bazzani, Marco Cristani, Vittorio Murino
ICIP1
2010 Multiple-Shot Person Re-identification by HPE Signature
abstract
In this paper, we propose a novel appearance-based method for person re-identification, that condenses a set of frames of the same individual into a highly informative signature, called Histogram Plus Epitome, HPE. It incorporates complementary global and local statistical descriptions of the human appearance, focusing on the overall chromatic content, via histograms representation, and on the presence of recurrent local patches, via epitome estimation. The matching of HPEs provides optimal performances against low resolution, occlusions, pose and illumination variations, defining novel state-of-the-art results on all the datasets considered.
Loris Bazzani, Marco Cristani, Alessandro Perina, Michela Farenzena, Vittorio Murino
ICPR1
2009 Online subjective feature selection for occlusion management in tracking applications
abstract
Most of the state-of-the-art tracking algorithms are prone to error when dealing with occlusions, especially when the involved moving objects are hardly discernible in appearance. In this paper, we propose a multi-object particle filtering tracking framework particularly suited to manage the occlusion problem. The presented solution consists in the introduction of a online subjective feature selection mechanism, which highlights and employs the most discriminant features characterizing a single object with respect to the neighbouring objects. The policy adopted fits formally in the observation step of the particle filtering process, it is effective and not computationally costly. Trials carried out on illustrative synthetic data and on recent challenging benchmark sequences report compelling performances and encourage further development of the technique.
Loris Bazzani, Marco Cristani, Manuele Bicego, Vittorio Murino
ICIP1