EDBT 2026 Demo / reviewers in the wild / expert
Niki Martinel
dblp:56/10105
· DBLP profile ↗
43ranked-venue papers
15as first author
15since 2021 · last 2026
0000-0002-6962-8643ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 19 · 8 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adapt-PEFT: Adaptive Parameter Efficient Fine Tuning for Underwater Image Enhancement
Sameer Malik, Niki Martinel |
ICPR (11) | 2 |
| 2026 | H3D-MarNet: Wavelet-Guided Dual-Path Learning for Metal Artifact Suppression and CT Modality Transformation for Radiotherapy Workflows
Mubashara Rehman, Niki Martinel, Michele Avanzo, Riccardo Spizzo, Christian Micheloni |
ICPR (10) | 2 |
| 2026 | Pyramidal anomaly detection with state-space modelsabstractRecent advancements in CNNs and transformer-based methods have significantly improved multiclass anomaly detection; however, accurately localizing small anomalies in industrial images remains challenging. Although promising, CNNs suffer in capturing long-range dependencies and the effectiveness of the transformers is hindered by their quadratic computational costs. This paper proposes a novel State Space Model (SSMs)-based multi-class anomaly detection method, termed Pyramidal Anomaly Detection with State-Space Models (PAD-SSM), which offers state-of-the-art performance with linear complexity. The main contribution of our work is the Pyramidal Scanning Approach (PSA), which performs the scanning hierarchically. This enables more localized and scale-aware analysis. PSA operates in a pyramid-like style, recursively partitions the image into equal non-overlapping patches across multiple scales. The SSM is then applied independently to each patch. This approach enables detailed, multi-scale, and high-resolution analysis, which is crucial for detecting small anomalies in industrial images. We integrate the PSA block with a pre-trained encoder for multi-scale feature extraction, a feature adapter to transform the encoded features to a target domain, and a synthetic anomaly generator that introduces noise at the feature level, further enhancing detection capabilities. Rigorous evaluations on the MVTec-AD and VisA datasets demonstrate that our method achieves state-of-the-art mean AU-ROC of 98.7, AP of 99.6, and F 1 max of 98.1 on MVTec-AD benchmark while using the lowest number of FLOPS (e.g., 10.3% fewer FLOPs than the next best method). Nasar Iqbal, Niki Martinel |
Comput. Vis. Image Underst. | 2 |
| 2025 | Sliced Vision Transformers for Fine-Grained Anomaly Detection and LocalizationabstractAccurate detection of small and narrow-shaped defects in industrial imaging is crucial for precisely identifying and localizing anomalies. Vision Transformer (ViT)-based image anomaly detection and localization networks have exhibited remarkable performance improvements in recent years, but their conventional square patch embedding may not be optimal for fine-grained anomaly detection, where anomalies typically exhibit elongated or irregular shapes. To address this problem, we introduce a novel approach that leverages sliced-shaped patches instead of conventional square patches in Vision Transformer (ViT). This approach improves the spatial resolution and ensures more detailed feature representations. State-of-the-art results on two existing industrial anomaly detection benchmarks show that our model effectively captures morphological details and spatial dependencies, thus demonstrating its ability to capture intricate anomaly patterns. Nasar Iqbal, Mattia Zanier, Marco Vernier, Christian Micheloni, Niki Martinel |
AVSS | 5 |
| 2025 | Nutrition Prediction from Food Images Using Foundation ModelsabstractWe consider the problem of nutrition prediction from food images. That is, given an RGB image of a dish composed of several food categories, the goal is to predict its mass, caloric intake, and the macronutrient masses such as fats, carbohydrates, and proteins. In this work, we first identify that the nutrition prediction problem can be considered a special case of a more general problem: the prediction of food category volumes. If such quantities are available (i.e., estimated), we can estimate the nutrition values using a linear mapping. Leveraging such a result, we propose a framework for nutrition prediction based on food volume estimation. It consists of a modular pipeline composed of (i) a food category recognition module, (ii) a semantic segmentation module, and (iii) a depth estimation module–which are common computer vision tasks. Finally, we study the zero-shot performance of vision-language foundation models applied for the aforementioned tasks. The code is provided at www.github.com/vitaly-emelianov/nutrition-fm. Vitalii Emelianov 0003, Niki Martinel |
ICME | 2 |
| 2025 | Pyramid-based Mamba Multi-class Unsupervised Anomaly DetectionabstractRecent advances in convolutional neural networks (CNNs) and transformer-based methods have improved anomaly detection and localization, but challenges persist in precisely localizing small anomalies. While CNNs face limitations in capturing long-range dependencies, transformer architectures often suffer from substantial computational overheads. We introduce a state space model (SSM)-based Pyramidal Scanning Strategy (PSS) for multi-class anomaly detection and localization-a novel approach designed to address the challenge of small anomaly localization. Our method captures fine-grained details at multiple scales by integrating the PSS with a pre-trained encoder for multi-scale feature extraction and a feature-level synthetic anomaly generator. An improvement of +1% AP for multi-class anomaly localization and a +1% increase in AU-PRO on MVTec benchmark demonstrate our method’s superiority in precise anomaly localization across diverse industrial scenarios. The code is available at https://github.com/iqbalmlpuniud/Pyramid_Mamba. Nasar Iqbal, Niki Martinel |
ICME | 2 |
| 2025 | Neural Additive Adapters for Interpretable Nutrition PredictionabstractWe study how large vision models (LVMs) can predict food nutrition through lightweight and interpretable adapters---the machine learning modules the predictions of which could be understood by humans. We introduce novel nutrition adapters that use features extracted by pre-trained LVMs and output the so-called nutrition maps. Nutrition maps indicate the concentration of nutrition values per each image location. We use such an interpretable representation to obtain the nutrition targets as a sum of all nutrition concentrations on the maps. To understand our approach's generalization capability, we systematically analyze the behavior of our novel interpretable adapters leveraging different LVMs with different food image-nutrition datasets. Our lightweight approach delivers better or on-par performance than the state-of-the-art models on the Nutrition5k and the Nutritionverse-Real benchmarks. The code is provided at https://github.com/vitaly-emelianov/nutrition-adapters. Vitalii Emelianov 0001, Niki Martinel |
ACM Multimedia | 2 |
| 2025 | CE-VAE: Capsule Enhanced Variational AutoEncoder for Underwater Image EnhancementabstractUnmanned underwater image analysis for marine monitoring faces two key challenges: (i) degraded image quality due to light attenuation and (ii) hardware storage constraints limiting high-resolution image collection. Existing methods primarily address image enhancement with approaches that hinge on storing the full-size input. In contrast, we introduce the Capsule Enhanced Variational AutoEncoder (CE-VAE), a novel architecture designed to efficiently compress and enhance degraded underwater images. Our attention-aware image encoder can project the input image onto a latent space representation while being able to run online on a remote device. The only information that needs to be stored on the device or sent to a beacon is a compressed representation. There is a dual-decoder module that performs offline, full-size enhanced image generation. One branch reconstructs spatial details from the compressed latent space, while the second branch utilizes a capsule-clustering layer to capture entity-level structures and complex spatial relationships. This parallel decoding strategy enables the model to balance fine-detail preservation with context-aware enhancements. CE- VAE achieves state-of-the-art performance in underwater image enhancement on six benchmark datasets, providing up to 3 × higher compression efficiency than existing approaches. Code available at https://github.com/iN1k1/ce-vae-underwater-image-enhancement. Rita Pucci, Niki Martinel |
WACV | 2 |
| 2024 | MAR-DTN: Metal Artifact Reduction Using Domain Transformation Network for Radiotherapy Planning
Belén Serrano-Antón, Mubashara Rehman, Niki Martinel, Michele Avanzo, Riccardo Spizzo, Giuseppe Fanetti, Alberto P. Muñuzuri, Christian Micheloni |
ICPR (11) | 3 |
| 2024 | Tracking Skiers from the Top to the BottomabstractSkiing is a popular winter sport discipline with a long history of competitive events. In this domain, computer vision has the potential to enhance the understanding of athletes’ performance, but its application lags behind other sports due to limited studies and datasets. This paper makes a step forward in filling such gaps. A thorough investigation is performed on the task of skier tracking in a video capturing his/her complete performance. Obtaining continuous and accurate skier localization is preemptive for further higher-level performance analyses. To enable the study, the largest and most annotated dataset for computer vision in skiing, SkiTB, is introduced. Several visual object tracking algorithms, including both established methodologies and a newly introduced skier-optimized baseline algorithm, are tested using the dataset. The results provide valuable insights into the applicability of different tracking methods for vision-based skiing analysis. SkiTB, code, and results are available at https://machinelearning.uniud.it/datasets/skitb. Matteo Dunnhofer, Luca Sordi, Niki Martinel, Christian Micheloni |
WACV | 3 |
| 2024 | Lightweight Prompt Learning Implicit Degradation Estimation Network for Blind Super ResolutionabstractBlind image super-resolution (SR) aims to recover a high-resolution (HR) image from its low-resolution (LR) counterpart under the assumption of unknown degradations. Many existing blind SR methods rely on supervising ground-truth kernels referred to as explicit degradation estimators. However, it is very challenging to obtain the ground-truths for different degradations kernels. Moreover, most of these methods rely on heavy backbone networks, which demand extensive computational resources. Implicit degradation estimators do not require the availability of ground truth kernels, but they see a significant performance gap with the explicit degradation estimators due to such missing information. We present a novel approach that significantly narrows such a gap by means of a lightweight architecture that implicitly learns the degradation kernel with the help of a novel loss component. The kernel is exploited by a learnable Wiener filter that performs efficient deconvolution in the Fourier domain by deriving a closed-form solution. Inspired by prompt-based learning, we also propose a novel degradation-conditioned prompt layer that exploits the estimated kernel to drive the focus on the discriminative contextual information that guides the reconstruction process in recovering the latent HR image. Extensive experiments under different degradation settings demonstrate that our model, named PL-IDENet, yields PSNR and SSIM improvements of more than 0.4dB and 1.3%, and 1.4dB and 4.8% to the best implicit and explicit blind-SR method, respectively. These results are achieved while maintaining a substantially lower number of parameters/FLOPs (i.e., 25% and 68% fewer parameters than best implicit and explicit methods, respectively). Asif Hussain Khan, Christian Micheloni, Niki Martinel |
IEEE Trans. Image Process. | 3 |
| 2022 | Pro-CCaps: Progressively Teaching Colourisation to CapsulesabstractAutomatic image colourisation studies how to colourise greyscale images. Existing approaches exploit convolutional layers that extract image-level features learning the colourisation on the entire image, but miss entities-level ones due to pooling strategies. We believe that entity-level features are of paramount importance to deal with the intrinsic multimodality of the problem (i.e., the same object can have different colours, and the same colour can have different properties). Models based on capsule layers aim to identify entity-level features in the image from different points of view, but they do not keep track of global features.Our network architecture integrates entity-level features into the image-level features to generate a plausible image colourisation. We observed that results obtained with direct integration of such two representations are largely dominated by the image-level features, thus resulting in unsaturated colours for the entities. To limit such an issue, we propose a gradual growth of the reconstruction phase of the model while training. By advantaging of prior knowledge from each growing step, we obtain a stable collaboration between image-level and entity-level features that ultimately generates stable and vibrant colourisations. Experimental results on three benchmark datasets, and a user study, demonstrate that our approach has competitive performance with respect to the state-of-the-art and provides more consistent colourisation. Rita Pucci, Christian Micheloni, Gian Luca Foresti, Niki Martinel |
WACV | 4 |
| 2022 | Consistent attentive dual branch network for person re-identificationabstractAbstract Several recent person re-identification methods are focusing on learning discriminative representations by designing efficient metric learning loss functions. Other approaches design part based architectures to compute an informative descriptor based on local features from semantically coherent parts. Few efforts learn the relationship between distant similar regions and parts by adjusting them to their most feasible positions with the help of soft attention. However, they focus on calibrating distant similar parts features and ignore to learn the noise (blur) free and distinct feature representations as the person re-identification datasets contain degraded images. To tackle these issues, we propose a novel Consistent Attention Dual Branch Network (CadNet) that has ability to model long-range dependencies (correlations) between channels as well as feature maps. We adopt multiple classifiers trained to learn the most discriminative global features for a unique representation of a person. Correlation between channels are consistently computed by using channel attention mechanism to make the learned feature noise free and distict from noisy and blurry data. Feature correlations interpret the relationship between distant similarities in the images computed by the self attention mechanism. The proposed CadNet significantly enhances the performance with respect to the baseline on the person re-identification benchmarks. Asad Munir, Niki Martinel, Christian Micheloni |
Multim. Tools Appl. | 2 |
| 2022 | Lord of the Rings: Hanoi Pooling and Self-Knowledge Distillation for Fast and Accurate Vehicle ReidentificationabstractVehicle reidentification has seen increasing interest, thanks to its fundamental impact on intelligent surveillance systems and smart transportation. The visual data acquired from monitoring camera networks come with severe challenges, including occlusions, color and illumination changes, as well as orientation issues (a vehicle can be seen from the side/front/rear due to different camera viewpoints). To deal with such challenges, the community has spent much effort in learning robust feature representations that hinge on additional visual attributes and part-driven methods, but with the side effects of requiring extensive human annotation labor as well as increasing computational complexity. In this article, we propose an approach that learns a feature representation robust to vehicle orientation issues without the need for extra-labeled data and adding negligible computational overheads. The former objective is achieved through the introduction of a Hanoi pooling layer exploiting ring regions and the image pyramid approach yielding a multiscale representation of vehicle appearance. The latter is tackled by transferring the accuracy of a deep network to its first layers, thus reducing the inference effort by the early stop of a test example. This is obtained by means of a self-knowledge distillation framework encouraging multiexit network decisions to agree with each other. Results demonstrate that the proposed approach significantly improves the accuracy of early (i.e., very fast) exits while maintaining the same accuracy of a deep (slow) baseline. Moreover, our solution obtains the best existing performance on three benchmark datasets.11[Online]. Available:https://github.com/iN1k1/. Niki Martinel, Matteo Dunnhofer, Rita Pucci, Gian Luca Foresti, Christian Micheloni |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Oriented Splits Network to Distill Background for Vehicle Re-IdentificationabstractVehicle re-identification (re-id) is a challenging task due to the presence of high intra-class and low inter-class variations in the visual data acquired from monitoring camera networks. Unique and discriminative feature representations are needed to overcome the existence of several variations including color, illumination, orientation, background and occlusion. The orientations of the vehicles in the images make the learned models unable to learn multiple parts of the vehicle and relationship between them. The combination of global and partial features is one of the solutions to improve the discriminative learning of deep learning models. Leveraging on such solutions, we propose an Oriented Splits Network (OSN) for an end to end learning of multiple features along with global features to form a strong descriptor for vehicle re-identification. To capture the orientation variability of the vehicles, the proposed network introduces a partition of the images into several oriented stripes to obtain local descriptors for each part/region. Such a scheme is therefore exploited by a camera based feature distillation (CBD) training strategy to remove the background features. These are filtered out from oriented vehicles representations which yield to a much stronger unique representation of the vehicles. We perform experiments on two benchmark vehicle re-id datasets to verify the performance of the proposed approach which show that the proposed solution achieves better result with respect to the state of the art with margin. Asad Munir, Niki Martinel, Christian Micheloni |
AVSS | 2 |
| 2020 | Tracking-by-Trackers with a Distilled and Reinforced Model
Matteo Dunnhofer, Niki Martinel, Christian Micheloni |
ACCV (2) | 2 |
| 2020 | Multi Branch Siamese Network For Person Re-IdentificationabstractTo capture robust person features, learning discriminative, style and view invariant descriptors is a key challenge in person Re-Identification (re-id). Most deep Re-ID models learn single scale feature representation which are unable to grasp compact and style invariant representations. In this paper, we present a multi branch Siamese Deep Neural Network with multiple classifiers to overcome the above issues. The multi-branch learning of the network creates a stronger descriptor with fine-grained information from global features of a person. Camera to camera image translation is performed with generative adversarial network to generate diverse data and add style invariance in learned features. Experimental results on benchmark datasets demonstrate that the proposed method performs better than other state of the arts methods. Asad Munir, Niki Martinel, Christian Micheloni |
ICIP | 2 |
| 2020 | Self and Channel Attention Network for Person Re-IdentificationabstractRecent research has shown promising results for person re-identification by focusing on several trends. One is designing efficient metric learning loss functions such as triplet loss family to learn the most discriminative representations. The other is learning local features by designing part based architectures to form an informative descriptor from semantically coherent parts. Some efforts adjust distant outliers to their most similar positions by using soft attention and learn the relationship between distant similar features. However, only a few prior efforts focus on channel-wise dependencies and learn non-local sharp similar part features directly for the degraded data in the person re-identification task. In this paper, we propose a novel Self and Channel Attention Network (SCAN) to model long-range dependencies between channels and feature maps. We add multiple classifiers to learn discriminative global features by using classification loss. Self Attention (SA) module and Channel Attention (CA) module are introduced to model non-local and channel-wise dependencies in the learned features. Spectral normalization is applied to the whole network to stabilize the training process. Experimental results on the person re-identification benchmarks show the proposed components achieve significant improvement with respect to the baseline. Asad Munir, Niki Martinel, Christian Micheloni |
ICPR | 2 |
| 2020 | Fixed simplex coordinates for angular margin loss in CapsNetabstractA more stationary and discriminative embedding is necessary for robust classification of images. We focus our attention on the newel CapsNet model and we propose the angular margin loss function in composition with margin loss. We define a fixed classifier implemented with fixed weights vectors obtained by the vertex coordinates of a simplex polytope. The advantage of using simplex polytope is that we obtain the maximal symmetry for stationary features angularly centred. Each weight vector is to be considered as the centroid of a class in the dataset. The embedding of an image is obtained through the capsule network encoding phase, that is identified as digitcaps matrix. Based on the centroids from the simplex coordinates and the embedding from the model, we compute the angular distance between the image embedding and the centroid of the correspondent class of the image. We take this angular distance as angular margin loss. We keep the computation proposed for margin loss in the original architecture of CapsNet. We train the model to minimise the angular between the embedding and the centroid of the class and maximise the magnitude of the embedding for the predicted class. The experiments on different datasets demonstrate that the angular margin loss improves the capability of capsule networks with complex datasets. Rita Pucci, Christian Micheloni, Gian Luca Foresti, Niki Martinel |
ICPR | 4 |
| 2020 | Siam-U-Net: encoder-decoder siamese network for knee cartilage tracking in ultrasound images
Matteo Dunnhofer, Maria Antico, Fumio Sasazawa, Yu Takeda, Saskia Camps, Niki Martinel, Christian Micheloni, Gustavo Carneiro 0001, Davide Fontanarosa |
Medical Image Anal. | 6 |
| 2020 | Deep interactive encoding with capsule networks for image classification
Rita Pucci, Christian Micheloni, Gian Luca Foresti, Niki Martinel |
Multim. Tools Appl. | 4 |
| 2020 | Deep Pyramidal Pooling With Attention for Person Re-IdentificationabstractLearning discriminative, view-invariant and multi-scale representations of object appearance with different semantic levels is of paramount importance for person Re-Identification (ReID). Recently, the community has focused on learning deep Re-ID models to capture a single holistic representation. To improve the achieved results, additional visual attributes and object part-driven models have been considered, inevitably introducing additional human annotation labor or computational efforts. In this paper, we argue that pyramid-inspired methods capturing multi-scale information may overcome such requirements. Precisely, multi-scale pooled regions representing visual information of an object are integrated within a novel deep architecture factorizing them into discriminative features at multiple semantic levels. These are exploited through an attention mechanism later considered in an identification-similarity multi-task loss, trained by means of a curriculum learning strategy. Extensive results on three person ReID benchmarks demonstrate that better performance than existing methods are achieved. Code is available at https://github.com/iN1k1. Niki Martinel, Gian Luca Foresti, Christian Micheloni |
IEEE Trans. Image Process. | 1 |
| 2020 | A UAV Video Dataset for Mosaicking and Change Detection From Low-Altitude FlightsabstractIn recent years, the technology of small-scale unmanned aerial vehicles (UAVs) has steadily improved in terms of flight time, automatic control, and image acquisition. This has lead to the development of several applications for low-altitude tasks, such as vehicle tracking, person identification, and object recognition. These applications often require to stitch together several video frames to get a comprehensive view of large areas (mosaicking), or to detect differences between images or mosaics acquired at different times (change detection). However, the datasets used to test mosaicking and change detection algorithms are typically acquired at high-altitudes, thus ignoring the specific challenges of low-altitude scenarios. The purpose of this paper is to fill this gap by providing the UAV mosaicking and change detection dataset. It consists of 50 challenging aerial video sequences acquired at low-altitude in different environments with and without the presence of vehicles, persons, and objects, plus metadata and telemetry. In addition, this paper provides some performance metrics to evaluate both the quality of the obtained mosaics and the correctness of the detected changes. Finally, the results achieved by two baseline algorithms, one for mosaicking and one for detection, are presented. The aim is to provide a shared performance reference that can be used for comparison with future algorithms that will be tested on the dataset. Danilo Avola, Luigi Cinque, Gian Luca Foresti, Niki Martinel, Daniele Pannone, Claudio Piciarelli |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2019 | From person to group re-identification via unsupervised transfer of sparse features
Giuseppe Lisanti, Niki Martinel, Christian Micheloni, Alberto Del Bimbo, Gian Luca Foresti |
Image Vis. Comput. | 2 |
| 2019 | Distributed person re-identification through network-wise rank fusion consensus
Niki Martinel, Gian Luca Foresti, Christian Micheloni |
Pattern Recognit. Lett. | 1 |
| 2018 | Wide-Slice Residual Networks for Food RecognitionabstractImage-based food recognition pose new challenges for mainstream computer vision algorithms. Recent works in the field focused either on hand-crafted representations or on learning these by exploiting deep neural networks (DNN). Despite the success of DNN-based works, these exploit off-the-shelf deep architectures which are not cast to the specific food classification problem. We believe that better results can be obtained if the architecture is defined with respect to an analysis of the food composition. Following such an intuition, this work introduces a new deep scheme that is designed to handle the food structure. In particular, we focus on the vertical food traits that are common to a large number of categories (i.e., 15% of the whole data in current datasets). Towards the final objective, we first introduce a slice convolution block to capture such specific information. Then, we leverage on the recent success of deep residual blocks and combine those with the sliced convolution to produce the classification score. Extensive evaluations on three benchmark datasets demonstrated that our solution has better performance than existing approaches (e.g., a top-1 accuracy of 90.27% on the Food-101 dataset). Niki Martinel, Gian Luca Foresti, Christian Micheloni |
WACV | 1 |
| 2018 | Accelerated low-rank sparse metric learning for person re-identification
Niki Martinel |
Pattern Recognit. Lett. | 1 |
| 2017 | Aerial video surveillance system for small-scale UAV environment monitoringabstractChange detection algorithms are commonly used to detect novelties for surveillance purposes in public and private places equipped by static or Pan-Tilt-Zoom (PTZ) cameras. Often, these techniques are also used as prerequisite to support more complex algorithms, including event recognition, object classification, person re-identification, and many others. With regard to small-scale Unmanned Aerial Vehicles (UAVs) at low-altitude, the change detection techniques require further investigation. In fact, most of the works currently available in the literature process video sequences acquired at very high-altitude for large-scale operations, such as vegetation monitoring, mapping of buildings, and so on. In a wide range of application contexts that require, for example, frequent monitoring or high spatial resolution for detecting small objects, video sequences acquired at high-altitude are not suitable. This paper presents a change detection system based on histogram equalization and RGB-Local Binary Pattern (RGB-LBP) operator for monitoring of wide areas by small-scale UAVs at low-altitude. Extensive experimental results show the robustness of the proposed pipeline. These latter were performed by using challenging video sequences of the public UAV Mosaicking and Change Detection (UMCD) dataset and measured a set of well-known statistical metrics. Finally, a performance analysis of the proposed algorithm is also provided. Danilo Avola, Gian Luca Foresti, Niki Martinel, Christian Micheloni, Daniele Pannone, Claudio Piciarelli |
AVSS | 3 |
| 2017 | Group Re-identification via Unsupervised Transfer of Sparse Features EncodingabstractPerson re-identification is best known as the problem of associating a single person that is observed from one or more disjoint cameras. The existing literature has mainly addressed such an issue, neglecting the fact that people usually move in groups, like in crowded scenarios. We believe that the additional information carried by neighboring individuals provides a relevant visual context that can be exploited to obtain a more robust match of single persons within the group. Despite this, re-identifying groups of people compound the common single person re-identification problems by introducing changes in the relative position of persons within the group and severe self-occlusions. In this paper, we propose a solution for group re-identification that grounds on transferring knowledge from single person reidentification to group re-identification by exploiting sparse dictionary learning. First, a dictionary of sparse atoms is learned using patches extracted from single person images. Then, the learned dictionary is exploited to obtain a sparsity-driven residual group representation, which is finally matched to perform the re-identification. Extensive experiments on the i-LIDS groups and two newly collected datasets show that the proposed solution outperforms stateof-the-art approaches. Giuseppe Lisanti, Niki Martinel, Alberto Del Bimbo, Gian Luca Foresti |
ICCV | 2 |
| 2017 | Person Reidentification in a Distributed Camera Network FrameworkabstractPlenty of research has been conducted to obtain the best reidentification performance between a single camera-pairs. None of the current approaches has addressed the reidentification in a camera network by considering the network topology (i.e., the structure of the monitored environment). We introduce a distributed network person reidentification framework which introduces the following contributions. 1) a camera matching cost to measure the reidentification performance between nodes of the network and 2) a derivation of the distance vector algorithm which allows to learn the network topology thus to prioritize and limit the cameras inquired for the matching of the probe. Results on three benchmark datasets show that the network topology can be learned in an unsupervised fashion and network-wise reidentification performance improves. As a side effect, we obtain that the communication bandwidth usage is reduced. Niki Martinel, Gian Luca Foresti, Christian Micheloni |
IEEE Trans. Cybern. | 1 |
| 2017 | Discriminant Context Information Analysis for Post-Ranking Person Re-IdentificationabstractExisting approaches for person re-identification are mainly based on creating distinctive representations or on learning optimal metrics. The achieved results are then provided in the form of a list of ranked matching persons. It often happens that the true match is not ranked first but it is in the first positions. This is mostly due to the visual ambiguities shared between the true match and other "similar" persons. At the current state, there is a lack of a study of such visual ambiguities which limit the re-identification performance within the first ranks. We believe that an analysis of the similar appearances of the first ranks can be helpful in detecting, hence removing, such visual ambiguities. We propose to achieve such a goal by introducing an unsupervised post-ranking framework. Once the initial ranking is available, content and context sets are extracted. Then, these are exploited to remove the visual ambiguities and to obtain the discriminant feature space which is finally exploited to compute the new ranking. An in-depth analysis of the performance achieved on three public benchmark data sets support our believes. For every data set, the proposed method remarkably improves the first ranks results and outperforms the state-of-the-art approaches. Jorge García 0002, Niki Martinel, Alfredo Gardel Vicente, Ignacio Bravo Muñoz, Gian Luca Foresti, Christian Micheloni |
IEEE Trans. Image Process. | 2 |
| 2016 | Temporal Model Adaptation for Person Re-identification
Niki Martinel, Abir Das, Christian Micheloni, Amit K. Roy-Chowdhury |
ECCV (4) | 1 |
| 2016 | Distributed and Unsupervised Cost-Driven Person Re-IdentificationabstractThe problem of re-identify persons across single disjoint camera-pairs has received great attention from the community. Despite this, when the re-identification process has to be carried out on a large camera network a different approach has to be considered. In particular, existing approaches have neglected the importance of the network topology (i.e., the structure of the monitored environment) in such a process. To try filling such a gap, we propose a Distributed and Unsupervised Cost-Driven Person Re-Identification framework (DUPRe) which introduces the following contributions: (i) a camera matching cost to measure the re-identification performance between nodes of the network; (ii) a derivation of the distance vector algorithm which allows to learn the network topology hence to prioritize and limit the cameras inquired for the re-identification. Results on two benchmark datasets show that our solution brings to significant network-wise re-identification improvements. Niki Martinel, Gian Luca Foresti, Christian Micheloni |
ICPR | 1 |
| 2016 | A supervised extreme learning committee for food recognition
Niki Martinel, Claudio Piciarelli, Christian Micheloni |
Comput. Vis. Image Underst. | 1 |
| 2016 | Modeling feature distances by orientation driven classifiers for person re-identification
Jorge García 0002, Niki Martinel, Alfredo Gardel Vicente, Ignacio Bravo Muñoz, Gian Luca Foresti, Christian Micheloni |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | A pool of multiple person re-identification expertsabstractThe person re-identification problem, i.e. recognizing a person across non-overlapping cameras at different times and locations, is of fundamental importance for video surveillance applications. Due to pose variations, illumination conditions, background clutter, and occlusions, re-identify a person is an inherently difficult problem which is still far from being solved. In this work, inspired by the recent police lineup innovations, we propose a re-identification approach where Multiple Re-identification Experts (MuRE) are trained to reliably match new probes. The answers from all the experts are then combined to achieve a final decision. The proposed method has been evaluated on three datasets showing significant improvements over state-of-the-art approaches. Niki Martinel, Christian Micheloni, Gian Luca Foresti |
Pattern Recognit. Lett. | 1 |
| 2015 | Person Re-Identification Ranking Optimisation by Discriminant Context Information AnalysisabstractPerson re-identification is an open and challenging problem in computer vision. Existing re-identification approaches focus on optimal methods for features matching (e.g., metric learning approaches) or study the inter-camera transformations of such features. These methods hardly ever pay attention to the problem of visual ambiguities shared between the first ranks. In this paper, we focus on such a problem and introduce an unsupervised ranking optimization approach based on discriminant context information analysis. The proposed approach refines a given initial ranking by removing the visual ambiguities common to first ranks. This is achieved by analyzing their content and context information. Extensive experiments on three publicly available benchmark datasets and different baseline methods have been conducted. Results demonstrate a remarkable improvement in the first positions of the ranking. Regardless of the selected dataset, state-of-the-art methods are strongly outperformed by our method. Jorge García 0002, Niki Martinel, Christian Micheloni, Alfredo Gardel Vicente |
ICCV | 2 |
| 2015 | Re-Identification in the Function Space of Feature WarpsabstractPerson re-identification in a non-overlapping multicamera scenario is an open challenge in computer vision because of the large changes in appearances caused by variations in viewing angle, lighting, background clutter, and occlusion over multiple cameras. As a result of these variations, features describing the same person get transformed between cameras. To model the transformation of features, the feature space is nonlinearly warped to get the "warp functions". The warp functions between two instances of the same target form the set of feasible warp functions while those between instances of different targets form the set of infeasible warp functions. In this work, we build upon the observation that feature transformations between cameras lie in a nonlinear function space of all possible feature transformations. The space consisting of all the feasible and infeasible warp functions is the warp function space (WFS). We propose to learn a discriminating surface separating these two sets of warp functions in the WFS and to re-identify persons by classifying a test warp function as feasible or infeasible. Towards this objective, a Random Forest (RF) classifier is employed which effectively chooses the warp function components according to their importance in separating the feasible and the infeasible warp functions in the WFS. Extensive experiments on five datasets are carried out to show the superior performance of the proposed approach over state-of-the-art person re-identification methods. We show that our approach outperforms all other methods when large illumination variations are considered. At the same time it has been shown that our method reaches the best average performance over multiple combinations of the datasets, thus, showing that our method is not designed only to address a specific challenge posed by a particular dataset. Niki Martinel, Abir Das, Christian Micheloni, Amit K. Roy-Chowdhury |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Classification of Local Eigen-Dissimilarities for Person Re-IdentificationabstractThe task of re-identifying a person that moves across cameras fields-of-view is a challenge to the community known as the person re-identification problem. State-of-the art approaches are either based on direct modeling and matching of the human appearance or on machine learning-based techniques. In this work we introduce a novel approach that studies densely localized image dissimilarities in a low dimensional space and uses those to re-identify between persons in a supervised classification framework. To achieve the goal: i) we compute the localized image dissimilarity between a pair of images; ii) we learn the lower dimensional space of such localized image dissimilarities, known as the “local eigen-dissimilarities” (LEDs) space; iii) we train a binary classifier to discriminate between LEDs computed for a positive pair (images are for a same person) from the ones computed for a negative pair (images are for different persons). We show the competitive performance of our approach on two publicly available benchmark datasets. Niki Martinel, Christian Micheloni |
IEEE Signal Process. Lett. | 1 |
| 2015 | Kernelized Saliency-Based Person Re-Identification Through Multiple Metric LearningabstractPerson re-identification in a non-overlapping multi-camera scenario is an open and interesting challenge. While the task can hardly be completed by machines, we, as humans, are inherently able to sample those relevant persons' details that allow us to correctly solve the problem in a fraction of a second. Thus, knowing where a human might fixate to recognize a person is of paramount interest for re-identification. Inspired by the human gazing capabilities, we want to identify the salient regions of a person appearance to tackle the problem. Toward this objective, we introduce the following main contributions. A kernelized graph-based approach is used to detect the salient regions of a person appearance, later used as a weighting tool in the feature extraction process. The proposed person representation combines visual features either considering or not the saliency. These are then exploited in a pairwise-based multiple metric learning framework. Finally, the non-Euclidean metrics that have been separately learned for each feature are fused to re-identify a person. The proposed kernelized saliency-based person re-identification through multiple metric learning has been evaluated on four publicly available benchmark data sets to show its superior performance over the state-of-the-art approaches (e.g., it achieves a rank 1 correct recognition rate of 42.41% on the VIPeR data set). Niki Martinel, Christian Micheloni, Gian Luca Foresti |
IEEE Trans. Image Process. | 1 |
| 2014 | Person Orientation and Feature Distances Boost Re-identificationabstractMost of the open challenges in person re-identification arise from the large variations of human appearance and from the different camera views that may be involved, making pure feature matching an unreliable solution. To tackle these challenges state-of-the-art methods assume that a unique inter-camera transformation of features undergoes between two cameras. However, the combination of view points, scene illumination and photometric settings, etc., together with the appearance, pose and orientation of a person make the inter-camera transformation of features multi-modal. To address these challenges we introduce three main contributions. We propose a method to extract multiple frames of the same person with different orientation. We learn the pair wise feature dissimilarities space (PFDS) formed by the subspace of pair wise feature dissimilarities computed between images of persons with similar orientation and the subspace of pair wise feature dissimilarities computed between images of persons non-similar orientations. Finally, a classifier is trained to capture the multi-modal inter-camera transformation of pair wise images for each subspace. To validate the proposed approach we show the superior performance of our approach to state-of-the-art methods using two publicly available benchmark datasets. Jorge García 0002, Niki Martinel, Gian Luca Foresti, Alfredo Gardel Vicente, Christian Micheloni |
ICPR | 2 |
| 2014 | Camera Selection for Adaptive Human-Computer InterfaceabstractVideo analytics has become a very important topic in computer vision. This paper introduces advanced video analytics human-computer interfaces for a video surveillance system to ease the tasks of security operators. The visualization of the most relevant views is provided by the human-computer interface module that preemptively activates cameras that will probably cover the motion of interesting objects. Human-computer interaction principles have been considered to develop the novel user interface. Four prototypes have been designed and usability performance has been evaluated, exploiting standard methods. Results obtained from such evaluations show the efficiency of the novel information visualization technique. Niki Martinel, Christian Micheloni, Claudio Piciarelli, Gian Luca Foresti |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2013 | Robust Painting Recognition and Registration for Mobile Augmented RealityabstractIn this work we introduce a novel approach for painting recognition and registration for mobile Augmented Reality applications. To address the challenges of real-time painting recognition and registration we introduce three main contributions: i) A relevant painting region detector extracts the painting region from the given image. ii) Two local and global features are extracted from the relevant region to robustly match a painting database. iii) A RANSAC homography estimation method is used to overlay the additional content in an AR framework. Experiments have been carried out on a dataset built with publicly available images. Niki Martinel, Christian Micheloni, Gian Luca Foresti |
IEEE Signal Process. Lett. | 1 |