Christian Micheloni

dblp:28/38 · DBLP profile ↗
← Back
76ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0003-4503-7483ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 32 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 H3D-MarNet: Wavelet-Guided Dual-Path Learning for Metal Artifact Suppression and CT Modality Transformation for Radiotherapy Workflows
Mubashara Rehman, Niki Martinel, Michele Avanzo, Riccardo Spizzo, Christian Micheloni
ICPR (10)5
2026 Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
abstract
Episodic memory retrieval enables wearable cameras to recall objects or events previously observed in video. However, existing formulations assume an "offline" setting with full video access at query time, limiting their applicability in real-world scenarios with power and storage-constrained wearable devices. Towards more application-ready episodic memory systems, we introduce Online Visual Query 2D (OVQ2D), a task where models process video streams online, observing each frame only once, and retrieve object localizations using a compact memory instead of full video history. We address OVQ2D with ESOM (Egocentric Streaming Object Memory), a novel framework integrating an object discovery module, an object tracking module, and a memory module that find, track, and store spatiotemporal object information for efficient querying. Experiments on Ego4D demonstrate ESOM’s superiority over other online approaches, though OVQ2D remains challenging, with top performance at only 4% success. ESOM’s accuracy increases markedly with perfect object tracking (31.91%), discovery (40.55%), or both (81.92%), underscoring the need of applied research on these components.
Zaira Manigrasso, Matteo Dunnhofer, Antonino Furnari, Moritz Nottebaum, Antonio Finocchiaro, Davide Marana, Rosario Forte, Giovanni Maria Farinella, Christian Micheloni
WACV9
2026 Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones
abstract
Abstract Vision backbone networks play a central role in modern computer vision. Enhancing their efficiency directly benefits a wide range of downstream applications. To measure efficiency, many publications rely on MACs (Multiply Accumulate operations) as a predictor of execution time. In this paper, we experimentally demonstrate the shortcomings of such a metric, especially in the context of edge devices. By contrasting the MAC count and execution time of common architectural design elements, we identify key factors for efficient execution and provide insights to optimize backbone design. Based on these insights, we present LowFormer, a novel vision backbone family. LowFormer features a streamlined macro and micro design that includes Lowtention, a lightweight alternative to Multi-Head Self-Attention. Lowtention not only proves more efficient, but also enables superior results on ImageNet. Additionally, we present an edge GPU version of LowFormer, that can further improve upon its baseline’s speed on edge GPU and desktop GPU. We demonstrate the LowFormer’s wide applicability by evaluating it on smaller image classification datasets, as well as adapting it to several downstream tasks, such as object detection, semantic segmentation, image retrieval, and visual object tracking. LowFormer models consistently achieve remarkable speed-ups across various hardware platforms compared to recent state-of-the-art backbones. Code and models are publicly available at https://github.com/altair199797/LowFormer.
Moritz Nottebaum, Matteo Dunnhofer, Christian Micheloni
Int. J. Comput. Vis.3
2025 Sliced Vision Transformers for Fine-Grained Anomaly Detection and Localization
abstract
Accurate detection of small and narrow-shaped defects in industrial imaging is crucial for precisely identifying and localizing anomalies. Vision Transformer (ViT)-based image anomaly detection and localization networks have exhibited remarkable performance improvements in recent years, but their conventional square patch embedding may not be optimal for fine-grained anomaly detection, where anomalies typically exhibit elongated or irregular shapes. To address this problem, we introduce a novel approach that leverages sliced-shaped patches instead of conventional square patches in Vision Transformer (ViT). This approach improves the spatial resolution and ensures more detailed feature representations. State-of-the-art results on two existing industrial anomaly detection benchmarks show that our model effectively captures morphological details and spatial dependencies, thus demonstrating its ability to capture intricate anomaly patterns.
Nasar Iqbal, Mattia Zanier, Marco Vernier, Christian Micheloni, Niki Martinel
AVSS4
2025 Is Tracking Really More Challenging in First Person Egocentric Vision?
abstract
Visual object tracking and segmentation are becoming fundamental tasks for understanding human activities in egocentric vision. Recent research has benchmarked state-of-the-art methods and concluded that first person egocentric vision presents challenges compared to previously studied domains. However, these claims are based on evaluations conducted across significantly different scenarios. Many of the challenging characteristics attributed to egocentric vision are also present in third person videos of human-object activities. This raises a critical question: how much of the observed performance drop stems from the unique first person viewpoint inherent to egocentric vision versus the domain of human-object activities? To address this question, we introduce a new benchmark study designed to disentangle such factors. Our evaluation strategy enables a more precise separation of challenges related to the first person perspective from those linked to the broader domain of human-object activity understanding. By doing so, we provide deeper insights into the true sources of difficulty in egocentric tracking and segmentation, facilitating more targeted advancements on this task.
Matteo Dunnhofer, Zaira Manigrasso, Christian Micheloni
ICCV3
2025 LowFormer: Hardware Efficient Design for Convolutional Transformer Backbones
abstract
Research in efficient vision backbones is evolving into models that are a mixture of convolutions and transformer blocks. A smart combination of both, architecture-wise and component-wise is mandatory to excel in the speed-accuracy trade-off. Most publications focus on maximizing accuracy and utilize MACs (multiply accumulate operations) as an efficiency metric. The latter however often do not measure accurately how fast a model actually is due to factors like memory access cost and degree of parallelism. We analyzed common modules and architectural design choices for backbones not in terms of MACs, but rather in actual throughput and latency, as the combination of the latter two is a better representation of the efficiency of models in real applications. We applied the conclusions taken from that analysis to create a recipe for increasing hardware-efficiency in macro design. Additionally we introduce a simple slimmed-down version of Multi-Head Self-Attention, that aligns with our analysis. We combine both macro and micro design to create a new family of hardware-efficient backbone networks called Low-Former. Low Former achieves a remarkable speedup in terms of throughput and latency, while achieving similar or better accuracy than current state-of-the-art efficient backbones. In order to prove the generalizability of our hardware-efficient design, we evaluate our method on GPU, mobile GPU and ARM CPU. We further show that the downstream tasks object detection and semantic segmentation profit from our hardware-efficient architecture. Code and models are available at https://github.com/altair199797/LowFormer.
Moritz Nottebaum, Matteo Dunnhofer, Christian Micheloni
WACV3
2024 MAR-DTN: Metal Artifact Reduction Using Domain Transformation Network for Radiotherapy Planning
Belén Serrano-Antón, Mubashara Rehman, Niki Martinel, Michele Avanzo, Riccardo Spizzo, Giuseppe Fanetti, Alberto P. Muñuzuri, Christian Micheloni
ICPR (11)8
2024 Tracking Skiers from the Top to the Bottom
abstract
Skiing is a popular winter sport discipline with a long history of competitive events. In this domain, computer vision has the potential to enhance the understanding of athletes’ performance, but its application lags behind other sports due to limited studies and datasets. This paper makes a step forward in filling such gaps. A thorough investigation is performed on the task of skier tracking in a video capturing his/her complete performance. Obtaining continuous and accurate skier localization is preemptive for further higher-level performance analyses. To enable the study, the largest and most annotated dataset for computer vision in skiing, SkiTB, is introduced. Several visual object tracking algorithms, including both established methodologies and a newly introduced skier-optimized baseline algorithm, are tested using the dataset. The results provide valuable insights into the applicability of different tracking methods for vision-based skiing analysis. SkiTB, code, and results are available at https://machinelearning.uniud.it/datasets/skitb.
Matteo Dunnhofer, Luca Sordi, Niki Martinel, Christian Micheloni
WACV4
2024 Visual tracking in camera-switching outdoor sport videos: Benchmark and baselines for skiing
abstract
Skiing is a globally popular winter sport discipline with a rich history of competitive events. This domain offers ample opportunities for the application of computer vision to enhance the understanding of athletes’ performances. However, this potential has remained relatively untapped in comparison to other sports, primarily due to the limited availability of dedicated research studies and datasets. The present paper takes a significant stride towards bridging these gaps. It conducts a comprehensive examination of skier appearance tracking in videos capturing their entire performance—an essential step for more advanced performance analyses. To implement this investigation, we introduce SkiTB, the largest and most annotated dataset tailored for computer vision applications in skiing. We subject a range of visual object tracking algorithms to rigorous testing, including both well-established methodologies and a novel skier-specific baseline algorithm. The results yield valuable insights into the suitability of various tracking techniques for vision-based skiing analysis and into the generalization of state-of-the-art algorithms to complex target behaviors and conditions set by winter outdoor environments. To foster further development, we make SkiTB, the associated code, and the obtained results accessible through https://machinelearning.uniud.it/datasets/skitb.
Matteo Dunnhofer, Christian Micheloni
Comput. Vis. Image Underst.2
2024 Lightweight Prompt Learning Implicit Degradation Estimation Network for Blind Super Resolution
abstract
Blind image super-resolution (SR) aims to recover a high-resolution (HR) image from its low-resolution (LR) counterpart under the assumption of unknown degradations. Many existing blind SR methods rely on supervising ground-truth kernels referred to as explicit degradation estimators. However, it is very challenging to obtain the ground-truths for different degradations kernels. Moreover, most of these methods rely on heavy backbone networks, which demand extensive computational resources. Implicit degradation estimators do not require the availability of ground truth kernels, but they see a significant performance gap with the explicit degradation estimators due to such missing information. We present a novel approach that significantly narrows such a gap by means of a lightweight architecture that implicitly learns the degradation kernel with the help of a novel loss component. The kernel is exploited by a learnable Wiener filter that performs efficient deconvolution in the Fourier domain by deriving a closed-form solution. Inspired by prompt-based learning, we also propose a novel degradation-conditioned prompt layer that exploits the estimated kernel to drive the focus on the discriminative contextual information that guides the reconstruction process in recovering the latent HR image. Extensive experiments under different degradation settings demonstrate that our model, named PL-IDENet, yields PSNR and SSIM improvements of more than 0.4dB and 1.3%, and 1.4dB and 4.8% to the best implicit and explicit blind-SR method, respectively. These results are achieved while maintaining a substantially lower number of parameters/FLOPs (i.e., 25% and 68% fewer parameters than best implicit and explicit methods, respectively).
Asif Hussain Khan, Christian Micheloni, Niki Martinel
IEEE Trans. Image Process.2
2023 Visual Object Tracking in First Person Vision
abstract
The understanding of human-object interactions is fundamental in First Person Vision (FPV). Visual tracking algorithms which follow the objects manipulated by the camera wearer can provide useful information to effectively model such interactions. In the last years, the computer vision community has significantly improved the performance of tracking algorithms for a large variety of target objects and scenarios. Despite a few previous attempts to exploit trackers in the FPV domain, a methodical analysis of the performance of state-of-the-art trackers is still missing. This research gap raises the question of whether current solutions can be used "off-the-shelf" or more domain-specific investigations should be carried out. This paper aims to provide answers to such questions. We present the first systematic investigation of single object tracking in FPV. Our study extensively analyses the performance of 42 algorithms including generic object trackers and baseline FPV-specific trackers. The analysis is carried out by focusing on different aspects of the FPV setting, introducing new performance measures, and in relation to FPV-specific tasks. The study is made possible through the introduction of TREK-150, a novel benchmark dataset composed of 150 densely annotated video sequences. Our results show that object tracking in FPV poses new challenges to current visual trackers. We highlight the factors causing such behavior and point out possible research directions. Despite their difficulties, we prove that trackers bring benefits to FPV downstream tasks requiring short-term object tracking. We expect that generic object tracking will gain popularity in FPV as new and FPV-specific methodologies are investigated. Supplementary Information: The online version contains supplementary material available at 10.1007/s11263-022-01694-6.
Matteo Dunnhofer, Antonino Furnari, Giovanni Maria Farinella, Christian Micheloni
Int. J. Comput. Vis.4
2022 Real Image Super-Resolution using GAN through modeling of LR and HR process
abstract
The current existing deep image super-resolution methods usually assume that a Low Resolution (LR) image is bicubicly downscaled of a High Resolution (HR) image. However, such an ideal bicubic downsampling process is different from the real LR degradations, which usually come from complicated combinations of different degradation processes, such as camera blur, sensor noise, sharpening artifacts, JPEG compression, and further image editing, and several times image transmission over the internet and unpredictable noises. It leads to the highly ill-posed nature of the inverse upscaling problem. To address these issues, we propose a GAN-based SR approach with learnable adaptive sinusoidal nonlinearities incorporated in LR and SR models by directly learn degradation distributions and then synthesize paired LR/HR training data to train the generalized SR model to real image degradations. We demonstrate the effectiveness of our proposed approach in quantitative and qualitative experiments.
Rao Muhammad Umer, Christian Micheloni
AVSS2
2022 CoCoLoT: Combining Complementary Trackers in Long-Term Visual Tracking
abstract
How to combine the complementary capabilities of an ensemble of different algorithms has been of central interest in visual object tracking. A significant progress on such a problem has been achieved, but considering short-term tracking scenarios. Instead, long-term tracking settings have been substantially ignored by the solutions. In this paper, we explicitly consider long-term tracking scenarios and provide a framework, named CoCoLoT, that combines the characteristics of complementary visual trackers to achieve enhanced long-term tracking performance. CoCoLoT perceives whether the trackers are following the target object through an online learned deep verification model, and accordingly activates a decision policy which selects the best performing tracker as well as it corrects the performance of the failing one. The proposed methodology is evaluated extensively and the comparison with several other solutions reveals that it competes favourably with the state-of-the-art on the most popular long-term visual tracking benchmarks.
Matteo Dunnhofer, Christian Micheloni
ICPR2
2022 Pro-CCaps: Progressively Teaching Colourisation to Capsules
abstract
Automatic image colourisation studies how to colourise greyscale images. Existing approaches exploit convolutional layers that extract image-level features learning the colourisation on the entire image, but miss entities-level ones due to pooling strategies. We believe that entity-level features are of paramount importance to deal with the intrinsic multimodality of the problem (i.e., the same object can have different colours, and the same colour can have different properties). Models based on capsule layers aim to identify entity-level features in the image from different points of view, but they do not keep track of global features.Our network architecture integrates entity-level features into the image-level features to generate a plausible image colourisation. We observed that results obtained with direct integration of such two representations are largely dominated by the image-level features, thus resulting in unsaturated colours for the entities. To limit such an issue, we propose a gradual growth of the reconstruction phase of the model while training. By advantaging of prior knowledge from each growing step, we obtain a stable collaboration between image-level and entity-level features that ultimately generates stable and vibrant colourisations. Experimental results on three benchmark datasets, and a user study, demonstrate that our approach has competitive performance with respect to the state-of-the-art and provides more consistent colourisation.
Rita Pucci, Christian Micheloni, Gian Luca Foresti, Niki Martinel
WACV2
2022 Combining complementary trackers for enhanced long-term visual object tracking
Matteo Dunnhofer, Kristian Simonato, Christian Micheloni
Image Vis. Comput.3
2022 Consistent attentive dual branch network for person re-identification
abstract
Abstract Several recent person re-identification methods are focusing on learning discriminative representations by designing efficient metric learning loss functions. Other approaches design part based architectures to compute an informative descriptor based on local features from semantically coherent parts. Few efforts learn the relationship between distant similar regions and parts by adjusting them to their most feasible positions with the help of soft attention. However, they focus on calibrating distant similar parts features and ignore to learn the noise (blur) free and distinct feature representations as the person re-identification datasets contain degraded images. To tackle these issues, we propose a novel Consistent Attention Dual Branch Network (CadNet) that has ability to model long-range dependencies (correlations) between channels as well as feature maps. We adopt multiple classifiers trained to learn the most discriminative global features for a unique representation of a person. Correlation between channels are consistently computed by using channel attention mechanism to make the learned feature noise free and distict from noisy and blurry data. Feature correlations interpret the relationship between distant similarities in the images computed by the self attention mechanism. The proposed CadNet significantly enhances the performance with respect to the baseline on the person re-identification benchmarks.
Asad Munir, Niki Martinel, Christian Micheloni
Multim. Tools Appl.3
2022 Lord of the Rings: Hanoi Pooling and Self-Knowledge Distillation for Fast and Accurate Vehicle Reidentification
abstract
Vehicle reidentification has seen increasing interest, thanks to its fundamental impact on intelligent surveillance systems and smart transportation. The visual data acquired from monitoring camera networks come with severe challenges, including occlusions, color and illumination changes, as well as orientation issues (a vehicle can be seen from the side/front/rear due to different camera viewpoints). To deal with such challenges, the community has spent much effort in learning robust feature representations that hinge on additional visual attributes and part-driven methods, but with the side effects of requiring extensive human annotation labor as well as increasing computational complexity. In this article, we propose an approach that learns a feature representation robust to vehicle orientation issues without the need for extra-labeled data and adding negligible computational overheads. The former objective is achieved through the introduction of a Hanoi pooling layer exploiting ring regions and the image pyramid approach yielding a multiscale representation of vehicle appearance. The latter is tackled by transferring the accuracy of a deep network to its first layers, thus reducing the inference effort by the early stop of a test example. This is obtained by means of a self-knowledge distillation framework encouraging multiexit network decisions to agree with each other. Results demonstrate that the proposed approach significantly improves the accuracy of early (i.e., very fast) exits while maintaining the same accuracy of a deep (slow) baseline. Moreover, our solution obtains the best existing performance on three benchmark datasets.11[Online]. Available:https://github.com/iN1k1/.
Niki Martinel, Matteo Dunnhofer, Rita Pucci, Gian Luca Foresti, Christian Micheloni
IEEE Trans. Ind. Informatics5
2021 Oriented Splits Network to Distill Background for Vehicle Re-Identification
abstract
Vehicle re-identification (re-id) is a challenging task due to the presence of high intra-class and low inter-class variations in the visual data acquired from monitoring camera networks. Unique and discriminative feature representations are needed to overcome the existence of several variations including color, illumination, orientation, background and occlusion. The orientations of the vehicles in the images make the learned models unable to learn multiple parts of the vehicle and relationship between them. The combination of global and partial features is one of the solutions to improve the discriminative learning of deep learning models. Leveraging on such solutions, we propose an Oriented Splits Network (OSN) for an end to end learning of multiple features along with global features to form a strong descriptor for vehicle re-identification. To capture the orientation variability of the vehicles, the proposed network introduces a partition of the images into several oriented stripes to obtain local descriptors for each part/region. Such a scheme is therefore exploited by a camera based feature distillation (CBD) training strategy to remove the background features. These are filtered out from oriented vehicles representations which yield to a much stronger unique representation of the vehicles. We perform experiments on two benchmark vehicle re-id datasets to verify the performance of the proposed approach which show that the proposed solution achieves better result with respect to the state of the art with margin.
Asad Munir, Niki Martinel, Christian Micheloni
AVSS3
2020 Tracking-by-Trackers with a Distilled and Reinforced Model
Matteo Dunnhofer, Niki Martinel, Christian Micheloni
ACCV (2)3
2020 Multi Branch Siamese Network For Person Re-Identification
abstract
To capture robust person features, learning discriminative, style and view invariant descriptors is a key challenge in person Re-Identification (re-id). Most deep Re-ID models learn single scale feature representation which are unable to grasp compact and style invariant representations. In this paper, we present a multi branch Siamese Deep Neural Network with multiple classifiers to overcome the above issues. The multi-branch learning of the network creates a stronger descriptor with fine-grained information from global features of a person. Camera to camera image translation is performed with generative adversarial network to generate diverse data and add style invariance in learned features. Experimental results on benchmark datasets demonstrate that the proposed method performs better than other state of the arts methods.
Asad Munir, Niki Martinel, Christian Micheloni
ICIP3
2020 Self and Channel Attention Network for Person Re-Identification
abstract
Recent research has shown promising results for person re-identification by focusing on several trends. One is designing efficient metric learning loss functions such as triplet loss family to learn the most discriminative representations. The other is learning local features by designing part based architectures to form an informative descriptor from semantically coherent parts. Some efforts adjust distant outliers to their most similar positions by using soft attention and learn the relationship between distant similar features. However, only a few prior efforts focus on channel-wise dependencies and learn non-local sharp similar part features directly for the degraded data in the person re-identification task. In this paper, we propose a novel Self and Channel Attention Network (SCAN) to model long-range dependencies between channels and feature maps. We add multiple classifiers to learn discriminative global features by using classification loss. Self Attention (SA) module and Channel Attention (CA) module are introduced to model non-local and channel-wise dependencies in the learned features. Spectral normalization is applied to the whole network to stabilize the training process. Experimental results on the person re-identification benchmarks show the proposed components achieve significant improvement with respect to the baseline.
Asad Munir, Niki Martinel, Christian Micheloni
ICPR3
2020 Fixed simplex coordinates for angular margin loss in CapsNet
abstract
A more stationary and discriminative embedding is necessary for robust classification of images. We focus our attention on the newel CapsNet model and we propose the angular margin loss function in composition with margin loss. We define a fixed classifier implemented with fixed weights vectors obtained by the vertex coordinates of a simplex polytope. The advantage of using simplex polytope is that we obtain the maximal symmetry for stationary features angularly centred. Each weight vector is to be considered as the centroid of a class in the dataset. The embedding of an image is obtained through the capsule network encoding phase, that is identified as digitcaps matrix. Based on the centroids from the simplex coordinates and the embedding from the model, we compute the angular distance between the image embedding and the centroid of the correspondent class of the image. We take this angular distance as angular margin loss. We keep the computation proposed for margin loss in the original architecture of CapsNet. We train the model to minimise the angular between the embedding and the centroid of the class and maximise the magnitude of the embedding for the predicted class. The experiments on different datasets demonstrate that the angular margin loss improves the capability of capsule networks with complex datasets.
Rita Pucci, Christian Micheloni, Gian Luca Foresti, Niki Martinel
ICPR2
2020 Deep Iterative Residual Convolutional Network for Single Image Super-Resolution
abstract
Deep convolutional neural networks (CNNs) have recently achieved great success for single image super-resolution (SISR) task due to their powerful feature representation capabilities. The most recent deep learning based SISR methods focus on designing deeper / wider models to learn the non-linear mapping between low-resolution (LR) inputs and high-resolution (HR) outputs. These existing SR methods do not take into account the image observation (physical) model and thus require a large number of network's trainable parameters with a great volume of training data. To address these issues, we propose a deep Iterative Super-Resolution Residual Convolutional Network (ISRResCNet) that exploits the powerful image regularization and large-scale optimization techniques by training the deep network in an iterative manner with a residual learning approach. Extensive experimental results on various super-resolution benchmarks demonstrate that our method with a few trainable parameters improves the results for different scaling factors in comparison with the state-of-art methods.
Rao Muhammad Umer, Gian Luca Foresti, Christian Micheloni
ICPR3
2020 Siam-U-Net: encoder-decoder siamese network for knee cartilage tracking in ultrasound images
Matteo Dunnhofer, Maria Antico, Fumio Sasazawa, Yu Takeda, Saskia Camps, Niki Martinel, Christian Micheloni, Gustavo Carneiro 0001, Davide Fontanarosa
Medical Image Anal.7
2020 Deep interactive encoding with capsule networks for image classification
Rita Pucci, Christian Micheloni, Gian Luca Foresti, Niki Martinel
Multim. Tools Appl.2
2020 Adaptive neural tree exploiting expert nodes to classify high-dimensional data
Shadi Abpeikar, Mehdi Ghatee, Gian Luca Foresti, Christian Micheloni
Neural Networks4
2020 Deep Pyramidal Pooling With Attention for Person Re-Identification
abstract
Learning discriminative, view-invariant and multi-scale representations of object appearance with different semantic levels is of paramount importance for person Re-Identification (ReID). Recently, the community has focused on learning deep Re-ID models to capture a single holistic representation. To improve the achieved results, additional visual attributes and object part-driven models have been considered, inevitably introducing additional human annotation labor or computational efforts. In this paper, we argue that pyramid-inspired methods capturing multi-scale information may overcome such requirements. Precisely, multi-scale pooled regions representing visual information of an object are integrated within a novel deep architecture factorizing them into discriminative features at multiple semantic levels. These are exploited through an attention mechanism later considered in an identification-similarity multi-task loss, trained by means of a curriculum learning strategy. Extensive results on three person ReID benchmarks demonstrate that better performance than existing methods are achieved. Code is available at https://github.com/iN1k1.
Niki Martinel, Gian Luca Foresti, Christian Micheloni
IEEE Trans. Image Process.3
2019 From person to group re-identification via unsupervised transfer of sparse features
Giuseppe Lisanti, Niki Martinel, Christian Micheloni, Alberto Del Bimbo, Gian Luca Foresti
Image Vis. Comput.3
2019 Distributed person re-identification through network-wise rank fusion consensus
Niki Martinel, Gian Luca Foresti, Christian Micheloni
Pattern Recognit. Lett.3
2018 Wide-Slice Residual Networks for Food Recognition
abstract
Image-based food recognition pose new challenges for mainstream computer vision algorithms. Recent works in the field focused either on hand-crafted representations or on learning these by exploiting deep neural networks (DNN). Despite the success of DNN-based works, these exploit off-the-shelf deep architectures which are not cast to the specific food classification problem. We believe that better results can be obtained if the architecture is defined with respect to an analysis of the food composition. Following such an intuition, this work introduces a new deep scheme that is designed to handle the food structure. In particular, we focus on the vertical food traits that are common to a large number of categories (i.e., 15% of the whole data in current datasets). Towards the final objective, we first introduce a slice convolution block to capture such specific information. Then, we leverage on the recent success of deep residual blocks and combine those with the sliced convolution to produce the classification score. Extensive evaluations on three benchmark datasets demonstrated that our solution has better performance than existing approaches (e.g., a top-1 accuracy of 90.27% on the Food-101 dataset).
Niki Martinel, Gian Luca Foresti, Christian Micheloni
WACV3
2017 Aerial video surveillance system for small-scale UAV environment monitoring
abstract
Change detection algorithms are commonly used to detect novelties for surveillance purposes in public and private places equipped by static or Pan-Tilt-Zoom (PTZ) cameras. Often, these techniques are also used as prerequisite to support more complex algorithms, including event recognition, object classification, person re-identification, and many others. With regard to small-scale Unmanned Aerial Vehicles (UAVs) at low-altitude, the change detection techniques require further investigation. In fact, most of the works currently available in the literature process video sequences acquired at very high-altitude for large-scale operations, such as vegetation monitoring, mapping of buildings, and so on. In a wide range of application contexts that require, for example, frequent monitoring or high spatial resolution for detecting small objects, video sequences acquired at high-altitude are not suitable. This paper presents a change detection system based on histogram equalization and RGB-Local Binary Pattern (RGB-LBP) operator for monitoring of wide areas by small-scale UAVs at low-altitude. Extensive experimental results show the robustness of the proposed pipeline. These latter were performed by using challenging video sequences of the public UAV Mosaicking and Change Detection (UMCD) dataset and measured a set of well-known statistical metrics. Finally, a performance analysis of the proposed algorithm is also provided.
Danilo Avola, Gian Luca Foresti, Niki Martinel, Christian Micheloni, Daniele Pannone, Claudio Piciarelli
AVSS4
2017 An ADAS Design based on IoT V2X Communications to Improve Safety - Case Study and IoT Architecture Reference Model
Yakusheva Nadezda, Gian Luca Foresti, Christian Micheloni
VEHITS3
2017 Person Reidentification in a Distributed Camera Network Framework
abstract
Plenty of research has been conducted to obtain the best reidentification performance between a single camera-pairs. None of the current approaches has addressed the reidentification in a camera network by considering the network topology (i.e., the structure of the monitored environment). We introduce a distributed network person reidentification framework which introduces the following contributions. 1) a camera matching cost to measure the reidentification performance between nodes of the network and 2) a derivation of the distance vector algorithm which allows to learn the network topology thus to prioritize and limit the cameras inquired for the matching of the probe. Results on three benchmark datasets show that the network topology can be learned in an unsupervised fashion and network-wise reidentification performance improves. As a side effect, we obtain that the communication bandwidth usage is reduced.
Niki Martinel, Gian Luca Foresti, Christian Micheloni
IEEE Trans. Cybern.3
2017 Discriminant Context Information Analysis for Post-Ranking Person Re-Identification
abstract
Existing approaches for person re-identification are mainly based on creating distinctive representations or on learning optimal metrics. The achieved results are then provided in the form of a list of ranked matching persons. It often happens that the true match is not ranked first but it is in the first positions. This is mostly due to the visual ambiguities shared between the true match and other "similar" persons. At the current state, there is a lack of a study of such visual ambiguities which limit the re-identification performance within the first ranks. We believe that an analysis of the similar appearances of the first ranks can be helpful in detecting, hence removing, such visual ambiguities. We propose to achieve such a goal by introducing an unsupervised post-ranking framework. Once the initial ranking is available, content and context sets are extracted. Then, these are exploited to remove the visual ambiguities and to obtain the discriminant feature space which is finally exploited to compute the new ranking. An in-depth analysis of the performance achieved on three public benchmark data sets support our believes. For every data set, the proposed method remarkably improves the first ranks results and outperforms the state-of-the-art approaches.
Jorge García 0002, Niki Martinel, Alfredo Gardel Vicente, Ignacio Bravo Muñoz, Gian Luca Foresti, Christian Micheloni
IEEE Trans. Image Process.6
2016 Temporal Model Adaptation for Person Re-identification
Niki Martinel, Abir Das, Christian Micheloni, Amit K. Roy-Chowdhury
ECCV (4)3
2016 Mobile ocular biometrics in visible spectrum using local image descriptors: A preliminary study
abstract
Ocular biometrics refers to personal identification using iris, conjunctival vasculature, periocular or eye movements. Contrary to most of other biometric traits, ocular biometrics does not require high user cooperation and close capture distance. Biometrics is now adopted ubiquitously as an alternative to passwords on mobile devices. Especially, ocular biometrics in the visible spectrum has attracted a lot of attention owing to the fact that it can be acquired using the regular RGB cameras already available in all mobile devices. The use of local image descriptors (i.e., analysis of microtextural features) for ocular biometrics is gaining more and more popularity because of their compactness, computationally inexpensiveness, excellent performance and flexibility. In this work, we explore the possibility of performing large scale mobile ocular biometric recognition in the visible spectrum using local image descriptors. We design a weighted fusion scheme to combine the information originating from four different local descriptors. The experimental analysis of the devised scheme, on newly collected and publicly available large scale database using three different mobile devices, shows promising results.
Zahid Akhtar, Christian Micheloni, Gian Luca Foresti
ICIP2
2016 Distributed and Unsupervised Cost-Driven Person Re-Identification
abstract
The problem of re-identify persons across single disjoint camera-pairs has received great attention from the community. Despite this, when the re-identification process has to be carried out on a large camera network a different approach has to be considered. In particular, existing approaches have neglected the importance of the network topology (i.e., the structure of the monitored environment) in such a process. To try filling such a gap, we propose a Distributed and Unsupervised Cost-Driven Person Re-Identification framework (DUPRe) which introduces the following contributions: (i) a camera matching cost to measure the re-identification performance between nodes of the network; (ii) a derivation of the distance vector algorithm which allows to learn the network topology hence to prioritize and limit the cameras inquired for the re-identification. Results on two benchmark datasets show that our solution brings to significant network-wise re-identification improvements.
Niki Martinel, Gian Luca Foresti, Christian Micheloni
ICPR3
2016 A supervised extreme learning committee for food recognition
Niki Martinel, Claudio Piciarelli, Christian Micheloni
Comput. Vis. Image Underst.3
2016 Modeling feature distances by orientation driven classifiers for person re-identification
Jorge García 0002, Niki Martinel, Alfredo Gardel Vicente, Ignacio Bravo Muñoz, Gian Luca Foresti, Christian Micheloni
J. Vis. Commun. Image Represent.6
2016 A pool of multiple person re-identification experts
abstract
The person re-identification problem, i.e. recognizing a person across non-overlapping cameras at different times and locations, is of fundamental importance for video surveillance applications. Due to pose variations, illumination conditions, background clutter, and occlusions, re-identify a person is an inherently difficult problem which is still far from being solved. In this work, inspired by the recent police lineup innovations, we propose a re-identification approach where Multiple Re-identification Experts (MuRE) are trained to reliably match new probes. The answers from all the experts are then combined to achieve a final decision. The proposed method has been evaluated on three datasets showing significant improvements over state-of-the-art approaches.
Niki Martinel, Christian Micheloni, Gian Luca Foresti
Pattern Recognit. Lett.2
2015 Person Re-Identification Ranking Optimisation by Discriminant Context Information Analysis
abstract
Person re-identification is an open and challenging problem in computer vision. Existing re-identification approaches focus on optimal methods for features matching (e.g., metric learning approaches) or study the inter-camera transformations of such features. These methods hardly ever pay attention to the problem of visual ambiguities shared between the first ranks. In this paper, we focus on such a problem and introduce an unsupervised ranking optimization approach based on discriminant context information analysis. The proposed approach refines a given initial ranking by removing the visual ambiguities common to first ranks. This is achieved by analyzing their content and context information. Extensive experiments on three publicly available benchmark datasets and different baseline methods have been conducted. Results demonstrate a remarkable improvement in the first positions of the ranking. Regardless of the selected dataset, state-of-the-art methods are strongly outperformed by our method.
Jorge García 0002, Niki Martinel, Christian Micheloni, Alfredo Gardel Vicente
ICCV3
2015 Re-Identification in the Function Space of Feature Warps
abstract
Person re-identification in a non-overlapping multicamera scenario is an open challenge in computer vision because of the large changes in appearances caused by variations in viewing angle, lighting, background clutter, and occlusion over multiple cameras. As a result of these variations, features describing the same person get transformed between cameras. To model the transformation of features, the feature space is nonlinearly warped to get the "warp functions". The warp functions between two instances of the same target form the set of feasible warp functions while those between instances of different targets form the set of infeasible warp functions. In this work, we build upon the observation that feature transformations between cameras lie in a nonlinear function space of all possible feature transformations. The space consisting of all the feasible and infeasible warp functions is the warp function space (WFS). We propose to learn a discriminating surface separating these two sets of warp functions in the WFS and to re-identify persons by classifying a test warp function as feasible or infeasible. Towards this objective, a Random Forest (RF) classifier is employed which effectively chooses the warp function components according to their importance in separating the feasible and the infeasible warp functions in the WFS. Extensive experiments on five datasets are carried out to show the superior performance of the proposed approach over state-of-the-art person re-identification methods. We show that our approach outperforms all other methods when large illumination variations are considered. At the same time it has been shown that our method reaches the best average performance over multiple combinations of the datasets, thus, showing that our method is not designed only to address a specific challenge posed by a particular dataset.
Niki Martinel, Abir Das, Christian Micheloni, Amit K. Roy-Chowdhury
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 A neural tree for classification using convex objective function
Asha Rani 0005, Gian Luca Foresti, Christian Micheloni
Pattern Recognit. Lett.3
2015 Classification of Local Eigen-Dissimilarities for Person Re-Identification
abstract
The task of re-identifying a person that moves across cameras fields-of-view is a challenge to the community known as the person re-identification problem. State-of-the art approaches are either based on direct modeling and matching of the human appearance or on machine learning-based techniques. In this work we introduce a novel approach that studies densely localized image dissimilarities in a low dimensional space and uses those to re-identify between persons in a supervised classification framework. To achieve the goal: i) we compute the localized image dissimilarity between a pair of images; ii) we learn the lower dimensional space of such localized image dissimilarities, known as the “local eigen-dissimilarities” (LEDs) space; iii) we train a binary classifier to discriminate between LEDs computed for a positive pair (images are for a same person) from the ones computed for a negative pair (images are for different persons). We show the competitive performance of our approach on two publicly available benchmark datasets.
Niki Martinel, Christian Micheloni
IEEE Signal Process. Lett.2
2015 Kernelized Saliency-Based Person Re-Identification Through Multiple Metric Learning
abstract
Person re-identification in a non-overlapping multi-camera scenario is an open and interesting challenge. While the task can hardly be completed by machines, we, as humans, are inherently able to sample those relevant persons' details that allow us to correctly solve the problem in a fraction of a second. Thus, knowing where a human might fixate to recognize a person is of paramount interest for re-identification. Inspired by the human gazing capabilities, we want to identify the salient regions of a person appearance to tackle the problem. Toward this objective, we introduce the following main contributions. A kernelized graph-based approach is used to detect the salient regions of a person appearance, later used as a weighting tool in the feature extraction process. The proposed person representation combines visual features either considering or not the saliency. These are then exploited in a pairwise-based multiple metric learning framework. Finally, the non-Euclidean metrics that have been separately learned for each feature are fused to re-identify a person. The proposed kernelized saliency-based person re-identification through multiple metric learning has been evaluated on four publicly available benchmark data sets to show its superior performance over the state-of-the-art approaches (e.g., it achieves a rank 1 correct recognition rate of 42.41% on the VIPeR data set).
Niki Martinel, Christian Micheloni, Gian Luca Foresti
IEEE Trans. Image Process.2
2014 MoBio_LivDet: Mobile biometric liveness detection
abstract
Biometric authentication is now being used ubiquitously as an alternative to passwords on mobile devices. However, current biometric systems are vulnerable to simple spoofing attacks. Several liveness detection methods have been proposed to determine whether there is a live person or an artificial replica in front of the biometric sensor. Yet, the problem is unsolved due to hardship in finding discriminative and computationally inexpensive features for spoofing attacks. Moreover, previous liveness detection approaches are not explicitly aimed for mobile biometric, thus principally unsuited for portable devices. Therefore, we build a software-based multi-biometric prototype that detects face, iris and fingerprint spoofing attacks on mobile devices. We present MoBio_LivDet (Mobile Biometric Liveness Detection), a novel approach that analyzes local features and global structures of the biometric images using a set of low-level feature descriptors and decision level fusion. The system allows user to balance the security level (robustness against spoofing) and convenience that they want. The proposed method is highly fast, simple, efficient, robust and does not require user-cooperation, thus making it extremely apt for mobile devices. Experimental analysis on publicly available face, iris and fingerprint data sets with real spoofing attacks show promising results.
Zahid Akhtar, Christian Micheloni, Claudio Piciarelli, Gian Luca Foresti
AVSS2
2014 Person Orientation and Feature Distances Boost Re-identification
abstract
Most of the open challenges in person re-identification arise from the large variations of human appearance and from the different camera views that may be involved, making pure feature matching an unreliable solution. To tackle these challenges state-of-the-art methods assume that a unique inter-camera transformation of features undergoes between two cameras. However, the combination of view points, scene illumination and photometric settings, etc., together with the appearance, pose and orientation of a person make the inter-camera transformation of features multi-modal. To address these challenges we introduce three main contributions. We propose a method to extract multiple frames of the same person with different orientation. We learn the pair wise feature dissimilarities space (PFDS) formed by the subspace of pair wise feature dissimilarities computed between images of persons with similar orientation and the subspace of pair wise feature dissimilarities computed between images of persons non-similar orientations. Finally, a classifier is trained to capture the multi-modal inter-camera transformation of pair wise images for each subspace. To validate the proposed approach we show the superior performance of our approach to state-of-the-art methods using two publicly available benchmark datasets.
Jorge García 0002, Niki Martinel, Gian Luca Foresti, Alfredo Gardel Vicente, Christian Micheloni
ICPR5
2014 Camera Selection for Adaptive Human-Computer Interface
abstract
Video analytics has become a very important topic in computer vision. This paper introduces advanced video analytics human-computer interfaces for a video surveillance system to ease the tasks of security operators. The visualization of the most relevant views is provided by the human-computer interface module that preemptively activates cameras that will probably cover the motion of interesting objects. Human-computer interaction principles have been considered to develop the novel user interface. Four prototypes have been designed and usability performance has been evaluated, exploiting standard methods. Results obtained from such evaluations show the efficiency of the novel information visualization technique.
Niki Martinel, Christian Micheloni, Claudio Piciarelli, Gian Luca Foresti
IEEE Trans. Syst. Man Cybern. Syst.2
2013 Robust Painting Recognition and Registration for Mobile Augmented Reality
abstract
In this work we introduce a novel approach for painting recognition and registration for mobile Augmented Reality applications. To address the challenges of real-time painting recognition and registration we introduce three main contributions: i) A relevant painting region detector extracts the painting region from the given image. ii) Two local and global features are extracted from the relevant region to robustly match a painting database. iii) A RANSAC homography estimation method is used to overlay the additional content in an AR framework. Experiments have been carried out on a dataset built with publicly available images.
Niki Martinel, Christian Micheloni, Gian Luca Foresti
IEEE Signal Process. Lett.2
2012 A balanced neural tree for pattern classification
Christian Micheloni, Asha Rani 0005, Sanjeev Kumar 0001, Gian Luca Foresti
Neural Networks1
2011 AVSS 2011 demo session: Smart Resource-Aware Multi-Sensor Network
abstract
Summary form only given. The invited talks are: How to Compare Alternative Architectures by Radia Perlman of Intel; Portals 4: Enabling Application/Architecture Co-Design for High-Performance Interconnects by Ron Brightwell of Sandia National Laboratories; and Electronic-Photonic Integration within Switches and Routers by Mike Watts of MIT. Brief author biographies are also included.
Fadi Al Machot, Bernhard Dieber, Petra Hossl, Kyandoghere Kyamakya, Sabrina Londero, Christian Micheloni, Paolo Omero, Claudio Piciarelli, Bernhard Rinner, Carlo Tasso, Massimiliano Valotto
AVSS6
2011 Smart resource-aware multimedia sensor network for automatic detection of complex events
abstract
This paper presents a smart resource-aware multimedia sensor network. We illustrate a surveillance system which supports human operators, by automatically detecting the complex events and giving the possibility to recall the detected events and searching them in an intelligent search engine. Four subsystems have been implemented, the tracking and detection system, the network configuration system, the reasoning system and an advanced archiving system in an annotated multimedia database.
Fadi Al Machot, Carlo Tasso, Bernhard Dieber, Kyandoghere Kyamakya, Claudio Piciarelli, Christian Micheloni, Sabrina Londero, Massimiliano Valotto, Paolo Omero, Bernhard Rinner
AVSS6
2011 Tracking sound sources by means of HMM
abstract
Video-based surveillance systems may benefit from the integration with microphone arrays for the localization of sound events. Applying the sound localization techniques to the surveillance of large areas requires addressing some open issues, such as the non uniform resolution of the microphones-based localization systems. This paper presents a new method for tracking moving sound events based on an Hidden Markov Model (HMM), which exploits a priori information derived from medium and longterm observations of the monitored area. The results obtained with simulated trajectories show that the HMM-based tracker is able to significantly reduce the localization error. Applications can be found in surveillance systems for large areas, such as square, streets, or parking lots, where it is of interest the monitoring of moving vehicles and people.
Antonio Rodà, Christian Micheloni
AVSS2
2011 Resource-Aware Coverage and Task Assignment in Visual Sensor Networks
abstract
A visual sensor network (VSN) consists of a large amount of camera nodes which are able to process the captured image data locally and to extract the relevant information. The tight resource limitations in these networks of embedded sensors and processors represent a major challenge for the application development. In this paper, we focus on finding optimal VSN configurations which are basically given by: 1) the selection of cameras to sufficiently monitor the area of interest; 2) the setting of the cameras' frame rate and resolution to fulfill the quality of service requirements; and 3) the assignment of processing tasks to cameras to achieve all required monitoring activities. We formally specify this configuration problem and describe an efficient approximation method based on an evolutionary algorithm. We analyze our approximation method on three different scenarios and compare the predicted results with measurements on real implementations on a VSN platform. We finally combine our approximation method with an expectation-maximization algorithm for optimizing the coverage and resource allocation in VSN with pan-tilt-zoom camera nodes.
Bernhard Dieber, Christian Micheloni, Bernhard Rinner
IEEE Trans. Circuits Syst. Video Technol.2
2010 Human Action Recognition using a Hybrid NTLD Classifier
abstract
This work proposes a hybrid classifier to recognize human actions in different contexts. In particular, the proposed hybrid classifier (a neural tree with linear discriminant nodes NTLD), is a neural tree whose nodes can be either simple preceptrons or recursive fisher linear discriminant (RFLD) classifiers. A novel technique to substitute bad trained perceptron with more performant linear discriminators is introduced. For a given frame, geometrical features are extracted from the skeleton of the human blob (silhouette). These geometrical features are collected for a fixed number of consecutive frames to recognize the corresponding activity. The resulting feature vector is adopted as input to the NTLD classifier. The performance of the proposed classifier has been evaluated on two available databases.
Asha Rani 0005, Sanjeev Kumar 0001, Christian Micheloni, Gian Luca Foresti
AVSS3
2010 Stereo rectification of uncalibrated and heterogeneous images
Sanjeev Kumar 0001, Christian Micheloni, Claudio Piciarelli, Gian Luca Foresti
Pattern Recognit. Lett.2
2009 Stereo Localization Based on Network's Uncalibrated Camera Pairs
abstract
In this paper, a stereo framework for a robust real time localization of objects using networkpsilas camera pairs is presented. The stereo system contains a combination of static and pan-tilt-zoom (PTZ) cameras instead of traditional dual head mounted cameras. The proposed novelty consists in applying stereo vision to heterogeneous cameras belonging to a video-surveillance network. First, a look-up-table (LUT) is built with the rectification transformations computed for some predefined pan and tilt values. Then, the LUT is used to compute rectification transformations by means of neural networks for any arbitrary pan and tilt settings. Different zoom levels are compensated by resizing images according to their focal ratio and by applying zero padding. Localization of any object is made using its 3D position information obtained by a modified stereo concept. Experimental results are presented for the localization of moving objects in a parking lot scenario.
Sanjeev Kumar 0001, Christian Micheloni, Claudio Piciarelli, Gian Luca Foresti
AVSS2
2009 Stereo Localization Using Dual PTZ Cameras
Sanjeev Kumar 0001, Christian Micheloni, Claudio Piciarelli
CAIP2
2009 Exploiting temporal statistics for events analysis and understanding
Christian Micheloni, Lauro Snidaro, Gian Luca Foresti
Image Vis. Comput.1
2009 Active Tuning of Intrinsic Camera Parameters
abstract
In the last years, the research effort of the scientific community to study systems for ambient intelligence has been really strong. Usually, the systems developed so far base their analysis on images acquired by automatic cameras. In this paper, we propose a way to develop new smart systems that are able to actively decide both what to see and how to see it. In particular, the main idea is to tune the acquisition parameters on the basis of what the system desires to acquire. The regulation strategy is based on two camera parameters, focus and iris. It aims to identify an optimal sequence of steps to enhance the acquisition quality of an object of interest. To this end, a hierarchy of neural networks has been employed first to select which parameter must be regulated then to adjust it. The proposed solution can be applied to both static and moving cameras. The results show how the proposed technique can be applied to images acquired by a moving camera with zoom capabilities for surveillance purposes.
Christian Micheloni, Gian Luca Foresti
IEEE Trans Autom. Sci. Eng.1
2008 A security assistance system combining person tracking with chemical attributes and video event analysis
Christopher Becher, Gian Luca Foresti, Peter Kaul, Wolfgang Koch 0001, Frank P. Lorenz, Daniel Lubczyk, Christian Micheloni, Claudio Piciarelli, Konstantin Safenreiter, Carsten Siering, Macarena Varela, Siegfried R. Waldvogel, Monika Wieneke
FUSION7
2008 Support vector machines for robust trajectory clustering
abstract
Many event analysis systems are based on the detection of uncommon feature patterns that could be associated to anomalous events; the uncommon patterns are identified by comparison with a "normality model" describing the previously acquired data. In this work we propose an anomaly detection system based on trajectory clustering with single-class support vector machines. However, SVM parameter tuning would require an a-priori estimate of the number of outlier trajectories in the training data, which is unknown. We here propose a technique for automatic estimation of the number of outliers, thus avoiding the arbitrary choice of constant tuning parameters.
Claudio Piciarelli, Christian Micheloni, Gian Luca Foresti
ICIP2
2008 Anomalous trajectory patterns detection
abstract
In the field of event analysis, the detection of anomalous events has often been based on the creation of a model representing the most common patterns of activity detected within a monitored scene. This way, anomalous events can be identified by comparison with the model as patterns differing from typical events. In particular, trajectories of moving objects have often been used as a feature for anomalous event detection. In this paper we propose a combination of clustering and SVM techniques in order to automatically detect anomalous trajectories.
Claudio Piciarelli, Christian Micheloni, Gian Luca Foresti
ICPR2
2008 Adaptive video communication for an intelligent distributed system: Tuning sensors parameters for surveillance purposes
Christian Micheloni, Marco Lestuzzi, Gian Luca Foresti
Mach. Vis. Appl.1
2008 Trajectory-Based Anomalous Event Detection
abstract
During the last years, the task of automatic event analysis in video sequences has gained an increasing attention among the research community. The application domains are disparate, ranging from video surveillance to automatic video annotation for sport videos or TV shots. Whatever the application field, most of the works in event analysis are based on two main approaches: the former based on explicit event recognition, focused on finding high-level, semantic interpretations of video sequences, and the latter based on anomaly detection. This paper deals with the second approach, where the final goal is not the explicit labeling of recognized events, but the detection of anomalous events differing from typical patterns. In particular, the proposed work addresses anomaly detection by means of trajectory analysis, an approach with several application fields, most notably video surveillance and traffic monitoring. The proposed approach is based on single-class support vector machine (SVM) clustering, where the novelty detection SVM capabilities are used for the identification of anomalous trajectories. Particular attention is given to trajectory classification in absence ofaprioriinformation on the distribution of outliers. Experimental results prove the validity of the proposed approach.
Claudio Piciarelli, Christian Micheloni, Gian Luca Foresti
IEEE Trans. Circuits Syst. Video Technol.2
2007 Tuning Asymboost Cascades Improves Face Detection
abstract
The face detection problem is certainly one of the most studied topics in artificial vision. This interest raises from the conscience that this is a crucial step for every system that uses biometric information. Video surveillance and security systems, biometrics, HCI and multimedia applications are some examples of systems that exploit face localization to improve their robustness. AdaBoost and AsymBoost based classifiers are widely used to achieve high performances saving computational time. In this paper, a new reactive strategy to build a strong classifier cascade is provided; at each stage of the cascade a different tradeoff between accuracy and computational complexity is explored. The results will show that this method is effective, and propose a way to construct a rapid and robust multipose detector.
Ingrid Visentini, Christian Micheloni, Gian Luca Foresti
ICIP (4)2
2006 Sensor Bandwidth Assignment through Video Annotation
abstract
The state of the art of surveillance systems include a large set of techniques for both low level and high level tasks. In particular, the research community has witnessed in the last decade a high proliferation of techniques that span from object detection and tracking to object recogni- tion and event understanding. Although some techniques have been proven to be very effective those tasks cannot be considered solved. Although more effort is needed in the event analysis field, a new problem arises from the develop- ment of large scale networked surveillance systems: infor- mation sharing. The way information is shared between the nodes of the surveillance network today represents a key- point issue. To provide a first and novel solution to such a problem, we propose an innovative system architecture for a video surveillance system with distributed processing over multiple processing units and with distributed communica- tion over multiple heterogeneous channels (wireless, satel- lite, local IP networks, etc.). In particular, a new real-time technique for changing the video transmission parameters (e.g., frame rate, spatial/color resolution, etc.) according to the bandwidth available will be here presented.
Christian Micheloni, Lauro Snidaro, Ingrid Visentini, Gian Luca Foresti
AVSS1
2006 Real-time image processing for active monitoring of wide areas
Christian Micheloni, Gian Luca Foresti
J. Vis. Commun. Image Represent.1
2005 An integrated surveillance system for outdoor security
abstract
An integrated system for the detection, active tracking and recognition of people in wide outdoor environments is hereafter discussed. Specifically, a static sensor with a wide view is used to detect people inside the environment and to classify their behaviours. As outcome of anomalous activities, an active camera is selected to focus its attention on a particular person. Here, techniques for face detection are employed to determine a region of interest where to extract features used by a tracking algorithm for an autonomous gaze of a PTZ camera. Finally, a face recognition phase is considered to recognize the person of interest. Results show how the integrated system is able to detect, track and recognise people inside a tough environment such as a parking lot.
Christian Micheloni, Elena Salvador, Flavio Bigaran, Gian Luca Foresti
AVSS1
2005 Zoom on target while tracking
abstract
In this paper, the problem of continuous tracking of moving objects with a PTZ camera is addressed. In particular, the problem of tracking moving objects during zoom phases is solved by using a feature clustering technique. In order to adopt such a method, we need, first, a step where during tracking with a pan&tilt camera we can identify the mobile objects in the monitored scene. Therefore, a set of good trackable features belonging to the selected target is extracted. In this research, we adopt a feature clustering method that is able to discriminate between features associated with the background and features associated with different moving objects. As a result, for each moving object, we have a set of correctly tracked features that is used to track the objects. Experiments have been performed on outdoor environments where either people or vehicles have been tracked. The results highlight how such a technique can be included in a more complex system able to maintain targets in the field of view of the camera, and to zoom on an object of interest when desired.
Christian Micheloni, Gian Luca Foresti
ICIP (3)1
2005 Detecting moving people in video streams
Gian Luca Foresti, Christian Micheloni, Claudio Piciarelli
Pattern Recognit. Lett.2
2005 Video security for ambient intelligence
abstract
Moving toward the implementation of the intelligent building idea in the framework of ambient intelligence, a video security application for people detection, tracking, and counting in indoor environments is presented in this paper. In addition to security purposes, the system may be employed to estimate the number of accesses in public buildings, as well as the preferred followed routes. Computer vision techniques are used to analyze and process video streams acquired from multiple video cameras. Image segmentation is performed to detect moving regions and to calculate the number of people in the scene. Testing was performed on indoor video sequences with different illumination conditions.
Lauro Snidaro, Christian Micheloni, C. Chiavedale
IEEE Trans. Syst. Man Cybern. Part A2
2004 A new feature clustering method for object detection with an active camera
abstract
Feature based methods for ego-motion estimation are widely used in computer vision but they must deal with errors in feature tracking. In this paper, we propose a robust real-time method for ego-motion estimation by assuming an affine motion of the background from the previous to the current frame. A new clustering technique is applied on image's subareas to select in a fast and reliable way three features for the affine transform computation. The previous frame after being warped according to the computed affine transform is processed with the current frame by a change detection method in order to detect mobile objects. Results are presented in the context of a visual-based surveillance system for monitoring outdoor environments.
Christian Micheloni, Gian Luca Foresti, Flavio Alberti
ICIP1
2003 Fast Good Features Selection for Wide Area Monitoring
abstract
Recently the surveillance of wide areas has pointed the interest of the research community. The use of active vision seems to be the most effective solutions for these needs. Against the better acquiring resolution there is the problem of the apparent motion inducted by the camera motion known as ego-motion. Feature based methods for ego-motion estimation are widely used in computer vision but they deal with feature recovery and with errors in feature tracking. In this paper, we propose a fast method to extract and select new features during camera motion. This is achieved by adopting a reference map containing well trackable features that is updated at each frame by introducing new good features related to regions appearing in the current image. A new procedure is applied to reject badly tracked features. The current frame and the background after compensation are processed by a change detection method in order to locate mobile objects. Results are presented in the context of a visual-based surveillance system for monitoring outdoor environments.
Christian Micheloni, Gian Luca Foresti
AVSS1
2003 A robust face detection system for real environments
abstract
In this paper, a robust real-time face detection system based on the integration of different location methods is proposed. A hierarchical architecture composed of three levels is designed. At the first level, a change detection method is applied to detect blobs of moving objects (i.e., humans) in the scene. Then, the silhouette of each blob is analyzed to focalize the attention of the system on small image areas where the probability of finding human heads is high. At the second level, two different methods, i.e., the skin color and the principal component analysis, are applied to locate human faces. Finally, the higher level fuses the obtained location data to improve the face detection reliability. The system is tested in outdoor environments in the context of a video-based surveillance system.
Gian Luca Foresti, Christian Micheloni, Lauro Snidaro
ICIP (3)2
2002 Generalized neural trees for pattern classification
abstract
In this paper, a new neural tree (NT) model, the generalized NT (GNT), is presented. The main novelty of the GNT consists in the definition of a new training rule that performs an overall optimization of the tree. Each time the tree is increased by a new level, the whole tree is reevaluated. The training rule uses a weight correction strategy that takes into account the entire tree structure, and it applies a normalization procedure to the activation values of each node such that these values can be interpreted as a probability. The weight connection updating is calculated by minimizing a cost function, which represents a measure of the overall probability of correct classification. Significant results on both synthetic and real data have been obtained by comparing the classification performances among multilayer perceptrons (MLPs), NTs, and GNTs. In particular, the GNT model displays good classification performances for training sets having complex distributions. Moreover, its particular structure provides an easily probabilistic interpretation of the pattern classification task and allows growing small neural trees with good generalization properties.
Gian Luca Foresti, Christian Micheloni
IEEE Trans. Neural Networks2