Simone Bianco 0001

dblp:36/1511-1 · DBLP profile ↗
← Back
42ranked-venue papers
24as first author
17since 2021 · last 2026
0000-0002-7070-1545ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 12 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 13 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 On Using AI for EEG-Based BCI Applications: Problems, Current Challenges and Future Trends
Thomas Barbera, Jacopo Burger, Alessandro D'Amelio, Simone Zini, Simone Bianco 0001, Raffaella Lanzarotti, Paolo Napoletano, Giuseppe Boccignone, José Luis Contreras-Vidal
Int. J. Hum. Comput. Interact.5
2026 Decoupling spatial and spectral features for efficient hyperspectral image super-resolution
abstract
Abstract Hyperspectral image super-resolution (HSI-SR) aims to reconstruct hyperspectral images at high spatial resolution, starting from low-resolution inputs, while preserving both spatial details and spectral fidelity. In this work, we propose the Efficient Spatial-Spectral Processing Network (ESSPN), a lightweight deep learning architecture designed to address the challenges of HSI-SR in a computationally efficient way. ESSPN is built around a novel Spatial-Spectral Block (SSB) that separately models spatial structures and spectral correlations through residual convolutional and attention mechanisms. The network head incorporates an efficient upsampling module based on pixel shuffle decomposition to produce high-resolution outputs without interpolation artifacts. Extensive experiments on two publicly available datasets, i.e., ARAD1K and StereoMSI demonstrate that ESSPN achieves competitive or superior performance compared to state-of-the-art methods when evaluated at scale factors of $$\times $$ 4, $$\times $$ 6 and $$\times $$ 8. Notably, the model shows strong generalization across hyperspectral cameras with varying spectral responses, and across radiometric domains, covering both radiance and reflectance measurements, while requiring significantly fewer parameters and FLOPs compared to existing methods. These results position the proposed ESSPN as a practical and effective solution for high-quality hyperspectral image super-resolution in real-world applications.
Matteo Kolyszko, Marco Buzzelli, Simone Bianco 0001, Raimondo Schettini
Mach. Vis. Appl.3
2026 Cross-Camera Distracted Driver Classification Through Feature Disentanglement and Contrastive Learning
abstract
The classification of distracted drivers is pivotal for ensuring safe driving. Previous studies demonstrated the effectiveness of neural networks in automatically predicting driver distraction, fatigue, and potential hazards. However, recent research has uncovered a significant loss of accuracy in these models when applied to samples acquired under conditions that differ from the training data. In this paper, we introduce a robust model designed to withstand changes in camera position within the vehicle. Our Driver Behavior Monitoring Network (DBMNet) relies on a lightweight backbone and integrates a disentanglement module to discard camera view information from features, coupled with contrastive learning to enhance the encoding of various driver actions. Experiments conducted using a leave-one-camera-out protocol on the daytime and nighttime subsets of the 100-Driver dataset validate the effectiveness of our approach. Cross-dataset and cross-camera experiments conducted on three benchmark datasets, namely AUCDD-V1, EZZ2021 and SFD, demonstrate the superior generalization capabilities of the proposed method. Overall DBMNet achieves an improvement of 7% in Top-1 accuracy compared to existing efficient approaches. Moreover, a quantized version of the DBMNet and all considered methods has been deployed on a Coral Dev Board board. In this deployment scenario, DBMNet outperforms alternatives, achieving the lowest average error while maintaining a compact model size, low memory footprint, fast inference time, and minimal power consumption.
Luigi Celona, Simone Bianco 0001, Paolo Napoletano
IEEE Trans. Intell. Transp. Syst.2
2025 Federated Learning for Cross-Dataset Generalization in Litter Detection
abstract
The accumulation of litter in natural environments poses significant ecological and social challenges, motivating the development of automated solutions for litter detection. However, collecting and centrally aggregating large-scale annotated datasets for training object detectors often raises privacy and ownership concerns. In this work, we propose a Federated Learning (FL) framework to train a lightweight litter detection model based on the YOLO architecture, which enables collaborative model development without requiring centralized access to raw data. Each participating client locally trains the model on site-specific datasets collected in the wild, and only model updates are shared with a central server for aggregation. We compare and contrast different FL process configurations involving mixed and heterogeneous training datasets built starting from two commonly used benchmark datasets collected across different locations and having very different visual data distributions, i.e. TACO and PlastOPol. Experimental results show that the federated model, trained across these non-IID data distributions, achieves superior generalization in cross-dataset evaluation compared to the corresponding centrally trained models.
Luciano Baresi, Simone Bianco 0001, Livia Lestingi, Iyad Wehbe
ECAI2
2025 Enhancing color selectivity in foundation models for downstream color vision tasks
abstract
The emergence of computer vision foundation models, inspired by the success of task-agnostic pretrained representations in Natural Language Processing (NLP), is revolutionizing the field. These models produce features that excel in downstream tasks even without fine-tuning. Last year, DINOv2 emerged, surpassing previous state-of-the-art general-purpose features on computer vision benchmarks, both at the image and pixel levels. In this work, we focus on what type of color information is embedded in DINOv2 features, and to assess their performance in computer vision tasks where color is a critical cue—for instance, recognizing the color of vehicles for traffic monitoring, detecting skin tones in biometric applications, or assessing product color attributes in fashion and e-commerce. Furthermore, we also propose a training-free feature transformation that increases color selectivity in DINOv2 features, i.e. their ability to respond differently to various colors in an image, boosting the performance on several classes of the color vision tasks considered.
Simone Bianco 0001
Neurocomputing1
2025 Improving image captioning descriptiveness by ranking and LLM-based fusion
abstract
Abstract State-of-the-art (SoTA) image captioning models are often trained on the MicroSoft Common Objects in Context (MS-COCO) dataset, which contains human-annotated captions with an average length of approximately ten tokens. Although effective for general scene understanding, these short captions often fail to capture complex scenes and convey detailed information. Moreover, captioning models tend to exhibit bias toward the “average” caption, which captures only the more general aspects, thus overlooking finer details. In this paper, we present a novel approach to generate richer and more informative image captions by combining the captions generated from different SoTA captioning models. Our proposed method requires no additional model training: given an image, it leverages pretrained models from the literature to generate the initial captions, and then ranks them using a newly introduced image-text-based metric, which we name BLIPScore. Subsequently, the top two captions are fused using a Large Language Model to produce the final, more detailed description. Experimental results on the MS-COCO and Flickr30k test sets demonstrate the effectiveness of our approach in terms of caption-image alignment and hallucination reduction according to the ALOHa, CAPTURE, and Polos metrics. A subjective study lends additional support to these results, suggesting that the captions produced by our model are generally perceived as more consistent with human judgment. By combining the strengths of diverse SoTA models, our method enhances the quality and appeal of image captions, bridging the gap between automated systems and the rich and informative nature of human-generated descriptions. This advance enables the generation of more suitable captions for the training of both vision-language and captioning models.
Luigi Celona, Simone Bianco 0001, Marco Donzella, Paolo Napoletano
Neural Comput. Appl.2
2025 Uncertainty estimation in color constancy
Marco Buzzelli, Simone Bianco 0001
Pattern Recognit.2
2025 Robust camera-independent color chart localization using YOLO
abstract
Accurate color information plays a critical role in numerous computer vision tasks, with the Macbeth ColorChecker being a widely used reference target due to its colorimetrically characterized color patches. However, automating the precise extraction of color information in complex scenes remains a challenge. In this paper, we propose a novel method for the automatic detection and accurate extraction of color information from Macbeth ColorCheckers in challenging environments. Our approach involves two distinct phases: (i) a chart localization step using a deep learning model to identify the presence of the ColorChecker, and (ii) a consensus-based pose estimation and color extraction phase that ensures precise localization and description of individual color patches. We rigorously evaluate our method using the widely adopted NUS and ColorChecker datasets. Comparative results against state-of-the-art methods show that our method outperforms the best solution in the state of the art achieving about 5% improvement on the ColorChecker dataset and about 17% on the NUS dataset. Furthermore, the design of our approach enables it to handle the presence of multiple ColorCheckers in complex scenes. Code will be made available after pubblication at: https://github.com/LucaCogo/ColorChartLocalization . • A robust and precise method for Camera-Independent Color Chart Localization is proposed. • The method has two phases: a chart localization step, and a consensus-based pose estimation. • The method can handle the presence of multiple color targets in complex scenes. • The method outperforms the state-of-the-art solution by up to about 17% on standard datasets.
Luca Cogo, Marco Buzzelli, Simone Bianco 0001, Raimondo Schettini
Pattern Recognit. Lett.3
2025 A Convolutional Framework for Color Constancy
abstract
We introduce a convolutional framework (CF) for computational color constancy, building upon the established low-level image feature-based framework, which utilized simple image statistics for illuminant estimation. Our framework expands upon this through an end-to-end learnable neural architecture. This adaptation enables the learning and usage of advanced filters that are not restricted to Gaussian kernels operating on individual color channels, thus generalizing the capabilities of the original framework. Additionally, our general framework supports deeper convolutional architectures, thus increasing its computational power. It can also be efficiently applied to estimate multiple spatially varying illuminants within a single scene. Our experimental results on standard datasets demonstrate that the CF outperforms the best methods in the low-level framework, improving the illuminant estimation accuracy by up to 34% for single illuminant estimation and 30% for multiple illuminants estimation. Additionally, our framework exhibits superior performance even when the number of training images is reduced. Finally, we document the inference speedup of our implementation reaching up to $30\times $ , making the CF especially suitable for applications where efficiency is critical. Source code and trained models available at: https://github.com/MarcoBauzz/convolutional-color-constancy.
Marco Buzzelli, Simone Bianco 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Scalable Residual Laplacian Network for HEVC-compressed Video Restoration
abstract
We present a novel Convolutional Neural Network that exploits the Laplacian decomposition technique, which is typically used in traditional image processing, to restore videos compressed with the High-Efficiency Video Coding (HEVC) algorithm. The proposed method decomposes the compressed frames into multi-scale frequency bands using the Laplacian decomposition, it restores each band using the ad-hoc designed Multi-frame Residual Laplacian Network (MRLN), and finally recomposes the restored bands to obtain the restored frames. By leveraging the multi-scale frequency representation of compressed frames provided by the Laplacian decomposition, MRLN can effectively reduce the compression artifacts and restore the image details with a reduced computational cost. In addition, our method can be easily instantiated in various versions to control the tradeoff between efficiency and effectiveness, representing a versatile solution for scenarios with constrained computational resources. Experimental results on the MFQEv2 benchmark dataset show that our method achieves the state-of-the-art performance in HEVC-compressed video restoration with a lower model complexity and shorter runtime with respect to existing methods. The project page is available at https://github.com/claudiom4sir/LaplacianVCAR .
Claudio Rota, Marco Buzzelli, Simone Bianco 0001, Raimondo Schettini
ACM Trans. Multim. Comput. Commun. Appl.3
2024 COBOL: COmmunity-Based Organized Littering
abstract
Littering is a major problem that threatens the environment, society, and economy. Keep track, monitor and regularly clean littering sites can be a crucial problem that involves public authorities, municipalities, companies, and citizens. So far approaches have not well leveraged the knowledge and capabilities that derive from the federation of multiple communities, such as cities, public bodies, and organizations. In this paper, we describe the COBOL project, a National PRIN (Progetti di Rilevante Interesse Nazionale) PNRR (Piano Nazionale Ripresa e Resilienza) project funded by the Italian MUR (Ministero dell'Università e della Ricerca) in 2023. The project aims to definite a flexible framework for managing the waste disposal process through a federated learning architecture that collects and integrates the reports (e.g., annotated pictures and user feedback) shared by the communities involved in the waste disposal process. To deliver an advanced waste disposal service based on the direct participation of citizens, COBOL also integrates Model-Driven Engineering principles, Computer Vision techniques, and Self-Adaptation mechanisms. Early results show that reports can be effectively collected and processed with COBOL.
Luciano Baresi, Simone Bianco 0001, Amleto Di Salle, Ludovico Iovino, Leonardo Mariani, Daniela Micucci, Luciana Brasil Rebelo dos Santos, Maria Teresa Rossi, Raimondo Schettini
SEAA2
2024 Semi-supervised cross-lingual speech emotion recognition
abstract
Performance in Speech Emotion Recognition (SER) on a single language has increased greatly in the last few years thanks to the use of deep learning techniques. However, cross-lingual SER remains a challenge in real-world applications due to two main factors: the first is the big gap among the source and the target domain distributions; the second factor is the major availability of unlabeled utterances in contrast to the labeled ones for the new language. Taking into account previous aspects, we propose a Semi-Supervised Learning (SSL) method for cross-lingual emotion recognition when only few labeled examples in the target domain (i.e. the new language) are available. Our method is based on a Transformer and it adapts to the new domain by exploiting a pseudo-labeling strategy on the unlabeled utterances. In particular, the use of a hard and soft pseudo-labels approach is investigated. We thoroughly evaluate the performance of the proposed method in a speaker-independent setup on both the source and the new language and show its robustness across five languages belonging to different linguistic strains. The experimental findings indicate that the unweighted accuracy is increased by an average of 40% compared to state-of-the-art methods.
Mirko Agarla, Simone Bianco 0001, Luigi Celona, Paolo Napoletano, Alexey Petrovsky, Flavio Piccoli, Raimondo Schettini, Ivan Shanin
Expert Syst. Appl.2
2024 A RNN for Temporal Consistency in Low-Light Videos Enhanced by Single-Frame Methods
abstract
Low-light video enhancement (LLVE) has received little attention compared to low-light image enhancement (LLIE) mainly due to the lack of paired low-/normal-light video datasets. Consequently, a common approach to LLVE is to enhance each video frame individually using LLIE methods. However, this practice introduces temporal inconsistencies in the resulting video. In this work, we propose a recurrent neural network (RNN) that, given a low-light video and its per-frame enhanced version, produces a temporally consistent video preserving the underlying frame-based enhancement. We achieve this by training our network with a combination of a new forward-backward temporal consistency loss and a content-preserving loss. At inference time, we can use our trained network to correct videos processed by any LLIE method. Experimental results show that our method achieves the best trade-off between temporal consistency improvement and fidelity with the per-frame enhanced video, exhibiting a lower memory complexity and comparable time complexity with respect to other state-of-the-art methods for temporal consistency.
Claudio Rota, Marco Buzzelli, Simone Bianco 0001, Raimondo Schettini
IEEE Signal Process. Lett.3
2023 Unified Framework for Identity and Imagined Action Recognition From EEG Patterns
abstract
We present a unified deep learning framework for the recognition of user identity and the recognition of imagined actions, based on electroencephalography (EEG) signals, for application as a brain–computer interface. Our solution exploits a novel shifted subsampling preprocessing step as a form of data augmentation, and a matrix representation to encode the inherent local spatial relationships of multielectrode EEG signals. The resulting image-like data are then fed to a convolutional neural network to process the local spatial dependencies, and eventually analyzed through a bidirectional long-short term memory module to focus on temporal relationships. Our solution is compared against several methods in the state of the art, showing comparable or superior performance on different tasks. Specifically, we achieve accuracy levels above 90% both for action and user classification tasks. In terms of user identification, we reach 0.39% equal error rate in the case of known users and gestures, and 6.16% in the more challenging case of unknown users and gestures. Preliminary experiments are also conducted in order to direct future works toward everyday applications relying on a reduced set of EEG electrodes.
Marco Buzzelli, Simone Bianco 0001, Paolo Napoletano
IEEE Trans. Hum. Mach. Syst.2
2022 A Framework for Contrast Enhancement Algorithms Optimization
abstract
We present a general-purpose framework for the optimization of parametric contrast enhancement algorithms. We first define a regression module for image acceptability, which is based on deep neural features and which is trained on a large dataset of user-expressed preferences. This regression module is then used as the objective function of a Bayesian optimization process, guiding the search for the optimal parameters of a given contrast enhancement algorithm. In our experiments we optimize three different contrast enhancement algorithms of varying levels of complexity. The effectiveness of our optimization framework is experimentally confirmed by evaluating the output of the optimized contrast enhancement algorithms with respect to reference enhanced images.
Simone Zini, Marco Buzzelli, Simone Bianco 0001, Raimondo Schettini
ICIP3
2022 U-WeAr: User Recognition on Wearable Devices through Arm Gesture
abstract
The use of wearable devices equipped with inertial sensors has become increasingly pervasive. It has been widely demonstrated in the literature that inertial signals acquired by these sensors can be used by machine learning algorithms to predict actions performed and/or to recognize the identities of the person wearing the sensors. In this article, we present a hardware/software system for arm gesture recognition, identity recognition, and verification of a person based on inertial sensors. The hardware part is a custom wristband that consists of a computing unit, a wireless communication unit, and an inertial sensor. The software part is an algorithm based on recurrent neural networks that is able to process the signals coming from the sensor and to return a prediction. To validate the system, a dataset consisting of 25 symbols drawn with the arm is collected. These symbols are performed by 33 subjects. We conduct two evaluations: 1) performance evaluation for arm gesture recognition, user recognition and verification; and 2) usability assessment of the system. The performance of the three recognition tasks indicate that this system can be reliably applied in real environments with an accuracy above 96% for gesture recognition, an accuracy of about 85% for user identification, and an equal error rate of about 13% for user verification. The outcome of the usability test proves a great satisfaction from the users in terms of high simplicity in the use of the wristband and goodness of the machine learning predictions.
Simone Bianco 0001, Paolo Napoletano, Alberto Raimondi, Mirko Rima
IEEE Trans. Hum. Mach. Syst.1
2021 Disentangling Image distortions in deep feature space
Simone Bianco 0001, Luigi Celona, Paolo Napoletano
Pattern Recognit. Lett.1
2020 Consensus-driven illuminant estimation with GANs
abstract
We present a method for illuminant estimation that exploits a generative adversarial network architecture to generate a spatially-varying illuminant map. This map is then transformed by consensus into a global illuminant estimation, in the form of a single RGB triplet. To this end, different consensus strategies are designed and compared in this paper. The best solution won second place in the 2nd International Illumination Estimation Challenge, specifically for the indoor track.
Simone Bianco 0001, Raimondo Schettini
ICMV1
2020 Providing a Single Ground-Truth for Illuminant Estimation for the ColorChecker Dataset
abstract
The ColorChecker dataset is one of the most widely used image sets for evaluating and ranking illuminant estimation algorithms. However, this single set of images has at least 3 different sets of ground-truth (i.e., correct answers) associated with it. In the literature it is often asserted that one algorithm is better than another when the algorithms in question have been tuned and tested with the different ground-truths. In this short correspondence we present some of the background as to why the 3 existing ground-truths are different and go on to make a new single and recommended set of correct answers. Experiments reinforce the importance of this work in that we show that the total ordering of a set of algorithms may be reversed depending on whether we use the new or legacy ground-truth data.
Ghalia Hemrit, Graham D. Finlayson, Arjan Gijsenij, Peter V. Gehler, Simone Bianco 0001, Mark S. Drew, Brian V. Funt, Lilong Shi
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Personalized Image Enhancement Using Neural Spline Color Transforms
abstract
In this work we present SpliNet, a novel CNNbased method that estimates a global color transform for the enhancement of raw images. The method is designed to improve the perceived quality of the images by reproducing the ability of an expert in the field of photo editing. The transformation applied to the input image is found by a convolutional neural network specifically trained for this purpose. More precisely, the network takes as input a raw image and produces as output one set of control points for each of the three color channels. Then, the control points are interpolated with natural cubic splines and the resulting functions are globally applied to the values of the input pixels to produce the output image. Experimental results compare favorably against recent methods in the state of the art on the MIT-Adobe FiveK dataset. Furthermore, we also propose an extension of the SpliNet in which a single neural network is used to model the style of multiple reference retouchers by embedding them into a user space. The style of new users can be reproduced without retraining the network, after a quick modeling stage in which they are positioned in the user space on the basis of their preferences on a very small set of retouched images.
Simone Bianco 0001, Claudio Cusano, Flavio Piccoli, Raimondo Schettini
IEEE Trans. Image Process.1
2019 Quasi-Unsupervised Color Constancy
abstract
We present here a method for computational color constancy in which a deep convolutional neural network is trained to detect achromatic pixels in color images after they have been converted to grayscale. The method does not require any information about the illuminant in the scene and relies on the weak assumption, fulfilled by almost all images available on the web, that training images have been approximately balanced. Because of this requirement we define our method as quasi-unsupervised. After training, unbalanced images can be processed thanks to the preliminary conversion to grayscale of the input to the neural network. The results of an extensive experimentation demonstrate that the proposed method is able to outperform the other unsupervised methods in the state of the art being, at the same time, flexible enough to be supervisedly fine-tuned to reach performance comparable with those of the best supervised methods.
Simone Bianco 0001, Claudio Cusano
CVPR1
2019 Turning a Digital Camera into an Absolute 2D Tele-Colorimeter
abstract
Abstract We present a simple and effective technique for absolute colorimetric camera characterization, invariant to changes in exposure/aperture and scene irradiance, suitable in a wide range of applications including image‐based reflectance measurements, spectral pre‐filtering and spectral upsampling for rendering, to improve colour accuracy in high dynamic range imaging. Our method requires a limited number of acquisitions, an off‐the‐shelf target and a commonly available projector, used as a controllable light source, other than the reflected radiance to be known. The characterized camera can be effectively used as a 2D tele‐colorimeter, providing the user with an accurate estimate of the distribution of luminance and chromaticity in a scene, without requiring explicit knowledge of the incident lighting power spectra. We validate the approach by comparing our estimated absolute tristimulus values (XYZ data in ) with the measurements of a professional 2D tele‐colorimeter, for a set of scenes with complex geometry, spatially varying reflectance and light sources with very different spectral power distribution.
Giuseppe Claudio Guarnera, Simone Bianco 0001, Raimondo Schettini
Comput. Graph. Forum2
2019 Multitask painting categorization by deep multibranch neural network
Simone Bianco 0001, Davide Mazzini, Paolo Napoletano, Raimondo Schettini
Expert Syst. Appl.1
2019 A unifying representation for pixel-precise distance estimation
Simone Bianco 0001, Marco Buzzelli, Raimondo Schettini
Multim. Tools Appl.1
2018 Aesthetics Assessment of Images Containing Faces
abstract
Recent research has widely explored the problem of aesthetics assessment of images with generic content. However, few approaches have been specifically designed to predict the aesthetic quality of images containing human faces, which make up a massive portion of photos in the web. This paper introduces a method for aesthetic quality assessment of images with faces. We exploit three different Convolutional Neural Networks to encode information regarding perceptual quality, global image aesthetics, and facial attributes; then, a model is trained to combine these features to explicitly predict the aesthetics of images containing faces. Experimental results show that our approach outperforms existing methods for both binary, i.e. low/high, and continuous aesthetic score prediction on four different databases in the state-of-the-art.
Simone Bianco 0001, Luigi Celona, Raimondo Schettini
ICIP1
2018 Automated Pruning for Deep Neural Network Compression
abstract
In this work we present a method to improve the pruning step of the current state-of-the-art methodology to compress neural networks. The novelty of the proposed pruning technique is in its differentiability, which allows pruning to be performed during the backpropagation phase of the network training. This enables an end-to-end learning and strongly reduces the training time. The technique is based on a family of differentiable pruning functions and a new regularizer specifically designed to enforce pruning. The experimental results show that the joint optimization of both the thresholds and the network weights permits to reach a higher compression rate, reducing the number of weights of the pruned network by a further 14% to 33 % compared to the current state-of-the-art. Furthermore, we believe that this is the first study where the generalization capabilities in transfer learning tasks of the features extracted by a pruned network are analyzed. To achieve this goal, we show that the representations learned using the proposed pruning methodology maintain the same effectiveness and generality of those learned by the corresponding non-compressed network on a set of different recognition tasks.
Franco Manessi, Alessandro Rozza, Simone Bianco 0001, Paolo Napoletano, Raimondo Schettini
ICPR3
2017 Deep learning for logo recognition
Simone Bianco 0001, Marco Buzzelli, Davide Mazzini, Raimondo Schettini
Neurocomputing1
2017 Large Age-Gap face verification by feature injection in deep networks
Simone Bianco 0001
Pattern Recognit. Lett.1
2017 Combination of Video Change Detection Algorithms by Genetic Programming
abstract
Within the field of computer vision, change detection algorithms aim at automatically detecting significant changes occurring in a scene by analyzing the sequence of frames in a video stream. In this paper we investigate how state-of-the-art change detection algorithms can be combined and used to create a more robust algorithm leveraging their individual peculiarities. We exploited genetic programming (GP) to automatically select the best algorithms, combine them in different ways, and perform the most suitable post-processing operations on the outputs of the algorithms. In particular, algorithms' combination and post-processing operations are achieved with unary, binary and n-ary functions embedded into the GP framework. Using different experimental settings for combining existing algorithms we obtained different GP solutions that we termed In Unity There Is Strength. These solutions are then compared against state-of-the-art change detection algorithms on the video sequences and ground truth annotations of the ChangeDetection.net 2014 challenge. Results demonstrate that using GP, our solutions are able to outperform all the considered single state-of-the-art change detection algorithms, as well as other combination strategies. The performance of our algorithm are significantly different from those of the other state-of-the-art algorithms. This fact is supported by the statistical significance analysis conducted with the Friedman test and Wilcoxon rank sum post-hoc tests.
Simone Bianco 0001, Gianluigi Ciocca, Raimondo Schettini
IEEE Trans. Evol. Comput.1
2017 Single and Multiple Illuminant Estimation Using Convolutional Neural Networks
abstract
In this paper, we present a three-stage method for the estimation of the color of the illuminant in RAW images. The first stage uses a convolutional neural network that has been specially designed to produce multiple local estimates of the illuminant. The second stage, given the local estimates, determines the number of illuminants in the scene. Finally, local illuminant estimates are refined by non-linear local aggregation, resulting in a global estimate in case of single illuminant. An extensive comparison with both local and global illuminant estimation methods in the state of the art, on standard data sets with single and multiple illuminants, proves the effectiveness of our method.
Simone Bianco 0001, Claudio Cusano, Raimondo Schettini
IEEE Trans. Image Process.1
2016 Predicting Image Aesthetics with Deep Learning
Simone Bianco 0001, Luigi Celona, Paolo Napoletano, Raimondo Schettini
ACIVS1
2016 CURL: Image Classification using co-training and Unsupervised Representation Learning
Simone Bianco 0001, Gianluigi Ciocca, Claudio Cusano
Comput. Vis. Image Underst.1
2015 An interactive tool for manual, semi-automatic and automatic video annotation
Simone Bianco 0001, Gianluigi Ciocca, Paolo Napoletano, Raimondo Schettini
Comput. Vis. Image Underst.1
2015 Adaptive Skin Classification Using Face and Body Detection
abstract
In this paper, we propose a skin classification method exploiting faces and bodies automatically detected in the image, to adaptively initialize individual ad hoc skin classifiers. Each classifier is initialized by a face and body couple or by a single face, if no reliable body is detected. Thus, the proposed method builds an ad hoc skin classifier for each person in the image, resulting in a classifier less dependent from changes in skin color due to tan levels, races, genders, and illumination conditions. The experimental results on a heterogeneous data set of labeled images show that our proposal outperforms the state-of-the-art methods, and that this improvement is statistically significant.
Simone Bianco 0001, Francesca Gasparini, Raimondo Schettini
IEEE Trans. Image Process.1
2015 User Preferences Modeling and Learning for Pleasing Photo Collage Generation
abstract
In this article, we consider how to automatically create pleasing photo collages created by placing a set of images on a limited canvas area. The task is formulated as an optimization problem. Differently from existing state-of-the-art approaches, we here exploit subjective experiments to model and learn pleasantness from user preferences. To this end, we design an experimental framework for the identification of the criteria that need to be taken into account to generate a pleasing photo collage. Five different thematic photo datasets are used to create collages using state-of-the-art criteria. A first subjective experiment where several subjects evaluated the collages, emphasizes that different criteria are involved in the subjective definition of pleasantness. We then identify new global and local criteria and design algorithms to quantify them. The relative importance of these criteria are automatically learned by exploiting the user preferences, and new collages are generated. To validate our framework, we performed several psycho-visual experiments involving different users. The results shows that the proposed framework allows to learn a novel computational model which effectively encodes an inter-user definition of pleasantness. The learned definition of pleasantness generalizes well to new photo datasets of different themes and sizes not used in the learning. Moreover, compared with two state-of-the-art approaches, the collages created using our framework are preferred by the majority of the users.
Simone Bianco 0001, Gianluigi Ciocca
ACM Trans. Multim. Comput. Commun. Appl.1
2014 Adaptive Color Constancy Using Faces
abstract
In this work we design an adaptive color constancy algorithm that, exploiting the skin regions found in faces, is able to estimate and correct the scene illumination. The algorithm automatically switches from global to spatially varying color correction on the basis of the illuminant estimations on the different faces detected in the image. An extensive comparison with both global and local color constancy algorithms is carried out to validate the effectiveness of the proposed algorithm in terms of both statistical and perceptual significance on a large heterogeneous data set of RAW images containing faces.
Simone Bianco 0001, Raimondo Schettini
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Color constancy using faces
abstract
In this work, we investigate how illuminant estimation can be performed exploiting the color statistics extracted from the faces automatically detected in the image. The proposed method is based on two observations: first, skin colors tend to form a cluster in the color space, making it a cue to estimate the illuminant in the scene; second, many photographic images are portraits or contain people. The proposed method has been tested on a public dataset of images in RAW format, using both a manual and a real face detector. Experimental results demonstrate the effectiveness of our approach. The proposed method can be directly used in many digital still camera processing pipelines with an embedded face detector working on gray level images.
Simone Bianco 0001, Raimondo Schettini
CVPR1
2012 Sampling Optimization for Printer Characterization by Direct Search
abstract
Printer characterization usually requires many printer inputs and corresponding color measurements of the printed outputs. In this brief, a sampling optimization for printer characterization on the basis of direct search is proposed to maintain high color accuracy with a reduction in the number of characterization samples required. The proposed method is able to match a given level of color accuracy requiring, on average, a characterization set cardinality which is almost one-fourth of that required by the uniform sampling, while the best method in the state of the art needs almost one-third. The number of characterization samples required can be further reduced if the proposed algorithm is coupled with a sequential optimization method that refines the sample values in the device-independent color space. The proposed sampling optimization method is extended to deal with multiple substrates simultaneously, giving statistically better colorimetric accuracy (at the α = 0.05 significance level) than sampling optimization techniques in the state of the art optimized for each individual substrate, thus allowing use of a single set of characterization samples for multiple substrates.
Simone Bianco 0001, Raimondo Schettini
IEEE Trans. Image Process.1
2010 Genetic Algorithms for Training Data and Polynomial Optimization in Colorimetric Characterization of Scanners
Leonardo Vanneschi, Mauro Castelli, Simone Bianco 0001, Raimondo Schettini
EvoApplications (1)3
2010 Automatic color constancy algorithm selection and combination
Simone Bianco 0001, Gianluigi Ciocca, Claudio Cusano, Raimondo Schettini
Pattern Recognit.1
2009 Empirical modeling for colorimetric characterization of digital cameras
abstract
One of the most complete techniques that can be used to digitize an artefact is to generate its photo-textured 3D model, which combines high precision metrical information with a faithful color description of the object surfaces. In this work we focus on how to obtain reliable color information by colorimetrically characterizing the color sensors of the imaging device. To this end, a novel target based characterization procedure is proposed that exploits empirical polynomial modeling for colorimetric data estimation. Experimental results are reported and discussed.
Simone Bianco 0001, Raimondo Schettini, Leonardo Vanneschi
ICIP1
2008 Improving Color Constancy Using Indoor-Outdoor Image Classification
abstract
In this work, we investigate how illuminant estimation techniques can be improved, taking into account automatically extracted information about the content of the images. We considered indoor/outdoor classification because the images of these classes present different content and are usually taken under different illumination conditions. We have designed different strategies for the selection and the tuning of the most appropriate algorithm (or combination of algorithms) for each class. We also considered the adoption of an uncertainty class which corresponds to the images where the indoor/outdoor classifier is not confident enough. The illuminant estimation algorithms considered here are derived from the framework recently proposed by Van de Weijer and Gevers. We present a procedure to automatically tune the algorithms' parameters. We have tested the proposed strategies on a suitable subset of the widely used Funt and Ciurea dataset. Experimental results clearly demonstrate that classification based strategies outperform general purpose algorithms.
Simone Bianco 0001, Gianluigi Ciocca, Claudio Cusano, Raimondo Schettini
IEEE Trans. Image Process.1