João C. Neves 0001

dblp:71/9719 · also João C. R. Neves, João Carlos Raposo Neves, João Neves 0002 · DBLP profile ↗
← Back
26ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-0139-2213ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 7 since 2021Security and privacy · 8 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 FD-MAD: Frequency-Domain Residual Analysis for Face Morphing Attack Detection
Diogo J. Paulo, Hugo Proença 0001, João C. Neves 0001
FG3
2026 StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
abstract
Urban waste management remains a critical challenge for the development of smart cities. Despite the growing number of litter detection datasets, the problem of monitoring overflowing waste containers — particularly from images captured by garbage trucks — has received little attention. While existing datasets are valuable, they often lack annotations for specific container tracking or are captured in static, decontextualized environments, limiting their utility for real-world logistics. To address this gap, we present StreetView-Waste, a comprehensive dataset of urban scenes featuring litter and waste containers. The dataset supports three key evaluation tasks: (1) waste container detection, (2) waste container tracking, and (3) waste overflow segmentation. Alongside the dataset, we provide baselines for each task by benchmarking state-of-the-art models in object detection, tracking, and segmentation. Additionally, we enhance baseline performance by proposing two complementary strategies: a heuristic-based method for improved waste container tracking and a model-agnostic framework that leverages geometric priors to refine litter segmentation. Our experimental results show that while fine-tuned object detectors achieve reasonable performance in detecting waste containers, baseline tracking methods struggle to accurately estimate their number; however, our proposed heuristics reduce the mean absolute counting error by 79.6%. Similarly, while segmenting amorphous litter is challenging, our geometry-aware strategy improves segmentation [email protected] by 27% on lightweight models, demonstrating the value of multimodal inputs for this task. Ultimately, StreetView-Waste provides a challenging benchmark to encourage research into real-world perception systems for urban waste management.
Diogo J. Paulo, Hugo Proença 0001, João C. Neves 0001
WACV4
2026 Unsupervised contrastive analysis for anomaly detection in brain MRIs via conditional diffusion models
abstract
Contrastive Analysis (CA) detects anomalies by contrasting patterns unique to a target group (e.g., unhealthy subjects) from those in a background group (e.g., healthy subjects). In the context of brain MRIs, existing CA approaches rely on supervised contrastive learning or variational autoencoders (VAEs) using both healthy and unhealthy data, but such reliance on target samples is challenging in clinical settings. Unsupervised Anomaly Detection (UAD) learns a reference representation of healthy anatomy, eliminating the need for target samples. Deviations from this reference distribution can indicate potential anomalies. In this context, diffusion models have been increasingly adopted in UAD due to their superior performance in image generation compared to VAEs. Nonetheless, precisely reconstructing the anatomy of the brain remains a challenge. In this work, we bridge CA and UAD by reformulating contrastive analysis principles for the unsupervised setting. We propose an unsupervised framework to improve the reconstruction quality by training a self-supervised contrastive encoder on healthy images to extract meaningful anatomical features. These features are used to condition a diffusion model to reconstruct the healthy appearance of a given image, enabling interpretable anomaly localization via pixel-wise comparison. We validate our approach through a proof-of-concept on a facial image dataset and further demonstrate its effectiveness on four brain MRI datasets, outperforming baseline methods in anomaly localization on the NOVA benchmark. • Unsupervised framework enhancing reconstruction quality in brain MRIs. • Target-invariant contrastive encoder capturing meaningful anatomical features. • Conditional diffusion model to reconstruct the healthy appearance of a given image. • Outperforms competing methods in anomaly localization on the NOVA benchmark.
Cristiano Patrício, Carlo Alberto Barbano, Attilio Fiandrotti, Riccardo Renzulli, Marco Grangetto, Luís F. Teixeira 0001, João C. Neves 0001
Pattern Recognit. Lett.7
2025 Bias Analysis for Synthetic Face Detection: A Case Study of the Impact of Facial Attributes
abstract
Bias analysis for synthetic face detection is bound to become a critical topic in the coming years. Although many detection models have been developed and several datasets have been released to reliably identify synthetic content, one crucial aspect has been largely overlooked: these models and training datasets can be biased, leading to failures in detection for certain demographic groups and raising significant social, legal, and ethical issues. In this work, we introduce an evaluation framework to contribute to the analysis of bias of synthetic face detectors with respect to several facial attributes. This framework exploits synthetic data generation, with evenly distributed attribute labels, for mitigating any skew in the data that could otherwise influence the outcomes of bias analysis. We build on the proposed framework to provide an extensive case study of the bias level of five state-of-the-art detectors in synthetic datasets with 25 controlled facial attributes. While the results confirm that, in general, synthetic face detectors are biased towards the presence/absence of specific facial attributes, our study also sheds light on the origins of the observed bias through the analysis of the correlations with the balancing of facial attributes in the training sets of the detectors, and the analysis of detectors activation maps in image pairs with controlled attribute modifications.
Asmae Lamsaf, Lucia Cascone, Hugo Proença 0001, João C. Neves 0001
IJCB4
2025 VM-TAPS: View-specific Memory with Temporal and Scale Awareness Framework for Video-based Cross-View Person Re-Identification
abstract
Reliable aerial-ground video-based person re-identification (ReID) remains a challenge due to severe changes in data quality and features, such as viewpoint disparities, resolution drops, and cross-camera appearance inconsistency. This paper presents VM-TAPS, a lightweight and modular extension to the well-known TF-CLIP framework, designed to increase the robustness of ReID, without requiring end-to-end backbone retraining. When compared to its ancestor, VM-TAPS' novelties are five-fold: 1) View-Specific Processing Layers to normalize camera-dependent biases; 2) Scale-Aware Feature Adaptation for resolution-invariant feature fusion; 3) a View-Aware Memory Bank enabling long-range identity context; 4) a Motion Pattern Analyzer capturing temporal dynamics; and (5) Cross-View Interaction Modules that harmonize multi-view feature spaces. Despite adding fewer than two million parameters, VM-TAPS achieves +4.97% Rank-1 and +3.08% mAP gains over TF-CLIP on the challenging AG-VPReID2025 benchmark. At 80m and 120m altitudes, it sets a new performance baseline of 73.68%/75.73% and 69.45%/71.63% (Rank-1/mAP), respectively. All components are trained with frozen CLIP visual encoders in the early stages, enabling efficient and stable convergence. Our results support that the carefully disentanglement of viewpoint, scale, motion and memory factors substantially increases the robustness of cross-view ReID under real-world conditions.
Md. Rashidunnabi, Kailash A. Hambarde, João C. Neves 0001, Vasco Lopes, Hugo Proença 0001
IJCB3
2025 Beyond a Single Perspective: Neural Fusion of Lévy-Generated Super-Resolution Images for Robust Face Recognition
abstract
Despite significant advancements in facial recognition technology, these systems struggle in real-world surveillance scenarios, especially when dealing with low-resolution images. Super-resolution techniques have emerged as a natural solution to this issue, primarily using generative models such as generative adversarial networks or diffusion models. Despite their potential, these methods face critical limitations, including mode collapse and difficulty in preserving identity-specific features. To address these challenges, we propose Lévy Super-Resolution (LSR), a novel super-resolution approach leveraging Lévy processes as a noise source in a modified diffusion model. This stochastic process enhances diversity and mitigates mode collapse, enabling the generation of diverse super-resolved samples from a single low-resolution input. These samples are fused using an ensemble of neural networks trained with a triplet loss, producing robust face descriptors with reduced bias and improved identity preservation. To the best of our knowledge, LSR is the first method to utilize Lévy processes for super-resolution generation. We validate our approach on real and synthetic low-resolution samples from three standard face recognition datasets, and we demonstrate that our approach can surpass both general and face-specific state-of-the-art super-resolution (SR) methods. Our code is publicly available at https://github.com/marcelowds/lsr.
Marcelo dos Santos, João C. Neves 0001, David Menotti
ICMLA2
2025 Synthesizing multilevel abstraction ear sketches for enhanced biometric recognition
abstract
Sketch understanding poses unique challenges for general-purpose vision algorithms due to the sparse and semantically ambiguous nature of sketches. This paper introduces a novel approach to biometric recognition that leverages sketch-based representations of ears, a largely unexplored but promising area in biometric research. Specifically, we address the “ sketch-2-image ” matching problem by synthesizing ear sketches at multiple abstraction levels, achieved through a triplet-loss function adapted to integrate these levels. The abstraction level is determined by the number of strokes used, with fewer strokes reflecting higher abstraction. Our methodology combines sketch representations across abstraction levels to improve robustness and generalizability in matching. Extensive evaluations were conducted on four ear datasets (AMI, AWE, IITDII, and BIPLab) using various pre-trained neural network backbones, showing consistently superior performance over state-of-the-art methods. These results highlight the potential of ear sketch-based recognition, with cross-dataset tests confirming its adaptability to real-world conditions and suggesting applicability beyond ear biometrics. • Sketch-Based Datasets Expansion: Leveraging CLIPasso, we transformed ear images into sketches at various abstraction levels, preserving key features and introducing a novel data representation for biometric analysis. • Triplet-Loss Function Enhancement: Adapting the triplet-loss function to incorporate multiple abstraction levels significantly improves recognition performance over traditional methods. • Comparative Backbone Analysis: An exhaustive evaluation of different backbones highlights their effectiveness in sketch-based ear recognition, guiding advancements in biometric technologies. • Cross-Dataset Generalizability Tests: Training on combined datasets and testing on distinct ones validate our approach’s robustness and effectiveness against unseen data distributions.
David Freire-Obregón, João C. Neves 0001, Ziga Emersic, Blaz Meden, Modesto Castrillón-Santana, Hugo Proença 0001
Image Vis. Comput.2
2024 CFC-ATE: Causal Feature Construction via Average Treatment Effect
abstract
Dimensionality reduction is a crucial step in data preprocessing, particularly for high-dimensional datasets, where the excessive number of features increases the risk of overfitting in machine learning models. Traditional dimensionality reduction methods rely on statistical associations or the relative position of the feature embeddings in the hyper-space to map original features to a compact subspace that preserves the most relevant information of the data. However, these methods fail to capture the causal relationships among variables during the transfor-mation process, leading to a loss of structural coherence of the data in low-dimensional spaces. By employing causal discovery and causal inference, it is possible to simplify these problems, effectively merging critical features while reducing both complex-ity and dimensionality. Our paper introduces a novel approach, Causal Feature Construction via Average Treatment Effect (CFC-ATE), which leverages causal discovery and inference to create more interpretable and reliable features for predictive modeling. Our methodology consists of the following phases: i) leveraging the causal structure of data through the inference of the causal graph. ii) transforming features through the use of the average treatment effect conditioned on the causal structure of the data. The experiments on diverse real-world datasets and synthetic datasets demonstrate the effectiveness of CFC-ATE in improving model performance by comparing it with three methods of feature selection and three benchmark dimensionality reduction techniques.
Asmae Lamsaf, Hugo Proença 0001, João C. Neves 0001
ICMLA3
2023 WildFruiP: Estimating Fruit Physicochemical Parameters from Images Captured in the Wild
Diogo J. Paulo, Cláudia M. B. Neves, Dulcineia Ferreira Wessel, João C. Neves 0001
CIARP4
2023 Zero-shot face recognition: Improving the discriminability of visual face features using a Semantic-Guided Attention Model
abstract
Zero-shot learning enables the recognition of classes not seen during training through the use of semantic information comprising a visual description of the class either in textual or attribute form. Despite the advances in the performance of zero-shot learning methods, most of the works do not explicitly exploit the correlation between the visual attributes of the image and their corresponding semantic attributes for learning discriminative visual features. In this paper, we introduce an attention-based strategy for deriving features from the image regions regarding the most prominent attributes of the image class. In particular, we train a Convolutional Neural Network (CNN) for image attribute prediction and use a gradient-weighted method for deriving the attention activation maps of the most salient image attributes. These maps are then incorporated into the feature extraction process of Zero-Shot Learning (ZSL) approaches for improving the discriminability of the features produced through the implicit inclusion of semantic information. For experimental validation, the performance of state-of-the-art ZSL methods was determined using features with and without the proposed attention model. Surprisingly, we discover that the proposed strategy degrades the performance of ZSL methods in classical ZSL datasets (AWA2), but it can significantly improve performance when using face datasets. Our experiments show that these results are a consequence of the interpretability of the dataset attributes, suggesting that existing ZSL datasets attributes are, in most cases, difficult to be identifiable in the image. Source code is available at https://github.com/CristianoPatricio/SGAM.
Cristiano Patrício, João C. Neves 0001
Expert Syst. Appl.2
2022 Generative Adversarial Graph Convolutional Networks for Human Action Synthesis
abstract
Synthesising the spatial and temporal dynamics of the human body skeleton remains a challenging task, not only in terms of the quality of the generated shapes, but also of their diversity, particularly to synthesise realistic body movements of a specific action (action conditioning). In this paper, we propose Kinetic-GAN, a novel architecture that leverages the benefits of Generative Adversarial Networks and Graph Convolutional Networks to synthesise the kinetics of the human body. The proposed adversarial architecture can condition up to 120 different actions over local and global body movements while improving sample quality and diversity through latent space disentanglement and stochastic variations. Our experiments were carried out in three well-known datasets, where Kinetic-GAN notably surpasses the state-of-the-art methods in terms of distribution quality metrics while having the ability to synthesise more than one order of magnitude regarding the number of different actions. Our code and models are publicly available at https://github.com/DegardinBruno/Kinetic-GAN.
Bruno Degardin, João C. Neves 0001, Vasco Lopes, João Brito, Ehsan Yaghoubi, Hugo Proença 0001
WACV2
2021 ZSpeedL - Evaluating the Performance of Zero-Shot Learning Methods using Low-Power Devices
abstract
The recognition of unseen objects from a semantic representation or textual description, usually denoted as zero-shot learning, is more prone to be used in real-world scenarios when compared to traditional object recognition. Nevertheless, no work has evaluated the feasibility of deploying zero-shot learning approaches in these scenarios, particularly when using low-power devices. In this paper, we provide the first benchmark on the inference time of zero-shot learning, comprising an evaluation of state-of-the-art approaches regarding their speed/accuracy trade-off. An analysis to the processing time of the different phases of the ZSL inference stage reveals that visual feature extraction is the major bottleneck in this paradigm, but, we show that lightweight networks can dramatically reduce the overall inference time without reducing the accuracy obtained by the de facto ResNet101 architecture. Also, this benchmark evaluates how different ZSL approaches perform in low-power devices, and how the visual feature extraction phase could be optimized in this hardware. To foster the research and deployment of ZSL systems capable of operating in real-world scenarios, we release the evaluation framework used in this benchmark(https://github.com/CristianoPatricio/zsl-methods).
Cristiano Patrício, João C. Neves 0001
AVSS2
2020 All-in-one "HairNet": A Deep Neural Model for Joint Hair Segmentation and Characterization
abstract
The hair appearance is among the most valuable soft biometric traits when performing human recognition at-a-distance. Even in degraded data, the hair's appearance is instinctively used by humans to distinguish between individuals. In this paper we propose a multi-task deep neural model capable of segmenting the hair region, while also inferring the hair color, shape and style, all from in-the-wild images. Our main contributions are two-fold: 1) the design of an all-in-one neural network, based on depthwise separable convolutions to extract the features; and 2) the use convolutional feature masking layer as an attention mechanism that enforces the analysis only within the `hair' regions. In a conceptual perspective, the strength of our model is that the segmentation mask is used by the other tasks to perceive - at feature-map level - only the regions relevant to the attribute characterization task. This paradigm allows the network to analyze features from nonrectangular areas of the input data, which is particularly important, considering the irregularity of hair regions. Our experiments showed that the proposed approach reaches a hair segmentation performance comparable to the state-of-the-art, having as main advantage the fact of performing multiple levels of analysis in a single-shot paradigm.
Diana Borza, Ehsan Yaghoubi, João C. Neves 0001, Hugo Proença 0001
IJCB3
2020 An attention-based deep learning model for multiple pedestrian attributes recognition
Ehsan Yaghoubi, Diana Borza, João C. Neves 0001, Aruna Kumar, Hugo Proença 0001
Image Vis. Comput.3
2019 "A Leopard Cannot Change Its Spots": Improving Face Recognition Using 3D-Based Caricatures
abstract
Caricatures refer to a representation of a person, in which the distinctive features are deliberately exaggerated, with several studies showing that humans perform better at recognizing people from caricatures than using original images. Inspired by this observation, this paper introduces the first fully automated caricature-based face recognition approach capable of working with data acquired in the wild. Our approach leverages the 3D face structure from a single 2D image and compares it with a reference model for obtaining a compact representation of face features deviations. This descriptor is subsequently deformed using a “measure locally, weight globally” strategy to resemble the caricature drawing process. The deformed deviations are incorporated in the 3D model using the Laplacian mesh deformation algorithm, and the 2D face caricature image is obtained by projecting the deformed model in the original camera view. To demonstrate the advantages of caricature-based face recognition, we train the VGG-face network from scratch using either original face images (baseline) or caricatured images and use these models for extracting face descriptors from the LFW, IJB-A, and MegaFace data sets. The experiments show an increase in the recognition accuracy when using caricatures rather than original images. Moreover, our approach achieves competitive results with the state-of-the-art face recognition methods, even without explicitly tuning the network for any of the evaluation sets.
João C. Neves 0001, Hugo Proença 0001
IEEE Trans. Inf. Forensics Secur.1
2019 A Reminiscence of "Mastermind": Iris/Periocular Biometrics by "In-Set" CNN Iterative Analysis
abstract
Convolutional neural networks (CNNs) have emerged as the most popular classification models in biometrics research. Under the discriminative paradigm of pattern recognition, CNNs are used typically in one of two ways: (1) verification mode (“ are samples from the same person? ”), where pairs of images are provided to the network to distinguish between genuine and impostor instances and (2) identification mode (“ whom is this sample from? ”), where appropriate feature representations that map images to identities are found. This paper postulates a novel mode for using CNNs in biometric identification, by learning models that answer the question “ is the query's identity among this set? ”. The insight is a reminiscence of the classical Mastermind game: by iteratively analyzing the network responses when multiple random samples of k gallery elements are compared to the query, we obtain weakly correlated matching scores that, altogether, provide solid cues to infer the most likely identity. In this setting, identification is regarded as a variable selection and regularization problem, with sparse linear regression techniques being used to infer the matching probability with respect to each gallery identity. As main strength, this strategy is highly robust to outlier matching scores, which are known to be a primary error source in biometric recognition. Our experiments were carried out in full versions of two well-known irises near-infrared (CASIA-IrisV4-Thousand) and periocular visible wavelength (UBIRIS.v2) datasets, and confirm that recognition performance can be solidly boosted-up by the proposed algorithm, when compared with the traditional working modes of CNNs in biometrics.
Hugo Proença 0001, João C. Neves 0001
IEEE Trans. Inf. Forensics Secur.2
2018 Deep-PRWIS: Periocular Recognition Without the Iris and Sclera Using Deep Learning Frameworks
abstract
This paper is based on a disruptive hypothesis for periocular biometrics-in visible-light data, the recognition performance is optimized when the components inside the ocular globe (the iris and the sclera) are simply discarded, and the recognizer's response is exclusively based on the information from the surroundings of the eye. As a major novelty, we describe a processing chain based on convolution neural networks (CNNs) that defines the regions-of-interest in the input data that should be privileged in an implicit way, i.e., without masking out any areas in the learning/test samples. By using an ocular segmentation algorithm exclusively in the learning data, we separate the ocular from the periocular parts. Then, we produce a large set of “multi-class” artificial samples, by interchanging the periocular and ocular parts from different subjects. These samples are used for data augmentation purposes and feed the learning phase of the CNN, always considering as label the ID of the periocular part. This way, for every periocular region, the CNN receives multiple samples of different ocular classes, forcing it to conclude that such regions should not be considered in its response. During the test phase, samples are provided without any segmentation mask and the network naturally disregards the ocular components, which contributes for improvements in performance. Our experiments were carried out in full versions of two widely known data sets (UBIRIS.v2 and FRGC) and show that the proposed method consistently advances the state-of-the-art performance in the closed-world setting, reducing the EERs in about 82% (UBIRIS.v2) and 85% (FRGC) and improving the Rank-1 over 41% (UBIRIS.v2) and 12% (FRGC).
Hugo Proença 0001, João C. Neves 0001
IEEE Trans. Inf. Forensics Secur.2
2017 IRINA: Iris Recognition (Even) in Inaccurately Segmented Data
abstract
The effectiveness of current iris recognition systems depends on the accurate segmentation and parameterisation of the iris boundaries, as failures at this point misalign the coefficients of the biometric signatures. This paper describes IRINA, an algorithm for Iris Recognition that is robust against INAccurately segmented samples, which makes it a good candidate to work in poor-quality data. The process is based in the concept of corresponding patch between pairs of images, that is used to estimate the posterior probabilities that patches regard the same biological region, even in case of segmentation errors and non-linear texture deformations. Such information enables to infer a free-form deformation field (2D registration vectors) between images, whose first and second-order statistics provide effective biometric discriminating power. Extensive experiments were carried out in four datasets (CASIA-IrisV3-Lamp, CASIA-IrisV4-Lamp, CASIA-IrisV4-Thousand and WVU) and show that IRINA not only achieves state-of-the-art performance in good quality data, but also handles effectively severe segmentation errors and large differences in pupillary dilation/constriction.
Hugo Proença 0001, João C. Neves 0001
CVPR2
2017 Exploiting Data Redundancy for Error Detection in Degraded Biometric Signatures Resulting From in the Wild Environments
abstract
An error-correcting code (ECC) is a process of adding redundant data to a message, such that it can be recovered by a receiver even if a number of errors are introduced in transmission. Inspired by the principles of ECC, we introduce a method capable of detecting degraded features in biometric signatures by exploiting feature correlation. The main novelty is that, unlike existing biometric cryptosystems, the proposed method works directly on the biometric signature. Our approach performs a redundancy analysis of non-degraded data to build an undirected graphical model (Markov Random Field), whose energy minimization determines the sequence of degraded components of the biometric sample. Experiments carried out in different biometric traits ascertain the improvements attained when disregarding degraded features during the matching phase. Also, we stress that the proposed method is general enough to work in different classification methods, such as CNNs.
João C. Neves 0001, Hugo Proença 0001
FG1
2017 Soft Biometrics: Globally Coherent Solutions for Hair Segmentation and Style Recognition Based on Hierarchical MRFs
abstract
Markov Random Fields (MRFs) are a popular tool in many computer vision problems and faithfully model a broad range of local dependencies. However, rooted in the Hammersley-Clifford theorem, they face serious difficulties in enforcing the global coherence of the solutions without using too high order cliques that reduce the computational effectiveness of the inference phase. Having this problem in mind, we describe a multi-layered (hierarchical) architecture for MRFs that is based exclusively in pairwise connections and typically produces globally coherent solutions, with 1) one layer working at the local (pixel) level, modeling the interactions between adjacent image patches; and 2) a complementary layer working at the object (hypothesis) level pushing toward globally consistent solutions. During optimization, both layers interact into an equilibrium state that not only segments the data, but also classifies it. The proposed MRF architecture is particularly suitable for problems that deal with biological data (e.g., biometrics), where the reasonability of the solutions can be objectively measured. As test case, we considered the problem of hair / facial hair segmentation and labeling, which are soft biometric labels useful for human recognition in-the-wild. We observed performance levels close to the state-of-the-art at a much lower computational cost, both in the segmentation and classification (labeling) tasks.
Hugo Proença 0001, João C. Neves 0001
IEEE Trans. Inf. Forensics Secur.2
2016 Visible-wavelength iris/periocular imaging and recognition surveillance environments
Hugo Proença 0001, João C. Neves 0001
Image Vis. Comput.2
2016 Joint Head Pose/Soft Label Estimation for Human Recognition In-The-Wild
abstract
Soft biometrics have been emerging to complement other traits and are particularly useful for poor quality data. In this paper, we propose an efficient algorithm to estimate human head poses and to infer soft biometric labels based on the 3D morphology of the human head. Starting by considering a set of pose hypotheses, we use a learning set of head shapes synthesized from anthropometric surveys to derive a set of 3D head centroids that constitutes a metric space. Next, representing queries by sets of 2D head landmarks, we use projective geometry techniques to rank efficiently the joint 3D head centroids/pose hypotheses according to their likelihood of matching each query. The rationale is that the most likely hypotheses are sufficiently close to the query, so a good solution can be found by convex energy minimization techniques. Once a solution has been found, the 3D head centroid and the query are assumed to have similar morphology, yielding the soft label. Our experiments point toward the usefulness of the proposed solution, which can improve the effectiveness of face recognizers and can also be used as a privacy-preserving solution for biometric recognition in public environments.
Hugo Proença 0001, João C. Neves 0001, Silvio Barra, Tiago Marques, Juan Carlos Moreno
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Dynamic camera scheduling for visual surveillance in crowded scenes using Markov random fields
abstract
The use of pan-tilt-zoom (PTZ) cameras for capturing high-resolution data of human-beings is an emerging trend in surveillance systems. However, this new paradigm entails additional challenges, such as camera scheduling, that can dramatically affect the performance of the system. In this paper, we present a camera scheduling approach capable of determining - in real-time - the sequence of acquisitions that maximizes the number of different targets obtained, while minimizing the cumulative transition time. Our approach models the problem as an undirected graphical model (Markov random field, MRF), which energy minimization can approximate the shortest tour to visit the maximum number of targets. A comparative analysis with the state-of-the-art camera scheduling methods evidences that our approach is able to improve the observation rate while maintaining a competitive tour time.
João C. Neves 0001, Hugo Proença 0001
AVSS1
2015 Do we need a perfect ground-truth for benchmarking Internet traffic classifiers?
abstract
The classification of Internet traffic using supervised or semi-supervised statistical learning techniques, both for anomaly detection and identification of Internet applications, has been impaired by difficulties in obtaining a reliable ground-truth, required both to train the classifier and to evaluate its performance. A perfect ground-truth is increasingly difficult, or sometimes impossible, to obtain due to the growing percentage of cyphered traffic, the sophistication of network attacks, and the constant updates of Internet applications. In this paper, we study the impact of the ground-truth on training the classifier and estimating its performance measures. We show both theoretically and through simulation that ground-truth imperfections can severely bias the performance estimates. We then propose a latent class model that overcomes this problem by combining estimates of several classifiers over the same dataset. The model is evaluated using a high-quality dataset that includes the most representative Internet applications and network attacks. The results show that our latent class model produces very good performance estimates under mild levels of ground-truth imperfection, and can thus be used to correctly benchmark Internet traffic classifiers when only an imperfect ground-truth is available.
Maria Rosário de Oliveira, João C. Neves 0001, Rui Valadas, Paulo Salvador 0001
INFOCOM2
2015 Face recognition: handling data misalignments implicitly by fusion of sparse representations
abstract
Sparse representations for classification (SRC) are considered a relevant advance to the biometrics field, but are particularly sensitive to data misalignments. In previous studies, such misalignments were compensated for by finding appropriate geometric transforms between the elements in the dictionary and the query image, which is costly in terms of computational burden. This study describes an algorithm that compensates for data misalignments in SRC in an implicit way, that is, without finding/applying any geometric transform at every recognition attempt. The authors' study is based on three concepts: (i) sparse representations; (ii) projections on orthogonal subspaces; and (iii) discriminant locality preserving with maximum margin projections. When compared with the classical SRC algorithm, apart from providing slightly better performance, the proposed method is much more robust against global/local data misalignments. In addition, it attains performance close to the state‐of‐the‐art algorithms at a much lower computational cost, offering a potential solution for real‐time scenarios and large‐scale applications.
Hugo Proença 0001, João C. Neves 0001, Juan Carlos Briceño
IET Comput. Vis.2
2014 Segmenting the periocular region using a hierarchical graphical model fed by texture / shape information and geometrical constraints
abstract
Using the periocular region for biometric recognition is an interesting possibility: this area of the human body is highly discriminative among subjects and relatively stable in appearance. In this paper, the main idea is that improved solutions for defining the periocular region-of-interest and better pose / gaze estimates can be obtained by segmenting (labelling) all the components in the periocular vicinity. Accordingly, we describe an integrated algorithm for labelling the periocular region, that uses a unique model to discriminate between seven components in a single-shot: iris, sclera, eyelashes, eyebrows, hair, skin and glasses. Our solution fuses texture / shape descriptors and geometrical constraints to feed a two-layered graphical model (Markov Random Field), which energy minimization provides a robust solution against uncontrolled lighting conditions and variations in subjects pose and gaze.
Hugo Proença 0001, João C. Neves 0001, Gil Melfe Mateus Santos
IJCB2