Vasco Lopes

dblp:220/8448 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-5577-1094ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
abstract
Video Anomaly Understanding (VAU) is a novel task focused on describing unusual occurrences in videos. Despite growing interest, the evaluation of VAU remains an open challenge. Existing benchmarks rely on n-gram-based metrics (e.g., BLEU, ROUGE-L) or LLM-based evaluation. The first fails to capture the rich, free-form, and visually grounded nature of LVLM responses, while the latter focuses on assessing language quality over factual relevance, often resulting in subjective judgments that are misaligned with human perception. In this work, we address this issue by proposing FineVAU, a new benchmark for VAU that shifts the focus towards rich, fine-grained and domain-specific understanding of anomalous videos. We formulate VAU as a three-fold problem, with the goal of comprehensively understanding key descriptive elements of anomalies in video: events (What), participating entities (Who) and location (Where). Our benchmark introduces a) FV-Score, a novel, human-aligned evaluation metric that assesses the presence of critical visual elements in LVLM answers, providing interpretable, fine-grained feedback; and b) FineW³, a novel, comprehensive dataset curated through a structured and fully automatic procedure that augments existing human annotations with high quality, fine-grained visual information. Human evaluation reveals that our proposed metric has a superior alignment with human perception of anomalies in comparison to current approaches. Detailed experiments on FineVAU unveil critical limitations in LVLM's ability to perceive anomalous events that require spatial and fine-grained temporal understanding, despite strong performance on coarse grain, static information, and events with strong visual cues.
João Pereira 0007, Vasco Lopes, João Neves 0006, David Semedo
AAAI2
2026 ExplainablePR: Periocular recognition with interpretability in sight
abstract
Understanding and justifying the decisions generated by automated recognizers has been a focal point of biometrics research, driven by the need for systems that are not only effective but transparent and accountable. Clear explanations improve user trust and system credibility, while addressing the growing concerns regarding the role of AI in our daily lives. In this work, we propose a framework that integrates biometric recognition with visual interpretations of the features that contribute the most to a match/non-match decision. Our method leverages adversarial generative models to create a set composed exclusively of ``genuine" image pairs. From these, the most similar candidates to a given query are identified and, assuming enough similarity in phase between the query and the retrieved pairs, the pixel-wise differences highlight the image regions that supported the decision. Comparative evaluations against established interpretability techniques (SHAP, LIME, and Saliency Maps) as well as other state-of-the-art fine-grained visual recognizers demonstrate that our framework delivers intuitive visual explanations without compromising recognition performance. These findings underscore our method's potential to enhance the transparency and credibility of biometric systems while maintaining high accuracy.
João Brito, Vasco Lopes, Bruno Degardin, Hugo Proença 0001
Image Vis. Comput.2
2025 VM-TAPS: View-specific Memory with Temporal and Scale Awareness Framework for Video-based Cross-View Person Re-Identification
abstract
Reliable aerial-ground video-based person re-identification (ReID) remains a challenge due to severe changes in data quality and features, such as viewpoint disparities, resolution drops, and cross-camera appearance inconsistency. This paper presents VM-TAPS, a lightweight and modular extension to the well-known TF-CLIP framework, designed to increase the robustness of ReID, without requiring end-to-end backbone retraining. When compared to its ancestor, VM-TAPS' novelties are five-fold: 1) View-Specific Processing Layers to normalize camera-dependent biases; 2) Scale-Aware Feature Adaptation for resolution-invariant feature fusion; 3) a View-Aware Memory Bank enabling long-range identity context; 4) a Motion Pattern Analyzer capturing temporal dynamics; and (5) Cross-View Interaction Modules that harmonize multi-view feature spaces. Despite adding fewer than two million parameters, VM-TAPS achieves +4.97% Rank-1 and +3.08% mAP gains over TF-CLIP on the challenging AG-VPReID2025 benchmark. At 80m and 120m altitudes, it sets a new performance baseline of 73.68%/75.73% and 69.45%/71.63% (Rank-1/mAP), respectively. All components are trained with frozen CLIP visual encoders in the early stages, enabling efficient and stable convergence. Our results support that the carefully disentanglement of viewpoint, scale, motion and memory factors substantially increases the robustness of cross-view ReID under real-world conditions.
Md. Rashidunnabi, Kailash A. Hambarde, João C. Neves 0001, Vasco Lopes, Hugo Proença 0001
IJCB4
2025 Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding
abstract
Large Vision-Language Models (LVLMs) demonstrate remarkable performance in short-video tasks such as video question answering, but struggle in long-video understanding. The linear frame sampling strategy, conventionally used by LVLMs, fails to account for the non-linear distribution of key events in video data, often introducing redundant or irrelevant information in longer contexts while risking the omission of critical events in shorter ones. To address this, we propose Self-ReS, a non-linear spatiotemporal self-reflective sampling method that dynamically selects key video fragments based on user prompts. Unlike prior approaches, Self-ReS leverages the inherently sparse attention maps of LVLMs to define reflection tokens, enabling relevance-aware token selection without requiring additional training or external modules. Experiments demonstrate that Self-ReS can be seamlessly integrated into strong base LVLMs, improving long-video task accuracy and achieving up to 46% faster inference speed within the same GPU memory budget. The code can be found here: https://github.com/jaca-pereira/Self-ReS.
João Pereira 0007, Vasco Lopes, David Semedo, João Neves 0006
ICME2
2025 Toward Less Constrained Macro-Neural Architecture Search
abstract
Networks found with neural architecture search (NAS) achieve the state-of-the-art performance in a variety of tasks, out-performing human-designed networks. However, most NAS methods heavily rely on human-defined assumptions that constrain the search: architecture's outer skeletons, number of layers, parameter heuristics, and search spaces. In addition, common search spaces consist of repeatable modules (cells) instead of fully exploring the architecture's search space by designing entire architectures (macro-search). Imposing such constraints requires deep human expertise and restricts the search to predefined settings. In this article, we propose less constrained macro-neural architecture search (LCMNAS), a method that pushes NAS to less constrained search spaces by performing macro-search without relying on predefined heuristics or bounded search spaces. LCMNAS introduces three components for the NAS pipeline: 1) a method that leverages information about well-known architectures to autonomously generate complex search spaces based on weighted directed graphs (WDGs) with hidden properties; 2) an evolutionary search strategy that generates complete architectures from scratch; and 3) a mixed-performance estimation approach that combines information about architectures at the initialization stage and lower fidelity estimates to infer their trainability and capacity to model complex functions. We present experiments in 14 different datasets showing that LCMNAS is capable of generating both cell and macro-based architectures with minimal GPU computation and state-of-the-art results. Moreover, we conduct extensive studies on the importance of different NAS components in both cell and macro-based settings. The code for reproducibility is publicly available at https://github.com/VascoLopes/LCMNAS.
Vasco Lopes, Luís A. Alexandre
IEEE Trans. Neural Networks Learn. Syst.1
2024 Guided evolutionary neural architecture search with efficient performance estimation
abstract
Neural Architecture Search (NAS) methods have been successfully applied to image tasks with excellent results. However, NAS methods are often complex and tend to converge to local minima as soon as generated architectures yield good results. This paper proposes GEA, a novel approach for guided NAS. GEA guides the evolution by exploring the search space by generating and evaluating several architectures in each generation at initialisation stage using a zero-proxy estimator, where only the highest-scoring architecture is trained and kept for the next generation. Subsequently, GEA continuously extracts knowledge about the search space without increased complexity by generating several off-springs from an existing architecture at each generation. Moreover, GEA forces exploitation of the most performant architectures by descendant generation while simultaneously driving exploration through parent mutation and favouring younger architectures to the detriment of older ones. Experimental results demonstrate the effectiveness of the proposed method, and extensive ablation studies evaluate the importance of different parameters. Results show that GEA achieves competitive results on all data sets of NAS-Bench-101, NAS-Bench-201 and TransNAS-Bench-101 benchmarks, as well as in the DARTS search space.
Vasco Lopes, Miguel Santos, Bruno Degardin, Luís A. Alexandre
Neurocomputing1
2024 Manas: multi-agent neural architecture search
Vasco Lopes, Fabio Maria Carlucci, Pedro M. Esperança, Marco Singh, Antoine Yang, Victor Gabillon, Zewei Chen, Jun Wang 0012
Mach. Learn.1
2023 Are neural architecture search benchmarks well designed? A deeper look into operation importance
abstract
Neural Architecture Search (NAS) benchmarks significantly improved the capability of developing and comparing NAS methods while at the same time drastically reduced the computational overhead by providing meta-information of trained neural networks. However, tabular benchmarks have several drawbacks that can hinder fair comparisons and provide unreliable results. These usually focus on a small pool of operations in heavily constrained search spaces – usually cell-based neural networks with pre-defined outer-skeletons. In this work, we conducted an empirical analysis of the widely used NAS-Bench-101, NAS-Bench-201 and TransNAS-Bench-101 benchmarks in terms of their generability and how different operations influence the performance of the generated architectures. We found that only a subset of the operation pool is required to generate architectures close to the upper-bound of the performance range. More, the performance distribution is negatively skewed, with many architectures clustered near the upper accuracy bound. Further experiments revealed that convolution layers have the highest impact on the architecture's performance and that specific combinations of operations favor top-scoring architectures. Overall, our results demonstrate the need for benchmarks with greater operation diversity and less constrained search spaces. We provide suggestions for improving future benchmark design and evaluation of NAS methods when using existing benchmarks. The code used to conduct the evaluations is available at https://github.com/VascoLopes/NAS-Benchmark-Evaluation.
Vasco Lopes, Bruno Degardin, Luís A. Alexandre
Inf. Sci.1
2023 ATOM: Self-supervised human action recognition using atomic motion representation learning
abstract
Self-supervised learning (SSL) is a promising method for gaining perception and common sense from unlabelled data. Existing approaches to analyzing human body skeletons address the problem similar to SSL models for image and video understanding, but pixel data is far more challenging than coordinates. This paper presents ATOM, an SSL model designed for skeleton-based data analysis. Unlike video-based SSL approaches, ATOM leverages atomic movements within skeleton actions to achieve a more fine-grained representation. The proposed architecture predicts the action order at the frame level, leading to improved perceptions and representations of each action. ATOM outperforms state-of-the-art approaches in two well-known datasets (NTU RGB + D and NTU-120 RGB + D), and its weight transferability enables performance improvements on supervised and semi-supervised tasks, up to 4.4% (3.3% p.p.) and 14.1% (6.3% p.p.), respectively, in Top-1 Accuracy.
Bruno Degardin, Vasco Lopes, Hugo Proença 0001
Image Vis. Comput.2
2022 Controlling Robots using Image Analysis and a Consortium Blockchain
abstract
Blockchain is a disruptive technology, normally used within financial applications, however, it can be very beneficial in certain robotic contexts, such as when an immutable register of events is required. Among the several properties of Blockchain that can be useful within robotic environments, we find not just immutability but also data decentralization, irreversibility, accessibility and non-repudiation. In this paper, we propose an architecture that uses blockchain as a ledger, and smart-contracts for robotic control by using oracles to process data. We show how to register events in a secure way, how it is possible to use smart-contracts to control robots and how to interface with external algorithms for image analysis. The proposed architecture is modular and can be used in multiple contexts such as in manufacturing, network control, robot control, and others, since it is easy to integrate, adapt, maintain and extend to new domains, only requiring new tailored smart-contracts.
Vasco Lopes, Luís A. Alexandre, Nuno Pereira 0003
ICARCV1
2022 Generative Adversarial Graph Convolutional Networks for Human Action Synthesis
abstract
Synthesising the spatial and temporal dynamics of the human body skeleton remains a challenging task, not only in terms of the quality of the generated shapes, but also of their diversity, particularly to synthesise realistic body movements of a specific action (action conditioning). In this paper, we propose Kinetic-GAN, a novel architecture that leverages the benefits of Generative Adversarial Networks and Graph Convolutional Networks to synthesise the kinetics of the human body. The proposed adversarial architecture can condition up to 120 different actions over local and global body movements while improving sample quality and diversity through latent space disentanglement and stochastic variations. Our experiments were carried out in three well-known datasets, where Kinetic-GAN notably surpasses the state-of-the-art methods in terms of distribution quality metrics while having the ability to synthesise more than one order of magnitude regarding the number of different actions. Our code and models are publicly available at https://github.com/DegardinBruno/Kinetic-GAN.
Bruno Degardin, João C. Neves 0001, Vasco Lopes, João Brito, Ehsan Yaghoubi, Hugo Proença 0001
WACV3
2021 EPE-NAS: Efficient Performance Estimation Without Training for Neural Architecture Search
Vasco Lopes, Saeid Alirezazadeh, Luís A. Alexandre
ICANN (5)1
2021 An AutoML-based Approach to Multimodal Image Sentiment Analysis
abstract
Sentiment analysis is a research topic focused on analysing data to extract information related to the sentiment that it causes. Applications of sentiment analysis are wide, ranging from recommendation systems, and marketing to customer satisfaction. Recent approaches evaluate textual content using Machine Learning techniques that are trained over large corpora. However, as social media grown, other data types emerged in large quantities, such as images. Sentiment analysis in images has shown to be a valuable complement to textual data since it enables the inference of the underlying message polarity by creating context and connections. Multimodal sentiment analysis approaches intend to leverage information of both textual and image content to perform an evaluation. Despite recent advances, current solutions still flounder in combining both image and textual information to classify social media data, mainly due to subjectivity, inter-class homogeneity and fusion data differences. In this paper, we propose a method that combines both textual and image individual sentiment analysis into a final fused classification based on AutoML, that performs a random search to find the best model. Our method achieved state-of-the-art performance in the B-T4SA dataset, with 95.19% accuracy.
Vasco Lopes, António Gaspar, Luís A. Alexandre, João Cordeiro
IJCNN1
2021 REGINA - Reasoning Graph Convolutional Networks in Human Action Recognition
Bruno Degardin, Vasco Lopes, Hugo Proença 0001
IEEE Trans. Inf. Forensics Secur.2
2020 Auto-Classifier: A Robust Defect Detector Based on an AutoML Head
Vasco Lopes, Luís A. Alexandre
ICONIP (1)1