EDBT 2026 Demo / reviewers in the wild / expert
Vitor Albiero
dblp:202/4816 · also Vítor Albiero
· DBLP profile ↗
13ranked-venue papers
8as first author
7since 2021 · last 2027
0000-0002-8607-9914ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 5 since 2021Security and privacy · 5 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Evaluating AI models' capability to automate voice phishing attacksabstractVoice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults’ susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play.AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the “relative-in-distress” category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-automated voice phishing. Caller persuasiveness was the strongest predictor of compliance and certain models (most notably Sesame) achieved ratings comparable to human voices, or sometimes even slightly surpassing them. Our economic analysis suggests that while human-operated vishing is unprofitable at US wages, AI-powered vishing appears to be economically viable for several models. The primary risk of present-day AI-enabled vishing thus lies in the economics of automation rather than novel or “superhuman” persuasive techniques, though these cannot be ruled out for future systems. This raises significant concerns for the design of AI systems, consumer protection, and model release policies. Fred Heiding, Claudio Mayrink Verdun, Simon Lermen, Andrew Kao, Vitor Albiero, Lauren Deason, Irina-Elena Veliche, Christine Lehane |
Expert Syst. Appl. | 5 |
| 2025 | Automated Red Teaming with GOAT: the Generative Offensive Agent TesterabstractRed teaming aims to assess how large language models (LLMs) can produce content that violates norms, policies, and rules set forth during their safety training. However, most existing automated methods in literature are not representative of the way common users exploit the multi-turn conversational nature of AI models. While manual testing addresses this gap, it is an inefficient and often expensive process. To address these limitations, we introduce the Generative Offensive Agent Tester (GOAT), an automated agentic red teaming system that simulates plain language adversarial conversations while leveraging multiple adversarial prompting techniques to identify vuLnerabilities in LLMs. We instantiate GOAT with 7 red teaming attacks by prompting a general purpose model in a way that encourages reasoning through the choices of methods available, the current target model’s response, and the next steps. Our approach is designed to be extensible and efficient, allowing human testers to focus on exploring new areas of risk while automation covers the scaled adversarial stress-testing of known risk territory. We present the design and evaluation of GOAT, demonstrating its effectiveness in identifying vulnerabilities in state-of-the-art LLMs, with an ASR@10 of 96% against smaller models such as Llama 3.1 8B, and 91% against Llama 3.1 70B and 94% for GPT-4o when evaluated against larger models on the JailbreakBench dataset. Maya Pavlova, Erik Brinkman, Krithika Iyer, Vitor Albiero, Joanna Bitton, Hailey Nguyen, Cristian Canton, Ivan Evtimov, Aaron Grattafiori |
ICML | 4 |
| 2023 | CAST: Conditional Attribute Subsampling Toolkit for Fine-grained EvaluationabstractThorough evaluation is critical for developing models that are fair and robust. In this work, we describe the Conditional Attribute Subsampling Toolkit (CAST) for selecting data subsets for fine-grained scientific evaluations. Our toolkit efficiently filters data given an arbitrary number of conditions for metadata attributes. The purpose of the toolkit is to allow researchers to easily to evaluate models on targeted test distributions. The functionality of CAST is demonstrated on the WebFace42M face Recognition dataset. We calculate over 50 attributes for this dataset including race, image quality, facial features, and accessories. Using our toolkit, we create over a hundred test sets conditioned on one or multiple attributes. Results are presented for subsets of various demographics and image quality ranges. Using eleven different subsets, we build a face recognition 1:1 verification benchmark called C11 that exclusively contains pairs that are near the decision threshold. Evaluation on C11 with state-of-the-art methods demonstrates the suitability of the proposed benchmark. The toolkit is publicly available at https://github.com/WesRobbins/CAST. Wes Robbins, Steven Zhou, Aman Bhatta, Chad Mello, Vitor Albiero, Kevin W. Bowyer, Terrance E. Boult |
WACV | 5 |
| 2022 | Face Regions Impact Recognition Accuracy Differently Across DemographicsabstractVariation in face recognition accuracy across demographic groups has attracted attention from news media, civil liberties advocates and academic researchers. The problem is challenging, in that both the impostor distribution (matches across different people) and the genuine distribution (matches across same people) may vary across demographic groups. Simple answers such as balancing the number of subjects and images in the training data do not have a substantial impact on demographic accuracy disparities. We present the first investigation into whether parts of the face - such as eyes, nose, mouth - show the same accuracy differences across demographic groups as are seen with matching the whole face. We show that matching focused on different parts of the face may result in opposite accuracy differences across demographics. For example, using the eye region for face matching results in Caucasian males having a better impostor distribution (lower similarity scores) than Caucasian females, but using the nose regionfor face matching results in Caucasian females having a better impostor distribution. We also show that it is possible to select face region(s) that effectively minimize the difference in the impostor or genuine distributions across at least some demographics. Our results suggest that a new pathway to reducing accuracy disparity across demographic groups may be to weight the parts of the face differently in matching. Vitor Albiero, Kevin W. Bowyer, Michael C. King |
IJCB | 1 |
| 2022 | Gendered Differences in Face Recognition Accuracy Explained by Hairstyles, Makeup, and Facial MorphologyabstractMedia reports have accused face recognition of being “biased”, “sexist” and “racist”. There is consensus in the research literature that face recognition accuracy is lower for females, who often have both a higher false match rate and a higher false non-match rate. However, there is little published research aimed at identifying the cause of lower accuracy for females. For instance, the 2019 Face Recognition Vendor Test that documents lower female accuracy across a broad range of algorithms and datasets also lists “Analyze cause and effect” under the heading “What we did not do”. We present the first experimental analysis to identify major causes of lower face recognition accuracy for females on datasets where previous research has observed this result. Controlling for equal amount of visible face in the test images mitigates the apparent higher false non-match rate for females. Additional analysis shows that makeup-balanced datasets further improves females to achieve lower false non-match rates. Finally, a clustering experiment suggests that images of two different females are inherently more similar than of two different males, potentially accounting for a difference in false match rates. Vitor Albiero, Kai Zhang 0052, Michael C. King, Kevin W. Bowyer |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | img2pose: Face Alignment and Detection via 6DoF, Face Pose EstimationabstractWe propose real-time, six degrees of freedom (6DoF), 3D face pose estimation without face detection or landmark localization. We observe that estimating the 6DoF rigid transformation of a face is a simpler problem than facial landmark detection, often used for 3D face alignment. In addition, 6DoF offers more information than face bounding box labels. We leverage these observations to make multiple contributions: (a) We describe an easily trained, efficient, Faster R-CNN–based model which regresses 6DoF pose for all faces in the photo, without preliminary face detection. (b) We explain how pose is converted and kept consistent between the input photo and arbitrary crops created while training and evaluating our model. (c) Finally, we show how face poses can replace detection bounding box training labels. Tests on AFLW2000-3D and BIWI show that our method runs at real-time and outperforms state of the art (SotA) face pose estimators. Remarkably, our method also surpasses SotA models of comparable complexity on the WIDER FACE detection benchmark, despite not been optimized on bounding box labels. Vitor Albiero, Xi Yin 0001, Guan Pang, Tal Hassner |
CVPR | 1 |
| 2021 | Does Face Recognition Error Echo Gender Classification Error?abstractThis paper is the first to explore the question of whether images that are classified incorrectly by a face analytics algorithm (e.g., gender classification) are any more or less likely to participate in an image pair that results in a face recognition error. We analyze results from three different gender classification algorithms (one open-source and two commercial), and two face recognition algorithms (one open-source and one commercial), on image sets representing four demographic groups (African-American female and male, Caucasian female and male). For impostor image pairs, our results show that pairs in which one image has a gender classification error have a better impostor distribution than pairs in which both images have correct gender classification, and so are less likely to generate a false match error. For genuine image pairs, our results show that individuals whose images have a mix of correct and incorrect gender classification have a worse genuine distribution (increased false non-match rate) compared to individuals whose images consistently have correct gender classification. Thus, compared to images that generate correct gender classification, images with gender classification error have a lower false match rate and a higher false non-match rate. Vitor Albiero, Michael C. King, Kevin W. Bowyer |
IJCB | 2 |
| 2020 | Is Face Recognition Sexist? No, Gendered Hairstyles and Biology Are
Vitor Albiero, Kevin W. Bowyer |
BMVC | 1 |
| 2020 | Identity Document to Selfie Face Matching Across AdolescenceabstractMatching live images (“selfies”) to images from ID documents is a problem that can arise in various applications. A challenging instance of the problem arises when the face image on the ID document is from early adolescence and the live image is from later adolescence. We explore this problem using a private dataset called Chilean Young Adult (CHIYA) dataset, where we match live face images taken at age 18-19 to face images on scanned ID documents created at ages 9 to 18. State-of-the-art deep learning face matchers (e.g., ArcFace) have relatively poor accuracy for document-to-selfie face matching. To achieve higher accuracy, we fine-tune the best available open-source model with triplet loss for a few-shot learning. Experiments show that our approach achieves higher accuracy than the DocFace+ model recently developed for this problem. Our fine-tuned model was able to improve the true acceptance rate for the most difficult (largest age span) subset from 62.92% to 96.67% at a false acceptance rate of 0.01%. Our fine-tuned model is available for use by other researchers. Vitor Albiero, Nisha Srinivas, Esteban Villalobos, Jorge Perez-Facuse, Roberto Rosenthal, Domingo Mery, Karl Ricanek, Kevin W. Bowyer |
IJCB | 1 |
| 2020 | How Does Gender Balance In Training Data Affect Face Recognition Accuracy?abstractDeep learning methods have greatly increased the accuracy of face recognition, but an old problem still persists: accuracy is usually higher for men than women. It is often speculated that lower accuracy for women is caused by under-representation in the training data. This work investigates female under-representation in the training data is truly the cause of lower accuracy for females on test data. Using a state-of-the-art deep CNN, three different loss functions, and two training datasets, we train each on seven subsets with different male/female ratios, totaling forty two trainings, that are tested on three different datasets. Results show that (1) gender balance in the training data does not translate into gender balance in the test accuracy, (2) the “gender gap” in test accuracy is not minimized by a gender-balanced training set, but by a training set with more male images than female images, and (3) training to minimize the accuracy gap does not result in highest female, male or average accuracy. Vitor Albiero, Kai Zhang 0052, Kevin W. Bowyer |
IJCB | 1 |
| 2020 | Does Face Recognition Accuracy Get Better With Age? Deep Face Matchers Say NoabstractPrevious studies generally agree that face recognition accuracy is higher for older persons than for younger persons. But most previous studies were before the wave of deep learning matchers, and most considered accuracy only in terms of the verification rate for genuine pairs. This paper investigates accuracy for age groups 16-29, 30-49 and 50-70, using three modern deep CNN matchers, and considers differences in the impostor and genuine distributions as well as verification rates and ROC curves. We find that accuracy is lower for older persons and higher for younger persons. In contrast, a pre deep learning matcher on the same dataset shows the traditional result ofhigher accuracy for older persons, although its overall accuracy is much lower than that of the deep learning matchers. Comparing the impostor and genuine distributions, we conclude that impostor scores have a larger effect than genuine scores in causing lower accuracy for the older age group. We also investigate the effects of training data across the age groups. Our results show that fine-tuning the deep CNN models on additional images ofolder persons actually lowers accuracy for the older age group. Also, we fine-tune and train from scratch two models using age-balanced training datasets, and these results also show lower accuracy for older age group. These results argue that the lower accuracy for the older age group is not due to imbalance in the original training data. Vitor Albiero, Kevin W. Bowyer, Kushal Vangara, Michael C. King |
WACV | 1 |
| 2018 | Multi-Label Action Unit Detection on Multiple Head Poses with Dynamic Region LearningabstractThis paper presents a multi-label Action Unit (AU) detection method applied on multi-pose facial images. Action Unit detection on multiple head poses is an issue that robust AU detectors must deal with, as it is uncommon for a person to maintain always the same pose when displaying facial expressions. To this end, this work proposes a region learning approach, that dynamically creates regions of interest inside a convolutional neural network (CNN) using facial landmark points. The dynamic region learning (DRL) ensures that each AU is in the center of the region, and also follows the head pose movement. The DRL is built on top of the VGG-Face network, and transfer-learning is used to start the training. The experiments were conducted on the Facial Expression Recognition and Analysis Challenge (FERA 2017) database, which contains nine different head poses. The results show that the dynamic region learning is able to adapt to the nine poses in the database, improving the state-of-the-art with an an average F1-score of 0.582. Vitor Albiero, Olga R. P. Bellon, Luciano Silva |
ICIP | 1 |
| 2017 | AUMPNet: Simultaneous Action Units Detection and Intensity Estimation on Multipose Facial Images Using a Single Convolutional Neural NetworkabstractThis paper presents an unified convolutional neural network (CNN), named AUMPNet, to perform both Action Units (AUs) detection and intensity estimation on facial images with multiple poses. Although there are a variety of methods in the literature designed for facial expression analysis, only few of them can handle head pose variations. Therefore, it is essential to develop new models to work on non-frontal face images, for instance, those obtained from unconstrained environments. In order to cope with problems raised by pose variations, an unique CNN, based on region and multitask learning, is proposed for both AU detection and intensity estimation tasks. Also, the available head pose information was added to the multitask loss as a constraint to the network optimization, pushing the network towards learning better representations. As opposed to current approaches that require ad hoc models for every single AU in each task, the proposed network simultaneously learns AU occurrence and intensity levels for all AUs. The AUMPNet was evaluated on an extended version of the BP4D-Spontaneous database, which was synthesized into nine different head poses and made available to FG 2017 Facial Expression Recognition and Analysis Challenge (FERA 2017) participants. The achieved results surpass the FERA 2017 baseline, using the challenge metrics, for AU detection by 0.054 in F1-score and 0.182 in ICC(3, 1) for intensity estimation. Julio Cesar Batista, Vitor Albiero, Olga R. P. Bellon, Luciano Silva |
FG | 2 |