VLDB 2026 Research / reviewers in the wild / expert
Cristian Canton
dblp:43/950 · also Cristian Canton Ferrer, Cristian Canton-Ferrer
· DBLP profile ↗
22ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-3189-5498ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automated Red Teaming with GOAT: the Generative Offensive Agent TesterabstractRed teaming aims to assess how large language models (LLMs) can produce content that violates norms, policies, and rules set forth during their safety training. However, most existing automated methods in literature are not representative of the way common users exploit the multi-turn conversational nature of AI models. While manual testing addresses this gap, it is an inefficient and often expensive process. To address these limitations, we introduce the Generative Offensive Agent Tester (GOAT), an automated agentic red teaming system that simulates plain language adversarial conversations while leveraging multiple adversarial prompting techniques to identify vuLnerabilities in LLMs. We instantiate GOAT with 7 red teaming attacks by prompting a general purpose model in a way that encourages reasoning through the choices of methods available, the current target model’s response, and the next steps. Our approach is designed to be extensible and efficient, allowing human testers to focus on exploring new areas of risk while automation covers the scaled adversarial stress-testing of known risk territory. We present the design and evaluation of GOAT, demonstrating its effectiveness in identifying vulnerabilities in state-of-the-art LLMs, with an ASR@10 of 96% against smaller models such as Llama 3.1 8B, and 91% against Llama 3.1 70B and 94% for GPT-4o when evaluated against larger models on the JailbreakBench dataset. Maya Pavlova, Erik Brinkman, Krithika Iyer, Vitor Albiero, Joanna Bitton, Hailey Nguyen, Cristian Canton, Ivan Evtimov, Aaron Grattafiori |
ICML | 7 |
| 2024 | Data-Driven but Privacy-Conscious: Pedestrian Dataset De-Identification via Full-Body Person SynthesisabstractThe advent of data-driven technology solutions is accompanied by an increasing concern with data privacy. This is of particular importance for human-centered image recognition tasks, such as pedestrian detection, re-identification, and tracking. To highlight the importance of privacy issues and motivate future research, we motivate and introduce the Pedestrian Dataset De-Identification (PDI) task. PDI evaluates the degree of de-identification and downstream task training performance for a given de-identification method. As a first baseline, we propose IncogniMOT, a two-stage full-body de-identification pipeline based on image synthesis via generative adversarial networks. The first stage replaces target pedestrians with synthetic identities. To improve downstream task performance, we then apply stage two which blends and adapts the synthetic image parts into the data. To demonstrate the effectiveness of IncogniMOT, we generate a fully de-identified version of the MOT17 pedestrian tracking dataset and analyze its application as training data for pedestrian re-identification, detection, and tracking models. Furthermore, we show how our data is able to narrow the synthetic-to-real performance gap in a privacy-conscious manner. Maxim Maximov, Tim Meinhardt, Zoë Papakipos, Caner Hazirbas, Cristian Canton, Laura Leal-Taixé |
FG | 5 |
| 2024 | Fairness-Aware Meta-Learning via Nash BargainingabstractTo address issues of group-level fairness in machine learning, it is natural to adjust model parameters based on specific fairness objectives over a sensitive-attributed validation set. Such an adjustment procedure can be cast within a meta-learning framework. However, naive integration of fairness goals via meta-learning can cause hypergradient conflicts for subgroups, resulting in unstable convergence and compromising model performance and fairness. To navigate this issue, we frame the resolution of hypergradient conflicts as a multi-player cooperative bargaining game. We introduce a two-stage meta-learning framework in which the first stage involves the use of a Nash Bargaining Solution (NBS) to resolve hypergradient conflicts and steer the model toward the Pareto front, and the second stage optimizes with respect to specific fairness goals.
Our method is supported by theoretical results, notably a proof of the NBS for gradient aggregation free from linear independence assumptions, a proof of Pareto improvement, and a proof of monotonic improvement in validation loss. We also show empirical effects across various fairness objectives in six key fairness datasets and two image classification tasks. Yi Zeng 0005, Xuelin Yang, Cristian Canton, Ming Jin 0002, Michael I. Jordan, Ruoxi Jia 0001 |
NeurIPS | 4 |
| 2023 | A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies OthersabstractMachine learning models have been found to learn shortcuts—unintended decision rules that are unable to generalize—undermining models' reliability. Previous works address this problem under the tenuous assumption that only a single shortcut exists in the training data. Real-world images are rife with multiple visual cues from background to texture. Key to advancing the reliability of vision systems is understanding whether existing methods can overcome multiple shortcuts or struggle in a Whac-A-Mole game, i.e., where mitigating one shortcut amplifies reliance on others. To address this shortcoming, we propose two benchmarks: 1) UrbanCars, a dataset with precisely controlled spurious cues, and 2) ImageNet-W, an evaluation set based on ImageNet for watermark, a shortcut we discovered affects nearly every modern vision model. Along with texture and background, ImageNet-W allows us to study multiple shortcuts emerging from training on natural images. We find computer vision models, including large foundation models—regardless of training set, architecture, and supervision—struggle when multiple shortcuts are present. Even methods explicitly designed to combat shortcuts struggle in a Whac-A-Mole dilemma. To tackle this challenge, we propose Last Layer Ensemble, a simple-yet-effective method to mitigate multiple shortcuts without Whac-A-Mole behavior. Our results surface multi-shortcut mitigation as an overlooked challenge critical to advancing the reliability of vision systems. The datasets and code are released: https://github.com/facebookresearch/Whac-A-Mole. Zhiheng Li 0002, Ivan Evtimov, Albert Gordo, Caner Hazirbas, Tal Hassner, Cristian Canton, Chenliang Xu, Mark Ibrahim |
CVPR | 6 |
| 2023 | VPA: Fully Test-Time Visual Prompt AdaptationabstractTextual prompt tuning has demonstrated significant performance improvements in adapting natural language processing models to a variety of downstream tasks by treating hand-engineered prompts as trainable parameters. Inspired by the success of textual prompting, several studies have investigated the efficacy of visual prompt tuning. In this work, we present Visual Prompt Adaptation (VPA), the first framework that generalizes visual prompting with test-time adaptation. VPA introduces a small number of learnable tokens, enabling fully test-time and storage-efficient adaptation without necessitating source-domain information. We examine our VPA design under diverse adaptation settings, encompassing single-image, batched-image, and pseudo-label adaptation. We evaluate VPA on multiple tasks, including out-of-distribution (OOD) generalization, corruption robustness, and domain adaptation. Experimental results reveal that VPA effectively enhances OOD generalization by 3.3% across various models, surpassing previous test-time approaches. Furthermore, we show that VPA improves corruption robustness by 6.5% compared to strong baselines. Finally, we demonstrate that VPA also boosts domain adaptation performance by relatively 5.2%. Our VPA also exhibits marked effectiveness in improving the robustness of zero-shot recognition for vision-language models. Mark Ibrahim, Melissa Hall, Ivan Evtimov, Z. Morley Mao, Cristian Canton, Caner Hazirbas |
ACM Multimedia | 6 |
| 2022 | Generating High Fidelity Data from Low-density Regions using Diffusion ModelsabstractOur work focuses on addressing sample deficiency from low-density regions of data manifold in common image datasets. We leverage diffusion process based generative models to synthesize novel images from low-density regions. We observe that uniform sampling from diffusion models predominantly samples from high-density regions of the data manifold. Therefore, we modify the sampling process to guide it towards low-density regions while simulta-neously maintaining the fidelity of synthetic data. We rigorously demonstrate that our process successfully generates novel high fidelity samples from low-density regions. We further examine generated samples and show that the model does not memorize low-density data and indeed learns to generate novel samples from low-density regions. Vikash Sehwag, Caner Hazirbas, Albert Gordo, Firat Ozgenel, Cristian Canton |
CVPR | 5 |
| 2022 | Pixinwav: Residual Steganography for Hiding Pixels in AudioabstractSteganography comprises the mechanics of hiding data in a host media that may be publicly available. While previous works focused on unimodal setups (e.g., hiding images in images, or hiding audio in audio), PixInWav targets the multimodal case of hiding images in audio. To this end, we propose a novel residual architecture operating on top of short-time discrete cosine transform (STDCT) audio spectrograms. Among our results, we find that the residual steganography setup we propose allows an encoding of the hidden image that is independent from the host audio without compromising quality. Accordingly, while previous works require both host and hidden signals to hide a signal, PixInWav can encode images offline—which can be later hidden, in a residual fashion, into any audio signal. Margarita Geleta, Cristina Punti, Kevin McGuinness, Jordi Pons, Cristian Canton, Xavier Giró-i-Nieto |
ICASSP | 5 |
| 2022 | Affective Signals in a Social Media Recommender SystemabstractPeople come to social media to satisfy a variety of needs, such as being informed, entertained and inspired, or connected to their friends and community. Hence, to design a ranking function that gives useful and personalized post recommendations, it would be helpful to be able to predict the affective response a user may have to a post (e.g., entertained, informed, angered). This paper describes the challenges and solutions we developed to apply Affective Computing to social media recommendation systems. Jane Dwivedi-Yu, Yi-Chia Wang, Lijing Qin, Cristian Canton, Alon Y. Halevy |
KDD | 4 |
| 2021 | Deep Poisoning: Towards Robust Image Data Sharing against Visual DisclosureabstractDue to respectively limited training data, different entities addressing the same vision task based on certain sensitive images may not train a robust deep network. This paper introduces a new vision task where various entities share task-specific image data to enlarge each other's training data volume without visually disclosing sensitive contents (e.g. illegal images). Then, we present a new structure-based training regime to enable different entities learn task-specific and reconstruction-proof image representations for image data sharing. Specifically, each entity learns a private Deep Poisoning Module (DPM) and insert it to a pre-trained deep network, which is designed to perform the specific vision task. The DPM deliberately poisons convolutional image features to prevent image reconstructions, while ensuring that the altered image data is functionally equivalent to the non-poisoned data for the specific vision task. Given this equivalence, the poisoned features shared from one entity could be used by another entity for further model refinement. Experimental results on image classification prove the efficacy of the proposed method. Hao Guo 0002, Brian Dolhansky, Eric Hsin, Phong Dinh, Cristian Canton, Song Wang 0002 |
WACV | 5 |
| 2018 | Eye In-Painting With Exemplar Generative Adversarial NetworksabstractThis paper introduces a novel approach to in-painting where the identity of the object to remove or change is preserved and accounted for at inference time: Exemplar GANs (ExGANs). ExGANs are a type of conditional GAN that utilize exemplar information to produce high-quality, personalized in-painting results. We propose using exemplar information in the form of a reference image of the region to in-paint, or a perceptual code describing that object. Unlike previous conditional GAN formulations, this extra information can be inserted at multiple points within the adversarial network, thus increasing its descriptive power. We show that ExGANs can produce photo-realistic personalized in-painting results that are both perceptually and semantically plausible by applying them to the task of closed-to-open eye in-painting in natural pictures. A new benchmark dataset is also introduced for the task of eye inpainting for future comparisons. Brian Dolhansky, Cristian Canton |
CVPR | 2 |
| 2018 | Towards social pattern characterization in egocentric photo-streams
Maedeh Aghaei, Mariella Dimiccoli, Cristian Canton, Petia Radeva |
Comput. Vis. Image Underst. | 3 |
| 2016 | Emotion recognition in the wild from videos using imagesabstractThis paper presents the implementation details of the proposed solution to the Emotion Recognition in the Wild 2016 Challenge, in the category of video-based emotion recognition. The proposed approach takes the video stream from the audio-video trimmed clips provided by the challenge as input and produces the emotion label corresponding to this video sequence. This output is encoded as one out of seven classes: the six basic emotions (Anger, Disgust, Fear, Happiness, Sad, Surprise) and Neutral. Overall, the system consists of several pipelined modules: face detection, image pre-processing, deep feature extraction, feature encoding and, finally, an SVM classification. Sarah Adel Bargal, Emad Barsoum, Cristian Canton, Cha Zhang |
ICMI | 3 |
| 2016 | Training deep networks for facial expression recognition with crowd-sourced label distributionabstractCrowd sourcing has become a widely adopted scheme to collect ground truth labels. However, it is a well-known problem that these labels can be very noisy. In this paper, we demonstrate how to learn a deep convolutional neural network (DCNN) from noisy labels, using facial expression recognition as an example. More specifically, we have 10 taggers to label each input image, and compare four different approaches to utilizing the multiple labels: majority voting, multi-label learning, probabilistic label drawing, and cross-entropy loss. We show that the traditional majority voting scheme does not perform as well as the last two approaches that fully leverage the label distribution. An enhanced FER+ data set with multiple labels for each face image will also be shared with the research community. Emad Barsoum, Cha Zhang, Cristian Canton, Zhengyou Zhang |
ICMI | 3 |
| 2011 | Human motion capture using scalable body models
Cristian Canton, Josep R. Casas, Montse Pardàs |
Comput. Vis. Image Underst. | 1 |
| 2009 | Towards a low cost multi-camera marker based human motion capture systemabstractThis paper presents a low cost real-time alternative to available commercial human motion capture systems. First, a set of distinguishable markers are placed on several human body landmarks and the scene is captured by a number of calibrated and synchronized cameras. In order to establish a physical relation among markers, a human body model (HBM) is defined. Markers are detected on all camera views and delivered as the input of an annealed particle filter scheme where every particle encodes an instance of the pose of the HBM to be estimated. Likelihood between particles and input data is performed through the generalized symmetric epipolar distance and kinematic constrains are enforced in the propagation step towards avoiding impossible poses. Tests over the HumanEva annotated dataset yield quantitative results showing the effectiveness of the proposed algorithm. Results over sequences involving fast and complex motions are also presented. Cristian Canton, Josep R. Casas, Montse Pardàs |
ICIP | 1 |
| 2009 | Improving detection of acoustic events using audiovisual data and feature level fusionabstractThe detection of the acoustic events (AEs) that are naturally \nproduced in a meeting room may help to describe the human \nand social activity that takes place in it. When applied to \nspontaneous recordings, the detection of AEs from only audio \ninformation shows a large amount of errors, which are mostly due to temporal overlapping of sounds. In this paper, a system to detect and recognize AEs using both audio and video information is presented. A feature-level fusion strategy is used, and the structure of the HMM-GMM based system considers each class separately and uses a one-against-all strategy for training. Experime ntal AED results with a new and rather spontaneous dataset are presented which show the advantage of the proposed approach. Taras Butko, Cristian Canton, Carlos Segura, Xavier Giró-i-Nieto, Climent Nadeu, Javier Hernando, Josep R. Casas |
INTERSPEECH | 2 |
| 2008 | Audio-driven human body motion analysis and synthesisabstractThis paper presents a framework for audio-driven human body motion analysis and synthesis. We address the problem in the context of a dance performance, where gestures and movements of the dancer are mainly driven by a musical piece and characterized by the repetition of a set of dance figures. The system is trained in a supervised manner using the multiview video recordings of the dancer. The human body posture is extracted from multiview video information without any human intervention using a novel marker-based algorithm based on annealing particle filtering. Audio is analyzed to extract beat and tempo information. The joint analysis of audio and motion features provides a correlation model that is then used to animate a dancing avatar when driven with any musical piece of the same genre. Results are provided showing the effectiveness of the proposed algorithm. Ferda Ofli, Cristian Canton, Joëlle Tilmanne, Yasemin Demir, Elif Bozkurt, Yücel Yemez, Engin Erzin, A. Murat Tekalp |
ICASSP | 2 |
| 2008 | Particle filtering and sparse sampling for multi-person 3D trackingabstractThis paper presents a new approach to the problem of simultaneous tracking of several people in low resolution sequences from multiple calibrated cameras. Redundancy among cameras is exploited to generate a discrete 3D colored representation of the scene. Two Monte Carlo based schemes adapted to the incoming 3D discrete data are introduced. First, a particle filtering technique is proposed relying on a volume likelihood function taking into account both occupancy and color information. Sparse sampling is presented as an alternative based on a sampling of the surface voxels in order to estimate the centroid of the tracked people. In this case, the likelihood function is based on local neighborhoods computations thus decreasing the computational load of the algorithm. A discrete 3D re-sampling procedure is introduced to drive these samples along time. Multiple targets are tracked by means of multiple filters and interaction among them is modeled through a 3D blocking scheme. Tests over annotated databases yield quantitative results showing the effectiveness of the proposed algorithms in indoor scenarios. Cristian Canton, Josep R. Casas, Montse Pardàs, Rossella Sblendido |
ICIP | 1 |
| 2008 | Fusion of audio and video modalities for detection of acoustic eventsabstractDetection of acoustic events (AED) that take place in a meeting-room environment becomes a difficult task when signals show a large proportion of temporal overlap of sounds, like in seminar-type data, where the acoustic events often occur simultaneously with speech. Whenever the event that produces the sound is related to a given position or movement, video signals may be a useful additional source of information for AED. In this work, we aim at improving the AED accuracy by using two complementary audio-based AED systems, built with SVM and HMM classifiers, and also a video-based AED system, which employs the output of a 3D video tracking algorithm to improve detection of steps. Fuzzy integral is used to fuse the outputs of the three classification systems in two stages. Experimental results using the CLEAR'07 evaluation data show that the detection rate increases by fusing the two audio information sources, and it is further improved by including video information. Taras Butko, Andriy Temko, Climent Nadeu, Cristian Canton |
INTERSPEECH | 4 |
| 2007 | Multi-Person 3D Tracking with Particle Filters on VoxelsabstractThis paper presents a new approach to the problem of simultaneous tracking of several people in low resolution sequences from multiple calibrated cameras. Spatial redundancy is exploited to generate a discrete 3D binary representation of the scene. A particle filtering scheme adapted to the incoming 3D discrete data is proposed. A volume likelihood function and a discrete 3D re-sampling procedure are introduced to evaluate and drive particles. Multiple targets are tracked by means of multiple particle filters and interaction among them is modeled through a 3D blocking scheme. Test over annotated databases yield quantitative results showing the effectiveness of the proposed algorithm in the indoor scenarios. Adolfo López, Cristian Canton, Josep R. Casas |
ICASSP (1) | 2 |
| 2007 | Multimodal Head Orientation Towards Attention Tracking in SmartroomsabstractThis paper presents a multimodal approach to head pose estimation and 3D gaze orientation of individuals in a SmartRoom environment equipped with multiple cameras and microphones. We first introduce the two monomodal approaches as reference. In video, we estimate head orientation from color information by exploiting spatial redundancy among cameras. Audio information is processed to estimate the direction of the voice produced by a speaker making use of the directivity characteristics of the head radiation pattern. Two multimodal information fusion schemes working at data and decision levels are analyzed in terms of accuracy and robustness of the estimation. Experimental results conducted over the CLEAR evaluation database are reported and the comparison of the proposed multimodal head pose estimation algorithms with the reference monomodal approaches proves the effectiveness of the proposed approach. Carlos Segura, Cristian Canton, Alberto Abad, Josep R. Casas, Javier Hernando |
ICASSP (2) | 2 |
| 2005 | Fusion of multiple viewpoint information towards 3D face robust orientation detectionabstractThis paper presents a novel approach to the problem of determining head pose estimation and face 3D orientation of several people in low resolution sequences from multiple calibrated cameras. Spatial redundancy is exploited and the heads of people in the scene are detected and geometrically approximated by an ellipsoid using a voxel reconstruction and a moment analysis method. Skin patches from each detected head are located in each camera view. Data fusion is performed by back-projecting skin patches from single images onto the estimated 3D head model, thus providing a synthetic reconstruction of the head appearance. Finally, these data are processed in a pattern analysis framework thus giving a reliable and robust estimation of face orientation. Tracking over time is performed by Kalman filtering. Results are provided showing the effectiveness of the proposed algorithm in a Smart-Room scenario. Cristian Canton, Josep R. Casas, Montse Pardàs |
ICIP (2) | 1 |