EDBT 2026 Demo / reviewers in the wild / expert
Celso de Melo
dblp:84/4437 · also Celso M. de Melo, Celso Miguel de Melo
· DBLP profile ↗
67ranked-venue papers
26as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 15 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 4 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 20 · 14 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PromptGAR: Flexible Promptive Group Activity RecognitionabstractWe present PromptGAR, a novel framework for Group Activity Recognition (GAR) that offering both input flexibility and high recognition accuracy. The existing approaches suffer from limited real-world applicability due to their reliance on full prompt annotations, fixed number of frames and instances, and the lack of actor consistency. To bridge the gap, we proposed PromptGAR, which is the first GAR model to provide input flexibility across prompts, frames, and instances without the need for retraining. We leverage diverse visual prompts—like bounding boxes, skeletal keypoints, and instance identities—by unifying them as point prompts. A recognition decoder then cross-updates class and prompt tokens for enhanced performance. To ensure actor consistency for extended activity durations, we also introduce a relative instance attention mechanism that directly encodes instance identities. Comprehensive evaluations demonstrate that PromptGAR achieves competitive performances both on full prompts and partial prompt inputs, establishing its effectiveness on input flexibility and generalization ability for real-world applications. See the project page for more results: https://jinzhangyu.github.io/projects/PromptGAR/ Zhangyu Jin, Andrew Feng, Ankur Chemburkar, Celso de Melo |
WACV | 4 |
| 2026 | Referring Change Detection in Remote Sensing ImageryabstractChange detection in remote sensing imagery is essential for applications such as urban planning, environmental monitoring, and disaster management. Traditional change detection methods typically identify all changes between two temporal images without distinguishing the types of transitions, which can lead to results that may not align with specific user needs. Although semantic change detection methods have attempted to address this by categorizing changes into predefined classes, these methods rely on rigid class definitions and fixed model architectures, making it difficult to mix datasets with different label sets or reuse models across tasks, as the output channels are tightly coupled with the number and type of semantic classes. To overcome these limitations, we introduce Referring Change Detection (RCD), which leverages natural language prompts to detect specific classes of changes in remote sensing images. By integrating language understanding with visual analysis, our approach allows users to specify the exact type of change they are interested in. However, training models for RCD is challenging due to the limited availability of annotated data and severe class imbalance in existing datasets. To address this, we propose a two-stage framework consisting of (I) RCDNet, a cross-modal fusion network designed for referring change detection, and (II) RCDGen, a diffusion-based synthetic data generation pipeline that produces realistic post-change images and change maps for a specified category using only pre-change image, without relying on semantic segmentation masks and thereby significantly lowering the barrier to scalable data creation. Experiments across multiple datasets show that our framework enables scalable and targeted change detection. Code will be made publicly available on Github. Yilmaz Korkmaz, Jay N. Paranjape, Celso de Melo, Vishal M. Patel |
WACV | 3 |
| 2026 | F-ViTA: Foundation Model Guided Visible-to-Infrared TranslationabstractThermal imaging is crucial for scene understanding, particularly in low-light and nighttime conditions. However, collecting large thermal datasets is costly and labor-intensive due to the specialized equipment required for infrared image capture. To address this challenge, researchers have explored visible-to-thermal image translation. Most existing methods rely on Generative Adversarial Networks (GANs) or Diffusion Models (DMs), treating the task as a style transfer problem. As a result, these approaches attempt to learn both the modality distribution shift and underlying physical principles from limited training data. In this paper, we propose F-ViTA, a novel style transfer approach that leverages the general world knowledge embedded in foundation models to guide the diffusion process for improved translation. Specifically, we condition an In-structPix2Pix Diffusion Model with zero-shot masks and labels from foundation models such as SAM and Grounded DINO. This allows the model to learn meaningful correlations between scene objects and their thermal signatures in infrared imagery. Extensive experiments on five public datasets demonstrate that F-ViTA outperforms state-of-the-art (SOTA) methods. Furthermore, our model generalizes well to out-of-distribution (OOD) scenarios and can generate Long-Wave Infrared (LWIR), Mid-Wave Infrared (MWIR), and Near-Infrared (NIR) translations from the same visible image. Code: Released post-review. Jay N. Paranjape, Celso de Melo, Vishal M. Patel |
WACV | 2 |
| 2025 | Guarding Barlow Twins Against Overfitting with Mixed SamplesabstractSelf-supervised Learning (SSL) aims to learn transferable feature representations for downstream applications without relying on labeled data. The Barlow Twins algorithm, renowned for its widespread adoption and straightforward implementation compared to its counterparts like contrastive learning methods, minimizes feature redundancy while maximizing invariance to common corruptions. Optimizing for the above objective forces the network to learn useful representations, while avoiding noisy or constant features, resulting in improved downstream task performance with limited adaptation. Despite Barlow Twins’ proven effectiveness in pre-training, the underlying SSL objective can inadvertently cause feature overfitting due to the lack of strong interaction between the samples unlike the contrastive learning approaches. From our experiments, we observe that optimizing for the Barlow Twins objective doesn’t necessarily guarantee sustained improvements in representation quality beyond a certain pre-training phase, and can potentially degrade downstream performance on some datasets. To address this challenge, we introduce Mixed Barlow Twins, which aims to improve sample interaction during Barlow Twins training via linearly interpolated samples. This results in an additional regularization term to the original Barlow Twins objective, assuming linear interpolation in the input space translates to linearly interpolated features in the feature space. Pre-training with this regularization effectively mitigates feature over-fitting and further enhances the downstream performance on CIFAR-10, CIFAR-100, TinyImageNet, STL-10, and ImageNet datasets. Wele Gedara Chaminda Bandara, Celso de Melo, Vishal M. Patel |
AVSS | 2 |
| 2025 | Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language ModelsabstractAutomatic target recognition (ATR) is crucial for safety-critical tasks such as navigation and surveillance, particularly in demanding military scenarios with unfamiliar terrains, harsh environments, and unseen object categories. Conventional and open-world object detectors often fail in these contexts, lacking exposure to such novel conditions. Meanwhile, Large Vision-Language Models (LVLMs) exhibit zero-shot recognition capabilities across diverse settings yet struggle with precise localization. We address these limitations by combining the localization strength of open-world detectors with the recognition confidence of LVLMs, creating a robust pipeline for zero-shot ATR in novel domains and classes. Our study compares several LVLMs on underrepresented military vehicles, examining factors such as distance range, modality, and prompting strategies. These findings provide insights for developing more reliable ATR systems in uncharted environments. Yasiru Ranasinghe, Vibashan VS, James Uplinger, Celso de Melo, Vishal M. Patel |
AVSS | 4 |
| 2025 | A Bayesian Model of Mind Reading from Decisions and Emotions in Social Dilemmas
Kazunori Terada, Celso de Melo, Francisco C. Santos, Jonathan Gratch |
CogSci | 2 |
| 2025 | SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal ModelsabstractHumans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (LMMs), however, lack of this capability of 3D spatial reasoning. This limitation stems from the scarcity of 3D training data and the bias in current model designs toward 2D data. In this paper, we systematically study the impact of 3D-informed data, architecture, and training setups, introducing SpatialLLM, a large multi-modal model with advanced 3D spatial reasoning abilities. To address data limitations, we develop two types of 3D-informed training datasets: (1) 3D-informed probing data focused on object’s 3D location and orientation, and (2) 3D-informed conversation data for complex spatial relationships. Notably, we are the first to curate VQA data that incorporate 3D orientation relationships on real images. Furthermore, we systematically integrate these two types of training data with the architectural and training designs of LMMs, providing a roadmap for optimal design aimed at achieving superior 3D reasoning capabilities. Our SpatialLLM advances machines toward highly capable 3D-informed reasoning, surpassing GPT-4o performance by 8.7%. Our systematic empirical design and the resulting findings offer valuable insights for future research in this direction. Our project page is available at this link. Wufei Ma, Luoxin Ye, Celso de Melo, Alan L. Yuille, Jieneng Chen |
CVPR | 3 |
| 2025 | Video-ColBERT: Contextualized Late Interaction for Text-to-Video RetrievalabstractIn this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval, our approach, Video-ColBERT, introduces a simple and efficient mechanism for fine-grained similarity assessment between queries and videos. Video-ColBERT is built upon three main components: a fine-grained spatial and temporal token-wise interaction, query and visual expansions, and a dual sigmoid loss during training. We find that this interaction and training paradigm leads to strong individual, yet compatible, representations for encoding video content. These representations lead to increases in performance on common text-to-video retrieval benchmarks compared to other bi-encoder methods. Arun V. Reddy, Alexander Martin 0006, Eugene Yang 0001, Andrew Yates, Kate Sanders 0002, Kenton Murray, Reno Kriz, Celso de Melo, Benjamin Van Durme, Rama Chellappa |
CVPR | 8 |
| 2025 | Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Mutimodal ModelsabstractAlthough large multimodal models (LMMs) have demonstrated remarkable capabilities in visual scene interpretation and reasoning, their capacity for complex and precise 3-dimensional spatial reasoning remains uncertain. Existing benchmarks focus predominantly on 2D spatial understanding and lack a framework to comprehensively evaluate 6D spatial reasoning across varying complexities. To address this limitation, we present Spatial457, a scalable and unbiased synthetic dataset designed with 4 key capability for spatial reasoning: multi-object recognition, 2D location, 3D location, and 3D orientation. We develop a cascading evaluation structure, constructing 7 question types across 5 difficulty levels that range from basic single object recognition to our new proposed complex 6D spatial reasoning tasks. We evaluated various large multimodal models (LMMs) on Spatial457, observing a general decline in performance as task complexity increases, particularly in 3D reasoning and 6D spatial tasks. To quantify these challenges, we introduce the Relative Performance Dropping Rate (RPDR), highlighting key weaknesses in 3D reasoning capabilities. Leveraging the unbiased attribute design of our dataset, we also uncover prediction biases across different attributes, with similar patterns observed in real-world image settings.1The code is released in https://github.com/XingruiWang/Spatial457. Xingrui Wang, Wufei Ma, Tiezheng Zhang, Celso de Melo, Jieneng Chen, Alan L. Yuille |
CVPR | 4 |
| 2025 | AeroGen: Ground-to-Air Generalization for Action RecognitionabstractWe address the problem of action recognition from aerial views using only ground-based videos for training. Due to the viewpoint-induced domain shift, models trained solely on ground videos exhibit significant performance degradation when naively applied to aerial videos. To mitigate this performance gap, we introduce a domain generalization technique that addresses the viewpoint-induced domain shift. Our method uses available real ground videos to generate additional synthetic training data from both ground and air viewpoints, for improving generalization to aerial video-based action recognition. Specifically, we perform 3D human mesh estimation from the ground videos, then render synthetic videos from alternate viewpoints with additional appearance randomizations. To further align the ground-air and syntheticreal domains, we propose a Dual Domain Alignment loss by enforcing consistency in predictions between the original ground videos and augmented videos from each domain. In order to facilitate research on the problem of ground-to-air generalization for human action recognition, we also create a benchmark by combining parts of the NTU-60, UAV-Human and NEC-DRONE datasets. We demonstrate the effectiveness of our approach on these new benchmarks along with the existing RoCoG-Ground $\rightarrow$ RoCoG-Air benchmark, and also perform extensive ablations. Ketul Shah, Anshul Shah 0001, Arun V. Reddy, Aniket Roy, Celso de Melo, Rama Chellappa |
FG | 5 |
| 2025 | Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak SupervisionabstractDetecting vehicles in aerial imagery is a critical task with applications in traffic monitoring, urban planning, and defense intelligence. Deep learning methods have provided state-of-the-art (SOTA) results for this application. However, a significant challenge arises when models trained on data from one geographic region fail to generalize effectively to other areas. Variability in factors such as environmental conditions, urban layouts, road networks, vehicle types, and image acquisition parameters (e.g., resolution, lighting, and angle) leads to domain shifts that degrade model performance. This paper proposes a novel method that uses generative AI to synthesize high-quality aerial images and their labels, improving detector training through data augmentation. Our key contribution is the development of a multi-stage, multi-modal knowledge transfer framework utilizing fine-tuned latent diffusion models (LDMs) to mitigate the distribution gap between the source and target environments. Extensive experiments across diverse aerial imagery domains show consistent performance improvements in AP50 over supervised learning on source domain data, weakly supervised adaptation methods, unsupervised domain adaptation methods, and open-set object detectors by 4-23%, 6-10%, 7-40%, and more than 50%, respectively. Furthermore, we introduce two newly annotated aerial datasets from New Zealand and Utah to support further research in this field. Project page is available at: https://humansensinglab.github.io/AGenDA Minhyek Jeon, Shuowen Hu, Zheyang Qin, Shayok Chakraborty, Stanislav Panev, Celso de Melo, Fernando De la Torre |
ICCV | 7 |
| 2025 | 3DSRBENCH: A Comprehensive 3D Spatial Reasoning Benchmarkabstract3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their applicability to a broader range of areas, such as autonomous navigation, robotics, and AR/VR. While large multi-modal models (LMMs) have achieved remarkable progress in a wide range of image and video understanding tasks, their capabilities to perform 3D spatial reasoning on diverse natural images are less studied. In this work we present the first comprehensive 3D spatial reasoning benchmark, 3DSRBench, with 2,772 manually annotated visual question-answer pairs across 12 question types. We conduct robust and thorough evaluation of 3D spatial reasoning abilities by balancing data distribution and adopting a novel FlipEval strategy. To further study the robustness of 3D spatial reasoning w.r.t. camera 3D viewpoints, our 3DSRBench includes two subsets with 3D spatial reasoning questions on paired images with common and uncommon viewpoints. We benchmark a wide range of open-sourced and proprietary LMMs, uncovering their limitations in various aspects of 3D awareness, such as height, orientation, location, and multi-object reasoning, as well as their degraded performance on images from uncommon 6D viewpoints. Our 3DSRBench provide valuable findings and insights about future development of LMMs with strong spatial reasoning abilities. Our project page is available at https://3dsrbench.github.io/. Wufei Ma, Guofeng Zhang 0025, Yu-Cheng Chou, Jieneng Chen, Celso de Melo, Alan L. Yuille |
ICCV | 6 |
| 2025 | Shuffle PatchMix Augmentation with Confidence-Margin Weighted Pseudo-Labels for Enhanced Source-Free Domain AdaptationabstractThis work investigates Source-Free Domain Adaptation (SFDA), where a model adapts to a target domain without access to source data. A new augmentation technique, Shuffle PatchMix (SPM), and a novel reweighting strategy are introduced to enhance performance. SPM shuffles and blends image patches to generate diverse and challenging augmentations, while the reweighting strategy prioritizes reliable pseudo-labels to mitigate label noise. These techniques are particularly effective on smaller datasets like PACS, where overfitting and pseudo-label noise pose greater risks. State-of-the-art results are achieved on three major benchmarks: PACS, VisDA-C, and DomainNet-126. Notably, on PACS, improvements of 7.3% (79.4% to 86.7%) and 7.2% are observed in single-target and multi-target settings, respectively, while gains of 2.8% and 0.7% are attained on DomainNet-126 and VisDA-C. This combination of advanced augmentation and robust pseudo-label reweighting establishes a new benchmark for SFDA. The code is available at: https://github.com/PrasannaPulakurthi/SPM. Prasanna Reddy Pulakurthi, Majid Rabbani, Jamison Heard, Sohail A. Dianat, Celso de Melo, Raghuveer M. Rao |
ICIP | 5 |
| 2025 | Texture- and Shape-Based Adversarial Attacks for Overhead Image Vehicle DetectionabstractDetecting vehicles in aerial images is difficult due to complex backgrounds, small object sizes, shadows, and occlusions. Although recent deep learning advancements have improved object detection, these models remain susceptible to adversarial attacks (AAs), challenging their reliability. Traditional AA strategies often ignore practical implementation constraints. Our work proposes realistic and practical constraints on texture (lowering resolution, limiting modified areas, and color ranges) and analyzes the impact of shape modifications on attack performance. We conducted extensive experiments with three object detector architectures, demonstrating the performance-practicality trade-off: more practical modifications tend to be less effective, and vice versa. We release both code and data to support reproducibility at https://github.com/humansensinglab/texture-shape-adversarial-attacks. Mikael Yeghiazaryan, Sai Abhishek Si Namburu, Emily Kim, Stanislav Panev, Celso de Melo, Fernando De la Torre, Jessica K. Hodgins |
ICIP | 5 |
| 2025 | ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and ExecutionabstractRobotic planning and execution in open-world environments is a complex problem due to the vast state spaces and high variability of task embodiment. Recent advances in perception algorithms, combined with Large Language Models (LLMs) for planning, offer promising solutions to these challenges, as the common sense reasoning capabilities of LLMs provide a strong heuristic for efficiently searching the action space. However, prior work fails to address the possibility of hallucinations from LLMs, which results in failures to execute the planned actions largely due to logical fallacies at high-or low-levels. To contend with automation failure due to such hallucinations, we introduce ConceptAgent, a natural language-driven robotic platform designed for task execution in unstructured environments. With a focus on scalability and reliability of LLM-based planning in complex state and action spaces, we present innovations designed to limit these shortcomings, including 1) Predicate Grounding to prevent and recover from infeasible actions, and 2) an embodied version of LLM-guided Monte Carlo Tree Search with self reflection. ConceptAgent combines these planning enhancements with dynamic language aligned 3d scene graphs, and large multi-modal pretrained models to perceive, localize, and interact with its environment, enabling reliable task completion. In simulation experiments, ConceptAgent achieved a 19% task completion rate across three room layouts and 30 easy level embodied tasks outperforming other state-of-the-art LLM-driven reasoning baselines that scored 10.26% and 8.11% on the same benchmark. Additionally, ablation studies on moderate to hard embodied tasks revealed a 20% increase in task completion from the baseline agent to the fully enhanced ConceptAgent, highlighting the individual and combined contributions of Predicate Grounding and LLM-guided Tree Search to enable more robust automation in complex state and action spaces. Additionally, in real-world mobile manipulation trials, conducted in randomized, low-clutter environments, a ConceptAgent-driven Spot robot achieved a 40% task completion rate, demonstrating the performance of our perception system in real-world scenarios. Corban Rivera, Grayson Byrd, William Paul, Tyler Feldman, Meghan Booker, Emma Holmes, David Handelman, Bethany Kemp, Andrew Badger, Aurora Schmidt, Krishna Murthy Jatavallabhula, Celso de Melo, Seenivasan Lalithkumar, Mathias Unberath, Rama Chellappa |
ICRA | 12 |
| 2025 | Aha! - Predicting What Matters Next: Online Highlight Detection Without Looking AheadabstractReal-time understanding of continuous video streams is essential for intelligent agents operating in high-stakes environments, including autonomous vehicles, surveillance drones, and disaster response robots. Yet, most existing video understanding and highlight detection methods assume access to the entire video during inference, making them unsuitable for online or streaming scenarios. In particular, current models optimize for offline summarization, failing to support step-by-step reasoning needed for real-time decision-making.
We introduce Aha, an autoregressive highlight detection framework that predicts the relevance of each video frame against a task described in natural language. Without accessing future video frames, Aha utilizes a multimodal vision-language model and lightweight, decoupled heads trained on a large, curated dataset of human-centric video labels. To enable scalability, we introduce the Dynamic SinkCache mechanism that achieves constant memory usage across infinite-length streams without degrading performance on standard benchmarks. This encourages the hidden representation to capture high-level task objectives, enabling effective frame-level rankings for informativeness, relevance, and uncertainty with respect to the natural language task. Aha achieves state-of-the-art (SOTA) performance on highlight detection benchmarks, surpassing even prior offline, full-context approaches and video-language models by +5.9\% on TVSum and +8.3\% on Mr.Hisum in mAP (mean Average Precision). We explore Aha’s potential for real-world robotics applications given a task-oriented natural language input and a continuous, robot-centric video. Both experiments demonstrate Aha's potential effectiveness as a real-time reasoning module for downstream planning and long-horizon understanding. Aiden Chang, Celso de Melo, Stephanie M. Lukin |
NeurIPS | 2 |
| 2025 | SpatialReasoner: Towards Explicit and Generalizable 3D Spatial ReasoningabstractDespite recent advances on multi-modal models, 3D spatial reasoning remains a challenging task for state-of-the-art open-source and proprietary models. Recent studies explore data-driven approaches and achieve enhanced spatial reasoning performance by fine-tuning models on 3D-related visual question-answering data. However, these methods typically perform spatial reasoning in an implicit manner and often fail on questions that are trivial to humans, even with long chain-of-thought reasoning. In this work, we introduce SpatialReasoner, a novel large vision-language model (LVLM) that addresses 3D spatial reasoning with explicit 3D representations shared between multiple stages--3D perception, computation, and reasoning. Explicit 3D representations provide a coherent interface that supports advanced 3D spatial reasoning and improves the generalization ability to novel question types. Furthermore, by analyzing the explicit 3D representations in multi-step reasoning traces of SpatialReasoner, we study the factual errors and identify key shortcomings of current LVLMs. Results show that our SpatialReasoner achieves improved performance on a variety of spatial reasoning benchmarks, outperforming Gemini 2.0 by 9.2% on 3DSRBench, and generalizes better when evaluating on novel 3D spatial reasoning questions. Our study bridges the 3D parsing capabilities of prior visual foundation models with the powerful reasoning abilities of large language models, opening new directions for 3D spatial reasoning. Wufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang, Celso de Melo, Jianwen Xie, Alan L. Yuille |
NeurIPS | 5 |
| 2025 | Bisecle: Binding and Separation in Continual Learning for Video Language UnderstandingabstractFrontier vision-language models (VLMs) have made remarkable improvements in video understanding tasks. However, real-world videos typically exist as continuously evolving data streams (e.g., dynamic scenes captured by wearable glasses), necessitating models to continually adapt to shifting data distributions and novel scenarios. Considering the prohibitive computational costs of fine-tuning models on new tasks, usually, a small subset of parameters is updated while the bulk of the model remains frozen. This poses new challenges to existing continual learning frameworks in the context of large multimodal foundation models, i.e., catastrophic forgetting and update conflict. While the foundation models struggle with parameter-efficient continual learning, the hippocampus in the human brain has evolved highly efficient mechanisms for memory formation and consolidation. Inspired by the rapid **Bi**nding and pattern **se**paration mechanisms in the hippocampus, in this work, we propose **Bisecle** for video-language **c**ontinual **le**arning, where a multi-directional supervision module is used to capture more cross-modal relationships and a contrastive prompt learning scheme is designed to isolate task-specific knowledge to facilitate efficient memory storage. Binding and separation processes further strengthen the ability of VLMs to retain complex experiences, enabling robust and efficient continual learning in video understanding tasks. We perform a thorough evaluation of the proposed Bisecle, demonstrating its ability to mitigate forgetting and enhance cross-task generalization on several VideoQA benchmarks. Xiaoqian Hu, Hao Xue 0001, Celso de Melo, Flora D. Salim |
NeurIPS | 4 |
| 2025 | A Mamba-Based Siamese Network for Remote Sensing Change DetectionabstractChange detection in remote sensing images is an essential tool for analyzing a region at different times. It finds varied applications in monitoring environmental changes, man-made changes as well as corresponding decisionmaking and prediction of future trends. Deep learning methods like Convolutional Neural Networks (CNNs) and Transformers have achieved remarkable success in detecting significant changes, given two images at different times. In this paper, we propose a Mamba-based Change Detector (M-CD) that segments out the regions of interest even better. Mamba-based architectures demonstrate lineartime training capabilities and an improved receptive field over transformers. Our experiments on four widely used change detection datasets demonstrate significant improvements over existing state-of-the-art (SOTA) methods. Code: https://github.com/JayParanjape/M-CD Jay N. Paranjape, Celso de Melo, Vishal M. Patel |
WACV | 2 |
| 2024 | Entropic Open-Set Active LearningabstractActive Learning (AL) aims to enhance the performance of deep models by selecting the most informative samples for annotation from a pool of unlabeled data. Despite impressive performance in closed-set settings, most AL methods fail in real-world scenarios where the unlabeled data contains unknown categories. Recently, a few studies have attempted to tackle the AL problem for the open-set setting. However, these methods focus more on selecting known samples and do not efficiently utilize unknown samples obtained during AL rounds. In this work, we propose an Entropic Open-set AL (EOAL) framework which leverages both known and unknown distributions effectively to select informative samples during AL rounds. Specifically, our approach employs two different entropy scores. One measures the uncertainty of a sample with respect to the known-class distributions. The other measures the uncertainty of the sample with respect to the unknown-class distributions. By utilizing these two entropy scores we effectively separate the known and unknown samples from the unlabeled data resulting in better sampling. Through extensive experiments, we show that the proposed method outperforms existing state-of-the-art methods on CIFAR-10, CIFAR-100, and TinyImageNet datasets. Code is available at https://github.com/bardisafa/EOAL. Bardia Safaei 0002, Vibashan VS, Celso de Melo, Vishal M. Patel |
AAAI | 3 |
| 2024 | Emotional Expression Help Regulate the Appropriate Level of Cooperation with AgentsabstractPeople often anthropomorphize agents and show social concern for the agents' goals. Whereas this can be useful to build human-agent cooperation in some settings, in others it can be counterproductive - e.g., when people risk themselves to help a robot. A mechanism, thus, is needed to regulate how much cooperation people show towards agents, according to the context. Here, we show that emotion expressions can be a powerful mechanism to help people identify the appropriate level of cooperation given the situation. In the present study, participants (n=379) engaged in a 20-round iterated prisoner's dilemma game with agents that showed emotional expressions that reflected a preference for maximizing its own interests versus maximizing the participants' interests. Accordingly, the results showed that participants focused significantly more on their own interests when facing the agent that expressed emotions favoring the participants' outcome; moreover, this treatment was more successful in steering the participants' focus to their own interests than showing no emotion. These findings reveal that, in addition to helping build cooperation, as shown in prior work, emotion expression can play a central role in mitigating some negative consequences of anthropomorphizing agents. Ryoya Ito, Celso de Melo, Jonathan Gratch, Kazunori Terada |
ACII | 2 |
| 2024 | People Negotiate Better with Emotional Human-Like Virtual Agents Than Android RobotsabstractEmotional expressions serve as important communicative tools in human negotiations, and prior work has shown that artificial agents can use synthetic expressions to enhance negotiation outcomes and to train negotiation skills. These prior findings have focused on virtual agents and little is known about the effect of expressions when negotiating with physical robots. Therefore, in this study, we compared how participants negotiated with emotionally expressive virtual agents and android robots. Participants$(\mathrm{n}={82})$, as a proposer, played a nonverbal version of a four-issue ultimatum bargaining game with a counterpart who was either a virtual agent or an android robot. Before negotiating, participants observed their counterpart's emotional reactions to potential deals. The results showed that participants were better able to estimate the preferences of virtual counterparts compared with robotic counterpart, and thereby achieve better win-win solutions. We find this effect was mediated by uncanniness: participants found the emotional robot to be uncanny, and this undermined their ability to extract information from the robot's expressions. We discuss theoretical mplications for our understanding of human-robot negotiation and practical implications for the design of effective robot negotiators. Motoaki Sato, Takahisa Uchida, Yuichiro Yoshikawa, Celso de Melo, Jonathan Gratch, Kazunori Terada |
ACII | 4 |
| 2024 | Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-TrainingabstractIn this work, we tackle the problem of unsupervised domain adaptation (UDA) for video action recognition. Our approach, which we call UNITE, uses an image teacher model to adapt a video student model to the target domain. UNITE first employs self-supervised pretraining to promote discriminative feature learning on target domain videos using a teacher-guided masked distillation objective. We then perform self-training on masked target data, using the video student model and image teacher model together to generate improved pseudolabels for unlabeled target videos. Our self-training process successfully leverages the strengths of both models to achieve strong transfer performance across domains. We evaluate our approach on multiple video domain adaptation benchmarks and observe significant improvements upon previously reported results. Arun V. Reddy, William Paul, Corban Rivera, Ketul Shah, Celso de Melo, Rama Chellappa |
CVPR | 5 |
| 2024 | ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and PlanningabstractFor robots to perform a wide variety of tasks, they require a 3D representation of the world that is semantically rich, yet compact and efficient for task-driven perception and planning. Recent approaches have attempted to leverage features from large vision-language models to encode semantics in 3D representations. However, these approaches tend to produce maps with per-point feature vectors, which do not scale well in larger environments, nor do they contain semantic spatial relationships between entities in the environment, which are useful for downstream planning. In this work, we propose ConceptGraphs, an open-vocabulary graph-structured representation for 3D scenes. ConceptGraphs is built by leveraging 2D foundation models and fusing their output to 3D by multi-view association. The resulting representations generalize to novel semantic classes, without the need to collect large 3D datasets or finetune models. We demonstrate the utility of this representation through a number of downstream planning tasks that are specified through abstract (language) prompts and require complex reasoning over spatial and semantic concepts. To explore the full scope of our experiments and results, we encourage readers to visit our project webpage. Qiao Gu, Alihusein Kuwajerwala, Sacha Morin, Krishna Murthy Jatavallabhula, Bipasha Sen, Aditya Agarwal, Corban Rivera, William Paul, Kirsty Ellis, Rama Chellappa, Chuang Gan 0001, Celso de Melo, Josh Tenenbaum, Antonio Torralba 0001, Florian Shkurti, Liam Paull |
ICRA | 12 |
| 2024 | ViLCo-Bench: VIdeo Language COntinual learning BenchmarkabstractVideo language continual learning involves continuously adapting to information from video and text inputs, enhancing a model’s ability to handle new tasks while retaining prior knowledge. This field is a relatively under-explored area, and establishing appropriate datasets is crucial for facilitating communication and research in this field. In this study, we present the first dedicated benchmark, ViLCo-Bench, designed to evaluate continual learning models across a range of video-text tasks. The dataset comprises ten-minute-long videos and corresponding language queries collected from publicly available datasets. Additionally, we introduce a novel memory-efficient framework that incorporates self-supervised learning and mimics long-term and short-term memory effects. This framework addresses challenges including memory complexity from long video clips, natural language complexity from open queries, and text-video misalignment. We posit that ViLCo-Bench, with greater complexity compared to existing continual learning benchmarks, would serve as a critical tool for exploring the video-language domain, extending beyond conventional class-incremental tasks, and addressing complex and limited annotation issues. The curated data, evaluations, and our novel method are available at https://github.com/cruiseresearchgroup/ViLCo. Tianqi Tang 0002, Shohreh Deldari, Hao Xue 0001, Celso de Melo, Flora D. Salim |
NeurIPS | 4 |
| 2024 | Exploring the Impact of Rendering Method and Motion Quality on Model Performance when Using Multi-view Synthetic Data for Action RecognitionabstractThis paper explores the use of synthetic data in a human action recognition (HAR) task to avoid the challenges of obtaining and labeling real-world datasets. We introduce a new dataset suite comprising five datasets, eleven common human activities, three synchronized camera views (aerial and ground) in three outdoor environments, and three visual domains (real and two synthetic). For the synthetic data, two rendering methods (standard computer graphics and neural rendering) and two sources of human motions (motion capture and video-based motion reconstruction) were employed. We evaluated each dataset type by training popular activity recognition models and comparing the performance on the real test data. Our results show that synthetic data achieve slightly lower accuracy (4–8 %) than real data. On the other hand, a model pre-trained on synthetic data and fine-tuned on limited real data surpasses the performance of either domain alone. Standard computer graphics (CG)-rendered data delivers better performance than the data generated from the neural-based rendering method. The results suggest that the quality of the human motions in the training data also affects the test results: motion capture delivers higher test accuracy. Additionally, a model trained on CG aerial view synthetic data exhibits greater robustness against camera viewpoint changes than one trained on real data. See the project page: http://humansensinglab.github.io/REMAG/ Stanislav Panev, Emily Kim, Sai Abhishek Si Namburu, Desislava Nikolova, Celso de Melo, Fernando De la Torre, Jessica K. Hodgins |
WACV | 5 |
| 2023 | STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action RecognitionabstractWe study the problem of human action recognition using motion capture (MoCap) sequences. Unlike existing techniques that take multiple manual steps to derive standard-ized skeleton representations as model input, we propose a novel Spatial-Temporal Mesh Transformer (STMT) to directly model the mesh sequences. The model uses a hierarchical transformer with intra-frame off-set attention and inter-frame self-attention. The attention mechanism allows the model to freely attend between any two vertex patches to learn nonlocal relationships in the spatial-temporal domain. Masked vertex modeling and future frame prediction are used as two self-supervised tasks to fully activate the bi-directional and auto-regressive attention in our hierarchical transformer. The proposed method achieves state-of-the-art performance compared to skeleton-based and point-cloud-based models on common MoCap benchmarks. Code is available at https://github.com/zgzxy001/STMT. Po-Yao Huang 0001, Junwei Liang 0001, Celso de Melo, Alex Hauptmann 0001 |
CVPR | 4 |
| 2023 | Open-Set Automatic Target RecognitionabstractAutomatic Target Recognition (ATR) is a category of computer vision algorithms which attempts to recognize targets on data obtained from different sensors. ATR algorithms are extensively used in real-world scenarios such as military and surveillance applications. Existing ATR algorithms are developed for traditional closed-set methods where training and testing have the same class distribution. Thus, these algorithms have not been robust to unknown classes not seen during the training phase, limiting their utility in real-world applications. To this end, we propose an Open-set Automatic Target Recognition framework where we enable open-set recognition capability for ATR algorithms. In addition, we introduce a plugin Category-aware Binary Classifier (CBC) module to effectively tackle unknown classes seen during inference. The proposed CBC module can be easily integrated with any existing ATR algorithms and can be trained in an end-to-end manner. Experimental results show that the proposed approach outperforms many open-set methods on the DSIAC and CIFAR-10 datasets. To the best of our knowledge, this is the first work to address the open-set classification problem for ATR algorithms. Source code is available at: https://github.com/bardisafa/Open-set-ATR. Bardia Safaei 0002, Vibashan VS, Celso de Melo, Shuowen Hu, Vishal M. Patel |
ICASSP | 3 |
| 2023 | Synthetic-to-Real Domain Adaptation for Action Recognition: A Dataset and Baseline PerformancesabstractHuman action recognition is a challenging problem, particularly when there is high variability in factors such as subject appearance, backgrounds and viewpoint. While deep neural networks (DNNs) have been shown to perform well on action recognition tasks, they typically require large amounts of high-quality labeled data to achieve robust performance across a variety of conditions. Synthetic data has shown promise as a way to avoid the substantial costs and potential ethical concerns associated with collecting and labeling enormous amounts of data in the real-world. However, synthetic data may differ from real data in important ways. This phenomenon, known as domain shift, can limit the utility of synthetic data in robotics applications. To mitigate the effects of domain shift, substantial effort is being dedicated to the development of domain adaptation (DA) techniques. Yet, much remains to be understood about how best to develop these techniques. In this paper, we introduce a new dataset called Robot Control Gestures (RoCoG-v2). The dataset is composed of both real and synthetic videos from seven gesture classes, and is intended to support the study of synthetic-to-real domain shift for video-based action recognition. Our work expands upon existing datasets by focusing the action classes on gestures for human-robot teaming, as well as by enabling investigation of domain shift in both ground and aerial views. We present baseline results using state-of-the-art action recognition and domain adaptation algorithms and offer initial insight on tackling the synthetic-to-real and ground-to-air domain shifts. Instructions on accessing the dataset can be found at https://github.com/reddyav1/RoCoG-v2. Arun V. Reddy, Ketul Shah, William Paul, Rohita Mocharla, Judy Hoffman, Kapil D. Katyal, Dinesh Manocha, Celso de Melo, Rama Chellappa |
ICRA | 8 |
| 2023 | AZTR: Aerial Video Action Recognition with Auto Zoom and Temporal ReasoningabstractWe propose a novel approach for aerial video action recognition. Our method is designed for videos captured using UAVs and can run on edge or mobile devices. We present a learning-based approach that uses customized auto zoom to automatically identify the human target and scale it appropriately. This makes it easier to extract the key features and reduces the computational overhead. We also present an efficient temporal reasoning algorithm to capture the action information along the spatial and temporal domains within a controllable computational cost. Our approach has been implemented and evaluated both on the desktop with high-end GPUs and on the low power Robotics RB5 Platform for robots and drones. In practice, we achieve$6.1-7.4 \%$improvement over SOTA in Top-1 accuracy on the RoCoG-v2 dataset, 8.3-10.4% improvement on the UAV-Human dataset and 3.2% improvement on the Drone Action dataset. Xijun Wang 0002, Ruiqi Xian, Tianrui Guan, Celso de Melo, Stephen M. Nogar, Aniket Bera, Dinesh Manocha |
ICRA | 4 |
| 2023 | Multi-View Action Recognition using Contrastive LearningabstractIn this work, we present a method for RGB-based action recognition using multi-view videos. We present a supervised contrastive learning framework to learn a feature embedding robust to changes in viewpoint, by effectively leveraging multi-view data. We use an improved supervised contrastive loss and augment the positives with those coming from synchronized viewpoints. We also propose a new approach to use classifier probabilities to guide the selection of hard negatives in the contrastive loss, to learn a more discriminative representation. Negative samples from confusing classes based on posterior are weighted higher. We also show that our method leads to better domain generalization compared to the standard supervised training based on synthetic multi-view data. Extensive experiments on real (NTU-60, NTU-120, NUMA) and synthetic (RoCoG) data demonstrate the effectiveness of our approach. Ketul Shah, Anshul Shah 0001, Chun Pong Lau 0001, Celso de Melo, Rama Chellappa |
WACV | 4 |
| 2023 | Social Functions of Machine Emotional ExpressionsabstractVirtual humans and social robots frequently generate behaviors that human observers naturally see as expressing emotion. In this review article, we highlight that these expressions can have important benefits for human–machine interaction. We first summarize the psychological findings on how emotional expressions achieve important social functions in human relationships and highlight that artificial emotional expressions can serve analogous functions in human–machine interaction. We then review computational methods for determining what expressions make sense to generate within the context of interaction and how to realize those expressions across multiple modalities, such as facial expressions, voice, language, and touch. The use of synthetic expressions raises a number of ethical concerns, and we conclude with a discussion of principles to achieve the benefits of machine emotion in ethical ways. Celso de Melo, Jonathan Gratch, Stacy Marsella, Catherine Pelachaud |
Proc. IEEE | 1 |
| 2022 | The influence of emotional expressions of an industrial robot on human collaborative decision-makingabstractIn recent years, robots have been equipped with the ability to express emotions and have begun building social relationships with people. However, the significance and effectiveness of incorporating emotion in industrial robots, which have a strong instrumental nature, is not fully understood. We investigated how emotional expressions of an industrial robot influence human collaborative decision-making. The participants (n=52), in a laboratory experiment, engaged in a dessert survival task with an arm robot in a 2 (emotion expression: present vs. absent) × 2 (competence: high vs. low) between-participants study. Emotion was expressed using color through a LED strip of lights - e.g., anger was conveyed by flashing red. The results showed that emotion expression and competence did not influence the final agreement and, in fact, emotion expressions made the interaction longer, emphasizing the difficulty in communicating emotion and the reason for those expressions. We discuss lessons learnt and provide insight on improving the value of emotion expression in industrial robots. Koki Usui, Kazunori Terada, Celso de Melo |
ACII | 3 |
| 2022 | Not Just Streaks: Towards Ground Truth for Single Image Deraining
Yunhao Ba, Howard Zhang, Ethan Yang, Akira Suzuki 0002, Arnold Pfahnl, Chethan Chinder Chandrappa, Celso de Melo, Suya You, Stefano Soatto, Alex Wong 0001, Achuta Kadambi |
ECCV (7) | 7 |
| 2021 | Introduction to the Special Issue On Computational Modelling of EmotionabstractThe papers in this special issue focus on computational modeling of emotion recognition. Emotions play a pervasive role in personal, social, and professional life. As artificially intelligent systems become pervasive in our lives, it is important that these systems are able to understand emotion in humans and simulate the function of emotion to be effective in their interactions with people. Computational models of emotion contribute towards this goal by, on the one hand, serving as a means to test emotion theories and help understand the function of emotion and, on the other, as the end in itself by simulating appropriate emotion and its downstream consequences – such as expressions of emotion – in computational agents. This special issue presents a critical overview of this cross-disciplinary field, with contributions from some of the leading scholars in cognitive psychology and affective computing, focusing both on theory and practice. Celso de Melo, Dean Petters, Joel Parthemore, David C. Moffat, Christian Becker-Asano |
IEEE Trans. Affect. Comput. | 1 |
| 2020 | Vision-Based Gesture Recognition in Human-Robot Teams Using Synthetic DataabstractBuilding successful collaboration between humans and robots requires efficient, effective, and natural communication. Here we study a RGB-based deep learning approach for controlling robots through gestures (e.g., "follow me"). To address the challenge of collecting high-quality annotated data from human subjects, synthetic data is considered for this domain. We contribute a dataset of gestures that includes real videos with human subjects and synthetic videos from our custom simulator. A solution is presented for gesture recognition based on the state-of-the-art I3D model. Comprehensive testing was conducted to optimize the parameters for this model. Finally, to gather insight on the value of synthetic data, several experiments are described that systematically study the properties of synthetic data (e.g., gesture variations, character variety, generalization to new gestures). We discuss practical implications for the design of effective human-robot collaboration and the usefulness of synthetic data for deep learning. Celso de Melo, Brandon Rothrock, Prudhvi Gurram, Oytun Ulutan, B. S. Manjunath |
IROS | 1 |
| 2020 | Reducing Task Load with an Embodied Intelligent Virtual Assistant for Improved Performance in Collaborative Decision MakingabstractCollaboration in a group has the potential to achieve more effective solutions for challenging problems, but collaboration per se is not an easy task, rather a stressful burden if the collaboration partners do not communicate well with each other. While Intelligent Virtual Assistants (IVAs), such as Amazon Alexa, are becoming part of our daily lives, there are increasing occurrences in which we collaborate with such IVAs for our daily tasks. Although IVAs can provide important support to users, the limited verbal interface in the current state of IVAs lacks the ability to provide effective non-verbal social cues, which is critical for improving collaborative performance and reducing task load.In this paper, we investigate the effects of IVA embodiment on collaborative decision making. In a within-subjects study, participants performed a desert survival task in three conditions: (1) performing the task alone, (2) working with a disembodied voice assistant, and (3) working with an embodied assistant. Our results show that both assistant conditions led to higher performance over when performing the task alone, but interestingly the reported task load with the embodied assistant was significantly lower than with the disembodied voice assistant. We discuss the findings with implications for effective and efficient collaborations with IVAs while also emphasizing the increased social presence and richness of the embodied assistant. Kangsoo Kim, Celso de Melo, Nahal Norouzi, Gerd Bruder, Greg Welch |
VR | 2 |
| 2019 | Toward a Unified Theory of Learned Trust in Interpersonal and Human-Machine InteractionsabstractA proposal for a unified theory of learned trust implemented in a cognitive architecture is presented. The theory is instantiated as a computational cognitive model of learned trust that integrates several seemingly unrelated categories of findings from the literature on interpersonal and human-machine interactions and makes unintuitive predictions for future studies. The model relies on a combination of learning mechanisms to explain a variety of phenomena such as trust asymmetry, the higher impact of early trust breaches, the black-hat/white-hat effect, the correlation between trust and cognitive ability, and the higher resilience of interpersonal as compared to human-machine trust. In addition, the model predicts that trust decays in the absence of evidence of trustworthiness or untrustworthiness. The implications of the model for the advancement of the theory on trust are discussed. Specifically, this work suggests two more trust antecedents on the trustor's side: perceived trust necessity and cognitive ability to detect cues of trustworthiness. Ion Juvina, Michael G. Collins, Othalia Larue, William G. Kennedy, Ewart de Visser, Celso de Melo |
ACM Trans. Interact. Intell. Syst. | 6 |
| 2018 | Social decisions and fairness change when people's interests are represented by autonomous agents
Celso de Melo, Stacy Marsella, Jonathan Gratch |
Auton. Agents Multi Agent Syst. | 1 |
| 2016 | Neurophysiological Effects of Negotiation Framing
Peter Khooshabeh, Rebecca Lin, Celso de Melo, Jonathan Gratch, Brett Ouimette, Jim Blascovich |
CogSci | 3 |
| 2016 | People Do Not Feel Guilty About Exploiting MachinesabstractGuilt and envy play an important role in social interaction. Guilt occurs when individuals cause harm to others or break social norms. Envy occurs when individuals compare themselves unfavorably to others and desire to benefit from the others’ advantage. In both cases, these emotions motivate people to act and change the status quo: following guilt, people try to make amends for the perceived transgression, and following envy, people try to harm envied others. In this article, we present two experiments that study participants’ experience of guilt and envy when engaging in social decision making with machines and humans. The results showed that, though experiencing the same level of envy, people felt considerably less guilt with machines than with humans. These effects occurred both with subjective and behavioral measures of guilt and envy, and in three different economic games: public goods, ultimatum, and dictator game. This poses an important challenge for human-computer interaction because, as shown here, it leads people to systematically exploit machines, when compared to humans. We discuss theoretical and practical implications for the design of human-machine interaction systems that hope to achieve the kind of efficiency -- cooperation, fairness, reciprocity, etc. -- we see in human-human interaction. Celso de Melo, Stacy Marsella, Jonathan Gratch |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2015 | People show envy, not guilt, when making decisions with machinesabstractResearch shows that people consistently reach more efficient solutions than those predicted by standard economic models, which assume people are selfish. Artificial intelligence, in turn, seeks to create machines that can achieve these levels of efficiency in human-machine interaction. However, as reinforced in this paper, people's decisions are systematically less efficient - i.e., less fair and favorable - with machines than with humans. To understand the cause of this bias, we resort to a well-known experimental economics model: Fehr and Schmidt's inequity aversion model. This model accounts for people's aversion to disadvantageous outcome inequality (envy) and aversion to advantageous outcome inequality (guilt). We present an experiment where participants engaged in the ultimatum and dictator games with human or machine counterparts. By fitting this data to Fehr and Schmidt's model, we show that people acted as if they were just as envious of humans as of machines; but, in contrast, people showed less guilt when making unfavorable decisions to machines. This result, thus, provides critical insight into this bias people show, in economic settings, in favor of humans. We discuss implications for the design of machines that engage in social decision making with humans. Celso de Melo, Jonathan Gratch |
ACII | 1 |
| 2015 | Beyond Believability: Quantifying the Differences Between Real and Virtual Humans
Celso de Melo, Jonathan Gratch |
IVA | 1 |
| 2015 | Physiological evidence for a dual process model of the social effects of emotion in computers
Ahyoung Choi, Celso de Melo, Peter Khooshabeh, Woontack Woo, Jonathan Gratch |
Int. J. Hum. Comput. Stud. | 2 |
| 2015 | Humans versus Computers: Impact of Emotion Expressions on People's Decision MakingabstractRecent research in perception and theory of mind reveals that people show different behavior and lower activation of brain regions associated with mentalizing (i.e., the inference of other's mental states) when engaged in decision making with computers, when compared to humans. These findings are important for affective computing because they suggest people's decisions might be influenced differently according to whether they believe emotional expressions shown in computers are being generated by algorithms or humans. To test this, we had people engage in a social dilemma (Experiment 1) or negotiation (Experiment 2) with virtual humans that were either perceived to be agents (i.e., controlled by computers) or avatars (i.e., controlled by humans). The results showed that such perceptions have a deep impact on people's decisions: in Experiment 1, people cooperated more with virtual humans that showed cooperative facial displays (e.g., joy after mutual cooperation) than competitive displays (e.g., joy when the participant was exploited) but, the effect was stronger with avatars (d = .601) than with agents (d = .360); in Experiment 2, people conceded more to angry than neutral virtual humans but, again, the effect was much stronger with avatars (d = 1.162) than with agents (d = .066). Participants also showed less anger towards avatars and formed more positive impressions of avatars when compared to agents. Celso de Melo, Jonathan Gratch, Peter J. Carnevale |
IEEE Trans. Affect. Comput. | 1 |
| 2014 | The Importance of Cognition and Affect for Artificially Intelligent Decision MakersabstractAgency - the capacity to plan and act - and experience - the capacity to sense and feel - are two critical aspects that determine whether people will perceive non-human entities, such as autonomous agents, to have a mind. There is evidence that the absence of either can reduce cooperation. We present an experiment that tests the necessity of both for cooperation with agents. In this experiment we manipulated people's perceptions about the cognitive and affective abilities of agents, when engaging in the ultimatum game. The results indicated that people offered more money to agents that were perceived to make decisions according to their intentions (high agency), rather than randomly (low agency). Additionally, the results showed that people offered more money to agents that expressed emotion (high experience), when compared to agents that did not (low experience). We discuss the implications of this agency-experience theoretical framework for the design of artificially intelligent decision makers. Celso de Melo, Jonathan Gratch, Peter J. Carnevale |
AAAI | 1 |
| 2014 | Social Categorization and Cooperation between Humans and Computers
Celso de Melo, Peter J. Carnevale, Jonathan Gratch |
CogSci | 1 |
| 2013 | The Effect of Agency on the Impact of Emotion Expressions on People's Decision MakingabstractRecent research in neuroeconomics reveals that people show different behavior and lower activation of brain regions associated with mentalizing (i.e., the inference of other's mental states) when engaged in decision making tasks with a computer, when compared to a human. These findings are important for affective computing because they suggest people's decision making might be influenced differently according to whether they believe the emotional expressions shown by a computer are being generated by a computer algorithm or a human. To test this, we had people engage in a social dilemma (Experiment 1) or a negotiation (Experiment 2) with virtual humans that were either agents (i.e., controlled by computers) or avatars (i.e., controlled by humans). The results show a clear agency effect: in Experiment 1, people cooperated more with virtual humans that showed facial cooperative displays (e.g., joy after mutual cooperation) rather than competitive displays (e.g., joy when the participant was exploited) but, the effect was only significant with avatars, in Experiment 2, people conceded more to an angry than a neutral virtual human but, once again, the effect was only significant with avatars. Celso de Melo, Jonathan Gratch, Peter J. Carnevale |
ACII | 1 |
| 2013 | Cooperative Strategies with Incongruent Facial Expressions Cause Cardiovascular Threat
Peter Khooshabeh, Celso de Melo, Brooks Volkman, Jonathan Gratch, Jim Blascovich, Peter J. Carnevale |
CogSci | 2 |
| 2012 | Reverse appraisal: The importance of appraisals for the effect of emotion displays on people's decision making in a social dilemma
Celso de Melo, Jonathan Gratch, Peter J. Carnevale, Stephen Read |
CogSci | 1 |
| 2012 | The Effect of Virtual Agents' Emotion Displays and Appraisals on People's Decision Making in Negotiation
Celso de Melo, Peter J. Carnevale, Jonathan Gratch |
IVA | 1 |
| 2012 | Affective engagement to emotional facial expressions of embodied social agents in a decision-making gameabstractABSTRACT Previous research illustrates that people can be influenced by the emotional displays of computer‐generated agents. What is less clear is if these influences arise from cognitive or affective process (i.e., do people use agent displays as information or do they provoke user emotions). To unpack these processes, we examine the decisions and physiological reactions of participants (heart rate and electrodermal activity) when engaged in a decision task (prisoner's dilemma game) with emotionally expressive agents. Our results replicate findings that people's decisions are influenced by such emotional displays, but these influences differ depending on the extent to which these displays provoke an affective response. Specifically, we show that an individual difference known as electrodermal lability predicts the extent to whether people will engage affectively or strategically with such agents, thereby better predicting their decisions. We discuss implications for designing agent facial expressions to enhance social interaction between humans and agents. Copyright © 2012 John Wiley & Sons, Ltd. Ahyoung Choi, Celso de Melo, Woontack Woo, Jonathan Gratch |
Comput. Animat. Virtual Worlds | 2 |
| 2011 | The Influence of Emotion Expression on Perceptions of Trustworthiness in NegotiationabstractWhen interacting with computer agents, people make inferences about various characteristics of these agents, such as their reliability and trustworthiness. These perceptions are significant, as they influence people's behavior towards the agents, and may foster or inhibit repeated interactions between them. In this paper we investigate whether computer agents can use the expression of emotion to influence human perceptions of trustworthiness. In particular, we study human-computer interactions within the context of a negotiation game, in which players make alternating offers to decide on how to divide a set of resources. A series of negotiation games between a human and several agents is then followed by a "trust game." In this game people have to choose one among several agents to interact with, as well as how much of their resources they will trust to it. Our results indicate that, among those agents that displayed emotion, those whose expression was in accord with their actions (strategy) during the negotiation game were generally preferred as partners in the trust game over those whose emotion expressions and actions did not mesh. Moreover, we observed that when emotion does not carry useful new information, it fails to strongly influence human decision-making behavior in a negotiation setting. Dimitrios Antos, Celso de Melo, Jonathan Gratch, Barbara J. Grosz |
AAAI | 2 |
| 2011 | A Computer Model of the Interpersonal Effect of Emotion Displayed in a Social Dilemma
Celso de Melo, Peter J. Carnevale, Dimitrios Antos, Jonathan Gratch |
ACII (1) | 1 |
| 2011 | Reverse Appraisal: Inferring from Emotion Displays who is the Cooperator and the Competitor in a Social Dilemma
Celso de Melo, Peter J. Carnevale, Jonathan Gratch |
CogSci | 1 |
| 2010 | Evolving Expression of Emotions Through Color in Virtual Humans Using Genetic Algorithms
Celso de Melo, Jonathan Gratch |
ICCC | 1 |
| 2010 | The Influence of Emotions in Embodied Agents on Human Decision-Making
Celso de Melo, Peter J. Carnevale, Jonathan Gratch |
IVA | 1 |
| 2010 | Real-time expression of affect through respirationabstractAbstract Affect has been shown to influence respiration in people. This paper takes this insight and proposes a real‐time model to express affect through respiration in virtual humans. Fourteen affective states are explored: excitement, relaxation, focus, pain, relief, boredom, anger, fear, panic, disgust, surprise, startle, sadness, and joy. Specific respiratory patterns are described from the literature for each of these affective states. Then, a real‐time model of respiration is proposed that uses morphing to animate breathing and provides parameters to control respiration rate, respiration depth and the respiration cycle curve. These parameters are used to implement the respiratory patterns. Finally, a within‐subjects study is described where subjects are asked to classify videos of the virtual human expressing each affective state with or without the specific respiratory patterns. The study was presented to 41 subjects and the results show that the model improved perception of excitement, pain, relief, boredom, anger, fear, panic, disgust, and startle. Copyright © 2010 John Wiley & Sons, Ltd. Celso de Melo, Patrick G. Kenny, Jonathan Gratch |
Comput. Animat. Virtual Worlds | 1 |
| 2009 | Creative expression of emotions in virtual humansabstract'Works of art (…) can be expressive of human qualities: one of the most characteristic and pervasive features of art is that percepts (lines, colors, progressions of musical tones) can be and are suffused with affect.' Celso de Melo, Jonathan Gratch |
FDG | 1 |
| 2009 | Expression of Emotions Using Wrinkles, Blushing, Sweating and Tears
Celso de Melo, Jonathan Gratch |
IVA | 1 |
| 2009 | Expression of Moral Emotions in Cooperating Agents
Celso de Melo, Jonathan Gratch |
IVA | 1 |
| 2008 | Evolving Expression of Emotions in Virtual Humans Using Lights and Pixels
Celso de Melo, Jonathan Gratch |
IVA | 1 |
| 2007 | Expression of Emotions in Virtual Humans Using Lights, Shadows, Composition and Filters
Celso de Melo, Ana Paiva 0001 |
ACII | 1 |
| 2006 | A Story About Gesticulation Expression
Celso de Melo, Ana Paiva 0001 |
IVA | 1 |
| 2006 | Storytelling - The Difference Between Fantasy and Reality
Guilherme Raimundo, João P. Cabral, Celso de Melo, Luís C. Oliveira, Ana Paiva 0001 |
IVA | 3 |
| 2006 | Multimodal expression in virtual humansabstractAbstract This work proposes a real‐time virtual human multimodal expression model. Five modalities explore the affordances of the body: deterministic, non‐deterministic, gesticulation, facial, and vocal expression. Deterministic expression is keyframe body animation. Non‐deterministic expression is robotics‐based procedural body animation. Vocal expression is voice synthesis, through Festival, and parameterization, through SABLE. Facial expression is lip‐synch and emotion expression through a parametric muscle‐based face model. Inspired by psycholinguistics, gesticulation expression is unconventional, idiosyncratic, and unconscious hand gestures animation described as sequences of Portuguese Sign Language hand shapes, positions and orientations. Inspired by the arts, one modality goes beyond the body to explore the affordances of the environment and express emotions through camera, lights, and music. To control multimodal expression, this work proposes a high‐level integrated synchronized markup language—expressive markup language. Finally, three studies, involving a total of 197 subjects, evaluated the model in storytelling contexts and produced promising results. Copyright © 2006 John Wiley & Sons, Ltd. Celso de Melo, Ana Paiva 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2005 | Environment Expression: Expressing Emotions Through Cameras, Lights and Music
Celso de Melo, Ana Paiva 0001 |
ACII | 1 |