Kaixin Ma

dblp:203/9347 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
22since 2021 · last 2025
0000-0001-7414-5673ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 COLUMBUS: Evaluating COgnitive Lateral Understanding Through Multiple-Choice reBUSes
abstract
While visual question-answering (VQA) benchmarks have catalyzed the development of reasoning techniques, they have focused on vertical thinking. Effective problem-solving also necessitates lateral thinking, which remains understudied in AI and has not been used to test visual perception systems. To bridge this gap, we formulate visual lateral thinking as a multiple-choice question-answering task and describe a three-step taxonomy-driven methodology for instantiating task examples. Then, we develop COLUMBUS, a synthetic benchmark that applies the task pipeline to create QA sets with text and icon rebus puzzles based on publicly available collections of compounds and common phrases. COLUMBUS comprises over 1,000 puzzles, each with four answer candidates. While the SotA vision language models (VLMs) achieve decent performance, our evaluation demonstrates a substantial gap between humans and models. VLMs benefit from human-curated descriptions but struggle to self-generate such representations at the right level of abstraction.
Koen Kraaijveld, Yifan Jiang 0001, Kaixin Ma, Filip Ilievski
AAAI3
2025 OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
abstract
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Hongming Zhang, Tianqing Fang, Zhenzhong Lan, Dong Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Tianqing Fang, Zhen-Zhong Lan, Dong Yu 0001
ACL (1)3
2025 WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World Model
abstract
Agent self-improvement, where agents autonomously train their underlying Large Language Model (LLM) on self-sampled trajectories, shows promising results but often stagnates in web environments due to limited exploration and under-utilization of pretrained web knowledge.To improve the performance of self-improvement, we propose a novel framework that introduces a co-evolving World Model LLM.This world model predicts the next observation based on the current observation and action within the web environment.The World Model serves dual roles: (1) as a virtual web server generating self-instructed training data to continuously refine the agent's policy, and (2) as an imagination engine during inference, enabling look-ahead simulation to guide action selection for the agent LLM.Experiments in real-world web environments (Mind2Web-Live, WebVoyager, and GAIAweb) show a 10% performance gain over existing self-evolving agents, demonstrating the efficacy and generalizability of our approach, without using any distillation from more powerful close-sourced models 1 .
Tianqing Fang, Hongming Zhang 0009, Zhisong Zhang, Kaixin Ma, Wenhao Yu 0002, Haitao Mi, Dong Yu 0001
EMNLP4
2025 Retrieval-augmented GUI Agents with Generative Guidelines
abstract
GUI agents powered by vision-language models (VLMs) show promise in automating complex digital tasks.However, their effectiveness in real-world applications is often limited by scarce training data and the inherent complexity of these tasks, which frequently require longtailed knowledge covering rare, unseen scenarios.We propose RAG-GUI , a lightweight VLM that leverages web tutorials at inference time.RAG-GUI is first warm-started via supervised finetuning (SFT) and further refined through self-guided rejection sampling finetuning (RSF).Designed to be model-agnostic, RAG-GUI functions as a generic plug-in that enhances any VLM-based agent.Evaluated across three distinct tasks, it consistently outperforms baseline agents and surpasses other inference baselines by 2.6% to 13.3% across two model sizes, demonstrating strong generalization and practical plug-and-play capabilities in real-world scenarios.
Ran Xu 0002, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Joyce C. Ho, Carl Yang 0001, Dong Yu 0001
EMNLP2
2025 DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
abstract
Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) have demonstrated impressive language/vision reasoning abilities, igniting the recent trend of building agents for targeted applications such as shopping assistants or AI software engineers. Recently, many data science benchmarks have been proposed to investigate their performance in the data science domain. However, existing data science benchmarks still fall short when compared to real-world data science applications due to their simplified settings. To bridge this gap, we introduce DSBench, a comprehensive benchmark designed to evaluate data science agents with realistic tasks. This benchmark includes 466 data analysis tasks and 74 data modeling tasks, sourced from Eloquence and Kaggle competitions. DSBench offers a realistic setting by encompassing long contexts, multimodal task backgrounds, reasoning with large data files and multi-table structures, and performing end-to-end data modeling tasks. Our evaluation of state-of-the-art LLMs, LVLMs, and agents shows that they struggle with most tasks, with the best agent solving only 34.12% of data analysis tasks and achieving a 34.74% Relative Performance Gap (RPG). These findings underscore the need for further advancements in developing more practical, intelligent, and autonomous data science agents.
Liqiang Jing, Zhehui Huang, Xiaoyang Wang 0001, Wenlin Yao, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Xinya Du, Dong Yu 0001
ICLR6
2025 RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
abstract
Large Language Models (LLMs) excel in code generation yet struggle with modern AI software engineering tasks. Unlike traditional function-level or file-level coding tasks, AI software engineering requires not only basic coding proficiency but also advanced skills in managing and interacting with code repositories. However, existing methods often overlook the need for repository-level code understanding, which is crucial for accurately grasping the broader context and developing effective solutions. On this basis, we present RepoGraph, a plug-in module that manages a repository-level structure for modern AI software engineering solutions. RepoGraph offers the desired guidance and serves as a repository-wide navigation for AI software engineers. We evaluate RepoGraph on the SWE-bench by plugging it into four different methods of two lines of approaches, where RepoGraph substantially boosts the performance of all systems, leading to a new state-of-the-art among open-source frameworks. Our analyses also demonstrate the extensibility and flexibility of RepoGraph by testing on another repo-level coding benchmark, CrossCodeEval. Our code is available at https://github.com/ozyyshr/RepoGraph.
Siru Ouyang, Wenhao Yu 0002, Kaixin Ma, Zilin Xiao, Zhihan Zhang 0001, Mengzhao Jia, Jiawei Han 0001, Hongming Zhang 0009, Dong Yu 0001
ICLR3
2024 WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
abstract
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, Dong Yu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Yong Dai 0001, Hongming Zhang 0009, Zhen-Zhong Lan, Dong Yu 0001
ACL (1)3
2024 Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models
abstract
Retrieval-augmented language model (RALM) represents a significant advancement in mitigating factual hallucination by leveraging external knowledge sources.However, the reliability of the retrieved information is not always guaranteed, and the retrieval of irrelevant data can mislead the response generation.Moreover, standard RALMs frequently neglect their intrinsic knowledge due to the interference from retrieved information.In instances where the retrieved information is irrelevant, RALMs should ideally utilize their intrinsic knowledge or, in the absence of both intrinsic and retrieved knowledge, opt to respond with "unknown" to avoid hallucination.In this paper, we introduces CHAIN-OF-NOTE (CON), a novel approach to improve robustness of RALMs in facing noisy, irrelevant documents and in handling unknown scenarios.The core idea of CON is to generate sequential reading notes for each retrieved document, enabling a thorough evaluation of their relevance to the given question and integrating this information to formulate the final answer.Our experimental results show that GPT-4, when equipped with CON, outperforms the CHAIN-OF-THOUGHT approach.Besides, we utilized GPT-4 to create 10K CON data, subsequently trained on LLaMa-2 7B model.Our experiments across four open-domain QA benchmarks show that fine-tuned RALMs equipped with CON significantly outperform standard fine-tuned RALMs.
Wenhao Yu 0002, Hongming Zhang 0009, Xiaoman Pan, Peixin Cao, Kaixin Ma, Hongwei Wang 0001, Dong Yu 0001
EMNLP5
2024 Dense X Retrieval: What Retrieval Granularity Should We Use?
abstract
Dense retrieval has become a prominent method to obtain relevant context or world knowledge in open-domain NLP tasks.When we use a learned dense retriever on a retrieval corpus at inference time, an often-overlooked design choice is the retrieval unit in which the corpus is indexed, e.g.document, passage, or sentence.We discover that the retrieval unit choice significantly impacts the performance of both retrieval and downstream tasks.Distinct from the typical approach of using passages or sentences, we introduce a novel retrieval unit, proposition, for dense retrieval.Propositions are defined as atomic expressions within text, each encapsulating a distinct factoid and presented in a concise, self-contained natural language format.We conduct an empirical comparison of different retrieval granularity.Our experiments reveal that indexing a corpus by fine-grained units such as propositions significantly outperforms passage-level units in retrieval tasks.Moreover, constructing prompts with fine-grained retrieved units for retrievalaugmented language models improves the performance of downstream QA tasks given a specific computation budget.
Hongwei Wang 0010, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Dong Yu 0001
EMNLP5
2024 MARVEL: Multidimensional Abstraction and Reasoning through Visual Evaluation and Learning
abstract
While multi-modal large language models (MLLMs) have shown significant progress across popular visual reasoning benchmarks, whether they possess abstract visual reasoning abilities remains an open question. Similar to the Sudoku puzzles, abstract visual reasoning (AVR) problems require finding high-level patterns (e.g., repetition constraints on numbers) that control the input shapes (e.g., digits) in a specific task configuration (e.g., matrix). However, existing AVR benchmarks only consider a limited set of patterns (addition, conjunction), input shapes (rectangle, square), and task configurations (3 × 3 matrices). And they fail to capture all abstract reasoning patterns in human cognition necessary for addressing real-world tasks, such as geometric properties and object boundary understanding in real-world navigation. To evaluate MLLMs’ AVR abilities systematically, we introduce MARVEL founded on the core knowledge system in human cognition, a multi-dimensional AVR benchmark with 770 puzzles composed of six core knowledge patterns, geometric and abstract shapes, and five different task configurations. To inspect whether the model performance is grounded in perception or reasoning, MARVEL complements the standard AVR question with perception questions in a hierarchical evaluation framework. We conduct comprehensive experiments on MARVEL with ten representative MLLMs in zero-shot and few-shot settings. Our experiments reveal that all MLLMs show near-random performance on MARVEL, with significant performance gaps (40%) compared to humans across all patterns and task configurations. Further analysis of perception questions reveals that MLLMs struggle to comprehend the visual features (near-random performance). Although closed-source MLLMs, such as GPT-4V, show a promising understanding of reasoning patterns (on par with humans) after adding textual descriptions, this advantage is hindered by their weak perception abilities. We release our entirecode and dataset at https://github.com/1171-jpg/MARVEL_AVR.
Yifan Jiang 0001, Jiarui Zhang 0002, Kexuan Sun 0002, Zhivar Sourati, Kian Ahrabian, Kaixin Ma, Filip Ilievski, Jay Pujara
NeurIPS6
2023 Chain-of-Skills: A Configurable Model for Open-Domain Question Answering
abstract
The retrieval model is an indispensable component for real-world knowledge-intensive tasks, e.g., open-domain question answering (ODQA).As separate retrieval skills are annotated for different datasets, recent work focuses on customized methods, limiting the model transferability and scalability.In this work, we propose a modular retriever where individual modules correspond to key skills that can be reused across datasets.Our approach supports flexible skill configurations based on the target domain to boost performance.To mitigate task interference, we design a novel modularization parameterization inspired by sparse Transformer.We demonstrate that our model can benefit from self-supervised pretraining on Wikipedia and fine-tuning using multiple ODQA datasets, both in a multi-task fashion.Our approach outperforms recent self-supervised retrievers in zero-shot evaluations and achieves state-ofthe-art fine-tuned retrieval performance on NQ, HotpotQA and OTT-QA.
Kaixin Ma, Hao Cheng 0002, Yu Zhang 0044, Xiaodong Liu 0002, Eric Nyberg, Jianfeng Gao 0001
ACL (1)1
2023 Transferring Procedural Knowledge Across Commonsense Tasks
abstract
Stories about everyday situations are an essential part of human communication, motivating the need to develop AI agents that can reliably understand these stories. Despite the long list of supervised methods for story completion and procedural understanding, current AI fails to generalize its procedural reasoning to unseen stories. This paper is based on the hypothesis that the generalization can be improved by associating downstream prediction with fine-grained modeling and the abstraction of procedural knowledge in stories. To test this hypothesis, we design LEAP: a comprehensive framework that reasons over stories by jointly considering their (1) overall plausibility, (2) conflict sentence pairs, and (3) participant physical states. LEAP integrates state-of-the-art modeling architectures, training regimes, and augmentation strategies based on natural and synthetic stories. To address the lack of densely annotated training data on participants and their physical states, we devise a robust automatic labeler based on semantic parsing and few-shot prompting with large language models. Our experiments with in- and out-of-domain tasks reveal insights into the interplay of architectures, training regimes, and augmentation strategies. LEAP’s labeler consistently improves performance on out-of-domain datasets, while our case studies show that the dense annotation supports explainability.
Yifan Jiang 0001, Filip Ilievski, Kaixin Ma
ECAI3
2023 BRAINTEASER: Lateral Thinking Puzzles for Large Language Models
abstract
The success of language models has inspired the NLP community to attend to tasks that require implicit and complex reasoning, relying on human-like commonsense mechanisms.While such vertical thinking tasks have been relatively popular, lateral thinking puzzles have received little attention.To bridge this gap, we devise BRAINTEASER: a multiple-choice Question Answering task designed to test the model's ability to exhibit lateral thinking and defy default commonsense associations.We design a three-step procedure for creating the first lateral thinking benchmark, consisting of data collection, distractor generation, and generation of reconstruction examples, leading to 1,100 puzzles with high-quality annotations.To assess the consistency of lateral reasoning by models, we enrich BRAINTEASER based on a semantic and contextual reconstruction of its questions.Our experiments with state-ofthe-art instruction-and commonsense language models reveal a significant gap between human and model performance, which is further widened when consistency across reconstruction formats is considered.We make all of our code and data available to stimulate work on developing and evaluating lateral thinking models.
Yifan Jiang 0001, Filip Ilievski, Kaixin Ma, Zhivar Sourati
EMNLP3
2023 Knowledge-enhanced Agents for Interactive Text Games
abstract
Communication via natural language is a key aspect of machine intelligence, and it requires computational models to learn and reason about world concepts, with varying levels of supervision. Significant progress has been made on fully-supervised non-interactive tasks, such as question-answering and procedural text understanding. Yet, various sequential interactive tasks, as in text-based games, have revealed limitations of existing approaches in terms of coherence, contextual awareness, and their ability to learn effectively from the environment. In this paper, we propose a knowledge-injection framework for improved functional grounding of agents in text-based games. Specifically, we consider two forms of domain knowledge that we inject into learning-based agents: memory of previous correct actions and affordances of relevant objects in the environment. Our framework supports two representative model classes: reinforcement learning agents and language model agents. Furthermore, we devise multiple injection strategies for the above domain knowledge types and agent architectures, including injection via knowledge graphs and augmentation of the existing input encoding strategies. We experiment with four models on the 10 tasks in the ScienceWorld text-based game environment, to illustrate the impact of knowledge injection on various model configurations and challenging task settings. Our findings provide crucial insights into the interplay between task properties, model architectures, and domain knowledge for interactive contexts.
Prateek Chhikara, Jiarui Zhang 0002, Filip Ilievski, Jonathan Francis, Kaixin Ma
K-CAP5
2023 A Study of Situational Reasoning for Traffic Understanding
abstract
Intelligent Traffic Monitoring (ITMo) technologies hold the potential for improving road safety/security and for enabling smart city infrastructure. Understanding traffic situations requires a complex fusion of perceptual information with domain-specific and causal commonsense knowledge. Whereas prior work has provided benchmarks and methods for traffic monitoring, it remains unclear whether models can effectively align these information sources and reason in novel scenarios. To address this assessment gap, we devise three novel text-based tasks for situational reasoning in the traffic domain: i) BDD-QA, which evaluates the ability of Language Models (LMs) to perform situational decision-making, ii) TV-QA, which assesses LMs' abilities to reason about complex event causality, and iii) HDT-QA, which evaluates the ability of models to solve human driving exams. We adopt four knowledge-enhanced methods that have shown generalization capability across language reasoning tasks in prior work, based on natural language inference, commonsense knowledge-graph self-supervision, multi-QA joint training, and dense retrieval of domain information. We associate each method with a relevant knowledge source, including knowledge graphs, relevant benchmarks, and driving manuals. In extensive experiments, we benchmark various knowledge-aware methods against the three datasets, under zero-shot evaluation; we provide in-depth analyses of model performance on data partitions and examine model predictions categorically, to yield useful insights on traffic understanding, given different background knowledge and reasoning strategies.
Jiarui Zhang 0002, Filip Ilievski, Kaixin Ma, Aravinda Kollaa, Jonathan Francis, Alessandro Oltramari
KDD3
2022 Open Domain Question Answering with A Unified Knowledge Interface
abstract
The retriever-reader framework is popular for open-domain question answering (ODQA) due to its ability to use explicit knowledge.Although prior work has sought to increase the knowledge coverage by incorporating structured knowledge beyond text, accessing heterogeneous knowledge sources through a unified interface remains an open question.While data-to-text generation has the potential to serve as a universal interface for data and text, its feasibility for downstream tasks remains largely unknown.In this work, we bridge this gap and use the data-to-text method as a means for encoding structured knowledge for ODQA.Specifically, we propose a verbalizer-retrieverreader framework for ODQA over data and text where verbalized tables from Wikipedia and graphs from Wikidata are used as augmented knowledge sources.We show that our Unified Data and Text QA, UDT-QA, can effectively benefit from the expanded knowledge index, leading to large gains over textonly baselines.Notably, our approach sets the single-model state-of-the-art on Natural Questions.Furthermore, our analyses indicate that verbalized knowledge is preferred for answer reasoning for both adapted and hot-swap settings.
Kaixin Ma, Hao Cheng 0002, Xiaodong Liu 0003, Eric Nyberg, Jianfeng Gao 0001
ACL (1)1
2022 Coalescing Global and Local Information for Procedural Text Understanding
abstract
Procedural text understanding is a challenging language reasoning task that requires models to track entity states across the development of a narrative. We identify three core aspects required for modeling this task, namely the local and global view of the inputs, as well as the global view of outputs. Prior methods have considered a subset of these aspects, which leads to either low precision or low recall. In this paper, we propose a new model Coalescing Global and Local Information (CGLI), which builds entity- and timestep-aware input representations (local input) considering the whole context (global input), and we jointly model the entity states with a structured prediction objective (global output). Thus, CGLI simultaneously optimizes for both precision and recall. Moreover, we extend CGLI with additional output layers and integrate it into a story reasoning framework. Extensive experiments on a popular procedural text understanding dataset show that our model achieves state-of-the-art results, while experiments on a story reasoning benchmark show the positive impact of our model on downstream reasoning.
Kaixin Ma, Filip Ilievski, Jonathan Francis, Eric Nyberg, Alessandro Oltramari
COLING1
2021 Knowledge-driven Data Construction for Zero-shot Evaluation in Commonsense Question Answering
abstract
Recent developments in pre-trained neural language modeling have led to leaps in accuracy on common-sense question-answering benchmarks. However, there is increasing concern that models overfit to specific tasks, without learning to utilize external knowledge or perform general semantic reasoning. In contrast, zero-shot evaluations have shown promise as a more robust measure of a model’s general reasoning abilities. In this paper, we propose a novel neuro-symbolic framework for zero-shot question answering across commonsense tasks. Guided by a set of hypotheses, the framework studies how to transform various pre-existing knowledge resources into a form that is most effective for pre-training models. We vary the set of language models, training regimes, knowledge sources, and data generation strategies, and measure their impact across tasks. Extending on prior work, we devise and compare four constrained distractor-sampling strategies. We provide empirical results across five commonsense question-answering tasks with data generated from five external knowledge resources. We show that, while an individual knowledge graph is better suited for specific tasks, a global knowledge graph brings consistent gains across different tasks. In addition, both preserving the structure of the task as well as generating fair and informative questions help language models learn more effectively.
Kaixin Ma, Filip Ilievski, Jonathan Francis, Yonatan Bisk, Eric Nyberg, Alessandro Oltramari
AAAI1
2021 Exploring Strategies for Generalizable Commonsense Reasoning with Pre-trained Models
abstract
Commonsense reasoning benchmarks have been largely solved by fine-tuning language models.The downside is that fine-tuning may cause models to overfit to task-specific data and thereby forget their knowledge gained during pre-training.Recent works only propose lightweight model updates as models may already possess useful knowledge from past experience, but a challenge remains in understanding what parts and to what extent models should be refined for a given task.In this paper, we investigate what models learn from commonsense reasoning datasets.We measure the impact of three different adaptation methods on the generalization and accuracy of models.Our experiments with two models show that fine-tuning performs best, by learning both the content and the structure of the task, but suffers from overfitting and limited generalization to novel answers.We observe that alternative adaptation methods like prefixtuning have comparable accuracy, but generalize better to unseen answers and are more robust to adversarial splits.
Kaixin Ma, Filip Ilievski, Jonathan Francis, Satoru Ozaki, Eric Nyberg, Alessandro Oltramari
EMNLP (1)1
2021 Audio-Visual Event Recognition Through the Lens of Adversary
abstract
As audio/visual classification models are widely deployed for sensitive tasks like content filtering at scale, it is critical to understand their robustness along with improving the accuracy. This work aims to study several key questions related to multimodal learning through the lens of adversarial noises: 1) The trade-off between early/middle/late fusion affecting its robustness and accuracy 2) How does different frequency/time domain features contribute to the robustness? 3) How does different neural modules contribute against the adversarial noise? In our experiment, we construct adversarial examples to attack state-of-the-art neural models trained on Google AudioSet.[1]1We compare how much attack potency in terms of adversarial perturbation of size using different Lpnorms we would need to "deactivate" the victim model. Using adversarial noise to ablate multimodal models, we are able to provide insights into what is the best potential fusion strategy to balance the model parameters/accuracy and robustness trade-off, and distinguish the robust features versus the non-robust features that various neural networks model tend to learn.
Juncheng Li 0001, Kaixin Ma, Shuhui Qu, Po-Yao Huang 0001, Florian Metze
ICASSP2
2021 PTWA: Pre-Training with Word Attention for Chinese Named Entity Recognition
abstract
Recently, the character-based model that incorporates potential word information has proven effective for Chinese named entity recognition (NER). However, due to the independence of the pre-trained character model and the lexicon, it will cause the embedding space to be misaligned and cannot be combined well. Chinese pre-trained encoders usually process text as characters. It ignores the information carried by the larger granular information, so the encoder cannot easily adapt to certain character combinations. Because large-grained information is ignored and Chinese does not have clear character boundaries, this will lead to the loss of important semantic information, which is an important problem for Chinese. In this paper, we propose PTWA: pre-training with word attention for Chinese named entity recognition. PTWA uses multi-head word attention to form a word vector from multiple word vectors, and proposes a word length prediction task to better integrate the word vector into pre-training. With the powerful capabilities of the transformer, PTWA can explicitly make full use of potential word information without adding an external lexicon, and can coexist with pre-trained models that implicitly use word information (such as BERT-WWM, and ERNIE). Experiments conducted on four Chinese NER datasets show that the performance of PTWA is better than other word-word models and Chinese pre-training models.
Kaixin Ma, Tiejun Zhao, Jiyun Zhou
IJCNN1
2021 Dimensions of commonsense knowledge
Filip Ilievski, Alessandro Oltramari, Kaixin Ma, Deborah L. McGuinness, Pedro A. Szekely
Knowl. Based Syst.3
2019 ElderReact: A Multimodal Dataset for Recognizing Emotional Response in Aging Adults
abstract
Automatic emotion recognition plays a critical role in technologies such as intelligent agents and social robots and is increasingly being deployed in applied settings such as education and healthcare. Most research to date has focused on recognizing the emotional expressions of young and middle-aged adults and, to a lesser extent, children and adolescents. Very few studies have examined automatic emotion recognition in older adults (i.e., elders), which represent a large and growing population worldwide. Given that aging causes many changes in facial shape and appearance and has been found to alter patterns of nonverbal behavior, there is strong reason to believe that automatic emotion recognition systems may need to be developed specifically (or augmented) for the elder population. To promote and support this type of research, we introduce a newly collected multimodal dataset of elders reacting to emotion elicitation stimuli. Specifically, it contains 1323 video clips of 46 unique individuals with human annotations of six discrete emotions: anger, disgust, fear, happiness, sadness, and surprise as well as valence. We present a detailed analysis of the most indicative features for each emotion. We also establish several baselines using unimodal and multimodal features on this dataset. Finally, we show that models trained on dataset of another age group do not generalize well on elders.
Kaixin Ma, Xinru Yang, Jeffrey M. Girard, Louis-Philippe Morency
ICMI1
2018 Challenging Reading Comprehension on Daily Conversation: Passage Completion on Multiparty Dialog
abstract
Kaixin Ma, Tomasz Jurczyk, Jinho D. Choi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Kaixin Ma, Tomasz Jurczyk, Jinho D. Choi
NAACL-HLT1