Jiaao Chen

dblp:230/3663 · DBLP profile ↗
← Back
24ranked-venue papers
13as first author
20since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 12 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Decoding Susceptibility: Modeling Misbelief to Misinformation Through a Computational Approach
abstract
Susceptibility to misinformation describes the degree of belief in unverifiable claims, a latent aspect of individuals' mental processes that is not observable.Existing susceptibility studies heavily rely on self-reported beliefs, which can be subject to bias, expensive to collect, and challenging to scale for downstream applications.To address these limitations, in this work, we propose a computational approach to efficiently model users' latent susceptibility levels.As shown in previous work, susceptibility is influenced by various factors (e.g., demographic factors, political ideology), and directly influences people's reposting behavior on social media.To represent the underlying mental process, our susceptibility modeling incorporates these factors as inputs, guided by the supervision of people's sharing behavior.Using COVID-19 as a testbed, our experiments demonstrate a significant alignment between the susceptibility scores estimated by our computational modeling and human judgments, confirming the effectiveness of this latent modeling approach.Furthermore, we apply our model to annotate susceptibility scores on a large-scale dataset and analyze the relationships between susceptibility with various factors.Our analysis reveals that political leanings and other psychological factors exhibit varying degrees of association with susceptibility to COVID-19 misinformation, and shows that susceptibility is unevenly distributed across different professional and geographical backgrounds. 1
Mingyu Derek Ma, Wenna Qin, Azure Zhou, Jiaao Chen, Wei Wang 0010, Diyi Yang
EMNLP5
2024 DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
abstract
Large language models (LLMs) have achieved remarkable performance in various evaluation benchmarks. However, concerns are raised about potential data contamination in their considerable volume of training corpus. Moreover, the static nature and fixed complexity of current benchmarks may inadequately gauge the advancing capabilities of LLMs. In this paper, we introduce DyVal, a general and flexible protocol for dynamic evaluation of LLMs. Based on our framework, we build graph-informed DyVal by leveraging the structural advantage of directed acyclic graphs to dynamically generate evaluation samples with controllable complexities. DyVal generates challenging evaluation sets on reasoning tasks including mathematics, logical reasoning, and algorithm problems. We evaluate various LLMs ranging from Flan-T5-large to GPT-3.5-Turbo and GPT-4. Experiments show that LLMs perform worse in DyVal-generated evaluation samples with different complexities, highlighting the significance of dynamic evaluation. We also analyze the failure cases and results of different prompting methods. Moreover, DyVal-generated samples are not only evaluation sets, but also helpful data for fine-tuning to improve the performance of LLMs on existing benchmarks. We hope that DyVal can shed light on future evaluation research of LLMs. Code is available at: https://github.com/microsoft/promptbench.
Kaijie Zhu, Jiaao Chen, Jindong Wang 0001, Neil Zhenqiang Gong, Diyi Yang, Xing Xie 0001
ICLR2
2024 DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning Graph
abstract
The current paradigm of evaluating Large Language Models (LLMs) through static benchmarks comes with significant limitations, such as vulnerability to data contamination and a lack of adaptability to the evolving capabilities of LLMs. Therefore, evaluation methods that can adapt and generate evaluation data with controlled complexity are urgently needed. In this work, we introduce Dynamic Evaluation of LLMs via Adaptive Reasoning Graph Evolvement (DARG) to dynamically extend current benchmarks with controlled complexity and diversity. Specifically, we first extract the reasoning graphs of data points in current benchmarks and then perturb the reasoning graphs to generate novel testing data. Such newly generated test samples can have different levels of complexity while maintaining linguistic diversity similar to the original benchmarks. We further use a code-augmented LLM to ensure the label correctness of newly generated data. We apply our DARG framework to diverse reasoning tasks in four domains with 15 state-of-the-art LLMs. Experimental results show that almost all LLMs experience a performance decrease with increased complexity and certain LLMs exhibit significant drops. Additionally, we find that LLMs exhibit more biases when being evaluated via the data generated by DARG with higher complexity levels. These observations provide useful insights into how to dynamically and adaptively evaluate LLMs.
Zhehao Zhang 0001, Jiaao Chen, Diyi Yang
NeurIPS2
2024 AoI-Aware Resource Allocation for C-V2X Networks via Multi-Agent Reinforcement Learning with Attention
abstract
This paper delves into the dynamic resource allocation challenges for spectrum sharing in Vehicle-to-Everything (V2X) communication, influenced by the time-varying channel conditions. We concentrate on the dual goals of sub-channel and power allocation in complex V2X environment, underscoring the pivotal role that the Age of Information (AoI) plays in preserving the reliability of safety-critical data across Vehicle-to-Vehicle (V2V) links. To address the complex interplay between minimizing AoI for V2V links and maximizing the overall capacity for vehicle-to-infrastructure (V2I) links, we introduce a novel approach grounded in multi-agent reinforcement learning. This strategy enables adaptive learning in response to V2X rapidly changing channel conditions, with V2V links conceptualized as agents. These agents employ an actor network to select actions and a critic network to evaluate those actions through Q-values, incorporating observations, actions, and individual contributions via attention mechanisms. Our method is further refined by incorporating maximum entropy to enhance action exploration. Through extensive simulations, we demonstrate that our algorithm allows agents to effectively sample and assess the states of their counterparts, leading to optimized decision-making processes.
Jiaao Chen, Qiang Wang 0007
VTC Fall1
2024 Federated Multi-Agent Deep Reinforcement Learning Approach for Resource Allocation in Platoon-Based NR-V2X
abstract
Platoon-based vehicular network in NR-V2X has been considered as a promising technology to assist reducing traffic congestion, saving vehicle fuel, and enhancing driving experience. Resource allocation is the basis for ensuring stable and safety vehicular networks. In this paper, we propose a Distributed Resource Allocation algorithm using Federated Multi agent Deep Reinforcement Learning (DRAFRL), which mathematically utilize the federated averaging (FedAvg) mechanism to reduce the variance between agents and achieve better transmission performance. The proposed algorithm consists of four steps: Firstly, each agent updates local model by deep deterministic policy gradient (DDPG) algorithm. Secondly, the agents upload local model parameters to the base station (BS) for federated aggregation. Thirdly, the BS performs weight aggregation using the FedAvg method and updates the global model. Finally, the BS distributes the optimized global model parameters to each agent. The simulation results show that the proposed algorithm outperforms other baseline algorithms while reducing the variance between agents by 93.5% and 99.1% compared with two baselines.
Qiang Wang 0007, Jiaao Chen, Wenqi Zhang 0002, Chen Sun 0006
VTC Spring3
2024 Can Large Language Models Transform Computational Social Science?
abstract
Abstract Large language models (LLMs) are capable of successfully performing many language processing tasks zero-shot (without training data). If zero-shot LLMs can also reliably classify and explain social phenomena like persuasiveness and political ideology, then LLMs could augment the computational social science (CSS) pipeline in important ways. This work provides a road map for using LLMs as CSS tools. Towards this end, we contribute a set of prompting best practices and an extensive evaluation pipeline to measure the zero-shot performance of 13 language models on 25 representative English CSS benchmarks. On taxonomic labeling tasks (classification), LLMs fail to outperform the best fine-tuned models but still achieve fair levels of agreement with humans. On free-form coding tasks (generation), LLMs produce explanations that often exceed the quality of crowdworkers’ gold references. We conclude that the performance of today’s LLMs can augment the CSS research pipeline in two ways: (1) serving as zero-shot data annotators on human annotation teams, and (2) bootstrapping challenging creative generation tasks (e.g., explaining the underlying attributes of a text). In summary, LLMs are posed to meaningfully participate in social science analysis in partnership with humans.
Caleb Ziems, William Barr Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang 0001, Diyi Yang
Comput. Linguistics4
2023 Compositional Data Augmentation for Abstractive Conversation Summarization
abstract
Recent abstractive conversation summarization systems generally rely on large-scale datasets with annotated summaries.However, collecting and annotating these conversations can be a time-consuming and labor-intensive task.To address this issue, in this work, we present a sub-structure level compositional data augmentation method, COMPO, for generating diverse and high-quality pairs of conversations and summaries.Specifically, COMPO first extracts conversation structures like topic splits and action triples as basic units.Then we organize these semantically meaningful conversation snippets compositionally to create new training instances.Additionally, we explore noise-tolerant settings in both self-training and joint-training paradigms to make the most of these augmented samples.Our experiments on benchmark datasets, SAMSum and Di-alogSum, show that COMPO substantially outperforms prior baseline methods by achieving a nearly 10% increase of ROUGE scores with limited data.We have publically released our code at https://github.com/ ozyyshr/Compo.
Siru Ouyang, Jiaao Chen, Jiawei Han 0001, Diyi Yang
ACL (1)2
2023 Unlearn What You Want to Forget: Efficient Unlearning for LLMs
abstract
Large language models (LLMs) have achieved significant progress from pre-training on and memorizing a wide range of textual data, however, this process might suffer from privacy issues and violations of data protection regulations.As a result, the ability to easily remove data related to individual users from such models while not deteriorating their predictive quality after the removal becomes increasingly important.To address these issues, in this work, we propose an efficient unlearning framework that could efficiently update LLMs without having to retrain the whole model after data removals, by introducing lightweight unlearning layers learned with a selective teacher-student objective into the transformers.In addition, we introduce a fusion mechanism to effectively combine different unlearning layers that learns to forget different sets of data to handle a sequence of forgetting operations.Experiments on classification and generation tasks demonstrate the effectiveness of our proposed methods compared to the state-of-the-art baselines 1 .
Jiaao Chen, Diyi Yang
EMNLP1
2023 A Cheaper and Better Diffusion Language Model with Soft-Masked Noise
abstract
Diffusion models that are based on iterative denoising have been recently proposed and leveraged in various generation tasks like image generation.Whereas, as a way inherently built for continuous data, existing diffusion models still have some limitations in modeling discrete data, e.g., languages.For example, the generally used Gaussian noise can not handle the discrete corruption well, and the objectives in continuous spaces fail to be stable for textual data in the diffusion process especially when the dimension is high.To alleviate these issues, we introduce a novel diffusion model for language modeling, Masked-Diffusion LM, with lower training cost and better performances, inspired by linguistic features in languages.Specifically, we design a linguistic-informed forward process which adds corruptions to the text through strategically soft-masking to better noise the textual data.Also, we directly predict the categorical distribution with cross-entropy loss function in every diffusion step to connect the continuous space and discrete space in a more efficient and straightforward way.Through experiments on 5 controlled generation tasks, we demonstrate that our Masked-Diffusion LM can achieve better generation quality than the state-of-the-art diffusion models with better efficiency.Code is available at https://github. com/SALT-NLP/Masked_Diffusioin_LM.
Jiaao Chen, Aston Zhang, Mu Li 0003, Alexander J. Smola, Diyi Yang
EMNLP1
2023 Is ChatGPT a General-Purpose Natural Language Processing Task Solver?
abstract
Spurred by advancements in scale, large language models (LLMs) have demonstrated the ability to perform a variety of natural language processing (NLP) tasks zero-shot-i.e., without adaptation on downstream data.Recently, the debut of ChatGPT 1 has drawn a great deal of attention from the natural language processing (NLP) community due to the fact that it can generate high-quality responses to human input and self-correct previous mistakes based on subsequent conversations.However, it is not yet known whether ChatGPT can serve as a generalist model that can perform many NLP tasks zero-shot.In this work, we empirically analyze the zero-shot learning ability of ChatGPT by evaluating it on 20 popular NLP datasets covering 7 representative task categories.With extensive empirical studies, we demonstrate both the effectiveness and limitations of the current version of ChatGPT.We find that ChatGPT performs well on many tasks favoring reasoning capabilities (e.g., arithmetic reasoning) while it still faces challenges when solving specific tasks such as sequence tagging.We additionally provide in-depth analysis through qualitative case studies.
Chengwei Qin, Aston Zhang, Zhuosheng Zhang 0001, Jiaao Chen, Michihiro Yasunaga, Diyi Yang
EMNLP4
2023 Parameter-Efficient Fine-Tuning Design Spaces
Jiaao Chen, Aston Zhang, Xingjian Shi, Mu Li 0003, Alexander J. Smola, Diyi Yang
ICLR1
2023 Where Does Your News Come From? Predicting Information Pathways in Social Media
abstract
As social networks become further entrenched in modern society, it becomes increasingly important to understand and predict how information (e.g., news coverage of a given event) is propagated across social media (i.e., information pathway), which helps the understandings of the impact of real-world information. Thus, in this paper, we propose a novel task, Information Pathway Prediction (IPP), which depicts the propagation paths of a given passage as a community tree (rooted at the information source) on constructed community interaction graphs where we first aggregate individual users into communities formed around news sources and influential users, and then elucidate the patterns of information dissemination across media based on such community nodes. We argue that this is an important and useful task because, on one hand, community-level interactions offer more stability than those at the user level; on the other hand, individual users are often influenced by their community, and modeling community-level information propagation will help the traditional link-prediction problem. To tackle the IPP task, we introduce Lightning, a novel content-aware link prediction GNN model and demonstrate using a large Twitter dataset consisting of all COVID related tweets that Lightning outperforms state-of-the-art link prediction baselines by a significant margin.
Alexander K. Taylor 0002, Nuan Wen, Po-Nien Kung, Jiaao Chen, Violet Peng, Wei Wang 0010
SIGIR4
2023 An Empirical Survey of Data Augmentation for Limited Data Learning in NLP
abstract
Abstract NLP has achieved great progress in the past decade through the use of neural models and large labeled datasets. The dependence on abundant data prevents NLP models from being applied to low-resource settings or novel tasks where significant time, money, or expertise is required to label massive amounts of textual data. Recently, data augmentation methods have been explored as a means of improving data efficiency in NLP. To date, there has been no systematic empirical overview of data augmentation for NLP in the limited labeled data setting, making it difficult to understand which methods work in which settings. In this paper, we provide an empirical survey of recent progress on data augmentation for NLP in the limited labeled data setting, summarizing the landscape of methods (including token-level augmentations, sentence-level augmentations, adversarial augmentations, and hidden-space augmentations) and carrying out experiments on 11 datasets covering topics/news classification, inference tasks, paraphrasing tasks, and single-sentence tasks. Based on the results, we draw several conclusions to help practitioners choose appropriate augmentations in different settings and discuss the current challenges and future directions for limited data learning in NLP.
Jiaao Chen, Derek Tam, Colin Raffel, Mohit Bansal, Diyi Yang
Trans. Assoc. Comput. Linguistics1
2022 VALUE: Understanding Dialect Disparity in NLU
abstract
English Natural Language Understanding (NLU) systems have achieved great performances and even outperformed humans on benchmarks like GLUE and SuperGLUE.However, these benchmarks contain only textbook Standard American English (SAE).Other dialects have been largely overlooked in the NLP community.This leads to biased and inequitable NLU systems that serve only a sub-population of speakers.To understand disparities in current models and to facilitate more dialect-competent NLU systems, we introduce the VernAcular Language Understanding Evaluation (VALUE) benchmark, a challenging variant of GLUE that we created with a set of lexical and morphosyntactic transformation rules.In this initial release (V.1), we construct rules for 11 features of African American Vernacular English (AAVE), and we recruit fluent AAVE speakers to validate each feature transformation via linguistic acceptability judgments in a participatory design manner.Experiments show that these new dialectal features can lead to a drop in model performance.
Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, Diyi Yang
ACL (1)2
2022 When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain
abstract
Raj Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah, Wendi Du, Sudheer Chava, Natraj Raman, Charese Smiley, Jiaao Chen, Diyi Yang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Raj Sanjay Shah, Kunal Chawla, Dheeraj Eidnani, Agam Shah, Wendi Du, Sudheer Chava, Natraj Raman, Charese Smiley, Jiaao Chen, Diyi Yang
EMNLP9
2021 Weakly-Supervised Hierarchical Models for Predicting Persuasive Strategies in Good-faith Textual Requests
abstract
Modeling persuasive language has the potential to better facilitate our decision-making processes. Despite its importance, computational modeling of persuasion is still in its infancy, largely due to the lack of benchmark datasets that can provide quantitative labels of persuasive strategies to expedite this line of research. To this end, we introduce a large-scale multi-domain text corpus for modeling persuasive strategies in good-faith text requests. Moreover, we design a hierarchical weakly-supervised latent variable model that can leverage partially labeled data to predict such associated persuasive strategies for each sentence, where the supervision comes from both the overall document-level labels and very limited sentence-level labels. Experimental results showed that our proposed method outperformed existing semi-supervised baselines significantly. We have publicly released our code at https://github.com/GT-SALT/Persuasion_Strategy_WVAE.
Jiaao Chen, Diyi Yang
AAAI1
2021 HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better Generalizability
abstract
Jiaao Chen, Dinghan Shen, Weizhu Chen, Diyi Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jiaao Chen, Dinghan Shen, Weizhu Chen, Diyi Yang
ACL/IJCNLP (1)1
2021 Simple Conversational Data Augmentation for Semi-supervised Abstractive Dialogue Summarization
abstract
Abstractive conversation summarization has received growing attention while most current state-of-the-art summarization models heavily rely on human-annotated summaries.To reduce the dependence on labeled summaries, in this work, we present a simple yet effective set of Conversational Data Augmentation (CODA) methods for semisupervised abstractive conversation summarization, such as random swapping/deletion to perturb the discourse relations inside conversations, dialogue-acts-guided insertion to interrupt the development of conversations, and conditional-generation-based substitution to substitute utterances with their paraphrases generated based on the conversation context.To further utilize unlabeled conversations, we combine CODA with two-stage noisy selftraining where we first pre-train the summarization model on unlabeled conversations with pseudo summaries and then fine-tune it on labeled conversations.Experiments conducted on the recent conversation summarization datasets demonstrate the effectiveness of our methods over several state-of-the-art data augmentation baselines.
Jiaao Chen, Diyi Yang
EMNLP (1)1
2021 Structure-Aware Abstractive Conversation Summarization via Discourse and Action Graphs
abstract
Abstractive conversation summarization has received much attention recently.However, these generated summaries often suffer from insufficient, redundant, or incorrect content, largely due to the unstructured and complex characteristics of human-human interactions.To this end, we propose to explicitly model the rich structures in conversations for more precise and accurate conversation summarization, by first incorporating discourse relations between utterances and action triples ("WHO-DOING-WHAT") in utterances through structured graphs to better encode conversations, and then designing a multi-granularity decoder to generate summaries by combining all levels of information.Experiments show that our proposed models outperform state-of-theart methods and generalize well in other domains in terms of both automatic evaluations and human judgments.We have publicly released our code at https://github.com/ GT-SALT/Structure-Aware-BART.
Jiaao Chen, Diyi Yang
NAACL-HLT1
2021 Continual Learning for Text Classification with Information Disentanglement Based Regularization
abstract
Yufan Huang, Yanzhe Zhang, Jiaao Chen, Xuezhi Wang, Diyi Yang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yufan Huang, Jiaao Chen, Xuezhi Wang 0002, Diyi Yang
NAACL-HLT3
2020 MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification
abstract
This paper presents MixText, a semisupervised learning method for text classification, which uses our newly designed data augmentation method called TMix.TMix creates a large amount of augmented training samples by interpolating text in hidden space.Moreover, we leverage recent advances in data augmentation to guess low-entropy labels for unlabeled data, hence making them as easy to use as labeled data.By mixing labeled, unlabeled and augmented data, MixText significantly outperformed current pre-trained and fined-tuned models and other state-ofthe-art semi-supervised learning methods on several text classification benchmarks.The improvement is especially prominent when supervision is extremely limited.
Jiaao Chen, Diyi Yang
ACL1
2020 Local Additivity Based Data Augmentation for Semi-supervised NER
abstract
Named Entity Recognition (NER) is one of the first stages in deep language understanding yet current NER models heavily rely on humanannotated data.In this work, to alleviate the dependence on labeled data, we propose a Local Additivity based Data Augmentation (LADA) method for semi-supervised NER, in which we create virtual samples by interpolating sequences close to each other.Our approach has two variations: Intra-LADA and Inter-LADA, where Intra-LADA performs interpolations among tokens within one sentence, and Inter-LADA samples different sentences to interpolate.Through linear additions between sampled training data, LADA creates an infinite amount of labeled data and improves both entity and context learning.We further extend LADA to the semi-supervised setting by designing a novel consistency loss for unlabeled data.Experiments conducted on two NER benchmarks demonstrate the effectiveness of our methods over several strong baselines.We have publicly released our code at
Jiaao Chen, Zhenghui Wang, Diyi Yang
EMNLP (1)1
2020 Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue Summarization
abstract
Text summarization is one of the most challenging and interesting problems in NLP.Although much attention has been paid to summarizing structured text like news reports or encyclopedia articles, summarizing conversations-an essential part of humanhuman/machine interaction where most important pieces of information are scattered across various utterances of different speakersremains relatively under-investigated.This work proposes a multi-view sequence-tosequence model by first extracting conversational structures of unstructured daily chats from different views to represent conversations and then utilizing a multi-view decoder to incorporate different views to generate dialogue summaries.Experiments on a large-scale dialogue summarization corpus demonstrated that our methods significantly outperformed previous state-of-the-art models via both automatic evaluations and human judgment.We also discussed specific challenges that current approaches faced with this task.We have publicly released our code at https://github.com/GT-SALT/ Multi-View-Seq2Seq. ConversationTopic View Stage View James: Hey!I have
Jiaao Chen, Diyi Yang
EMNLP (1)1
2019 Incorporating Structured Commonsense Knowledge in Story Completion
abstract
The ability to select an appropriate story ending is the first step towards perfect narrative comprehension. Story ending prediction requires not only the explicit clues within the context, but also the implicit knowledge (such as commonsense) to construct a reasonable and consistent story. However, most previous approaches do not explicitly use background commonsense knowledge. We present a neural story ending selection model that integrates three types of information: narrative sequence, sentiment evolution and commonsense knowledge. Experiments show that our model outperforms state-ofthe-art approaches on a public dataset, ROCStory Cloze Task (Mostafazadeh et al. 2017), and the performance gain from adding the additional commonsense knowledge is significant.
Jiaao Chen, Jianshu Chen
AAAI1