Qingyang Wu

dblp:242/8977 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Mamba-Enhanced Text-Audio-Video Alignment Network for Emotion Recognition in Conversations
Xiaomao Fan, Qingyang Wu, Xiaojiang Peng, Ye Li 0002
ADMA (3)3
2025 CausalEval: Towards Better Causal Reasoning in Language Models
abstract
Longxuan Yu, Delin Chen, Siheng Xiong, Qingyang Wu, Dawei Li, Zhikai Chen, Xiaoze Liu, Liangming Pan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Longxuan Yu, Delin Chen, Siheng Xiong, Qingyang Wu, Xiaoze Liu, Liangming Pan
NAACL (Long Papers)4
2025 Enhancing domain adaptation for plant diseases detection through Masked Image Consistency in Multi-Granularity Alignment
Guinan Guo, Songning Lai, Qingyang Wu, Yuntao Shou, Wenxu Shi
Expert Syst. Appl.3
2024 LIONs: An Empirically Optimized Approach to Align Language Models
abstract
Alignment is a crucial step to enhance the instruction-following and conversational abilities of language models.Despite many recent work proposing new algorithms, datasets, and training pipelines, there is a lack of comprehensive studies measuring the impact of various design choices throughout the whole training process.We first conduct a rigorous analysis over a three-stage training pipeline consisting of supervised fine-tuning, offline preference learning, and online preference learning.We have found that using techniques like sequence packing, loss masking in SFT, increasing the preference dataset size in DPO, and online DPO training can significantly improve the performance of language models.We then train from Gemma-2b-base and LLama-3-8b-base, and find that our best models exceed the performance of the official instruct models tuned with closed-source data and algorithms.Our code and models can be found at https://github.
Xiao Yu 0011, Qingyang Wu, Yu Li 0013, Zhou Yu 0005
EMNLP2
2024 DECOR: Improving Coherence in L2 English Writing with a Novel Benchmark for Incoherence Detection, Reasoning, and Rewriting
abstract
Coherence in writing, an aspect that secondlanguage (L2) English learners often struggle with, is crucial in assessing L2 English writing.Existing automated writing evaluation systems primarily use basic surface linguistic features to detect coherence in writing.However, little effort has been made to correct the detected incoherence, which could significantly benefit L2 language learners seeking to improve their writing.To bridge this gap, we introduce DECOR, a novel benchmark that includes expert annotations for detecting incoherence in L2 English writing, identifying the underlying reasons, and rewriting the incoherent sentences.To our knowledge, DECOR is the first coherence assessment dataset specifically designed for improving L2 English writing, featuring pairs of original incoherent sentences alongside their expert-rewritten counterparts.Additionally, we fine-tuned models to automatically detect and rewrite incoherence in student essays.We find that incorporating specific reasons for incoherence during fine-tuning consistently improves the quality of the rewrites, achieving a result that is favored in both automatic and human evaluations.1
Xuanming Zhang, Anthony Diaz, Zixun Chen, Qingyang Wu, Kun Qian 0016, Erik Voss, Zhou Yu 0005
EMNLP4
2024 kNN-ICL: Compositional Task-Oriented Parsing Generalization with Nearest Neighbor In-Context Learning
abstract
Wenting Zhao, Ye Liu, Yao Wan, Yibo Wang, Qingyang Wu, Zhongfen Deng, Jiangshu Du, Shuaiqi Liu, Yunlong Xu, Philip Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Wenting Zhao 0006, Ye Liu 0006, Yao Wan 0001, Yibo Wang 0001, Qingyang Wu, Zhongfen Deng, Jiangshu Du, Shuaiqi Liu 0002, Philip S. Yu
NAACL-HLT5
2023 Large GAN Is All You Need
Kai Liu 0034, Qingyang Wu, Mengkun Xie
CGI (1)2
2023 GLIGEN: Open-Set Grounded Text-to-Image Generation
abstract
Large-scale text-to-image diffusion models have made amazing advances. However, the status quo is to use text input alone, which can impede controllability. In this work, we propose GLIGEN, Grounded-Language-to-Image Generation, a novel approach that builds upon and extends the functionality of existing pre-trained text-to-image diffusion models by enabling them to also be conditioned on grounding inputs. To preserve the vast concept knowledge of the pre-trained model, we freeze all of its weights and inject the grounding information into new trainable layers via a gated mechanism. Our model achieves open-world grounded text2img generation with caption and bounding box condition inputs, and the grounding ability generalizes well to novel spatial configurations and concepts. GLIGEN's zero-shot performance on COCO and LVIS outperforms existing supervised layout-to-image baselines by a large margin.
Qingyang Wu, Fangzhou Mu, Jianfeng Gao 0001, Chunyuan Li, Yong Jae Lee
CVPR3
2023 KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords Learning
abstract
In task-oriented dialogs (TOD), reinforcement learning (RL) algorithms train a model to directly optimize response for task-related metrics.However, RL needs to perform exploration, which can be time-consuming due to the slow auto-regressive sequence generation process.We investigate an approach to create a more efficient RL-based algorithm to improve TOD performance in an offline setting.First, we use a faster generation procedure that samples from independent next-word distributions after training the language model (LM) with supervised learning.We then introduce a finegrained reward function to help the model focus on learning key information in a dialog, by measuring the importance and semantic closeness of each generated token.Experiments on the MultiWoZ dataset show our new training algorithm, Keywords Reinforcement Learning with Next-word Sampling (KRLS), achieves state-of-the-art performance on the end-to-end response generation task, with a 15% training time reduction compared to a standard RL algorithm using auto-regressive generation 1 .
Xiao Yu 0011, Qingyang Wu, Kun Qian 0016, Zhou Yu 0005
EMNLP2
2023 ARNOLD: A Benchmark for Language-Grounded Task Learning With Continuous States in Realistic 3D Scenes
abstract
Understanding the continuous states of objects is essential for task learning and planning in the real world. However, most existing task learning benchmarks assume discrete (e.g., binary) object goal states, which poses challenges for the learning of complex tasks and transferring learned policy from simulated environments to the real world. Furthermore, state discretization limits a robot’s ability to follow human instructions based on the grounding of actions and states. To tackle these challenges, we present ARNOLD, a benchmark that evaluates language-grounded task learning with continuous states in realistic 3D scenes. ARNOLD is comprised of 8 language-conditioned tasks that involve understanding object states and learning policies for continuous goals. To promote language-instructed learning, we provide expert demonstrations with template-generated language descriptions. We assess task performance by utilizing the latest language-conditioned policy learning models. Our results indicate that current models for language-conditioned manipulations continue to experience significant challenges in novel goal-state generalizations, scene generalizations, and object generalizations. These findings highlight the need to develop new algorithms that address this gap and underscore the potential for further research in this area. Project website: https://arnold-benchmark.github.io.
Jiangyong Huang, Xiaofeng Gao 0002, Qingyang Wu, Wensi Ai, Demetri Terzopoulos, Song-Chun Zhu, Baoxiong Jia, Siyuan Huang 0001
ICCV6
2023 Visual Instruction Tuning
abstract
Instruction tuning large language models (LLMs) using machine-generated instruction-following data has been shown to improve zero-shot capabilities on new tasks, but the idea is less explored in the multimodal field. We present the first attempt to use language-only GPT-4 to generate multimodal language-image instruction-following data. By instruction tuning on such generated data, we introduce LLaVA: Large Language and Vision Assistant, an end-to-end trained large multimodal model that connects a vision encoder and an LLM for general-purpose visual and language understanding. To facilitate future research on visual instruction following, we construct two evaluation benchmarks with diverse and challenging application-oriented tasks. Our experiments show that LLaVA demonstrates impressive multimodal chat abilities, sometimes exhibiting the behaviors of multimodal GPT-4 on unseen images/instructions, and yields a 85.1% relative score compared with GPT-4 on a synthetic multimodal instruction-following dataset. When fine-tuned on Science QA, the synergy of LLaVA and GPT-4 achieves a new state-of-the-art accuracy of 92.53%. We make GPT-4 generated visual instruction tuning data, our model, and code publicly available.
Chunyuan Li, Qingyang Wu, Yong Jae Lee
NeurIPS3
2023 DiactTOD: Learning Generalizable Latent Dialogue Acts for Controllable Task-Oriented Dialogue Systems
abstract
Dialogue act annotations are important to improve response generation quality in taskoriented dialogue systems.However, it can be challenging to use dialogue acts to control response generation in a generalizable way because different datasets and tasks may have incompatible annotations.While alternative methods that utilize latent action spaces or reinforcement learning do not require explicit annotations, they may lack interpretability or face difficulties defining task-specific rewards.In this work, we present a novel end-to-end latent dialogue act model (DiactTOD) that represents dialogue acts in a latent space.Diact-TOD, when pre-trained on a large corpus, is able to predict and control dialogue acts to generate controllable responses using these latent representations in a zero-shot fashion.Our approach demonstrates state-of-the-art performance across a wide range of experimental settings on the MultiWOZ dataset, including zeroshot, few-shot, and full data fine-tuning with both end-to-end and policy optimization configurations.
Qingyang Wu, James Gung, Raphael Shu
SIGDIAL1
2022 DG2: Data Augmentation Through Document Grounded Dialogue Generation
abstract
Collecting data for training dialog systems can be extremely expensive due to the involvement of human participants and the need for extensive annotation.Especially in documentgrounded dialog systems, human experts need to carefully read the unstructured documents to answer the users' questions.As a result, existing document-grounded dialog datasets are relatively small-scale and obstruct the effective training of dialogue systems.In this paper, we propose an automatic data augmentation technique grounded on documents through a generative dialogue model.The dialogue model consists of a user bot and agent bot that can synthesize diverse dialogues given an input document, which are then used to train a downstream model.When supplementing the original dataset, our method achieves significant improvement over traditional data augmentation methods.We also achieve competitive performance in the low-resource setting.
Qingyang Wu, Song Feng 0002, Derek Chen, Sachindra Joshi, Luis A. Lastras
SIGDIAL1
2021 Perception Score: A Learned Metric for Open-ended Text Generation Evaluation
abstract
Automatic evaluation for open-ended natural language generation tasks remains a challenge. We propose a learned evaluation metric: Perception Score. It utilizes a pre-trained model and considers context information for conditional generation. Perception Score assigns a holistic score along with the uncertainty measurement. We conduct experiments on three open-ended conditional generation tasks and two open-ended unconditional generation tasks. Perception Score achieves state-of-the-art results on all the tasks consistently in terms of correlation with human evaluation scores.
Qingyang Wu
AAAI2
2021 TextGAIL: Generative Adversarial Imitation Learning for Text Generation
abstract
Generative Adversarial Networks (GANs) for text generation have recently received many criticisms, as they perform worse than their MLE counterparts. We suspect previous text GANs' inferior performance is due to the lack of a reliable guiding signal in their discriminators. To address this problem, we propose a generative adversarial imitation learning framework for text generation that uses large pre-trained language models to provide more reliable reward guidance. As previous text GANs suffer from high variance of gradients, we apply contrastive discriminator, and proximal policy optimization (PPO) to stabilize and improve text generation performance. For evaluation, we conduct experiments on a diverse set of unconditional and conditional text generation tasks. Experimental results show that TextGAIL achieves better performance in terms of both quality and diversity than the MLE baseline. We also validate our intuition that TextGAIL's discriminator demonstrates the capability of providing reasonable rewards with an additional task.
Qingyang Wu, Lei Li 0005
AAAI1
2021 Alternating Recurrent Dialog Model with Large-scale Pre-trained Language Models
abstract
Existing dialog system models require extensive human annotations and are difficult to generalize to different tasks.The recent success of large pre-trained language models has suggested the effectiveness of incorporating language priors in down-stream NLP tasks.However, how much pre-trained language models can help dialog response generation is still under exploration.In this paper, we propose a simple, general, and effective framework: Alternating Recurrent Dialog Model (ARDM) 1 .ARDM models each speaker separately and takes advantage of large pre-trained language models.It requires no supervision from human annotations such as belief states or dialog acts to achieve effective conversations.ARDM outperforms or is on par with the state-of-theart methods on two popular task-oriented dialog datasets: CamRest676 and MultiWOZ.Moreover, we can generalize ARDM to more challenging, non-collaborative tasks such as persuasion.In the PersuasionForGood task, ARDM is capable of generating human-like responses to persuade people to donate to a charity.
Qingyang Wu, Yichi Zhang 0001, Yu Li 0013, Zhou Yu 0005
EACL1
2020 Importance-Aware Learning for Neural Headline Editing
abstract
Many social media news writers are not professionally trained. Therefore, social media platforms have to hire professional editors to adjust amateur headlines to attract more readers. We propose to automate this headline editing process through neural network models to provide more immediate writing support for these social media news writers. To train such a neural headline editing model, we collected a dataset which contains articles with original headlines and professionally edited headlines. However, it is expensive to collect a large number of professionally edited headlines. To solve this low-resource problem, we design an encoder-decoder model which leverages large scale pre-trained language models. We further improve the pre-trained model's quality by introducing a headline generation task as an intermediate task before the headline editing task. Also, we propose Self Importance-Aware (SIA) loss to address the different levels of editing in the dataset by down-weighting the importance of easily classified tokens and sentences. With the help of Pre-training, Adaptation, and SIA, the model learns to generate headlines in the professional editor's style. Experimental results show that our method significantly improves the quality of headline editing comparing against previous methods.
Qingyang Wu, Lei Li 0005, Hao Zhou 0012
AAAI1