Wenqing Chen

dblp:31/2740 · DBLP profile ↗
← Back
37ranked-venue papers
7as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 5 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding
abstract
Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). However, whether these models can achieve human-level auditory perception, particularly in terms of their ability to comprehend latent intentions and implicit emotions in real-world spoken language, remains underexplored. To this end, we introduce the Human-level Perception in Spoken Speech Understanding (HPSU), a new benchmark for fully evaluating the human-level perceptual and understanding capabilities of Speech LLMs. HPSU comprises over 20,000 expert-validated spoken language understanding samples in English and Chinese. It establishes a comprehensive evaluation framework by encompassing a spectrum of tasks, ranging from basic speaker attribute recognition to complex inference of latent intentions and implicit emotions. To address the issues of data scarcity and high cost of manual annotation in real-world scenarios, we developed a semi-automatic annotation process. This process fuses audio, textual, and visual information to enable precise speech understanding and labeling, thus enhancing both annotation efficiency and quality. We systematically evaluate various open-source and proprietary Speech LLMs. The results demonstrate that even top-performing models still fall considerably short of human capabilities in understanding genuine spoken interactions. Consequently, HPSU will be useful for guiding the development of Speech LLMs toward human-level perception and cognition.
Peiji Yang, Yicheng Zhong, Jianxing Yu, Zhisheng Wang 0001, Zihao Gou, Wenqing Chen, Jian Yin 0001
AAAI7
2026 Functional consistency of LLM code embeddings: A self-evolving data synthesis framework for benchmarking
Zhuohao Li, Wenqing Chen, Jianxing Yu, Zhichao Lu
Expert Syst. Appl.2
2026 Accelerated Optimization of Large Mixture-of-Experts Models by Density-Aware Multi-Stage Learning
abstract
This article aims to speed up the training of large neural networks with the Mixture-of-Experts (MoE) structure. Training MoE often needs a lot of computing resources due to its large scale. Traditional acceleration methods either degrade prediction performance or rely on dedicated hardware with additional resources, but the resources are usually limited in real applications.One solution is to resort to new optimization strategies, such as learning from easy to hard by multiple stages. However, existing strategies are designed mainly for networks with a serial structure, but MoE has multiple expert networks working in parallel. They employ an identical learning plan for all experts, ignoring that each expert's learning domain and speed differ, resulting in some experts being over-learned while others being under-learned. This mismatch will make it hard for experts to train together, harming training efficiency. To address this problem, we propose a new training acceleration framework. It can customize an effective learning plan for each expert by considering their training progress, avoiding blindly searching in a huge parameter space. In detail, we first design a multi-stage planner that starts with optimizing a subpart of the network and then scales it up to retrain until it expands to an entire network. It uses the density function to assess the knowledge gained by the expert in each stage, giving priority to the experts who learn faster to increase the training scale, so as to boost convergence. Afterward, we exploit the growth operator to add the expert training scale of the next stage. In each stage, the network would converge to some locally optimal values. That can provide a better initialization to train the next stage more easily, since the time and data required for training from scratch are greatly reduced. To alleviate the gradient vanishing problem caused by network growth, we develop a scheduler to dynamically adjust the learning rate. Extensive experiments are conducted to validate the effectiveness of our method. The results show that we can obtain more than 25% training acceleration on average.
Jianxing Yu, Haowei Jiang, Huaijie Zhu, Wenqing Chen, Yanghui Rao, Qinliang Su, Jian Yin 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 A Federated Adaptive Large Language Model Fine-Tuning Framework for Software Development
abstract
Large Language Models (LLMs) have achieved remarkable progress in code intelligence tasks, significantly en hancing the efficiency of software development. However, several challenges remain. First, fine-tuning LLMs for specific tasks requires a large amount of task-specific labeled data, which is often costly and time-consuming to acquire. Second, due to the sensitivity of code data, high-quality internal datasets from different organizations cannot be directly shared or combined for fine tuning. Moreover, variations in programming languages across organizations can introduce interference during the fine-tuning process. To address these challenges, we propose F-CodeLLM, a federated adaptive large language model fine-tuning framework designed for real-world software development scenarios. To the best of our knowledge, this is the first approach to apply federated learning to the fine-tuning of code LLMs, enabling collaborative model optimization while preserving the privacy of each organization's code data. We design an efficient LLM fine tuning method to mitigate the computational and communication overhead associated with collaborative fine-tuning. Experimental results demonstrate that F-CodeLLM effectively allows LLMs to learn from each organization's dataset, achieving performance comparable to centralized fine-tuning. Furthermore, F-CodeLLM is well-suited for multilingual data environments, as it can lever age shared knowledge across programming languages to enhance performance within individual language domains. Our code is publicly available at: https://github.com/AAnony/F-CodeLLM.
Jianguo Chen 0001, Zeju Cai, Wenqing Chen, Zibin Zheng, Philip S. Yu
IEEE Trans. Serv. Comput.3
2025 Mitigating Social Bias in Large Language Models: A Multi-Objective Approach Within a Multi-Agent Framework
abstract
Natural language processing (NLP) has seen remarkable advancements with the development of large language models (LLMs). Despite these advancements, LLMs often produce socially biased outputs. Recent studies have mainly addressed this problem by prompting LLMs to behave ethically, but this approach results in unacceptable performance degradation. In this paper, we propose a multi-objective approach within a multi-agent framework (MOMA) to mitigate social bias in LLMs without significantly compromising their performance. The key idea of MOMA involves deploying multiple agents to perform causal interventions on bias-related contents of the input questions, breaking the shortcut connection between these contents and the corresponding answers. Unlike traditional debiasing techniques leading to performance degradation, MOMA substantially reduces bias while maintaining accuracy in downstream tasks. Our experiments conducted in two datasets and two models demonstrate that MOMA reduces bias scores by up to 87.7%, with only a marginal performance degradation of up to 6.8% in the BBQ dataset. Additionally, it significantly enhances the multi-objective metric icat in the StereoSet dataset by up to 58.1%.
Zhenjie Xu, Wenqing Chen, Xuanying Li, Zhixuan Chu, Kui Ren 0001, Zibin Zheng, Zhichao Lu
AAAI2
2025 Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving
abstract
World models envision potential future states based on various ego actions. They embed extensive knowledge about the driving environment, facilitating safe and scalable autonomous driving. Most existing methods primarily focus on either data generation or the pretraining paradigms of world models. Unlike the aforementioned prior works, we propose Drive-OccWorld, which adapts a vision-centric 4D forecasting world model to end-to-end planning for autonomous driving. Specifically, we first introduce a semantic and motion-conditional normalization in the memory module, which accumulates semantic and dynamic information from historical BEV embeddings. These BEV features are then conveyed to the world decoder for future occupancy and flow forecasting, considering both geometry and spatiotemporal modeling. Additionally, we propose injecting flexible action conditions, such as velocity, steering angle, trajectory, and commands, into the world model to enable controllable generation and facilitate a broader range of downstream applications. Furthermore, we explore integrating the generative capabilities of the 4D world model with end-to-end planning, enabling continuous forecasting of future states and the selection of optimal trajectories using an occupancy-based cost function. Extensive experiments on the nuScenes dataset demonstrate that our method can generate plausible and controllable 4D occupancy, opening new avenues for driving world generation and end-to-end planning.
Yu Yang 0001, Jianbiao Mei, Yukai Ma, Siliang Du, Wenqing Chen, Yijie Qian, Yong Liu 0007
AAAI5
2025 Answering Complex Geographic Questions by Adaptive Reasoning with Visual Context and External Commonsense Knowledge
abstract
Fan Li, Jianxing Yu, Jielong Tang, Wenqing Chen, Hanjiang Lai, Yanghui Rao, Jian Yin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jianxing Yu, Jielong Tang, Wenqing Chen, Hanjiang Lai, Yanghui Rao, Jian Yin 0001
ACL (1)4
2025 Generating Commonsense Reasoning Questions with Controllable Complexity through Multi-step Structural Composition
abstract
This paper studies the task of generating commonsense reasoning questions (QG) with desired difficulty levels. Compared to traditional shallow questions that can be solved by simple term matching, ours are more challenging. Our answering process requires reasoning over multiple contextual and commonsense clues. That involves advanced comprehension skills, such as abstract semantics learning and missing knowledge inference. Existing work mostly learns to map the given text into questions, lacking a mechanism to control results with the desired complexity. To address this problem, we propose a novel controllable framework. We first derive contextual and commonsense clues involved in reasoning questions from the text. These clues are used to create simple sub-questions. We then aggregate multiple sub-questions to compose complex ones under the guidance of prior reasoning structures. By iterating this process, we can compose a complex QG task based on a series of smaller and simpler QG subtasks. Each subtask serves as a building block for a larger one. Each composition corresponds to an increase in the reasoning step. Moreover, we design a voting verifier to ensure results’ validity from multiple views, including answer consistency, reasoning difficulty, and context correlation. Finally, we can learn the optimal QG model to yield thought-provoking results. Evaluations on two typical datasets validate our method.
Jianxing Yu, Shiqi Wang 0016, Hanjiang Lai, Wenqing Chen, Yanghui Rao, Qinliang Su, Jian Yin 0001
COLING4
2025 Asking Diversified Reasonable Questions with External Commonsense Knowledge to Infer Inconsistency for Multi-modal Clickbait Detection
Jianxing Yu, Shiqi Wang 0016, Huaijie Zhu, Libin Zheng 0001, Wenqing Chen, Jian Yin 0001
DASFAA (2)6
2025 Emotion-Based Conversational Recommendation by Inferring Implicit Users' Preferences from Their Subjective Claims
Xuanming Zhang, Yonghe Lu, Jianxing Yu, Huaijie Zhu, Wei Liu 0061, Wenqing Chen, Jian Yin 0001
DASFAA (5)6
2025 Eliciting Implicit Acoustic Styles from Open-domain Instructions to Facilitate Fine-grained Controllable Generation of Speech
abstract
This paper focuses on generating speech with the acoustic style that meets users' needs based on their open-domain instructions.To control the style, early work mostly relies on predefined rules or templates.The control types and formats are fixed in a closed domain, making it hard to meet users' diverse needs.One solution is to resort to instructions in free text to guide the generation.Current work mainly studies the instructions that clearly specify the acoustic styles, such as low pitch and fast speed.However, the instructions are complex, some even vague and abstract, such as "Generate a voice of a woman who is heartbroken due to a breakup."It is hard to infer this implicit style by traditional matching-based methods.To address this problem, we propose a new controllable model.It first utilizes multimodal LLMs with knowledge-augmented techniques to infer the desired speech style from the instructions.The powerful language understanding ability of LLMs can help us better elicit the implicit style factors from the instruction.By using these factors as a control condition, we design a diffusion-based generator adept at finely adjusting speech details, enabling higher flexibility to meet complex users' needs.Next, we verify the output speech from three aspects, i.e., consistency of decoding state, mel-spectrogram, and instruction style.This verified feedback can inversely optimize the generator.Extensive experiments are conducted on three popular datasets.The results show the effectiveness and good controllability of our approach.
Jianxing Yu, Zihao Gou, Zhisheng Wang 0001, Peiji Yang, Wenqing Chen, Jian Yin 0001
EMNLP6
2025 Detecting Violations of Physical Common Sense in Images: A Challenge Dataset and Effective Model
abstract
Vision-language models (VLMs) have achieved remarkable success in various vision-language tasks, such as image captioning and visual question answering. However, these models often lack physical common sense, frequently failing to identify visually evident violations of common physical principles. Therefore, evaluating the VLMs' understanding of physical common sense is essential, which has not yet been systematically explored in existing research. To fill this gap, we introduce PhyVIB (Physical Common Sense Violation Image Benchmark). This novel benchmark consists of 16,000 images across eight categories, aiming to systematically assess the VLMs' capability to detect violations of physical common sense in images. Our evaluations show that even the state-of-the-art VLMs perform poorly on PhyVIB, highlighting a significant area for improvement. In response, we propose PhyDetector, a two-stage fine-tuning framework to enhance the VLMs' capability to detect violations of physical common sense. The first stage involves supervised fine-tuning, which equips the VLM with essential concepts related to visual physical anomalies. The second stage utilizes group relative policy optimization to enhance the VLM's multimodal reasoning capability on physical plausibility. Experimental results show that the model fine-tuned with PhyDetector can significantly outperform the state-of-the-art VLMs in physical common sense understanding. Our artifacts are available at https://github.com/ZitongWang018/PhyVIB.
Weibin Wu 0002, Zitong Wang 0007, Zhengjie Luo, Wenqing Chen, Zibin Zheng
ACM Multimedia4
2025 Towards an understanding of large language models in software engineering tasks
Zibin Zheng, Kaiwen Ning, Qingyuan Zhong, Jiachi Chen, Wenqing Chen, Lianghong Guo, Yanlin Wang 0001
Empir. Softw. Eng.5
2024 CoSTV: Accelerating Code Search with Two-Stage Paradigm and Vector Retrieval
abstract
Given a query in natural language, code search is designed to search the corresponding target code from a code base, which can accelerate the software development process. Recent pre-trained code models based on deep learning can capture the semantic connection between programming language and natural language, generating more accurate vector representations for codes and queries, significantly improving the matching accuracy between programming language and natural language. However, in recent years, most research on code search only focuses on improving the accuracy of code search while neglecting the importance of efficiency. In this paper, we propose a novel code search framework CoSTV to speed up the code search process. CoSTV employs a two-stage paradigm to combine the advantages of both bi-encoder and cross-encoder in terms of efficiency and accuracy, decoupling the code search procedure into recall and re-rank stages. Specifically, we introduce a vector retrieval system, program simplification, and knowledge distillation approaches to substantially accelerate code search while retaining parallel accuracy. In the recall stage, CoSTV utilizes a bi-encoder code search model and vector retrieval engine to rapidly recall highly relevant code candidates. In the re-rank stage, CoSTvemploys a cross-encoder-based code search model, program simplification, and model distillation to enhance the precision of code search. Extensive experiments conducted on the CodeSearchNet dataset indicate that compared with previous code search baselines, CoSTV can reduce the time of code search by 79.1 % while improving the accuracy of code search by 7.93 % on average.
Dewu Zheng, Yanlin Wang 0001, Wenqing Chen, Jiachi Chen, Zibin Zheng
APSEC3
2024 Unlock the Potential of Counterfactually-Augmented Data in Out-Of-Distribution Generalization
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin
Expert Syst. Appl.2
2023 Preference-Controlled Multi-Objective Reinforcement Learning for Conditional Text Generation
abstract
Conditional text generation is to generate text sequences conditioning on linguistic or non-linguistic data. The main line of existing work proposed deterministic models to improve the fidelity of the generated text but often ignored the diversity. Another line relied on conditional variational auto-encoders (CVAEs), which increased the diversity over their deterministic backbones. However, CVAEs regard diversity as an implicit objective and may not be optimal. In this paper, we raise two questions: i) Can diversity be further improved with an explicit objective? ii) Since fidelity and diversity are two conflicting objectives, how can we obtain different multi-objective optimal solutions according to user preferences? To answer question i), we propose a multi-objective reinforcement learning (MORL) method which explicitly takes CIDEr and Self-CIDEr scores as the fidelity-oriented and diversity-oriented rewards respectively. To answer question ii), we propose a preference-controlled MORL method, which can obtain infinite multi-objective optimal solutions by tuning the preference variable. We conduct extensive experiments on paraphrasing and image captioning tasks, which show that in the fidelity-diversity trade-off space, our model outperforms both deterministic and CVAE-based baselines.
Wenqing Chen, Jidong Tian, Caoyun Fan, Hao He 0007, Yaohui Jin
AAAI1
2023 Latent Constraints on Unsupervised Text-Graph Alignment with Information Asymmetry
abstract
Unsupervised text-graph alignment (UTGA) is a fundamental task that bidirectionally generates texts and graphs without parallel data. Most available models of UTGA suffer from information asymmetry, a common phenomenon that texts and graphs include additional information invisible to each other. On the one hand, these models fail to supplement asymmetric information effectively due to the lack of ground truths. On the other hand, it is challenging to indicate asymmetric information with explicit indicators because it cannot be decoupled from the data directly. To address the challenge posed by information asymmetry, we propose the assumption that asymmetric information is encoded in unobservable latent variables and only affects the one-way generation processes. These latent variables corresponding to asymmetric information should obey prior distributions recovered approximately from original data. Therefore, we first propose a taxonomy of the latent variable that classifies the latent variable into transferrable (TV) and non-transferable (NTV) variables and further distinguish NTV as the dependent variable (DV) and the independent variable (IV). Next, we propose three latent VAE-based regularizations on TV, DV, and IV to constrain their distributions to well-designed prior distributions to introduce asymmetric information into models and enhance the preservation of shared contents. Finally, we impose the three proposed constraints on a cycle-consistent learning framework, back-translation (BT), named ConstrainedBT. Experimental results on three UTGA tasks demonstrate the effectiveness of ConstrainedBT on the information-asymmetric challenge.
Jidong Tian, Wenqing Chen, Caoyun Fan, Hao He 0007, Yaohui Jin
AAAI2
2023 Chain-of-Thought Tuning: Masked Language Models can also Think Step By Step in Natural Language Understanding
abstract
Chain-of-Thought (CoT) is a technique that guides Large Language Models (LLMs) to decompose complex tasks into multi-step reasoning through intermediate steps in natural language form.Briefly, CoT enables LLMs to think step by step.However, although many Natural Language Understanding (NLU) tasks also require thinking step by step, LLMs perform less well than small-scale Masked Language Models (MLMs).To migrate CoT from LLMs to MLMs, we propose Chain-of-Thought Tuning (CoTT), a two-step reasoning framework based on prompt tuning, to implement step-by-step thinking for MLMs on NLU tasks.From the perspective of CoT, CoTT's two-step framework enables MLMs to implement task decomposition; CoTT's prompt tuning allows intermediate steps to be used in natural language form.Thereby, the success of CoT can be extended to NLU tasks through MLMs.To verify the effectiveness of CoTT, we conduct experiments on two NLU tasks: hierarchical classification and relation extraction, and the results show that CoTT outperforms baselines and achieves state-of-the-art performance.
Caoyun Fan, Jidong Tian, Wenqing Chen, Hao He 0007, Yaohui Jin
EMNLP4
2023 Improving the out-of-Distribution Generalization Capability of Language Models: Counterfactually-Augmented Data is not Enough
abstract
Counterfactually-Augmented Data (CAD) has the potential to improve language models’ Out-Of-Distribution (OOD) generalization capability, as CAD induces language models to exploit causal features and exclude spurious correlations. However, the empirical results of OOD generalization on CAD are not as efficient as expected. In this paper, we attribute the inefficiency to Myopia Phenomenon caused by CAD: language models only focus on causal features that are edited in the augmentation and exclude other non-edited causal features. As a result, the potential of CAD is not fully exploited. Based on the structural properties of CAD, we design two additional constraints to help language models extract more complete causal features contained in CAD, thus improving the OOD generalization capability. We evaluate our method on two tasks: Sentiment Analysis and Natural Language Inference, and the experimental results demonstrate that our method could unlock CAD’s potential and improve language models’ OOD generalization capability.
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin
ICASSP2
2023 You Augment Me: Exploring ChatGPT-based Data Augmentation for Semantic Code Search
abstract
Code search plays a crucial role in software development, enabling developers to retrieve and reuse code using natural language queries. While the performance of code search models improves with an increase in high-quality data, obtaining such data can be challenging and expensive. Recently, large language models (LLMs) such as ChatGPT have made remarkable progress in both natural and programming language understanding and generation, offering user-friendly interaction via simple prompts. Inspired by these advancements, we propose a novel approach ChatDANCE, which utilizes high-quality and diverse augmented data generated by a large language model and leverages a filtering mechanism to eliminate low-quality augmentations. Specifically, we first propose a set of ChatGPT prompting rules that are specifically designed for source code and queries. Then, we leverage ChatGPT to rewrite code and queries based on the according prompts and then propose a filtering mechanism which trains a cross-encoder from the backbone model UniXcoder to filter out code and query pairs with low matching scores. Finally, we re-train the backbone model using the obtained high-quality augmented data. Experimental results show that ChatDANCE achieves state-of-the-art performance, improving the best baseline by 13.2% (R@1) and 7% (MRR). Surprisingly, we find that this augment-filter-retrain strategy enables the backbone model (UniXcoder) to self-grow. Moreover, extensive experiments show the effectiveness of each component and ChatDANCE has stable performance under different hyperparameter settings. In addition, we conduct qualitative and quantitative analyses to investigate why ChatDANCE works well and find that it learns a more uniform distribution of representations and effectively aligns the code and query spaces. We have made the code and data anonymously available at https://anonymous.4open.science/r/ChatDANCE.
Yanlin Wang 0001, Lianghong Guo, Ensheng Shi, Wenqing Chen, Jiachi Chen, Wanjun Zhong, Hui Li 0057, Hongyu Zhang 0002, Ziyu Lyu, Zibin Zheng
ICSME4
2023 Accurate use of label dependency in multi-label text classification through the lens of causality
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin
Appl. Intell.2
2023 A Recommendation System of Personalized Resource Reliability for Online Teaching System under Large-scale User Access
Wenqing Chen
Mob. Networks Appl.1
2022 Weakly Supervised Neural Symbolic Learning for Cognitive Tasks
abstract
Despite the recent success of end-to-end deep neural networks, there are growing concerns about their lack of logical reasoning abilities, especially on cognitive tasks with perception and reasoning processes. A solution is the neural symbolic learning (NeSyL) method that can effectively utilize pre-defined logic rules to constrain the neural architecture making it perform better on cognitive tasks. However, it is challenging to apply NeSyL to these cognitive tasks because of the lack of supervision, the non-differentiable manner of the symbolic system, and the difficulty to probabilistically constrain the neural network. In this paper, we propose WS-NeSyL, a weakly supervised neural symbolic learning model for cognitive tasks with logical reasoning. First, WS-NeSyL employs a novel back search algorithm to sample the possible reasoning process through logic rules. This sampled process can supervise the neural network as the pseudo label. Based on this algorithm, we can backpropagate gradients to the neural network of WS-NeSyL in a weakly supervised manner. Second, we introduce a probabilistic logic regularization into WS-NeSyL to help the neural network learn probabilistic logic. To evaluate WS-NeSyL, we have conducted experiments on three cognitive datasets, including temporal reasoning, handwritten formula recognition, and relational reasoning datasets. Experimental results show that WS-NeSyL not only outperforms the end-to-end neural model but also beats the state-of-the-art neural symbolic learning models.
Jidong Tian, Wenqing Chen, Liqiang Xiao, Hao He 0007, Yaohui Jin
AAAI3
2022 MaxGNR: A Dynamic Weight Strategy via Maximizing Gradient-to-Noise Ratio for Multi-task Learning
Caoyun Fan, Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin
ACCV (1)2
2022 To What Extent Do Natural Language Understanding Datasets Correlate to Logical Reasoning? A Method for Diagnosing Logical Reasoning
abstract
Reasoning and knowledge-related skills are considered as two fundamental skills for natural language understanding (NLU) tasks such as machine reading comprehension (MRC) and natural language inference (NLI). However, it is not clear to what extent an NLU task defined on a dataset correlates to a specific NLU skill. On the one hand, evaluating the correlation requires an understanding of the significance of the NLU skill in a dataset. Significance judges whether a dataset includes sufficient material to help the model master this skill. On the other hand, it is also necessary to evaluate the dependence of the task on the NLU skill. Dependence is a measure of how much the task defined on a dataset depends on the skill. In this paper, we propose a systematic method to diagnose the correlations between an NLU dataset and a specific skill, and then take a fundamental reasoning skill, logical reasoning, as an example for analysis. The method adopts a qualitative indicator to indicate the significance while adopting a quantitative indicator to measure the dependence. We perform diagnosis on 8 MRC datasets (including two types) and 3 NLI datasets and acquire intuitively reasonable results. We then perform the analysis to further understand the results and the proposed indicators. Based on the analysis, although the diagnostic method has some limitations, it is still an effective method to perform a basic diagnosis of the correlation between the dataset and logical reasoning skill, which also can be generalized to other NLU skills.
Jidong Tian, Wenqing Chen, Caoyun Fan, Hao He 0007, Yaohui Jin
COLING3
2021 De-Confounded Variational Encoder-Decoder for Logical Table-to-Text Generation
abstract
Wenqing Chen, Jidong Tian, Yitian Li, Hao He, Yaohui Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wenqing Chen, Jidong Tian, Hao He 0007, Yaohui Jin
ACL/IJCNLP (1)1
2021 Diagnosing the First-Order Logical Reasoning Ability Through LogicNLI
abstract
Recently, language models (LMs) have achieved significant performance on many NLU tasks, which has spurred widespread interest for their possible applications in the scientific and social area.However, LMs have faced much criticism of whether they are truly capable of reasoning in NLU.In this work, we propose a diagnostic method for first-order logic (FOL) reasoning with a new proposed benchmark, LogicNLI.LogicNLI is an NLI-style dataset that effectively disentangles the target FOL reasoning from commonsense inference and can be used to diagnose LMs from four perspectives: accuracy, robustness, generalization, and traceability.Experiments on BERT, RoBERTa, and XLNet, have uncovered the weaknesses of these LMs on FOL reasoning, which motivates future exploration to enhance the reasoning ability.
Jidong Tian, Wenqing Chen, Liqiang Xiao, Hao He 0007, Yaohui Jin
EMNLP (1)3
2021 Dependent Multi-Task Learning with Causal Intervention for Image Captioning
abstract
Recent work for image captioning mainly followed an extract-then-generate paradigm, pre-extracting a sequence of object-based features and then formulating image captioning as a single sequence-to-sequence task. Although promising, we observed two problems in generated captions: 1) content inconsistency where models would generate contradicting facts; 2) not informative enough where models would miss parts of important information. From a causal perspective, the reason is that models have captured spurious statistical correlations between visual features and certain expressions (e.g., visual features of "long hair" and "woman"). In this paper, we propose a dependent multi-task learning framework with the causal intervention (DMTCI). Firstly, we involve an intermediate task, bag-of-categories generation, before the final task, image captioning. The intermediate task would help the model better understand the visual features and thus alleviate the content inconsistency problem. Secondly, we apply Pearl's do-calculus on the model, cutting off the link between the visual features and possible confounders and thus letting models focus on the causal visual features. Specifically, the high-frequency concept set is considered as the proxy confounders where the real confounders are inferred in the continuous space. Finally, we use a multi-agent reinforcement learning (MARL) strategy to enable end-to-end training and reduce the inter-task error accumulations. The extensive experiments show that our model outperforms the baseline models and achieves competitive performance with state-of-the-art models.
Wenqing Chen, Jidong Tian, Caoyun Fan, Hao He 0007, Yaohui Jin
IJCAI1
2020 A Semantically Consistent and Syntactically Variational Encoder-Decoder Framework for Paraphrase Generation
abstract
Paraphrase generation aims to generate semantically consistent sentences with different syntactic realizations.Most of the recent studies rely on the typical encoder-decoder framework where the generation process is deterministic.However, in practice, the ability to generate multiple syntactically different paraphrases is important.Recent work proposed to cooperate variational inference on a target-related latent variable to introduce the diversity.But the latent variable may be contaminated by the semantic information of other unrelated sentences, and in turn, change the conveyed meaning of generated paraphrases.In this paper, we propose a semantically consistent and syntactically variational encoder-decoder framework, which uses adversarial learning to ensure the syntactic latent variable be semantic-free.Moreover, we adopt another discriminator to improve the word-level and sentence-level semantic consistency.So the proposed framework can generate multiple semantically consistent and syntactically different paraphrases.The experiments show that our model outperforms the baseline models on the metrics based on both n-gram matching and semantic similarity, and our model can generate multiple different paraphrases by assembling different syntactic variables.
Wenqing Chen, Jidong Tian, Liqiang Xiao, Hao He 0007, Yaohui Jin
COLING1
2020 Exploring Logically Dependent Multi-task Learning with Causal Inference
abstract
Previous studies have shown that hierarchical multi-task learning (MTL) can utilize task dependencies by stacking encoders and outperform democratic MTL.However, stacking encoders only considers the dependencies of feature representations and ignores the label dependencies in logically dependent tasks.Furthermore, how to properly utilize the labels remains an issue due to the cascading errors between tasks.In this paper, we view logically dependent MTL from the perspective of causal inference and suggest a mediation assumption instead of the confounding assumption in conventional MTL models.We propose a model including two key mechanisms: label transfer (LT) for each task to utilize the labels of all its lower-level tasks, and Gumbel sampling (GS) to deal with cascading errors.In the field of causal inference, GS in our model is essentially a counterfactual reasoning process, trying to estimate the causal effect between tasks and utilize it to improve MTL.We conduct experiments on two English datasets and one Chinese dataset.Experiment results show that our model achieves state-of-the-art on six out of seven subtasks and improves predictions' consistency.
Wenqing Chen, Jidong Tian, Liqiang Xiao, Hao He 0007, Yaohui Jin
EMNLP (1)1
2018 Learning What to Share: Leaky Multi-Task Network for Text Classification
abstract
Neural network based multi-task learning has achieved great success on many NLP problems, which focuses on sharing knowledge among tasks by linking some layers to enhance the performance. However, most existing approaches suffer from the interference between tasks because they lack of selection mechanism for feature sharing. In this way, the feature spaces of tasks may be easily contaminated by helpless features borrowed from others, which will confuse the models for making correct prediction. In this paper, we propose a multi-task convolutional neural network with the Leaky Unit, which has memory and forgetting mechanism to filter the feature flows between tasks. Experiments on five different datasets for text classification validate the benefits of our approach.
Liqiang Xiao, Honglun Zhang, Wenqing Chen, Yongkun Wang, Yaohui Jin
COLING3
2018 MCapsNet: Capsule Network for Text with Multi-Task Learning
abstract
Multi-task learning has an ability to share the knowledge among related tasks and implicitly increase the training data.However, it has long been frustrated by the interference among tasks.This paper investigates the performance of capsule network for text, and proposes a capsule-based multi-task learning architecture, which is unified, simple and effective.With the advantages of capsules for feature clustering, proposed task routing algorithm can cluster the features for each task in the network, which helps reduce the interference among tasks.Experiments on six text classification datasets demonstrate the effectiveness of our models and their characteristics for feature clustering.
Liqiang Xiao, Honglun Zhang, Wenqing Chen, Yongkun Wang, Yaohui Jin
EMNLP3
2018 Multi-Task Label Embedding for Text Classification
abstract
Multi-task learning in text classification leverages implicit correlations among related tasks to extract common features and yield performance gains.However, a large body of previous work treats labels of each task as independent and meaningless one-hot vectors, which cause a loss of potential label information.In this paper, we propose Multi-Task Label Embedding to convert labels in text classification into semantic vectors, thereby turning the original tasks into vector matching tasks.Our model utilizes semantic correlations among tasks and makes it convenient to scale or transfer when new tasks are involved.Extensive experiments on five benchmark datasets for text classification show that our model can effectively improve the performances of related tasks with semantic representations of labels and additional information from each other.
Honglun Zhang, Liqiang Xiao, Wenqing Chen, Yongkun Wang, Yaohui Jin
EMNLP3
2018 Transformable Convolutional Neural Network for Text Classification
abstract
Convolutional neural networks (CNNs) have shown their promising performance for natural language processing tasks, which extract n-grams as features to represent the input. However, n-gram based CNNs are inherently limited to fixed geometric structure and cannot proactively adapt to the transformations of features. In this paper, we propose two modules to provide CNNs with the flexibility for complex features and the adaptability for transformation, namely, transformable convolution and transformable pooling. Our method fuses dynamic and static deviations to redistribute the sampling locations, which can capture both current and global transformations. Our modules can be easily integrated by other models to generate new transformable networks. We test proposed modules on two state-of-the-art models, and the results demonstrate that our modules can effectively adapt to the feature transformation in text classification.
Liqiang Xiao, Honglun Zhang, Wenqing Chen, Yongkun Wang, Yaohui Jin
IJCAI3
2018 Generative Warfare Nets: Ensemble via Adversaries and Collaborators
abstract
Generative Adversarial Nets are a powerful method for training generative models of complex data, where a Generator and a Discriminator confront with each other and get optimized in a two-player minmax manner. In this paper, we propose the Generative Warfare Nets (GWN) that involve multiple generators and multiple discriminators from two sides to exploit the advantages of Ensemble Learning. We maintain the authorities for the generators and the discriminators to enhance inter-side interactions, and utilize the mechanisms of imitation and innovation to model intra-side interactions among the generators, where they can not only learn from but also compete with each other. Extensive experiments on three natural image datasets show that GWN can achieve state-of-the-art Inception scores and produce diverse high-quality synthetic results.
Honglun Zhang, Liqiang Xiao, Wenqing Chen, Yongkun Wang, Yaohui Jin
IJCAI3
2017 The Data and Science behind GrabShare Carpooling
abstract
As internet-based ride hailing platforms such as Grab, Lyft, Uber become more and more popular, the demand for pooling services have increased sharply in recent years. In December 2016, Grab first rolled out its own pooling (ride-sharing) service "GrabShare" in Singapore. By providing considerable discount to passengers along with good matching performance to drivers, GrabShare enjoyed a rapid growth in Singapore to reach two million accumulative rides within the first two months. The success in Singapore motivates its expansion to other cities such as Manila, Jakarta, Kuala Lumpur, Ho Chi Minh City and so on [1]. In this article we explicitly present how the GrabShare algorithm was developed from a data perspective and how various ways of formulating the problem can have different impact to Grab, passengers, and drivers. Specifically, we start from a data perspective on how the GrabShare matching problem can be tackled by formulating an optimization model and how the optimization challenges are solved. Subsequently, we also present our critical data insights on improving user experience in the design of GrabShare and its continuous improvement for diverse markets.
Muchen Tang, Serene Ow, Wenqing Chen, Kong-wei Lye, Yaozhang Pan
DSAA3
2005 Virtual MIMO Protocol Based on Clustering for Wireless Sensor Network
abstract
The wireless nature of the medium combined with energy constraints pose big challenges on the design of energy efficient and reliable protocols for wireless sensor networks (WSN). In this paper, the virtual MIMO scheme is incorporated with the multi-hop networking for energy saving in wireless sensor network (WSN). An energy consumption model is developed to investigate the energy saving performance by incorporating the virtual MlMO scheme with the multi-hop networking. Then, an optimization model is developed to find the optimal number of virtual nodes and the optimal number of hops. Based on the model, a virtual MlMO protocol based on clustering is developed. Simulation results show that the protocol can save energy significantly.
Wenqing Chen, Changchun Xu, Kezhong Liu, Zongkai Yang
ISCC1