Qianglong Chen

dblp:277/9817 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0002-7845-1544ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 VideoPro: Adaptive Program Reasoning for Long Video Understanding
abstract
Chenglin Li, Feng Han, Yikun Wang, Ruilin Li, Shuai Dong, Haowen Hou, Haitao Li, Qianglong Chen, Feng Tao, Jingqi Tong, Yin Zhang, Jiaqi Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yikun Wang 0001, Haowen Hou, Qianglong Chen, Jingqi Tong, Yin Zhang 0006
ACL (1)8
2026 Simple-VGC: Enhancing Visual Grounding in Multimodal Reasoning via Adaptive Tool Composition
abstract
Ye Wang, Qianglong Chen, Siyuan Wang, Zejun Li, Shijie Guo, Zhirui Zhang, Zhongyu Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qianglong Chen, Shijie Guo, Zhirui Zhang, Zhongyu Wei
ACL (1)2
2026 NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI Tasks
abstract
Zihan Zheng, Tianle Cui, Taoran Wang, Fengtao Wang, Jiahui Pan, Lewei He, Qianglong Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zihan Zheng, Tianle Cui, Taoran Wang, Fengtao Wang, Jiahui Pan 0003, Lewei He, Qianglong Chen
ACL (1)7
2026 Exp2RL: Enhancing LLM Agent Training with Expert Experiences
Yicheng Li 0005, Qianglong Chen, Zhirui Zhang, Yin Zhang 0006
KSEM (3)2
2025 PlanningArena: A Modular Benchmark for Multidimensional Evaluation of Planning and Tool Learning
abstract
One of the research focuses of large language models (LLMs) is the ability to generate action plans.Recent studies have revealed that the performance of LLMs can be significantly improved by integrating external tools.Based on this, we propose a benchmark framework called PlanningArena, which aims to simulate real application scenarios and provide a series of apps and API tools that may be involved in the actual planning process.This framework adopts a modular task structure and combines user portrait analysis to evaluate the ability of LLMs in correctly selecting tools, logical reasoning in complex scenarios, and parsing user information.In addition, we deeply diagnose the task execution effect of LLMs from both macro and micro levels.The experimental results show that even the most outstanding GPT-4o and DeepSeekV3 models only achieved a total score of 56.5% and 41.9% in PlanningArena, respectively, indicating that current LLMs still face challenges in logical reasoning, context memory, and tool calling when dealing with different structures, scenarios, and their complexity.Through this benchmark, we further explore the path to optimize LLMs to perform planning tasks.
Zihan Zheng, Tianle Cui, Chuwen Xie, Jiahui Pan 0003, Qianglong Chen, Lewei He
ACL (1)5
2025 Towards Faithful Multi-step Reasoning through Fine-Grained Causal-aware Attribution Reasoning Distillation
abstract
Despite the remarkable reasoning capabilities demonstrated by large language models (LLM), the substantial computational overhead limits their practices. Some efforts have been directed toward distilling multi-step reasoning capabilities into smaller models through chain-of-thought (CoT). While CoT facilitates multi-step reasoning, the dependencies between reasoning steps are not always clearly discernible, which may lead to inconsistent reasoning. In this paper, we introduce fine-grained attribution reasoning distillation (FARD), which incorporates grounded citations to consolidate the relationships between reasoning steps. Specifically, FARD distills attribution reasoning rationales from LLMs to substitute CoT reasonings, which clarifies the dependencies among reasoning steps. Besides, we regularize the model’s attention pattern by leveraging the causal dependencies between reasoning steps, thereby enhancing the consistency of reasoning. Grounded attribution reasoning also enhances interpretability and verifiability, thereby facilitating faithful reasoning. We evaluate FARD on mathematical and general reasoning benchmarks. The experimental results indicate that FARD outperforms CoT distillation methods in mathematical reasoning, demonstrating its effectiveness. Furthermore, the small models trained with FARD have shown outstanding performance in out-of-distribution reasoning, proving strong generalization capabilities.
Jingchang Chen, Zhongjie Wang 0003, Guo Tang, Qianglong Chen, Ming Liu 0004, Bing Qin 0001
COLING5
2025 Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
abstract
The development of reasoning capabilities represents a critical frontier in large language models (LLMs) research, where reinforcement learning (RL) and process reward models (PRMs) have emerged as predominant methodological frameworks. Contrary to conventional wisdom, empirical evidence from DeepSeek-R1 demonstrates that pure RL training focused on mathematical problem-solving can progressively enhance reasoning abilities without PRM integration, challenging the perceived necessity of process supervision. In this study, we conduct a systematic investigation of the relationship between RL training and PRM capabilities. Our findings demonstrate that problem-solving proficiency and process supervision capabilities represent complementary dimensions of reasoning that co-evolve synergistically during pure RL training. In particular, current PRMs underperform simple baselines like majority voting when applied to state-of-the-art models such as DeepSeek-R1 and QwQ-32B. To address this limitation, we propose Self-PRM, an introspective framework in which models autonomously evaluate and rerank their generated solutions through self-reward mechanisms. Although Self-PRM consistently improves the accuracy of the benchmark (particularly with larger sample sizes), analysis exposes persistent challenges: The approach exhibits low precision (<10\%) on difficult problems, frequently misclassifying flawed solutions as valid. These analyses underscore the need for combined training with process supervision and continued RL scaling to enhance reward alignment and introspective accuracy. We hope these findings provide actionable insights for building more reliable and self-aware complex reasoning models.
Zhangyin Feng, Qianglong Chen, Ning Lu 0006, Yongqian Li, Siqi Cheng, Shuangmu Peng, Duyu Tang, Shengcai Liu, Zhirui Zhang
NeurIPS2
2025 Learning to break: Knowledge-enhanced reasoning in multi-agent debate system
Haotian Wang 0007, Xiyuan Du, Weijiang Yu, Qianglong Chen, Kun Zhu 0025, Lian Yan, Yi Guan
Neurocomputing4
2025 Beyond decomposition: Hierarchical dependency management in multi-document question answering
abstract
Abstract When using retrieval‐augmented generation (RAG) to handle multi‐document question answering (MDQA) tasks, it is beneficial to decompose complex queries into multiple simpler ones to enhance retrieval results. However, previous strategies always employ a one‐shot approach of question decomposition, overlooking subquestions dependency problem and failing to ensure that the derived subqueries are single‐hop. To overcome this challenge, we introduce a novel framework called DSRC‐QCS. Decompose‐solve‐renewal‐cycle (DSRC) is an iterative multi‐hop question processing module. The key idea of DSRC involves using a unique symbol to achieve hierarchical dependency management and employing a cyclical process of question decomposition, solving, and renewal to continuously generate and resolve all single‐hop subquestions. Query‐chain selector (QCS) functions as a voting mechanism that effectively utilizes the reasoning process of DSRC to assess and select solutions. We compare DSRC‐QCS against five RAG approaches across three datasets and three LLMs. DSRC‐QCS demonstrates superior performance. Compared to the Direct Retrieval method, DSRC‐QCS improves the average F1 score by 17.36% with Alpaca‐7b, 10.83% with LLaMa2‐Chat‐7b, and 11.88% with GPT‐3.5‐Turbo. We also conduct ablation studies to validate the performance of both DSRC and QCS and explore factors influencing the effectiveness of DSRC. We have included all prompts in the Appendix.
Xiaoyan Zheng, Qianglong Chen, Yin Zhang 0006
J. Assoc. Inf. Sci. Technol.3
2025 A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
abstract
The emergence of large language models (LLMs) has marked a significant breakthrough in natural language processing (NLP), fueling a paradigm shift in information acquisition. Nevertheless, LLMs are prone to hallucination, generating plausible yet nonfactual content. This phenomenon raises significant concerns over the reliability of LLMs in real-world information retrieval (IR) systems and has attracted intensive research to detect and mitigate such hallucinations. Given the open-ended general-purpose attributes inherent to LLMs, LLM hallucinations present distinct challenges that diverge from prior task-specific models. This divergence highlights the urgency for a nuanced understanding and comprehensive overview of recent advances in LLM hallucinations. In this survey, we begin with an innovative taxonomy of hallucination in the era of LLM and then delve into the factors contributing to hallucinations. Subsequently, we present a thorough overview of hallucination detection methods and benchmarks. Our discussion then transfers to representative methodologies for mitigating LLM hallucinations. Additionally, we delve into the current limitations faced by retrieval-augmented LLMs in combating hallucinations, offering insights for developing more robust IR systems. Finally, we highlight the promising research directions on LLM hallucinations, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations.
Lei Huang 0021, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang 0007, Qianglong Chen, Weihua Peng, Bing Qin 0001, Ting Liu 0001
ACM Trans. Inf. Syst.7
2024 An Information Bottleneck Perspective for Effective Noise Filtering on Retrieval-Augmented Generation
abstract
Kun Zhu, Xiaocheng Feng, Xiyuan Du, Yuxuan Gu, Weijiang Yu, Haotian Wang, Qianglong Chen, Zheng Chu, Jingchang Chen, Bing Qin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Kun Zhu 0025, Xiyuan Du, Yuxuan Gu 0004, Weijiang Yu, Haotian Wang 0007, Qianglong Chen, Jingchang Chen, Bing Qin 0001
ACL (1)7
2024 BeamAggR: Beam Aggregation Reasoning over Multi-source Knowledge for Multi-hop Question Answering
abstract
Zheng Chu, Jingchang Chen, Qianglong Chen, Haotian Wang, Kun Zhu, Xiyuan Du, Weijiang Yu, Ming Liu, Bing Qin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jingchang Chen, Qianglong Chen, Haotian Wang 0007, Kun Zhu 0025, Xiyuan Du, Weijiang Yu, Ming Liu 0004, Bing Qin 0001
ACL (1)3
2024 TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
abstract
Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Haotian Wang, Ming Liu, Bing Qin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jingchang Chen, Qianglong Chen, Weijiang Yu, Haotian Wang 0007, Ming Liu 0004, Bing Qin 0001
ACL (1)3
2024 Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future
abstract
Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, Ting Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He 0014, Haotian Wang 0007, Weihua Peng, Ming Liu 0004, Bing Qin 0001, Ting Liu 0001
ACL (1)3
2024 Divide-and-Conquer Meets Consensus: Unleashing the Power of Functions in Code Generation
abstract
Despite recent progress made by large language models in code generation, they still struggle with programs that meet complex requirements. Recent work utilizes plan-and-solve decomposition to decrease the complexity and leverage self-tests to refine the generated program. Yet, planning deep-inside requirements in advance can be challenging, and the tests need to be accurate to accomplish self-improvement. To this end, we propose FunCoder, a code generation framework incorporating the divide-and-conquer strategy with functional consensus. Specifically, FunCoder recursively branches off sub-functions as smaller goals during code generation, represented by a tree hierarchy. These sub-functions are then composited to attain more complex objectives. Additionally, we designate functions via a consensus formed by identifying similarities in program behavior, mitigating error propagation. FunCoder outperforms state-of-the-art methods by +9.8% on average in HumanEval, MBPP, xCodeEval and MATH with GPT-3.5 and GPT-4. Moreover, our method demonstrates superiority on smaller models: With FunCoder, StableCode-3b surpasses GPT-3.5 by +18.6% and achieves 97.7% of GPT-4's performance on HumanEval. Further analysis reveals that our proposed dynamic function decomposition is capable of handling complex requirements, and the functional consensus prevails over self-testing in correctness evaluation.
Jingchang Chen, Hongxuan Tang, Qianglong Chen, Zekun Wang 0001, Ming Liu 0004, Bing Qin 0001
NeurIPS4
2023 Task Difficulty Aware Parameter Allocation & Regularization for Lifelong Learning
abstract
Parameter regularization or allocation methods are effective in overcoming catastrophic forgetting in lifelong learning. However, they solve all tasks in a sequence uniformly and ignore the differences in the learning difficulty of different tasks. So parameter regularization methods face significant forgetting when learning a new task very different from learned tasks, and parameter allocation methods face unnecessary parameter overhead when learning simple tasks. In this paper, we propose the Parameter Allocation & Regularization (PAR), which adaptively select an appropriate strategy for each task from parameter allocation and regularization based on its learning difficulty. A task is easy for a model that has learned tasks related to it and vice versa. We propose a divergence estimation method based on the Nearest-Prototype distance to measure the task relatedness using only features of the new task. Moreover, we propose a time-efficient relatedness-aware sampling-based architecture search strategy to reduce the parameter overhead for allocation. Experimental results on multiple benchmarks demonstrate that, compared with SOTAs, our method is scalable and significantly reduces the model's redundancy while improving the model's performance. Further qualitative analysis indicates that PAR obtains reasonable task-relatedness.
Wenjin Wang 0003, Yunqing Hu, Qianglong Chen, Yin Zhang 0006
CVPR3
2022 Continual Few-shot Intent Detection
abstract
Intent detection is at the core of task-oriented dialogue systems. Existing intent detection systems are typically trained with a large amount of data over a predefined set of intent classes. However, newly emerged intents in multiple domains are commonplace in the real world. And it is time-consuming and impractical for dialogue systems to re-collect enough annotated data and re-train the model. These limitations call for an intent detection system that could continually recognize new intents with very few labeled examples. In this work, we study the Continual Few-shot Intent Detection (CFID) problem and construct a benchmark consisting of nine tasks with multiple domains and imbalanced classes. To address the key challenges of (a) catastrophic forgetting during continuous learning and (b) negative knowledge transfer across tasks, we propose the Prefix-guided Lightweight Encoder (PLE) with three auxiliary strategies, namely Pseudo Samples Replay (PSR), Teacher Knowledge Transfer (TKT) and Dynamic Weighting Replay (DWR). Extensive experiments demonstrate the effectiveness and efficiency of our method in preventing catastrophic forgetting and encouraging positive knowledge transfer across tasks.
Guodun Li, Yuchen Zhai, Qianglong Chen, Ji Zhang 0011, Yin Zhang 0006
COLING3
2022 DictBERT: Dictionary Description Knowledge Enhanced Language Model Pre-training via Contrastive Learning
abstract
Although pre-trained language models (PLMs) have achieved state-of-the-art performance on various natural language processing (NLP) tasks, they are shown to be lacking in knowledge when dealing with knowledge driven tasks. Despite the many efforts made for injecting knowledge into PLMs, this problem remains open. To address the challenge, we propose DictBERT, a novel approach that enhances PLMs with dictionary knowledge which is easier to acquire than knowledge graph (KG). During pre-training, we present two novel pre-training tasks to inject dictionary knowledge into PLMs via contrastive learning: dictionary entry prediction and entry description discrimination. In fine-tuning, we use the pre-trained DictBERT as a plugin knowledge base (KB) to retrieve implicit knowledge for identified entries in an input sequence, and infuse the retrieved knowledge into the input to enhance its representation via a novel extra-hop attention mechanism. We evaluate our approach on a variety of knowledge driven and language understanding tasks, including NER, relation extraction, CommonsenseQA, OpenBookQA and GLUE. Experimental results demonstrate that our model can significantly improve typical PLMs: it gains a substantial improvement of 0.5%, 2.9%, 9.0%, 7.1% and 3.3% on BERT-large respectively, and is also effective on RoBERTa-large.
Qianglong Chen, Feng-Lin Li, Guohai Xu, Ming Yan 0008, Ji Zhang 0011, Yin Zhang 0006
IJCAI1
2022 mmLayout: Multi-grained MultiModal Transformer for Document Understanding
abstract
Recent efforts of multimodal Transformers have improved Visually Rich Document Understanding (VrDU) tasks via incorporating visual and textual information. However, existing approaches mainly focus on fine-grained elements such as words and document image patches, making it hard for them to learn from coarse-grained elements, including natural lexical units like phrases and salient visual regions like prominent image regions. In this paper, we attach more importance to coarse-grained elements containing high-density information and consistent semantics, which are valuable for document understanding. At first, a document graph is proposed to model complex relationships among multi-grained multimodal elements, in which salient visual regions are detected by a cluster-based method. Then, a multi-grained multimodal Transformer called mmLayout is proposed to incorporate coarse-grained information into existing pre-trained fine-grained multimodal Transformers based on the graph. In mmLayout, coarse-grained information is aggregated from fine-grained, and then, after further processing, is fused back into fine-grained for final prediction. Furthermore, common sense enhancement is introduced to exploit the semantic information of natural lexical units. Experimental results on four tasks, including information extraction and document question answering, show that our method can improve the performance of multimodal Transformers based on fine-grained elements and achieve better performance with fewer parameters. Qualitative analyses show that our method can capture consistent semantics in coarse-grained elements.
Wenjin Wang 0003, Zhengjie Huang, Qianglong Chen, Qiming Peng, Yinxu Pan, Weichong Yin, Shikun Feng, Yu Sun 0029, Dianhai Yu, Yin Zhang 0006
ACM Multimedia4
2022 Rethinking the Value of Gazetteer in Chinese Named Entity Recognition
Qianglong Chen, Xiangji Zeng, Jiangang Zhu, Yin Zhang 0006, Bojia Lin, Yang Yang 0012, Daxin Jiang
NLPCC (1)1
2021 KACE: Generating Knowledge Aware Contrastive Explanations for Natural Language Inference
abstract
Qianglong Chen, Feng Ji, Xiangji Zeng, Feng-Lin Li, Ji Zhang, Haiqing Chen, Yin Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Qianglong Chen, Xiangji Zeng, Feng-Lin Li, Ji Zhang 0011, Haiqing Chen, Yin Zhang 0006
ACL/IJCNLP (1)1
2021 K-AID: Enhancing Pre-trained Language Models with Domain Knowledge for Question Answering
abstract
Knowledge enhanced pre-trained language models (K-PLMs) are shown to be effective for many public tasks in the literature, but few of them have been successfully applied in practice. To address this problem, we propose K-AID, a systematic approach that includes a low-cost knowledge acquisition process for acquiring domain knowledge, an effective knowledge infusion module for improving model performance, and a knowledge distillation component for reducing the model size and deploying K-PLMs on resource-restricted devices (e.g., CPU) for real-world application. Importantly, instead of capturing entity knowledge like the majority of existing K-PLMs, our approach captures relational knowledge, which contributes to better improving sentence-level text classification and text matching tasks that play a key role in question answering (QA). We conducted a set of experiments on five text classification tasks and three text matching tasks from three domains, namely E-commerce, Government, and Film&TV, and performed online A/B tests in E-commerce. Experimental results show that our approach is able to achieve substantial improvement on sentence-level question answering tasks and bring beneficial business value in industrial settings.
Fu Sun, Feng-Lin Li, Qianglong Chen, Xingyi Cheng, Ji Zhang 0011
CIKM4
2020 Improving Commonsense Question Answering by Graph-based Iterative Retrieval over Multiple Knowledge Sources
abstract
In order to facilitate natural language understanding, the key is to engage commonsense or background knowledge.However, how to engage commonsense effectively in question answering systems is still under exploration in both research academia and industry.In this paper, we propose a novel question-answering method by integrating multiple knowledge sources, i.e.Con-ceptNet, Wikipedia, and the Cambridge Dictionary, to boost the performance.More concretely, we first introduce a novel graph-based iterative knowledge retrieval module, which iteratively retrieves concepts and entities related to the given question and its choices from multiple knowledge sources.Afterward, we use a pre-trained language model to encode the question, retrieved knowledge and choices, and propose an answer choice-aware attention mechanism to fuse all hidden representations of the previous modules.Finally, the linear classifier for specific tasks is used to predict the answer.Experimental results on the CommonsenseQA dataset show that our method significantly outperforms other competitive methods and achieves the new state-ofthe-art.In addition, further ablation studies demonstrate the effectiveness of our graph-based iterative knowledge retrieval module and the answer choice-aware attention module in retrieving and synthesizing background knowledge from multiple knowledge sources.
Qianglong Chen, Haiqing Chen, Yin Zhang 0006
COLING1