Chen Huang 0006

dblp:05/8125-6 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-3542-7085ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models
abstract
Pengfeng Li, Chen Huang, Chaoqun Hao, Hongyao Chen, Xiao-Yong Wei, Wenqiang Lei, See-Kiong Ng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pengfeng Li, Chen Huang 0006, Chaoqun Hao, Hongyao Chen, Xiaoyong Wei, Wenqiang Lei, See-Kiong Ng
ACL (1)2
2026 E2EDev: Benchmarking Large Language Models in End-to-End Software Development Task
abstract
The rapid advancement in large language models (LLMs) has demonstrated significant potential in End-to-End Software Development (E2ESD).However, existing E2ESD benchmarks are limited by coarse-grained requirement specifications and unreliable evaluation protocols, hindering a true understanding of current framework capabilities.To address these limitations, we present E2EDev, a novel benchmark grounded in the principles of Behavior-Driven Development (BDD) to assess whether the generated software meets user needs through mimicking real user interactions.E2EDev comprises (i) a fine-grained set of user requirements for each target software project (ii) multiple BDD test scenarios with corresponding Python step implementations for each requirement, and (iii) a fully automated testing pipeline built on the Behave framework.By evaluating various E2ESD frameworks and LLM backbones with E2EDev, our analysis reveals a persistent struggle to effectively solve these tasks, underscoring the critical need for more effective and cost-efficient E2ESD solutions.Our codebase and benchmark are available at https://github.com/ SCUNLP/E2EDev.
Chen Huang 0006, Zhizhao Guan, Wenqiang Lei, Yang Deng 0002
ACL (1)2
2026 METRO: Towards Strategy Induction from Expert Dialogue Transcripts for Non-collaborative Dialogues
abstract
Developing non-collaborative dialogue agents traditionally requires the manual, unscalable codification of expert strategies.We propose METRO, a method that leverages large language models to autonomously induce both strategy actions and planning logic directly from raw transcripts.METRO formalizes expert knowledge into a Strategy Forest, a hierarchical structure that captures both shortterm responses (nodes) and long-term strategic foresight (branches).Experimental results across two benchmarks show that METRO demonstrates promising performance, outperforming existing methods by an average of 9%-10%.Our further analysis not only reveals the success behind METRO (strategic behavioral diversity and foresight), but also demonstrates its robust cross-task transferability.This offers new insights into building non-collaborative agents in a cost-effective and scalable way.Our code is available at
Haofu Yang, Jiaji Liu, Chen Huang 0006, Faguo Wu, Wenqiang Lei, See-Kiong Ng
ACL (1)3
2026 EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
abstract
Yifan Zhang, Chen Huang, Yueke Zhang, Jiahao Zhang, Toby Jia-Jun Li, Collin McMillan, Kevin Leach, Yu Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yifan Zhang 0013, Chen Huang 0006, Yueke Zhang, Toby Jia-Jun Li, Collin McMillan, Kevin Leach, Yu Huang 0015
ACL (1)2
2026 Co-Matching: Towards Human-Model Collaborative Legal Case Matching
abstract
Recent efforts have aimed to improve AI models in legal case matching by integrating legal domain knowledge. However, successful legal case matching requires the tacit knowledge of legal practitioners, which is difficult to verbalize and encode into models. This emphasizes the crucial role of involving legal practitioners in high-stakes legal case matching. To address this, we propose a collaborative matching framework called Co-Matching , which encourages both the model and the legal practitioner to participate in the matching process, integrating tacit knowledge. Unlike existing methods that rely solely on the model, Co-Matching allows both the legal practitioner and the model to determine key sentences and then combine them probabilistically. Co-Matching introduces a method called ProtoEM to estimate human decision uncertainty, facilitating the probabilistic combination. Experimental results demonstrate that Co-Matching consistently outperforms existing legal case matching methods, delivering significant performance improvements over human- and model-based matching in isolation (on average, +5.51% and +8.71%, respectively). Further analysis shows that Co-Matching also ensures better human–model collaboration effectiveness. Our study represents an effort in human–model collaboration for the legal case matching task, marking a milestone for future collaborative matching studies.
Chen Huang 0006, Yang Deng 0002, Wenqiang Lei, Jiancheng Lv 0001, Tat-Seng Chua
ACM Trans. Inf. Syst.1
2025 LEGEND: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
abstract
The success of the reward model in distinguishing between responses with subtle safety differences depends critically on the high-quality preference dataset, which should capture the fine-grained nuances of harmful and harmless responses. This motivates the need to develop the datasets involving preference margins, which accurately quantify how harmless one response is compared to another. In this paper, we take the first step to propose an effective and cost-efficient framework to promote the margin-enhanced preference dataset development. Our framework, Legend, Leverages rEpresentation enGineering to annotate preferENce Datasets. It constructs the specific direction within the LLM's embedding space that represents safety. By leveraging this safety direction, Legend can then leverage the semantic distances of paired responses along this direction to annotate margins automatically. We experimentally demonstrate our effectiveness in both reward modeling and harmless alignment for LLMs. Legend also stands out for its efficiency, requiring only the inference time rather than additional training. This efficiency allows for easier implementation and scalability, making Legend particularly valuable for practical applications in aligning LLMs with safe conversations.
Duanyu Feng, Bowen Qin, Chen Huang 0006, Youcheng Huang, Zheng Zhang 0006, Wenqiang Lei
AAAI3
2025 How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond
abstract
With the advancement of large language models (LLMs), intelligent models have evolved from mere tools to autonomous agents with their own goals and strategies for cooperating with humans.This evolution has birthed a novel paradigm in NLP, i.e., human-model cooperation, that has yielded remarkable progress in numerous NLP tasks in recent years.In this paper, we take the first step to present a thorough review of human-model cooperation, exploring its principles, formalizations, and open challenges.In particular, we introduce a new taxonomy that provides a unified perspective to summarize existing approaches.Also, we discuss potential frontier areas and their corresponding challenges.We regard our work as an entry point, paving the way for more breakthrough research in this regard.
Chen Huang 0006, Yang Deng 0002, Wenqiang Lei, Jiancheng Lv 0001, Tat-Seng Chua, Jimmy Huang 0001
ACL (1)1
2025 Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
abstract
Understanding the inner workings of Large Language Models (LLMs) is a critical research frontier.Prior work has shown that a single LLM's concept representations can be captured as steering vectors (SVs), enabling the control of LLM behavior (e.g., towards generating harmful content).This paper takes a novel approach by exploring the intricate relationships between representations of concepts across different LLMs, drawing an intriguing parallel to the Plato's Allegory of the Cave.In particular, we introduce a linear transformation method to bridge these representations and present three key findings: 1) The representations of a same concept in different LLMs can be effectively aligned using simple linear transformations, enabling efficient cross-model transfer and behavioral control via SVs.2) This linear transformation generalizes across multiple concepts, facilitating alignment and control of SVs representing different concepts across LLMs.3) A weakto-strong transferability exists between LLMs, whereby SVs extracted from smaller LLMs can effectively control behaviors of larger LLMs. 1
Youcheng Huang, Chen Huang 0006, Duanyu Feng, Wenqiang Lei, Jiancheng Lv 0001
ACL (1)2
2025 Can Large Language Models Understand Internet Buzzwords Through User-Generated Content
abstract
The massive user-generated content (UGC) available in Chinese social media is giving rise to the possibility of studying internet buzzwords.In this paper, we study if large language models (LLMs) can generate accurate definitions for these buzzwords based on UGC as examples.Our work serves a threefold contribution.First, we introduce CHEER, the first dataset of Chinese internet buzzwords, each annotated with a definition and relevant UGC.Second, we propose a novel method, called RESS, to effectively steer the comprehending process of LLMs to produce more accurate buzzword definitions, mirroring the skills of human language learning.Third, with CHEER, we benchmark the strengths and weaknesses of various off-the-shelf definition generation methods and our RESS.Our benchmark demonstrates the effectiveness of RESS while revealing crucial shared challenges: over-reliance on prior exposure, underdeveloped inferential abilities, and difficulty identifying high-quality UGC to facilitate comprehension.We believe our work lays the groundwork for future advancements in LLM-based definition generation.Our dataset and code are available at https://github.com/SCUNLP/Buzzword.
Chen Huang 0006, Junkai Luo, Xinzuo Wang, Wenqiang Lei, Jiancheng Lv 0001
ACL (1)1
2025 ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive Programming
abstract
Xinwei Yang, Zhaofeng Liu, Chen Huang, Jiashuai Zhang, Tong Zhang, Yifan Zhang, Wenqiang Lei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhaofeng Liu, Chen Huang 0006, Jiashuai Zhang, Yifan Zhang 0013, Wenqiang Lei
ACL (1)3
2025 GraphOTTER: Evolving LLM-based Graph Reasoning for Complex Table Question Answering
abstract
Complex Table Question Answering involves providing accurate answers to specific questions based on intricate tables that exhibit complex layouts and flexible header locations. Despite considerable progress having been made in the LLM era, the reasoning processes of existing methods are often implicit, feeding the entire table into prompts, making it difficult to effectively filter out irrelevant information in the table. To this end, we propose GraphOTTER that explicitly establishes the reasoning process to pinpoint the correct answers. In particular, GraphOTTER leverages a graph-based representation, transforming the complex table into an undirected graph. It then conducts step-by-step reasoning on the graph, with each step guided by a set of pre-defined intermediate reasoning actions. As such, it constructs a clear reasoning path and effectively identifies the answer to a given question. Comprehensive experiments on two benchmark datasets and two LLM backbones demonstrate the effectiveness of GraphOTTER. Further analysis indicates that its success may be attributed to the ability to efficiently filter out irrelevant information, thereby focusing the reasoning process on the most pertinent data. Our code and experimental datasets are available at https://github.com/JDing0521/GraphOTTER.
Qianlong Li, Chen Huang 0006, Yuanxin Xiang, Deng Xiong, Wenqiang Lei
COLING2
2025 ReMindRAG: Low-Cost LLM-Guided Knowledge Graph Traversal for Efficient RAG
abstract
Knowledge graphs (KGs), with their structured representation capabilities, offer promising avenue for enhancing Retrieval Augmented Generation (RAG) systems, leading to the development of KG-RAG systems. Nevertheless, existing methods often struggle to achieve effective synergy between system effectiveness and cost efficiency, leading to neither unsatisfying performance nor excessive LLM prompt tokens and inference time. To this end, this paper proposes REMINDRAG, which employs an LLM-guided graph traversal featuring node exploration, node exploitation, and, most notably, memory replay, to improve both system effectiveness and cost efficiency. Specifically, REMINDRAG memorizes traversal experience within KG edge embeddings, mirroring the way LLMs "memorize" world knowledge within their parameters, but in a train-free manner. We theoretically and experimentally confirm the effectiveness of REMINDRAG, demonstrating its superiority over existing baselines across various benchmark datasets and LLM backbones. Our code is available at https://github.com/kilgrims/ReMindRAG.
Yikuan Hu, Jifeng Zhu, Lanrui Tang, Chen Huang 0006
NeurIPS4
2024 Towards Equipping Transformer with the Ability of Systematic Compositionality
abstract
One of the key factors in language productivity and human cognition is the ability of Systematic Compositionality, which refers to understanding composed, unseen examples of seen primitives. However, recent evidence reveals that the Transformers have difficulty in generalizing the composed context based on the seen primitives. To this end, we take the first step to propose a compositionality-aware Transformer called CAT and two novel pre-training tasks to facilitate the systematic compositionality. We tentatively provide a successful implementation of a multi-layer CAT on the basis of the especially popular BERT. The experimental results demonstrate that CAT outperforms baselines on compositionality-aware tasks with minimal impact on effectiveness on standardized language understanding tasks.
Chen Huang 0006, Peixin Qin, Wenqiang Lei, Jiancheng Lv 0001
AAAI1
2024 ARAIDA: Analogical Reasoning-Augmented Interactive Data Annotation
abstract
Human annotation is a time-consuming task that requires a significant amount of effort.To address this issue, interactive data annotation utilizes an annotation model to provide suggestions for humans to approve or correct.However, annotation models trained with limited labeled data are prone to generating incorrect suggestions, leading to extra human correction effort.To tackle this challenge, we propose ARAIDA, an analogical reasoning-based approach that enhances automatic annotation accuracy in the interactive data annotation setting and reduces the need for human corrections.ARAIDA involves an error-aware integration strategy that dynamically coordinates an annotation model and a k-nearest neighbors (KNN) model, giving more importance to KNN's predictions when predictions from the annotation model are deemed inaccurate.Empirical studies demonstrate that ARAIDA is adaptable to different annotation tasks and models.On average, it reduces human correction labor by 11.02% compared to vanilla interactive data annotation methods.
Chen Huang 0006, Yiping Jin, Ilija Ilievski, Wenqiang Lei, Jiancheng Lv 0001
ACL (1)1
2024 CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language Models
abstract
Tong Zhang, Peixin Qin, Yang Deng, Chen Huang, Wenqiang Lei, Junhong Liu, Dingnan Jin, Hongru Liang, Tat-Seng Chua. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Peixin Qin, Yang Deng 0002, Chen Huang 0006, Wenqiang Lei, Dingnan Jin, Hongru Liang, Tat-Seng Chua
ACL (1)4
2024 Strength Lies in Differences! Improving Strategy Planning for Non-collaborative Dialogues via Diversified User Simulation
abstract
We investigate non-collaborative dialogue agents, which are expected to engage in strategic conversations with diverse users, for securing a mutual agreement that leans favorably towards the system's objectives.This poses two main challenges for existing dialogue agents:1) The inability to integrate user-specific characteristics into the strategic planning, and 2) The difficulty of training strategic planners that can be generalized to diverse users.To address these challenges, we propose TRIP to enhance the capability in tailored strategic planning, incorporating a user-aware strategic planning module and a population-based training paradigm.Through experiments on benchmark non-collaborative dialogue tasks, we demonstrate the effectiveness of TRIP in catering to diverse users.
Chen Huang 0006, Yang Deng 0002, Hongru Liang, Zujie Wen, Wenqiang Lei, Tat-Seng Chua
EMNLP2
2024 Cross-Space Adaptive Filter: Integrating Graph Topology and Node Attributes for Alleviating the Over-smoothing Problem
abstract
The vanilla Graph Convolutional Network (GCN) uses a low-pass filter to extract low-frequency signals from graph topology, which may lead to the over-smoothing problem when GCN goes deep. To this end, various methods have been proposed to create an adaptive filter by incorporating an extra filter (e.g., a high-pass filter) extracted from the graph topology. However, these methods heavily rely on topological information and ignore the node attribute space, which severely sacrifices the expressive power of the deep GCNs, especially when dealing with disassortative graphs. In this paper, we propose a cross-space adaptive filter, called CSF, to produce the adaptive-frequency information extracted from both the topology and attribute spaces. Specifically, we first derive a tailored attribute-based high-pass filter that can be interpreted theoretically as a minimizer for semi-supervised kernel ridge regression. Then, we cast the topology-based low-pass filter as a Mercer's kernel within the context of GCNs. This serves as a foundation for combining it with the attribute-based filter to capture the adaptive-frequency information. Finally, we derive the cross-space filter via an effective multiple-kernel learning strategy, which unifies the attribute-based high-pass filter and the topology-based low-pass filter. This helps to address the over-smoothing problem while maintaining effectiveness. Extensive experiments demonstrate that CSF not only successfully alleviates the over-smoothing problem but also promotes the effectiveness of the node classification task. Our code is available at https://github.com/huangzichun/Cross-Space-Adaptive-Filter.
Chen Huang 0006, Haoyang Li 0001, Yifan Zhang 0013, Wenqiang Lei, Jiancheng Lv 0001
WWW1
2023 TRAVEL: Tag-Aware Conversational FAQ Retrieval via Reinforcement Learning
abstract
Efficiently retrieving FAQ questions that match users' intent is essential for online customer service.Existing methods aim to fully utilize the dynamic conversation context to enhance the semantic association between the user query and FAQ questions.However, the conversation context contains noise, e.g., users may click questions they don't like, leading to inaccurate semantics modeling.To tackle this, we introduce tags of FAQ questions, which can help us eliminate irrelevant information.We later integrate them into a reinforcement learning framework and minimize the negative impact of irrelevant information in the dynamic conversation context.We experimentally demonstrate our efficiency and effectiveness on conversational FAQ retrieval compared to other baselines.
Dingnan Jin, Chen Huang 0006, Wenqiang Lei
EMNLP3
2023 Reduce Human Labor On Evaluating Conversational Information Retrieval System: A Human-Machine Collaboration Approach
abstract
Evaluating conversational information retrieval (CIR) systems is a challenging task that requires a significant amount of human labor for annotation.It is imperative to invest significant effort into researching more labor-effective methods for evaluating CIR systems.To touch upon this challenge, we take the first step to involve active testing in CIR evaluation and propose a novel method called HumCoE.It strategically selects a few data for human annotation and then calibrates the evaluation results to eliminate evaluation biases.As such, it makes an accurate evaluation of the CIR system at low human labor.We experimentally reveal that it consumes less than 1% of human labor and achieves a consistency rate of 95%-99% with human evaluation results.This emphasizes the superiority of our method.
Chen Huang 0006, Peixin Qin, Wenqiang Lei, Jiancheng Lv 0001
EMNLP1