Zhi Chen 0006

dblp:05/1539-6 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
14since 2021 · last 2025
0000-0003-4180-8455ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2025 What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best Practices
abstract
Zhi Chen, Qiguang Chen, Libo Qin, Qipeng Guo, Haijun Lv, Yicheng Zou, Hang Yan, Kai Chen, Dahua Lin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhi Chen 0006, Qiguang Chen, Libo Qin 0001, Qipeng Guo, Haijun Lv, Yicheng Zou, Hang Yan 0001, Kai Chen 0026, Dahua Lin
ACL (1)1
2025 Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law
abstract
Scaling law builds the relationship between training computation and validation loss, enabling researchers to effectively predict the loss trending of models across different levels of computation. However, a gap still remains between validation loss and the model’s downstream capabilities, making it untrivial to apply scaling law to direct performance prediction for downstream tasks. The loss typically represents a cumulative penalty for predicted tokens, which are implicitly considered to have equal importance. Nevertheless, our studies have shown evidence that when considering different training data distributions, we cannot directly model the relationship between downstream capability and computation or token loss. To bridge the gap between validation loss and downstream task capabilities, in this work, we introduce Capability Salience Vector, which decomposes the overall loss and assigns different importance weights to tokens to assess a specific meta-capability, aligning the validation loss with downstream task performance in terms of the model’s capabilities. Experiments on various popular benchmarks demonstrate that our proposed Capability Salience Vector could significantly improve the predictability of language model performance on downstream tasks.
Qiming Ge, Shuhao Xing, Songyang Gao, Yunhua Zhou, Yicheng Zou, Songyang Zhang 0001, Zhi Chen 0006, Hang Yan 0001, Qi Zhang 0001, Qipeng Guo, Kai Chen 0026
ACL (1)7
2025 LESA: Learnable LLM Layer Scaling-Up
abstract
Training Large Language Models (LLMs) from scratch requires immense computational resources, making it prohibitively expensive.Model scaling-up offers a promising solution by leveraging the parameters of smaller models to create larger ones.However, existing depth scaling-up methods rely on empirical heuristic rules for layer duplication, which result in poorer initialization and slower convergence during continual pre-training.We propose LESA, a novel learnable method for depth scaling-up.By concatenating parameters from each layer and applying Singular Value Decomposition, we uncover latent patterns between layers, suggesting that inter-layer parameters can be learned.LESA uses a neural network to predict the parameters inserted between adjacent layers, enabling better initialization and faster training.Experiments show that LESA outperforms existing baselines, achieving superior performance with less than half the computational cost during continual pre-training.Extensive analyses demonstrate its effectiveness across different model sizes and tasks. 1
Zouying Cao, Xinbei Ma, Yao Yao 0008, Zhi Chen 0006, Libo Qin 0001, Hai Zhao 0001
ACL (1)5
2025 Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
abstract
Large Vision-Language Models (LVLMs) have achieved significant success in multimodal tasks, with multimodal chain-of-thought (MCoT) further enhancing performance and interpretability. Recent MCoT methods fall into two categories: (i) Textual-MCoT (T-MCoT), which takes multimodal input and produces textual output; and (ii) Interleaved-MCoT (I-MCoT), which generates interleaved image-text outputs. Despite advances in both approaches, the mechanisms driving these improvements are not fully understood. To fill this gap, we first reveal that MCoT boosts LVLMs by incorporating $\textit{visual thoughts}$, which convey image information to the reasoning process regardless of the MCoT format, depending only on clarity and conciseness of expression. Furthermore, to explore visual thoughts systematically, we define four distinct forms of visual thought expressions and analyze them comprehensively. Our findings demonstrate that these forms differ in clarity and conciseness, yielding varying levels of MCoT improvement. Additionally, we explore the internal nature of visual thoughts, finding that visual thoughts serve as intermediaries between the input image and reasoning to deeper transformer layers, enabling more advanced visual information transmission. We hope that the visual thoughts can inspire further breakthroughs for future MCoT research.
Zihui Cheng, Qiguang Chen, Xiao Xu 0005, Jiaqi Wang 0012, Weiyun Wang, Hao Fei 0003, Yidong Wang 0003, Alex Jinpeng Wang, Zhi Chen 0006, Wanxiang Che, Libo Qin 0001
NeurIPS9
2025 Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal Activations
abstract
Instruction tuning improves the ability of large language models (LLMs) to follow diverse human instructions, but achieving strong performance on specific target tasks remains challenging. A critical bottleneck is selecting the most relevant data to maximize task-specific performance. Existing data selection approaches include unstable influence-based methods and more stable distribution alignment methods, the latter of which critically rely on the underlying sample representation. In practice, most distribution alignment methods, from shallow features (e.g., BM25) to neural embeddings (e.g., BGE, LLM2Vec), may fail to capture how the model internally processes samples. To bridge this gap, we adopt a model-centric strategy in which each sample is represented by its neuronal activation pattern in the model, directly reflecting internal computation. However, directly using raw neuron activations leads to spurious similarity between unrelated samples due to neuron polysemanticity, where a single neuron may respond to multiple, unrelated concepts. To address this, we employ sparse autoencoders to disentangle polysemantic activations into sparse, monosemantic representations, and introduce a dedicated similarity metric for this space to better identify task-relevant data. Comprehensive experiments across multiple instruction datasets, models, tasks, and selection ratios show that our approach consistently outperforms existing data selection baselines in both stability and task-specific performance.
Gonghu Shang, Zhi Chen 0006, Libo Qin 0001, Yijie Luo, Hongshen Xu, Shuai Fan 0005, Kai Yu 0004, Lu Chen 0002
NeurIPS3
2024 M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
abstract
Multi-modal Chain-of-Thought (MCoT) requires models to leverage knowledge from both textual and visual modalities for step-bystep reasoning, which gains increasing attention.Nevertheless, the current MCoT benchmark still faces some challenges: (1) absence of visual modal reasoning, (2) single-step visual modal reasoning, and (3) Domain missing, thereby hindering the development of MCoT.Motivated by this, we introduce a novel benchmark (M 3 CoT) to address the above challenges, advancing the multi-domain, multi-step, and multi-modal CoT.Additionally, we conduct a thorough evaluation involving abundant MCoT approaches on Vision Large Language Models (VLLMs).In addition, we highlight that the current VLLMs still struggle to correctly reason in M 3 CoT and there remains a large gap between existing VLLMs and human performance in M 3 CoT, despite their superior results on previous MCoT benchmarks.To our knowledge, we take the first meaningful step toward the multi-domain, multi-step, and multi-modal scenario in MCoT.We hope that M 3 CoT can serve as a valuable resource, providing a pioneering foundation in multi-domain, multi-step, multi-modal chain-of-thought research.
Qiguang Chen, Libo Qin 0001, Zhi Chen 0006, Xiao Xu 0005, Wanxiang Che
ACL (1)4
2024 Linear Alignment: A Closed-form Solution for Aligning Human Preferences without Tuning and Feedback
abstract
The success of AI assistants based on Language Models (LLMs) hinges on Reinforcement Learning from Human Feedback (RLHF) to comprehend and align with user intentions. However, traditional alignment algorithms, such as PPO, are hampered by complex annotation and training requirements. This reliance limits the applicability of RLHF and hinders the development of professional assistants tailored to diverse human preferences. In this work, we introduce Linear Alignment, a novel algorithm that aligns language models with human preferences in one single inference step, eliminating the reliance on data annotation and model training. Linear alignment incorporates a new parameterization for policy optimization under divergence constraints, which enables the extraction of optimal policy in a closed-form manner and facilitates the direct estimation of the aligned response. Extensive experiments on both general and personalized preference datasets demonstrate that linear alignment significantly enhances the performance and efficiency of LLM alignment across diverse scenarios.
Songyang Gao, Qiming Ge, Shihan Dou, Junjie Ye 0005, Xiao Wang 0001, Yicheng Zou, Zhi Chen 0006, Hang Yan 0001, Qi Zhang 0001, Dahua Lin
ICML9
2024 What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
abstract
Recently, rapid advancements in Multi-Modal In-Context Learning (MM-ICL) have achieved notable success, which is capable of achieving superior performance across various tasks without requiring additional parameter tuning. However, the underlying rules for the effectiveness of MM-ICL remain under-explored. To fill this gap, this work aims to investigate the research question: "_What factors affect the performance of MM-ICL?_" To this end, we investigate extensive experiments on the three core steps of MM-ICL including demonstration retrieval, demonstration ordering, and prompt construction using 6 vision large language models and 20 strategies. Our findings highlight (1) the necessity of a multi-modal retriever for demonstration retrieval, (2) the importance of intra-demonstration ordering over inter-demonstration ordering, and (3) the enhancement of task comprehension through introductory instructions in prompts. We hope this study can serve as a foundational guide for optimizing MM-ICL strategies in future research.
Libo Qin 0001, Qiguang Chen, Hao Fei 0003, Zhi Chen 0006, Min Li 0007, Wanxiang Che
NeurIPS4
2023 OPAL: Ontology-Aware Pretrained Language Model for End-to-End Task-Oriented Dialogue
abstract
Abstract This paper presents an ontology-aware pretrained language model (OPAL) for end-to-end task-oriented dialogue (TOD). Unlike chit-chat dialogue models, task-oriented dialogue models fulfill at least two task-specific modules: Dialogue state tracker (DST) and response generator (RG). The dialogue state consists of the domain-slot-value triples, which are regarded as the user’s constraints to search the domain-related databases. The large-scale task-oriented dialogue data with the annotated structured dialogue state usually are inaccessible. It prevents the development of the pretrained language model for the task-oriented dialogue. We propose a simple yet effective pretraining method to alleviate this problem, which consists of two pretraining phases. The first phase is to pretrain on large-scale contextual text data, where the structured information of the text is extracted by the information extracting tool. To bridge the gap between the pretraining method and downstream tasks, we design two pretraining tasks: ontology-like triple recovery and next-text generation, which simulates the DST and RG, respectively. The second phase is to fine-tune the pretrained model on the TOD data. The experimental results show that our proposed method achieves an exciting boost and obtains competitive performance even without any TOD data on CamRest676 and MultiWOZ benchmarks.
Zhi Chen 0006, Yuncong Liu, Lu Chen 0002, Su Zhu, Mengyue Wu, Kai Yu 0004
Trans. Assoc. Comput. Linguistics1
2022 AdapterShare: Task Correlation Modeling with Adapter Differentiation
abstract
Thanks to the development of pre-trained language models, multitask learning (MTL) methods have achieved great success in natural language understanding.However, current MTL methods pay more attention to task selection or model design to fuse as much knowledge as possible, while the intrinsic task correlation is often neglected.It is important to learn sharing strategies among multiple tasks rather than sharing everything.In this paper, we propose AdapterShare, an adapter differentiation method to explicitly model task correlation among multiple tasks.AdapterShare is automatically learned based on the gradients on tiny held-out validation data.Compared to single-task learning and fully shared MTL methods, our proposed method obtains obvious performance improvements.Compared to the existing MTL method AdapterFusion, AdapterShare achieves an absolute average improvement of 1.90 points on five dialogue understanding tasks and 2.33 points on NLU tasks.Our implementation is available at https:// github.com/microsoft/ContextualSP.
Zhi Chen 0006, Bei Chen 0008, Lu Chen 0002, Kai Yu 0004, Jian-Guang Lou
EMNLP1
2022 UniDU: Towards A Unified Generative Dialogue Understanding Framework
abstract
With the development of pre-trained language models, remarkable success has been witnessed in dialogue understanding (DU).However, current DU approaches usually employ independent models for each distinct DU task without considering shared knowledge across different DU tasks.In this paper, we propose a unified generative dialogue understanding framework, named UniDU, to achieve effective information exchange across diverse DU tasks.Here, we reformulate all DU tasks into a unified promptbased generative model paradigm.More importantly, a novel model-agnostic multi-task training strategy (MATS) is introduced to dynamically adapt the weights of diverse tasks for best knowledge sharing during training, based on the nature and available data of each task.Experiments on ten DU datasets covering five fundamental DU tasks show that the proposed UniDU framework largely outperforms task-specific well-designed methods on all tasks.MATS also reveals the knowledgesharing structure of these tasks.Finally, UniDU obtains promising performance in the unseen dialogue domain, showing the great potential for generalization.
Zhi Chen 0006, Lu Chen 0002, Bei Chen 0008, Libo Qin 0001, Yuncong Liu, Su Zhu, Jian-Guang Lou, Kai Yu 0004
SIGDIAL1
2021 LGESQL: Line Graph Enhanced Text-to-SQL Model with Mixed Local and Non-Local Relations
abstract
Ruisheng Cao, Lu Chen, Zhi Chen, Yanbin Zhao, Su Zhu, Kai Yu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ruisheng Cao, Lu Chen 0002, Zhi Chen 0006, Yanbin Zhao, Su Zhu, Kai Yu 0004
ACL/IJCNLP (1)3
2021 ShadowGNN: Graph Projection Neural Network for Text-to-SQL Parser
abstract
Zhi Chen, Lu Chen, Yanbin Zhao, Ruisheng Cao, Zihan Xu, Su Zhu, Kai Yu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Zhi Chen 0006, Lu Chen 0002, Yanbin Zhao, Ruisheng Cao, Su Zhu, Kai Yu 0004
NAACL-HLT1
2021 Few-Shot NLU with Vector Projection Distance and Abstract Triangular CRF
Su Zhu, Lu Chen 0002, Ruisheng Cao, Zhi Chen 0006, Qingliang Miao, Kai Yu 0004
NLPCC (1)4
2020 Semi-Supervised Text Simplification with Back-Translation and Asymmetric Denoising Autoencoders
abstract
Text simplification (TS) rephrases long sentences into simplified variants while preserving inherent semantics. Traditional sequence-to-sequence models heavily rely on the quantity and quality of parallel sentences, which limits their applicability in different languages and domains. This work investigates how to leverage large amounts of unpaired corpora in TS task. We adopt the back-translation architecture in unsupervised machine translation (NMT), including denoising autoencoders for language modeling and automatic generation of parallel data by iterative back-translation. However, it is non-trivial to generate appropriate complex-simple pair if we directly treat the set of simple and complex corpora as two different languages, since the two types of sentences are quite similar and it is hard for the model to capture the characteristics in different types of sentences. To tackle this problem, we propose asymmetric denoising methods for sentences with separate complexity. When modeling simple and complex sentences with autoencoders, we introduce different types of noise into the training process. Such a method can significantly improve the simplification performance. Our model can be trained in both unsupervised and semi-supervised manner. Automatic and human evaluations show that our unsupervised model outperforms the previous systems, and with limited supervision, our model can perform competitively with multiple state-of-the-art simplification systems.
Yanbin Zhao, Lu Chen 0002, Zhi Chen 0006, Kai Yu 0004
AAAI3
2020 Neural Graph Matching Networks for Chinese Short Text Matching
abstract
Chinese short text matching usually employs word sequences rather than character sequences to get better performance.However, Chinese word segmentation can be erroneous, ambiguous or inconsistent, which consequently hurts the final matching performance.To address this problem, we propose neural graph matching networks, a novel sentence matching framework capable of dealing with multi-granular input information.Instead of a character sequence or a single word sequence, paired word lattices formed from multiple word segmentation hypotheses are used as input and the model learns a graph representation according to an attentive graph matching mechanism.Experiments on two Chinese datasets show that our models outperform the state-of-the-art short text matching models.
Lu Chen 0002, Yanbin Zhao, Boer Lyu, Lesheng Jin, Zhi Chen 0006, Su Zhu, Kai Yu 0004
ACL5
2020 Line Graph Enhanced AMR-to-Text Generation with Mix-Order Graph Attention Networks
abstract
Efficient structure encoding for graphs with labeled edges is an important yet challenging point in many graph-based models.This work focuses on AMR-to-text generation -A graph-to-sequence task aiming to recover natural language from Abstract Meaning Representations (AMR).Existing graph-to-sequence approaches generally utilize graph neural networks as their encoders, which have two limitations: 1) The message propagation process in AMR graphs is only guided by the firstorder adjacency information.2) The relationships between labeled edges are not fully considered.In this work, we propose a novel graph encoding framework which can effectively explore the edge relations.We also adopt graph attention networks with higherorder neighborhood information to encode the rich structure in AMR graphs.Experiment results show that our approach obtains new state-of-the-art performance on English AMR benchmark datasets.The ablation analyses also demonstrate that both edge relations and higher-order information are beneficial to graph-to-sequence modeling.
Yanbin Zhao, Lu Chen 0002, Zhi Chen 0006, Ruisheng Cao, Su Zhu, Kai Yu 0004
ACL3
2020 Memory Attention Neural Network for Multi-domain Dialogue State Tracking
Zhi Chen 0006, Lu Chen 0002, Su Zhu, Kai Yu 0004
NLPCC (1)2
2020 Distributed Structured Actor-Critic Reinforcement Learning for Universal Dialogue Management
abstract
Traditional dialogue policy needs to be trained independently for each dialogue task. In this work, we aim to solve a collection of independent dialogue tasks using a unified dialogue agent. The unified policy is parallelly trained using the conversation data from all of the distributed dialogue tasks. However, there are two key challenges:(1) the design of a unified dialogue model to adapt to different dialogue tasks; (2) finding a robust reinforcement learning method to keep the efficiency and the stability of the training process. Here we propose a novel structured actor-critic approach to implement structured deep reinforcement learning (DRL), which not only can learn parallelly from data of different dialogue tasks but also achieves stable and sample-efficient learning. We demonstrate the effectiveness of the proposed approach on 18 tasks of PyDial benchmark. The results show that our method is able to achieve state-of-the-art performance.
Zhi Chen 0006, Lu Chen 0002, Kai Yu 0004
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 AgentGraph: Toward Universal Dialogue Management With Structured Deep Reinforcement Learning
abstract
Dialogue policy plays an important role in task-oriented spoken dialogue systems. It determines how to respond to users. The recently proposed deep reinforcement learning (DRL) approaches have been used for policy optimization. However, these deep models are still challenging for two reasons: first, many DRL-based policies are not sample efficient; and second, most models do not have the capability of policy transfer between different domains. In this paper, we propose a universal framework, AgentGraph, to tackle these two problems. The proposed AgentGraph is the combination of graph neural network (GNN) based architecture and DRL-based algorithm. It can be regarded as one of the multi-agent reinforcement learning approaches. Each agent corresponds to a node in a graph, which is defined according to the dialogue domain ontology. When making a decision, each agent can communicate with its neighbors on the graph. Under AgentGraph framework, we further propose dual GNN-based dialogue policy, which implicitly decomposes the decision in each turn into a high-level global decision and a low-level local decision. Experiments show that AgentGraph models significantly outperform traditional reinforcement learning approaches on most of the 18 tasks of the PyDial benchmark. Moreover, when transferred from the source task to a target task, these models not only have acceptable initial performance but also converge much faster on the target task.
Lu Chen 0002, Zhi Chen 0006, Bowen Tan, Sishan Long, Milica Gasic, Kai Yu 0004
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Policy Adaptation for Deep Reinforcement Learning-Based Dialogue Management
abstract
Policy optimization is the core part of statistical dialogue management. Deep reinforcement learning has been successfully used for dialogue policy optimization for a static pre-defined domain. However, when the domain changes dynamically, e.g. a new previously unseen concept (or slot) which can be then used as a database search constraint is added, or the policy for one domain is transferred to another domain, the dialogue state space and action sets both will change. Therefore, the model structures for different domains have to be different. This makes dialogue policy adaptation/transfer challenging. Here a multi -agent dialogue policy (MADP) is proposed to tackle these problems. MADP consists of some slot-dependent agents (S-Agents) and a slot-independent agent (G-Agent). S-Agents have shared parameters in addition to private parameters for each one. During policy transfer, the shared parameters in S-Agents and the parameters in G-Agent can be directly transferred to the agents in extended/new domain. Simulation experiments showed that MADP can significantly speed up the policy learning and facilitate policy adaptation.
Lu Chen 0002, Zhi Chen 0006, Bowen Tan, Milica Gasic, Kai Yu 0004
ICASSP3