Libo Qin 0001

dblp:213/9481 · DBLP profile ↗
← Back
77ranked-venue papers
24as first author
68since 2021 · last 2026
0000-0002-3619-675XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 61 · 20 first-author · 52 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 9 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Significant Gains in Reasoning Efficiency in Large Language Models
abstract
Recent advancements in large language models (LLMs) have greatly improved their ability to perform complex reasoning tasks through Long Chain-of-Thought (CoT). However, this approach often results in substantial redundancy, impairing computational efficiency and causing significant delays in real-time applications. To improve efficiency, current methods often rely on human-defined difficulty priors, which do not align with the LLM's self-awared difficulty, leading to inefficiencies. In this paper, we introduce the Dynamic Reasoning-Boundary Self-Awareness Framework (DR. SAF), which enables LLMs to dynamically assess and adjust their reasoning depth in response to problem complexity. DR. SAF integrates three key components: Boundary Self-Awareness Alignment, Adaptive Reward Management, and a Boundary Preservation Mechanism. These components allow models to optimize their reasoning processes, balancing efficiency and accuracy without compromising performance. Our experimental results demonstrate that DR. SAF achieves a 49.27% reduction in total response tokens with minimal loss in accuracy. The framework also delivers a 6.59x gain in token efficiency and a 5x reduction in training time, making it well-suited to resource-limited settings. During extreme training, DR. SAF can even surpass traditional instruction-based models in token efficiency with more than 16% accuracy improvement.
Qiguang Chen, Dengyun Peng, Huikang Su, Jiannan Guan, Libo Qin 0001, Wanxiang Che
AAAI6
2026 Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
abstract
Large Language Models (LLMs) excel in reasoning tasks requiring a single correct answer, but they perform poorly in multi-solution tasks that require generating comprehensive and diverse answers. We attribute this limitation to reasoning overconfidence: a tendency to express undue certainty in an incomplete solution set. To examine the effect, we introduce MuSoBench, a benchmark of multi-solution problems. Experiments show that the conventional short chain-of-thought (Short-CoT) prompting paradigm exhibits pronounced overconfidence, whereas the emerging long chain-of-thought (Long-CoT) approach mitigates it through iterative exploration and self-reflection. We further characterise observable behaviours and influential factors. To probe the underlying cause, we propose the cognitive-rigidity hypothesis, which posits that overconfidence arises when the reasoning process prematurely converges on a narrow set of thought paths. An attention-entropy analysis offers preliminary support for this view. These findings provide tools for assessing the completeness of LLM reasoning and highlight the need to move evaluation beyond single-answer accuracy toward comprehensive exploration.
Jiannan Guan, Qiguang Chen, Libo Qin 0001, Dengyun Peng, Liangyu Huo, Wanxiang Che
AAAI3
2026 Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
abstract
Recently, Interleaved-modal Chain-of-Thought (ICoT) reasoning has achieved remarkable success by leveraging both multimodal inputs and outputs, attracting increasing attention. While achieving promising performance, current ICoT methods still suffer from two major limitations: (1) Static Visual Thought Positioning, which statically inserts visual information at fixed steps, resulting in inefficient and inflexible reasoning; and (2) Broken Visual Thought Representation, which involves discontinuous and semantically incoherent visual tokens. To address these limitations, we introduce Interleaved-modal Chain-of-Thought reasoning with Dynamic and Precise Visual Thoughts (DaP-ICoT), which incorporates two key components: (1) Dynamic Visual Thought Integration adaptively introduces visual inputs based on reasoning needs, reducing redundancy and improving efficiency. (2) Precise Visual Thought Guidance ensures visual semantically coherent and contextually aligned representations. Experiments across multiple benchmarks and models demonstrate that DaP-ICoT achieves state-of-the-art performance. In addition, DaP-ICoT significantly reduces the number of inserted images, leading to a 72.6% decrease in token consumption, enabling more efficient ICoT reasoning.
Yongheng Zhang 0001, Qiguang Chen, Libo Qin 0001
AAAI6
2026 OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Models
abstract
Qiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu, Yi Yang, Yizhuo Li, Jingqi Tong, Xiachong Feng, Libo Qin, Wanxiang Che. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qiguang Chen, Chengyu Luan, Qiming Yu, Yizhuo Li 0007, Jingqi Tong, Xiachong Feng, Libo Qin 0001, Wanxiang Che
ACL (1)9
2026 Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play
abstract
Xiachong Feng, Deyi Yin, Xiaocheng Feng, Yi Jiang, Libo Qin, Yangfan Ye, Lei Huang, Weitao Ma, Qiming Li, Yuxuan Gu, Bing Qin, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiachong Feng, Deyi Yin, Libo Qin 0001, Yangfan Ye, Lei Huang 0021, Weitao Ma, Yuxuan Gu 0004, Bing Qin 0001, Lingpeng Kong
ACL (1)5
2026 Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
abstract
Qiming Li, Xiaocheng Feng, Yixuan Ma, Ruihan Chen, Zihe Tong, Zekai Ye, Xiachong Feng, Libo Qin, Haoyu Ren, Kun Chen, Yunfei Lu, Dandan Tu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yixuan Ma, Ruihan Chen 0001, Zihe Tong, Zekai Ye, Xiachong Feng, Libo Qin 0001, Yunfei Lu, Dandan Tu, Bing Qin 0001
ACL (1)8
2026 Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
abstract
Chenyuan Zhang, Qiguang Chen, Xie Chen, Zhuotao Tian, Bowen Xing, Meishan Zhang, Libo Qin, Baotian Hu, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qiguang Chen, Xie Chen 0001, Zhuotao Tian, Meishan Zhang, Libo Qin 0001, Baotian Hu, Min Zhang 0005
ACL (1)7
2026 Large language models meet NLP: a survey
abstract
Abstract While large language models (LLMs) like ChatGPT have shown impressive capabilities in Natural Language Processing (NLP) tasks, a systematic investigation of their potential in this field remains largely unexplored. This study aims to address this gap by exploring the following questions. (1) How are LLMs currently applied to NLP tasks in the literature ? (2) Have traditional NLP tasks already been solved with LLMs ? (3) What is the future of the LLMs for NLP ? To answer these questions, we take the first step to provide a comprehensive overview of LLMs in NLP. Specifically, we first introduce a unified taxonomy including (1) parameter-frozen paradigm and (2) parameter-tuning paradigm to offer a unified perspective for understanding the current progress of LLMs in NLP. Furthermore, we summarize the new frontiers and the corresponding challenges, aiming to inspire further groundbreaking advancements. We hope this work offers valuable insights into {the potential and limitations} of LLMs, while also serving as a practical guide for building effective LLMs in NLP.
Libo Qin 0001, Qiguang Chen, Xiachong Feng, Yang Wu 0010, Yongheng Zhang 0001, Min Li 0007, Wanxiang Che, Philip S. Yu
Frontiers Comput. Sci.1
2026 DXA-Net: Dual-Task Cross-Lingual Alignment Network for Zero-Shot Cross-Lingual Spoken Language Understanding
abstract
The state-of-the-art zero-shot cross-lingual spoken language understanding (SLU) model utilizes cross-lingual unsupervised contrastive learning to achieve multilingual semantics alignment. While existing methods have achieved promising results, they still have two issues limiting cross-lingual knowledge transfer: (1) dual-task correlative knowledge is not explicitly modeled and transferred to target languages; (2) the semantics differences among samples are ignored, and the contrastive semantics knowledge is not transferred to target languages. In this paper, we propose a dual-task cross-lingual alignment network (DXA-Net), which makes the first attempt to tackle zero-shot cross-lingual SLU based on the prompt-tuning paradigm. To solve the first issue, we propose the co-guiding prompt, which allows the model to conditionally generate one task's label based on another one's. To solve the second issue, we propose the intent/slot contrastive prompt to teach the model to discriminate whether a pair of samples have the same or similar labels. Additionally, we propose multilingual semantics contrastive prompt to enhance multilingual semantics alignment. Experiments on the benchmark show that our model achieves new state-of-the-art performance on nine languages.
Libo Qin 0001, Zhihong Zhu 0001, Zhou Yu 0005, Ivor W. Tsang
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 S HARING B EYOND D ECISION : Deep Collaboration between Large Language Models via Representation Ensemble
abstract
Abstract Large Language Models (LLMs) exhibit unique strengths arising from differences in model architecture, training data, and strategies. Ensemble learning has been explored to leverage these complementary strengths through decision-level sharing (i.e.,Decision Ensemble), which combines the predictions from multiple LLMs. However, such methods integrate only shallow decisions and overlook the exchange of deeper levels of information within the internal representations of LLMs, such as problem understanding, world knowledge, and latent reasoning patterns. In this work, we propose Representation Ensemble (RISE), a novel ensemble framework that enables cross-LLM representation sharing for richer information exchange. To address challenges of representation-level interaction caused by layer misalignment and latent-space incompatibility across LLMs, we introduce a representation alignment method based on relational similarity measures and an orthogonal latent-space transformation. Experimental results show that (1) RISE achieves performance competitive with existing decision ensemble methods, and (2) RISE is strongly complementary to decision ensemble, with their combination boosting collaboration gains by 14%–41%. Finally, we further compare ensemble of small LLMs to a single larger LLM and to model merging and composition approaches, and find that ensemble learning consistently generalizes well without additional training.
Yichong Huang, Jinlan Fu, Xiachong Feng, Baohang Li, Zekai Ye, Libo Qin 0001, Hao Fei 0001, See-Kiong Ng, Bing Qin 0001
Trans. Assoc. Comput. Linguistics7
2025 Divide-Solve-Combine: An Interpretable and Accurate Prompting Framework for Zero-shot Multi-Intent Detection
abstract
Zero-shot multi-intent detection is capable of capturing multiple intents within a single utterance without any training data, which gains increasing attention. Building on the success of large language models (LLM), dominant approaches in the literature explore prompting techniques to enable zero-shot multi-intent detection. While significant advancements have been witnessed, the existing prompting approaches still face two major issues: lacking explicit reasoning and lacking interpretability. Therefore, in this paper, we introduce a Divide-Solve-Combine Prompting (DSCP) to address the above issues. Specifically, DSCP explicitly decomposes multi-intent detection into three components including (1) single-intent division prompting is utilized to decompose an input query into distinct sub-sentences, each containing a single intent; (2) intent-by-intent solution prompting is applied to solve each sub-sentence recurrently; and (3) multi-intent combination prompting is employed for combining each sub-sentence result to obtain the final multi-intent result. By decomposition, DSCP allows the model to track the explicit reasoning process and improve the interpretability. In addition, we propose an interactive divide-solve-combine prompting (Inter-DSCP) to naturally capture the interaction capabilities of large language models. Experimental results on two standard multi-intent benchmarks (i.e., MixATIS and MixSNIPS) reveal that both DSCP and Inter-DSCP obtain substantial improvements over baselines, achieving superior performance and higher interpretability.
Libo Qin 0001, Qiguang Chen, Jingxuan Zhou, Hao Fei 0003, Wanxiang Che, Min Li 0007
AAAI1
2025 CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) have recently demonstrated amazing success in multi-modal tasks, including advancements in Multi-modal Chain-of-Thought (MCoT) reasoning. Despite these successes, current benchmarks still follow a traditional paradigm with multi-modal input and text-modal output, which leads to significant drawbacks such as missing visual operations and vague expressions. Motivated by this, we introduce a novel Chain of Multi-modal Thought (CoMT) benchmark to address these limitations. Different from the traditional MCoT benchmark, CoMT requires both multi-modal input and multi-modal reasoning output, aiming to mimic human-like reasoning that inherently integrates visual operation. Specifically, CoMT consists of four categories: (1) Visual Creation, (2) Visual Deletion, (3) Visual Update, and (4) Visual Selection to comprehensively explore complex visual operations and concise expression in real scenarios. We evaluate various LVLMs and strategies on CoMT, revealing some key insights into the capabilities and limitations of the current approaches. We hope that CoMT can inspire more research on introducing multi-modal generation into the reasoning process.
Zihui Cheng, Qiguang Chen, Hao Fei 0003, Wanxiang Che, Min Li 0007, Libo Qin 0001
AAAI8
2025 EXCGEC: A Benchmark for Edit-Wise Explainable Chinese Grammatical Error Correction
abstract
Existing studies explore the explainability of Grammatical Error Correction (GEC) in a limited scenario, where they ignore the interaction between corrections and explanations and have not established a corresponding comprehensive benchmark. To bridge the gap, this paper first introduces the task of EXplainable GEC (EXGEC), which focuses on the integral role of correction and explanation tasks. To facilitate the task, we propose EXCGEC, a tailored benchmark for Chinese EXGEC consisting of 8,216 explanation-augmented samples featuring the design of hybrid edit-wise explanations. We then benchmark several series of LLMs in multi-task learning settings, including post-explaining and pre-explaining. To promote the development of the task, we also build a comprehensive evaluation suite by leveraging existing automatic metrics and conducting human evaluation experiments to demonstrate the human consistency of the automatic metrics for free-text explanations. Our experiments reveal the effectiveness of evaluating free-text explanations using traditional metrics like METEOR and ROUGE, and the inferior performance of multi-task models compared to the pipeline solution, indicating its challenges to establish positive effects in learning both tasks.
Jingheng Ye, Shang Qin, Xuxin Cheng, Libo Qin 0001, Hai-Tao Zheng 0002, Ying Shen 0001, Peng Xing, Zishan Xu
AAAI5
2025 CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models
abstract
Investigating hallucination issues in large language models (LLMs) within cross-lingual and cross-modal scenarios can greatly advance the large-scale deployment in real-world applications.Nevertheless, the current studies are limited to a single scenario, either cross-lingual or cross-modal, leaving a gap in the exploration of hallucinations in the joint cross-lingual and cross-modal scenarios.Motivated by this, we introduce a novel joint Cross-lingual and Crossmodal Hallucinations benchmark (CCHall) to fill this gap.Specifically, CCHall simultaneously incorporates both cross-lingual and cross-modal hallucination scenarios, which can be used to assess the cross-lingual and crossmodal capabilities of LLMs.Furthermore, we conduct a comprehensive evaluation on CCHall, exploring both mainstream opensource and closed-source LLMs.The experimental results highlight that current LLMs still struggle with CCHall.We hope CCHall can serve as a valuable resource to assess LLMs in joint cross-lingual and cross-modal scenarios.
Yongheng Zhang 0001, Ruoxi Zhou, Qiguang Chen, Hao Fei 0003, Wenpeng Lu, Libo Qin 0001
ACL (1)7
2025 What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best Practices
abstract
Zhi Chen, Qiguang Chen, Libo Qin, Qipeng Guo, Haijun Lv, Yicheng Zou, Hang Yan, Kai Chen, Dahua Lin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhi Chen 0006, Qiguang Chen, Libo Qin 0001, Qipeng Guo, Haijun Lv, Yicheng Zou, Hang Yan 0001, Kai Chen 0026, Dahua Lin
ACL (1)3
2025 CSTree-SRI: Introspection-Driven Cognitive Semantic Tree for Multi-Turn Question Answering over Extra-Long Contexts
abstract
Zhaowen Wang, Xiang Wei, Kangshao Du, Yiting Zhang, Libo Qin, Yingjie Xia, Li Kuang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kangshao Du, Libo Qin 0001, Yingjie Xia, Li Kuang
ACL (1)5
2025 LESA: Learnable LLM Layer Scaling-Up
abstract
Training Large Language Models (LLMs) from scratch requires immense computational resources, making it prohibitively expensive.Model scaling-up offers a promising solution by leveraging the parameters of smaller models to create larger ones.However, existing depth scaling-up methods rely on empirical heuristic rules for layer duplication, which result in poorer initialization and slower convergence during continual pre-training.We propose LESA, a novel learnable method for depth scaling-up.By concatenating parameters from each layer and applying Singular Value Decomposition, we uncover latent patterns between layers, suggesting that inter-layer parameters can be learned.LESA uses a neural network to predict the parameters inserted between adjacent layers, enabling better initialization and faster training.Experiments show that LESA outperforms existing baselines, achieving superior performance with less than half the computational cost during continual pre-training.Extensive analyses demonstrate its effectiveness across different model sizes and tasks. 1
Zouying Cao, Xinbei Ma, Yao Yao 0008, Zhi Chen 0006, Libo Qin 0001, Hai Zhao 0001
ACL (1)6
2025 CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning
abstract
Yangfan Ye, Xiaocheng Feng, Zekun Yuan, Xiachong Feng, Libo Qin, Lei Huang, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Xiaohui Yan, Duyu Tang, Dandan Tu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yangfan Ye, Zekun Yuan, Xiachong Feng, Libo Qin 0001, Lei Huang 0021, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001
ACL (1)5
2025 CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention
abstract
Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal abilities but remain prone to multilingual object hallucination, with a higher likelihood of generating responses inconsistent with the visual input when utilizing queries in non-English languages compared to English. Most existing approaches to address these rely on pretraining or fine-tuning, which are resource-intensive. In this paper, inspired by observing the disparities in cross-modal attention patterns across languages, we propose Cross-Lingual Attention Intervention for Mitigating multilingual object hallucination (CLAIM) in LVLMs, a novel near training-free method by aligning attention patterns. CLAIM first identifies language-specific cross-modal attention heads, then estimates language shift vectors from English to the target language, and finally intervenes in the attention outputs during inference to facilitate cross-lingual visual perception capability alignment. Extensive experiments demonstrate that CLAIM achieves an average improvement of 13.56% (up to 30% in Spanish) on the POPE and 21.75% on the hallucination subsets of the MME benchmark across various languages. Further analysis reveals that multilingual attention divergence is most prominent in intermediate layers, highlighting their critical role in multilingual scenarios.
Zekai Ye, Libo Qin 0001, Yichong Huang, Baohang Li, Kui Jiang, Yang Xiang 0003, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001
ACL (1)4
2025 An Efficient Task-Oriented Dialogue Policy: Evolutionary Reinforcement Learning Injected by Elite Individuals
abstract
Deep Reinforcement Learning (DRL) is widely used in task-oriented dialogue systems to optimize dialogue policy, but it struggles to balance exploration and exploitation due to the high dimensionality of state and action spaces.This challenge often results in local optima or poor convergence.Evolutionary Algorithms (EAs) have been proven to effectively explore the solution space of neural networks by maintaining population diversity.Inspired by this, we innovatively combine the global search capabilities of EA with the local optimization of DRL to achieve a balance between exploration and exploitation.Nevertheless, the inherent flexibility of natural language in dialogue tasks complicates this direct integration, leading to prolonged evolutionary times.Thus, we further propose an elite individual injection mechanism to enhance EA's search efficiency by adaptively introducing best-performing individuals into the population.Experiments across four datasets show that our approach significantly improves the balance between exploration and exploitation, boosting performance.Moreover, the effectiveness of the EII mechanism in reducing exploration time has been demonstrated, achieving an efficient integration of EA and DRL on task-oriented dialogue policy tasks.
Libo Qin 0001, Shihan Wang 0001
ACL (1)3
2025 LAMA-AD: Label-Aware Multi-Agent Alzheimer's Disease Diagnosis with Counterfactual Reasoning
abstract
The early and accurate diagnosis of Alzheimer's Disease (AD) is essential for effective intervention, yet remains challenging due to the complexity of clinical data. While large language models (LLMs) have become a promising avenue for medical diagnostics, existing approaches often overlook the implicit medical knowledge embedded in diagnostic labels, limiting diagnostic reliability and interpretability. Motivated by this, we introduce Label-Aware Multi-Agent Alzheimer's Disease Diagnosis with Counterfactual Reasoning (LAMA-AD), a labelaware multi-agent framework that explicitly incorporates clinical priors into the diagnostic reasoning process and integrates both factual and counterfactual reasoning. Specifically, LAMA-AD consists of two components: (1) Label-Aware Multi-Agent Diagnosis, which deploys multiple independent analyst agents, each guided by distinct label-aware prior knowledge, to explore diverse diagnostic hypotheses in parallel. Each agent provides supporting facts for its reasoning process, thereby enhancing both the depth and robustness of the analysis. (2) Counterfactual Reasoningsystematically assesses conflicting hypotheses to yield stable, interpretable, and highly accurate diagnostic conclusions. Experimental verification confirms LAMA-AD's superiority over established methods. Analysis further shows it generates semantically rich, logically consistent explanations.
Chunlin Lu, Yongheng Zhang 0001, Libo Qin 0001
BIBM5
2025 Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits
abstract
The Myers-Briggs Type Indicator (MBTI) is one of the most influential personality theories reflecting individual differences in thinking, feeling, and behaving. MBTI personality detection has garnered considerable research interest and has evolved significantly over the years. However, this task tends to be overly optimistic, as it currently does not align well with the natural distribution of population personality traits. Specifically, the self-reported labels in existing datasets result in data quality issues and the hard labels fail to capture the full range of population personality distributions. In this paper, we identify the task by constructing MBTIBench, the first manually annotated MBTI personality detection dataset with soft labels, under the guidance of psychologists. Our experimental results confirm that soft labels can provide more benefits to other psychological tasks than hard labels. We highlight the polarized predictions and biases in LLMs as key directions for future research.
Bohan Li 0010, Jiannan Guan, Longxu Dou, Yunlong Feng, Dingzirui Wang, Yang Xu 0049, Enbo Wang, Qiguang Chen, Bichen Wang, Xiao Xu 0005, Libo Qin 0001, Qingfu Zhu, Wanxiang Che
COLING12
2025 CroPrompt: Cross-task Interactive Prompting for Zero-shot Spoken Language Understanding
abstract
Slot filling and intent detection are two highly correlated tasks in spoken language understanding (SLU). Recent SLU research attempts to explore zero-shot prompting techniques in large language models to alleviate the data scarcity problem. Nevertheless, the existing prompting work ignores the cross-task interaction information for SLU, which leads to sub-optimal performance. To solve this problem, we present the pioneering work of Cross-task Interactive Prompting (CroPrompt) for SLU, which enables the model to interactively leverage the information exchange across the correlated tasks in SLU. Additionally, we further introduce a multi-task self-consistency mechanism to mitigate the error propagation caused by the intent information injection. We conduct extensive experiments on the standard SLU benchmark and the results reveal that CroPrompt consistently outperforms the existing prompting approaches. In addition, the multi-task self-consistency mechanism can effectively ease the error propagation issue, thereby enhancing the performance. We hope this work can inspire more research on cross-task prompting for SLU.
Libo Qin 0001, Fuxuan Wei, Qiguang Chen, Jingxuan Zhou, Shijue Huang, Jiasheng Si, Wenpeng Lu, Wanxiang Che
ICASSP1
2025 MdCoT: Medical Diagnosis Chain-of-Thought with Self-Diagnostic Refinement for Alzheimer's Disease
abstract
Accurate diagnosis of Alzheimer’s disease is vital for effective treatment. Recently, a mainstream approach emerged that uses large language models (LLMs) to generate interpretable diagnosis. However, these methods face two significant challenges: (1) Incomplete Utilization of Raw Image Information: Converting images into text for LLMs input leads to the loss of the original image information. (2) Hallucination in LLMs: Diagnostic biases from LLMs hallucinations are neglected. To address these challenges, we propose a Medical Diagnosis Chain-of-Thought with Self-Diagnostic Refinement (MdCoT) framework. The MdCoT framework consists of two core modules: (1) Multimodal AD Diagnostic Chain-of-Thought: By including raw images as input, multimodal large language models (MLLMs) are allowed to capture fine-grained features, and (2) Self-Diagnostic Refinement: By guiding MLLMs to self-check and refine the diagnostic results, MdCoT achieves to mitigate hallucination. Experiments on widely used benchmarks demonstrate that MdCoT surpasses all baselines. Besides, extensive auxiliary experiments demonstrate the superiority of MdCoT.
Chunlin Lu, Yongheng Zhang 0001, Peng Wang 0168, Wenpeng Lu, Libo Qin 0001
ICME5
2025 Improving Consistency Identification in Task-oriented Dialogue Through Multi-Agent Collaboration
abstract
Consistency identification in task-oriented dialog (CI-ToD) typically consists of three sub-tasks: User Query Inconsistency (QI) identification, Dialogue History Inconsistency (HI) identification, and Knowledge Base Inconsistency (KBI) identification, which aim to determine inconsistent relationships between system response and user query, dialogue history, and knowledge base. Previous approaches focus on the exploration of deep learning models for CI-ToD. While these models achieve remarkable progress, they still rely on large amounts of labeled data, which is hard to achieve in real-world scenarios. Motivated by this, in the paper, we aim to explore large language models for CI-ToD, which do not require any training data. In addition, we further introduce a multi-agent collaboration framework (MAC-CIToD) to model the interaction across three sub-tasks in CI-ToD, including (1) Full Connection paradigm, (2) Cycle Connection paradigm, and (3) Central Connection paradigm, which effectively builds interaction across QI, HI, and KBI. Experiments on the standard benchmark reveal that our framework achieves superior performance. Additionally, we compare MAC-CIToD with the most advanced trained approaches and find that its zero-shot performance on most metrics even surpasses that of models after training on the CI-ToD dataset.
Ruoxi Zhou, Qiguang Chen, Xiao Xu 0005, Hao Fei 0003, Dagang Li 0001, Wanxiang Che, Libo Qin 0001
IJCAI9
2025 MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models
abstract
Multimodal planning capabilities refer to the ability to predict, reason, and design steps for task execution with multimodal context, which is essential for complex reasoning and decision-making across multiple steps. However, current benchmarks face two key challenges: (1) they cannot directly assess multimodal real-world planning capabilities, and (2) they lack constraints or implicit constraints across modalities. To address these issues, we introduce Multimodal Planning with Complex Constraints (MPCC), the first benchmark to systematically evaluate MLLMs' ability to handle multimodal constraints in planning. To address the first challenge, MPCC focuses on three real-world tasks: Flight Planning, Calendar Planning, and Meeting Planning. To solve the second challenge, we introduce complex constraints (e.g. budget, temporal, and spatial) in these tasks, with graded difficulty levels (EASY, MEDIUM, HARD) to separate constraint complexity from search space expansion. Experiments on 13 advanced MLLMs reveal significant challenges: closed-source models achieve only 21.3% feasible plans, while open-source models average below 11%. Additionally, we observe that MLLMs are highly sensitive to constraint complexity and that traditional multimodal prompting strategies fail in multi-constraint scenarios. Our work formalizes multimodal constraints in planning, provides a rigorous evaluation framework, and highlights the need for advancements in constraint-aware reasoning for real-world MLLM applications.
Yiyan Ji, Qiguang Chen, Chengyue Wu, Libo Qin 0001, Wanxiang Che
ACM Multimedia5
2025 ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
abstract
Video understanding plays a vital role in bridging low-level visual signals with high-level cognitive reasoning, and is fundamental to applications such as autonomous driving, embodied AI, and the broader pursuit of AGI. The rapid development of large language models (LLMs), particularly those utilizing Chain-of-Thought (CoT) technology, has significantly advanced video reasoning capabilities. However, current approaches primarily depend on textual information for reasoning, overlooking the visual modality in the actual video reasoning process. In contrast, humans naturally re-examine visual content while reasoning. Motivated by this, we introduce a novel video reasoning paradigm: Video-Text Interleaved CoT (ViTCoT), which facilitates more intuitive and cognitively aligned reasoning. To the end, first, we construct the Video-Text Interleaved Benchmark (ViTIB), which is created using MLLMs for key-video selection and manually verified. Furthermore, we extensively explore the potential of the ViTCoT paradigm in the video understanding field. Extensive experiments demonstrate that ViTCoT significantly enhances performance compared to the traditional text-only CoT paradigm and effectively activates more neuron values in MLLMs.
Yongheng Zhang 0001, Ruihan Tao, Qiguang Chen, Hao Fei 0001, Wanxiang Che, Libo Qin 0001
ACM Multimedia7
2025 Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring
abstract
Honglin Mu, Han He, Yuxin Zhou, Yunlong Feng, Yang Xu, Libo Qin, Xiaoming Shi, Zeming Liu, Xudong Han, Qi Shi, Qingfu Zhu, Wanxiang Che. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Honglin Mu, Han He, Yunlong Feng, Yang Xu 0049, Libo Qin 0001, Zeming Liu, Qi Shi 0002, Qingfu Zhu, Wanxiang Che
NAACL (Long Papers)6
2025 Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
abstract
Large Vision-Language Models (LVLMs) have achieved significant success in multimodal tasks, with multimodal chain-of-thought (MCoT) further enhancing performance and interpretability. Recent MCoT methods fall into two categories: (i) Textual-MCoT (T-MCoT), which takes multimodal input and produces textual output; and (ii) Interleaved-MCoT (I-MCoT), which generates interleaved image-text outputs. Despite advances in both approaches, the mechanisms driving these improvements are not fully understood. To fill this gap, we first reveal that MCoT boosts LVLMs by incorporating $\textit{visual thoughts}$, which convey image information to the reasoning process regardless of the MCoT format, depending only on clarity and conciseness of expression. Furthermore, to explore visual thoughts systematically, we define four distinct forms of visual thought expressions and analyze them comprehensively. Our findings demonstrate that these forms differ in clarity and conciseness, yielding varying levels of MCoT improvement. Additionally, we explore the internal nature of visual thoughts, finding that visual thoughts serve as intermediaries between the input image and reasoning to deeper transformer layers, enabling more advanced visual information transmission. We hope that the visual thoughts can inspire further breakthroughs for future MCoT research.
Zihui Cheng, Qiguang Chen, Xiao Xu 0005, Jiaqi Wang 0012, Weiyun Wang, Hao Fei 0003, Yidong Wang 0003, Alex Jinpeng Wang, Zhi Chen 0006, Wanxiang Che, Libo Qin 0001
NeurIPS11
2025 Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal Activations
abstract
Instruction tuning improves the ability of large language models (LLMs) to follow diverse human instructions, but achieving strong performance on specific target tasks remains challenging. A critical bottleneck is selecting the most relevant data to maximize task-specific performance. Existing data selection approaches include unstable influence-based methods and more stable distribution alignment methods, the latter of which critically rely on the underlying sample representation. In practice, most distribution alignment methods, from shallow features (e.g., BM25) to neural embeddings (e.g., BGE, LLM2Vec), may fail to capture how the model internally processes samples. To bridge this gap, we adopt a model-centric strategy in which each sample is represented by its neuronal activation pattern in the model, directly reflecting internal computation. However, directly using raw neuron activations leads to spurious similarity between unrelated samples due to neuron polysemanticity, where a single neuron may respond to multiple, unrelated concepts. To address this, we employ sparse autoencoders to disentangle polysemantic activations into sparse, monosemantic representations, and introduce a dedicated similarity metric for this space to better identify task-relevant data. Comprehensive experiments across multiple instruction datasets, models, tasks, and selection ratios show that our approach consistently outperforms existing data selection baselines in both stability and task-specific performance.
Gonghu Shang, Zhi Chen 0006, Libo Qin 0001, Yijie Luo, Hongshen Xu, Shuai Fan 0005, Kai Yu 0004, Lu Chen 0002
NeurIPS4
2025 Noise-Robustness Through Noise: A Framework combining Asymmetric LoRA with Poisoning MoE
abstract
Current parameter-efficient fine-tuning methods for adapting pre-trained language models to downstream tasks are susceptible to interference from noisy data. Conventional noise-handling approaches either rely on laborious data pre-processing or employ model architecture modifications prone to error accumulation. In contrast to existing noise-process paradigms, we propose a noise-robust adaptation method via asymmetric LoRA poisoning experts (LoPE), a novel framework that enhances model robustness to noise only with generated noisy data. Drawing inspiration from the mixture-of-experts architecture, LoPE strategically integrates a dedicated poisoning expert in an asymmetric LoRA configuration. Through a two-stage paradigm, LoPE performs noise injection on the poisoning expert during fine-tuning to enhance its noise discrimination and processing ability. During inference, we selectively mask the dedicated poisoning expert to leverage purified knowledge acquired by normal experts for noise-robust output. Extensive experiments demonstrate that LoPE achieves strong performance and robustness purely through the low-cost noise injection, which completely eliminates the requirement of data cleaning.
Zhaokun Wang, Jinyu Guo, Jingwen Pu, Lingfeng Chen, Hongli Pu, Jie Ou, Libo Qin 0001, Wenhong Tian
NeurIPS7
2025 MvDDI: A Multi-view Interaction Framework for Few-Shot Drug-Drug Interaction
Zihao Mao, Qiguang Chen, Yongheng Zhang 0001, Ruoxi Zhou, Peng Wang 0168, Libo Qin 0001
NLPCC (2)8
2025 MPFToD: a modularized pre-training framework for consistency identification in task-oriented dialogue
Libo Qin 0001, Shijue Huang, Qiguang Chen, Qian Liu 0033, Wanxiang Che, Ruifeng Xu 0001
Frontiers Comput. Sci.1
2025 Manager: Aggregating Insights From Unimodal Experts in Two-Tower VLMs and MLLMs
abstract
Two-Tower Vision–Language Models (VLMs) have demonstrated strong performance across various downstream VL tasks. While BridgeTower further enhances performance by building bridges between encoders, it(i)suffers from ineffective layer-by-layer utilization of unimodal representations,(ii)restricts the flexible exploitation of different levels of unimodal semantic knowledge, and(iii)is limited to the evaluation on traditional low-resolution datasets only with the Two-Tower VLM architecture. In this work, we propose Manager, a lightweight, efficient and effective plugin that adaptively aggregates insights from different levels of pre-trained unimodal experts to facilitate more comprehensive VL alignment and fusion. First, under the Two-Tower VLM architecture, we introduce ManagerTower, a novel VLM that introduces the manager in each cross-modal layer. Whether with or without VL pre-training, ManagerTower outperforms previous strong baselines and achieves superior performance on 4 downstream VL tasks. Moreover, we extend our exploration to the latest Multimodal Large Language Model (MLLM) architecture.We demonstrate that LLaVA-OV-Manager significantly boosts the zero-shot performance of LLaVA-OV across different categories of capabilities, images, and resolutions on 20 downstream datasets, whether the multi-grid algorithm is enabled or not. In-depth analysis reveals that both our manager and the multi-grid algorithm can be viewed as a plugin that improves the visual representation by capturing more diverse visual details from two orthogonal perspectives (depth and width). Their synergy can mitigate the semantic ambiguity caused by the multi-grid algorithm and further improve performance. Code and models are available at https://github.com/LooperXX/ManagerTower.
Xiao Xu 0005, Libo Qin 0001, Wanxiang Che, Min-Yen Kan
IEEE Trans. Circuits Syst. Video Technol.2
2025 CommentAgent: a LLM-powered agent framework for automated comment generation and opinion understanding
Jingyun Sun, Songhua Yu, Libo Qin 0001
J. Supercomput.4
2025 S3 Agent: Unlocking the Power of VLLM for Zero-Shot Multi-Modal Sarcasm Detection
abstract
Multi-modal sarcasm detection involves determining whether a given multi-modal input conveys sarcastic intent by analyzing the underlying sentiment. Recently, vision large language models have shown remarkable success on various of multi-modal tasks. Inspired by this, we systematically investigate the impact of vision large language models in zero-shot multi-modal sarcasm detection task. Furthermore, to capture different perspectives of sarcastic expressions, we propose a multi-view agent framework, S 3 Agent, designed to enhance zero-shot multi-modal sarcasm detection by leveraging three critical perspectives: superficial expression , semantic information , and sentiment expression . Our experiments on the MMSD2.0 dataset, which involves six models and four prompting strategies, demonstrate that our approach achieves state-of-the-art performance. Our method achieves an average improvement of 13.2% in accuracy. Moreover, we evaluate our method on the text-only sarcasm detection task, where it also surpasses baseline approaches.
Peng Wang 0168, Yongheng Zhang 0001, Hao Fei 0001, Qiguang Chen, Jiasheng Si, Wenpeng Lu, Min Li 0007, Libo Qin 0001
ACM Trans. Multim. Comput. Commun. Appl.9
2024 M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
abstract
Multi-modal Chain-of-Thought (MCoT) requires models to leverage knowledge from both textual and visual modalities for step-bystep reasoning, which gains increasing attention.Nevertheless, the current MCoT benchmark still faces some challenges: (1) absence of visual modal reasoning, (2) single-step visual modal reasoning, and (3) Domain missing, thereby hindering the development of MCoT.Motivated by this, we introduce a novel benchmark (M 3 CoT) to address the above challenges, advancing the multi-domain, multi-step, and multi-modal CoT.Additionally, we conduct a thorough evaluation involving abundant MCoT approaches on Vision Large Language Models (VLLMs).In addition, we highlight that the current VLLMs still struggle to correctly reason in M 3 CoT and there remains a large gap between existing VLLMs and human performance in M 3 CoT, despite their superior results on previous MCoT benchmarks.To our knowledge, we take the first meaningful step toward the multi-domain, multi-step, and multi-modal scenario in MCoT.We hope that M 3 CoT can serve as a valuable resource, providing a pioneering foundation in multi-domain, multi-step, multi-modal chain-of-thought research.
Qiguang Chen, Libo Qin 0001, Zhi Chen 0006, Xiao Xu 0005, Wanxiang Che
ACL (1)2
2024 Self-chats from Large Language Models Make Small Emotional Support Chatbot Better
abstract
Large Language Models (LLMs) have shown strong generalization abilities to excel in various tasks, including emotion support conversations.However, deploying such LLMs like GPT-3 (175B parameters) is resource-intensive and challenging at scale.In this study, we utilize LLMs as "Counseling Teacher" to enhance smaller models' emotion support response abilities, significantly reducing the necessity of scaling up model size.To this end, we first introduce an iterative expansion framework, aiming to prompt the large teacher model to curate an expansive emotion support dialogue dataset.This curated dataset, termed ExTES, encompasses a broad spectrum of scenarios and is crafted with meticulous strategies to ensure its quality and comprehensiveness.Based on this, we then devise a Diverse Response Inpainting (DRI) mechanism to harness the teacher model to produce multiple diverse responses by filling in the masked conversation context.This richness and variety serve as instructive examples, providing a robust foundation for finetuning smaller student models.Experiments across varied scenarios reveal that the teacherstudent scheme with DRI notably improves the response abilities of smaller models, even outperforming the teacher model in some cases.The dataset and codes are available 1 .
Zhonghua Zheng, Lizi Liao, Yang Deng 0002, Libo Qin 0001, Liqiang Nie
ACL (1)4
2024 A Two-Stage Framework with Self-Supervised Distillation for Cross-Domain Text Classification
abstract
Cross-domain text classification is a crucial task as it enables models to adapt to a target domain that lacks labeled data. It leverages or reuses rich labeled data from the different but related source domain(s) and unlabeled data from the target domain. To this end, previous work focuses on either extracting domain-invariant features or task-agnostic features, ignoring domain-aware features that may be present in the target domain and could be useful for the downstream task. In this paper, we propose a two-stage framework for cross-domain text classification. In the first stage, we finetune the model with mask language modeling (MLM) and labeled data from the source domain. In the second stage, we further fine-tune the model with self-supervised distillation (SSD) and unlabeled data from the target domain. We evaluate its performance on a public cross-domain text classification benchmark and the experiment results show that our method achieves new state-of-the-art results for both single-source domain adaptations (94.17% +1.03%) and multi-source domain adaptations (95.09% +1.34%).
Yunlong Feng, Bohan Li 0010, Libo Qin 0001, Xiao Xu 0005, Wanxiang Che
LREC/COLING3
2024 Improving Language Model Reasoning with Self-motivated Learning
abstract
Large-scale high-quality training data is important for improving the performance of models. After trained with data that has rationales (reasoning steps), models gain reasoning capability. However, the dataset with high-quality rationales is relatively scarce due to the high annotation cost. To address this issue, we propose Self-motivated Learning framework. The framework motivates the model itself to automatically generate rationales on existing datasets. Based on the inherent rank from correctness across multiple rationales, the model learns to generate better rationales, leading to higher reasoning capability. Specifically, we train a reward model with the rank to evaluate the quality of rationales, and improve the performance of reasoning through reinforcement learning. Experiment results of Llama2 7B on multiple reasoning datasets show that our method significantly improves the reasoning ability of models, even outperforming InstructGPT in some datasets.
Yunlong Feng, Yang Xu 0049, Libo Qin 0001, Yasheng Wang, Wanxiang Che
LREC/COLING3
2024 Python is Not Always the Best Choice: Embracing Multilingual Program of Thoughts
abstract
Program of Thoughts (PoT) is an approach characterized by its executable intermediate steps, which ensure the accuracy of the logical calculations in the reasoning process.Currently, PoT primarily uses Python.However, relying solely on a single language may result in suboptimal solutions and overlook the potential benefits of other programming languages.In this paper, we conduct comprehensive experiments on the programming languages used in PoT and find that no single language consistently delivers optimal performance across all tasks and models.The effectiveness of each language varies depending on the specific scenarios.Inspired by this, we propose a task and model agnostic approach called MultiPoT, which harnesses strength and diversity from various languages.Experimental results reveal that it significantly outperforms Python Self-Consistency.Furthermore, it achieves comparable or superior performance compared to the best monolingual PoT in almost all tasks across all models.In particular, MultiPoT achieves more than 4.6% improvement on average on ChatGPT (gpt-3.5-turbo-0701) 1 .
Xianzhen Luo, Qingfu Zhu, Libo Qin 0001, Qing Yang 0033, Dongliang Xu, Wanxiang Che
EMNLP4
2024 GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization
abstract
Yangfan Ye, Xiachong Feng, Xiaocheng Feng, Weitao Ma, Libo Qin, Dongliang Xu, Qing Yang, Hongtao Liu, Bing Qin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yangfan Ye, Xiachong Feng, Weitao Ma, Libo Qin 0001, Dongliang Xu, Qing Yang 0033, Hongtao Liu 0008, Bing Qin 0001
EMNLP5
2024 Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale
abstract
In recent years, Large Language Models (LLMs) have made significant strides towards Artificial General Intelligence.However, training these models from scratch requires substantial computational resources and vast amounts of text data.In this paper, we explore an alternative approach to constructing an LLM for a new language by continually pre-training (CPT) from existing pre-trained LLMs, instead of using randomly initialized parameters.Based on parallel experiments on 40 model sizes ranging from 40M to 5B parameters, we find that 1) CPT converges faster and saves significant resources in a scalable manner; 2) CPT adheres to an extended scaling law derived from Hoffmann et al. ( 2022) with a joint data-parameter scaling term; 3) The compute-optimal data-parameter allocation for CPT markedly differs based on our estimated scaling factors; 4) The effectiveness of transfer at scale is influenced by training duration and linguistic properties, while robust to data replaying, a method that effectively mitigates catastrophic forgetting in CPT.We hope our findings provide deeper insights into the transferability of LLMs at scale for the research community.
Wenzhen Zheng, Wenbo Pan 0001, Libo Qin 0001, Li Yue 0010
EMNLP4
2024 SDIF-DA: A Shallow-to-Deep Interaction Framework with Data Augmentation for Multi-Modal Intent Detection
abstract
Multi-modal intent detection aims to utilize various modalities to understand the user’s intentions, which is essential for the deployment of dialogue systems in real-world scenarios. The two core challenges for multi-modal intent detection are (1) how to effectively align and fuse different features of modalities and (2) the limited labeled multi-modal intent training data. In this work, we introduce a shallow-to-deep interaction framework with data augmentation (SDIF-DA) to address the above challenges. Firstly, SDIF-DA leverages a shallow-to-deep interaction module to progressively and effectively align and fuse features across text, video, and audio modalities. Secondly, we propose a ChatGPT-based data augmentation approach to automatically augment sufficient training data. Experimental results demonstrate that SDIF-DA can effectively align and fuse multi-modal features by achieving state-of-the-art performance. In addition, extensive analyses show that the introduced data augmentation approach can successfully distill knowledge from the large language model.
Shijue Huang, Libo Qin 0001, Geng Tu, Ruifeng Xu 0001
ICASSP2
2024 Pro-HAN: A Heterogeneous Graph Attention Network for Profile-based Spoken Language Understanding
abstract
Recently, Profile-based Spoken Language Understanding (SLU) has gained increasing attention, which aims to incorporate various types of supplementary profile information (i.e., Knowledge Graph, User Profile, Context Awareness) to eliminate the prevalent ambiguities in user utterances. However, existing approaches can only separately model different profile information, without considering their interrelationships or excluding irrelevant and conflicting information within them. To address the above issues, we introduce a Heterogeneous Graph Attention Network to perform reasoning across multiple Profile information, called Pro-HAN. Specifically, we design three types of edges, denoted as intra-Pro, inter-Pro, and utterance-Pro, to capture interrelationships among multiple Pros. We establish a new state-of-the-art on the ProSLU dataset, with an improvement of approximately 8% across all three metrics. Further analysis experiments also confirm the effectiveness of our method in modeling multi-source profile information.
Dechuan Teng, Chunlin Lu, Xiao Xu 0005, Wanxiang Che, Libo Qin 0001
ICASSP5
2024 LabCLIP: Label-Enhanced Clip for Improving Zero-Shot Text Classification
abstract
Zero-shot text classification aims to handle the text classification task without any annotated training data, which can greatly alleviate the data scarcity problem. Current dominant approaches follow a novel text-image matching paradigm, reformulating zero-shot text classification into a text-image matching problem, which can capture the visual image information and show promising performance. Nevertheless, existing text-image matching approaches solely focus on the visual image information, ignoring the semantic knowledge embedded in the text labels. To address the challenge, in the work, we present a label-enhanced CLIP framework (Lab-CLIP) for zero-shot text classification to consider both the visual image and text label semantic information simultaneously. Specifically, LabCLIP first converts the label into the corresponding image, and then injects the text label into the corresponding label image to explicitly capture the label semantic knowledge. We conduct experiments on 8 publicly available zero-shot text classification datasets and experimental results indicate that LabCLIP outperforms previous approaches on all datasets (with 4.3% improvement on average). In addition, we provide extensive analysis on exploring how to effectively incorporate the text label information.
Yongheng Zhang 0001, Peng Wang 0168, Qiguang Chen, Jingxuan Zhou, Yongmei Michelle Wang, Min Li 0007, Libo Qin 0001
ICASSP7
2024 Decoupling Breaks Data Barriers: A Decoupled Pre-training Framework for Multi-intent Spoken Language Understanding
Libo Qin 0001, Qiguang Chen, Jingxuan Zhou, Qinzheng Li, Chunlin Lu, Wanxiang Che
IJCAI1
2024 What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
abstract
Recently, rapid advancements in Multi-Modal In-Context Learning (MM-ICL) have achieved notable success, which is capable of achieving superior performance across various tasks without requiring additional parameter tuning. However, the underlying rules for the effectiveness of MM-ICL remain under-explored. To fill this gap, this work aims to investigate the research question: "_What factors affect the performance of MM-ICL?_" To this end, we investigate extensive experiments on the three core steps of MM-ICL including demonstration retrieval, demonstration ordering, and prompt construction using 6 vision large language models and 20 strategies. Our findings highlight (1) the necessity of a multi-modal retriever for demonstration retrieval, (2) the importance of intra-demonstration ordering over inter-demonstration ordering, and (3) the enhancement of task comprehension through introductory instructions in prompts. We hope this study can serve as a foundational guide for optimizing MM-ICL strategies in future research.
Libo Qin 0001, Qiguang Chen, Hao Fei 0003, Zhi Chen 0006, Min Li 0007, Wanxiang Che
NeurIPS1
2024 Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought
abstract
Chain-of-Thought (CoT) reasoning has emerged as a promising approach for enhancing the performance of large language models (LLMs) on complex reasoning tasks. Recently, a series of studies attempt to explain the mechanisms underlying CoT, aiming to deepen the understanding of its efficacy. Nevertheless, the existing research faces two major challenges: (1) a lack of quantitative metrics to assess CoT capabilities and (2) a dearth of guidance on optimizing CoT performance. Motivated by this, in this work, we introduce a novel reasoning boundary framework (RBF) to address these challenges. To solve the lack of quantification, we first define a reasoning boundary (RB) to quantify the upper-bound of CoT and establish a combination law for RB, enabling a practical quantitative approach applicable to various real-world CoT tasks. To address the lack of optimization, we propose three categories of RBs. We further optimize these categories with combination laws focused on RB promotion and reasoning path optimization for CoT improvement. Through extensive experiments on 27 models and 5 tasks, the study validates the existence and rationality of the proposed framework. Furthermore, it explains the effectiveness of 10 CoT strategies and guides optimization from two perspectives. We hope this work can provide a comprehensive understanding of the boundaries and optimization strategies for reasoning in LLMs. Our code and data are available at https://github.com/LightChen233/reasoning-boundary.
Qiguang Chen, Libo Qin 0001, Jiaqi Wang 0012, Jingxuan Zhou, Wanxiang Che
NeurIPS2
2023 Towards Complex Scenarios: Building End-to-End Task-Oriented Dialogue System across Multiple Knowledge Bases
abstract
With the success of the sequence-to-sequence model, end-to-end task-oriented dialogue systems (EToDs) have obtained remarkable progress. However, most existing EToDs are limited to single KB settings where dialogues can be supported by a single KB, which is still far from satisfying the requirements of some complex applications (multi-KBs setting). In this work, we first empirically show that the existing single-KB EToDs fail to work on multi-KB settings that require models to reason across various KBs. To solve this issue, we take the first step to consider the multi-KBs scenario in EToDs and introduce a KB-over-KB Heterogeneous Graph Attention Network (KoK-HAN) to facilitate model to reason over multiple KBs. The core module is a triple-connection graph interaction layer that can model different granularity levels of interaction information across different KBs (i.e., intra-KB connection, inter-KB connection and dialogue-KB connection). Experimental results confirm the superiority of our model for multiple KBs reasoning.
Libo Qin 0001, Zhouyang Li, Qiying Yu, Lehan Wang, Wanxiang Che
AAAI1
2023 Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across Languages
abstract
Chain-of-thought (CoT) is capable of eliciting models to explicitly generate reasoning paths, thus promoting reasoning accuracy and attracting increasing attention.Specifically, zeroshot CoT achieves remarkable improvements in a wide range of reasoning tasks by simply instructing the LLM with the prompt "Let's think step by step!".Despite the success of zero-shot CoT, the existing zero-shot prompting techniques remain limited to a single language, making it challenging to generalize to other languages and hindering global development.In this work, we introduce cross-lingual prompting (CLP), aiming to improve zero-shot CoT reasoning across languages.Specifically, CLP consists of two main components: (1) cross-lingual alignment prompting and (2) taskspecific solver prompting.The cross-lingual alignment prompting is responsible for aligning representations across different languages, whereas the task-specific solver prompting is used to generate the final chain of thoughts and results for the reasoning task.In addition, we further introduce cross-lingual self-consistent prompting (CLSP) to ensemble different reasoning paths across languages.Our experimental evaluations on several benchmarks demonstrate that CLP and CLSP significantly outperform the existing prompting methods and achieve state-of-the-art performance.We hope this work will inspire further breakthroughs in cross-lingual CoT.Answer: 48.After understanding, you should act as an expert in arithmetic reasoning in German… Verstanden! Hier ist die Lösung für die gegebene Aufgabe: 1. Da Jiesha 20 Jahre alt ist, ist Qioung 20-5... 2. Da Mary zwei Jahre jünger als Qioung ist... Antwort: 48.Sorry, I am an AI language model and I can only communicate in English.May I assist you with anything else?Let's understand the task in German step-by-step!German Prompting Please act as an expert in multi-lingual understanding in Chinese.Request:玛丽比琼小两岁,琼比杰莎大五岁。如果杰莎 20 岁,这 三个女孩的年龄总和是多少? Let's understand the task in English step-by-step!After understanding, you should act as an expert in arithmetic reasoning in English.Let's resolve the task you understand above step-by-step!Finally, you should format your answer as 'Answer: [num]'.Sure! Let's solve the task step-by-step: 1.According to the first sentence, Mary is two years younger than Joan.Let's suppose Joan's age is J, then Mary's age is J -2. 2. According to the second sentence, Joan is five years older than Jessie.If Jessie's age is x, then Joan's age is x + 5... Answer: 68.
Libo Qin 0001, Qiguang Chen, Fuxuan Wei, Shijue Huang, Wanxiang Che
EMNLP1
2023 End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future Directions
abstract
End-to-end task-oriented dialogue (EToD) can directly generate responses in an end-to-end fashion without modular training, which attracts escalating popularity.The advancement of deep neural networks, especially the successful use of large pre-trained models, has further led to significant progress in EToD research in recent years.In this paper, we present a thorough review and provide a unified perspective to summarize existing approaches as well as recent trends to advance the development of EToD research.The contributions of this paper can be summarized: (1) First survey: to our knowledge, we take the first step to present a thorough survey of this research field; (2) New taxonomy: we first introduce a unified perspective for EToD, including (i) Modularly EToD and (ii) Fully EToD; (3) New Frontiers: we discuss some potential frontier areas as well as the corresponding challenges, hoping to spur breakthrough research in EToD field; (4) Abundant resources: we build a public website 1 , where EToD researchers could directly access the recent progress.We hope this work can serve as a thorough reference for the EToD research community.EToD Modularly EToD ( §3.1)
Libo Qin 0001, Wenbo Pan 0001, Qiguang Chen, Lizi Liao, Zhou Yu 0005, Yue Zhang 0004, Wanxiang Che, Min Li 0007
EMNLP1
2023 Improving Few-Shot and Zero-Shot Entity Linking with Coarse-to-Fine Lexicon-Based Retriever
Shijue Huang, Libo Qin 0001, Ruifeng Xu 0001
NLPCC (3)3
2023 Modularized Pre-Training for End-to-End Task-Oriented Dialogue
abstract
Pre-training forend-to-endtask-orienteddialoguesystems (EToDs) is a challenging task due to its unique knowledge base query (accuracy) need and lack of sufficient training data (fluency). In this paper, we try to mitigate the above challenges by introducing a modularized pre-training framework for EToDs, which achieves to effectively improve both accuracy and fluency of EToDs through a pre-training paradigm. The core insight is a modular design by decomposing EToDs into ageneration (fluency)module and aknowledge-retriever (accuracy)module, which allows us to optimize each module by pre-training these two sub-modules with different well-designed pre-training tasks, respectively. In addition, such a modularized paradigm enables us to make full use of large amounts of KB-free dialogue corpus for the pre-traininggenerationmodule, which can alleviate the insufficient training problem. Furthermore, we introduce a newconsistency-guideddata augmentation (CGDA) strategy to cope with the data scarcity problem to better pre-train theknowledge-retrievermodule. Finally, we fine-tune the pre-trainedgenerationmodule andknowledge-retrievermodule jointly. Experimental results on three datasets show that our model achieve superior performance in terms of both fluency and accuracy. To our knowledge, this is the first work to explore modularized pre-training methods for EToDs.
Libo Qin 0001, Xiao Xu 0005, Lehan Wang, Yue Zhang 0004, Wanxiang Che
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Text Is No More Enough! A Benchmark for Profile-Based Spoken Language Understanding
abstract
Current researches on spoken language understanding (SLU) heavily are limited to a simple setting: the plain text-based SLU that takes the user utterance as input and generates its corresponding semantic frames (e.g., intent and slots). Unfortunately, such a simple setting may fail to work in complex real-world scenarios when an utterance is semantically ambiguous, which cannot be achieved by the text-based SLU models. In this paper, we first introduce a new and important task, Profile-based Spoken Language Understanding (ProSLU), which requires the model that not only relies on the plain text but also the supporting profile information to predict the correct intents and slots. To this end, we further introduce a large-scale human-annotated Chinese dataset with over 5K utterances and their corresponding supporting profile information (Knowledge Graph (KG), User Profile (UP), Context Awareness (CA)). In addition, we evaluate several state-of-the-art baseline models and explore a multi-level knowledge adapter to effectively incorporate profile information. Experimental results reveal that all existing text-based SLU models fail to work when the utterances are semantically ambiguous and our proposed framework can effectively fuse the supporting information for sentence-level intent detection and token-level slot filling. Finally, we summarize key challenges and provide new points for future directions, which hopes to facilitate the research.
Xiao Xu 0005, Libo Qin 0001, Kaiji Chen, Guoxing Wu, Linlin Li 0001, Wanxiang Che
AAAI2
2022 GL-CLeF: A Global-Local Contrastive Learning Framework for Cross-lingual Spoken Language Understanding
abstract
Libo Qin, Qiguang Chen, Tianbao Xie, Qixin Li, Jian-Guang Lou, Wanxiang Che, Min-Yen Kan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Libo Qin 0001, Qiguang Chen, Tianbao Xie, Qixin Li, Jian-Guang Lou, Wanxiang Che, Min-Yen Kan
ACL (1)1
2022 CGIM: A Cycle Guided Interactive Learning Model for Consistency Identification in Task-oriented Dialogue
abstract
Consistency identification in task-oriented dialog (CI-ToD) usually consists of three subtasks, aiming to identify inconsistency between current system response and current user response, dialog history and the corresponding knowledge base. This work aims to solve CI-ToD task by introducing an explicit interaction paradigm, Cycle Guided Interactive learning Model (CGIM), which achieves to make information exchange explicitly from all the three tasks. Specifically, CGIM relies on two core insights, referred to as guided multi-head attention module and cycle interactive mechanism, that collaborate from each other. On the one hand, each two tasks are linked with the guided multi-head attention module, aiming to explicitly model the interaction across two related tasks. On the other hand, we further introduce cycle interactive mechanism that focuses on facilitating model to exchange information among the three correlated sub-tasks via a cycle interaction manner. Experimental results on CI-ToD benchmark show that our model achieves the state-of-the-art performance, pushing the overall score to 56.3% (5.0% point absolute improvement). In addition, we find that CGIM is robust to the initial task flow order.
Libo Qin 0001, Qiguang Chen, Tianbao Xie, Qian Liu 0033, Shijue Huang, Wanxiang Che, Zhou Yu 0005
COLING1
2022 UniDU: Towards A Unified Generative Dialogue Understanding Framework
abstract
With the development of pre-trained language models, remarkable success has been witnessed in dialogue understanding (DU).However, current DU approaches usually employ independent models for each distinct DU task without considering shared knowledge across different DU tasks.In this paper, we propose a unified generative dialogue understanding framework, named UniDU, to achieve effective information exchange across diverse DU tasks.Here, we reformulate all DU tasks into a unified promptbased generative model paradigm.More importantly, a novel model-agnostic multi-task training strategy (MATS) is introduced to dynamically adapt the weights of diverse tasks for best knowledge sharing during training, based on the nature and available data of each task.Experiments on ten DU datasets covering five fundamental DU tasks show that the proposed UniDU framework largely outperforms task-specific well-designed methods on all tasks.MATS also reveals the knowledgesharing structure of these tasks.Finally, UniDU obtains promising performance in the unseen dialogue domain, showing the great potential for generalization.
Zhi Chen 0006, Lu Chen 0002, Bei Chen 0008, Libo Qin 0001, Yuncong Liu, Su Zhu, Jian-Guang Lou, Kai Yu 0004
SIGDIAL4
2022 Towards Event-level Causal Relation Identification
abstract
Existing methods usually identify causal relations between events at the mention-level, which takes each event mention pair as a separate input. As a result, they either suffer from conflicts among causal relations predicted separately or require a set of additional constraints to resolve such conflicts. We propose to study this task in a more realistic setting, where event-level causality identification can be made. The advantage is two folds: 1) with modeling different mentions of an event as a single unit, no more conflicts among predicted results, without any extra constraints; 2) with the use of diverse knowledge sources (e.g., co-occurrence and coreference relations), a rich graph-based event structure can be induced from the document for supporting event-level causal inference. Graph convolutional network is used to encode such structural information, which aims to capture the local and non-local dependencies among nodes. Results show that our model achieves the best performance under both mention- and event-level settings, outperforming a number of strong baselines by at least 2.8% on F1 score.
Chuang Fan, Daoxing Liu, Libo Qin 0001, Yue Zhang 0004, Ruifeng Xu 0001
SIGIR3
2022 Multi-domain Spoken Language Understanding Using Domain- and Task-aware Parameterization
abstract
Spoken language understanding (SLU) has been addressed as a supervised learning problem, where a set of training data is available for each domain. However, annotating data for a new domain can be both financially costly and non-scalable. One existing approach solves the problem by conducting multi-domain learning where parameters are shared for joint training across domains, which is domain-agnostic and task-agnostic . In the article, we propose to improve the parameterization of this method by using domain-specific and task-specific model parameters for fine-grained knowledge representation and transfer. Experiments on five domains show that our model is more effective for multi-domain SLU and obtain the best results. In addition, we show its transferability when adapting to a new domain with little data, outperforming the prior best model by 12.4%. Finally, we explore the strong pre-trained model in our framework and find that the contributions from our framework do not fully overlap with contextualized word representations (RoBERTa).
Libo Qin 0001, Fuxuan Wei, Minheng Ni, Yue Zhang 0004, Wanxiang Che, Yangming Li, Ting Liu 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2021 Co-GAT: A Co-Interactive Graph Attention Network for Joint Dialog Act Recognition and Sentiment Classification
abstract
In a dialog system, dialog act recognition and sentiment classification are two correlative tasks to capture speakers’ intentions, where dialog act and sentiment can indicate the explicit and the implicit intentions separately. The dialog context information (contextual information) and the mutual interaction information are two key factors that contribute to the two related tasks. Unfortunately, none of the existing approaches consider the two important sources of information simultaneously. In this paper, we propose a Co-Interactive Graph Attention Network (Co-GAT) to jointly perform the two tasks. The core module is a proposed co-interactive graph interaction layer where a cross-utterances connection and a cross-tasks connection are constructed and iteratively updated with each other, achieving to consider the two types of information simultaneously. Experimental results on two public datasets show that our model successfully captures the two sources of information and achieve the state-of-the-art performance. In addition, we find that the contributions from the contextual and mutual interaction information do not fully overlap with contextualized word representations (BERT, Roberta, XLNet).
Libo Qin 0001, Zhouyang Li, Wanxiang Che, Minheng Ni, Ting Liu 0001
AAAI1
2021 Language Model as an Annotator: Exploring DialoGPT for Dialogue Summarization
abstract
Xiachong Feng, Xiaocheng Feng, Libo Qin, Bing Qin, Ting Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xiachong Feng, Libo Qin 0001, Bing Qin 0001, Ting Liu 0001
ACL/IJCNLP (1)3
2021 GL-GIN: Fast and Accurate Non-Autoregressive Model for Joint Multiple Intent Detection and Slot Filling
abstract
Libo Qin, Fuxuan Wei, Tianbao Xie, Xiao Xu, Wanxiang Che, Ting Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Libo Qin 0001, Fuxuan Wei, Tianbao Xie, Xiao Xu 0005, Wanxiang Che, Ting Liu 0001
ACL/IJCNLP (1)1
2021 Don't be Contradicted with Anything! CI-ToD: Towards Benchmarking Consistency for Task-oriented Dialogue System
abstract
Consistency Identification has obtained remarkable success on open-domain dialogue, which can be used for preventing inconsistent response generation.However, in contrast to the rapid development in open-domain dialogue, few efforts have been made to the task-oriented dialogue direction.In this paper, we argue that consistency problem is more urgent in task-oriented domain.To facilitate the research, we introduce CI-ToD, a novel dataset for Consistency Identification in Taskoriented Dialog system.In addition, we not only annotate the single label to enable the model to judge whether the system response is contradictory, but also provide more finegrained labels (i.e., Dialogue History Inconsistency, User Query Inconsistency and Knowledge Base Inconsistency) to encourage model to know what inconsistent sources lead to it.Empirical results show that state-of-the-art methods only achieve 51.3%, which is far behind the human performance of 93.2%, indicating that there is ample room for improving consistency identification ability.Finally, we conduct exhaustive experiments and qualitative analysis to comprehend key challenges and provide guidance for future directions.All datasets and models are publicly available at https://github.com/yizhen20133868/CI-ToD. * Email corresponding. User: Give me directions to the closest grocery store.System: There is a whole foods 2 miles away and their address is 880_ames_ct.User: I need a route that avoids all heavy traffic.System: mandarin roots is located at 271 springer street.
Libo Qin 0001, Tianbao Xie, Shijue Huang, Qiguang Chen, Xiao Xu 0005, Wanxiang Che
EMNLP (1)1
2021 A Co-Interactive Transformer for Joint Slot Filling and Intent Detection
abstract
Intent detection and slot filling are two main tasks for building a spoken language understanding (SLU) system. The two tasks are closely related and the information of one task can benefit the other. Previous studies either implicitly model the two tasks with multi-task framework or only explicitly consider the single information flow from intent to slot. None of the prior approaches model the bidirectional connection between the two tasks simultaneously in a unified framework. In this paper, we propose a Co-Interactive Transformer which considers the cross-impact between the two tasks. Instead of adopting the self-attention mechanism in vanilla Transformer, we propose a co-interactive module to consider the cross-impact by building a bidirectional connection between the two related tasks, where slot and intent can be able to attend on the corresponding mutual information. The experimental results on two public datasets show that our model achieves the state-of-the-art performance.
Libo Qin 0001, Tailu Liu, Wanxiang Che, Bingbing Kang, Sendong Zhao, Ting Liu 0001
ICASSP1
2021 Injecting Word Information with Multi-Level Word Adapter for Chinese Spoken Language Understanding
abstract
In this paper, we improve Chinese spoken language understanding (SLU) by injecting word information. Previous studies on Chinese SLU do not consider the word information, failing to detect word boundaries that are beneficial for intent detection and slot filling. To address this issue, we propose a multi-level word adapter to inject word information for Chinese SLU, which consists of (1) sentence-level word adapter, which directly fuses the sentence representations of the word information and character information to perform intent detection and (2) character-level word adapter, which is applied at each character for selectively controlling weights on word information as well as character information. Experimental results on two Chinese SLU datasets show that our model can capture useful word information and achieve state-of-the-art performance.
Dechuan Teng, Libo Qin 0001, Wanxiang Che, Sendong Zhao, Ting Liu 0001
ICASSP2
2021 A Survey on Spoken Language Understanding: Recent Advances and New Frontiers
abstract
Spoken Language Understanding (SLU) aims to extract the semantics frame of user queries, which is a core component in a task-oriented dialog system. With the burst of deep neural networks and the evolution of pre-trained language models, the research of SLU has obtained significant breakthroughs. However, there remains a lack of a comprehensive survey summarizing existing approaches and recent trends, which motivated the work presented in this article. In this paper, we survey recent advances and new frontiers in SLU. Specifically, we give a thorough review of this research field, covering different aspects including (1) new taxonomy: we provide a new perspective for SLU filed, including single model vs. joint model, implicit joint modeling vs. explicit joint modeling in joint model, non pre-trained paradigm vs. pretrained paradigm; (2) new frontiers: some emerging areas in complex SLU as well as the corresponding challenges; (3) abundant open-source resources: to help the community, we have collected, organized the related papers, baseline projects and leaderboard on a public website where SLU researchers could directly access to the recent progress. We hope that this survey can shed a light on future research in SLU field.
Libo Qin 0001, Tianbao Xie, Wanxiang Che, Ting Liu 0001
IJCAI1
2021 Knowing Where to Leverage: Context-Aware Graph Convolutional Network With an Adaptive Fusion Layer for Contextual Spoken Language Understanding
abstract
Spoken language understanding (SLU) systems aim to understand users’ utterance, which is a key component of task-oriented dialogue systems. In this paper, we focus on improving the contextual SLU. The contextual SLU systems mainly focus on how to effectively incorporate dialog context information (contextual information). The existing approaches all use the same contextual information to guide slot filling at all tokens, which may inject the irrelevant information and result in ambiguity. To tackle this problem, we propose a context-aware graph convolutional network (GCN) with an adaptive fusion layer for contextual SLU. The context-aware GCN is proposed to automatically aggregate the contextual information, which frees our model from the manually designed heuristic aggregation function. Meanwhile, an adaptive fusion layer is applied at each token to dynamically incorporate relevant contextual information, which achieves a fine-grained contextual information transfer to guide the token-level slot filling. Experiments on the Simulated Dialog Dataset show that our model achieves state-of-the-art performance and outperforms other previous methods by a large margin (+3.67% on Sim-R, +4.18% on Sim-M and +3.75% on Overall dataset). In addition, we explore and analyze the pre-trained model (i.e., BERT) in our framework. We show that incorporating BERT brings a large improvement in low-resource setting.
Libo Qin 0001, Wanxiang Che, Minheng Ni, Yangming Li, Ting Liu 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Span-Based Neural Buffer: Towards Efficient and Effective Utilization of Long-Distance Context for Neural Sequence Models
abstract
Neural sequence model, though widely used for modeling sequential data such as the language model, has sequential recency bias (Kuncoro et al. 2018) to the local context, limiting its full potential to capture long-distance context. To address this problem, this paper proposes augmenting sequence models with a span-based neural buffer that efficiently represents long-distance context, allowing a gate policy network to make interpolated predictions from both the neural buffer and the underlying sequence model. Training this policy network to utilize long-distance context is however challenging due to the simple sentence dominance problem (Marvin and Linzen 2018). To alleviate this problem, we propose a novel training algorithm that combines an annealed maximum likelihood estimation with an intrinsic reward-driven reinforcement learning. Sequence models with the proposed span-based neural buffer significantly improve the state-of-the-art perplexities on the benchmark Penn Treebank and WikiText-2 datasets to 43.9 and 35.2 respectively. We conduct extensive analysis and confirm that the proposed architecture and the training algorithm both contribute to the improvements.
Yangming Li, Kaisheng Yao, Libo Qin 0001, Shuang Peng 0009, Xiaolong Li 0005
AAAI3
2020 DCR-Net: A Deep Co-Interactive Relation Network for Joint Dialog Act Recognition and Sentiment Classification
abstract
In dialog system, dialog act recognition and sentiment classification are two correlative tasks to capture speakers' intentions, where dialog act and sentiment can indicate the explicit and the implicit intentions separately (Kim and Kim 2018). Most of the existing systems either treat them as separate tasks or just jointly model the two tasks by sharing parameters in an implicit way without explicitly modeling mutual interaction and relation. To address this problem, we propose a Deep Co-Interactive Relation Network (DCR-Net) to explicitly consider the cross-impact and model the interaction between the two tasks by introducing a co-interactive relation layer. In addition, the proposed relation layer can be stacked to gradually capture mutual knowledge with multiple steps of interaction. Especially, we thoroughly study different relation layers and their effects. Experimental results on two public datasets (Mastodon and Dailydialog) show that our model outperforms the state-of-the-art joint model by 4.3% and 3.4% in terms of F1 score on dialog act recognition task, 5.7% and 12.4% on sentiment classification respectively. Comprehensive analysis empirically verifies the effectiveness of explicitly modeling the relation between the two tasks and the multi-steps interaction mechanism. Finally, we employ the Bidirectional Encoder Representation from Transformer (BERT) in our framework, which can further boost our performance in both tasks.
Libo Qin 0001, Wanxiang Che, Yangming Li, Minheng Ni, Ting Liu 0001
AAAI1
2020 Slot-consistent NLG for Task-oriented Dialogue Systems with Iterative Rectification Network
abstract
Data-driven approaches using neural networks have achieved promising performances in natural language generation (NLG).However, neural generators are prone to make mistakes, e.g., neglecting an input slot value and generating a redundant slot value.Prior works refer this to hallucination phenomenon.In this paper, we study slot consistency for building reliable NLG systems with all slot values of input dialogue act (DA) properly generated in output sentences.We propose Iterative Rectification Network (IRN) for improving general NLG systems to produce both correct and fluent responses.It applies a bootstrapping algorithm to sample training candidates and uses reinforcement learning to incorporate discrete reward related to slot inconsistency into training.Comprehensive studies have been conducted on multiple benchmark datasets, showing that the proposed methods have significantly reduced the slot error rate (ERR) for all strong baselines.Human evaluations also have confirmed its effectiveness.
Yangming Li, Kaisheng Yao, Libo Qin 0001, Wanxiang Che, Xiaolong Li 0005, Ting Liu 0001
ACL3
2020 Dynamic Fusion Network for Multi-Domain End-to-end Task-Oriented Dialog
abstract
Recent studies have shown remarkable success in end-to-end task-oriented dialog system. However, most neural models rely on large training data, which are only available for a certain number of task domains, such as navigation and scheduling. This makes it difficult to scalable for a new domain with limited labeled data. However, there has been relatively little research on how to effectively use data from all domains to improve the performance of each domain and also unseen domains. To this end, we investigate methods that can make explicit use of domain knowledge and introduce a shared-private network to learn shared and specific knowledge. In addition, we propose a novel Dynamic Fusion Network (DF-Net) which automatically exploit the relevance between the target domain and each domain. Results show that our models outperforms existing methods on multi-domain dialogue, giving the state-of-the-art in the literature. Besides, with little training data, we show its transferability by outperforming prior best model by 13.9% on average.
Libo Qin 0001, Xiao Xu 0005, Wanxiang Che, Yue Zhang 0004, Ting Liu 0001
ACL1
2020 Dialogue State Induction Using Neural Latent Variable Models
abstract
Dialogue state modules are a useful component in a task-oriented dialogue system. Traditional methods find dialogue states by manually labeling training corpora, upon which neural models are trained. However, the labeling process can be costly, slow, error-prone, and more importantly, cannot cover the vast range of domains in real-world dialogues for customer service. We propose the task of dialogue state induction, building two neural latent variable models that mine dialogue states automatically from unlabeled customer service dialogue records. Results show that the models can effectively find meaningful dialogue states. In addition, equipped with induced dialogue states, a state-of-the-art dialogue system gives better performance compared with not using a dialogue state module.
Qingkai Min, Libo Qin 0001, Zhiyang Teng, Xiao Liu 0029, Yue Zhang 0004
IJCAI2
2020 CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual NLP
abstract
Multi-lingual contextualized embeddings, such as multilingual-BERT (mBERT), have shown success in a variety of zero-shot cross-lingual tasks. However, these models are limited by having inconsistent contextualized representations of subwords across different languages. Existing work addresses this issue by bilingual projection and fine-tuning technique. We propose a data augmentation framework to generate multi-lingual code-switching data to fine-tune mBERT, which encourages model to align representations from source and multiple target languages once by mixing their context information. Compared with the existing work, our method does not rely on bilingual sentences for training, and requires only one training process for multiple target languages. Experimental results on five tasks with 19 languages show that our method leads to significantly improved performances for all the tasks compared with mBERT.
Libo Qin 0001, Minheng Ni, Yue Zhang 0004, Wanxiang Che
IJCAI1
2019 A Stack-Propagation Framework with Token-Level Intent Detection for Spoken Language Understanding
abstract
Libo Qin, Wanxiang Che, Yangming Li, Haoyang Wen, Ting Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Libo Qin 0001, Wanxiang Che, Yangming Li, Haoyang Wen, Ting Liu 0001
EMNLP/IJCNLP (1)1
2019 Entity-Consistent End-to-end Task-Oriented Dialogue System with KB Retriever
abstract
Libo Qin, Yijia Liu, Wanxiang Che, Haoyang Wen, Yangming Li, Ting Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Libo Qin 0001, Wanxiang Che, Haoyang Wen, Yangming Li, Ting Liu 0001
EMNLP/IJCNLP (1)1
2018 Sequence-to-Sequence Learning for Task-oriented Dialogue with Dialogue State Representation
abstract
Classic pipeline models for task-oriented dialogue system require explicit modeling the dialogue states and hand-crafted action spaces to query a domain-specific knowledge base. Conversely, sequence-to-sequence models learn to map dialogue history to the response in current turn without explicit knowledge base querying. In this work, we propose a novel framework that leverages the advantages of classic pipeline and sequence-to-sequence models. Our framework models a dialogue state as a fixed-size distributed representation and use this representation to query a knowledge base via an attention mechanism. Experiment on Stanford Multi-turn Multi-domain Task-oriented Dialogue Dataset shows that our framework significantly outperforms other sequence-to-sequence based baseline models on both automatic and human evaluation.
Haoyang Wen, Wanxiang Che, Libo Qin 0001, Ting Liu 0001
COLING4