EDBT 2026 Demo / reviewers in the wild / expert
Wanxiang Che
dblp:98/4640
· DBLP profile ↗
177ranked-venue papers
10as first author
100since 2021 · last 2026
0000-0002-3907-0335ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 155 · 10 first-author · 79 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 28 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 10 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Significant Gains in Reasoning Efficiency in Large Language ModelsabstractRecent advancements in large language models (LLMs) have greatly improved their ability to perform complex reasoning tasks through Long Chain-of-Thought (CoT). However, this approach often results in substantial redundancy, impairing computational efficiency and causing significant delays in real-time applications. To improve efficiency, current methods often rely on human-defined difficulty priors, which do not align with the LLM's self-awared difficulty, leading to inefficiencies. In this paper, we introduce the Dynamic Reasoning-Boundary Self-Awareness Framework (DR. SAF), which enables LLMs to dynamically assess and adjust their reasoning depth in response to problem complexity. DR. SAF integrates three key components: Boundary Self-Awareness Alignment, Adaptive Reward Management, and a Boundary Preservation Mechanism. These components allow models to optimize their reasoning processes, balancing efficiency and accuracy without compromising performance. Our experimental results demonstrate that DR. SAF achieves a 49.27% reduction in total response tokens with minimal loss in accuracy. The framework also delivers a 6.59x gain in token efficiency and a 5x reduction in training time, making it well-suited to resource-limited settings. During extreme training, DR. SAF can even surpass traditional instruction-based models in token efficiency with more than 16% accuracy improvement. Qiguang Chen, Dengyun Peng, Huikang Su, Jiannan Guan, Libo Qin 0001, Wanxiang Che |
AAAI | 7 |
| 2026 | Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution TasksabstractLarge Language Models (LLMs) excel in reasoning tasks requiring a single correct answer, but they perform poorly in multi-solution tasks that require generating comprehensive and diverse answers. We attribute this limitation to reasoning overconfidence: a tendency to express undue certainty in an incomplete solution set. To examine the effect, we introduce MuSoBench, a benchmark of multi-solution problems. Experiments show that the conventional short chain-of-thought (Short-CoT) prompting paradigm exhibits pronounced overconfidence, whereas the emerging long chain-of-thought (Long-CoT) approach mitigates it through iterative exploration and self-reflection. We further characterise observable behaviours and influential factors. To probe the underlying cause, we propose the cognitive-rigidity hypothesis, which posits that overconfidence arises when the reasoning process prematurely converges on a narrow set of thought paths. An attention-entropy analysis offers preliminary support for this view. These findings provide tools for assessing the completeness of LLM reasoning and highlight the need to move evaluation beyond single-answer accuracy toward comprehensive exploration. Jiannan Guan, Qiguang Chen, Libo Qin 0001, Dengyun Peng, Liangyu Huo, Wanxiang Che |
AAAI | 8 |
| 2026 | Judge Q: Trainable Queries for Optimized Information Retention in KV Cache EvictionabstractLarge language models (LLMs) utilize key-value (KV) cache to store historical information during sequence processing. The size of KV cache grows linearly as the length of the sequence extends, which seriously affects memory usage and decoding efficiency. Current methods for KV cache eviction typically utilize the last window from the pre-filling phase as queries to compute the KV importance scores for eviction. Although this scheme is simple to implement, it tends to overly focus on local information, potentially leading to the neglect or omission of crucial global information. To mitigate this issue, we propose **Judge Q**, a novel training method which incorporates a soft token list. This method only tunes the model’s embedding layer at a low training cost. By concatenating the soft token list at the end of the input sequence, we train these tokens' attention map to the original input sequence to align with that of the actual decoded tokens. In this way, the queries corresponding to the soft tokens can effectively capture global information and better evaluate the importance of the keys and values within the KV cache, thus maintaining decoding quality when KV cache is evicted. Under the same eviction budget, our method exhibits less performance degradation compared to existing eviction approaches. We validate our approach through experiments conducted on models such as Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3, using benchmarks including LongBench, RULER, and Needle-in-a-Haystack. Results indicate an improvement of approximately 1 point on the LongBench and over 3 points on RULER. This proposed methodology can be seamlessly integrated into existing open-source models with minimal training overhead, thereby enhancing performance in KV cache eviction scenarios. Yuzhuang Xu, Shiyu Ji, Yang Xu 0049, Qingfu Zhu, Wanxiang Che |
AAAI | 7 |
| 2026 | CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy AnalysisabstractLarge Language Models (LLMs) with Mixture-of-Experts (MoE) architectures are distinguished by their strong performance scaling with increasing parameters across a wide range of tasks, yet they also suffer from substantial computational and storage overheads. Notably, the performance gains of MoE models do not scale proportionally with the growth in expert parameters. While prior works attempt to reduce parameters via expert-level pruning, merging, or decomposition, they still suffer from challenges in both performance and computational efficiency. In this paper, we address these challenges by introducing micro-expert as a finer-grained compression unit that spans across matrices. We first establish a more fundamental perspective, viewing MoE layers as mixtures of micro-experts, and present CAMERA, a lightweight and training-free framework for identifying micro-expert redundancy. Our analysis uncovers significant variance in micro-expert contributions during decoding. Based on this insight, we further propose CAMERA-P, a structured micro-expert pruning framework, and CAMERA-Q, a mixed-precision quantization idea designed for micro-experts. Extensive experiments on nine downstream tasks show that CAMERA-P consistently outperforms strong baselines under pruning ratios ranging from 20% to 60%. Furthermore, CAMERA-Q achieves superior results under aggressive 2-bit quantization, surpassing existing matrix- and channel-level ideas. Notably, our method enables complete micro-expert analysis of Qwen2-57B-A14B in less than 5 minutes on a single NVIDIA A100-40GB GPU. Yuzhuang Xu, Xu Han 0007, Yuanchi Zhang, Shiyu Ji, Qingfu Zhu, Wanxiang Che |
AAAI | 8 |
| 2026 | Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational CapabilitiesabstractRecent advancements in Large Reasoning Models (LRMs), such as OpenAI's o1/o3 and DeepSeek-R1, have demonstrated remarkable performance in specialized reasoning tasks through human-like deliberative thinking and long chain-of-thought reasoning. However, our systematic evaluation across various model families (DeepSeek, Qwen, and LLaMA) and scales (7B to 32B) reveals that acquiring these deliberative reasoning capabilities significantly reduces the foundational capabilities of LRMs, including notable declines in helpfulness and harmlessness, alongside substantially increased inference costs. Importantly, we demonstrate that adaptive reasoning---employing modes like Zero-Thinking, Less-Thinking, and Summary-Thinking---can effectively alleviate these drawbacks. Our empirical insights underline the critical need for developing more versatile LRMs capable of dynamically allocating inference-time compute according to specific task characteristics. Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng 0002, Xuda Zhi, Yongbo Huang, Wanxiang Che, Ting Liu 0001, Bing Qin 0001 |
AAAI | 10 |
| 2026 | OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language ModelsabstractQiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu, Yi Yang, Yizhuo Li, Jingqi Tong, Xiachong Feng, Libo Qin, Wanxiang Che. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qiguang Chen, Chengyu Luan, Qiming Yu, Yizhuo Li 0007, Jingqi Tong, Xiachong Feng, Libo Qin 0001, Wanxiang Che |
ACL (1) | 10 |
| 2026 | When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action ModelsabstractVision-Language-Action (VLA) models have shown strong performance in language-conditioned robotic manipulation, yet their robustness to linguistic variation remains poorly understood. In this work, We present the first systematic multilingual evaluation of VLA models by translating the LIBERO benchmark into ten languages, revealing severe performance degradation under non-English instructions, with success rates dropping by 30–50%. Through fine-grained analysis of task executions, we find that language influence is highly non-uniform across steps: certain steps exhibit strong language dependence and dominate overall task failure, while others are largely language-agnostic. Based on this insight, we propose a step-wise inference-time intervention that aligns representations according to step language sensitivity, substantially improving performance under linguistic variation. Our results indicate that language robustness in VLA models is fundamentally a step-wise control problem, highlighting the importance of temporally structured analysis for reliable embodied agents. Tianhao Niu, Qingfu Zhu, Wanxiang Che |
ACL (1) | 5 |
| 2026 | Scaling Laws for Code: A More Data-Hungry RegimeabstractXianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang, Houyi Li, Siming Huang, YuanTao Fan, Wanxiang Che. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang, Houyi Li, Siming Huang, YuanTao Fan, Wanxiang Che |
ACL (1) | 8 |
| 2026 | Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive DashboardsabstractTianhao Niu, Ziyu Han, Qiguang Chen, Shiqi Zhou, Baocai Shan, Hengjie Fang, Qingfu Zhu, Wanxiang Che. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianhao Niu, Ziyu Han, Qiguang Chen, Shiqi Zhou, Baocai Shan, Hengjie Fang, Qingfu Zhu, Wanxiang Che |
ACL (1) | 8 |
| 2026 | Program of actions (PoA): Interactive semantic action parser as more powerful dialogue state tracker through code generation
Dechuan Teng, Wanxiang Che |
Expert Syst. Appl. | 3 |
| 2026 | Large language models meet NLP: a surveyabstractAbstract While large language models (LLMs) like ChatGPT have shown impressive capabilities in Natural Language Processing (NLP) tasks, a systematic investigation of their potential in this field remains largely unexplored. This study aims to address this gap by exploring the following questions. (1) How are LLMs currently applied to NLP tasks in the literature ? (2) Have traditional NLP tasks already been solved with LLMs ? (3) What is the future of the LLMs for NLP ? To answer these questions, we take the first step to provide a comprehensive overview of LLMs in NLP. Specifically, we first introduce a unified taxonomy including (1) parameter-frozen paradigm and (2) parameter-tuning paradigm to offer a unified perspective for understanding the current progress of LLMs in NLP. Furthermore, we summarize the new frontiers and the corresponding challenges, aiming to inspire further groundbreaking advancements. We hope this work offers valuable insights into {the potential and limitations} of LLMs, while also serving as a practical guide for building effective LLMs in NLP. Libo Qin 0001, Qiguang Chen, Xiachong Feng, Yang Wu 0010, Yongheng Zhang 0001, Min Li 0007, Wanxiang Che, Philip S. Yu |
Frontiers Comput. Sci. | 8 |
| 2026 | ColdChat: benchmarking large language model personalization using limited real-user interaction history
Zheni Zeng, Zhiyuan Liu 0001, Wanxiang Che |
Frontiers Comput. Sci. | 4 |
| 2026 | The gains do not make up for the losses: a comprehensive evaluation for safety alignment of large language models via machine unlearningabstractAbstract Machine Unlearning (MU) has emerged as a promising technique for aligning large language models (LLMs) with safety requirements to steer them forgetting specific harmful contents. Despite the significant progress in previous studies, we argue that the current evaluation criteria, which solely focus on safety evaluation, are actually impractical and biased , leading to concerns about the true effectiveness of MU techniques. To address this, we propose to comprehensively evaluate LLMs after MU from three aspects: safety, over-safety, and general utility. Specifically, a novel benchmark M u B ench with 18 related datasets is first constructed, where the safety is measured with both vanilla harmful inputs and 10 types of jailbreak attacks. Furthermore, we examine whether MU introduces side effects, focusing on over-safety and utility-loss. Extensive experiments are performed on 3 popular LLMs with 7 recent MU methods. The results highlight a challenging trilemma in safety alignment without side effects, indicating that there is still considerable room for further exploration. M u B ench serves as a comprehensive benchmark, fostering future research on MU for safety alignment of LLMs. Weixiang Zhao, Yulin Hu, Xingyu Sui, Zhuojun Li, Yang Deng 0002, Bing Qin 0001, Wanxiang Che |
Frontiers Comput. Sci. | 8 |
| 2025 | Divide-Solve-Combine: An Interpretable and Accurate Prompting Framework for Zero-shot Multi-Intent DetectionabstractZero-shot multi-intent detection is capable of capturing multiple intents within a single utterance without any training data, which gains increasing attention. Building on the success of large language models (LLM), dominant approaches in the literature explore prompting techniques to enable zero-shot multi-intent detection. While significant advancements have been witnessed, the existing prompting approaches still face two major issues: lacking explicit reasoning and lacking interpretability. Therefore, in this paper, we introduce a Divide-Solve-Combine Prompting (DSCP) to address the above issues. Specifically, DSCP explicitly decomposes multi-intent detection into three components including (1) single-intent division prompting is utilized to decompose an input query into distinct sub-sentences, each containing a single intent; (2) intent-by-intent solution prompting is applied to solve each sub-sentence recurrently; and (3) multi-intent combination prompting is employed for combining each sub-sentence result to obtain the final multi-intent result. By decomposition, DSCP allows the model to track the explicit reasoning process and improve the interpretability. In addition, we propose an interactive divide-solve-combine prompting (Inter-DSCP) to naturally capture the interaction capabilities of large language models. Experimental results on two standard multi-intent benchmarks (i.e., MixATIS and MixSNIPS) reveal that both DSCP and Inter-DSCP obtain substantial improvements over baselines, achieving superior performance and higher interpretability. Libo Qin 0001, Qiguang Chen, Jingxuan Zhou, Hao Fei 0003, Wanxiang Che, Min Li 0007 |
AAAI | 6 |
| 2025 | CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have recently demonstrated amazing success in multi-modal tasks, including advancements in Multi-modal Chain-of-Thought (MCoT) reasoning. Despite these successes, current benchmarks still follow a traditional paradigm with multi-modal input and text-modal output, which leads to significant drawbacks such as missing visual operations and vague expressions. Motivated by this, we introduce a novel Chain of Multi-modal Thought (CoMT) benchmark to address these limitations. Different from the traditional MCoT benchmark, CoMT requires both multi-modal input and multi-modal reasoning output, aiming to mimic human-like reasoning that inherently integrates visual operation. Specifically, CoMT consists of four categories: (1) Visual Creation, (2) Visual Deletion, (3) Visual Update, and (4) Visual Selection to comprehensively explore complex visual operations and concise expression in real scenarios. We evaluate various LVLMs and strategies on CoMT, revealing some key insights into the capabilities and limitations of the current approaches. We hope that CoMT can inspire more research on introducing multi-modal generation into the reasoning process. Zihui Cheng, Qiguang Chen, Hao Fei 0003, Wanxiang Che, Min Li 0007, Libo Qin 0001 |
AAAI | 6 |
| 2025 | CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding EvaluationabstractDespite the rapid development of Chinese vision-language models (VLMs), most existing Chinese vision-language (VL) datasets are constructed on Western-centric images from existing English VL datasets. The cultural bias in the images makes these datasets unsuitable for evaluating VLMs in Chinese culture. To remedy this issue, we present a new Chinese Vision-Language Understanding Evaluation (CVLUE) benchmark dataset, where the selection of object categories and images is entirely driven by Chinese native speakers, ensuring that the source images are representative of Chinese culture. The benchmark contains four distinct VL tasks ranging from image-text retrieval to visual question answering, visual grounding and visual dialogue. We present a detailed statistical analysis of CVLUE and provide a baseline performance analysis with several open-source multilingual VLMs on CVLUE and its English counterparts to reveal their performance gap between English and Chinese. Our in-depth category-level analysis reveals a lack of Chinese cultural knowledge in existing VLMs. We also find that fine-tuning on Chinese culture-related VL datasets effectively enhances VLMs' understanding of Chinese culture. Yuxuan Wang 0001, Fei Yu 0012, Zhiguo Wan, Wanxiang Che, Hongyang Chen 0001 |
AAAI | 7 |
| 2025 | Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population TraitsabstractThe Myers-Briggs Type Indicator (MBTI) is one of the most influential personality theories reflecting individual differences in thinking, feeling, and behaving. MBTI personality detection has garnered considerable research interest and has evolved significantly over the years. However, this task tends to be overly optimistic, as it currently does not align well with the natural distribution of population personality traits. Specifically, the self-reported labels in existing datasets result in data quality issues and the hard labels fail to capture the full range of population personality distributions. In this paper, we identify the task by constructing MBTIBench, the first manually annotated MBTI personality detection dataset with soft labels, under the guidance of psychologists. Our experimental results confirm that soft labels can provide more benefits to other psychological tasks than hard labels. We highlight the polarized predictions and biases in LLMs as key directions for future research. Bohan Li 0010, Jiannan Guan, Longxu Dou, Yunlong Feng, Dingzirui Wang, Yang Xu 0049, Enbo Wang, Qiguang Chen, Bichen Wang, Xiao Xu 0005, Libo Qin 0001, Qingfu Zhu, Wanxiang Che |
COLING | 15 |
| 2025 | MURRE: Multi-Hop Table Retrieval with Removal for Open-Domain Text-to-SQLabstractThe open-domain text-to-SQL task aims to retrieve question-relevant tables from massive databases and generate SQL. However, the performance of current methods is constrained by single-hop retrieval, and existing multi-hop retrieval of open-domain question answering is not directly applicable due to the tendency to retrieve tables similar to the retrieved ones but irrelevant to the question. Since the questions in text-to-SQL usually contain all required information, while previous multi-hop retrieval supplements the questions with retrieved documents. Therefore, we propose the multi-hop table retrieval with removal (MURRE), which removes previously retrieved information from the question to guide the retriever towards unretrieved relevant tables. Our experiments on two open-domain text-to-SQL datasets demonstrate an average improvement of 5.7% over the previous state-of-the-art results. Xuanliang Zhang, Dingzirui Wang, Longxu Dou, Qingfu Zhu, Wanxiang Che |
COLING | 5 |
| 2025 | Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code GenerationabstractChart2Code has recently received significant attention in the multimodal community due to its potential to reduce the burden of visualization and promote a more detailed understanding of charts.However, existing Chart2Coderelated training datasets suffer from at least one of the following issues: (1) limited scale, (2) limited type coverage, and ( 3) inadequate complexity.To address these challenges, we seek more diverse sources that better align with real-world user distributions and propose dual data synthesis pipelines: (1) Synthesize based on online plotting code.(2) Synthesize based on the chart images in the academic paper.We create a large-scale Chart2Code training dataset Chart2Code53, including 53 chart types, 130K Chart-code pairs based on the pipeline.Experimental results demonstrate that even with few parameters, the model finetuned on Chart2Code53 achieves state-ofthe-art performance on multiple Chart2Code benchmarks within open-source models 1 . Tianhao Niu, Yiming Cui 0001, Baoxin Wang, Xiao Xu 0005, Qingfu Zhu, Dayong Wu, Shijin Wang 0001, Wanxiang Che |
EMNLP | 9 |
| 2025 | Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo QueryabstractLarge language models (LLMs) rely on keyvalue cache (KV cache) to accelerate decoding by reducing redundant computations.However, the KV cache memory usage grows substantially with longer text sequences, posing challenges for efficient deployment.Existing KV cache eviction methods prune tokens using prefilling-stage attention scores, causing inconsistency with actual inference queries, especially under tight memory budgets.In this paper, we propose Lookahead Q-Cache (LAQ), a novel eviction framework that generates lowcost pseudo lookahead queries to better approximate the true decoding-stage queries.By using these lookahead queries as the observation window for importance estimation, LAQ achieves more consistent and accurate KV cache eviction aligned with real inference scenarios.Experimental results on LongBench and Needlein-a-Haystack benchmarks show that LAQ outperforms existing methods across various budget levels, achieving a 1 ∼ 4 point improvement on LongBench under limited cache budget.Moreover, LAQ is complementary to existing approaches and can be flexibly combined to yield further improvements. Shiyu Ji, Yuzhuang Xu, Yang Xu 0049, Qingfu Zhu, Wanxiang Che |
EMNLP | 7 |
| 2025 | RoT: Enhancing Table Reasoning with Iterative Row-Wise TraversalsabstractThe table reasoning task, crucial for efficient data acquisition, aims to answer questions based on the given table .Recently, reasoning large language models (RLLMs) with Long Chain-of-Thought (Long CoT) significantly enhance reasoning capabilities, leading to brilliant performance on table reasoning.However, Long CoT suffers from high cost for training and exhibits low reliability due to table content hallucinations.Therefore, we propose Rowof-Thought (ROT), which performs iteratively row-wise table traversal, allowing for reasoning extension and reflection-based refinement at each traversal.Scaling reasoning length by rowwise traversal and leveraging reflection capabilities of LLMs, ROT is training-free.The sequential traversal encourages greater attention to the table, thus reducing hallucinations.Experiments show that ROT, using non-reasoning models, outperforms RLLMs by an average of 4.3%, and achieves state-of-the-art results on WikiTableQuestions and TableBench with comparable models, proving its effectiveness.Also, ROT outperforms Long CoT with fewer reasoning tokens, indicating higher efficiency. Xuanliang Zhang, Dingzirui Wang, Keyan Xu, Qingfu Zhu, Wanxiang Che |
EMNLP | 5 |
| 2025 | CroPrompt: Cross-task Interactive Prompting for Zero-shot Spoken Language UnderstandingabstractSlot filling and intent detection are two highly correlated tasks in spoken language understanding (SLU). Recent SLU research attempts to explore zero-shot prompting techniques in large language models to alleviate the data scarcity problem. Nevertheless, the existing prompting work ignores the cross-task interaction information for SLU, which leads to sub-optimal performance. To solve this problem, we present the pioneering work of Cross-task Interactive Prompting (CroPrompt) for SLU, which enables the model to interactively leverage the information exchange across the correlated tasks in SLU. Additionally, we further introduce a multi-task self-consistency mechanism to mitigate the error propagation caused by the intent information injection. We conduct extensive experiments on the standard SLU benchmark and the results reveal that CroPrompt consistently outperforms the existing prompting approaches. In addition, the multi-task self-consistency mechanism can effectively ease the error propagation issue, thereby enhancing the performance. We hope this work can inspire more research on cross-task prompting for SLU. Libo Qin 0001, Fuxuan Wei, Qiguang Chen, Jingxuan Zhou, Shijue Huang, Jiasheng Si, Wenpeng Lu, Wanxiang Che |
ICASSP | 8 |
| 2025 | Improving Consistency Identification in Task-oriented Dialogue Through Multi-Agent CollaborationabstractConsistency identification in task-oriented dialog (CI-ToD) typically consists of three sub-tasks: User Query Inconsistency (QI) identification, Dialogue History Inconsistency (HI) identification, and Knowledge Base Inconsistency (KBI) identification, which aim to determine inconsistent relationships between system response and user query, dialogue history, and knowledge base. Previous approaches focus on the exploration of deep learning models for CI-ToD. While these models achieve remarkable progress, they still rely on large amounts of labeled data, which is hard to achieve in real-world scenarios. Motivated by this, in the paper, we aim to explore large language models for CI-ToD, which do not require any training data. In addition, we further introduce a multi-agent collaboration framework (MAC-CIToD) to model the interaction across three sub-tasks in CI-ToD, including (1) Full Connection paradigm, (2) Cycle Connection paradigm, and (3) Central Connection paradigm, which effectively builds interaction across QI, HI, and KBI. Experiments on the standard benchmark reveal that our framework achieves superior performance. Additionally, we compare MAC-CIToD with the most advanced trained approaches and find that its zero-shot performance on most metrics even surpasses that of models after training on the CI-ToD dataset. Ruoxi Zhou, Qiguang Chen, Xiao Xu 0005, Hao Fei 0003, Dagang Li 0001, Wanxiang Che, Libo Qin 0001 |
IJCAI | 8 |
| 2025 | Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection LearningabstractEmpowering large language models (LLMs) with effective tool utilization capabilities is crucial for enabling AI agents to solve complex problems. However, current models face two major limitations: (1) unreliable tool planning and invocation due to low-quality instruction datasets (e.g., widespread hallucinated API calls), and (2) weak tool reflection abilities (over 90% of errors cannot be corrected) resulting from static imitation learning. To address these critical limitations, we propose Tool-MVR, a novel Tool-Augmented LLM that achieves comprehensive System 2 reasoning through two key innovations. Specifically, we first introduce Multi-Agent Meta-Verification (MAMV), a systematic pipeline that rigorously validates APIs, queries, and reasoning trajectories to construct ToolBench-V, a new high-quality instruction dataset that addresses the limitation of unreliable tool planning and invocation. Second, we propose Exploration-based Reflection Learning (EXPLORE), which enhances tool reflection capabilities by leveraging tool feedback through a dynamic "Error → Reflection → Correction" learning paradigm, resulting in our reflection dataset ToolBench-R and addressing the critical weakness in tool reflection. Finally, we obtain Tool-MVR by finetuning open-source LLMs (e.g., Qwen-7B) on both ToolBench-V and ToolBench-R. Our experiments demonstrate that Tool-MVR achieves state-of-the-art performance on StableToolBench, surpassing both ToolLLM (by 23.9%) and GPT-4 (by 15.3%) while reducing API calls by 31.4%, with strong generalization capabilities across unseen tools and scenarios. Additionally, on our proposed RefineToolBench, the first benchmark specifically designed to evaluate tool reflection capabilities. Tool-MVR achieves a 58.9% error correction rate, significantly outperforming ToolLLM's 9.1%. Zhiyuan Ma 0006, Jiayu Liu 0001, Xianzhen Luo, Zhenya Huang, Qingfu Zhu, Wanxiang Che |
KDD (2) | 6 |
| 2025 | MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language ModelsabstractMultimodal planning capabilities refer to the ability to predict, reason, and design steps for task execution with multimodal context, which is essential for complex reasoning and decision-making across multiple steps. However, current benchmarks face two key challenges: (1) they cannot directly assess multimodal real-world planning capabilities, and (2) they lack constraints or implicit constraints across modalities. To address these issues, we introduce Multimodal Planning with Complex Constraints (MPCC), the first benchmark to systematically evaluate MLLMs' ability to handle multimodal constraints in planning. To address the first challenge, MPCC focuses on three real-world tasks: Flight Planning, Calendar Planning, and Meeting Planning. To solve the second challenge, we introduce complex constraints (e.g. budget, temporal, and spatial) in these tasks, with graded difficulty levels (EASY, MEDIUM, HARD) to separate constraint complexity from search space expansion. Experiments on 13 advanced MLLMs reveal significant challenges: closed-source models achieve only 21.3% feasible plans, while open-source models average below 11%. Additionally, we observe that MLLMs are highly sensitive to constraint complexity and that traditional multimodal prompting strategies fail in multi-constraint scenarios. Our work formalizes multimodal constraints in planning, provides a rigorous evaluation framework, and highlights the need for advancements in constraint-aware reasoning for real-world MLLM applications. Yiyan Ji, Qiguang Chen, Chengyue Wu, Libo Qin 0001, Wanxiang Che |
ACM Multimedia | 6 |
| 2025 | ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language ModelsabstractVideo understanding plays a vital role in bridging low-level visual signals with high-level cognitive reasoning, and is fundamental to applications such as autonomous driving, embodied AI, and the broader pursuit of AGI. The rapid development of large language models (LLMs), particularly those utilizing Chain-of-Thought (CoT) technology, has significantly advanced video reasoning capabilities. However, current approaches primarily depend on textual information for reasoning, overlooking the visual modality in the actual video reasoning process. In contrast, humans naturally re-examine visual content while reasoning. Motivated by this, we introduce a novel video reasoning paradigm: Video-Text Interleaved CoT (ViTCoT), which facilitates more intuitive and cognitively aligned reasoning. To the end, first, we construct the Video-Text Interleaved Benchmark (ViTIB), which is created using MLLMs for key-video selection and manually verified. Furthermore, we extensively explore the potential of the ViTCoT paradigm in the video understanding field. Extensive experiments demonstrate that ViTCoT significantly enhances performance compared to the traditional text-only CoT paradigm and effectively activates more neuron values in MLLMs. Yongheng Zhang 0001, Ruihan Tao, Qiguang Chen, Hao Fei 0001, Wanxiang Che, Libo Qin 0001 |
ACM Multimedia | 6 |
| 2025 | Stealthy Jailbreak Attacks on Large Language Models via Benign Data MirroringabstractHonglin Mu, Han He, Yuxin Zhou, Yunlong Feng, Yang Xu, Libo Qin, Xiaoming Shi, Zeming Liu, Xudong Han, Qi Shi, Qingfu Zhu, Wanxiang Che. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Honglin Mu, Han He, Yunlong Feng, Yang Xu 0049, Libo Qin 0001, Zeming Liu, Qi Shi 0002, Qingfu Zhu, Wanxiang Che |
NAACL (Long Papers) | 12 |
| 2025 | Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-ThoughtabstractLarge Vision-Language Models (LVLMs) have achieved significant success in multimodal tasks, with multimodal chain-of-thought (MCoT) further enhancing performance and interpretability. Recent MCoT methods fall into two categories: (i) Textual-MCoT (T-MCoT), which takes multimodal input and produces textual output; and (ii) Interleaved-MCoT (I-MCoT), which generates interleaved image-text outputs. Despite advances in both approaches, the mechanisms driving these improvements are not fully understood. To fill this gap, we first reveal that MCoT boosts LVLMs by incorporating $\textit{visual thoughts}$, which convey image information to the reasoning process regardless of the MCoT format, depending only on clarity and conciseness of expression. Furthermore, to explore visual thoughts systematically, we define four distinct forms of visual thought expressions and analyze them comprehensively. Our findings demonstrate that these forms differ in clarity and conciseness, yielding varying levels of MCoT improvement. Additionally, we explore the internal nature of visual thoughts, finding that visual thoughts serve as intermediaries between the input image and reasoning to deeper transformer layers, enabling more advanced visual information transmission. We hope that the visual thoughts can inspire further breakthroughs for future MCoT research. Zihui Cheng, Qiguang Chen, Xiao Xu 0005, Jiaqi Wang 0012, Weiyun Wang, Hao Fei 0003, Yidong Wang 0003, Alex Jinpeng Wang, Zhi Chen 0006, Wanxiang Che, Libo Qin 0001 |
NeurIPS | 10 |
| 2025 | When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual ReasonersabstractMultilingual reasoning remains a significant challenge for large language models (LLMs), with performance disproportionately favoring high-resource languages. Drawing inspiration from cognitive neuroscience, which suggests that human reasoning functions largely independently of language processing, we hypothesize that LLMs similarly encode reasoning and language as separable components that can be disentangled to enhance multilingual reasoning. To evaluate this, we perform a causal intervention by ablating language-specific representations at inference time. Experiments on 10 open-weight LLMs spanning 11 typologically diverse languages show that this language-specific ablation consistently boosts multilingual reasoning performance. Layer-wise analyses further confirm that language and reasoning representations can be effectively disentangled throughout the model, yielding improved multilingual reasoning capabilities, while preserving top-layer language features remains essential for maintaining linguistic fidelity. Compared to post-training methods such as supervised fine-tuning or reinforcement learning, our training-free language-reasoning disentanglement achieves comparable or superior results with minimal computational overhead. These findings shed light on the internal mechanisms underlying multilingual reasoning in LLMs and suggest a lightweight and interpretable strategy for improving cross-lingual generalization. Weixiang Zhao, Jiahe Guo, Yang Deng 0002, Tongtong Wu, Wenxuan Zhang 0001, Yulin Hu, Xingyu Sui, Wanxiang Che, Bing Qin 0001, Tat-Seng Chua, Ting Liu 0001 |
NeurIPS | 9 |
| 2025 | Towards few-shot mixed-type dialogue generation
Zeming Liu, Haifeng Wang 0001, Zeyang Lei, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
Sci. China Inf. Sci. | 6 |
| 2025 | Large language models meet text-centric multimodal sentiment analysis: a survey
Hao Yang 0066, Yang Wu 0010, Shilong Wang 0003, Zongyang Ma, Wanxiang Che, Shijin Wang 0001, Si Wei, Bing Qin 0001 |
Sci. China Inf. Sci. | 8 |
| 2025 | From unimodal to multimodal: a framework for generating high-quality multimodal emotional chit-chat dialogue
Hao Yang 0066, Yang Wu 0010, Jianhua Yuan, Wanxiang Che, Shijin Wang 0001, Si Wei, Bing Qin 0001 |
Sci. China Inf. Sci. | 6 |
| 2025 | MPFToD: a modularized pre-training framework for consistency identification in task-oriented dialogue
Libo Qin 0001, Shijue Huang, Qiguang Chen, Qian Liu 0033, Wanxiang Che, Ruifeng Xu 0001 |
Frontiers Comput. Sci. | 5 |
| 2025 | A survey of table reasoning with large language models
Xuanliang Zhang, Dingzirui Wang, Longxu Dou, Qingfu Zhu, Wanxiang Che |
Frontiers Comput. Sci. | 5 |
| 2025 | Against The Achilles' Heel: A Survey on Red Teaming for Generative ModelsabstractGenerative models are rapidly gaining popularity and being integrated into everyday applications, raising concerns over their safe use as various vulnerabilities are exposed. In light of this, the field of red teaming is undergoing fast-paced growth, highlighting the need for a comprehensive survey covering the entire pipeline and addressing emerging topics. Our extensive survey, which examines over 120 papers, introduces a taxonomy of fine-grained attack strategies grounded in the inherent capabilities of language models. Additionally, we have developed the “searcher” framework to unify various automatic red teaming approaches. Moreover, our survey covers novel areas including multimodal attacks and defenses, risks around LLM-based agents, overkill of harmless queries, and the balance between harmlessness and helpfulness. Warning: This paper contains examples that may be offensive, harmful, or biased. Lizhi Lin, Honglin Mu, Zenan Zhai, Minghan Wang, Yuxia Wang 0003, Renxi Wang, Wanxiang Che, Timothy Baldwin, Haonan Li 0002 |
J. Artif. Intell. Res. | 9 |
| 2025 | CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMsabstractAbstract Powerful large language models (LLMs) are increasingly expected to be deployed with lower computational costs, enabling their capabilities on resource-constrained devices. Post-training quantization (PTQ) has emerged as a star approach to achieve this ambition, with best methods compressing weights to less than 2 bit on average. In this paper, we propose Channel-Relaxed Vector Quantization (CRVQ), a novel technique that significantly improves the performance of PTQ baselines at the cost of only minimal additional bits. This state-of-the-art extreme compression method achieves its results through two key innovations: (1) carefully selecting and reordering a very small subset of critical weight channels, and (2) leveraging extended codebooks to relax the constraint of critical channels. With our method, we demonstrate a 38.9% improvement over the current strongest sub-2-bit PTQ baseline, enabling nearer lossless 1-bit compression. Furthermore, our approach offers flexible customization of quantization bit-width and performance, providing a wider range of deployment options for diverse hardware platforms. Code and checkpoints are available at https://github.com/xuyuzhuang11/CRVQ. Yuzhuang Xu, Shiyu Ji, Qingfu Zhu, Wanxiang Che |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Manager: Aggregating Insights From Unimodal Experts in Two-Tower VLMs and MLLMsabstractTwo-Tower Vision–Language Models (VLMs) have demonstrated strong performance across various downstream VL tasks. While BridgeTower further enhances performance by building bridges between encoders, it(i)suffers from ineffective layer-by-layer utilization of unimodal representations,(ii)restricts the flexible exploitation of different levels of unimodal semantic knowledge, and(iii)is limited to the evaluation on traditional low-resolution datasets only with the Two-Tower VLM architecture. In this work, we propose Manager, a lightweight, efficient and effective plugin that adaptively aggregates insights from different levels of pre-trained unimodal experts to facilitate more comprehensive VL alignment and fusion. First, under the Two-Tower VLM architecture, we introduce ManagerTower, a novel VLM that introduces the manager in each cross-modal layer. Whether with or without VL pre-training, ManagerTower outperforms previous strong baselines and achieves superior performance on 4 downstream VL tasks. Moreover, we extend our exploration to the latest Multimodal Large Language Model (MLLM) architecture.We demonstrate that LLaVA-OV-Manager significantly boosts the zero-shot performance of LLaVA-OV across different categories of capabilities, images, and resolutions on 20 downstream datasets, whether the multi-grid algorithm is enabled or not. In-depth analysis reveals that both our manager and the multi-grid algorithm can be viewed as a plugin that improves the visual representation by capturing more diverse visual details from two orthogonal perspectives (depth and width). Their synergy can mitigate the semantic ambiguity caused by the multi-grid algorithm and further improve performance. Code and models are available at https://github.com/LooperXX/ManagerTower. Xiao Xu 0005, Libo Qin 0001, Wanxiang Che, Min-Yen Kan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Augmentation with Neighboring Information for Conversational RecommendationabstractConversational recommender systems (CRSs) suggest items to users by understanding their needs and preferences from natural language conversations. While users can freely express preferences, modeling needs and preferences solely from users’ conversations is challenging due to the sparsity of the available information. Prior work introduces external resources to enrich information expressed in conversations. Obtaining such resources is challenging and not always effective. Can learning intrinsic relations among conversations and items enhance information without the use of external resources? Inspired by collaborative filtering, we propose to use so-called neighboring relations within training data, i.e., relations between conversations, items, and similar conversations and items, to enhance our algorithmic understanding of CRSs. We propose a neighboring relations enhanced conversational recommender system (NR-CRS) and study how neighboring relations improve CRSs from two angles: (i) We mine preference information from neighboring conversations to enhance the modeling of user representations and learning of user preferences. (ii) We generate negative samples based on neighboring items to extend the data available for training CRSs. Experiments on the ReDial dataset show that neighboring relations enhanced conversational recommender system (NR-CRS) outperforms the state-of-the-art baseline by 11.3–20.6% regarding recommendation performance while generating informative and diverse responses. We also assess the capabilities of large language models (i.e., Llama 2, Llama 3, and Chinese-Alpaca2) for CRSs. While the generated responses exhibit enhanced fluency and informativeness, recommending target items with LLMs remains challenging; we recommend that LLMs be used as a decoding base for NR-CRS to generate relevant and informative responses. Yuanxing Liu 0001, Jiahuan Pei, Weinan Zhang 0003, Ming Li 0068, Wanxiang Che, Maarten de Rijke |
ACM Trans. Inf. Syst. | 5 |
| 2024 | Semantic-Guided Generative Image Augmentation Method with Diffusion Models for Image ClassificationabstractExisting image augmentation methods consist of two categories: perturbation-based methods and generative methods. Perturbation-based methods apply pre-defined perturbations to augment an original image, but only locally vary the image, thus lacking image diversity. In contrast, generative methods bring more image diversity in the augmented images but may not preserve semantic consistency, thus may incorrectly change the essential semantics of the original image. To balance image diversity and semantic consistency in augmented images, we propose SGID, a Semantic-guided Generative Image augmentation method with Diffusion models for image classification. Specifically, SGID employs diffusion models to generate augmented images with good image diversity. More importantly, SGID takes image labels and captions as guidance to maintain semantic consistency between the augmented and original images. Experimental results show that SGID outperforms the best augmentation baseline by 1.72% on ResNet-50 (from scratch), 0.33% on ViT (ImageNet-21k), and 0.14% on CLIP-ViT (LAION-2B). Moreover, SGID can be combined with other image augmentation baselines and further improves the overall performance. We demonstrate the semantic consistency and image diversity of SGID through quantitative human and automated evaluations, as well as qualitative case studies. Bohan Li 0010, Xiao Xu 0005, Yutai Hou, Yunlong Feng, Xuanliang Zhang, Qingfu Zhu, Wanxiang Che |
AAAI | 9 |
| 2024 | Exploring Equation as a Better Intermediate Meaning Representation for Numerical Reasoning of Large Language ModelsabstractNumerical reasoning is a vital capability for natural language processing models to understand and process numerical information in real-world scenarios. Most current methods first generate the Intermediate Meaning Representations (IMRs) of questions and then generate answers. Current SOTA methods generate programs as IMRs with large language models (LLMs). Intuitively, equations have fewer restrictions and closer semantics to the question than programs, leading to higher generation accuracy. However, current LLMs generate equations worse than programs, where we assume that the equation data is rare in pre-training data compared to programs. So in this paper, we try to use equations as IMRs to solve the numerical reasoning task by addressing two problems: (1) Theoretically, how to prove that the equation is an IMR with higher generation accuracy than programs; (2) Empirically, how to improve the generation accuracy of equations with LLMs. For the first problem, we propose and prove a proposition to theoretically compare the generation accuracy of different IMRs. For the second problem, we present a method called Boosting Numerical ReasonIng by Decomposing the Generation of Equations Bridge, which can improve the accuracy of LLMs in generating equations as IMRs by reducing the tendency of generating constant expressions and programs. Our method improves the performance by 2.2%, 0.9%, and 1.7% on GSM8K, SVAMP, and Algebra datasets compared to the previous state-of-the-art methods under the single reasoning path setting. Our code and prompts are available at https://github.com/zirui-HIT/Bridge_for_Numerical_Reasoning}. Dingzirui Wang, Longxu Dou, Junyu Zeng, Wanxiang Che |
AAAI | 5 |
| 2024 | Exploring Hybrid Question Answering via Program-based PromptingabstractQuestion answering over heterogeneous data requires reasoning over diverse sources of data, which is challenging due to the large scale of information and organic coupling of heterogeneous data.Various approaches have been proposed to address these challenges.One approach involves training specialized retrievers to select relevant information, thereby reducing the input length.Another approach is to transform diverse modalities of data into a single modality, simplifying the task difficulty and enabling more straightforward processing.In this paper, we propose HPROPRO, a novel program-based prompting framework for the hybrid question answering task.HPRO-PRO follows the code generation and execution paradigm.In addition, HPROPRO integrates various functions to tackle the hybrid reasoning scenario.Specifically, HPROPRO contains function declaration and function implementation to perform hybrid information-seeking over data from various sources and modalities, which enables reasoning over such data without training specialized retrievers or performing modal transformations.Experimental results on two typical hybrid question answering benchmarks HybridQA and MultiModalQA demonstrate the effectiveness of HPROPRO: it surpasses all baseline systems and achieves the best performances in the few-shot settings on both datasets 1 . Qi Shi 0002, Qingfu Zhu, Wanxiang Che, Ting Liu 0001 |
ACL (1) | 5 |
| 2024 | M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-ThoughtabstractMulti-modal Chain-of-Thought (MCoT) requires models to leverage knowledge from both textual and visual modalities for step-bystep reasoning, which gains increasing attention.Nevertheless, the current MCoT benchmark still faces some challenges: (1) absence of visual modal reasoning, (2) single-step visual modal reasoning, and (3) Domain missing, thereby hindering the development of MCoT.Motivated by this, we introduce a novel benchmark (M 3 CoT) to address the above challenges, advancing the multi-domain, multi-step, and multi-modal CoT.Additionally, we conduct a thorough evaluation involving abundant MCoT approaches on Vision Large Language Models (VLLMs).In addition, we highlight that the current VLLMs still struggle to correctly reason in M 3 CoT and there remains a large gap between existing VLLMs and human performance in M 3 CoT, despite their superior results on previous MCoT benchmarks.To our knowledge, we take the first meaningful step toward the multi-domain, multi-step, and multi-modal scenario in MCoT.We hope that M 3 CoT can serve as a valuable resource, providing a pioneering foundation in multi-domain, multi-step, multi-modal chain-of-thought research. Qiguang Chen, Libo Qin 0001, Zhi Chen 0006, Xiao Xu 0005, Wanxiang Che |
ACL (1) | 6 |
| 2024 | Enhancing Numerical Reasoning with the Guidance of Reliable Reasoning ProcessesabstractNumerical reasoning is an essential ability for NLP systems to handle numeric information.Recent research indicates that fine-tuning a small-scale model to learn generating reasoning processes alongside answers can significantly enhance performance.However, current methods have the limitation that most methods generate reasoning processes with large language models (LLMs), which are "unreliable" since such processes could contain information unrelated to the answer.To address this limitation, we introduce Enhancing NumeriCal reasOning with Reliable procEsses (ENCORE), which derives the reliable reasoning process by decomposing the answer formula, ensuring which fully supports the answer.Nevertheless, models could lack enough data to learn the reasoning process generation adequately, since our method generates only one single reasoning process for one formula.To overcome this difficulty, we present a series of pre-training tasks to help models learn the reasoning process generation with synthesized data.The experiments show that ENCORE yields improvement on all five experimental datasets with an average of 1.8%, proving the effectiveness of our method 1 .* Corresponding author. 1 Our code is released in link. 2 For the sake of conciseness in this paper, we collectively refer to these elements as formulas. Dingzirui Wang, Longxu Dou, Xuanliang Zhang, Qingfu Zhu, Wanxiang Che |
ACL (1) | 5 |
| 2024 | SAPT: A Shared Attention Framework for Parameter-Efficient Continual Learning of Large Language ModelsabstractWeixiang Zhao, Shilong Wang, Yulin Hu, Yanyan Zhao, Bing Qin, Xuanyu Zhang, Qing Yang, Dongliang Xu, Wanxiang Che. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Weixiang Zhao, Shilong Wang 0003, Yulin Hu, Bing Qin 0001, Qing Yang 0033, Dongliang Xu, Wanxiang Che |
ACL (1) | 9 |
| 2024 | A Two-Stage Framework with Self-Supervised Distillation for Cross-Domain Text ClassificationabstractCross-domain text classification is a crucial task as it enables models to adapt to a target domain that lacks labeled data. It leverages or reuses rich labeled data from the different but related source domain(s) and unlabeled data from the target domain. To this end, previous work focuses on either extracting domain-invariant features or task-agnostic features, ignoring domain-aware features that may be present in the target domain and could be useful for the downstream task. In this paper, we propose a two-stage framework for cross-domain text classification. In the first stage, we finetune the model with mask language modeling (MLM) and labeled data from the source domain. In the second stage, we further fine-tune the model with self-supervised distillation (SSD) and unlabeled data from the target domain. We evaluate its performance on a public cross-domain text classification benchmark and the experiment results show that our method achieves new state-of-the-art results for both single-source domain adaptations (94.17% +1.03%) and multi-source domain adaptations (95.09% +1.34%). Yunlong Feng, Bohan Li 0010, Libo Qin 0001, Xiao Xu 0005, Wanxiang Che |
LREC/COLING | 5 |
| 2024 | Improving Language Model Reasoning with Self-motivated LearningabstractLarge-scale high-quality training data is important for improving the performance of models. After trained with data that has rationales (reasoning steps), models gain reasoning capability. However, the dataset with high-quality rationales is relatively scarce due to the high annotation cost. To address this issue, we propose Self-motivated Learning framework. The framework motivates the model itself to automatically generate rationales on existing datasets. Based on the inherent rank from correctness across multiple rationales, the model learns to generate better rationales, leading to higher reasoning capability. Specifically, we train a reward model with the rank to evaluate the quality of rationales, and improve the performance of reasoning through reinforcement learning. Experiment results of Llama2 7B on multiple reasoning datasets show that our method significantly improves the reasoning ability of models, even outperforming InstructGPT in some datasets. Yunlong Feng, Yang Xu 0049, Libo Qin 0001, Yasheng Wang, Wanxiang Che |
LREC/COLING | 5 |
| 2024 | Beyond Static Evaluation: A Dynamic Approach to Assessing AI Assistants' API Invocation CapabilitiesabstractWith the rise of Large Language Models (LLMs), AI assistants’ ability to utilize tools, especially through API calls, has advanced notably. This progress has necessitated more accurate evaluation methods. Many existing studies adopt static evaluation, where they assess AI assistants’ API call based on pre-defined dialogue histories. However, such evaluation method can be misleading, as an AI assistant might fail in generating API calls from preceding human interaction in real cases. Instead of the resource-intensive method of direct human-machine interactions, we propose Automated Dynamic Evaluation (AutoDE) to assess an assistant’s API call capability without human involvement. In our framework, we endeavor to closely mirror genuine human conversation patterns in human-machine interactions, using a LLM-based user agent, equipped with a user script to ensure human alignment. Experimental results highlight that AutoDE uncovers errors overlooked by static evaluations, aligning more closely with human assessment. Testing four AI assistants using our crafted benchmark, our method further mirrored human evaluation compared to conventional static evaluations. Honglin Mu, Yang Xu 0049, Yunlong Feng, Yutai Hou, Wanxiang Che |
LREC/COLING | 7 |
| 2024 | LM-Combiner: A Contextual Rewriting Model for Chinese Grammatical Error CorrectionabstractOver-correction is a critical problem in Chinese grammatical error correction (CGEC) task. Recent work using model ensemble methods based on voting can effectively mitigate over-correction and improve the precision of the GEC system. However, these methods still require the output of several GEC systems and inevitably lead to reduced error recall. In this light, we propose the LM-Combiner, a rewriting model that can directly modify the over-correction of GEC system outputs without a model ensemble. Specifically, we train the model on an over-correction dataset constructed through the proposed K-fold cross inference method, which allows it to directly generate filtered sentences by combining the original and the over-corrected text. In the inference stage, we directly take the original sentences and the output results of other systems as input and then obtain the filtered sentences through LM-Combiner. Experiments on the FCGEC dataset show that our proposed method effectively alleviates the over-correction of the original system (+18.2 Precision) while ensuring the error recall remains unchanged. Besides, we find that LM-Combiner still has a good rewriting performance even with small parameters and few training data, and thus can cost-effectively mitigate the over-correction of black-box GEC systems (e.g., ChatGPT). Baoxin Wang, Dayong Wu, Wanxiang Che |
LREC/COLING | 5 |
| 2024 | A Survey on Natural Language Processing for ProgrammingabstractNatural language processing for programming aims to use NLP techniques to assist programming. It is increasingly prevalent for its effectiveness in improving productivity. Distinct from natural language, a programming language is highly structured and functional. Constructing a structure-based representation and a functionality-oriented algorithm is at the heart of program understanding and generation. In this paper, we conduct a systematic review covering tasks, datasets, evaluation methods, techniques, and models from the perspective of the structure-based and functionality-oriented property, aiming to understand the role of the two properties in each component. Based on the analysis, we illustrate unexplored areas and suggest potential directions for future work. Qingfu Zhu, Xianzhen Luo, Fang Liu 0032, Wanxiang Che |
LREC/COLING | 5 |
| 2024 | Python is Not Always the Best Choice: Embracing Multilingual Program of ThoughtsabstractProgram of Thoughts (PoT) is an approach characterized by its executable intermediate steps, which ensure the accuracy of the logical calculations in the reasoning process.Currently, PoT primarily uses Python.However, relying solely on a single language may result in suboptimal solutions and overlook the potential benefits of other programming languages.In this paper, we conduct comprehensive experiments on the programming languages used in PoT and find that no single language consistently delivers optimal performance across all tasks and models.The effectiveness of each language varies depending on the specific scenarios.Inspired by this, we propose a task and model agnostic approach called MultiPoT, which harnesses strength and diversity from various languages.Experimental results reveal that it significantly outperforms Python Self-Consistency.Furthermore, it achieves comparable or superior performance compared to the best monolingual PoT in almost all tasks across all models.In particular, MultiPoT achieves more than 4.6% improvement on average on ChatGPT (gpt-3.5-turbo-0701) 1 . Xianzhen Luo, Qingfu Zhu, Libo Qin 0001, Qing Yang 0033, Dongliang Xu, Wanxiang Che |
EMNLP | 8 |
| 2024 | Make Some Noise: Unlocking Language Model Parallel Inference Capability through Noisy TrainingabstractYixuan Wang, Xianzhen Luo, Fuxuan Wei, Yijun Liu, Qingfu Zhu, Xuanyu Zhang, Qing Yang, Dongliang Xu, Wanxiang Che. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Xianzhen Luo, Fuxuan Wei, Qingfu Zhu, Qing Yang 0033, Dongliang Xu, Wanxiang Che |
EMNLP | 9 |
| 2024 | Pro-HAN: A Heterogeneous Graph Attention Network for Profile-based Spoken Language UnderstandingabstractRecently, Profile-based Spoken Language Understanding (SLU) has gained increasing attention, which aims to incorporate various types of supplementary profile information (i.e., Knowledge Graph, User Profile, Context Awareness) to eliminate the prevalent ambiguities in user utterances. However, existing approaches can only separately model different profile information, without considering their interrelationships or excluding irrelevant and conflicting information within them. To address the above issues, we introduce a Heterogeneous Graph Attention Network to perform reasoning across multiple Profile information, called Pro-HAN. Specifically, we design three types of edges, denoted as intra-Pro, inter-Pro, and utterance-Pro, to capture interrelationships among multiple Pros. We establish a new state-of-the-art on the ProSLU dataset, with an improvement of approximately 8% across all three metrics. Further analysis experiments also confirm the effectiveness of our method in modeling multi-source profile information. Dechuan Teng, Chunlin Lu, Xiao Xu 0005, Wanxiang Che, Libo Qin 0001 |
ICASSP | 4 |
| 2024 | Decoupling Breaks Data Barriers: A Decoupled Pre-training Framework for Multi-intent Spoken Language Understanding
Libo Qin 0001, Qiguang Chen, Jingxuan Zhou, Qinzheng Li, Chunlin Lu, Wanxiang Che |
IJCAI | 6 |
| 2024 | What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationabstractRecently, rapid advancements in Multi-Modal In-Context Learning (MM-ICL) have achieved notable success, which is capable of achieving superior performance across various tasks without requiring additional parameter tuning. However, the underlying rules for the effectiveness of MM-ICL remain under-explored. To fill this gap, this work aims to investigate the research question: "_What factors affect the performance of MM-ICL?_" To this end, we investigate extensive experiments on the three core steps of MM-ICL including demonstration retrieval, demonstration ordering, and prompt construction using 6 vision large language models and 20 strategies. Our findings highlight (1) the necessity of a multi-modal retriever for demonstration retrieval, (2) the importance of intra-demonstration ordering over inter-demonstration ordering, and (3) the enhancement of task comprehension through introductory instructions in prompts. We hope this study can serve as a foundational guide for optimizing MM-ICL strategies in future research. Libo Qin 0001, Qiguang Chen, Hao Fei 0003, Zhi Chen 0006, Min Li 0007, Wanxiang Che |
NeurIPS | 6 |
| 2024 | Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-ThoughtabstractChain-of-Thought (CoT) reasoning has emerged as a promising approach for enhancing the performance of large language models (LLMs) on complex reasoning tasks. Recently, a series of studies attempt to explain the mechanisms underlying CoT, aiming to deepen the understanding of its efficacy. Nevertheless, the existing research faces two major challenges: (1) a lack of quantitative metrics to assess CoT capabilities and (2) a dearth of guidance on optimizing CoT performance. Motivated by this, in this work, we introduce a novel reasoning boundary framework (RBF) to address these challenges. To solve the lack of quantification, we first define a reasoning boundary (RB) to quantify the upper-bound of CoT and establish a combination law for RB, enabling a practical quantitative approach applicable to various real-world CoT tasks. To address the lack of optimization, we propose three categories of RBs. We further optimize these categories with combination laws focused on RB promotion and reasoning path optimization for CoT improvement. Through extensive experiments on 27 models and 5 tasks, the study validates the existence and rationality of the proposed framework. Furthermore, it explains the effectiveness of 10 CoT strategies and guides optimization from two perspectives. We hope this work can provide a comprehensive understanding of the boundaries and optimization strategies for reasoning in LLMs. Our code and data are available at https://github.com/LightChen233/reasoning-boundary. Qiguang Chen, Libo Qin 0001, Jiaqi Wang 0012, Jingxuan Zhou, Wanxiang Che |
NeurIPS | 5 |
| 2024 | OneBit: Towards Extremely Low-bit Large Language ModelsabstractModel quantification uses low bit-width values to represent the weight matrices of existing models to be quantized, which is a promising approach to reduce both storage and computational overheads of deploying highly anticipated LLMs. However, current quantization methods suffer severe performance degradation when the bit-width is extremely reduced, and thus focus on utilizing 4-bit or 8-bit values to quantize models. This paper boldly quantizes the weight matrices of LLMs to 1-bit, paving the way for the extremely low bit-width deployment of LLMs. For this target, we introduce a 1-bit model compressing framework named OneBit, including a novel 1-bit parameter representation method to better quantize LLMs as well as an effective parameter initialization method based on matrix decomposition to improve the convergence speed of the quantization framework. Sufficient experimental results indicate that OneBit achieves good performance (at least 81% of the non-quantized performance on LLaMA models) with robust training processes when only using 1-bit weight matrices. Yuzhuang Xu, Xu Han 0007, Zonghan Yang, Shuo Wang 0013, Qingfu Zhu, Zhiyuan Liu 0001, Wanxiang Che |
NeurIPS | 8 |
| 2023 | Towards Complex Scenarios: Building End-to-End Task-Oriented Dialogue System across Multiple Knowledge BasesabstractWith the success of the sequence-to-sequence model, end-to-end task-oriented dialogue systems (EToDs) have obtained remarkable progress. However, most existing EToDs are limited to single KB settings where dialogues can be supported by a single KB, which is still far from satisfying the requirements of some complex applications (multi-KBs setting). In this work, we first empirically show that the existing single-KB EToDs fail to work on multi-KB settings that require models to reason across various KBs. To solve this issue, we take the first step to consider the multi-KBs scenario in EToDs and introduce a KB-over-KB Heterogeneous Graph Attention Network (KoK-HAN) to facilitate model to reason over multiple KBs. The core module is a triple-connection graph interaction layer that can model different granularity levels of interaction information across different KBs (i.e., intra-KB connection, inter-KB connection and dialogue-KB connection). Experimental results confirm the superiority of our model for multiple KBs reasoning. Libo Qin 0001, Zhouyang Li, Qiying Yu, Lehan Wang, Wanxiang Che |
AAAI | 5 |
| 2023 | BridgeTower: Building Bridges between Encoders in Vision-Language Representation LearningabstractVision-Language (VL) models with the Two-Tower architecture have dominated visual-language representation learning in recent years. Current VL models either use lightweight uni-modal encoders and learn to extract, align and fuse both modalities simultaneously in a deep cross-modal encoder, or feed the last-layer uni-modal representations from the deep pre-trained uni-modal encoders into the top cross-modal encoder. Both approaches potentially restrict vision-language representation learning and limit model performance. In this paper, we propose BridgeTower, which introduces multiple bridge layers that build a connection between the top layers of uni-modal encoders and each layer of the cross-modal encoder. This enables effective bottom-up cross-modal alignment and fusion between visual and textual representations of different semantic levels of pre-trained uni-modal encoders in the cross-modal encoder. Pre-trained with only 4M images, BridgeTower achieves state-of-the-art performance on various downstream vision-language tasks. In particular, on the VQAv2 test-std set, BridgeTower achieves an accuracy of 78.73%, outperforming the previous state-of-the-art model METER by 1.09% with the same pre-training data and almost negligible additional parameters and computational costs. Notably, when further scaling the model, BridgeTower achieves an accuracy of 81.15%, surpassing models that are pre-trained on orders-of-magnitude larger datasets. Code and checkpoints are available at https://github.com/microsoft/BridgeTower. Xiao Xu 0005, Chenfei Wu, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan 0001 |
AAAI | 5 |
| 2023 | MultiSpider: Towards Benchmarking Multilingual Text-to-SQL Semantic ParsingabstractText-to-SQL semantic parsing is an important NLP task, which facilitates the interaction between users and the database. Much recent progress in text-to-SQL has been driven by large-scale datasets, but most of them are centered on English. In this work, we present MultiSpider, the largest multilingual text-to-SQL semantic parsing dataset which covers seven languages (English, German, French, Spanish, Japanese, Chinese, and Vietnamese). Upon MultiSpider we further identify the lexical and structural challenges of text-to-SQL (caused by specific language properties and dialect sayings) and their intensity across different languages. Experimental results under various settings (zero-shot, monolingual and multilingual) reveal a 6.1% absolute drop in accuracy in non-English languages. Qualitative and quantitative analyses are conducted to understand the reason for the performance drop of each language. Besides the dataset, we also propose a simple schema augmentation framework SAVe (Schema-Augmentation-with-Verification), which significantly boosts the overall performance by about 1.8% and closes the 29.5% performance gap across languages. Longxu Dou, Yan Gao 0002, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Jian-Guang Lou |
AAAI | 5 |
| 2023 | ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation LearningabstractXiao Xu, Bei Li, Chenfei Wu, Shao-Yen Tseng, Anahita Bhiwandiwalla, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Xiao Xu 0005, Chenfei Wu, Shao-Yen Tseng, Anahita Bhiwandiwalla, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan 0001 |
ACL (1) | 8 |
| 2023 | Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across LanguagesabstractChain-of-thought (CoT) is capable of eliciting models to explicitly generate reasoning paths, thus promoting reasoning accuracy and attracting increasing attention.Specifically, zeroshot CoT achieves remarkable improvements in a wide range of reasoning tasks by simply instructing the LLM with the prompt "Let's think step by step!".Despite the success of zero-shot CoT, the existing zero-shot prompting techniques remain limited to a single language, making it challenging to generalize to other languages and hindering global development.In this work, we introduce cross-lingual prompting (CLP), aiming to improve zero-shot CoT reasoning across languages.Specifically, CLP consists of two main components: (1) cross-lingual alignment prompting and (2) taskspecific solver prompting.The cross-lingual alignment prompting is responsible for aligning representations across different languages, whereas the task-specific solver prompting is used to generate the final chain of thoughts and results for the reasoning task.In addition, we further introduce cross-lingual self-consistent prompting (CLSP) to ensemble different reasoning paths across languages.Our experimental evaluations on several benchmarks demonstrate that CLP and CLSP significantly outperform the existing prompting methods and achieve state-of-the-art performance.We hope this work will inspire further breakthroughs in cross-lingual CoT.Answer: 48.After understanding, you should act as an expert in arithmetic reasoning in German… Verstanden! Hier ist die Lösung für die gegebene Aufgabe: 1. Da Jiesha 20 Jahre alt ist, ist Qioung 20-5... 2. Da Mary zwei Jahre jünger als Qioung ist... Antwort: 48.Sorry, I am an AI language model and I can only communicate in English.May I assist you with anything else?Let's understand the task in German step-by-step!German Prompting Please act as an expert in multi-lingual understanding in Chinese.Request:玛丽比琼小两岁,琼比杰莎大五岁。如果杰莎 20 岁,这 三个女孩的年龄总和是多少? Let's understand the task in English step-by-step!After understanding, you should act as an expert in arithmetic reasoning in English.Let's resolve the task you understand above step-by-step!Finally, you should format your answer as 'Answer: [num]'.Sure! Let's solve the task step-by-step: 1.According to the first sentence, Mary is two years younger than Joan.Let's suppose Joan's age is J, then Mary's age is J -2. 2. According to the second sentence, Joan is five years older than Jessie.If Jessie's age is x, then Joan's age is x + 5... Answer: 68. Libo Qin 0001, Qiguang Chen, Fuxuan Wei, Shijue Huang, Wanxiang Che |
EMNLP | 5 |
| 2023 | End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future DirectionsabstractEnd-to-end task-oriented dialogue (EToD) can directly generate responses in an end-to-end fashion without modular training, which attracts escalating popularity.The advancement of deep neural networks, especially the successful use of large pre-trained models, has further led to significant progress in EToD research in recent years.In this paper, we present a thorough review and provide a unified perspective to summarize existing approaches as well as recent trends to advance the development of EToD research.The contributions of this paper can be summarized: (1) First survey: to our knowledge, we take the first step to present a thorough survey of this research field; (2) New taxonomy: we first introduce a unified perspective for EToD, including (i) Modularly EToD and (ii) Fully EToD; (3) New Frontiers: we discuss some potential frontier areas as well as the corresponding challenges, hoping to spur breakthrough research in EToD field; (4) Abundant resources: we build a public website 1 , where EToD researchers could directly access the recent progress.We hope this work can serve as a thorough reference for the EToD research community.EToD Modularly EToD ( §3.1) Libo Qin 0001, Wenbo Pan 0001, Qiguang Chen, Lizi Liao, Zhou Yu 0005, Yue Zhang 0004, Wanxiang Che, Min Li 0007 |
EMNLP | 7 |
| 2023 | ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai |
ICDAR (2) | 18 |
| 2023 | Improving Domain Generalization for Sound Classification with Sparse Frequency-Regularized TransformerabstractSound classification models’ performance suffers from generalizing on out-of-distribution (OOD) data. Numerous methods have been proposed to help the model generalize. However, most either introduce inference overheads or focus on long-lasting CNN-variants, while Transformers has been proven to outperform CNNs on numerous natural language processing and computer vision tasks. We propose FRITO, an effective regularization technique on Transformer’s self-attention, to improve the model’s generalization ability by limiting each sequence position’s attention receptive field along the frequency dimension on the spectrogram. Experiments show that our method helps Transformer models achieve SOTA generalization performance on TAU 2020 and Nsynth datasets while saving 20% inference time. Honglin Mu, Wentian Xia, Wanxiang Che |
ICME | 3 |
| 2023 | BEATs: Audio Pre-Training with Acoustic TokenizersabstractWe introduce a self-supervised learning (SSL) framework BEATs for general audio representation pre-training, where we optimize an acoustic tokenizer and an audio SSL model by iterations. Unlike the previous audio SSL models that employ reconstruction loss for pre-training, our audio SSL model is trained with the discrete label prediction task, where the labels are generated by a semantic-rich acoustic tokenizer. We propose an iterative pipeline to jointly optimize the tokenizer and the pre-trained model, aiming to abstract high-level semantics and discard the redundant details for audio. The experimental results demonstrate our acoustic tokenizers can generate discrete labels with rich audio semantics and our audio SSL models achieve state-of-the-art (SOTA) results across various audio classification benchmarks, even outperforming previous models that use more training data and model parameters significantly. Specifically, we set a new SOTA mAP 50.6% on AudioSet-2M without using any external data, and 98.1% accuracy on ESC-50. The code and pre-trained models are available at https://aka.ms/beats. Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Shujie Liu 0001, Daniel Tompkins, Zhuo Chen 0006, Wanxiang Che, Xiangzhan Yu, Furu Wei |
ICML | 7 |
| 2023 | MetricPrompt: Prompting Model as a Relevance Metric for Few-shot Text ClassificationabstractPrompting methods have shown impressive performance in a variety of text mining tasks and applications, especially few-shot ones. Despite the promising prospects, the performance of prompting model largely depends on the design of prompt template and verbalizer. In this work, we propose MetricPrompt, which eases verbalizer design difficulty by reformulating few-shot text classification task into text pair relevance estimation task. MetricPrompt adopts prompting model as the relevance metric, further bridging the gap between Pre-trained Language Model's (PLM) pre-training objective and text classification task, making possible PLM's smooth adaption. Taking a training sample and a query one simultaneously, MetricPrompt captures cross-sample relevance information for accurate relevance estimation. We conduct experiments on three widely used text classification datasets across four few-shot settings. Results show that MetricPrompt outperforms manual verbalizer and other automatic verbalizer design methods across all few-shot settings, achieving new state-of-the-art (SOTA) performance. Hongyuan Dong, Weinan Zhang 0003, Wanxiang Che |
KDD | 3 |
| 2023 | TiBERT: A Non-autoregressive Pre-trained Model for Text Editing
Baoxin Wang, Ziyue Wang 0002, Wanxiang Che, Dayong Wu, Shijin Wang 0001 |
NLPCC (3) | 3 |
| 2023 | U-NEED: A Fine-grained Dataset for User Needs-Centric E-commerce Conversational RecommendationabstractConversational recommender systems ( CRS s) aim to understand the information needs and preferences expressed in a dialogue to recommend suitable items to the user. Most of the existing conversational recommendation datasets are synthesized or simulated with crowdsourcing, which has a large gap with real-world scenarios. To bridge the gap, previous work contributes a dataset E-ConvRec, based on pre-sales dialogues between users and customer service staff in E-commerce scenarios. However, E-ConvRec only supplies coarse-grained annotations and general tasks for making recommendations in pre-sales dialogues. Different from it, we use real user needs as a clue to explore the E-commerce conversational recommendation in complex pre-sales dialogues, namely user needs-centric E-commerce conversational recommendation (UNECR). Yuanxing Liu 0001, Weinan Zhang 0003, Baohua Dong, Yan Fan 0004, Ziyu Zhuang, Hengbin Cui, Yongbin Li 0001, Wanxiang Che |
SIGIR | 11 |
| 2023 | Combating with extremely noisy samples in weakly supervised slot filling for automatic diagnosis
Wanxiang Che |
Frontiers Comput. Sci. | 2 |
| 2023 | Modularized Pre-Training for End-to-End Task-Oriented DialogueabstractPre-training forend-to-endtask-orienteddialoguesystems (EToDs) is a challenging task due to its unique knowledge base query (accuracy) need and lack of sufficient training data (fluency). In this paper, we try to mitigate the above challenges by introducing a modularized pre-training framework for EToDs, which achieves to effectively improve both accuracy and fluency of EToDs through a pre-training paradigm. The core insight is a modular design by decomposing EToDs into ageneration (fluency)module and aknowledge-retriever (accuracy)module, which allows us to optimize each module by pre-training these two sub-modules with different well-designed pre-training tasks, respectively. In addition, such a modularized paradigm enables us to make full use of large amounts of KB-free dialogue corpus for the pre-traininggenerationmodule, which can alleviate the insufficient training problem. Furthermore, we introduce a newconsistency-guideddata augmentation (CGDA) strategy to cope with the data scarcity problem to better pre-train theknowledge-retrievermodule. Finally, we fine-tune the pre-trainedgenerationmodule andknowledge-retrievermodule jointly. Experimental results on three datasets show that our model achieve superior performance in terms of both fluency and accuracy. To our knowledge, this is the first work to explore modularized pre-training methods for EToDs. Libo Qin 0001, Xiao Xu 0005, Lehan Wang, Yue Zhang 0004, Wanxiang Che |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Graph-Grounded Goal Planning for Conversational RecommendationabstractConversational recommendation casts the recommendation problem as a dialog-based interactive task, which could acquire user interest more efficiently and effectively by allowing users to express what they like. In this work, we move a step towards a new conversational recommendation task that is more suitable for real-world applications. In this task, the recommender proactively and naturally lead a dialog from non-recommendation content to approach an item being of interest to users, and allow users to ask questions for better support of user decisions. The challenge of this task lies in how to effectively control the dialog flow to complete the recommendation while appropriately responding to user utterances. To address this challenge, we first construct a Chinese recommendation dialog dataset DuRecDial. We then propose a two-stage Multi-Goal driven Conversation Generation framework, MGCG. In particular, the goal planning module leverages the global graph structure information and local goal-sequence information to effectively control the dialog flow step by step. The goal-guided responding module can produce an in-depth dialog about each goal by fully exploiting hierarchical goal information for response retrieval or generation. Results on DuRecDial demonstrate that MGCG can lead the dialog more proactively and naturally, and complete the recommendation task more effectively. Zeming Liu, Hao Liu 0026, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2022 | Text Is No More Enough! A Benchmark for Profile-Based Spoken Language UnderstandingabstractCurrent researches on spoken language understanding (SLU) heavily are limited to a simple setting: the plain text-based SLU that takes the user utterance as input and generates its corresponding semantic frames (e.g., intent and slots). Unfortunately, such a simple setting may fail to work in complex real-world scenarios when an utterance is semantically ambiguous, which cannot be achieved by the text-based SLU models. In this paper, we first introduce a new and important task, Profile-based Spoken Language Understanding (ProSLU), which requires the model that not only relies on the plain text but also the supporting profile information to predict the correct intents and slots. To this end, we further introduce a large-scale human-annotated Chinese dataset with over 5K utterances and their corresponding supporting profile information (Knowledge Graph (KG), User Profile (UP), Context Awareness (CA)). In addition, we evaluate several state-of-the-art baseline models and explore a multi-level knowledge adapter to effectively incorporate profile information. Experimental results reveal that all existing text-based SLU models fail to work when the utterances are semantically ambiguous and our proposed framework can effectively fuse the supporting information for sentence-level intent detection and token-level slot filling. Finally, we summarize key challenges and provide new points for future directions, which hopes to facilitate the research. Xiao Xu 0005, Libo Qin 0001, Kaiji Chen, Guoxing Wu, Linlin Li 0001, Wanxiang Che |
AAAI | 6 |
| 2022 | GL-CLeF: A Global-Local Contrastive Learning Framework for Cross-lingual Spoken Language UnderstandingabstractLibo Qin, Qiguang Chen, Tianbao Xie, Qixin Li, Jian-Guang Lou, Wanxiang Che, Min-Yen Kan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Libo Qin 0001, Qiguang Chen, Tianbao Xie, Qixin Li, Jian-Guang Lou, Wanxiang Che, Min-Yen Kan |
ACL (1) | 6 |
| 2022 | Simple and Effective Graph-to-Graph Annotation ConversionabstractAnnotation conversion is an effective way to construct datasets under new annotation guidelines based on existing datasets with little human labour. Previous work has been limited in conversion between tree-structured datasets and mainly focused on feature-based models which are not easily applicable to new conversions. In this paper, we propose two simple and effective graph-to-graph annotation conversion approaches, namely Label Switching and Graph2Graph Linear Transformation, which use pseudo data and inherit parameters to guide graph conversions respectively. These methods are able to deal with conversion between graph-structured annotations and require no manually designed features. To verify their effectiveness, we manually construct a graph-structured parallel annotated dataset and evaluate the proposed approaches on it as well as other existing parallel annotated datasets. Experimental results show that the proposed approaches outperform strong baselines with higher conversion score. To further validate the quality of converted graphs, we utilize them to train the target parser and find graphs generated by our approaches lead to higher parsing score than those generated by the baselines. Yuxuan Wang 0001, Zhilin Lei 0001, Yuqiu Ji, Wanxiang Che |
COLING | 4 |
| 2022 | MetaPrompting: Learning to Learn Better PromptsabstractPrompting method is regarded as one of the crucial progress for few-shot nature language processing. Recent research on prompting moves from discrete tokens based “hard prompts” to continuous “soft prompts”, which employ learnable vectors as pseudo prompt tokens and achieve better performance. Though showing promising prospects, these soft-prompting methods are observed to rely heavily on good initialization to take effect. Unfortunately, obtaining a perfect initialization for soft prompts requires understanding of inner language models working and elaborate design, which is no easy task and has to restart from scratch for each new task. To remedy this, we propose a generalized soft prompting method called MetaPrompting, which adopts the well-recognized model-agnostic meta-learning algorithm to automatically find better prompt initialization that facilitates fast adaptation to new prompting tasks. Extensive experiments show MetaPrompting tackles soft prompt initialization problem and brings significant improvement on three different datasets (over 7 points improvement in accuracy for 1-shot setting), achieving new state-of-the-art performance. Yutai Hou, Hongyuan Dong, Bohan Li 0010, Wanxiang Che |
COLING | 5 |
| 2022 | CGIM: A Cycle Guided Interactive Learning Model for Consistency Identification in Task-oriented DialogueabstractConsistency identification in task-oriented dialog (CI-ToD) usually consists of three subtasks, aiming to identify inconsistency between current system response and current user response, dialog history and the corresponding knowledge base. This work aims to solve CI-ToD task by introducing an explicit interaction paradigm, Cycle Guided Interactive learning Model (CGIM), which achieves to make information exchange explicitly from all the three tasks. Specifically, CGIM relies on two core insights, referred to as guided multi-head attention module and cycle interactive mechanism, that collaborate from each other. On the one hand, each two tasks are linked with the guided multi-head attention module, aiming to explicitly model the interaction across two related tasks. On the other hand, we further introduce cycle interactive mechanism that focuses on facilitating model to exchange information among the three correlated sub-tasks via a cycle interaction manner. Experimental results on CI-ToD benchmark show that our model achieves the state-of-the-art performance, pushing the overall score to 56.3% (5.0% point absolute improvement). In addition, we find that CGIM is robust to the initial task flow order. Libo Qin 0001, Qiguang Chen, Tianbao Xie, Qian Liu 0033, Shijue Huang, Wanxiang Che, Zhou Yu 0005 |
COLING | 6 |
| 2022 | CCTC: A Cross-Sentence Chinese Text Correction Dataset for Native SpeakersabstractThe Chinese text correction (CTC) focuses on detecting and correcting Chinese spelling errors and grammatical errors. Most existing datasets of Chinese spelling check (CSC) and Chinese grammatical error correction (GEC) are focused on a single sentence written by Chinese-as-a-second-language (CSL) learners. We find that errors caused by native speakers differ significantly from those produced by non-native speakers. These differences make it inappropriate to use the existing test sets directly to evaluate text correction systems for native speakers. Some errors also require the cross-sentence information to be identified and corrected. In this paper, we propose a cross-sentence Chinese text correction dataset for native speakers. Concretely, we manually annotated 1,500 texts written by native speakers. The dataset consists of 30,811 sentences and more than 1,000,000 Chinese characters. It contains four types of errors: spelling errors, redundant words, missing words, and word ordering errors. We also test some state-of-the-art models on the dataset. The experimental results show that even the model with the best performance is 20 points lower than humans, which indicates that there is still much room for improvement. We hope that the new dataset can fill the gap in cross-sentence text correction for native Chinese speakers. Baoxin Wang, Xingyi Duan, Dayong Wu, Wanxiang Che, Zhigang Chen 0003 |
COLING | 4 |
| 2022 | Adaptive Unsupervised Self-training for Disfluency DetectionabstractSupervised methods have achieved remarkable results in disfluency detection. However, in real-world scenarios, human-annotated data is difficult to obtain. Recent works try to handle disfluency detection with unsupervised self-training, which can exploit existing large-scale unlabeled data efficiently. However, their self-training-based methods suffer from the problems of selection bias and error accumulation. To tackle these problems, we propose an adaptive unsupervised self-training method for disfluency detection. Specifically, we re-weight the importance of each training example according to its grammatical feature and prediction confidence. Experiments on the Switchboard dataset show that our method improves 2.3 points over the current SOTA unsupervised method. Moreover, our method is competitive with the SOTA supervised method. Shaolei Wang, Wanxiang Che |
COLING | 4 |
| 2022 | Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic KnowledgeabstractLongxu Dou, Yan Gao, Xuqi Liu, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Min-Yen Kan, Jian-Guang Lou. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Longxu Dou, Yan Gao 0002, Xuqi Liu, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Min-Yen Kan, Jian-Guang Lou |
EMNLP | 6 |
| 2022 | Interactive Gated Decoder for Machine Reading ComprehensionabstractOwing to the availability of various large-scale Machine Reading Comprehension ( MRC ) datasets, building an effective model to extract passage spans for question answering has been well studied in previous works. However, in reality, there are some questions that cannot be answered through the passage information, which brings more challenges to this task. In this article, we propose an Interactive Gated Decoder ( IG Decoder ), which focuses on modeling the interactions between the answer span prediction and no-answer prediction with a gating mechanism. We also propose a simple but effective approach for automatically generating pseudo training data, which aims to enrich the training data of the unanswerable questions. Experimental results on popular benchmark SQuAD 2.0 and NewsQA show that the proposed approaches yield consistent improvements over traditional BERT-large and strong ALBERT-xxlarge baseline systems. We also provide detailed ablations of the proposed method and error analysis on hard samples, which could be helpful in future research. Yiming Cui 0001, Wanxiang Che, Ziqing Yang 0001, Ting Liu 0001, Bing Qin 0001, Shijin Wang 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Multi-domain Spoken Language Understanding Using Domain- and Task-aware ParameterizationabstractSpoken language understanding (SLU) has been addressed as a supervised learning problem, where a set of training data is available for each domain. However, annotating data for a new domain can be both financially costly and non-scalable. One existing approach solves the problem by conducting multi-domain learning where parameters are shared for joint training across domains, which is domain-agnostic and task-agnostic . In the article, we propose to improve the parameterization of this method by using domain-specific and task-specific model parameters for fine-grained knowledge representation and transfer. Experiments on five domains show that our model is more effective for multi-domain SLU and obtain the best results. In addition, we show its transferability when adapting to a new domain with little data, outperforming the prior best model by 12.4%. Finally, we explore the strong pre-trained model in our framework and find that the contributions from our framework do not fully overlap with contextualized word representations (RoBERTa). Libo Qin 0001, Fuxuan Wei, Minheng Ni, Yue Zhang 0004, Wanxiang Che, Yangming Li, Ting Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2022 | Combining Self-supervised Learning and Active Learning for Disfluency DetectionabstractSpoken language is fundamentally different from the written language in that it contains frequent disfluencies or parts of an utterance that are corrected by the speaker. Disfluency detection (removing these disfluencies) is desirable to clean the input for use in downstream NLP tasks. Most existing approaches to disfluency detection heavily rely on human-annotated data, which is scarce and expensive to obtain in practice. To tackle the training data bottleneck, in this work, we investigate methods for combining self-supervised learning and active learning for disfluency detection. First, we construct large-scale pseudo training data by randomly adding or deleting words from unlabeled data and propose two self-supervised pre-training tasks: (i) a tagging task to detect the added noisy words and (ii) sentence classification to distinguish original sentences from grammatically incorrect sentences. We then combine these two tasks to jointly pre-train a neural network. The pre-trained neural network is then fine-tuned using human-annotated disfluency detection training data. The self-supervised learning method can capture task-special knowledge for disfluency detection and achieve better performance when fine-tuning on a small annotated dataset compared to other supervised methods. However, limited in that the pseudo training data are generated based on simple heuristics and cannot fully cover all the disfluency patterns, there is still a performance gap compared to the supervised models trained on the full training dataset. We further explore how to bridge the performance gap by integrating active learning during the fine-tuning process. Active learning strives to reduce annotation costs by choosing the most critical examples to label and can address the weakness of self-supervised learning with a small annotated dataset. We show that by combining self-supervised learning with active learning, our model is able to match state-of-the-art performance with just about 10% of the original training data on both the commonly used English Switchboard test set and a set of in-house annotated Chinese data. Shaolei Wang, Wanxiang Che, Sendong Zhao, Ting Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2022 | Teaching Machines to Read, Answer and ExplainabstractWith various Pre-trained Language Models (PLMs) blooming, Machine Reading Comprehension (MRC) systems have embraced significant improvements on various benchmarks and even surpassed human performances. However, most existing works only focus on the accuracy of the answer predictions and neglect the importance of the explanations for the prediction, which is a big obstacle when utilizing these models in real-life applications to convince humans. This paper proposes a novel unsupervised self-explainable framework, called Recursive Dynamic Gating (RDG), for the machine reading comprehension task. The main idea is that the proposed system tries to use less passage information and achieves similar results to the system that uses the whole passage, while the filtered passage is used as text explanations. We carried out experiments on three multiple-choice MRC datasets (including English and Chinese) and found that the proposed system can not only achieve better performance in answer prediction but also provide informative explanations compared to the attention mechanism. Yiming Cui 0001, Ting Liu 0001, Wanxiang Che, Zhigang Chen 0003, Shijin Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Understanding Patient Query With Weak Supervision From Doctor ResponseabstractCurrently, the need for high-quality dialogue systems that assist users to conduct self-diagnosis is rapidly increasing. Slot filling for automatic diagnosis, which converts medical queries into structured representations, plays an important role in diagnostic dialogue systems. However, the lack of high-quality datasets limits the performance of slot filling. While medical communities like AskAPatient usually have multiple rounds of diagnostic dialogue containing colloquial input and professional responses from doctors. Therefore, the data of diagnostic dialogue in medical communities can be utilized to solve the main challenges in slot filling. This paper proposes a two-step training framework to make full use of these unlabeled dialogue data in medical communities. To promote further researches, we provide a Chinese dataset with 2,652 annotated samples and a large amount of unlabeled samples. Experimental results on the dataset demonstrate the effectiveness of the proposed method with an increase of 6.32% in Micro F1 and 8.20% in Macro F1 on average over strong baselines. Sendong Zhao, Yuxuan Wang 0001, Xi Chen 0003, Yefeng Zheng 0001, Wanxiang Che |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | C2C-GenDA: Cluster-to-Cluster Generation for Data Augmentation of Slot FillingabstractSlot filling, a fundamental module of spoken language understanding, often suffers from insufficient quantity and diversity of training data. To remedy this, we propose a novel Cluster-to-Cluster generation framework for Data Augmentation (DA), named C2C-GenDA. It enlarges the training set by reconstructing existing utterances into alternative expressions while keeping semantic. Different from previous DA works that reconstruct utterances one by one independently, C2C-GenDA jointly encodes multiple existing utterances of the same semantics and simultaneously decodes multiple unseen expressions. Jointly generating multiple new utterances allows to consider the relations between generated instances and encourages diversity. Besides, encoding multiple existing utterances endows C2C with a wider view of existing expressions, helping to reduce generation that duplicates existing data. Experiments on ATIS and Snips datasets show that instances augmented by C2C-GenDA improve slot filling by 7.99 (11.9%↑) and 5.76 (13.6%↑) F-scores respectively, when there are only hundreds of training utterances. Code: https://github.com/Sanyuan-Chen/C2C-DA. Yutai Hou, Sanyuan Chen, Wanxiang Che, Ting Liu 0001 |
AAAI | 3 |
| 2021 | Few-shot Learning for Multi-label Intent DetectionabstractIn this paper, we study the few-shot multi-label classification for user intent detection. For multi-label intent detection, state-of-the-art work estimates label-instance relevance scores and uses a threshold to select multiple associated intent labels. To determine appropriate thresholds with only a few examples, we first learn universal thresholding experience on data-rich domains, and then adapt the thresholds to certain few-shot domains with a calibration based on nonparametric learning. For better calculation of label-instance relevance score, we introduce label name embedding as anchor points in representation space, which refines representations of different classes to be well-separated from each other. Experiments on two datasets show that the proposed model significantly outperforms strong baselines in both one-shot and five-shot settings. Yutai Hou, Yongkui Lai, Yushan Wu, Wanxiang Che, Ting Liu 0001 |
AAAI | 4 |
| 2021 | Co-GAT: A Co-Interactive Graph Attention Network for Joint Dialog Act Recognition and Sentiment ClassificationabstractIn a dialog system, dialog act recognition and sentiment classification are two correlative tasks to capture speakers’ intentions, where dialog act and sentiment can indicate the explicit and the implicit intentions separately. The dialog context information (contextual information) and the mutual interaction information are two key factors that contribute to the two related tasks. Unfortunately, none of the existing approaches consider the two important sources of information simultaneously. In this paper, we propose a Co-Interactive Graph Attention Network (Co-GAT) to jointly perform the two tasks. The core module is a proposed co-interactive graph interaction layer where a cross-utterances connection and a cross-tasks connection are constructed and iteratively updated with each other, achieving to consider the two types of information simultaneously. Experimental results on two public datasets show that our model successfully captures the two sources of information and achieve the state-of-the-art performance. In addition, we find that the contributions from the contextual and mutual interaction information do not fully overlap with contextualized word representations (BERT, Roberta, XLNet). Libo Qin 0001, Zhouyang Li, Wanxiang Che, Minheng Ni, Ting Liu 0001 |
AAAI | 3 |
| 2021 | Discovering Dialog Structure Graph for Coherent Dialog GenerationabstractJun Xu, Zeyang Lei, Haifeng Wang, Zheng-Yu Niu, Hua Wu, Wanxiang Che. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
ACL/IJCNLP (1) | 6 |
| 2021 | GL-GIN: Fast and Accurate Non-Autoregressive Model for Joint Multiple Intent Detection and Slot FillingabstractLibo Qin, Fuxuan Wei, Tianbao Xie, Xiao Xu, Wanxiang Che, Ting Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Libo Qin 0001, Fuxuan Wei, Tianbao Xie, Xiao Xu 0005, Wanxiang Che, Ting Liu 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingabstractYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, Lidong Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yang Xu 0049, Yiheng Xu, Tengchao Lv, Lei Cui 0001, Furu Wei, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Wanxiang Che, Min Zhang 0005, Lidong Zhou |
ACL/IJCNLP (1) | 10 |
| 2021 | Consistency Regularization for Cross-Lingual Fine-TuningabstractBo Zheng, Li Dong, Shaohan Huang, Wenhui Wang, Zewen Chi, Saksham Singhal, Wanxiang Che, Ting Liu, Xia Song, Furu Wei. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bo Zheng 0010, Li Dong 0004, Shaohan Huang, Wenhui Wang 0003, Zewen Chi, Saksham Singhal, Wanxiang Che, Ting Liu 0001, Furu Wei |
ACL/IJCNLP (1) | 7 |
| 2021 | DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational RecommendationabstractIn this paper, we provide a bilingual parallel human-to-human recommendation dialog dataset (DuRecDial 2.0) to enable researchers to explore a challenging task of multilingual and cross-lingual conversational recommendation.The difference between DuRecDial 2.0 and existing conversational recommendation datasets is that the data item (Profile, Goal, Knowledge, Context, Response) in DuRecDial 2.0 is annotated in two languages, both English and Chinese, while other datasets are built with the setting of a single language.We collect 8.2k dialogs aligned across English and Chinese languages (16.5k dialogs and 255k utterances in total) that are annotated by crowdsourced workers with strict quality control procedure.We then build monolingual, multilingual, and cross-lingual conversational recommendation baselines on DuRecDial 2.0.Experiment results show that the use of additional English data can bring performance improvement for Chinese conversational recommendation, indicating the benefits of DuRecDial 2.0.Finally, this dataset provides a challenging testbed for future studies of monolingual, multilingual, and cross-lingual conversational recommendation. 1 Zeming Liu, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
EMNLP (1) | 5 |
| 2021 | Don't be Contradicted with Anything! CI-ToD: Towards Benchmarking Consistency for Task-oriented Dialogue SystemabstractConsistency Identification has obtained remarkable success on open-domain dialogue, which can be used for preventing inconsistent response generation.However, in contrast to the rapid development in open-domain dialogue, few efforts have been made to the task-oriented dialogue direction.In this paper, we argue that consistency problem is more urgent in task-oriented domain.To facilitate the research, we introduce CI-ToD, a novel dataset for Consistency Identification in Taskoriented Dialog system.In addition, we not only annotate the single label to enable the model to judge whether the system response is contradictory, but also provide more finegrained labels (i.e., Dialogue History Inconsistency, User Query Inconsistency and Knowledge Base Inconsistency) to encourage model to know what inconsistent sources lead to it.Empirical results show that state-of-the-art methods only achieve 51.3%, which is far behind the human performance of 93.2%, indicating that there is ample room for improving consistency identification ability.Finally, we conduct exhaustive experiments and qualitative analysis to comprehend key challenges and provide guidance for future directions.All datasets and models are publicly available at https://github.com/yizhen20133868/CI-ToD. * Email corresponding. User: Give me directions to the closest grocery store.System: There is a whole foods 2 miles away and their address is 880_ames_ct.User: I need a route that avoids all heavy traffic.System: mandarin roots is located at 271 springer street. Libo Qin 0001, Tianbao Xie, Shijue Huang, Qiguang Chen, Xiao Xu 0005, Wanxiang Che |
EMNLP (1) | 6 |
| 2021 | Allocating Large Vocabulary Capacity for Cross-Lingual Language Model Pre-TrainingabstractCompared to monolingual models, crosslingual models usually require a more expressive vocabulary to represent all languages adequately.We find that many languages are under-represented in recent cross-lingual language models due to the limited vocabulary capacity.To this end, we propose an algorithm VOCAP to determine the desired vocabulary capacity of each language.However, increasing the vocabulary size significantly slows down the pre-training speed.In order to address the issues, we propose k-NN-based target sampling to accelerate the expensive softmax.Our experiments show that the multilingual vocabulary learned with VOCAP benefits cross-lingual language model pre-training.Moreover, k-NN-based target sampling mitigates the side-effects of increasing the vocabulary size while achieving comparable performance and faster pre-training speed.The code and the pretrained multilingual vocabularies are available at https://github. com/bozheng-hit/VoCapXLM. Bo Zheng 0010, Li Dong 0004, Shaohan Huang, Saksham Singhal, Wanxiang Che, Ting Liu 0001, Furu Wei |
EMNLP (1) | 5 |
| 2021 | A Co-Interactive Transformer for Joint Slot Filling and Intent DetectionabstractIntent detection and slot filling are two main tasks for building a spoken language understanding (SLU) system. The two tasks are closely related and the information of one task can benefit the other. Previous studies either implicitly model the two tasks with multi-task framework or only explicitly consider the single information flow from intent to slot. None of the prior approaches model the bidirectional connection between the two tasks simultaneously in a unified framework. In this paper, we propose a Co-Interactive Transformer which considers the cross-impact between the two tasks. Instead of adopting the self-attention mechanism in vanilla Transformer, we propose a co-interactive module to consider the cross-impact by building a bidirectional connection between the two related tasks, where slot and intent can be able to attend on the corresponding mutual information. The experimental results on two public datasets show that our model achieves the state-of-the-art performance. Libo Qin 0001, Tailu Liu, Wanxiang Che, Bingbing Kang, Sendong Zhao, Ting Liu 0001 |
ICASSP | 3 |
| 2021 | Injecting Word Information with Multi-Level Word Adapter for Chinese Spoken Language UnderstandingabstractIn this paper, we improve Chinese spoken language understanding (SLU) by injecting word information. Previous studies on Chinese SLU do not consider the word information, failing to detect word boundaries that are beneficial for intent detection and slot filling. To address this issue, we propose a multi-level word adapter to inject word information for Chinese SLU, which consists of (1) sentence-level word adapter, which directly fuses the sentence representations of the word information and character information to perform intent detection and (2) character-level word adapter, which is applied at each character for selectively controlling weights on word information as well as character information. Experimental results on two Chinese SLU datasets show that our model can capture useful word information and achieve state-of-the-art performance. Dechuan Teng, Libo Qin 0001, Wanxiang Che, Sendong Zhao, Ting Liu 0001 |
ICASSP | 3 |
| 2021 | A Survey on Spoken Language Understanding: Recent Advances and New FrontiersabstractSpoken Language Understanding (SLU) aims to extract the semantics frame of user queries, which is a core component in a task-oriented dialog system. With the burst of deep neural networks and the evolution of pre-trained language models, the research of SLU has obtained significant breakthroughs. However, there remains a lack of a comprehensive survey summarizing existing approaches and recent trends, which motivated the work presented in this article. In this paper, we survey recent advances and new frontiers in SLU. Specifically, we give a thorough review of this research field, covering different aspects including (1) new taxonomy: we provide a new perspective for SLU filed, including single model vs. joint model, implicit joint modeling vs. explicit joint modeling in joint model, non pre-trained paradigm vs. pretrained paradigm; (2) new frontiers: some emerging areas in complex SLU as well as the corresponding challenges; (3) abundant open-source resources: to help the community, we have collected, organized the related papers, baseline projects and leaderboard on a public website where SLU researchers could directly access to the recent progress. We hope that this survey can shed a light on future research in SLU field. Libo Qin 0001, Tianbao Xie, Wanxiang Che, Ting Liu 0001 |
IJCAI | 3 |
| 2021 | Coherent Dialog Generation with Query GraphabstractLearning to generate coherent and informative dialogs is an enduring challenge for open-domain conversation generation. Previous work leverage knowledge graph or documents to facilitate informative dialog generation, with little attention on dialog coherence. In this article, to enhance multi-turn open-domain dialog coherence, we propose to leverage a new knowledge source, web search session data, to facilitate hierarchical knowledge sequence planning, which determines a sketch of a multi-turn dialog. Specifically, we formulate knowledge sequence planning or dialog policy learning as a graph grounded Reinforcement Learning (RL) problem. To this end, we first build a two-level query graph with queries as utterance-level vertices and their topics (entities in queries) as topic-level vertices. We then present a two-level dialog policy model that plans a high-level topic sequence and a low-level query sequence over the query graph to guide a knowledge aware response generator. In particular, to foster forward-looking knowledge planning decisions for better dialog coherence, we devise a heterogeneous graph neural network to incorporate neighbouring vertex information, or possible future RL action information, into each vertex (as an RL action) representation. Experiment results on two benchmark dialog datasets demonstrate that our framework can outperform strong baselines in terms of dialog coherence, informativeness, and engagingness. Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Jizhou Huang, Ting Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2021 | Pre-Training With Whole Word Masking for Chinese BERTabstractBidirectional Encoder Representations from Transformers (BERT) has shown marvelous improvements across various NLP tasks, and its consecutive variants have been proposed to further improve the performance of the pre-trained language models. In this paper, we aim to first introduce the whole word masking (wwm) strategy for Chinese BERT, along with a series of Chinese pre-trained language models. Then we also propose a simple but effective model called MacBERT, which improves upon RoBERTa in several ways. Especially, we propose a new masking strategy called MLM as correction (Mac). To demonstrate the effectiveness of these models, we create a series of Chinese pre-trained language models as our baselines, including BERT, RoBERTa, ELECTRA, RBT, etc. We carried out extensive experiments on ten Chinese NLP tasks to evaluate the created Chinese pre-trained language models as well as the proposed MacBERT. Experimental results show that MacBERT could achieve state-of-the-art performances on many NLP tasks, and we also ablate details with several findings that may help future research. We open-source our pre-trained language models for further facilitating our research community. Yiming Cui 0001, Wanxiang Che, Ting Liu 0001, Bing Qin 0001, Ziqing Yang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Knowing Where to Leverage: Context-Aware Graph Convolutional Network With an Adaptive Fusion Layer for Contextual Spoken Language UnderstandingabstractSpoken language understanding (SLU) systems aim to understand users’ utterance, which is a key component of task-oriented dialogue systems. In this paper, we focus on improving the contextual SLU. The contextual SLU systems mainly focus on how to effectively incorporate dialog context information (contextual information). The existing approaches all use the same contextual information to guide slot filling at all tokens, which may inject the irrelevant information and result in ambiguity. To tackle this problem, we propose a context-aware graph convolutional network (GCN) with an adaptive fusion layer for contextual SLU. The context-aware GCN is proposed to automatically aggregate the contextual information, which frees our model from the manually designed heuristic aggregation function. Meanwhile, an adaptive fusion layer is applied at each token to dynamically incorporate relevant contextual information, which achieves a fine-grained contextual information transfer to guide the token-level slot filling. Experiments on the Simulated Dialog Dataset show that our model achieves state-of-the-art performance and outperforms other previous methods by a large margin (+3.67% on Sim-R, +4.18% on Sim-M and +3.75% on Overall dataset). In addition, we explore and analyze the pre-trained model (i.e., BERT) in our framework. We show that incorporating BERT brings a large improvement in low-resource setting. Libo Qin 0001, Wanxiang Che, Minheng Ni, Yangming Li, Ting Liu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Discriminative Sentence Modeling for Story Ending PredictionabstractStory Ending Prediction is a task that needs to select an appropriate ending for the given story, which requires the machine to understand the story and sometimes needs commonsense knowledge. To tackle this task, we propose a new neural network called Diff-Net for better modeling the differences of each ending in this task. The proposed model could discriminate two endings in three semantic levels: contextual representation, story-aware representation, and discriminative representation. Experimental results on the Story Cloze Test dataset show that the proposed model siginificantly outperforms various systems by a large margin, and detailed ablation studies are given for better understanding our model. We also carefully examine the traditional and BERT-based models on both SCT v1.0 and v1.5 with interesting findings that may potentially help future studies. Yiming Cui 0001, Wanxiang Che, Weinan Zhang 0003, Ting Liu 0001, Shijin Wang 0001 |
AAAI | 2 |
| 2020 | DCR-Net: A Deep Co-Interactive Relation Network for Joint Dialog Act Recognition and Sentiment ClassificationabstractIn dialog system, dialog act recognition and sentiment classification are two correlative tasks to capture speakers' intentions, where dialog act and sentiment can indicate the explicit and the implicit intentions separately (Kim and Kim 2018). Most of the existing systems either treat them as separate tasks or just jointly model the two tasks by sharing parameters in an implicit way without explicitly modeling mutual interaction and relation. To address this problem, we propose a Deep Co-Interactive Relation Network (DCR-Net) to explicitly consider the cross-impact and model the interaction between the two tasks by introducing a co-interactive relation layer. In addition, the proposed relation layer can be stacked to gradually capture mutual knowledge with multiple steps of interaction. Especially, we thoroughly study different relation layers and their effects. Experimental results on two public datasets (Mastodon and Dailydialog) show that our model outperforms the state-of-the-art joint model by 4.3% and 3.4% in terms of F1 score on dialog act recognition task, 5.7% and 12.4% on sentiment classification respectively. Comprehensive analysis empirically verifies the effectiveness of explicitly modeling the relation between the two tasks and the multi-steps interaction mechanism. Finally, we employ the Bidirectional Encoder Representation from Transformer (BERT) in our framework, which can further boost our performance in both tasks. Libo Qin 0001, Wanxiang Che, Yangming Li, Minheng Ni, Ting Liu 0001 |
AAAI | 2 |
| 2020 | Understanding Medical Conversations with Scattered Keyword Attention and Weak Supervision from ResponsesabstractIn this work, we consider the medical slot filling problem, i.e., the problem of converting medical queries into structured representations which is a challenging task. We analyze the effectiveness of two points: scattered keywords in user utterances and weak supervision with responses. We approach the medical slot filling as a multi-label classification problem with label-embedding attentive model to pay more attention to scattered medical keywords and learn the classification models by weak-supervision from responses. To evaluate the approaches, we annotate a medical slot filling data and collect a large scale unlabeled data. The experiments demonstrate that these two points are promising to improve the task. Haifeng Hu 0009, Wanxiang Che, Zhongqian Sun, Ting Liu 0001, Junzhou Huang |
AAAI | 3 |
| 2020 | Multi-Task Self-Supervised Learning for Disfluency DetectionabstractMost existing approaches to disfluency detection heavily rely on human-annotated data, which is expensive to obtain in practice. To tackle the training data bottleneck, we investigate methods for combining multiple self-supervised tasks-i.e., supervised tasks where data can be collected without manual labeling. First, we construct large-scale pseudo training data by randomly adding or deleting words from unlabeled news data, and propose two self-supervised pre-training tasks: (i) tagging task to detect the added noisy words. (ii) sentence classification to distinguish original sentences from grammatically-incorrect sentences. We then combine these two tasks to jointly train a network. The pre-trained network is then fine-tuned using human-annotated disfluency detection training data. Experimental results on the commonly used English Switchboard test set show that our approach can achieve competitive performance compared to the previous systems (trained using the full dataset) by using less than 1% (1000 sentences) of the training data. Our method trained on the full dataset significantly outperforms previous methods, reducing the error by 21% on English Switchboard. Shaolei Wang, Wanxiang Che, Qi Liu 0049, Pengda Qin, Ting Liu 0001, William Yang Wang |
AAAI | 2 |
| 2020 | Knowledge Graph Grounded Goal Planning for Open-Domain Conversation GenerationabstractPrevious neural models on open-domain conversation generation have no effective mechanisms to manage chatting topics, and tend to produce less coherent dialogs. Inspired by the strategies in human-human dialogs, we divide the task of multi-turn open-domain conversation generation into two sub-tasks: explicit goal (chatting about a topic) sequence planning and goal completion by topic elaboration. To this end, we propose a three-layer Knowledge aware Hierarchical Reinforcement Learning based Model (KnowHRL). Specifically, for the first sub-task, the upper-layer policy learns to traverse a knowledge graph (KG) in order to plan a high-level goal sequence towards a good balance between dialog coherence and topic consistency with user interests. For the second sub-task, the middle-layer policy and the lower-layer one work together to produce an in-depth multi-turn conversation about a single topic with a goal-driven generation mechanism. The capability of goal-sequence planning enables chatbots to conduct proactive open-domain conversations towards recommended topics, which has many practical applications. Experiments demonstrate that our model outperforms state of the art baselines in terms of user-interest consistency, dialog coherence, and knowledge accuracy. Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
AAAI | 5 |
| 2020 | Few-shot Slot Tagging with Collapsed Dependency Transfer and Label-enhanced Task-adaptive Projection NetworkabstractIn this paper, we explore the slot tagging with only a few labeled support sentences (a.k.a.few-shot).Few-shot slot tagging faces a unique challenge compared to the other fewshot classification problems as it calls for modeling the dependencies between labels.But it is hard to apply previously learned label dependencies to an unseen domain, due to the discrepancy of label sets.To tackle this, we introduce a collapsed dependency transfer mechanism into the conditional random field (CRF) to transfer abstract label dependency patterns as transition scores.In the few-shot setting, the emission score of CRF can be calculated as a word's similarity to the representation of each label.To calculate such similarity, we propose a Label-enhanced Task-Adaptive Projection Network (L-TapNet) based on the stateof-the-art few-shot classification model -Tap-Net, by leveraging label name semantics in representing labels.Experimental results show that our model significantly outperforms the strongest few-shot learning baseline by 14.64 F1 scores in the one-shot setting. 1 Yutai Hou, Wanxiang Che, Yongkui Lai, Zhihan Zhou 0001, Han Liu 0001, Ting Liu 0001 |
ACL | 2 |
| 2020 | Slot-consistent NLG for Task-oriented Dialogue Systems with Iterative Rectification NetworkabstractData-driven approaches using neural networks have achieved promising performances in natural language generation (NLG).However, neural generators are prone to make mistakes, e.g., neglecting an input slot value and generating a redundant slot value.Prior works refer this to hallucination phenomenon.In this paper, we study slot consistency for building reliable NLG systems with all slot values of input dialogue act (DA) properly generated in output sentences.We propose Iterative Rectification Network (IRN) for improving general NLG systems to produce both correct and fluent responses.It applies a bootstrapping algorithm to sample training candidates and uses reinforcement learning to incorporate discrete reward related to slot inconsistency into training.Comprehensive studies have been conducted on multiple benchmark datasets, showing that the proposed methods have significantly reduced the slot error rate (ERR) for all strong baselines.Human evaluations also have confirmed its effectiveness. Yangming Li, Kaisheng Yao, Libo Qin 0001, Wanxiang Che, Xiaolong Li 0005, Ting Liu 0001 |
ACL | 4 |
| 2020 | Towards Conversational Recommendation over Multi-Type DialogsabstractWe focus on the study of conversational recommendation in the context of multi-type dialogs, where the bots can proactively and naturally lead a conversation from a nonrecommendation dialog (e.g., QA) to a recommendation dialog, taking into account user's interests and feedback.To facilitate the study of this task, we create a human-to-human Chinese dialog dataset DuRecDial (about 10k dialogs, 156k utterances), which contains multiple sequential dialogs for every pair of a recommendation seeker (user) and a recommender (bot).In each dialog, the recommender proactively leads a multi-type dialog to approach recommendation targets and then makes multiple recommendations with rich interaction behavior.This dataset allows us to systematically investigate different parts of the overall problem, e.g., how to naturally lead a dialog, how to interact with users for recommendation.Finally we establish baseline results on DuRecDial for future studies. 1 Zeming Liu, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001 |
ACL | 5 |
| 2020 | Dynamic Fusion Network for Multi-Domain End-to-end Task-Oriented DialogabstractRecent studies have shown remarkable success in end-to-end task-oriented dialog system. However, most neural models rely on large training data, which are only available for a certain number of task domains, such as navigation and scheduling. This makes it difficult to scalable for a new domain with limited labeled data. However, there has been relatively little research on how to effectively use data from all domains to improve the performance of each domain and also unseen domains. To this end, we investigate methods that can make explicit use of domain knowledge and introduce a shared-private network to learn shared and specific knowledge. In addition, we propose a novel Dynamic Fusion Network (DF-Net) which automatically exploit the relevance between the target domain and each domain. Results show that our models outperforms existing methods on multi-domain dialogue, giving the state-of-the-art in the literature. Besides, with little training data, we show its transferability by outperforming prior best model by 13.9% on average. Libo Qin 0001, Xiao Xu 0005, Wanxiang Che, Yue Zhang 0004, Ting Liu 0001 |
ACL | 3 |
| 2020 | Conversational Graph Grounded Policy Learning for Open-Domain Conversation GenerationabstractTo address the challenge of policy learning in open-domain multi-turn conversation, we propose to represent prior information about dialog transitions as a graph and learn a graph grounded dialog policy, aimed at fostering a more coherent and controllable dialog.To this end, we first construct a conversational graph (CG) from dialog corpora, in which there are vertices to represent "what to say" and "how to say", and edges to represent natural transition between a message (the last utterance in a dialog context) and its response.We then present a novel CG grounded policy learning framework that conducts dialog flow planning by graph traversal, which learns to identify a what-vertex and a how-vertex from the CG at each turn to guide response generation.In this way, we effectively leverage the CG to facilitate policy learning as follows: (1) it enables more effective long-term reward design, (2) it provides high-quality candidate actions, and (3) it gives us more control over the policy.Results on two benchmark corpora demonstrate the effectiveness of this framework. Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001 |
ACL | 5 |
| 2020 | Document Modeling with Graph Attention Networks for Multi-grained Machine Reading ComprehensionabstractNatural Questions is a new challenging machine reading comprehension benchmark with two-grained answers, which are a long answer (typically a paragraph) and a short answer (one or more entities inside the long answer).Despite the effectiveness of existing methods on this benchmark, they treat these two sub-tasks individually during training while ignoring their dependencies.To address this issue, we present a novel multi-grained machine reading comprehension framework that focuses on modeling documents at their hierarchical nature, which are different levels of granularity: documents, paragraphs, sentences, and tokens.We utilize graph attention networks to obtain different levels of representations so that they can be learned simultaneously.The long and short answers can be extracted from paragraphlevel representation and token-level representation, respectively.In this way, we can model the dependencies between the two-grained answers to provide evidence for each other.We jointly train the two sub-tasks, and our experiments show that our approach significantly outperforms previous systems at both long and short answer criteria. Bo Zheng 0010, Haoyang Wen, Yaobo Liang, Nan Duan 0001, Wanxiang Che, Daxin Jiang, Ming Zhou 0001, Ting Liu 0001 |
ACL | 5 |
| 2020 | A Sentence Cloze Dataset for Chinese Machine Reading ComprehensionabstractOwing to the continuous efforts by the Chinese NLP community, more and more Chinese machine reading comprehension datasets become available.To add diversity in this area, in this paper, we propose a new task called Sentence Cloze-style Machine Reading Comprehension (SC-MRC).The proposed task aims to fill the right candidate sentence into the passage that has several blanks.We built a Chinese dataset called CMRC 2019 to evaluate the difficulty of the SC-MRC task.Moreover, to add more difficulties, we also made fake candidates that are similar to the correct ones, which requires the machine to judge their correctness in the context.The proposed dataset contains over 100K blanks (questions) within over 10K passages, which was originated from Chinese narrative stories.To evaluate the dataset, we implement several baseline systems based on the pre-trained models, and the results show that the stateof-the-art model still underperforms human performance by a large margin.We release the dataset and baseline system to further facilitate our community. Yiming Cui 0001, Ting Liu 0001, Ziqing Yang 0001, Zhipeng Chen 0001, Wanxiang Che, Shijin Wang 0001 |
COLING | 6 |
| 2020 | Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less ForgettingabstractDeep pretrained language models have achieved great success in the way of pretraining first and then fine-tuning.But such a sequential transfer learning paradigm often confronts the catastrophic forgetting problem and leads to sub-optimal performance.To fine-tune with less forgetting, we propose a recall and learn mechanism, which adopts the idea of multi-task learning and jointly learns pretraining tasks and downstream tasks.Specifically, we propose a Pretraining Simulation mechanism to recall the knowledge from pretraining tasks without data, and an Objective Shifting mechanism to focus the learning on downstream tasks gradually.Experiments show that our method achieves state-of-the-art performance on the GLUE benchmark.Our method also enables BERT-base to achieve better performance than directly fine-tuning of BERT-large.Further, we provide the open-source RECADAM optimizer, which integrates the proposed mechanisms into Adam optimizer, to facility the NLP community. Sanyuan Chen, Yutai Hou, Yiming Cui 0001, Wanxiang Che, Ting Liu 0001, Xiangzhan Yu |
EMNLP (1) | 4 |
| 2020 | Combining Self-Training and Self-Supervised Learning for Unsupervised Disfluency DetectionabstractMost existing approaches to disfluency detection heavily rely on human-annotated corpora, which is expensive to obtain in practice.There have been several proposals to alleviate this issue with, for instance, self-supervised learning techniques, but they still require humanannotated corpora.In this work, we explore the unsupervised learning paradigm which can potentially work with unlabeled text corpora that are cheaper and easier to obtain.Our model builds upon the recent work on Noisy Student Training, a semi-supervised learning approach that extends the idea of self-training.Experimental results on the commonly used English Switchboard test set show that our approach achieves competitive performance compared to the previous state-of-the-art supervised systems using contextualized word embeddings (e.g.BERT and ELECTRA). Shaolei Wang, Wanxiang Che, Ting Liu 0001 |
EMNLP (1) | 3 |
| 2020 | CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual NLPabstractMulti-lingual contextualized embeddings, such as multilingual-BERT (mBERT), have shown success in a variety of zero-shot cross-lingual tasks. However, these models are limited by having inconsistent contextualized representations of subwords across different languages. Existing work addresses this issue by bilingual projection and fine-tuning technique. We propose a data augmentation framework to generate multi-lingual code-switching data to fine-tune mBERT, which encourages model to align representations from source and multiple target languages once by mixing their context information. Compared with the existing work, our method does not rely on bilingual sentences for training, and requires only one training process for multiple target languages. Experimental results on five tasks with 19 languages show that our method leads to significantly improved performances for all the tasks compared with mBERT. Libo Qin 0001, Minheng Ni, Yue Zhang 0004, Wanxiang Che |
IJCAI | 4 |
| 2020 | Enhancing Dialog Coherence with Event Graph Grounded Content PlanningabstractHow to generate informative, coherent and sustainable open-domain conversations is a non-trivial task. Previous work on knowledge grounded conversation generation focus on improving dialog informativeness with little attention on dialog coherence. In this paper, to enhance multi-turn dialog coherence, we propose to leverage event chains to help determine a sketch of a multi-turn dialog. We first extract event chains from narrative texts and connect them as a graph. We then present a novel event graph grounded Reinforcement Learning (RL) framework. It conducts high-level response content (simply an event) planning by learning to walk over the graph, and then produces a response conditioned on the planned content. In particular, we devise a novel multi-policy decision making mechanism to foster a coherent dialog with both appropriate content ordering and high contextual relevance. Experimental results indicate the effectiveness of this framework in terms of dialog coherence and informativeness. Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
IJCAI | 6 |
| 2020 | Keywords Generation Improves E-Commerce Session-based RecommendationabstractBy exploring fine-grained user behaviors, session-based recommendation predicts a user’s next action from short-term behavior sessions. Most of previous work learns about a user’s implicit behavior by merely taking the last click action as the supervision signal. However, in e-commerce scenarios, large-scale products with elusive click behaviors make such task challenging because of the low inclusiveness problem, i.e., many relevant products that satisfy the user’s shopping intention are neglected by recommenders. Since similar products with different IDs may share the same intention, we argue that the textual information (e.g., keywords of product titles) from sessions can be used as additional supervision signals to tackle above problem through learning more shared intention within similar products. Therefore, to improve the performance of e-commerce session-based recommendation, we explicitly infer the user’s intention by generating keywords entirely from the click sequence in the current session. Yuanxing Liu 0001, Zhaochun Ren, Weinan Zhang 0003, Wanxiang Che, Ting Liu 0001, Dawei Yin 0001 |
WWW | 4 |
| 2020 | Deep Contextualized Word Embeddings for Universal Dependency ParsingabstractDeep contextualized word embeddings (Embeddings from Language Model, short for ELMo), as an emerging and effective replacement for the static word embeddings, have achieved success on a bunch of syntactic and semantic NLP problems. However, little is known about what is responsible for the improvements. In this article, we focus on the effect of ELMo for a typical syntax problem—universal POS tagging and dependency parsing. We incorporate ELMo as additional word embeddings into the state-of-the-art POS tagger and dependency parser, and it leads to consistent performance improvements. Experimental results show the model using ELMo outperforms the state-of-the-art baseline by an average of 0.91 for POS tagging and 1.11 for dependency parsing. Further analysis reveals that the improvements mainly result from the ELMo’s better abstraction ability on the out-of-vocabulary (OOV) words, and the character-level word representation in ELMo contributes a lot to the abstraction. Based on ELMo’s advantage on OOV, experiments that simulate low-resource settings are conducted and the results show that deep contextualized word embeddings are effective for data-insufficient tasks where the OOV problem is severe. Wanxiang Che, Yuxuan Wang 0001, Bo Zheng 0010, Bing Qin 0001, Ting Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2020 | Exploring Segment Representations for Neural Semi-Markov Conditional Random FieldsabstractMany problems in natural language processing (NLP) can be cast as the problem of segmenting a sequence. In this article, we combine the semi-Markov conditional random fields (semi-CRF) with neural networks to solve NLP segmentation problems. We focus on the segment representation in neural semi-CRF which is important to the performance. Based on our preliminary work in Liu et al.[1], we represent a segment by both encoding the subsequence and embedding the segment string. We conduct a systematic study of the utility of various components in subsequence encoding and propose a method of constructing and deriving segment string embeddings. Extensive experiments on three typical segmentation problems, namely, shallow syntax parsing, named entity recognition, and Chinese word segmentation are conducted. The results show that we can achieve equally-performed subsequence encoding with a three times faster concatenation network compared to previous work. The results also show that the segment string embeddings help our neural semi-CRF model to achieve a macro-averaged error reduction of 13.15% over a strong baseline using deep contextualized embeddings and bidirectional long-short-term memory CRF, which also show the usefulness of semi-CRF even with contextualized embeddings. These results are competitive with the state-of-the-art segmentation systems. Wanxiang Che, Bing Qin 0001, Ting Liu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Generating Natural Language Adversarial Examples through Probability Weighted Word SaliencyabstractWe address the problem of adversarial attacks on text classification, which is rarely studied comparing to attacks on image classification.The challenge of this task is to generate adversarial examples that maintain lexical correctness, grammatical correctness and semantic similarity.Based on the synonyms substitution strategy, we introduce a new word replacement order determined by both the word saliency and the classification probability, and propose a greedy algorithm called probability weighted word saliency (PWWS) for text adversarial attack.Experiments on three popular datasets using convolutional as well as LSTM models show that PWWS reduces the classification accuracy to the most extent, and keeps a very low word substitution rate.A human evaluation study shows that our generated adversarial examples maintain the semantic similarity well and are hard for humans to perceive.Performing adversarial training using our perturbed datasets improves the robustness of the models.At last, our method also exhibits a good transferability on the generated adversarial examples. Shuhuai Ren, Yihe Deng, Kun He 0001, Wanxiang Che |
ACL (1) | 4 |
| 2019 | Cross-Lingual Machine Reading ComprehensionabstractYiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Shijin Wang, Guoping Hu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yiming Cui 0001, Wanxiang Che, Ting Liu 0001, Bing Qin 0001, Shijin Wang 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | A Span-Extraction Dataset for Chinese Machine Reading ComprehensionabstractYiming Cui, Ting Liu, Wanxiang Che, Li Xiao, Zhipeng Chen, Wentao Ma, Shijin Wang, Guoping Hu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yiming Cui 0001, Ting Liu 0001, Wanxiang Che, Zhipeng Chen 0001, Shijin Wang 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Stack-Propagation Framework with Token-Level Intent Detection for Spoken Language UnderstandingabstractLibo Qin, Wanxiang Che, Yangming Li, Haoyang Wen, Ting Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Libo Qin 0001, Wanxiang Che, Yangming Li, Haoyang Wen, Ting Liu 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Entity-Consistent End-to-end Task-Oriented Dialogue System with KB RetrieverabstractLibo Qin, Yijia Liu, Wanxiang Che, Haoyang Wen, Yangming Li, Ting Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Libo Qin 0001, Wanxiang Che, Haoyang Wen, Yangming Li, Ting Liu 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Cross-Lingual BERT Transformation for Zero-Shot Dependency ParsingabstractYuxuan Wang, Wanxiang Che, Jiang Guo, Yijia Liu, Ting Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yuxuan Wang 0001, Wanxiang Che, Ting Liu 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | A Key-Phrase Aware End2end Neural Response Generation Model
Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
NLPCC (2) | 5 |
| 2018 | A Neural Transition-Based Approach for Semantic Dependency Graph ParsingabstractSemantic dependency graph has been recently proposed as an extension of tree-structured syntactic or semantic representation for natural language sentences. It particularly features the structural property of multi-head, which allows nodes to have multiple heads, resulting in a directed acyclic graph(DAG) parsing problem. Yet most statistical parsers focused exclusively on shallow bi-lexical tree structures, DAG parsing remains under-explored. In this paper, we propose a neural transition-based parser, using a variant of list-based arc-eager transition algorithm for dependency graph parsing. Particularly, two non-trivial improvements are proposed for representing the key components of the transition system, to better capture the semantics of segments and internal sub-graph structures. We test our parser on the SemEval-2016 Task 9 dataset (Chinese) and the SemEval-2015 Task 18 dataset (English). On both benchmark datasets, we obtain superior or comparable results to the best performing systems. Our parser can be further improved with a simple ensemble mechanism, resulting in the state-of-the-art performance. Yuxuan Wang 0001, Wanxiang Che, Ting Liu 0001 |
AAAI | 2 |
| 2018 | Distilling Knowledge for Search-based Structured PredictionabstractMany natural language processing tasks can be modeled into structured prediction and solved as a search problem.In this paper, we distill an ensemble of multiple models trained with different initialization into a single model.In addition to learning to match the ensemble's probability output on the reference states, we also use the ensemble to explore the search space and learn from the encountered states in the exploration.Experimental results on two typical search-based structured prediction tasks -transition-based dependency parsing and neural machine translation show that distillation can effectively improve the single model's performance and the final model achieves improvements of 1.32 in LAS and 2.65 in BLEU score on these two tasks respectively over strong baselines and it outperforms the greedy structured prediction models in previous literatures. Wanxiang Che, Huaipeng Zhao, Bing Qin 0001, Ting Liu 0001 |
ACL (1) | 2 |
| 2018 | Sequence-to-Sequence Data Augmentation for Dialogue Language UnderstandingabstractIn this paper, we study the problem of data augmentation for language understanding in task-oriented dialogue system. In contrast to previous work which augments an utterance without considering its relation with other utterances, we propose a sequence-to-sequence generation based data augmentation framework that leverages one utterance’s same semantic alternatives in the training data. A novel diversity rank is incorporated into the utterance representation to make the model produce diverse utterances and these diversely augmented utterances help to improve the language understanding module. Experimental results on the Airline Travel Information System dataset and a newly created semantic frame annotation on Stanford Multi-turn, Multi-domain Dialogue Dataset show that our framework achieves significant improvements of 6.38 and 10.04 F-scores respectively when only a training set of hundreds utterances is represented. Case studies also confirm that our method generates diverse utterances. Yutai Hou, Wanxiang Che, Ting Liu 0001 |
COLING | 3 |
| 2018 | Sequence-to-Sequence Learning for Task-oriented Dialogue with Dialogue State RepresentationabstractClassic pipeline models for task-oriented dialogue system require explicit modeling the dialogue states and hand-crafted action spaces to query a domain-specific knowledge base. Conversely, sequence-to-sequence models learn to map dialogue history to the response in current turn without explicit knowledge base querying. In this work, we propose a novel framework that leverages the advantages of classic pipeline and sequence-to-sequence models. Our framework models a dialogue state as a fixed-size distributed representation and use this representation to query a knowledge base via an attention mechanism. Experiment on Stanford Multi-turn Multi-domain Task-oriented Dialogue Dataset shows that our framework significantly outperforms other sequence-to-sequence based baseline models on both automatic and human evaluation. Haoyang Wen, Wanxiang Che, Libo Qin 0001, Ting Liu 0001 |
COLING | 3 |
| 2018 | An AMR Aligner Tuned by Transition-based ParserabstractIn this paper, we propose a new rich resource enhanced AMR aligner which produces multiple alignments and a new transition system for AMR parsing along with its oracle parser.Our aligner is further tuned by our oracle parser via picking the alignment that leads to the highestscored achievable AMR graph.Experimental results show that our aligner outperforms the rule-based aligner in previous work by achieving higher alignment F1 score and consistently improving two open-sourced AMR parsers.Based on our aligner and transition system, we develop a transition-based AMR parser that parses a sentence into its AMR graph directly.An ensemble of our parsers with only words and POS tags as input leads to 68.4 Smatch F1 score, which outperforms the parser of Wang and Xue (2017). Wanxiang Che, Bo Zheng 0010, Bing Qin 0001, Ting Liu 0001 |
EMNLP | 2 |
| 2018 | Joint Extraction of Entities and Relations Based on a Novel Graph SchemeabstractBoth entity and relation extraction can benefit from being performed jointly, allowing each task to correct the errors of the other. Most existing neural joint methods extract entities and relations separately and achieve joint learning through parameter sharing, leading to a drawback that information between output entities and relations cannot be fully exploited. In this paper, we convert the joint task into a directed graph by designing a novel graph scheme and propose a transition-based approach to generate the directed graph incrementally, which can achieve joint learning through joint decoding. Our method can model underlying dependencies not only between entities and relations, but also between relations. Experiments on NewYork Times (NYT) corpora show that our approach outperforms the state-of-the-art methods. Shaolei Wang, Yue Zhang 0004, Wanxiang Che, Ting Liu 0001 |
IJCAI | 3 |
| 2018 | Parsing Tweets into Universal DependenciesabstractYijia Liu, Yi Zhu, Wanxiang Che, Bing Qin, Nathan Schneider, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Wanxiang Che, Bing Qin 0001, Nathan Schneider 0001, Noah A. Smith |
NAACL-HLT | 3 |
| 2017 | Transition-Based Disfluency Detection using LSTMsabstractWe model the problem of disfluency detection using a transition-based framework, which incrementally constructs and labels the disfluency chunk of input sentences using a set of transition actions without syntax information.Compared with sequence labeling methods, it can capture non-local chunk-level features; compared with joint parsing and disfluency detection methods, it is free for noise in syntax.Experiments show that our model achieves state-of-theart F-score on both the commonly used English Switchboard test set and a set of in-house annotated Chinese data. Shaolei Wang, Wanxiang Che, Yue Zhang 0004, Meishan Zhang, Ting Liu 0001 |
EMNLP | 2 |
| 2016 | A Representation Learning Framework for Multi-Source Transfer ParsingabstractCross-lingual model transfer has been a promising approach for inducing dependency parsers for low-resource languages where annotated treebanks are not available. The major obstacles for the model transfer approach are two-fold: 1. Lexical features are not directly transferable across languages; 2. Target language-specific syntactic structures are difficult to be recovered. To address these two challenges, we present a novel representation learning framework for multi-source transfer parsing. Our framework allows multi-source transfer parsing using full lexical features straightforwardly. By evaluating on the Google universal dependency treebanks (v2.0), our best models yield an absolute improvement of 6.53% in averaged labeled attachment score, as compared with delexicalized multi-source transfer models. We also significantly outperform the state-of-the-art transfer system proposed most recently. Wanxiang Che, David Yarowsky, Haifeng Wang 0001, Ting Liu 0001 |
AAAI | 2 |
| 2016 | A Universal Framework for Inductive Transfer Parsing across Multi-typed TreebanksabstractVarious treebanks have been released for dependency parsing. Despite that treebanks may belong to different languages or have different annotation schemes, they contain common syntactic knowledge that is potential to benefit each other. This paper presents a universal framework for transfer parsing across multi-typed treebanks with deep multi-task learning. We consider two kinds of treebanks as source: the multilingual universal treebanks and the monolingual heterogeneous treebanks. Knowledge across the source and target treebanks are effectively transferred through multi-level parameter sharing. Experiments on several benchmark datasets in various languages demonstrate that our approach can make effective use of arbitrary source treebanks to improve target parsing models. Wanxiang Che, Haifeng Wang 0001, Ting Liu 0001 |
COLING | 2 |
| 2016 | A Unified Architecture for Semantic Role Labeling and Relation ClassificationabstractThis paper describes a unified neural architecture for identifying and classifying multi-typed semantic relations between words in a sentence. We investigate two typical and well-studied tasks: semantic role labeling (SRL) which identifies the relations between predicates and arguments, and relation classification (RC) which focuses on the relation between two entities or nominals. While mostly studied separately in prior work, we show that the two tasks can be effectively connected and modeled using a general architecture. Experiments on CoNLL-2009 benchmark datasets show that our SRL models significantly outperform state-of-the-art approaches. Our RC models also yield competitive performance with the best published records. Furthermore, we show that the two tasks can be trained jointly with multi-task learning, resulting in additive significant improvements for SRL. Wanxiang Che, Haifeng Wang 0001, Ting Liu 0001, Jun Xu 0027 |
COLING | 2 |
| 2016 | A Neural Attention Model for Disfluency DetectionabstractIn this paper, we study the problem of disfluency detection using the encoder-decoder framework. We treat disfluency detection as a sequence-to-sequence problem and propose a neural attention-based model which can efficiently model the long-range dependencies between words and make the resulting sentence more likely to be grammatically correct. Our model firstly encode the source sentence with a bidirectional Long Short-Term Memory (BI-LSTM) and then use the neural attention as a pointer to select an ordered sub sequence of the input as the output. Experiments show that our model achieves the state-of-the-art f-score of 86.7% on the commonly used English Switchboard test set. We also evaluate the performance of our model on the in-house annotated Chinese data and achieve a significantly higher f-score compared to the baseline of CRF-based approach. Shaolei Wang, Wanxiang Che, Ting Liu 0001 |
COLING | 2 |
| 2016 | Exploring Segment Representations for Neural Segmentation Models
Wanxiang Che, Bing Qin 0001, Ting Liu 0001 |
IJCAI | 2 |
| 2016 | HC-Search for Incremental Parsing
Wanxiang Che, Bing Qin 0001, Ting Liu 0001 |
IJCAI | 2 |
| 2016 | A Distributed Representation-Based Framework for Cross-Lingual Transfer ParsingabstractThis paper investigates the problem of cross-lingual transfer parsing, aiming at inducing dependency parsers for low-resource languages while using only training data from a resource-rich language (e.g., English). Existing model transfer approaches typically don't include lexical features, which are not transferable across languages. In this paper, we bridge the lexical feature gap by using distributed feature representations and their composition. We provide two algorithms for inducing cross-lingual distributed representations of words, which map vocabularies from two different languages into a common vector space. Consequently, both lexical features and non-lexical features can be used in our model for cross-lingual transfer. Furthermore, our framework is flexible enough to incorporate additional useful features such as cross-lingual word clusters. Our combined contributions achieve an average relative error reduction of 10.9% in labeled attachment score as compared with the delexicalized parser, trained on English universal treebank and transferred to three other languages. It also significantly outperforms state-of-the-art delexicalized models augmented with projected cluster features on identical data. Finally, we demonstrate that our models can be further boosted with minimal supervision (e.g., 100 annotated sentences) from target languages, which is of great significance for practical usage. Wanxiang Che, David Yarowsky, Haifeng Wang 0001, Ting Liu 0001 |
J. Artif. Intell. Res. | 2 |
| 2015 | Cross-lingual Dependency Parsing Based on Distributed RepresentationsabstractJiang Guo, Wanxiang Che, David Yarowsky, Haifeng Wang, Ting Liu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Wanxiang Che, David Yarowsky, Haifeng Wang 0001, Ting Liu 0001 |
ACL (1) | 2 |
| 2015 | Transition-Based Syntactic LinearizationabstractSyntactic linearization algorithms take a bag of input words and a set of optional constraints, and construct an output sentence and its syntactic derivation simultaneously.The search problem is NP-hard, and the current best results are achieved by bottom-up bestfirst search.One drawback of the method is low efficiency; and there is no theoretical guarantee that a full sentence can be found within bounded time.We propose an alternative algorithm that constructs output structures from left to right using beam-search.The algorithm is based on incremental parsing algorithms.We extend the transition system so that word ordering is performed in addition to syntactic parsing, resulting in a linearization system that runs in guaranteed quadratic time.In standard evaluations, our system runs an order of magnitude faster than a state-of-the-art baseline using best-first search, with improved accuracies. Yue Zhang 0004, Wanxiang Che, Bing Qin 0001 |
HLT-NAACL | 3 |
| 2015 | Sentence Compression for Aspect-Based Sentiment AnalysisabstractSentiment analysis, which addresses the computational treatment of opinion, sentiment, and subjectivity in text, has received considerable attention in recent years. In contrast to the traditional coarse-grained sentiment analysis tasks, such as document-level sentiment classification, we are interested in the fine-grained aspect-based sentiment analysis that aims to identify aspects that users comment on and these aspects' polarities. Aspect-based sentiment analysis relies heavily on syntactic features. However, the reviews that this task focuses on are natural and spontaneous, thus posing a challenge to syntactic parsers. In this paper, we address this problem by proposing a framework of adding a sentiment sentence compression (Sent_Comp) step before performing the aspect-based sentiment analysis. Different from the previous sentence compression model for common news sentences, Sent_Comp seeks to remove the sentiment-unnecessary information for sentiment analysis, thereby compressing a complicated sentiment sentence into one that is shorter and easier to parse. We apply a discriminative conditional random field model, with certain special features, to automatically compress sentiment sentences. Using the Chinese corpora of four product domains, Sent_Comp significantly improves the performance of the aspect-based sentiment analysis. The features proposed for Sent_Comp, especially the potential semantic features, are useful for sentiment sentence compression. Wanxiang Che, Zhong Su, Ting Liu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Learning Semantic Hierarchies: A Continuous Vector Space ApproachabstractSemantic hierarchy construction aims to build structures of concepts linked by hypernym-hyponym (“is-a”) relations. A major challenge for this task is the automatic discovery of such relations. This paper proposes a novel and effective method for the construction of semantic hierarchies based on continuous vector representation of words, named word embeddings, which can be used to measure the semantic relationship between words. We identify whether a candidate word pair has hypernym-hyponym relation by using the word-embedding-based semantic projections between words and their hypernyms. Our result, an F-score of 73.74%, outperforms the state-of-the-art methods on a manually labeled test dataset. Moreover, combining our method with a previous manually built hierarchy extension method can further improve F-score to 80.29%. Ruiji Fu, Bing Qin 0001, Wanxiang Che, Haifeng Wang 0001, Ting Liu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Learning Semantic Hierarchies via Word EmbeddingsabstractSemantic hierarchy construction aims to build structures of concepts linked by hypernym-hyponym ("is-a") relations.A major challenge for this task is the automatic discovery of such relations.This paper proposes a novel and effective method for the construction of semantic hierarchies based on word embeddings, which can be used to measure the semantic relationship between words.We identify whether a candidate word pair has hypernym-hyponym relation by using the word-embedding-based semantic projections between words and their hypernyms.Our result, an F-score of 73.74%, outperforms the state-of-theart methods on a manually labeled test dataset.Moreover, combining our method with a previous manually-built hierarchy extension method can further improve Fscore to 80.29%. Ruiji Fu, Bing Qin 0001, Wanxiang Che, Haifeng Wang 0001, Ting Liu 0001 |
ACL (1) | 4 |
| 2014 | Character-Level Chinese Dependency ParsingabstractRecent work on Chinese analysis has led to large-scale annotations of the internal structures of words, enabling characterlevel analysis of Chinese syntactic structures.In this paper, we investigate the problem of character-level Chinese dependency parsing, building dependency trees over characters.Character-level information can benefit downstream applications by offering flexible granularities for word segmentation while improving wordlevel dependency parsing accuracies.We present novel adaptations of two major shift-reduce dependency parsing algorithms to character-level parsing.Experimental results on the Chinese Treebank demonstrate improved performances over word-based parsing methods. Meishan Zhang, Yue Zhang 0004, Wanxiang Che, Ting Liu 0001 |
ACL (1) | 3 |
| 2014 | A Semantics Oriented Grammar for Chinese Treebanking
Meishan Zhang, Yue Zhang 0004, Wanxiang Che, Ting Liu 0001 |
CICLing (1) | 3 |
| 2014 | Learning Sense-specific Word Embeddings By Exploiting Bilingual Resources
Wanxiang Che, Haifeng Wang 0001, Ting Liu 0001 |
COLING | 2 |
| 2014 | Jointly or Separately: Which is Better for Parsing Heterogeneous Dependencies?
Meishan Zhang, Wanxiang Che, Yanqiu Shao, Ting Liu 0001 |
COLING | 2 |
| 2014 | Sentence Compression for Target-Polarity Word Collocation Extraction
Wanxiang Che, Bing Qin 0001, Zhong Su, Ting Liu 0001 |
COLING | 2 |
| 2014 | Type-Supervised Domain Adaptation for Joint Segmentation and POS-TaggingabstractWe report an empirical investigation on type-supervised domain adaptation for joint Chinese word segmentation and POS-tagging, making use of domainspecific tag dictionaries and only unlabeled target domain data to improve target-domain accuracies, given a set of annotated source domain sentences.Previous work on POS-tagging of other languages showed that type-supervision can be a competitive alternative to tokensupervision, while semi-supervised techniques such as label propagation are important to the effectiveness of typesupervision.We report similar findings using a novel approach for joint Chinese segmentation and POS-tagging, under a cross-domain setting.With the help of unlabeled sentences and a lexicon of 3,000 words, we obtain 33% error reduction in target-domain tagging.In addition, combined type-and token-supervision can lead to improved cost-effectiveness. Meishan Zhang, Yue Zhang 0004, Wanxiang Che, Ting Liu 0001 |
EACL | 3 |
| 2014 | Revisiting Embedding Features for Simple Semi-supervised LearningabstractRecent work has shown success in using continuous word embeddings learned from unlabeled data as features to improve supervised NLP systems, which is regarded as a simple semi-supervised learning mechanism.However, fundamental problems on effectively incorporating the word embedding features within the framework of linear models remain.In this study, we investigate and analyze three different approaches, including a new proposed distributional prototype approach, for utilizing the embedding features.The presented approaches can be integrated into most of the classical linear models in NLP.Experiments on the task of named entity recognition show that each of the proposed approaches can better utilize the word embedding features, among which the distributional prototype approach performs the best.Moreover, the combination of the approaches provides additive improvements, outperforming the dense and continuous embedding features by nearly 2 points of F1 score. Wanxiang Che, Haifeng Wang 0001, Ting Liu 0001 |
EMNLP | 2 |
| 2014 | Domain Adaptation for CRF-based Chinese Word Segmentation using Free AnnotationsabstractSupervised methods have been the dominant approach for Chinese word segmentation.The performance can drop significantly when the test domain is different from the training domain.In this paper, we study the problem of obtaining partial annotation from freely available data to help Chinese word segmentation on different domains.Different sources of free annotations are transformed into a unified form of partial annotation and a variant CRF model is used to leverage both fully and partially annotated data consistently.Experimental results show that the Chinese word segmentation model benefits from free partially annotated data.On the SIGHAN Bakeoff 2010 data, we achieve results that are competitive to the best reported in the literature. Yue Zhang 0004, Wanxiang Che, Ting Liu 0001 |
EMNLP | 3 |
| 2014 | ReliAble dependency arc recognition
Wanxiang Che, Ting Liu 0001 |
Expert Syst. Appl. | 1 |
| 2014 | Joint Optimization for Chinese POS Tagging and Dependency ParsingabstractDependency parsing has gained more and more interest in natural language processing in recent years due to its simplicity and general applicability for diverse languages. Previous work demonstrates that part-of-speech (POS) is an indispensable feature in dependency parsing since pure lexical features suffer from serious data sparseness problem. However, due to little morphological changes, Chinese POS tagging has proven to be much more challenging than morphology-richer languages such as English (94% vs. 97% on POS tagging accuracy). This leads to severe error propagation for Chinese dependency parsing. Our experiments show that parsing accuracy drops by about 6% when replacing manual POS tags of the input sentence with automatic ones generated by a state-of-the-art statistical POS tagger. To address this issue, this paper proposes a solution by jointly optimizing POS tagging and dependency parsing in a unique model. We propose for our joint models several dynamic programming based decoding algorithms which can incorporate rich POS tagging and syntactic features. Then we present an effective pruning strategy to reduce the search space of candidate POS tags, leading to significant improvement of parsing speed. Experimental results on two Chinese data sets, i.e. Penn Chinese Treebank 5.1 and Penn Chinese Treebank 7, demonstrate that our joint models significantly improve both the state-of-the-art tagging and parsing accuracies. Detailed analysis shows that the joint method can help resolve syntax-sensitive POS ambiguities$\{{\ssr{NN}},{\ssr{VV}}\}$. In return, the POS tags become more reliable and helpful for parsing since the syntactic features are used in POS tagging. This is the fundamental reason for the performance improvement. Zhenghua Li, Min Zhang 0005, Wanxiang Che, Ting Liu 0001, Wenliang Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Effective Bilingual Constraints for Semi-Supervised Learning of Named Entity RecognizersabstractMost semi-supervised methods in Natural Language Processing capitalize on unannotated resources in a single language; however, information can be gained from using parallel resources in more than one language, since translations of the same utterance in different languages can help to disambiguate each other. We demonstrate a method that makes effective use of vast amounts of bilingual text (a.k.a. bitext) to improve monolingual systems. We propose a factored probabilistic sequence model that encourages both crosslanguage and intra-document consistency. A simple Gibbs sampling algorithm is introduced for performing approximate inference. Experiments on English-Chinese Named Entity Recognition (NER) using the OntoNotes dataset demonstrate that our method is significantly more accurate than state-ofthe- art monolingual CRF models in a bilingual test setting. Our model also improves on previous work by Burkett et al. (2010), achieving a relative error reduction of 10.8% and 4.5% in Chinese and English, respectively. Furthermore, by annotating a moderate amount of unlabeled bi-text with our bilingual model, and using the tagged data for uptraining, we achieve a 9.2% error reduction in Chinese over the state-ofthe- art Stanford monolingual NER system. Mengqiu Wang, Wanxiang Che, Christopher D. Manning |
AAAI | 2 |
| 2013 | Joint Word Alignment and Bilingual Named Entity Recognition Using Dual Decomposition
Mengqiu Wang, Wanxiang Che, Christopher D. Manning |
ACL (1) | 2 |
| 2013 | Chinese Parsing Exploiting Characters
Meishan Zhang, Yue Zhang 0004, Wanxiang Che, Ting Liu 0001 |
ACL (1) | 3 |
| 2013 | Convolution Neural Network for Relation Extraction
Wen-Han Chao, Wanxiang Che |
ADMA (2) | 4 |
| 2013 | Named Entity Recognition with Bilingual Constraints
Wanxiang Che, Mengqiu Wang, Christopher D. Manning, Ting Liu 0001 |
HLT-NAACL | 1 |
| 2012 | Exploiting Multiple Treebanks for Parsing with Quasi-synchronous Grammars
Zhenghua Li, Ting Liu 0001, Wanxiang Che |
ACL (1) | 3 |
| 2012 | A Separately Passive-Aggressive Training Algorithm for Joint POS Tagging and Dependency Parsing
Zhenghua Li, Min Zhang 0005, Wanxiang Che, Ting Liu 0001 |
COLING | 3 |
| 2012 | Stacking Heterogeneous Joint Models of Chinese POS Tagging and Dependency Parsing
Meishan Zhang, Wanxiang Che, Ting Liu 0001, Zhenghua Li |
COLING | 2 |
| 2011 | Joint Models for Chinese POS Tagging and Dependency Parsing
Zhenghua Li, Min Zhang 0005, Wanxiang Che, Ting Liu 0001, Wenliang Chen, Haizhou Li 0001 |
EMNLP | 3 |
| 2011 | Word Sense Disambiguation Corpora Acquisition via Confirmation Code
Wanxiang Che, Ting Liu 0001 |
IJCNLP | 1 |
| 2011 | A Graph-based Method for Entity Linking
Yuhang Guo 0001, Wanxiang Che, Ting Liu 0001, Sheng Li 0003 |
IJCNLP | 2 |
| 2011 | Improving Chinese POS Tagging with Dependency Parsing
Zhenghua Li, Wanxiang Che, Ting Liu 0001 |
IJCNLP | 2 |
| 2010 | Jointly Modeling WSD and SRL with Markov Logic
Wanxiang Che, Ting Liu 0001 |
COLING | 1 |
| 2010 | Improving Semantic Role Labeling with Word Sense
Wanxiang Che, Ting Liu 0001 |
HLT-NAACL | 1 |
| 2008 | A Cascaded Syntactic and Semantic Dependency Parsing System
Wanxiang Che, Zhenghua Li, Yuxuan Hu 0001, Bing Qin 0001, Ting Liu 0001, Sheng Li 0003 |
CoNLL | 1 |
| 2008 | Fast Computing Grammar-driven Convolution Tree Kernel for Semantic Role Labeling
Wanxiang Che, Min Zhang 0005, AiTi Aw, Chew Lim Tan, Ting Liu 0001, Sheng Li 0003 |
IJCNLP | 1 |
| 2008 | Using a Hybrid Convolution Tree Kernel for Semantic Role LabelingabstractAs a kind of Shallow Semantic Parsing, Semantic Role Labeling (SRL) is gaining more attention as it benefits a wide range of natural language processing applications. Given a sentence, the task of SRL is to recognize semantic arguments (roles) for each predicate (target verb or noun). Feature-based methods have achieved much success in SRL and are regarded as the state-of-the-art methods for SRL. However, these methods are less effective in modeling structured features. As an extension of feature-based methods, kernel-based methods are able to capture structured features more efficiently in a much higher dimension. Application of kernel methods to SRL has been achieved by selecting the tree portion of a predicate and one of its arguments as feature space, which is named as predicate-argument feature (PAF) kernel. The PAF kernel captures the syntactic tree structure features using convolution tree kernel, however, it does not distinguish between the path structure and the constituent structure. In this article, a hybrid convolution tree kernel is proposed to model different linguistic objects. The hybrid convolution tree kernel consists of two individual convolution tree kernels. They are a Path kernel, which captures predicate-argument link features, and a Constituent Structure kernel, which captures the syntactic structure features of arguments. Evaluations on the data sets of the CoNLL-2005 SRL shared task and the Chinese PropBank (CPB) show that our proposed hybrid convolution tree kernel statistically significantly outperforms the previous tree kernels. Moreover, in order to maximize the system performance, we present a composite kernel through combining our hybrid convolution tree kernel method with a feature-based method extended by the polynomial kernel. The experimental results show that the composite kernel achieves better performance than each of the individual methods and outperforms the best reported system on the CoNLL-2005 corpus when only one syntactic parser is used and on the CPB corpus when automated syntactic parse results and correct syntactic parse results are used respectively. Wanxiang Che, Min Zhang 0005, AiTi Aw, Chew Lim Tan, Ting Liu 0001, Sheng Li 0003 |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2008 | Semantic Role Labeling Using a Grammar-Driven Convolution Tree KernelabstractConvolution tree kernel has shown promising results in semantic role labeling (SRL). However, this kernel does not consider much linguistic knowledge in kernel design and only performs hard matching between subtrees. To overcome these constraints, this paper proposes a grammar-driven convolution tree kernel for SRL by introducing more linguistic knowledge. Compared with the standard convolution tree kernel, the proposed grammar-driven kernel has two advantages: 1) grammar-driven approximate substructure matching, and 2) grammar-driven approximate tree node matching. The two approximate matching mechanisms enable the proposed kernel to better explore linguistically motivated structured knowledge. Experiments on the CoNLL-2005 SRL shared task and the PropBank I corpus show that the proposed kernel outperforms the standard convolution tree kernel significantly. Moreover, we present a composite kernel to integrate a feature-based polynomial kernel and the proposed grammar-driven convolution tree kernel for SRL. Experimental results show that our composite kernel-based method significantly outperforms the previously best-reported ones. Min Zhang 0005, Wanxiang Che, Guodong Zhou 0001, AiTi Aw, Chew Lim Tan, Ting Liu 0001, Sheng Li 0003 |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | A Grammar-driven Convolution Tree Kernel for Semantic Role Classification
Min Zhang 0005, Wanxiang Che, AiTi Aw, Chew Lim Tan, Guodong Zhou 0001, Ting Liu 0001, Sheng Li 0003 |
ACL | 2 |
| 2006 | A Hybrid Convolution Tree Kernel for Semantic Role Labeling
Wanxiang Che, Min Zhang 0005, Ting Liu 0001, Sheng Li 0003 |
ACL | 1 |
| 2005 | Semantic Role Labeling System Using Maximum Entropy Classifier
Ting Liu 0001, Wanxiang Che, Sheng Li 0003, Yuxuan Hu 0001, Huaijun Liu |
CoNLL | 2 |