Tao Gui

dblp:135/6973 · DBLP profile ↗
← Back
113ranked-venue papers
10as first author
95since 2021 · last 2026
0000-0002-0059-0210ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 104 · 9 first-author · 90 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
abstract
Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12× faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency.
Xiaoran Fan, Yangfan Gao, Jingfei Xiong, Hang Yan 0001, Yifei Cao, Zhihao Zhang 0002, Zhiheng Xi, Yuhao Zhou 0005, Senjie Jin, Changhao Jiang, Junjie Ye 0005, Ming Zhang 0030, Zhenhua Han, Yunke Zhang, Demei Yan, Shaokang Dong, Tao Gui
AAAI22
2026 MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention Across Vision-Language Models
abstract
As vision-language models (VLMs) tackle increasingly complex and multimodal tasks, the rapid growth of Key-Value (KV) cache imposes significant memory and computational bottlenecks during inference. While Multi-Head Latent Attention (MLA) offers an effective means to compress the KV cache and accelerate inference, adapting existing VLMs to the MLA architecture without costly pretraining remains largely unexplored. In this work, we present \textbf{MHA2MLA-VLM}, a parameter-efficient and multimodal-aware framework for converting off-the-shelf VLMs to MLA. Our approach features two core techniques: (1) a modality-adaptive partial-RoPE strategy that supports both traditional and multimodal settings by selectively masking nonessential dimensions, and (2) a modality-decoupled low-rank approximation method that independently compresses the visual and textual KV spaces. Furthermore, we introduce parameter-efficient fine-tuning to minimize adaptation cost and demonstrate that minimizing output activation error, rather than parameter distance, substantially reduces performance loss. Extensive experiments on three representative VLMs show that MHA2MLA-VLM restores original model performance with minimal supervised data, significantly reduces KV cache footprint, and integrates seamlessly with KV quantization.
Xiaoran Fan, Lixing Shen, Tao Gui
AAAI5
2026 MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement Learning
abstract
Outcome-based reinforcement learning has made notable advances in training language models (LMs) for reasoning. However, without explicit incentives and controls, this paradigm has limitations and instability in eliciting high-quality reasoning trajectories with diverse actions—particularly for models whose pretraining lacked extensive reasoning-related data. To this end, we introduce MetaAct-RL, a new RL framework that frames LMs’ thinking as sequential decision making over meta-actions. In this framework, the model chooses and executes a high-level action at each step—such as forward reasoning, critique, or refinement—to gradually reach the correct answer. To encourage deeper exploration, richer action diversity, and to improve sampling efficiency in the RL optimization process, MetaAct-RL incorporates appropriate length-based reward and regularization, and a key-state restart mechanism. Extensive experiments across six benchmarks show that MetaAct-RL improves reasoning performance by 7.99 on Llama3.2-1B and 7.17 on Llama3.1-8B relative to vanilla RL method. Moreover, on the challenging AIME-2024, our method outperforms the vanilla RL by 7.5 with Qwen2.5-1.5B.
Zhiheng Xi, Yiwen Ding, Senjie Jin, Shichun Liu, Jixuan Huang, Dingwen Yang, Jiafu Tang, Boyang Hong, Junjie Ye 0005, Shihan Dou, Ming Zhang 0030, Jian Guan 0002, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
AAAI17
2026 OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
abstract
Deming Ding, Shichun Liu, Enhui Yang, Jiahang Lin, Ziying Chen, Shihan Dou, Honglin Guo, Weiyu Cheng, Pengyu Zhao, Chengjun Xiao, Qunhong Zeng, Qi Zhang, Xuanjing Huang, Qidi Xu, Tao Gui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Deming Ding, Shichun Liu, Enhui Yang, Jiahang Lin, Ziying Chen, Shihan Dou, Honglin Guo, Weiyu Cheng, Chengjun Xiao, Qunhong Zeng, Qi Zhang 0001, Xuanjing Huang 0001, Qidi Xu, Tao Gui
ACL (1)15
2026 Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
abstract
Xin Guo, Zhiheng Xi, Yiwen Ding, Yitao Zhai, Xiaowei Shi, Xunliang Cai, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhiheng Xi, Yiwen Ding, Yitao Zhai, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)7
2026 Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
abstract
Changhao Jiang, Ming Zhang, Yifei Cao, Junjie Ye, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Jiajun Sun, Yi Dong, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Changhao Jiang, Ming Zhang 0030, Yifei Cao, Junjie Ye 0005, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)13
2026 DARM: Distribution-Aware Reward Modeling by Alleviating Biases from Low Preference-Context Dependency Data
abstract
Shaofan Liu, Guoqiang Zhang, Shihan Dou, Huiyuan Zheng, Yiming Zhou, Junjie Ye, Shaowen Wang, Shichun Liu, Jiazheng Zhang, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shaofan Liu, Shihan Dou, Huiyuan Zheng, Junjie Ye 0005, Shichun Liu, Jiazheng Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)10
2026 Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
abstract
Binghai Wang, Yantao Liu, Yuxuan Liu, Tianyi Tang, Shenzhi Wang, Chang Gao, Chujie Zheng, Yichang Zhang, Le Yu, Shixuan Liu, Tao Gui, Qi Zhang, Xuanjing Huang, Bowen Yu, Fei Huang, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Binghai Wang, Yantao Liu, Shenzhi Wang, Chujie Zheng, Yichang Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Bowen Yu 0002, Fei Huang 0002, Junyang Lin
ACL (1)11
2026 Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
abstract
Search agents extend Large Language Models (LLMs) beyond static parametric knowledge by enabling access to up-to-date and longtail information unavailable during pretraining.While reinforcement learning has been widely adopted for training such agents, existing approaches face key limitations: process supervision often suffers from unstable value estimation, whereas outcome supervision struggles with credit assignment due to sparse, trajectory-level rewards.To bridge this gap, we propose Contribution-Weighted GRPO (CW-GRPO), a framework that integrates process supervision into group relative policy optimization.Instead of directly optimizing process rewards, CW-GRPO employs an LLM judge to assess the retrieval utility and reasoning correctness at each search round, producing per-round contribution weights.These weights are used to rescale outcome-based advantages along the trajectory, enabling finegrained credit assignment without sacrificing optimization stability.Experiments on multiple knowledge-intensive benchmarks show that CW-GRPO outperforms standard GRPO by 5.0% on Qwen3-8B and 6.3% on Qwen3-1.7B,leading to more effective search behaviors.Additional analysis reveals that successful trajectories exhibit concentrated contributions in specific rounds, providing empirical insight into search agent tasks.Our code is available at https://github.com/zsxmwjz/CW-GRPO.
Junzhe Wang 0001, Zhiheng Xi, Yajie Yang, Shihan Dou, Tao Gui, Qi Zhang 0001
ACL (1)6
2026 AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
abstract
Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhiheng Xi, Dingwen Yang, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang 0001, Zhonghang Lu, Jiazheng Zhang, Dingwei Zhu, Junzhe Wang 0001, Zhihao Zhang 0002, Yuming Yang 0001, Junjie Ye 0005, Minghe Gao, Dongrui Liu, Jiaming Ji, Tao Gui, Xuanjing Huang 0001
ACL (1)23
2026 Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment
abstract
Yuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuming Yang 0001, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao 0019, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)14
2026 LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
abstract
Ming Zhang, Yujiong Shen, Jingyi Deng, Yuhui Wang, Huayu Sha, Kexin Tan, Qiyuan Peng, Yue Zhang, Junzhe Wang, Shichun Liu, Yueyuan Huang, Jingqi Tong, Changhao Jiang, Yilong Wu, Zhihao Zhang, Mingqi Wu, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ming Zhang 0030, Yujiong Shen, Jingyi Deng, Huayu Sha, Kexin Tan, Qiyuan Peng, Yue Zhang 0004, Junzhe Wang 0001, Shichun Liu, Yueyuan Huang, Jingqi Tong, Changhao Jiang, Yilong Wu, Zhihao Zhang 0002, Mingqi Wu, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)20
2026 VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training
abstract
Dingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Guoqiang Zhang, Jiazheng Zhang, Junjie Ye, Mingxu Chai, Enyu Zhou, Ming Zhang, Yuhui Wang, Caishuang Huang, Chenhao Huang, Yunke Zhang, Yuran Wang, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Dingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Jiazheng Zhang, Junjie Ye 0005, Mingxu Chai, Enyu Zhou, Ming Zhang 0030, Caishuang Huang, Chenhao Huang, Yunke Zhang, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001
ACL (1)16
2026 AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
abstract
Despite rapid development, large language models (LLMs) still encounter challenges in multi-turn decision-making tasks (i.e., agent tasks) like web shopping and browser navigation, which require making a sequence of intelligent decisions based on environmental feedback. Previous work for LLM agents typically relies on elaborate prompt engineering or fine-tuning with expert trajectories to improve performance. In this work, we take a different perspective: we explore constructing process reward models (PRMs) to evaluate each decision and guide the agent's decision-making process. Unlike LLM reasoning, where each step is scored based on correctness, actions in agent tasks do not have a clear-cut correctness. Instead, they should be evaluated based on their proximity to the goal and the progress they have made. Building on this insight, we propose a re-defined PRM for agent tasks, named AgentPRM, to capture both the interdependence between sequential decisions and their contribution to the final goal. This enables better progress tracking and exploration-exploitation balance. To scalably obtain labeled data for training AgentPRM, we employ a Temporal Difference-based (TD-based) estimation method combined with Generalized Advantage Estimation (GAE), which proves more sample-efficient than prior methods. Extensive experiments across different agentic tasks show that AgentPRM is over 8× more compute-efficient than baselines, and it demonstrates robust improvement when scaling up test-time compute. Moreover, we perform detailed analyses to show how our method works and offer more insights, e.g., applying AgentPRM to the reinforcement learning of LLM agents.
Zhiheng Xi, Chenyang Liao, Zhihao Zhang 0002, Wenxiang Chen, Binghai Wang, Senjie Jin, Yuhao Zhou 0005, Jian Guan 0002, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
WWW12
2026 What is wrong with your code generated by large language models? An extensive study
Shihan Dou, Haoxiang Jia, Shenxi Wu, Huiyuan Zheng, Muling Wu, Yunbo Tao, Ming Zhang 0030, Mingxu Chai, Jessica Fan, Zhiheng Xi, Yueming Wu 0001, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001
Sci. China Inf. Sci.14
2025 Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning
abstract
The capability of the reward model (RM) is crucial for the success of Reinforcement Learning from Human Feedback (RLHF) in aligning with human preferences. However, as training progresses, the output space distribution of the policy model shifts. The RM, initially trained on responses sampled from the output distribution of the early policy model, gradually loses its ability to distinguish between responses from the newly shifted distribution. This issue is further compounded when the RM, trained on a specific data distribution, struggles to generalize to examples outside of that distribution. These two issues can be united as a challenge posed by the shifted distribution of the environment. To surmount this challenge, we introduce MetaRM, a novel method leveraging meta-learning to adapt the RM to the shifted environment distribution. MetaRM optimizes the RM in an alternating way, by preserving both the preferences of the original preference pairs, as well as maximizing discrimination power over new examples of the shifted distribution. Extensive experiments demonstrate that MetaRM can iteratively enhance the performance of human preference alignment by improving the RM's capacity to identify subtle differences in samples of shifted distributions.
Shihan Dou, Yan Liu 0002, Enyu Zhou, Songyang Gao, Tianlong Li, Limao Xiong, Haoxiang Jia, Junjie Ye 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
AAAI11
2025 Lost in the Context: Insufficient and Distracted Attention to Contexts in Preference Modeling
abstract
Shihan Dou, Jiayi Chen, Chenhao Huang, Feng Chen, Wei Chengzhi, Huiyuan Zheng, Shichun Liu, Yan Liu, Chenxiao Liu, Chao Xin, Lin Yan, Zongzhang Zhang, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shihan Dou, Chenhao Huang, Feng Chen 0042, Wei Chengzhi, Huiyuan Zheng, Shichun Liu, Yan Liu 0002, Chenxiao Liu, Chao Xin, Zongzhang Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)13
2025 CritiQ: Mining Data Quality Criteria from Human Preferences
abstract
Honglin Guo, Kai Lv, Qipeng Guo, Tianyi Liang, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun, Kai Chen, Xipeng Qiu, Tao Gui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Honglin Guo, Kai Lv 0001, Qipeng Guo, Tianyi Liang 0002, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun 0031, Kai Chen 0026, Xipeng Qiu, Tao Gui
ACL (1)11
2025 Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
abstract
Tao Ji, Bin Guo, Yuanbin Wu, Qipeng Guo, Shenlixing Shenlixing, Chenzhan Chenzhan, Xipeng Qiu, Qi Zhang, Tao Gui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yuanbin Wu, Qipeng Guo, Shenlixing Shenlixing, Chenzhan Chenzhan, Xipeng Qiu, Qi Zhang 0001, Tao Gui
ACL (1)9
2025 AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments
abstract
Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Xin Guo, Dingwen Yang, Chenyang Liao, Wei He, Songyang Gao, Lu Chen, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, Yu-Gang Jiang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang 0001, Dingwen Yang, Chenyang Liao, Wei He 0024, Songyang Gao, Lu Chen 0001, Yicheng Zou, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Zuxuan Wu, Yu-Gang Jiang 0001
ACL (1)15
2025 Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric
abstract
Data diversity is crucial for the instruction tuning of large language models. Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. However, the fundamental problem of precisely defining and measuring data diversity remains underexplored, limiting clear guidance for data engineering. To address this, we systematically analyze 11 existing diversity measurement methods by evaluating their correlation with model performance through extensive fine-tuning experiments. Our results indicate that a reliable diversity measure should properly account for both inter-sample differences and the information density in the sample space. Building on this, we propose NovelSum, a new diversity metric based on sample-level “novelty.” Experiments on both simulated and real-world data show that NovelSum accurately captures diversity variations and achieves a 0.97 correlation with instruction-tuned model performance, highlighting its value in guiding data engineering practices. With NovelSum as an optimization objective, we further develop a greedy, diversity-oriented data selection strategy that outperforms existing approaches, validating both the effectiveness and practical significance of our metric.
Yuming Yang 0001, Junjie Ye 0005, Shihan Dou, Xiao Wang 0042, Huijie Lv, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)8
2025 ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
abstract
Junjie Ye, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zehui Chen, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Siyu Yuan, Tao Gui, Qi Zhang, Xuanjing Huang, Jiecao Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Junjie Ye 0005, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Jiecao Chen
ACL (1)11
2025 Joining the Dots: Efficient Joins Over Parent-Child Ordered Postings Lists
Karthikeyan Ramasamy, Yupeng Fu, Jinny Jingyu Wang, Tejas Naik, George Zhai, Saurabh Kathpalia, Noah Schlager, Kamyar Arbabifard, Sandhya Sainath, Kunal Veera, Vivek Muniyandi, Nimish Sheth, Yiyu Pan, Tao Gui, Kaustubh Butte, Monika Agarwal, Aparajita Pandey
IEEE Big Data16
2025 Beyond Boundaries: Learning a Universal Entity Taxonomy across Datasets and Languages for Open Named Entity Recognition
abstract
Open Named Entity Recognition (NER), which involves identifying arbitrary types of entities from arbitrary domains, remains challenging for Large Language Models (LLMs). Recent studies suggest that fine-tuning LLMs on extensive NER data can boost their performance. However, training directly on existing datasets neglects their inconsistent entity definitions and redundant data, limiting LLMs to dataset-specific learning and hindering out-of-domain adaptation. To address this, we present B2NERD, a compact dataset designed to guide LLMs’ generalization in Open NER under a universal entity taxonomy. B2NERD is refined from 54 existing English and Chinese datasets using a two-step process. First, we detect inconsistent entity definitions across datasets and clarify them by distinguishable label names to construct a universal taxonomy of 400+ entity types. Second, we address redundancy using a data pruning strategy that selects fewer samples with greater category and semantic diversity. Comprehensive evaluation shows that B2NERD significantly enhances LLMs’ Open NER capabilities. Our B2NER models, trained on B2NERD, outperform GPT-4 by 6.8-12.0 F1 points and surpass previous methods in 3 out-of-domain benchmarks across 15 datasets and 6 languages. The data, models, and code are publicly available at https://github.com/UmeanNever/B2NER.
Yuming Yang 0001, Wantong Zhao, Caishuang Huang, Junjie Ye 0005, Xiao Wang 0042, Huiyuan Zheng, Xueying Xu, Kaixin Huang, Yunke Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
COLING12
2025 ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios
abstract
Existing evaluations of tool learning primarily focus on validating the alignment of selected tools for large language models (LLMs) with expected outcomes. However, these approaches rely on a limited set of scenarios where answers can be pre-determined. Furthermore, a sole emphasis on outcomes disregards the complex capabilities required for LLMs to effectively use tools. To tackle this issue, we propose ToolEyes, a fine-grained system tailored for the evaluation of the LLMs’ tool learning capabilities in authentic scenarios. The system meticulously examines seven real-world scenarios, analyzing five dimensions crucial to LLMs in tool learning: format alignment, intent comprehension, behavior planning, tool selection, and answer organization. Additionally, ToolEyes incorporates a tool library boasting approximately 600 tools, serving as an intermediary between LLMs and the physical world. Evaluations involving ten LLMs across three categories reveal a preference for specific scenarios and limited cognitive abilities in tool learning. Intriguingly, expanding the model size even exacerbates the hindrance to tool learning. The code and data are available at https://github.com/Junjie-Ye/ToolEyes.
Junjie Ye 0005, Songyang Gao, Caishuang Huang, Yilong Wu, Sixian Li, Xiaoran Fan, Shihan Dou, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
COLING11
2025 SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Models
abstract
The emergence of Vision Language Models (VLMs) has brought unprecedented advances in understanding multi-modal information. The combination of textual and visual semantics in VLMs is highly complex and diverse, making the safety alignment of these models challenging. Furthermore, due to the limited study on the safety alignment of VLMs, there is a lack of large-scale, high-quality datasets. To address these limitations, we propose a Safety Preference Alignment dataset for Vision Language Models named SPA-VL. In terms of breadth, SPA-VL covers 6 harmfulness domains, 13 categories, and 53 subcategories, and contains 100,788 samples of the quadruple (question, image, chosen response, rejected response). In terms of depth, the responses are collected from 12 open-source (e.g., QwenVL) and closed-source (e.g., Gemini) VLMs to ensure diversity. The construction of preference data is fully automated, and the experimental results indicate that models trained with alignment techniques on the SPA-VL dataset exhibit substantial improvements in harmlessness and helpfulness while maintaining core capabilities. SPA-VL, as a large-scale, high-quality, and diverse dataset, represents a significant milestone in ensuring that VLMs achieve both harmlessness and helpfulness.
Yongting Zhang, Lu Chen 0001, Guodong Zheng, Yifeng Gao 0002, Jinlan Fu, Zhenfei Yin, Senjie Jin, Yu Qiao 0001, Xuanjing Huang 0001, Feng Zhao 0004, Tao Gui
CVPR12
2025 Governance in Motion: Co-evolution of Constitutions and AI models for Scalable Safety
abstract
Chenhao Huang, Ziyu Shen, Yicong Ren, Huiyuan Zheng, Jiazheng Zhang, Mingxu Chai, Ming Zhang, Shihan Dou, Fan Mo, Jie Shi, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Chenhao Huang, Ziyu Shen, Yicong Ren, Huiyuan Zheng, Jiazheng Zhang, Mingxu Chai, Ming Zhang 0030, Shihan Dou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP11
2025 Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
abstract
Senjie Jin, Lu Chen, Zhiheng Xi, Yuhui Wang, Sirui Song, Yuhao Zhou, Xinbo Zhang, Peng Sun, Hong Lu, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Senjie Jin, Lu Chen 0001, Zhiheng Xi, Sirui Song, Yuhao Zhou 0005, Xinbo Zhang, Peng Sun 0006, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP10
2025 LoRACoE: Improving Large Language Model via Composition-based LoRA Expert
abstract
The Mixture of Experts (MoE) architecture improves large language models (LLMs) by utilizing sparsely activated expert sub-networks with a routing module, yet it typically demands high training cost.Previous work introduces parameter-efficient fine-tuning (PEFT) modules, e.g., LoRA, to achieve a lightweight MoE for efficiency.However, they construct static experts by manually splitting the LoRA parameters into fixed groups, which limits flexibility and dynamism.Furthermore, this manual partitioning also hinders the effective utilization of well-initialized LoRA modules.To tackl the challenges, we first delve into the parameter patterns in LoRA modules, revealing that there exists task-relevant parameters that are concentrated along the rank dimension.Based on this, we redesign the construction of experts and propose the LoRACoE (LoRA Composition of Experts) method.Specifically, when confronted with a task, it dynamically builds experts based on rank-level parameter composition, i.e., experts can flexibly combine rank-level parameters in LoRA module.Extensive experiments demonstrate that compared to other LoRA-based MoE methods, our method achieves better task performance across a broader range of tasks.
Zhiheng Xi, Zhihao Zhang 0002, Boyang Hong, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP5
2025 Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
abstract
Junjie Ye, Yuming Yang, Yang Nan, Shuo Li, Qi Zhang, Tao Gui, Xuanjing Huang, Peng Wang, Zhongchao Shi, Jianping Fan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Junjie Ye 0005, Yuming Yang 0001, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001, Peng Wang 0095, Zhongchao Shi, Jianping Fan 0007
EMNLP6
2025 Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
abstract
In the study of LLMs, sycophancy represents a prevalent hallucination that poses significant challenges to these models. Specifically, LLMs often fail to adhere to original correct responses, instead blindly agreeing with users' opinions, even when those opinions are incorrect or malicious. However, research on sycophancy in visual language models (VLMs) has been scarce. In this work, we extend the exploration of sycophancy from LLMs to VLMs, introducing the MM-SY benchmark to evaluate this phenomenon. We present evaluation results from multiple representative models, addressing the gap in sycophancy research for VLMs. To mitigate sycophancy, we propose a synthetic dataset for training and employ methods based on prompts, supervised fine-tuning, and DPO. Our experiments demonstrate that these methods effectively alleviate sycophancy in VLMs. Additionally, we probe VLMs to assess the semantic impact of sycophancy and analyze the attention distribution of visual tokens. Our findings indicate that the ability to prevent sycophancy is predominantly observed in higher layers of the model. The lack of attention to image knowledge in these higher layers may contribute to sycophancy, and enhancing image attention at high layers proves beneficial in mitigating this issue.
Xiaoran Fan, Linsheng Lu, Leyi Yang, Yuming Yang 0001, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ICLR11
2025 RMB: Comprehensively benchmarking reward models in LLM alignment
abstract
Reward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. However, the current evaluation of RMs may not directly correspond to their alignment performance due to the limited distribution of evaluation data and evaluation methods that are not closely related to alignment objectives. To address these limitations, we propose RMB, a comprehensive RM benchmark that covers over 49 real-world scenarios and includes both pairwise and Best-of-N (BoN) evaluations to better reflect the effectiveness of RMs in guiding alignment optimization. We demonstrate a positive correlation between our benchmark and the downstream alignment task performance. Based on our benchmark, we conduct extensive analysis on the state-of-the-art RMs, revealing their generalization defects that were not discovered by previous benchmarks, and highlighting the potential of generative RMs. Furthermore, we delve into open questions in reward models, specifically examining the effectiveness of majority voting for the evaluation of reward models and analyzing the impact factors of generative RMs, including the influence of evaluation criteria and instructing methods. We will release our evaluation code and datasets upon publication.
Enyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi, Shihan Dou, Rong Bao, Limao Xiong, Jessica Fan, Yurong Mou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ICLR12
2025 Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
abstract
Yiwen Ding, Zhiheng Xi, Wei He, Lizhuoyuan Lizhuoyuan, Yitao Zhai, Shi Xiaowei, Xunliang Cai, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yiwen Ding, Zhiheng Xi, Wei He 0024, Lizhuoyuan Lizhuoyuan, Yitao Zhai, Shi Xiaowei, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
NAACL (Long Papers)8
2025 Pre-Trained Policy Discriminators are General Reward Models
abstract
We offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named POLicy DiscriminAtive LeaRning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance. For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines. POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance—improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks. Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99. The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models.
Shihan Dou, Shichun Liu, Yuming Yang 0001, Yicheng Zou, Yunhua Zhou, Shuhao Xing, Chenhao Huang, Qiming Ge, Haijun Lv, Demin Song, Songyang Gao, Chengqi Lyu, Enyu Zhou, Honglin Guo, Zhiheng Xi, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Kai Chen 0026
NeurIPS18
2025 EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
abstract
We introduce EvaLearn, a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks, a critical, yet underexplored aspect of model potential. EvaLearn contains 648 challenging problems across six task types, grouped into 182 sequences, each sequence dedicated to one task type. Diverging from most existing benchmarks that evaluate models in parallel, EvaLearn requires models to solve problems sequentially, allowing them to leverage the experience gained from previous solutions. EvaLearn provides five comprehensive automated metrics to evaluate models and quantify their learning capability and efficiency. We extensively benchmark nine frontier models and observe varied performance profiles: some models, such as Claude-3.7-sonnet, start with moderate initial performance but exhibit strong learning ability, while some models struggle to benefit from experience and may even show negative transfer. Moreover, we investigate model performance under two learning settings and find that instance-level rubrics and teacher-model feedback further facilitate model learning. Importantly, we observe that current LLMs with stronger static abilities do not show a clear advantage in learning capability across all tasks, highlighting that EvaLearn evaluates a new dimension of model performance. We hope EvaLearn provides a novel evaluation perspective for assessing LLM potential and understanding the gap between models and human capabilities, promoting the development of deeper and more dynamic evaluation approaches. All datasets, the automatic evaluation framework, and the results studied in this paper are available in the supplementary materials.
Shihan Dou, Ming Zhang 0030, Chenhao Huang, Feng Chen 0042, Shichun Liu, Yan Liu 0002, Chenxiao Liu, Zongzhang Zhang, Tao Gui, Chao Xin, Wei Chengzhi, Qi Zhang 0001, Xuanjing Huang 0001
NeurIPS11
2025 INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
abstract
Large Multimodal Models (LMMs) have made significant breakthroughs with the advancement of instruction tuning. However, while existing models can understand images and videos at a holistic level, they still struggle with instance-level understanding that requires a more fine-grained comprehension and alignment. Instance-level understanding is crucial for LMMs, as it focuses on the specific elements that we are most interested in. Excitingly, existing works find that the state-of-the-art LMMs exhibit strong instance understanding capabilities when provided with explicit visual cues. Motivated by this, we proposed Inst-IT, a solution to enhance LMMs in Instance understanding via explicit visual prompt Instruction Tuning for instance guidance. Inst-IT consists of a benchmark to diagnose multimodal instance-level understanding, a large-scale instruction-tuning dataset, and a continuous instruction-tuning training paradigm to effectively enhance spatial-temporal instance understanding capabilities of existing LMMs. Experimental results show that, enhanced by Inst-IT, our models not only achieve outstanding performance on Inst-IT-Bench and other instance understanding benchmarks, but also demonstrate significant improvements across various generic image and video understanding benchmarks. This highlights that our method not only boosts instance-level understanding but also strengthens the overall capabilities of generic image and video comprehension.
Wujian Peng, Lingchen Meng, Yiweng Xie, Yang Liu 0003, Tao Gui, Hang Xu 0004, Xipeng Qiu, Zuxuan Wu, Yu-Gang Jiang 0001
NeurIPS6
2025 BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
abstract
In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 100k university-level questions drawn from 300 UNESCO-defined subjects, spanning diverse formats—multiple-choice, fill-in-the-blank, and open-ended QA—and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop, automated, and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20k high-quality instances to comprehensively assess LMMs’ knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 80k instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline BMMR-Verifier for accurate and fine-grained evaluation of LMMs’ reasoning. Extensive experiments reveal that (i) even SOTA models leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data and models, and we believe our work can offers valuable insights and contributions to the community.
Zhiheng Xi, Yutao Fan, Honglin Guo, Yufang Liu, Xiaoran Fan, Jingchao Ding, Wangmeng Zuo, Zhenfei Yin, Lei Bai 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
NeurIPS13
2025 Understanding Parametric and Contextual Knowledge Reconciliation within Large Language Models
abstract
Retrieval-Augmented Generation (RAG) provides additional contextual knowledge to complement the parametric knowledge in Large Language Models (LLMs). These two knowledge interweave to enhance the accuracy and timeliness of LLM responses. However, the internal mechanisms by which LLMs utilize these knowledge remain unclear. We propose modeling the forward propagation of knowledge as an entity flow, employing this framework to trace LLMs' internal behaviors when processing mixed-source knowledge. Linear probing utilizes a trainable linear classifier to detect specific attributes in hidden layers. However, once trained, a probe cannot adapt to dynamically specified entities. To address this challenge, we construct an entity-aware probe, which introduces special tokens to mark probing targets and employs a small trainable rank-8 lora update to process these special markers. We first verify this approach through an attribution experiment, demonstrating that it can accurately detect information about ad-hoc entities from complex hidden states. Next, we trace entity flows across layers to understand how LLMs reconcile conflicting knowledge internally. Our probing results reveal that contextual and parametric knowledge are routed between tokens through distinct sets of attention heads, supporting attention competition only within knowledge types. While conflicting knowledge maintains a residual presence across layers, aligned knowledge from multiple sources gradually accumulates, with the magnitude of this accumulation directly determining its influence on final outputs.
Jun Zhao 0019, Yongzhuo Yang, Jingqi Tong, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
NeurIPS7
2025 Improving RL Exploration for LLM Reasoning Through Retrospective Replay
Shihan Dou, Muling Wu, Tao Gui, Qi Zhang 0001
NLPCC (1)5
2025 Visual Sketchbook: Enhancing Chart-to-Code Generation via Reflective Refinement
Junzhe Wang 0001, Zhiheng Xi, Wei He 0024, Dingwei Zhu, Shihan Dou, Tao Gui, Qi Zhang 0001
NLPCC (2)6
2025 The rise and potential of large language model based agents: a survey
Zhiheng Xi, Wenxiang Chen, Wei He 0024, Yiwen Ding, Boyang Hong, Ming Zhang 0030, Junzhe Wang 0001, Senjie Jin, Enyu Zhou, Xiaoran Fan, Xiao Wang 0001, Limao Xiong, Yuhao Zhou 0005, Weiran Wang 0003, Changhao Jiang, Yicheng Zou, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang 0001, Qi Zhang 0001, Tao Gui
Sci. China Inf. Sci.28
2024 LLMEval: A Preliminary Study on How to Evaluate Large Language Models
abstract
Recently, the evaluation of Large Language Models has emerged as a popular area of research. The three crucial questions for LLM evaluation are ``what, where, and how to evaluate''. However, the existing research mainly focuses on the first two questions, which are basically what tasks to give the LLM during testing and what kind of knowledge it should deal with. As for the third question, which is about what standards to use, the types of evaluators, how to score, and how to rank, there hasn't been much discussion. In this paper, we analyze evaluation methods by comparing various criteria with both manual and automatic evaluation, utilizing onsite, crowd-sourcing, public annotators and GPT-4, with different scoring methods and ranking systems. We propose a new dataset, LLMEval and conduct evaluations on 20 LLMs. A total of 2,186 individuals participated, leading to the generation of 243,337 manual annotations and 57,511 automatic evaluation results. We perform comparisons and analyses of different settings and conduct 10 conclusions that can provide some insights for evaluating LLM in the future. The dataset and the results are publicly available at https://github.com/llmeval. The version with the appendix are publicly available at https://arxiv.org/abs/2312.07398.
Yue Zhang 0004, Ming Zhang 0030, Haipeng Yuan, Shichun Liu, Yongyao Shi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
AAAI6
2024 StepCoder: Improving Code Generation with Reinforcement Learning from Compiler Feedback
abstract
Shihan Dou, Yan Liu, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou, Tao Ji, Rui Zheng, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shihan Dou, Yan Liu 0002, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang 0001, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
ACL (1)15
2024 LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style Plugin
abstract
Shihan Dou, Enyu Zhou, Yan Liu, Songyang Gao, Wei Shen, Limao Xiong, Yuhao Zhou, Xiao Wang, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Jiang Zhu, Rui Zheng, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shihan Dou, Enyu Zhou, Yan Liu 0002, Songyang Gao, Limao Xiong, Yuhao Zhou 0005, Xiao Wang 0001, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)14
2024 Navigating the OverKill in Large Language Models
abstract
Chenyu Shi, Xiao Wang, Qiming Ge, Songyang Gao, Xianjun Yang, Tao Gui, Qi Zhang, Xuanjing Huang, Xun Zhao, Dahua Lin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Chenyu Shi, Xiao Wang 0042, Qiming Ge, Songyang Gao, Xianjun Yang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Dahua Lin
ACL (1)6
2024 Enhancing Contrastive Learning with Noise-Guided Attack: Towards Continual Relation Extraction in the Wild
abstract
The principle of continual relation extraction (CRE) involves adapting to emerging novel relations while preserving old knowledge.Existing CRE approaches excel in preserving old knowledge but falter when confronted with contaminated data streams, likely due to an artificial assumption of no annotation errors.Recognizing the prevalence of noisy labels in realworld datasets, we introduce a more practical learning scenario, termed as noisy-CRE.In response to this challenge, we propose a noiseresistant contrastive framework called Noiseguided Attack in Contrastive Learning (NaCL), aimed at learning incremental corrupted relations.Diverging from conventional approaches like sample discarding or relabeling in the presence of noisy labels, NaCL takes a transformative route by modifying the feature space through targeted attack.This attack aims to align the feature space with the provided, albeit inaccurate, labels, thereby enhancing contrastive representations.Extensive empirical validations demonstrate the consistent performance improvement of NaCL with increasing noise rates, surpassing state-of-the-art methods 1 .
Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)4
2024 ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages
abstract
Junjie Ye, Sixian Li, Guanyu Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Junjie Ye 0005, Sixian Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
ACL (1)8
2024 AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
abstract
Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Junqi Dai, Jiasheng Ye, Yunhua Zhou, Zhigeng Liu, Ruibin Yuan, Ge Zhang 0009, Linyang Li, Hang Yan 0001, Jie Fu 0001, Tao Gui, Tianxiang Sun, Yu-Gang Jiang 0001, Xipeng Qiu
ACL (1)13
2024 Unveiling Linguistic Regions in Large Language Models
abstract
Large Language Models (LLMs) have demonstrated considerable cross-lingual alignment and generalization ability.Current research primarily focuses on improving LLMs' crosslingual generalization capabilities.However, there is still a lack of research on the intrinsic mechanisms of how LLMs achieve crosslingual alignment.From the perspective of region partitioning, this paper conducts several investigations on the linguistic competence of LLMs.We discover a core region in LLMs that corresponds to linguistic competence, accounting for approximately 1% of the total model parameters.Removing this core region by setting parameters to zero results in a significant performance decrease across 30 different languages.Furthermore, this core region exhibits significant dimensional dependence, perturbations to even a single parameter on specific dimensions leading to a loss of linguistic competence.Moreover, we discover that distinct monolingual regions exist for different languages, and disruption to these specific regions substantially reduces the LLMs' proficiency in those corresponding languages.Our research also indicates that freezing the core linguistic region during further pre-training can mitigate the issue of catastrophic forgetting (CF), a common phenomenon observed during further pre-training of LLMs.Overall, exploring the LLMs' functional regions provides insights into the foundation of their intelligence 1 . * Equal contributions.† Corresponding authors. 1 Our code is released in https://github.com/ zzhang0179/Unveiling-Linguistic-Regions-in-LLMs.
Zhihao Zhang 0002, Jun Zhao 0019, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
ACL (1)4
2024 Domain Generalization via Causal Adjustment for Cross-Domain Sentiment Analysis
abstract
Domain adaption has been widely adapted for cross-domain sentiment analysis to transfer knowledge from the source domain to the target domain. Whereas, most methods are proposed under the assumption that the target (test) domain is known, making them fail to generalize well on unknown test data that is not always available in practice. In this paper, we focus on the problem of domain generalization for cross-domain sentiment analysis. Specifically, we propose a backdoor adjustment-based causal model to disentangle the domain-specific and domain-invariant representations that play essential roles in tackling domain shift. First, we rethink the cross-domain sentiment analysis task in a causal view to model the causal-and-effect relationships among different variables. Then, to learn an invariant feature representation, we remove the effect of domain confounders (e.g., domain knowledge) using the backdoor adjustment. A series of experiments over many homologous and diverse datasets show the great performance and robustness of our model by comparing it with the state-of-the-art domain generalization baselines.
Siyin Wang, Jie Zhou 0015, Qin Chen 0001, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
LREC/COLING5
2024 RoCoIns: Enhancing Robustness of Large Language Models through Code-Style Instructions
abstract
Large Language Models (LLMs) have showcased remarkable capabilities in following human instructions. However, recent studies have raised concerns about the robustness of LLMs for natural language understanding (NLU) tasks when prompted with instructions combining textual adversarial samples. In this paper, drawing inspiration from recent works that LLMs are sensitive to the design of the instructions, we utilize instructions in code style, which are more structural and less ambiguous, to replace typically natural language instructions. Through this conversion, we provide LLMs with more precise instructions and strengthen the robustness of LLMs. Moreover, under few-shot scenarios, we propose a novel method to compose in-context demonstrations using both clean and adversarial samples (adversarial context method) to further boost the robustness of the LLMs. Experiments on eight robustness datasets show that our method consistently outperforms prompting LLMs with natural language, for example, with gpt-3.5-turbo on average, our method achieves an improvement of 5.68% in test set accuracy and a reduction of 5.66 points in Attack Success Rate (ASR).
Yuansen Zhang, Xiao Wang 0001, Zhiheng Xi, Han Xia 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
LREC/COLING5
2024 Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean Signals
abstract
Deep neural networks (DNNs) are notoriously vulnerable to adversarial attacks that place carefully crafted perturbations on normal examples to fool DNNs. To better understand such attacks, a characterization of the features carried by adversarial examples is needed. In this paper, we tackle this challenge by inspecting the subspaces of sample features through spectral analysis. We first empirically show that the features of either clean signals or adversarial perturbations are redundant and span in low-dimensional linear subspaces respectively with minimal overlap, and the classical low-dimensional subspace projection can suppress perturbation features out of the subspace of clean signals. This makes it possible for DNNs to learn a subspace where only features of clean signals exist while those of perturbations are discarded, which can facilitate the distinction of adversarial examples. To prevent the residual perturbations that is inevitable in subspace learning, we propose an independence criterion to disentangle clean signals from perturbations. Experimental results show that the proposed strategy enables the model to inherently suppress adversaries, which not only boosts model robustness but also motivates new directions of effective adversarial defense.
Yuhao Zhou 0005, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
LREC/COLING4
2024 ORTicket: Let One Robust BERT Ticket Transfer across Different Tasks
abstract
Pretrained language models can be applied for various downstream tasks but are susceptible to subtle perturbations. Most adversarial defense methods often introduce adversarial training during the fine-tuning phase to enhance empirical robustness. However, the repeated execution of adversarial training hinders training efficiency when transitioning to different tasks. In this paper, we explore the transferability of robustness within subnetworks and leverage this insight to introduce a novel adversarial defense method ORTicket, eliminating the need for separate adversarial training across diverse downstream tasks. Specifically, (i) pruning the full model using the MLM task (the same task employed for BERT pretraining) yields a task-agnostic robust subnetwork(i.e., winning ticket in Lottery Ticket Hypothesis); and (ii) fine-tuning this subnetwork for downstream tasks. Extensive experiments demonstrate that our approach achieves comparable robustness to other defense methods while retaining the efficiency of traditional fine-tuning.This also confirms the significance of selecting MLM task for identifying the transferable robust subnetwork. Furthermore, our method is orthogonal to other adversarial training approaches, indicating the potential for further enhancement of model robustness.
Yuhao Zhou 0005, Wenxiang Chen, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
LREC/COLING5
2024 LONGAGENT: Achieving Question Answering for 128k-Token-Long Documents through Multi-Agent Collaboration
abstract
Large language models (LLMs) have achieved tremendous success in understanding language and processing text.However, questionanswering (QA) on lengthy documents faces challenges of resource constraints and a high propensity for errors, even for the most advanced models such as GPT-4 and Claude2.In this paper, we introduce LONGAGENT, a multi-agent collaboration method that enables efficient and effective QA over 128k-tokenlong documents.LONGAGENT adopts a divideand-conquer strategy, breaking down lengthy documents into shorter, more manageable text chunks.A leader agent comprehends the user's query and organizes the member agents to read their assigned chunks, reasoning a final answer through multiple rounds of discussion.Due to members' hallucinations, it's difficult to guarantee that every response provided by each member is accurate.To address this, we develop an inter-member communication mechanism that facilitates information sharing, allowing for the detection and mitigation of hallucinatory responses.Experimental results show that a LLaMA-2 7B driven by LONGAGENT can effectively support QA over 128k-token documents, achieving 16.42% and 1.63% accuracy gains over GPT-4 on singlehop and multi-hop QA settings, respectively.
Jun Zhao 0019, Can Zu, Xu Hao, Wei He 0024, Yiwen Ding, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP7
2024 Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning
abstract
Lu Chen, Rui Zheng, Binghai Wang, Senjie Jin, Caishuang Huang, Junjie Ye, Zhihao Zhang, Yuhao Zhou, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Lu Chen 0001, Binghai Wang, Senjie Jin, Caishuang Huang, Junjie Ye 0005, Zhihao Zhang 0002, Yuhao Zhou 0005, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP10
2024 RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning
abstract
Junjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Guanyu Li, Xiaoran Fan, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Junjie Ye 0005, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Xiaoran Fan, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
EMNLP9
2024 TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer Capabilities
abstract
Ming Zhang, Caishuang Huang, Yilong Wu, Shichun Liu, Huiyuan Zheng, Yurui Dong, Yujiong Shen, Shihan Dou, Jun Zhao, Junjie Ye, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Ming Zhang 0030, Caishuang Huang, Yilong Wu, Shichun Liu, Huiyuan Zheng, Yurui Dong 0001, Yujiong Shen, Shihan Dou, Jun Zhao 0019, Junjie Ye 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001
EMNLP12
2024 Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
abstract
Chong Zhang, Yi Tu, Yixi Zhao, Chenshu Yuan, Huan Chen, Yue Zhang, Mingxu Chai, Ya Guo, Huijia Zhu, Qi Zhang, Tao Gui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yixi Zhao, Chenshu Yuan, Huan Chen 0012, Yue Zhang 0073, Mingxu Chai, Huijia Zhu, Qi Zhang 0001, Tao Gui
EMNLP11
2024 Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs
abstract
Xin Zhou, Ping Nie, Yiwen Guo, Haojie Wei, Zhanqiu Zhang, Pasquale Minervini, Ruotian Ma, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Xin Zhou 0012, Ping Nie, Yiwen Guo, Haojie Wei, Zhanqiu Zhang, Pasquale Minervini, Ruotian Ma, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP8
2024 A Soft Contrastive Learning-Based Prompt Model for Few-Shot Sentiment Analysis
abstract
Few-shot text classification has attracted great interest in both academia and industry due to the lack of labeled data in many fields. Different from general text classification (e.g., topic classification), few-shot sentiment classification is more challenging because the semantic distances among the classes are more subtle. For instance, the semantic distances between the sentiment labels in a positive or negative polarity (e.g., "love" and "joy", "remorse" and "sadness") are close, while the distances are large for the sentiment labels in two opposite polarities (e.g., "love" and "sadness"). To address this problem, we propose a Soft Contrastive learning-based Prompt (SCP) model for few-shot sentiment analysis. First, we design a sentiment-aware chain of thought prompt module to guide the model to predict the sentiment from coarse grain to fine grain via a series of intermediate reasoning steps. Then, we propose a soft contrastive learning algorithm to take the correlation of the labels into account. A series of experiments on several sentiment analysis datasets show the great advantages of SCP by comparing it with SOTA baselines (e.g., ChatGPT).
Jie Zhou 0015, Jiabao Zhao, Siyin Wang, Haijun Shan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ICASSP6
2024 Improving Generalization of Alignment with Human Preferences through Group Invariant Learning
abstract
The success of AI assistants based on language models (LLMs) hinges crucially on Reinforcement Learning from Human Feedback (RLHF), which enables the generation of responses more aligned with human preferences. As universal AI assistants, there's a growing expectation for them to perform consistently across various domains. However, previous work shows that Reinforcement Learning (RL) often exploits shortcuts to attain high rewards and overlooks challenging samples. This focus on quick reward gains undermines both the stability in training and the model's ability to generalize to new, unseen data. In this work, we propose a novel approach that can learn a consistent policy via RL across various data groups or domains. Given the challenges associated with acquiring group annotations, our method automatically classifies data into different groups, deliberately maximizing performance variance. Then, we optimize the policy to perform well on challenging groups. Lastly, leveraging the established groups, our approach adaptively adjusts the exploration space, allocating more learning capacity to more challenging data and preventing the model from over-optimizing on simpler data. Experimental results indicate that our approach significantly enhances training stability and model generalization.
Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou 0005, Zhiheng Xi, Xiao Wang 0001, Haoran Huang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ICLR10
2024 Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
abstract
In this paper, we propose R$^3$: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the benefits of process supervision for large language models. The core challenge in applying RL to complex reasoning is to identify a sequence of actions that result in positive rewards and provide appropriate supervision for optimization. Outcome supervision provides sparse rewards for final results without identifying error locations, whereas process supervision offers step-wise rewards but requires extensive manual annotation. R$^3$ overcomes these limitations by learning from correct demonstrations. Specifically, R$^3$ progressively slides the start state of reasoning from a demonstration’s end to its beginning, facilitating easier model exploration at all stages. Thus, R$^3$ establishes a step-wise curriculum, allowing outcome supervision to offer step-level signals and precisely pinpoint errors. Using Llama2-7B, our method surpasses RL baseline on eight reasoning tasks by $4.1$ points on average. Notably, in program-based reasoning, 7B-scale models perform comparably to larger models or closed-source models with our R$^3$.
Zhiheng Xi, Wenxiang Chen, Boyang Hong, Senjie Jin, Wei He 0024, Yiwen Ding, Shichun Liu, Junzhe Wang 0001, Honglin Guo, Xiaoran Fan, Yuhao Zhou 0005, Shihan Dou, Xiao Wang 0001, Xinbo Zhang, Peng Sun 0006, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ICML19
2024 CausalAPM: Generalizable Literal Disentanglement for NLU Debiasing
Shihan Dou, Songyang Gao, Tao Gui, Qi Zhang 0001
NLPCC (1)3
2024 Visual Explanation for Open-Domain Question Answering With BERT
abstract
Open-domain question answering (OpenQA) is an essential but challenging task in natural language processing that aims to answer questions in natural language formats on the basis of large-scale unstructured passages. Recent research has taken the performance of benchmark datasets to new heights, especially when these datasets are combined with techniques for machine reading comprehension based on Transformer models. However, as identified through our ongoing collaboration with domain experts and our review of literature, three key challenges limit their further improvement: (i) complex data with multiple long texts, (ii) complex model architecture with multiple modules, and (iii) semantically complex decision process. In this paper, we present VEQA, a visual analytics system that helps experts understand the decision reasons of OpenQA and provides insights into model improvement. The system summarizes the data flow within and between modules in the OpenQA model as the decision process takes place at the summary, instance and candidate levels. Specifically, it guides users through a summary visualization of dataset and module response to explore individual instances with a ranking visualization that incorporates context. Furthermore, VEQA supports fine-grained exploration of the decision flow within a single module through a comparative tree visualization. We demonstrate the effectiveness of VEQA in promoting interpretability and providing insights into model enhancement through a case study and expert evaluation.
Zekai Shao 0001, Shuran Sun, Yuheng Zhao, Siyuan Wang 0025, Zhongyu Wei, Tao Gui, Cagatay Turkay, Siming Chen 0001
IEEE Trans. Vis. Comput. Graph.6
2023 Learning "O" Helps for Learning More: Handling the Unlabeled Entity Problem for Class-incremental NER
abstract
Ruotian Ma, Xuanting Chen, Zhang Lin, Xin Zhou, Junzhe Wang, Tao Gui, Qi Zhang, Xiang Gao, Yun Wen Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ruotian Ma, Xuanting Chen, Xin Zhou 0012, Junzhe Wang 0001, Tao Gui, Qi Zhang 0001, Xiang Gao 0017, Yun Wen Chen
ACL (1)6
2023 Actively Supervised Clustering for Open Relation Extraction
abstract
Jun Zhao, Yongxin Zhang, Qi Zhang, Tao Gui, Zhongyu Wei, Minlong Peng, Mingming Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jun Zhao 0019, Qi Zhang 0001, Tao Gui, Zhongyu Wei, Minlong Peng, Mingming Sun 0001
ACL (1)4
2023 Open Set Relation Extraction via Unknown-Aware Training
abstract
Jun Zhao, Xin Zhao, WenYu Zhan, Qi Zhang, Tao Gui, Zhongyu Wei, Yun Wen Chen, Xiang Gao, Xuanjing Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jun Zhao 0019, Wenyu Zhan, Qi Zhang 0001, Tao Gui, Zhongyu Wei, Yun Wen Chen, Xiang Gao 0017, Xuanjing Huang 0001
ACL (1)5
2023 RE-Matching: A Fine-Grained Semantic Matching Method for Zero-Shot Relation Extraction
abstract
Jun Zhao, WenYu Zhan, Xin Zhao, Qi Zhang, Tao Gui, Zhongyu Wei, Junzhe Wang, Minlong Peng, Mingming Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jun Zhao 0019, Wenyu Zhan, Qi Zhang 0001, Tao Gui, Zhongyu Wei, Junzhe Wang 0001, Minlong Peng, Mingming Sun 0001
ACL (1)5
2023 Towards Understanding Omission in Dialogue Summarization
abstract
Dialogue summarization aims to condense the lengthy dialogue into a concise summary, and has recently achieved significant progress.However, the result of existing methods is still far from satisfactory.Previous works indicated that omission is a major factor in affecting the quality of summarization, but few of them have further explored the omission problem, such as how omission affects summarization results and how to detect omission, which is critical for reducing omission and improving summarization quality.Moreover, analyzing and detecting omission relies on summarization datasets with omission labels (i.e., which dialogue utterances are omitted in the summarization), which are not available in the current literature.In this paper, we propose the OLDS dataset, which provides high-quality Omission Labels for Dialogue Summarization.By analyzing this dataset, we find that a large improvement in summarization quality can be achieved by providing ground-truth omission labels for the summarization model to recover omission information, which demonstrates the importance of omission detection for omission mitigation in dialogue summarization.Therefore, we formulate an omission detection task and demonstrate our proposed dataset can support the training and evaluation of this task well.We also call for research action on omission detection based on our proposed datasets.Our dataset and codes are publicly available 1 .
Yicheng Zou, Kaitao Song, Xu Tan 0003, Zhongkai Fu, Qi Zhang 0001, Dongsheng Li 0002, Tao Gui
ACL (1)7
2023 Correspondence Transformers with Asymmetric Feature Learning and Matching Flow Super-Resolution
abstract
This paper solves the problem of learning dense visual correspondences between different object instances of the same category with only sparse annotations. We decompose this pixel-level semantic matching problem into two easier ones: (i) First, local feature descriptors of source and target images need to be mapped into shared semantic spaces to get coarse matching flows. (ii) Second, matching flows in low resolution should be refined to generate accurate point-to-point matching results. We propose asymmetric feature learning and matching flow super-resolution based on vision transformers to solve the above problems. The asymmetric feature learning module exploits a biased cross-attention mechanism to encode token features of source images with their target counterparts. Then matching flow in low resolutions is enhanced by a super-resolution network to get accurate correspondences. Our pipeline is built upon vision transformers and can be trained in an end-to-end manner. Extensive experimental results on several popular benchmarks, such as PF-PASCAL, PF-WILLOW, and SPair-71 K, demonstrate that the proposed method can catch subtle semantic differences in pixels efficiently. Code is available on https://github.com/YXSUNMADMAX/ACTR.
Yixuan Sun, Dongyang Zhao, Zhangyue Yin, Tao Gui, Weifeng Ge
CVPR5
2023 Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction
abstract
Recent advances in multimodal pre-trained models have significantly improved information extraction from visually-rich documents (VrDs), in which named entity recognition (NER) is treated as a sequence-labeling task of predicting the BIO entity tags for tokens, following the typical setting of NLP.However, BIO-tagging scheme relies on the correct order of model inputs, which is not guaranteed in real-world NER on scanned VrDs where text are recognized and arranged by OCR systems.Such reading order issue hinders the accurate marking of entities by BIO-tagging scheme, making it impossible for sequencelabeling methods to predict correct named entities.To address the reading order issue, we introduce Token Path Prediction (TPP), a simple prediction head to predict entity mentions as token sequences within documents.Alternative to token classification, TPP models the document layout as a complete directed graph of tokens, and predicts token paths within the graph as entities.For better evaluation of VrD-NER systems, we also propose two revised benchmark datasets of NER on scanned documents which can reflect real-world scenarios.Experiment results demonstrate the effectiveness of our method, and suggest its potential to be a universal solution to various information extraction tasks on documents.
Huan Chen 0012, Jinyang Tang, Huijia Zhu, Qi Zhang 0001, Tao Gui
EMNLP8
2022 CQG: A Simple and Effective Controlled Generation Framework for Multi-hop Question Generation
abstract
Multi-hop question generation focuses on generating complex questions that require reasoning over multiple pieces of information of the input passage.Current models with state-of-the-art performance have been able to generate the correct questions corresponding to the answers.However, most models can not ensure the complexity of generated questions, so they may generate shallow questions that can be answered without multi-hop reasoning.To address this challenge, we propose the CQG, which is a simple and effective controlled framework.CQG employs a simple method to generate the multi-hop questions that contain key entities in multi-hop reasoning chains, which ensure the complexity and quality of the questions.In addition, we introduce a novel controlled Transformer-based decoder to guarantee that key entities appear in the questions.Experiment results show that our model greatly improves performance, which also outperforms the state-of-the-art model about 25% by 5 BLEU points on HotpotQA 1 .
Zichu Fei, Qi Zhang 0001, Tao Gui, Di Liang, Wei Wu 0014, Xuanjing Huang 0001
ACL (1)3
2022 Flooding-X: Improving BERT's Resistance to Adversarial Attacks via Loss-Restricted Fine-Tuning
abstract
Qin Liu, Rui Zheng, Bao Rong, Jingyi Liu, ZhiHua Liu, Zhanzhan Cheng, Liang Qiao, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Qin Liu 0010, Bao Rong, Zhanzhan Cheng, Liang Qiao 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)8
2022 MINER: Improving Out-of-Vocabulary Named Entity Recognition from an Information Theoretic Perspective
abstract
Xiao Wang, Shihan Dou, Limao Xiong, Yicheng Zou, Qi Zhang, Tao Gui, Liang Qiao, Zhanzhan Cheng, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Xiao Wang 0001, Shihan Dou, Limao Xiong, Yicheng Zou, Qi Zhang 0001, Tao Gui, Liang Qiao 0001, Zhanzhan Cheng, Xuanjing Huang 0001
ACL (1)6
2022 Robust Lottery Tickets for Pre-trained Language Models
abstract
Rui Zheng, Bao Rong, Yuhao Zhou, Di Liang, Sirui Wang, Wei Wu, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Bao Rong, Yuhao Zhou 0005, Di Liang, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
ACL (1)7
2022 LFKQG: A Controlled Generation Framework with Local Fine-tuning for Question Generation over Knowledge Bases
abstract
Question generation over knowledge bases (KBQG) aims at generating natural questions about a subgraph, which can be answered by a given answer entity. Existing KBQG models still face two main challenges: (1) Most models often focus on the most relevant part of the answer entity, while neglecting the rest of the subgraph. (2) There are a large number of out-of-vocabulary (OOV) predicates in real-world scenarios, which are hard to adapt for most KBQG models. To address these challenges, we propose LFKQG, a controlled generation framework for Question Generation over Knowledge Bases. (1) LFKQG employs a simple controlled generation method to generate the questions containing the critical entities in the subgraph, ensuring the question is relevant to the whole subgraph. (2) We propose an optimization strategy called local fine-tuning, which can make good use of the rich information hidden in the pre-trained model to improve the ability of the model to adapt the OOV predicates. Extensive experiments show that our method outperforms existing methods significantly on three widely-used benchmark datasets SimpleQuestion, PathQuestions, and WebQuestions.
Zichu Fei, Xin Zhou 0012, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
COLING3
2022 Causal Intervention Improves Implicit Sentiment Analysis
abstract
Despite having achieved great success for sentiment analysis, existing neural models struggle with implicit sentiment analysis. It is because they may latch onto spurious correlations (“shortcuts”, e.g., focusing only on explicit sentiment words), resulting in undermining the effectiveness and robustness of the learned model. In this work, we propose a CausaL intervention model for implicit sEntiment ANalysis using instrumental variable (CLEAN). We first review sentiment analysis from a causal perspective and analyze the confounders existing in this task. Then, we introduce instrumental variable to eliminate the confounding causal effects, thus extracting the pure causal effect between sentence and sentiment. We compare the proposed CLEAN with several strong baselines on both the general implicit sentiment analysis and aspect-based implicit sentiment analysis tasks. The results indicate the great advantages of our model and the efficacy of implicit sentiment reasoning.
Siyin Wang, Jie Zhou 0015, Changzhi Sun, Junjie Ye 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
COLING5
2022 Less Is Better: Recovering Intended-Feature Subspace to Robustify NLU Models
abstract
Datasets with significant proportions of bias present threats for training a trustworthy model on NLU tasks. Despite yielding great progress, current debiasing methods impose excessive reliance on the knowledge of bias attributes. Definition of the attributes, however, is elusive and varies across different datasets. In addition, leveraging these attributes at input level to bias mitigation may leave a gap between intrinsic properties and the underlying decision rule. To narrow down this gap and liberate the supervision on bias, we suggest extending bias mitigation into feature space. Therefore, a novel model, Recovering Intended-Feature Subspace with Knowledge-Free (RISK) is developed. Assuming that shortcut features caused by various biases are unintended for prediction, RISK views them as redundant features. When delving into a lower manifold to remove redundancies, RISK reveals that an extremely low-dimensional subspace with intended features can robustly represent the highly biased dataset. Empirical results demonstrate our model can consistently improve model generalization to out-of-distribution set, and achieves a new state-of-the-art performance.
Tao Gui
COLING2
2022 Read Extensively, Focus Smartly: A Cross-document Semantic Enhancement Method for Visual Documents NER
abstract
The introduction of multimodal information and pretraining technique significantly improves entity recognition from visually-rich documents. However, most of the existing methods pay unnecessary attention to irrelevant regions of the current document while ignoring the potentially valuable information in related documents. To deal with this problem, this work proposes a cross-document semantic enhancement method, which consists of two modules: 1) To prevent distractions from irrelevant regions in the current document, we design a learnable attention mask mechanism, which is used to adaptively filter redundant information in the current document. 2) To further enrich the entity-related context, we propose a cross-document information awareness technique, which enables the model to collect more evidence across documents to assist in prediction. The experimental results on two documents understanding benchmarks covering eight languages demonstrate that our method outperforms the SOTA methods.
Jun Zhao 0019, Wenyu Zhan, Tao Gui, Qi Zhang 0001, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu
COLING4
2022 PlugAT: A Plug and Play Module to Defend against Textual Adversarial Attack
abstract
Adversarial training, which minimizes the loss of adversarially perturbed examples, has received considerable attention. However, these methods require modifying all model parameters and optimizing the model from scratch, which is parameter inefficient and unfriendly to the already deployed models. As an alternative, we propose a pluggable defense module PlugAT, to provide robust predictions by adding a few trainable parameters to the model inputs while keeping the original model frozen. To reduce the potential side effects of using defense modules, we further propose a novel forgetting restricted adversarial training, which filters out bad adversarial examples that impair the performance of original ones. The PlugAT-equipped BERT model substantially improves robustness over several strong baselines on various text classification tasks, whilst training only 9.1% parameters. We observe that defense modules trained under the same model architecture have domain adaptation ability between similar text classification datasets.
Rong Bao, Qin Liu 0010, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Rui Xie 0005, Wei Wu 0014
COLING4
2022 Making Parameter-efficient Tuning More Efficient: A Unified Framework for Classification Tasks
abstract
Large pre-trained language models (PLMs) have demonstrated superior performance in industrial applications. Recent studies have explored parameter-efficient PLM tuning, which only updates a small amount of task-specific parameters while achieving both high efficiency and comparable performance against standard fine-tuning. However, all these methods ignore the inefficiency problem caused by the task-specific output layers, which is inflexible for us to re-use PLMs and introduces non-negligible parameters. In this work, we focus on the text classification task and propose plugin-tuning, a framework that further improves the efficiency of existing parameter-efficient methods with a unified classifier. Specifically, we re-formulate both token and sentence classification tasks into a unified language modeling task, and map label spaces of different tasks into the same vocabulary space. In this way, we can directly re-use the language modeling heads of PLMs, avoiding introducing extra parameters for different tasks. We conduct experiments on six classification benchmarks. The experimental results show that plugin-tuning can achieve comparable performance against fine-tuned PLMs, while further saving around 50% parameters on top of other parameter-efficient methods.
Xin Zhou 0012, Ruotian Ma, Yicheng Zou, Xuanting Chen, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Rui Xie 0005, Wei Wu 0014
COLING5
2022 ProofInfer: Generating Proof via Iterative Hierarchical Inference
abstract
Proof generation focuses on deductive reasoning: given a hypothesis and a set of theories, including some supporting facts and logical rules expressed in natural language, the model generates a proof tree indicating how to deduce the hypothesis from given theories.Current models with state-of-theart performance employ the stepwise method, linking an individual node to the proof step-bystep.However, these methods actually focus on generating several proof paths rather than a whole tree.To address this problem, we propose ProofInfer, which generates the proof tree via iterative hierarchical inference.At each step, ProofInfer generates the entire layer for proof tree, where all nodes in this layer are generated simultaneously.Since the conventional autoregressive generation architecture cannot simultaneously predict multiple nodes, ProofInfer employs text-to-text paradigm to avoid it.To this end, we propose a divideand-conquer algorithm to encode the proof tree as the plain text recursively without structure information loss.Experimental results show that ProofInfer significantly outperforms the state-of-the-art (SOTA) models on several widely-used datasets.In addition, ProofInfer still performs well with data-limited, achieving comparable performance to the SOTA models with only 40% of the training data. 1
Zichu Fei, Qi Zhang 0001, Xin Zhou 0012, Tao Gui, Xuanjing Huang 0001
EMNLP4
2022 Efficient Adversarial Training with Robust Early-Bird Tickets
abstract
Adversarial training is one of the most powerful methods to improve the robustness of pretrained language models (PLMs).However, this approach is typically more expensive than traditional fine-tuning because of the necessity to generate adversarial examples via gradient descent.Delving into the optimization process of adversarial training, we find that robust connectivity patterns emerge in the early training phase (typically 0.15 ∼ 0.3 epochs), far before parameters converge.Inspired by this finding, we dig out robust early-bird tickets (i.e., subnetworks) to develop an efficient adversarial training method: (1) searching for robust tickets with structured sparsity in the early stage; (2) fine-tuning robust tickets in the remaining time.To extract the robust tickets as early as possible, we design a ticket convergence metric to automatically terminate the searching process.Experiments show that the proposed efficient adversarial training method can achieve up to 7× ∼ 13× training speedups while maintaining comparable or even better robustness compared to the most competitive state-of-the-art adversarial training methods.
Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP3
2022 Cross-Linguistic Syntactic Difference in Multilingual BERT: How Good is It and How Does It Affect Transfer?
abstract
Multilingual BERT (mBERT) has demonstrated considerable cross-lingual syntactic ability, whereby it enables effective zero-shot cross-lingual transfer of syntactic knowledge.The transfer is more successful between some languages, but it is not well understood what leads to this variation and whether it fairly reflects difference between languages.In this work, we investigate the distributions of grammatical relations induced from mBERT in the context of 24 typologically different languages.We demonstrate that the distance between the distributions of different languages is highly consistent with the syntactic difference in terms of linguistic formalisms.Such difference learnt via self-supervision plays a crucial role in the zero-shot transfer performance and can be predicted by variation in morphosyntactic properties between languages.These results suggest that mBERT properly encodes languages in a way consistent with linguistic diversity and provide insights into the mechanism of crosslingual transfer.
Ningyu Xu, Tao Gui, Ruotian Ma, Qi Zhang 0001, Jingting Ye, Menghan Zhang, Xuanjing Huang 0001
EMNLP2
2022 TextFusion: Privacy-Preserving Pre-trained Model Inference via Token Fusion
abstract
Xin Zhou, Jinzhu Lu, Tao Gui, Ruotian Ma, Zichu Fei, Yuran Wang, Yong Ding, Yibo Cheung, Qi Zhang, Xuanjing Huang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xin Zhou 0012, Jinzhu Lu, Tao Gui, Ruotian Ma, Zichu Fei, Yibo Cheung, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP3
2022 Searching for Optimal Subword Tokenization in Cross-domain NER
abstract
Input distribution shift is one of the vital problems in unsupervised domain adaptation (UDA). The most popular UDA approaches focus on domain-invariant representation learning, trying to align the features from different domains into a similar feature distribution. However, these approaches ignore the direct alignment of input word distributions between domains, which is a vital factor in word-level classification tasks such as cross-domain NER. In this work, we shed new light on cross-domain NER by introducing a subword-level solution, X-Piece, for input word-level distribution shift in NER. Specifically, we re-tokenize the input words of the source domain to approach the target subword distribution, which is formulated and solved as an optimal transport problem. As this approach focuses on the input level, it can also be combined with previous DIRL methods for further improvement. Experimental results show the effectiveness of the proposed method based on BERT-tagger on four benchmark NER datasets. Also, the proposed method is proved to benefit DIRL methods such as DANN.
Ruotian Ma, Yiding Tan, Xin Zhou 0012, Xuanting Chen, Di Liang, Wei Wu 0014, Tao Gui
IJCAI8
2022 Template-free Prompt Tuning for Few-shot NER
abstract
Ruotian Ma, Xin Zhou, Tao Gui, Yiding Tan, Linyang Li, Qi Zhang, Xuanjing Huang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Ruotian Ma, Xin Zhou 0012, Tao Gui, Yiding Tan, Linyang Li, Qi Zhang 0001, Xuanjing Huang 0001
NAACL-HLT3
2022 Sentiment-aware multimodal pre-training for multimodal sentiment analysis
Junjie Ye 0005, Jie Zhou 0015, Rui Wang 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
Knowl. Based Syst.6
2022 Uncertainty-Aware Sequence Labeling
abstract
Conditional random fields (CRFs) have been widely used for sequence labeling tasks in the field of natural language processing. However, how to model both local and global dependencies among labels is not well solved yet. In this study, we introduce a novel two-stage label decoding method to better model the short- and long-term label dependencies, while being much more computationally efficient with the use of graphics processing units (GPUs). A base model is first used to propose draft labels, and then a novel two-stream self-attention model makes refinements on these draft predictions based on long-range label dependencies. Besides, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with high probabilities of being wrong, which helps to mitigate the error propagation. Not only can our method model sentence-level label dependencies, but it is also easily extended to document-level sequence labeling by querying and storing a key-value memory matrix with label co-occurrence relationships. The experimental results on both sentence-level and document-level sequence labeling benchmarks show that the proposed method outperforms existing label decoding methods while taking advantage of parallel computations on GPUs.
Jiacheng Ye, Xiaoqing Zheng, Tao Gui, Qi Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 SENT: Sentence-level Distant Relation Extraction via Negative Training
abstract
Ruotian Ma, Tao Gui, Linyang Li, Qi Zhang, Xuanjing Huang, Yaqian Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ruotian Ma, Tao Gui, Linyang Li, Qi Zhang 0001, Xuanjing Huang 0001, Yaqian Zhou 0001
ACL/IJCNLP (1)2
2021 A Unified Generative Framework for Various NER Subtasks
abstract
Hang Yan, Tao Gui, Junqi Dai, Qipeng Guo, Zheng Zhang, Xipeng Qiu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Hang Yan 0001, Tao Gui, Junqi Dai, Qipeng Guo, Zheng Zhang 0001, Xipeng Qiu
ACL/IJCNLP (1)2
2021 One2Set: Generating Diverse Keyphrases as a Set
abstract
Jiacheng Ye, Tao Gui, Yichao Luo, Yige Xu, Qi Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jiacheng Ye, Tao Gui, Yichao Luo, Yige Xu 0001, Qi Zhang 0001
ACL/IJCNLP (1)2
2021 Heterogeneous Graph Neural Networks for Keyphrase Generation
abstract
The encoder-decoder framework achieves stateof-the-art results in keyphrase generation (KG) tasks by predicting both present keyphrases that appear in the source document and absent keyphrases that do not.However, relying solely on the source document can result in generating uncontrollable and inaccurate absent keyphrases.To address these problems, we propose a novel graph-based method that can capture explicit knowledge from related references.Our model first retrieves some document-keyphrases pairs similar to the source document from a pre-defined index as references.Then a heterogeneous graph is constructed to capture relationships of different granularities between the source document and its references.To guide the decoding process, a hierarchical attention and copy mechanism is introduced, which directly copies appropriate words from both the source document and its references based on their relevance and significance.The experimental results on multiple KG benchmarks show that the proposed model achieves significant improvements against other baseline models, especially with regard to the absent keyphrase prediction.
Jiacheng Ye, Ruijian Cai, Tao Gui, Qi Zhang 0001
EMNLP (1)3
2021 A Relation-Oriented Clustering Method for Open Relation Extraction
abstract
The clustering-based unsupervised relation discovery method has gradually become one of the important methods of open relation extraction (OpenRE).However, high-dimensional vectors can encode complex linguistic information which leads to the problem that the derived clusters cannot explicitly align with the relational semantic classes.In this work, we propose a relationoriented clustering model and use it to identify the novel relations in the unlabeled data.Specifically, to enable the model to learn to cluster relational data, our method leverages the readily available labeled data of pre-defined relations to learn a relationoriented representation.We minimize distance between the instance with same relation by gathering the instances towards their corresponding relation centroids to form a cluster structure, so that the learned representation is cluster-friendly.To reduce the clustering bias on predefined classes, we optimize the model by minimizing a joint objective on both labeled and unlabeled data.Experimental results show that our method reduces the error rate by 29.2% and 15.7%, on two datasets respectively, compared with current SOTA methods.
Jun Zhao 0019, Tao Gui, Qi Zhang 0001, Yaqian Zhou 0001
EMNLP (1)2
2021 Low-Resource Dialogue Summarization with Domain-Agnostic Multi-Source Pretraining
abstract
With the rapid increase in the volume of dialogue data from daily life, there is a growing demand for dialogue summarization.Unfortunately, training a large summarization model is generally infeasible due to the inadequacy of dialogue data with annotated summaries.Most existing works for low-resource dialogue summarization directly pretrain models in other domains, e.g., the news domain, but they generally neglect the huge difference between dialogues and conventional articles.To bridge the gap between out-of-domain pretraining and indomain fine-tuning, in this work, we propose a multi-source pretraining paradigm to better leverage the external summary data.Specifically, we exploit large-scale in-domain nonsummary data to separately pretrain the dialogue encoder and the summary decoder.The combined encoder-decoder model is then pretrained on the out-of-domain summary data using adversarial critics, aiming to facilitate domain-agnostic summarization.The experimental results on two public datasets show that with only limited training data, our approach achieves competitive performance and generalizes well in different dialogue scenarios.
Yicheng Zou, Bolin Zhu, Xingwu Hu, Tao Gui, Qi Zhang 0001
EMNLP (1)4
2020 Constructing Multiple Tasks for Augmentation: Improving Neural Image Classification with K-Means Features
abstract
Multi-task learning (MTL) has received considerable attention, and numerous deep learning applications benefit from MTL with multiple objectives. However, constructing multiple related tasks is difficult, and sometimes only a single task is available for training in a dataset. To tackle this problem, we explored the idea of using unsupervised clustering to construct a variety of auxiliary tasks from unlabeled data or existing labeled data. We found that some of these newly constructed tasks could exhibit semantic meanings corresponding to certain human-specific attributes, but some were non-ideal. In order to effectively reduce the impact of non-ideal auxiliary tasks on the main task, we further proposed a novel meta-learning-based multi-task learning approach, which trained the shared hidden layers on auxiliary tasks, while the meta-optimization objective was to minimize the loss on the main task, ensuring that the optimizing direction led to an improvement on the main task. Experimental results across five image datasets demonstrated that the proposed method significantly outperformed existing single task learning, semi-supervised learning, and some data augmentation methods, including an improvement of more than 9% on the Omniglot dataset.
Tao Gui, Lizhi Qing, Qi Zhang 0001, Jiacheng Ye, Hang Yan 0001, Zichu Fei, Xuanjing Huang 0001
AAAI1
2020 Uncertainty-Aware Label Refinement for Sequence Labeling
abstract
Conditional random fields (CRF) for label decoding has become ubiquitous in sequence labeling tasks.However, the local label dependencies and inefficient Viterbi decoding have always been a problem to be solved.In this work, we introduce a novel two-stage label decoding framework to model long-term label dependencies, while being much more computationally efficient.A base model first predicts draft labels, and then a novel twostream self-attention model makes refinements on these draft predictions based on longrange label dependencies, which can achieve parallel decoding for a faster prediction.In addition, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with a high probability of being wrong, which can greatly assist in preventing error propagation.The experimental results on three sequence labeling benchmarks demonstrated that the proposed method not only outperformed the CRF-based methods but also greatly accelerated the inference process.* Both authors contributed equally.
Tao Gui, Jiacheng Ye, Qi Zhang 0001, Zhengyan Li, Zichu Fei, Yeyun Gong, Xuanjing Huang 0001
EMNLP (1)1
2020 Leveraging Document-Level Label Consistency for Named Entity Recognition
abstract
Document-level label consistency is an effective indicator that different occurrences of a particular token sequence are very likely to have the same entity types. Previous work focused on better context representations and used the CRF for label decoding. However, CRF-based methods are inadequate for modeling document-level label consistency. This work introduces a novel two-stage label refinement approach to handle document-level label consistency, where a key-value memory network is first used to record draft labels predicted by the base model, and then a multi-channel Transformer makes refinements on these draft predictions based on the explicit co-occurrence relationship derived from the memory network. In addition, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with a high probability of being wrong, which can greatly assist in preventing the incorrect refinement of correct draft labels. The experimental results on three named entity recognition benchmarks demonstrated that the proposed method significantly outperformed the state-of-the-art methods.
Tao Gui, Jiacheng Ye, Qi Zhang 0001, Yaqian Zhou 0001, Yeyun Gong, Xuanjing Huang 0001
IJCAI1
2019 Switch-LSTMs for Multi-Criteria Chinese Word Segmentation
abstract
Multi-criteria Chinese word segmentation is a promising but challenging task, which exploits several different segmentation criteria and mines their common underlying knowledge. In this paper, we propose a flexible multi-criteria learning for Chinese word segmentation. Usually, a segmentation criterion could be decomposed into multiple sub-criteria, which are shareable with other segmentation criteria. The process of word segmentation is a routing among these sub-criteria. From this perspective, we present Switch-LSTMs to segment words, which consist of several long short-term memory neural networks (LSTM), and a switcher to automatically switch the routing among these LSTMs. With these auto-switched LSTMs, our model provides a more flexible solution for multi-criteria CWS, which is also easy to transfer the learned knowledge to new criteria. Experiments show that our model obtains significant improvements on eight corpora with heterogeneous segmentation criteria, compared to the previous method and single-criterion learning.
Jingjing Gong, Xinchi Chen, Tao Gui, Xipeng Qiu
AAAI3
2019 Cooperative Multimodal Approach to Depression Detection in Twitter
abstract
The advent of social media has presented a promising new opportunity for the early detection of depression. To do so effectively, there are two challenges to overcome. The first is that textual and visual information must be jointly considered to make accurate inferences about depression. The second challenge is that due to the variety of content types posted by users, it is difficult to extract many of the relevant indicator texts and images. In this work, we propose the use of a novel cooperative multi-agent model to address these challenges. From the historical posts of users, the proposed method can automatically select related indicator texts and images. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods by a large margin (over 30% error reduction). In several experiments and examples, we also verify that the selected posts can successfully indicate user depression, and our model can obtained a robust performance in realistic scenarios.
Tao Gui, Qi Zhang 0001, Minlong Peng, Keyu Ding
AAAI1
2019 Long Short-Term Memory with Dynamic Skip Connections
abstract
In recent years, long short-term memory (LSTM) has been successfully used to model sequential data of variable length. However, LSTM can still experience difficulty in capturing long-term dependencies. In this work, we tried to alleviate this problem by introducing a dynamic skip connection, which can learn to directly connect two dependent words. Since there is no dependency information in the training data, we propose a novel reinforcement learning-based method to model the dependency relationship and connect dependent words. The proposed model computes the recurrent transition functions based on the skip connections, which provides a dynamic skipping advantage over RNNs that always tackle entire sentences sequentially. Our experimental results on three natural language processing tasks demonstrate that the proposed method can achieve better performance than existing methods. In the number prediction experiment, the proposed model outperformed LSTM with respect to accuracy by nearly 20%.
Tao Gui, Qi Zhang 0001, Lujun Zhao, Yaosong Lin, Minlong Peng, Jingjing Gong, Xuanjing Huang 0001
AAAI1
2019 Trainable Undersampling for Class-Imbalance Learning
abstract
Undersampling has been widely used in the class-imbalance learning area. The main deficiency of most existing undersampling methods is that their data sampling strategies are heuristic-based and independent of the used classifier and evaluation metric. Thus, they may discard informative instances for the classifier during the data sampling. In this work, we propose a meta-learning method built on the undersampling to address this issue. The key idea of this method is to parametrize the data sampler and train it to optimize the classification performance over the evaluation metric. We solve the non-differentiable optimization problem for training the data sampler via reinforcement learning. By incorporating evaluation metric optimization into the data sampling process, the proposed method can learn which instance should be discarded for the given classifier and evaluation metric. In addition, as a data level operation, this method can be easily applied to arbitrary evaluation metric and classifier, including non-parametric ones (e.g., C4.5 and KNN). Experimental results on both synthetic and realistic datasets demonstrate the effectiveness of the proposed method.
Minlong Peng, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001, Yu-Gang Jiang 0001, Keyu Ding
AAAI4
2019 A Lexicon-Based Graph Neural Network for Chinese NER
abstract
Tao Gui, Yicheng Zou, Qi Zhang, Minlong Peng, Jinlan Fu, Zhongyu Wei, Xuanjing Huang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tao Gui, Yicheng Zou, Qi Zhang 0001, Minlong Peng, Jinlan Fu, Zhongyu Wei, Xuanjing Huang 0001
EMNLP/IJCNLP (1)1
2019 CNN-Based Chinese NER with Lexicon Rethinking
abstract
Character-level Chinese named entity recognition (NER) that applies long short-term memory (LSTM) to incorporate lexicons has achieved great success. However, this method fails to fully exploit GPU parallelism and candidate lexicons can conflict. In this work, we propose a faster alternative to Chinese NER: a convolutional neural network (CNN)-based method that incorporates lexicons using a rethinking mechanism. The proposed method can model all the characters and potential words that match the sentence in parallel. In addition, the rethinking mechanism can address the word conflict by feeding back the high-level features to refine the networks. Experimental results on four datasets show that the proposed method can achieve better performance than both word-level and character-level baseline methods. In addition, the proposed method performs up to 3.21 times faster than state-of-the-art methods, while realizing better performance.
Tao Gui, Ruotian Ma, Qi Zhang 0001, Lujun Zhao, Yu-Gang Jiang 0001, Xuanjing Huang 0001
IJCAI1
2019 Learning Task-Specific Representation for Novel Words in Sequence Labeling
abstract
Word representation is a key component in neural-network-based sequence labeling systems. However, representations of unseen or rare words trained on the end task are usually poor for appreciable performance. This is commonly referred to as the out-of-vocabulary (OOV) problem. In this work, we address the OOV problem in sequence labeling using only training data of the task. To this end, we propose a novel method to predict representations for OOV words from their surface-forms (e.g., character sequence) and contexts. The method is specifically designed to avoid the error propagation problem suffered by existing approaches in the same paradigm. To evaluate its effectiveness, we performed extensive empirical studies on four part-of-speech tagging (POS) tasks and four named entity recognition (NER) tasks. Experimental results show that the proposed method can achieve better or competitive performance on the OOV problem compared with existing state-of-the-art methods.
Minlong Peng, Qi Zhang 0001, Tao Gui, Jinlan Fu, Xuanjing Huang 0001
IJCAI4
2019 Model the Long-Term Post History for Hashtag Recommendation
Minlong Peng, Qiyuan Bian, Qi Zhang 0001, Tao Gui, Jinlan Fu, Lanjun Zeng, Xuanjing Huang 0001
NLPCC (1)4
2019 Mention Recommendation in Twitter with Cooperative Multi-Agent Reinforcement Learning
abstract
In Twitter-like social networking services, the "@'' symbol can be used with the tweet to mention users whom the user wants to alert regarding the message. An automatic suggestion to the user of a small list of candidate names can improve communication efficiency. Previous work usually used several most recent tweets or randomly select historical tweets to make an inference about this preferred list of names. However, because there are too many historical tweets by users and a wide variety of content types, the use of several tweets cannot guarantee the desired results. In this work, we propose the use of a novel cooperative multi-agent approach to mention recommendation, which incorporates dozens of more historical tweets than earlier approaches. The proposed method can effectively select a small set of historical tweets and cooperatively extract relevant indicator tweets from both the user and mentioned users. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods.
Tao Gui, Qi Zhang 0001, Minlong Peng, Yunhua Zhou, Xuanjing Huang 0001
SIGIR1
2019 Adaptive Multi-Attention Network Incorporating Answer Information for Duplicate Question Detection
abstract
Community-based question answering (CQA), which provides a platform for people with diverse backgrounds to share information and knowledge, has become increasingly popular. With the accumulation of site data, methods to detect duplicate questions in CQA sites have attracted considerable attention. Existing methods typically use only questions to complete the task. However, the paired answers may also provide valuable information. In this paper, we propose an answer information- enhanced adaptive multi-attention network (AMAN) to perform this task. AMAN takes full advantage of the semantic information in the paired answers while alleviating the noise problem caused by adding the answers. To evaluate the proposed method, we use a CQADupStack set and the Quora question-pair dataset expanded with paired answers. Experimental results demonstrate that the proposed model can achieve state-of-the-art performance on the above two data sets.
Di Liang, Fubao Zhang, Qi Zhang 0001, Jinlan Fu, Minlong Peng, Tao Gui, Xuanjing Huang 0001
SIGIR7
2019 Implicit discourse relation detection using concatenated word embeddings and a gated relevance network
Jinlan Fu, Qi Zhang 0001, Jifan Chen, Minlong Peng, Tao Gui, Xipeng Qiu, Xuanjing Huang 0001
Sci. China Inf. Sci.5
2018 A Lexicon-Based Supervised Attention Model for Neural Sentiment Analysis
abstract
Attention mechanisms have been leveraged for sentiment classification tasks because not all words have the same importance. However, most existing attention models did not take full advantage of sentiment lexicons, which provide rich sentiment information and play a critical role in sentiment analysis. To achieve the above target, in this work, we propose a novel lexicon-based supervised attention model (LBSA), which allows a recurrent neural network to focus on the sentiment content, thus generating sentiment-informative representations. Compared with general attention models, our model has better interpretability and less noise. Experimental results on three large-scale sentiment classification datasets showed that the proposed method outperforms previous methods.
Yicheng Zou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001
COLING2
2018 Transferring from Formal Newswire Domain with Hypernet for Twitter POS Tagging
abstract
Part-of-Speech (POS) tagging for Twitter has received considerable attention in recent years.Because most POS tagging methods are based on supervised models, they usually require a large amount of labeled data for training.However, the existing labeled datasets for Twitter are much smaller than those for newswire text.Hence, to help POS tagging for Twitter, most domain adaptation methods try to leverage newswire datasets by learning the shared features between the two domains.However, from a linguistic perspective, Twitter users not only tend to mimic the formal expressions of traditional media, like news, but they also appear to be developing linguistically informal styles.Therefore, POS tagging for the formal Twitter context can be learned together with the newswire dataset, while POS tagging for the informal Twitter context should be learned separately.To achieve this task, in this work, we propose a hypernetworkbased method to generate different parameters to separately model contexts with different expression styles.Experimental results on three different datasets show that our approach achieves better performance than state-of-theart methods in most cases.
Tao Gui, Qi Zhang 0001, Jingjing Gong, Minlong Peng, Di Liang, Keyu Ding, Xuanjing Huang 0001
EMNLP1
2017 Part-of-Speech Tagging for Twitter with Adversarial Neural Networks
abstract
In this work, we study the problem of partof-speech tagging for Tweets.In contrast to newswire articles, Tweets are usually informal and contain numerous out-ofvocabulary words.Moreover, there is a lack of large scale labeled datasets for this domain.To tackle these challenges, we propose a novel neural network to make use of out-of-domain labeled data, unlabeled in-domain data, and labeled indomain data.Inspired by adversarial neural networks, the proposed method tries to learn common features through adversarial discriminator.In addition, we hypothesize that domain-specific features of target domain should be preserved in some degree.Hence, the proposed method adopts a sequence-to-sequence autoencoder to perform this task.Experimental results on three different datasets show that our method achieves better performance than state-of-the-art methods.
Tao Gui, Qi Zhang 0001, Haoran Huang, Minlong Peng, Xuanjing Huang 0001
EMNLP1
2014 Optical spectrally efficient FDM system for electrical and optical bandwidth saving
abstract
A newly proposed optical-spectrally efficient frequency division multiplexing (O-SEFDM) system reduces the required communication spectrum by employing non-orthogonal and overlapping sub-carriers. This results in higher spectral efficiency relative to an equivalent optical-orthogonal frequency division multiplexing (O-OFDM) delivering the same data rate. O-SEFDM technique can save spectrums in both the electrical and optical domains. However, due to the loss of orthogonality, detection of O-SEFDM signals becomes more complicated. In this work, we employ a hybrid soft Iterative Detection (ID) together with fixed sphere decoder (FSD), concurrently optimizing performance and complexity. We show that for Bandwidth Compression Factor (BCF) of up to 25 percent, we can achieve the same performance as O-OFDM. This verifies Mazo's rates of transmission 25 percent faster than the Nyquist rate. We report a 4QAM system occupying approximately the same bandwidth as that of an 8QAM with 1.6 dB improved error performance for the same transmission rate. The same system shows only minor power penalty (1 dB) relative to its 4QAM OFDM bit rate equivalent but with the advantage of 30% bandwidth saving. This study reports1optical and electrical bandwidth saving with minor error performance degradation and paves the way for practical.
Izzat Darwazeh, Tongyang Xu, Tao Gui, Yuan Bao
ICC3