EDBT 2026 Demo / reviewers in the wild / expert
Xuanjing Huang 0001
dblp:05/6735-1
· DBLP profile ↗
327ranked-venue papers
2as first author
173since 2021 · last 2026
0000-0001-9197-9426ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 288 · 1 first-author · 155 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 17 since 2021Databases, data management, data science and information retrieval · 37 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainable Synthetic Image Detection Through Diffusion Timestep EnsemblingabstractRecent advances in diffusion models have enabled the creation of deceptively real images, posing significant security risks when misused. In this study, we empirically show that different timesteps of DDIM inversion reveal varying subtle distinctions between synthetic and real images that are extractable for detection, taking the forms of such as Fourier power spectrum high-frequency discrepancies and inter-pixel variance distributions. Based on these observations, we propose a novel detection method named ESIDE that directly utilizes features of intermediately noised images by training an ensemble on multiple noised timesteps, circumventing the overtime of conventional reconstruction-based strategies. To enhance human comprehension, we introduce a metric-grounded explanation refinement module to identify and explain AI-generated flaws. Additionally, we present the benchmarks GenHard and GenExplain, offering detection samples of greater difficulty and high-quality rationales for fake images. Extensive experiments show that ESIDE achieves state-of-the-art performance with 98.91% and 95.89% detection accuracy on regular and challenging samples respectively, and demonstrates generalizability and robustness. Yixin Wu 0005, Feiran Zhang, Tianyuan Shi, Ruicheng Yin, Zhenghua Wang, Zhenliang Gan, Changze Lv, Xiaoqing Zheng, Xuanjing Huang 0001 |
AAAI | 10 |
| 2026 | MetaAct-RL: Training Language Models for Reasoning Through Meta-Action-Based Reinforcement LearningabstractOutcome-based reinforcement learning has made notable advances in training language models (LMs) for reasoning. However, without explicit incentives and controls, this paradigm has limitations and instability in eliciting high-quality reasoning trajectories with diverse actions—particularly for models whose pretraining lacked extensive reasoning-related data. To this end, we introduce MetaAct-RL, a new RL framework that frames LMs’ thinking as sequential decision making over meta-actions. In this framework, the model chooses and executes a high-level action at each step—such as forward reasoning, critique, or refinement—to gradually reach the correct answer. To encourage deeper exploration, richer action diversity, and to improve sampling efficiency in the RL optimization process, MetaAct-RL incorporates appropriate length-based reward and regularization, and a key-state restart mechanism. Extensive experiments across six benchmarks show that MetaAct-RL improves reasoning performance by 7.99 on Llama3.2-1B and 7.17 on Llama3.1-8B relative to vanilla RL method. Moreover, on the challenging AIME-2024, our method outperforms the vanilla RL by 7.5 with Qwen2.5-1.5B. Zhiheng Xi, Yiwen Ding, Senjie Jin, Shichun Liu, Jixuan Huang, Dingwen Yang, Jiafu Tang, Boyang Hong, Junjie Ye 0005, Shihan Dou, Ming Zhang 0030, Jian Guan 0002, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
AAAI | 19 |
| 2026 | OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic CodingabstractDeming Ding, Shichun Liu, Enhui Yang, Jiahang Lin, Ziying Chen, Shihan Dou, Honglin Guo, Weiyu Cheng, Pengyu Zhao, Chengjun Xiao, Qunhong Zeng, Qi Zhang, Xuanjing Huang, Qidi Xu, Tao Gui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Deming Ding, Shichun Liu, Enhui Yang, Jiahang Lin, Ziying Chen, Shihan Dou, Honglin Guo, Weiyu Cheng, Chengjun Xiao, Qunhong Zeng, Qi Zhang 0001, Xuanjing Huang 0001, Qidi Xu, Tao Gui |
ACL (1) | 13 |
| 2026 | Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancingabstractXin Guo, Zhiheng Xi, Yiwen Ding, Yitao Zhai, Xiaowei Shi, Xunliang Cai, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhiheng Xi, Yiwen Ding, Yitao Zhai, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 9 |
| 2026 | Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-TrainingabstractChanghao Jiang, Ming Zhang, Yifei Cao, Junjie Ye, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Jiajun Sun, Yi Dong, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Changhao Jiang, Ming Zhang 0030, Yifei Cao, Junjie Ye 0005, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 15 |
| 2026 | DARM: Distribution-Aware Reward Modeling by Alleviating Biases from Low Preference-Context Dependency DataabstractShaofan Liu, Guoqiang Zhang, Shihan Dou, Huiyuan Zheng, Yiming Zhou, Junjie Ye, Shaowen Wang, Shichun Liu, Jiazheng Zhang, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaofan Liu, Shihan Dou, Huiyuan Zheng, Junjie Ye 0005, Shichun Liu, Jiazheng Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 12 |
| 2026 | Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward ModelsabstractBinghai Wang, Yantao Liu, Yuxuan Liu, Tianyi Tang, Shenzhi Wang, Chang Gao, Chujie Zheng, Yichang Zhang, Le Yu, Shixuan Liu, Tao Gui, Qi Zhang, Xuanjing Huang, Bowen Yu, Fei Huang, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Binghai Wang, Yantao Liu, Shenzhi Wang, Chujie Zheng, Yichang Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Bowen Yu 0002, Fei Huang 0002, Junyang Lin |
ACL (1) | 13 |
| 2026 | AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World EnvironmentsabstractZhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhiheng Xi, Dingwen Yang, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang 0001, Zhonghang Lu, Jiazheng Zhang, Dingwei Zhu, Junzhe Wang 0001, Zhihao Zhang 0002, Yuming Yang 0001, Junjie Ye 0005, Minghe Gao, Dongrui Liu, Jiaming Ji, Tao Gui, Xuanjing Huang 0001 |
ACL (1) | 24 |
| 2026 | Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative AlignmentabstractYuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuming Yang 0001, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao 0019, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 16 |
| 2026 | LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language ModelsabstractMing Zhang, Yujiong Shen, Jingyi Deng, Yuhui Wang, Huayu Sha, Kexin Tan, Qiyuan Peng, Yue Zhang, Junzhe Wang, Shichun Liu, Yueyuan Huang, Jingqi Tong, Changhao Jiang, Yilong Wu, Zhihao Zhang, Mingqi Wu, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ming Zhang 0030, Yujiong Shen, Jingyi Deng, Huayu Sha, Kexin Tan, Qiyuan Peng, Yue Zhang 0004, Junzhe Wang 0001, Shichun Liu, Yueyuan Huang, Jingqi Tong, Changhao Jiang, Yilong Wu, Zhihao Zhang 0002, Mingqi Wu, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 22 |
| 2026 | VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information BottleneckabstractFeiran Zhang, Yixin Wu, Zhenghua Wang, Xiaohua Wang, Changze Lv, Xuanjing Huang, Xiaoqing Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Feiran Zhang, Yixin Wu 0005, Zhenghua Wang, Changze Lv, Xuanjing Huang 0001, Xiaoqing Zheng |
ACL (1) | 6 |
| 2026 | VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-TrainingabstractDingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Guoqiang Zhang, Jiazheng Zhang, Junjie Ye, Mingxu Chai, Enyu Zhou, Ming Zhang, Yuhui Wang, Caishuang Huang, Chenhao Huang, Yunke Zhang, Yuran Wang, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dingwei Zhu, Shihan Dou, Zhiheng Xi, Senjie Jin, Jiazheng Zhang, Junjie Ye 0005, Mingxu Chai, Enyu Zhou, Ming Zhang 0030, Caishuang Huang, Chenhao Huang, Yunke Zhang, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 19 |
| 2026 | AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and ProgressabstractDespite rapid development, large language models (LLMs) still encounter challenges in multi-turn decision-making tasks (i.e., agent tasks) like web shopping and browser navigation, which require making a sequence of intelligent decisions based on environmental feedback. Previous work for LLM agents typically relies on elaborate prompt engineering or fine-tuning with expert trajectories to improve performance. In this work, we take a different perspective: we explore constructing process reward models (PRMs) to evaluate each decision and guide the agent's decision-making process. Unlike LLM reasoning, where each step is scored based on correctness, actions in agent tasks do not have a clear-cut correctness. Instead, they should be evaluated based on their proximity to the goal and the progress they have made. Building on this insight, we propose a re-defined PRM for agent tasks, named AgentPRM, to capture both the interdependence between sequential decisions and their contribution to the final goal. This enables better progress tracking and exploration-exploitation balance. To scalably obtain labeled data for training AgentPRM, we employ a Temporal Difference-based (TD-based) estimation method combined with Generalized Advantage Estimation (GAE), which proves more sample-efficient than prior methods. Extensive experiments across different agentic tasks show that AgentPRM is over 8× more compute-efficient than baselines, and it demonstrates robust improvement when scaling up test-time compute. Moreover, we perform detailed analyses to show how our method works and offer more insights, e.g., applying AgentPRM to the reinforcement learning of LLM agents. Zhiheng Xi, Chenyang Liao, Zhihao Zhang 0002, Wenxiang Chen, Binghai Wang, Senjie Jin, Yuhao Zhou 0005, Jian Guan 0002, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
WWW | 14 |
| 2026 | What is wrong with your code generated by large language models? An extensive study
Shihan Dou, Haoxiang Jia, Shenxi Wu, Huiyuan Zheng, Muling Wu, Yunbo Tao, Ming Zhang 0030, Mingxu Chai, Jessica Fan, Zhiheng Xi, Yueming Wu 0001, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001 |
Sci. China Inf. Sci. | 17 |
| 2026 | SpikeBERT: A language spikformer learned from BERT with knowledge distillation
Changze Lv, Tianlong Li, Weiming Qiao, Muling Wu, Shihan Dou, Xiaoqing Zheng, Xuanjing Huang 0001 |
Neural Networks | 9 |
| 2025 | Alleviating Shifted Distribution in Human Preference Alignment through Meta-LearningabstractThe capability of the reward model (RM) is crucial for the success of Reinforcement Learning from Human Feedback (RLHF) in aligning with human preferences. However, as training progresses, the output space distribution of the policy model shifts. The RM, initially trained on responses sampled from the output distribution of the early policy model, gradually loses its ability to distinguish between responses from the newly shifted distribution. This issue is further compounded when the RM, trained on a specific data distribution, struggles to generalize to examples outside of that distribution. These two issues can be united as a challenge posed by the shifted distribution of the environment. To surmount this challenge, we introduce MetaRM, a novel method leveraging meta-learning to adapt the RM to the shifted environment distribution. MetaRM optimizes the RM in an alternating way, by preserving both the preferences of the original preference pairs, as well as maximizing discrimination power over new examples of the shifted distribution. Extensive experiments demonstrate that MetaRM can iteratively enhance the performance of human preference alignment by improving the RM's capacity to identify subtle differences in samples of shifted distributions. Shihan Dou, Yan Liu 0002, Enyu Zhou, Songyang Gao, Tianlong Li, Limao Xiong, Haoxiang Jia, Junjie Ye 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
AAAI | 13 |
| 2025 | COSEE: Consistency-Oriented Signal-Based Early Exiting via Calibrated Sample Weighting MechanismabstractEarly exiting is an effective paradigm for improving the inference efficiency of pre-trained language models (PLMs) by dynamically adjusting the number of executed layers for each sample. However, in most existing works, easy and hard samples are treated equally by each classifier during training, which neglects the test-time early exiting behavior, leading to inconsistency between training and testing. Although some methods have tackled this issue under a fixed speed-up ratio, the challenge of flexibly adjusting the speed-up ratio while maintaining consistency between training and testing is still under-explored. To bridge the gap, we propose a novel Consistency-Oriented Signal-based Early Exiting (COSEE) framework, which leverages a calibrated sample weighting mechanism to enable each classifier to emphasize the samples that are more likely to exit at that classifier under various acceleration scenarios. Extensive experiments on the GLUE benchmark demonstrate the effectiveness of our COSEE across multiple exiting signals and backbones, yielding a better trade-off between performance and efficiency. Jianing He, Qi Zhang 0020, Hongyun Zhang 0001, Xuanjing Huang 0001, Usman Naseem, Duoqian Miao 0001 |
AAAI | 4 |
| 2025 | Synergistic Multi-Agent Framework with Trajectory Learning for Knowledge-Intensive TasksabstractRecent advancements in Large Language Models (LLMs) have led to significant breakthroughs in various natural language processing tasks. However, generating factually consistent responses in knowledge-intensive scenarios remains a challenge due to issues such as hallucination, difficulty in acquiring long-tailed knowledge, and limited memory expansion. This paper introduces SMART, a novel multi-agent framework that leverages external knowledge to enhance the interpretability and factual consistency of LLM-generated responses. SMART comprises four specialized agents, each performing a specific sub-trajectory action to navigate complex knowledge-intensive tasks. We propose a multi-agent co-training paradigm, Long-Short Trajectory Learning, which ensures synergistic collaboration among agents while maintaining fine-grained execution by each agent. Extensive experiments on five knowledge-intensive tasks demonstrate SMART's superior performance compared to widely adopted knowledge internalization and knowledge enhancement methods. Our framework can extend beyond knowledge-intensive tasks to more complex scenarios. Shengbin Yue, Siyuan Wang 0025, Wei Chen 0088, Xuanjing Huang 0001, Zhongyu Wei |
AAAI | 4 |
| 2025 | EvoWiki: Evaluating LLMs on Evolving KnowledgeabstractWei Tang, Yixin Cao, Yang Deng, Jiahao Ying, Bo Wang, Yizhe Yang, Yuyue Zhao, Qi Zhang, Xuanjing Huang, Yu-Gang Jiang, Yong Liao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wei Tang 0015, Yixin Cao 0002, Yang Deng 0002, Jiahao Ying, Yizhe Yang, Yuyue Zhao, Qi Zhang 0001, Xuanjing Huang 0001, Yu-Gang Jiang 0001, Yong Liao 0003 |
ACL (1) | 9 |
| 2025 | Lost in the Context: Insufficient and Distracted Attention to Contexts in Preference ModelingabstractShihan Dou, Jiayi Chen, Chenhao Huang, Feng Chen, Wei Chengzhi, Huiyuan Zheng, Shichun Liu, Yan Liu, Chenxiao Liu, Chao Xin, Lin Yan, Zongzhang Zhang, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shihan Dou, Chenhao Huang, Feng Chen 0042, Wei Chengzhi, Huiyuan Zheng, Shichun Liu, Yan Liu 0002, Chenxiao Liu, Chao Xin, Zongzhang Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 15 |
| 2025 | VLSBench: Unveiling Visual Leakage in Multimodal SafetyabstractSafety concerns of Multimodal large language models (MLLMs) have gradually become an important problem in various applications.Surprisingly, previous works indicate a counterintuitive phenomenon that using textual unlearning to align MLLMs achieves comparable safety performances with MLLMs aligned with image-text pairs.To explain such a phenomenon, we discover a Visual Safety Information Leakage (VSIL) problem in existing multimodal safety benchmarks, i.e., the potentially risky content in the image has been revealed in the textual query.Thus, MLLMs can easily refuse these sensitive image-text pairs according to textual queries only, leading to unreliable cross-modality safety evaluation of MLLMs.To this end, we construct multimodal Visual Leakless Safety Bench (VLS-Bench) with 2.2k image-text pairs through an automated data pipeline.Experimental results indicate that VLSBench poses a significant challenge to both open-source and closesource MLLMs, e.g., LLaVA, Qwen2-VL and GPT-4o.Besides, we empirically compare textual and multimodal alignment methods on VLSBench and find that textual alignment is effective enough for multimodal safety scenarios with VSIL, while multimodal alignment is preferable for safety scenarios without VSIL.Code and data are released under https://github.com/AI45Lab/VLSBench. Xuhao Hu, Dongrui Liu, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2025 | HAF-RM: A Hybrid Alignment Framework for Reward Model TrainingabstractShujun Liu, Xiaoyu Shen, Yuhang Lai, Siyuan Wang, Shengbin Yue, Zengfeng Huang, Xuanjing Huang, Zhongyu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shujun Liu, Yuhang Lai, Siyuan Wang 0025, Shengbin Yue, Zengfeng Huang, Xuanjing Huang 0001, Zhongyu Wei |
ACL (1) | 7 |
| 2025 | Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and InferenceabstractSiyuan Wang, Dianyi Wang, Chengxing Zhou, Zejun Li, Zhihao Fan, Xuanjing Huang, Zhongyu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Siyuan Wang 0025, Dianyi Wang, Chengxing Zhou, Zhihao Fan, Xuanjing Huang 0001, Zhongyu Wei |
ACL (1) | 6 |
| 2025 | AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse EnvironmentsabstractZhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Xin Guo, Dingwen Yang, Chenyang Liao, Wei He, Songyang Gao, Lu Chen, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, Yu-Gang Jiang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang 0001, Dingwen Yang, Chenyang Liao, Wei He 0024, Songyang Gao, Lu Chen 0001, Yicheng Zou, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Zuxuan Wu, Yu-Gang Jiang 0001 |
ACL (1) | 18 |
| 2025 | Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable MetricabstractData diversity is crucial for the instruction tuning of large language models. Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. However, the fundamental problem of precisely defining and measuring data diversity remains underexplored, limiting clear guidance for data engineering. To address this, we systematically analyze 11 existing diversity measurement methods by evaluating their correlation with model performance through extensive fine-tuning experiments. Our results indicate that a reliable diversity measure should properly account for both inter-sample differences and the information density in the sample space. Building on this, we propose NovelSum, a new diversity metric based on sample-level “novelty.” Experiments on both simulated and real-world data show that NovelSum accurately captures diversity variations and achieves a 0.97 correlation with instruction-tuned model performance, highlighting its value in guiding data engineering practices. With NovelSum as an optimization objective, we further develop a greedy, diversity-oriented data selection strategy that outperforms existing approaches, validating both the effectiveness and practical significance of our metric. Yuming Yang 0001, Junjie Ye 0005, Shihan Dou, Xiao Wang 0042, Huijie Lv, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 10 |
| 2025 | ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool UseabstractJunjie Ye, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zehui Chen, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Siyu Yuan, Tao Gui, Qi Zhang, Xuanjing Huang, Jiecao Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Junjie Ye 0005, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Jiecao Chen |
ACL (1) | 13 |
| 2025 | Dynamic and Generalizable Process Reward ModelingabstractZhangyue Yin, Qiushi Sun, Zhiyuan Zeng, Qinyuan Cheng, Xipeng Qiu, Xuanjing Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhangyue Yin, Qiushi Sun, Zhiyuan Zeng 0004, Qinyuan Cheng, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 6 |
| 2025 | Prior-Fitted Networks Scale to Larger Datasets When Treated as Weak LearnersabstractPrior-Fitted Networks (PFNs) have recently been proposed to efficiently perform tabular classification tasks. Although they achieve good performance on small datasets, they encounter limitations with larger datasets. These limitations include significant memory consumption and increased computational complexity, primarily due to the impracticality of incorporating all training samples as inputs within these networks. To address these challenges, we investigate the fitting assumption for PFNs and input samples. Building on this understanding, we propose \emph{BoostPFN} designed to enhance the performance of these networks, especially for large-scale datasets. We also theoretically validate the convergence of BoostPFN and our empirical results demonstrate that the BoostPFN method can outperform standard PFNs with the same size of training samples in large datasets and achieve a significant acceleration in training times compared to other established baselines in the field, including widely-used Gradient Boosting Decision Trees (GBDTs), deep learning methods and AutoML systems. High performance is maintained for up to 50x of the pre-training size of PFNs, substantially extending the limit of training samples. Through this work, we address the challenges of efficiently handling large datasets via PFN-based models, paving the way for faster and more effective tabular data classification training and prediction process. Yuxin Wang 0005, Botian Jiang, David P. Wipf, Xuanjing Huang 0001, Xipeng Qiu |
AISTATS | 6 |
| 2025 | SpikeBERT: A Language Understanding Spiking Neural Network Learned from BERT with Knowledge Distillation
Changze Lv, Tianlong Li, Muling Wu, Shihan Dou, Xiaoqing Zheng, Xuanjing Huang 0001 |
CogSci | 8 |
| 2025 | Revisiting Jailbreaking for Large Language Models: A Representation Engineering PerspectiveabstractThe recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) when exposed to malicious inputs. While various defense strategies have been proposed to mitigate these threats, there has been limited research into the underlying mechanisms that make LLMs vulnerable to such attacks. In this study, we suggest that the self-safeguarding capability of LLMs is linked to specific activity patterns within their representation space. Although these patterns have little impact on the semantic content of the generated text, they play a crucial role in shaping LLM behavior under jailbreaking attacks. Our findings demonstrate that these patterns can be detected with just a few pairs of contrastive queries. Extensive experimentation shows that the robustness of LLMs against jailbreaking can be manipulated by weakening or strengthening these patterns. Further visual analysis provides additional evidence for our conclusions, providing new insights into the jailbreaking phenomenon. These findings highlight the importance of addressing the potential misuse of open-source LLMs within the community. Tianlong Li, Zhenghua Wang, Muling Wu, Shihan Dou, Changze Lv, Xiaoqing Zheng, Xuanjing Huang 0001 |
COLING | 9 |
| 2025 | Case2Code: Scalable Synthetic Data for Code GenerationabstractLarge Language Models (LLMs) have shown outstanding breakthroughs in code generation. Recent work improves code LLMs by training on synthetic data generated by some powerful LLMs, which can be challenging to scale due to the dependence on a teacher model and high generation costs. In this paper, we focus on synthesizing code data at scale and propose a Case2Code task by exploiting the expressiveness and correctness of programs. Case2Code is an inductive inference task that aims to infer underlying code implementations by observing input-output examples or program behaviors, By incorporating LLMs to generate program inputs, and executing the program with these inputs to obtain the program outputs, we can synthesize diverse and high-quality Case2Code data at scale for training and evaluating code LLMs. Experimental results show that case-to-code induction is challenging for current representative LLMs if they are untrained. Models trained with Case2Code improve performance not only on distribution case-to-code induction but also various coding-generation tasks, demonstrating the great potential of large-scale synthetic data and inductive learning. Yunfan Shao, Linyang Li, Yichuan Ma, Peiji Li, Demin Song, Qinyuan Cheng, Pengyu Wang 0006, Qipeng Guo, Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001, Dahua Lin |
COLING | 13 |
| 2025 | Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM EvaluationabstractThis paper presents a benchmark self-evolving framework to dynamically evaluate rapidly advancing Large Language Models (LLMs). We utilize a multi-agent system to reframe new evolving instances with high confidence that extend existing benchmarks. Towards a more scalable, robust and fine-grained evaluation, we implement six reframing operations to construct evolving instances testing LLMs against diverse queries, shortcut biases and probing their problem-solving sub-abilities. With this framework, we extend datasets across general and specific tasks, through various iterations. Experimental results show a performance decline in most LLMs against their original results under scalable and robust evaluations, offering a more accurate reflection of model capabilities alongside our fine-grained evaluation. Besides, our framework widens performance discrepancies both between different models and within the same model across various tasks, facilitating more informed model selection for specific tasks. We hope this framework contributes the research community for continuously evolving benchmarks alongside LLM development. Siyuan Wang 0025, Zhuohan Long, Zhihao Fan, Xuanjing Huang 0001, Zhongyu Wei |
COLING | 4 |
| 2025 | Beyond Boundaries: Learning a Universal Entity Taxonomy across Datasets and Languages for Open Named Entity RecognitionabstractOpen Named Entity Recognition (NER), which involves identifying arbitrary types of entities from arbitrary domains, remains challenging for Large Language Models (LLMs). Recent studies suggest that fine-tuning LLMs on extensive NER data can boost their performance. However, training directly on existing datasets neglects their inconsistent entity definitions and redundant data, limiting LLMs to dataset-specific learning and hindering out-of-domain adaptation. To address this, we present B2NERD, a compact dataset designed to guide LLMs’ generalization in Open NER under a universal entity taxonomy. B2NERD is refined from 54 existing English and Chinese datasets using a two-step process. First, we detect inconsistent entity definitions across datasets and clarify them by distinguishable label names to construct a universal taxonomy of 400+ entity types. Second, we address redundancy using a data pruning strategy that selects fewer samples with greater category and semantic diversity. Comprehensive evaluation shows that B2NERD significantly enhances LLMs’ Open NER capabilities. Our B2NER models, trained on B2NERD, outperform GPT-4 by 6.8-12.0 F1 points and surpass previous methods in 3 out-of-domain benchmarks across 15 datasets and 6 languages. The data, models, and code are publicly available at https://github.com/UmeanNever/B2NER. Yuming Yang 0001, Wantong Zhao, Caishuang Huang, Junjie Ye 0005, Xiao Wang 0042, Huiyuan Zheng, Xueying Xu, Kaixin Huang, Yunke Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 14 |
| 2025 | ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world ScenariosabstractExisting evaluations of tool learning primarily focus on validating the alignment of selected tools for large language models (LLMs) with expected outcomes. However, these approaches rely on a limited set of scenarios where answers can be pre-determined. Furthermore, a sole emphasis on outcomes disregards the complex capabilities required for LLMs to effectively use tools. To tackle this issue, we propose ToolEyes, a fine-grained system tailored for the evaluation of the LLMs’ tool learning capabilities in authentic scenarios. The system meticulously examines seven real-world scenarios, analyzing five dimensions crucial to LLMs in tool learning: format alignment, intent comprehension, behavior planning, tool selection, and answer organization. Additionally, ToolEyes incorporates a tool library boasting approximately 600 tools, serving as an intermediary between LLMs and the physical world. Evaluations involving ten LLMs across three categories reveal a preference for specific scenarios and limited cognitive abilities in tool learning. Intriguingly, expanding the model size even exacerbates the hindrance to tool learning. The code and data are available at https://github.com/Junjie-Ye/ToolEyes. Junjie Ye 0005, Songyang Gao, Caishuang Huang, Yilong Wu, Sixian Li, Xiaoran Fan, Shihan Dou, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
COLING | 12 |
| 2025 | SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language ModelsabstractThe emergence of Vision Language Models (VLMs) has brought unprecedented advances in understanding multi-modal information. The combination of textual and visual semantics in VLMs is highly complex and diverse, making the safety alignment of these models challenging. Furthermore, due to the limited study on the safety alignment of VLMs, there is a lack of large-scale, high-quality datasets. To address these limitations, we propose a Safety Preference Alignment dataset for Vision Language Models named SPA-VL. In terms of breadth, SPA-VL covers 6 harmfulness domains, 13 categories, and 53 subcategories, and contains 100,788 samples of the quadruple (question, image, chosen response, rejected response). In terms of depth, the responses are collected from 12 open-source (e.g., QwenVL) and closed-source (e.g., Gemini) VLMs to ensure diversity. The construction of preference data is fully automated, and the experimental results indicate that models trained with alignment techniques on the SPA-VL dataset exhibit substantial improvements in harmlessness and helpfulness while maintaining core capabilities. SPA-VL, as a large-scale, high-quality, and diverse dataset, represents a significant milestone in ensuring that VLMs achieve both harmlessness and helpfulness. Yongting Zhang, Lu Chen 0001, Guodong Zheng, Yifeng Gao 0002, Jinlan Fu, Zhenfei Yin, Senjie Jin, Yu Qiao 0001, Xuanjing Huang 0001, Feng Zhao 0004, Tao Gui |
CVPR | 10 |
| 2025 | Governance in Motion: Co-evolution of Constitutions and AI models for Scalable SafetyabstractChenhao Huang, Ziyu Shen, Yicong Ren, Huiyuan Zheng, Jiazheng Zhang, Mingxu Chai, Ming Zhang, Shihan Dou, Fan Mo, Jie Shi, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chenhao Huang, Ziyu Shen, Yicong Ren, Huiyuan Zheng, Jiazheng Zhang, Mingxu Chai, Ming Zhang 0030, Shihan Dou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 13 |
| 2025 | Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for ReasoningabstractSenjie Jin, Lu Chen, Zhiheng Xi, Yuhui Wang, Sirui Song, Yuhao Zhou, Xinbo Zhang, Peng Sun, Hong Lu, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Senjie Jin, Lu Chen 0001, Zhiheng Xi, Sirui Song, Yuhao Zhou 0005, Xinbo Zhang, Peng Sun 0006, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 12 |
| 2025 | LoRACoE: Improving Large Language Model via Composition-based LoRA ExpertabstractThe Mixture of Experts (MoE) architecture improves large language models (LLMs) by utilizing sparsely activated expert sub-networks with a routing module, yet it typically demands high training cost.Previous work introduces parameter-efficient fine-tuning (PEFT) modules, e.g., LoRA, to achieve a lightweight MoE for efficiency.However, they construct static experts by manually splitting the LoRA parameters into fixed groups, which limits flexibility and dynamism.Furthermore, this manual partitioning also hinders the effective utilization of well-initialized LoRA modules.To tackl the challenges, we first delve into the parameter patterns in LoRA modules, revealing that there exists task-relevant parameters that are concentrated along the rank dimension.Based on this, we redesign the construction of experts and propose the LoRACoE (LoRA Composition of Experts) method.Specifically, when confronted with a task, it dynamically builds experts based on rank-level parameter composition, i.e., experts can flexibly combine rank-level parameters in LoRA module.Extensive experiments demonstrate that compared to other LoRA-based MoE methods, our method achieves better task performance across a broader range of tasks. Zhiheng Xi, Zhihao Zhang 0002, Boyang Hong, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 7 |
| 2025 | SATER: A Self-Aware and Token-Efficient Approach to Routing and CascadingabstractLarge language models (LLMs) demonstrate remarkable performance across diverse tasks, yet their effectiveness frequently depends on costly commercial APIs or cloud services.Model selection thus entails a critical trade-off between performance and cost: high-performing LLMs typically incur substantial expenses, whereas budget-friendly small language models (SLMs) are constrained by limited capabilities.Current research primarily proposes two routing strategies: pre-generation routing and cascade routing.Both approaches have distinct characteristics, with cascade routing typically offering superior cost-effectiveness and accuracy despite its higher latency.To further address the limitations of both approaches, we introduce SATER, a dual-mode compatible approach that finetunes models through shortest-response preference optimization and a confidence-aware rejection mechanism.SATER significantly reduces redundant outputs and response times, while improving both the performance of pregeneration routing and the efficiency of cascade routing.Experiments across three SLMs and six datasets, varying in type and complexity, demonstrate that SATER achieves comparable performance while consistently reducing computational costs by over 50% and cascade latency by over 80%. Yuanzhe Shen, Yide Liu, Zisu Huang, Ruicheng Yin, Xiaoqing Zheng, Xuanjing Huang 0001 |
EMNLP | 6 |
| 2025 | Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter LevelsabstractJunjie Ye, Yuming Yang, Yang Nan, Shuo Li, Qi Zhang, Tao Gui, Xuanjing Huang, Peng Wang, Zhongchao Shi, Jianping Fan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Junjie Ye 0005, Yuming Yang 0001, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001, Peng Wang 0095, Zhongchao Shi, Jianping Fan 0007 |
EMNLP | 7 |
| 2025 | ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
Yuqian Fu, Bin Ren 0005, Guolei Sun, Biao Gong, Yanwei Fu 0001, Danda Pani Paudel, Xuanjing Huang 0001, Luc Van Gool |
ICCV | 8 |
| 2025 | Have the VLMs Lost Confidence? A Study of Sycophancy in VLMsabstractIn the study of LLMs, sycophancy represents a prevalent hallucination that poses significant challenges to these models. Specifically, LLMs often fail to adhere to original correct responses, instead blindly agreeing with users' opinions, even when those opinions are incorrect or malicious. However, research on sycophancy in visual language models (VLMs) has been scarce. In this work, we extend the exploration of sycophancy from LLMs to VLMs, introducing the MM-SY benchmark to evaluate this phenomenon. We present evaluation results from multiple representative models, addressing the gap in sycophancy research for VLMs. To mitigate sycophancy, we propose a synthetic dataset for training and employ methods based on prompts, supervised fine-tuning, and DPO. Our experiments demonstrate that these methods effectively alleviate sycophancy in VLMs. Additionally, we probe VLMs to assess the semantic impact of sycophancy and analyze the attention distribution of visual tokens. Our findings indicate that the ability to prevent sycophancy is predominantly observed in higher layers of the model. The lack of attention to image knowledge in these higher layers may contribute to sycophancy, and enhancing image attention at high layers proves beneficial in mitigating this issue. Xiaoran Fan, Linsheng Lu, Leyi Yang, Yuming Yang 0001, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICLR | 13 |
| 2025 | RMB: Comprehensively benchmarking reward models in LLM alignmentabstractReward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. However, the current evaluation of RMs may not directly correspond to their alignment performance due to the limited distribution of evaluation data and evaluation methods that are not closely related to alignment objectives. To address these limitations, we propose RMB, a comprehensive RM benchmark that covers over 49 real-world scenarios and includes both pairwise and Best-of-N (BoN) evaluations to better reflect the effectiveness of RMs in guiding alignment optimization.
We demonstrate a positive correlation between our benchmark and the downstream alignment task performance. Based on our benchmark, we conduct extensive analysis on the state-of-the-art RMs, revealing their generalization defects that were not discovered by previous benchmarks, and highlighting the potential of generative RMs. Furthermore, we delve into open questions in reward models, specifically examining the effectiveness of majority voting for the evaluation of reward models and analyzing the impact factors of generative RMs, including the influence of evaluation criteria and instructing methods. We will release our evaluation code and datasets upon publication. Enyu Zhou, Guodong Zheng, Binghai Wang, Zhiheng Xi, Shihan Dou, Rong Bao, Limao Xiong, Jessica Fan, Yurong Mou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICLR | 14 |
| 2025 | Dendritic Localized Learning: Toward Biologically Plausible AlgorithmabstractBackpropagation is the foundational algorithm for training neural networks and a key driver of deep learning’s success. However, its biological plausibility has been challenged due to three primary limitations: weight symmetry, reliance on global error signals, and the dual-phase nature of training, as highlighted by the existing literature. Although various alternative learning approaches have been proposed to address these issues, most either fail to satisfy all three criteria simultaneously or yield suboptimal results. Inspired by the dynamics and plasticity of pyramidal neurons, we propose Dendritic Localized Learning (DLL), a novel learning algorithm designed to overcome these challenges. Extensive empirical experiments demonstrate that DLL satisfies all three criteria of biological plausibility while achieving state-of-the-art performance among algorithms that meet these requirements. Furthermore, DLL exhibits strong generalization across a range of architectures, including MLPs, CNNs, and RNNs. These results, benchmarked against existing biologically plausible learning algorithms, offer valuable empirical insights for future research. We hope this study can inspire the development of new biologically plausible algorithms for training multilayer networks and advancing progress in both neuroscience and machine learning. Our code is available at https://github.com/Lvchangze/Dendritic-Localized-Learning. Changze Lv, Zhenghua Wang, Zhibo Xu, Di Yu 0001, Xin Du 0002, Xiaoqing Zheng, Xuanjing Huang 0001 |
ICML | 10 |
| 2025 | Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided SamplingabstractYiwen Ding, Zhiheng Xi, Wei He, Lizhuoyuan Lizhuoyuan, Yitao Zhai, Shi Xiaowei, Xunliang Cai, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yiwen Ding, Zhiheng Xi, Wei He 0024, Lizhuoyuan Lizhuoyuan, Yitao Zhai, Shi Xiaowei, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
NAACL (Long Papers) | 10 |
| 2025 | VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal ModelsabstractZejun Li, Ruipu Luo, Jiwen Zhang, Minghui Qiu, Xuanjing Huang, Zhongyu Wei. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ruipu Luo, Jiwen Zhang, Minghui Qiu, Xuanjing Huang 0001, Zhongyu Wei |
NAACL (Long Papers) | 5 |
| 2025 | AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive ScenariosabstractXinyi Mou, Jingcong Liang, Jiayu Lin, Xinnong Zhang, Xiawei Liu, Shiyue Yang, Rong Ye, Lei Chen, Haoyu Kuang, Xuanjing Huang, Zhongyu Wei. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xinyi Mou, Jingcong Liang, Xinnong Zhang, Xiawei Liu, Shiyue Yang, Rong Ye, Lei Chen 0082, Haoyu Kuang, Xuanjing Huang 0001, Zhongyu Wei |
NAACL (Long Papers) | 10 |
| 2025 | Pre-Trained Policy Discriminators are General Reward ModelsabstractWe offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named POLicy DiscriminAtive LeaRning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance.
For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines.
POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance—improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks.
Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99.
The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models. Shihan Dou, Shichun Liu, Yuming Yang 0001, Yicheng Zou, Yunhua Zhou, Shuhao Xing, Chenhao Huang, Qiming Ge, Haijun Lv, Demin Song, Songyang Gao, Chengqi Lyu, Enyu Zhou, Honglin Guo, Zhiheng Xi, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Kai Chen 0026 |
NeurIPS | 21 |
| 2025 | EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem SolvingabstractWe introduce EvaLearn, a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks, a critical, yet underexplored aspect of model potential. EvaLearn contains 648 challenging problems across six task types, grouped into 182 sequences, each sequence dedicated to one task type. Diverging from most existing benchmarks that evaluate models in parallel, EvaLearn requires models to solve problems sequentially, allowing them to leverage the experience gained from previous solutions. EvaLearn provides five comprehensive automated metrics to evaluate models and quantify their learning capability and efficiency. We extensively benchmark nine frontier models and observe varied performance profiles: some models, such as Claude-3.7-sonnet, start with moderate initial performance but exhibit strong learning ability, while some models struggle to benefit from experience and may even show negative transfer. Moreover, we investigate model performance under two learning settings and find that instance-level rubrics and teacher-model feedback further facilitate model learning. Importantly, we observe that current LLMs with stronger static abilities do not show a clear advantage in learning capability across all tasks, highlighting that EvaLearn evaluates a new dimension of model performance. We hope EvaLearn provides a novel evaluation perspective for assessing LLM potential and understanding the gap between models and human capabilities, promoting the development of deeper and more dynamic evaluation approaches. All datasets, the automatic evaluation framework, and the results studied in this paper are available in the supplementary materials. Shihan Dou, Ming Zhang 0030, Chenhao Huang, Feng Chen 0042, Shichun Liu, Yan Liu 0002, Chenxiao Liu, Zongzhang Zhang, Tao Gui, Chao Xin, Wei Chengzhi, Qi Zhang 0001, Xuanjing Huang 0001 |
NeurIPS | 16 |
| 2025 | Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object DetectionabstractCross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel objects with only a handful of labeled samples from previously unseen domains. While data augmentation and generative methods have shown promise in few-shot learning, their effectiveness for CD-FSOD remains unclear due to the need for both visual realism and domain alignment. Existing strategies, such as copy-paste augmentation and text-to-image generation, often fail to preserve the correct object category or produce backgrounds coherent with the target domain, making them non-trivial to apply directly to CD-FSOD. To address these challenges, we propose Domain-RAG, a training-free, retrieval-guided compositional image generation framework tailored for CD-FSOD. Domain-RAG consists of three stages: domain-aware background retrieval, domain-guided background generation, and foreground-background composition. Specifically, the input image is first decomposed into foreground and background regions. We then retrieve semantically and stylistically similar images to guide a generative model in synthesizing a new background, conditioned on both the original and retrieved contexts. Finally, the preserved foreground is composed with the newly generated domain-aligned background to form the generated image. Without requiring any additional supervision or training, Domain-RAG produces high-quality, domain-consistent samples across diverse tasks, including CD-FSOD, remote sensing FSOD, and camouflaged FSOD. Extensive experiments show consistent improvements over strong baselines and establish new state-of-the-art results. Codes will be released upon acceptance.The source code and instructions are available at https://github.com/LiYu0524/Domain-RAG. Yu Li 0007, Xingyu Qiu, Yuqian Fu, Tianwen Qian, Xu Zheng 0002, Danda Pani Paudel, Yanwei Fu 0001, Xuanjing Huang 0001, Luc Van Gool, Yu-Gang Jiang 0001 |
NeurIPS | 9 |
| 2025 | Toward Relative Positional Encoding in Spiking TransformersabstractSpiking neural networks (SNNs) are bio-inspired networks that mimic how neurons in the brain communicate through discrete spikes, which have great potential in various tasks due to their energy efficiency and temporal processing capabilities.
SNNs with self-attention mechanisms (spiking Transformers) have recently shown great advancements in various tasks, and inspired by traditional Transformers, several studies have demonstrated that spiking absolute positional encoding can help capture sequential relationships for input data, enhancing the capabilities of spiking Transformers for tasks such as sequential modeling and image classification. However, how to incorporate relative positional information into SNNs remains a challenge.
In this paper, we introduce several strategies to approximate relative positional encoding (RPE) in spiking Transformers while preserving the binary nature of spikes.
Firstly, we formally prove that encoding relative distances with Gray Code ensures that the binary representations of positional indices maintain a constant Hamming distance whenever their decimal values differ by a power of two, and we propose **Gray-PE** based on this property.
In addition, we propose another RPE method called **Log-PE**, which combines the logarithmic form of the relative distance matrix directly into the spiking attention map.
Furthermore, we extend our RPE methods to a two-dimensional form, making them suitable for processing image patches.
We evaluate our RPE methods on various tasks, including time series forecasting, text classification, and patch-based image classification, and the experimental results demonstrate a satisfying performance gain by incorporating our RPE methods across many architectures.
Our results provide fresh perspectives on designing spiking Transformers to advance their sequential modeling capability, thereby expanding their applicability across various domains.
Our code is available at https://github.com/microsoft/SeqSNN. Changze Lv, Yansen Wang, Yifei Shen 0004, Xiaoqing Zheng, Xuanjing Huang 0001, Dongsheng Li 0002 |
NeurIPS | 6 |
| 2025 | BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning DatasetabstractIn this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 100k university-level questions drawn from 300 UNESCO-defined subjects, spanning diverse formats—multiple-choice, fill-in-the-blank, and open-ended QA—and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop, automated, and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20k high-quality instances to comprehensively assess LMMs’ knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 80k instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline BMMR-Verifier for accurate and fine-grained evaluation of LMMs’ reasoning. Extensive experiments reveal that (i) even SOTA models leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data and models, and we believe our work can offers valuable insights and contributions to the community. Zhiheng Xi, Yutao Fan, Honglin Guo, Yufang Liu, Xiaoran Fan, Jingchao Ding, Wangmeng Zuo, Zhenfei Yin, Lei Bai 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
NeurIPS | 15 |
| 2025 | Understanding Parametric and Contextual Knowledge Reconciliation within Large Language ModelsabstractRetrieval-Augmented Generation (RAG) provides additional contextual knowledge to complement the parametric knowledge in Large Language Models (LLMs). These two knowledge interweave to enhance the accuracy and timeliness of LLM responses. However,
the internal mechanisms by which LLMs utilize these knowledge remain unclear. We propose modeling the forward propagation of knowledge as an entity flow, employing this framework to trace LLMs' internal behaviors when processing mixed-source knowledge. Linear probing utilizes a trainable linear classifier to detect specific attributes in hidden layers. However, once trained, a probe cannot adapt to dynamically specified entities. To address this challenge, we construct an entity-aware probe, which introduces special tokens to mark probing targets and employs a small trainable rank-8 lora update to process these special markers. We first verify this approach through an attribution experiment, demonstrating that it can accurately detect information about ad-hoc entities from complex hidden states. Next, we trace entity flows across layers to understand how LLMs reconcile conflicting knowledge internally. Our probing results reveal that contextual and parametric knowledge are routed between tokens through distinct sets of attention heads, supporting attention competition only within knowledge types. While conflicting knowledge maintains a residual presence across layers, aligned knowledge from multiple sources gradually accumulates, with the magnitude of this accumulation directly determining its influence on final outputs. Jun Zhao 0019, Yongzhuo Yang, Jingqi Tong, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
NeurIPS | 9 |
| 2025 | The dual-edged sword: artificial intelligence's evolving role in academic peer review
Xuanjing Huang 0001, Shihan Dou, Zhangyue Yin |
Sci. China Inf. Sci. | 1 |
| 2025 | The rise and potential of large language model based agents: a survey
Zhiheng Xi, Wenxiang Chen, Wei He 0024, Yiwen Ding, Boyang Hong, Ming Zhang 0030, Junzhe Wang 0001, Senjie Jin, Enyu Zhou, Xiaoran Fan, Xiao Wang 0001, Limao Xiong, Yuhao Zhou 0005, Weiran Wang 0003, Changhao Jiang, Yicheng Zou, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang 0001, Qi Zhang 0001, Tao Gui |
Sci. China Inf. Sci. | 26 |
| 2025 | SpikeCLIP: A contrastive language-image pretrained spiking neural network
Changze Lv, Tianlong Li, Yufei Gu, Jianhan Xu, Cenyuan Zhang, Muling Wu, Xiaoqing Zheng, Xuanjing Huang 0001 |
Neural Networks | 9 |
| 2025 | Efficient Link Prediction via GNN Layers Induced by Negative SamplingabstractGraph neural networks (GNNs) for link prediction can loosely be divided into two broad categories. First,node-wisearchitectures pre-compute individual embeddings for each node that are later combined by a simple decoder to make predictions. While extremely efficient at inference time, model expressiveness is limited such that isomorphic nodes contributing to candidate edges may not be distinguishable, compromising accuracy. In contrast,edge-wisemethods rely on the formation of edge-specific subgraph embeddings to enrich the representation of pair-wise relationships, disambiguating isomorphic nodes to improve accuracy, but with increased model complexity. To better navigate this trade-off, we propose a novel GNN architecture whereby theforward passexplicitly depends onbothpositive (as is typical) and negative (unique to our approach) edges to inform more flexible, yet still cheap node-wise embeddings. This is achieved by recasting the embeddings themselves as minimizers of a forward-pass-specific energy function that favors separation of positive and negative samples. Notably, this energy is distinct from the actual training loss shared by most existing link prediction models, where contrastive pairs only influence thebackward pass. As demonstrated by extensive empirical evaluations, the resulting architecture retains the inference speed of node-wise models, while producing competitive accuracy with edge-wise alternatives. Yuxin Wang 0005, Xiannian Hu, Xuanjing Huang 0001, Xipeng Qiu, David P. Wipf |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | LLMEval: A Preliminary Study on How to Evaluate Large Language ModelsabstractRecently, the evaluation of Large Language Models has emerged as a popular area of research. The three crucial questions for LLM evaluation are ``what, where, and how to evaluate''. However, the existing research mainly focuses on the first two questions, which are basically what tasks to give the LLM during testing and what kind of knowledge it should deal with. As for the third question, which is about what standards to use, the types of evaluators, how to score, and how to rank, there hasn't been much discussion. In this paper, we analyze evaluation methods by comparing various criteria with both manual and automatic evaluation, utilizing onsite, crowd-sourcing, public annotators and GPT-4, with different scoring methods and ranking systems. We propose a new dataset, LLMEval and conduct evaluations on 20 LLMs. A total of 2,186 individuals participated, leading to the generation of 243,337 manual annotations and 57,511 automatic evaluation results. We perform comparisons and analyses of different settings and conduct 10 conclusions that can provide some insights for evaluating LLM in the future. The dataset and the results are publicly available at https://github.com/llmeval. The version with the appendix are publicly available at https://arxiv.org/abs/2312.07398. Yue Zhang 0004, Ming Zhang 0030, Haipeng Yuan, Shichun Liu, Yongyao Shi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
AAAI | 8 |
| 2024 | StepCoder: Improving Code Generation with Reinforcement Learning from Compiler FeedbackabstractShihan Dou, Yan Liu, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou, Tao Ji, Rui Zheng, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shihan Dou, Yan Liu 0002, Haoxiang Jia, Enyu Zhou, Limao Xiong, Junjie Shan, Caishuang Huang, Xiao Wang 0001, Xiaoran Fan, Zhiheng Xi, Yuhao Zhou 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
ACL (1) | 16 |
| 2024 | LoRAMoE: Alleviating World Knowledge Forgetting in Large Language Models via MoE-Style PluginabstractShihan Dou, Enyu Zhou, Yan Liu, Songyang Gao, Wei Shen, Limao Xiong, Yuhao Zhou, Xiao Wang, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Jiang Zhu, Rui Zheng, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shihan Dou, Enyu Zhou, Yan Liu 0002, Songyang Gao, Limao Xiong, Yuhao Zhou 0005, Xiao Wang 0001, Zhiheng Xi, Xiaoran Fan, Shiliang Pu, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 16 |
| 2024 | Aligning Large Language Models with Human Preferences through Representation EngineeringabstractWenhao Liu, Xiaohua Wang, Muling Wu, Tianlong Li, Changze Lv, Zixuan Ling, Zhu JianHao, Cenyuan Zhang, Xiaoqing Zheng, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Muling Wu, Tianlong Li, Changze Lv, Zixuan Ling, Jianhao Zhu, Cenyuan Zhang, Xiaoqing Zheng, Xuanjing Huang 0001 |
ACL (1) | 10 |
| 2024 | Navigating the OverKill in Large Language ModelsabstractChenyu Shi, Xiao Wang, Qiming Ge, Songyang Gao, Xianjun Yang, Tao Gui, Qi Zhang, Xuanjing Huang, Xun Zhao, Dahua Lin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chenyu Shi, Xiao Wang 0042, Qiming Ge, Songyang Gao, Xianjun Yang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Dahua Lin |
ACL (1) | 8 |
| 2024 | F-Eval: Asssessing Fundamental Abilities with Refined Evaluation MethodsabstractYu Sun, Keyu Chen, Shujie Wang, Peiji Li, Qipeng Guo, Hang Yan, Xipeng Qiu, Xuanjing Huang, Dahua Lin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yu Sun 0031, Keyuchen Keyuchen, Peiji Li, Qipeng Guo, Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001, Dahua Lin |
ACL (1) | 8 |
| 2024 | Advancing Parameter Efficiency in Fine-tuning via Representation EditingabstractMuling Wu, Wenhao Liu, Xiaohua Wang, Tianlong Li, Changze Lv, Zixuan Ling, Zhu JianHao, Cenyuan Zhang, Xiaoqing Zheng, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Muling Wu, Tianlong Li, Changze Lv, Zixuan Ling, Jianhao Zhu, Cenyuan Zhang, Xiaoqing Zheng, Xuanjing Huang 0001 |
ACL (1) | 10 |
| 2024 | Enhancing Contrastive Learning with Noise-Guided Attack: Towards Continual Relation Extraction in the WildabstractThe principle of continual relation extraction (CRE) involves adapting to emerging novel relations while preserving old knowledge.Existing CRE approaches excel in preserving old knowledge but falter when confronted with contaminated data streams, likely due to an artificial assumption of no annotation errors.Recognizing the prevalence of noisy labels in realworld datasets, we introduce a more practical learning scenario, termed as noisy-CRE.In response to this challenge, we propose a noiseresistant contrastive framework called Noiseguided Attack in Contrastive Learning (NaCL), aimed at learning incremental corrupted relations.Diverging from conventional approaches like sample discarding or relabeling in the presence of noisy labels, NaCL takes a transformative route by modifying the feature space through targeted attack.This attack aims to align the feature space with the provided, albeit inaccurate, labels, thereby enhancing contrastive representations.Extensive empirical validations demonstrate the consistent performance improvement of NaCL with increasing noise rates, surpassing state-of-the-art methods 1 . Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 6 |
| 2024 | ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three StagesabstractJunjie Ye, Sixian Li, Guanyu Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Junjie Ye 0005, Sixian Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
ACL (1) | 9 |
| 2024 | Reasoning in Flux: Enhancing Large Language Models Reasoning through Uncertainty-aware Adaptive GuidanceabstractZhangyue Yin, Qiushi Sun, Qipeng Guo, Zhiyuan Zeng, Xiaonan Li, Junqi Dai, Qinyuan Cheng, Xuanjing Huang, Xipeng Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhangyue Yin, Qiushi Sun, Qipeng Guo, Zhiyuan Zeng 0004, Junqi Dai, Qinyuan Cheng, Xuanjing Huang 0001, Xipeng Qiu |
ACL (1) | 8 |
| 2024 | Unveiling Linguistic Regions in Large Language ModelsabstractLarge Language Models (LLMs) have demonstrated considerable cross-lingual alignment and generalization ability.Current research primarily focuses on improving LLMs' crosslingual generalization capabilities.However, there is still a lack of research on the intrinsic mechanisms of how LLMs achieve crosslingual alignment.From the perspective of region partitioning, this paper conducts several investigations on the linguistic competence of LLMs.We discover a core region in LLMs that corresponds to linguistic competence, accounting for approximately 1% of the total model parameters.Removing this core region by setting parameters to zero results in a significant performance decrease across 30 different languages.Furthermore, this core region exhibits significant dimensional dependence, perturbations to even a single parameter on specific dimensions leading to a loss of linguistic competence.Moreover, we discover that distinct monolingual regions exist for different languages, and disruption to these specific regions substantially reduces the LLMs' proficiency in those corresponding languages.Our research also indicates that freezing the core linguistic region during further pre-training can mitigate the issue of catastrophic forgetting (CF), a common phenomenon observed during further pre-training of LLMs.Overall, exploring the LLMs' functional regions provides insights into the foundation of their intelligence 1 . * Equal contributions.† Corresponding authors. 1 Our code is released in https://github.com/ zzhang0179/Unveiling-Linguistic-Regions-in-LLMs. Zhihao Zhang 0002, Jun Zhao 0019, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
ACL (1) | 5 |
| 2024 | DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive LearningabstractVision-and-Language navigation (VLN) requires an agent to navigate in unseen environment by following natural language instruction. For task completion, the agent needs to align and integrate various navigation modalities, including instruction, observation and navigation history. Existing works primarily concentrate on cross-modal attention at the fusion stage to achieve this objective. Nevertheless, modality features generated by disparate uni-encoders reside in their own spaces, leading to a decline in the quality of cross-modal fusion and decision. To address this problem, we propose a Dual-levEL AligNment (DELAN) framework by cross-modal contrastive learning. This framework is designed to align various navigation-related modalities before fusion, thereby enhancing cross-modal interaction and action decision-making. Specifically, we divide the pre-fusion alignment into dual levels: instruction-history level and landmark-observation level according to their semantic correlations. We also reconstruct a dual-level instruction for adaptation to the dual-level alignment. As the training signals for pre-fusion alignment are extremely limited, self-supervised contrastive learning strategies are employed to enforce the matching between different modalities. Our approach seamlessly integrates with the majority of existing models, resulting in improved navigation performance on various VLN benchmarks, including R2R, R4R, RxR and CVDN. Mengfei Du, Binhao Wu, Jiwen Zhang, Zhihao Fan, Ruipu Luo, Xuanjing Huang 0001, Zhongyu Wei |
LREC/COLING | 7 |
| 2024 | Multi-Objective Forward Reasoning and Multi-Reward Backward Refinement for Product Review SummarizationabstractProduct review summarization aims to generate a concise summary based on product reviews to facilitate purchasing decisions. This intricate task gives rise to three challenges in existing work: factual accuracy, aspect comprehensiveness, and content relevance. In this paper, we first propose an FB-Thinker framework to improve the summarization ability of LLMs with multi-objective forward reasoning and multi-reward backward refinement. To enable LLM with these dual capabilities, we present two Chinese product review summarization datasets, Product-CSum and Product-CSum-Cross, for both instruction-tuning and cross-domain evaluation. Specifically, these datasets are collected via GPT-assisted manual annotations from an online forum and public datasets. We further design an evaluation mechanism Product-Eval, integrating both automatic and human evaluation across multiple dimensions for product summarization. Experimental results show the competitiveness and generalizability of our proposed framework in the product review summarization tasks. Siyuan Wang 0025, Ruofei Lai, Xinyu Zhang 0018, Xuanjing Huang 0001, Zhongyu Wei |
LREC/COLING | 6 |
| 2024 | Domain Generalization via Causal Adjustment for Cross-Domain Sentiment AnalysisabstractDomain adaption has been widely adapted for cross-domain sentiment analysis to transfer knowledge from the source domain to the target domain. Whereas, most methods are proposed under the assumption that the target (test) domain is known, making them fail to generalize well on unknown test data that is not always available in practice. In this paper, we focus on the problem of domain generalization for cross-domain sentiment analysis. Specifically, we propose a backdoor adjustment-based causal model to disentangle the domain-specific and domain-invariant representations that play essential roles in tackling domain shift. First, we rethink the cross-domain sentiment analysis task in a causal view to model the causal-and-effect relationships among different variables. Then, to learn an invariant feature representation, we remove the effect of domain confounders (e.g., domain knowledge) using the backdoor adjustment. A series of experiments over many homologous and diverse datasets show the great performance and robustness of our model by comparing it with the state-of-the-art domain generalization baselines. Siyin Wang, Jie Zhou 0015, Qin Chen 0001, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
LREC/COLING | 6 |
| 2024 | PASUM: A Pre-training Architecture for Social Media User Modeling Based on Text GraphabstractModeling social media users is the core of social governance in the digital society. Existing works have incorporated different digital traces to better learn the representations of social media users, including text information encoded by pre-trained language models and social network information encoded by graph models. However, limited by overloaded text information and hard-to-collect social network information, they cannot utilize global text information and cannot be generalized without social relationships. In this paper, we propose a Pre-training Architecture for Social Media User Modeling based on Text Graph(PASUM). We aggregate all microblogs to represent social media users based on the text graph model and learn the mapping from microblogs to user representation. We further design inter-user and intra-user contrastive learning tasks to inject general structural information into the mapping. In different scenarios, we can represent users based on text, even without social network information. Experimental results on various downstream tasks demonstrate the effectiveness and superiority of our framework. Xinyi Mou, Lanqing Xue, Zhenzhe Ying, Weiqiang Wang 0002, Qi Zhang 0001, Xuanjing Huang 0001, Zhongyu Wei |
LREC/COLING | 7 |
| 2024 | Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language ModelsabstractRecent advancements in Chain-of-Thought prompting have facilitated significant breakthroughs for Large Language Models (LLMs) in complex reasoning tasks. Current research enhances the reasoning performance of LLMs by sampling multiple reasoning chains and ensembling based on the answer frequency. However, this approach fails in scenarios where the correct answers are in the minority. We identify this as a primary factor constraining the reasoning capabilities of LLMs, a limitation that cannot be resolved solely based on the predicted answers. To address this shortcoming, we introduce a hierarchical reasoning aggregation framework AoR (Aggregation of Reasoning), which selects answers based on the evaluation of reasoning chains. Additionally, AoR incorporates dynamic sampling, adjusting the number of reasoning chains in accordance with the complexity of the task. Experimental results on a series of complex reasoning tasks show that AoR outperforms prominent ensemble methods. Further analysis reveals that AoR not only adapts various LLMs but also achieves a superior performance ceiling when compared to current methods. Zhangyue Yin, Qiushi Sun, Qipeng Guo, Zhiyuan Zeng 0004, Tianxiang Sun, Qinyuan Cheng, Xiaofeng Mou, Xipeng Qiu, Xuanjing Huang 0001 |
LREC/COLING | 12 |
| 2024 | RoCoIns: Enhancing Robustness of Large Language Models through Code-Style InstructionsabstractLarge Language Models (LLMs) have showcased remarkable capabilities in following human instructions. However, recent studies have raised concerns about the robustness of LLMs for natural language understanding (NLU) tasks when prompted with instructions combining textual adversarial samples. In this paper, drawing inspiration from recent works that LLMs are sensitive to the design of the instructions, we utilize instructions in code style, which are more structural and less ambiguous, to replace typically natural language instructions. Through this conversion, we provide LLMs with more precise instructions and strengthen the robustness of LLMs. Moreover, under few-shot scenarios, we propose a novel method to compose in-context demonstrations using both clean and adversarial samples (adversarial context method) to further boost the robustness of the LLMs. Experiments on eight robustness datasets show that our method consistently outperforms prompting LLMs with natural language, for example, with gpt-3.5-turbo on average, our method achieves an improvement of 5.68% in test set accuracy and a reduction of 5.66 points in Attack Success Rate (ASR). Yuansen Zhang, Xiao Wang 0001, Zhiheng Xi, Han Xia 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
LREC/COLING | 7 |
| 2024 | Subspace Defense: Discarding Adversarial Perturbations by Learning a Subspace for Clean SignalsabstractDeep neural networks (DNNs) are notoriously vulnerable to adversarial attacks that place carefully crafted perturbations on normal examples to fool DNNs. To better understand such attacks, a characterization of the features carried by adversarial examples is needed. In this paper, we tackle this challenge by inspecting the subspaces of sample features through spectral analysis. We first empirically show that the features of either clean signals or adversarial perturbations are redundant and span in low-dimensional linear subspaces respectively with minimal overlap, and the classical low-dimensional subspace projection can suppress perturbation features out of the subspace of clean signals. This makes it possible for DNNs to learn a subspace where only features of clean signals exist while those of perturbations are discarded, which can facilitate the distinction of adversarial examples. To prevent the residual perturbations that is inevitable in subspace learning, we propose an independence criterion to disentangle clean signals from perturbations. Experimental results show that the proposed strategy enables the model to inherently suppress adversaries, which not only boosts model robustness but also motivates new directions of effective adversarial defense. Yuhao Zhou 0005, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
LREC/COLING | 6 |
| 2024 | ORTicket: Let One Robust BERT Ticket Transfer across Different TasksabstractPretrained language models can be applied for various downstream tasks but are susceptible to subtle perturbations. Most adversarial defense methods often introduce adversarial training during the fine-tuning phase to enhance empirical robustness. However, the repeated execution of adversarial training hinders training efficiency when transitioning to different tasks. In this paper, we explore the transferability of robustness within subnetworks and leverage this insight to introduce a novel adversarial defense method ORTicket, eliminating the need for separate adversarial training across diverse downstream tasks. Specifically, (i) pruning the full model using the MLM task (the same task employed for BERT pretraining) yields a task-agnostic robust subnetwork(i.e., winning ticket in Lottery Ticket Hypothesis); and (ii) fine-tuning this subnetwork for downstream tasks. Extensive experiments demonstrate that our approach achieves comparable robustness to other defense methods while retaining the efficiency of traditional fine-tuning.This also confirms the significance of selecting MLM task for identifying the transferable robust subnetwork. Furthermore, our method is orthogonal to other adversarial training approaches, indicating the potential for further enhancement of model robustness. Yuhao Zhou 0005, Wenxiang Chen, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
LREC/COLING | 7 |
| 2024 | LawLLM: Intelligent Legal System with Legal Reasoning and Verifiable Retrieval
Shengbin Yue, Shujun Liu, Chenchen Shen, Siyuan Wang 0025, Yun Song, Wei Chen 0088, Xuanjing Huang 0001, Zhongyu Wei |
DASFAA (5) | 11 |
| 2024 | LONGAGENT: Achieving Question Answering for 128k-Token-Long Documents through Multi-Agent CollaborationabstractLarge language models (LLMs) have achieved tremendous success in understanding language and processing text.However, questionanswering (QA) on lengthy documents faces challenges of resource constraints and a high propensity for errors, even for the most advanced models such as GPT-4 and Claude2.In this paper, we introduce LONGAGENT, a multi-agent collaboration method that enables efficient and effective QA over 128k-tokenlong documents.LONGAGENT adopts a divideand-conquer strategy, breaking down lengthy documents into shorter, more manageable text chunks.A leader agent comprehends the user's query and organizes the member agents to read their assigned chunks, reasoning a final answer through multiple rounds of discussion.Due to members' hallucinations, it's difficult to guarantee that every response provided by each member is accurate.To address this, we develop an inter-member communication mechanism that facilitates information sharing, allowing for the detection and mitigation of hallucinatory responses.Experimental results show that a LLaMA-2 7B driven by LONGAGENT can effectively support QA over 128k-token documents, achieving 16.42% and 1.63% accuracy gains over GPT-4 on singlehop and multi-hop QA settings, respectively. Jun Zhao 0019, Can Zu, Xu Hao, Wei He 0024, Yiwen Ding, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 9 |
| 2024 | Improving Discriminative Capability of Reward Models in RLHF Using Contrastive LearningabstractLu Chen, Rui Zheng, Binghai Wang, Senjie Jin, Caishuang Huang, Junjie Ye, Zhihao Zhang, Yuhao Zhou, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Lu Chen 0001, Binghai Wang, Senjie Jin, Caishuang Huang, Junjie Ye 0005, Zhihao Zhang 0002, Yuhao Zhou 0005, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 12 |
| 2024 | Searching for Best Practices in Retrieval-Augmented GenerationabstractXiaohua Wang, Zhenghua Wang, Xuan Gao, Feiran Zhang, Yixin Wu, Zhibo Xu, Tianyuan Shi, Zhengyuan Wang, Shizheng Li, Qi Qian, Ruicheng Yin, Changze Lv, Xiaoqing Zheng, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhenghua Wang, Feiran Zhang, Yixin Wu 0005, Zhibo Xu, Tianyuan Shi, Zhengyuan Wang, Shizheng Li, Ruicheng Yin, Changze Lv, Xiaoqing Zheng, Xuanjing Huang 0001 |
EMNLP | 14 |
| 2024 | RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool LearningabstractJunjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Guanyu Li, Xiaoran Fan, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Junjie Ye 0005, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Xiaoran Fan, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
EMNLP | 10 |
| 2024 | Explicit Memory Learning with Expectation MaximizationabstractLarge Language Models (LLMs) have revolutionized the landscape of natural language processing, demonstrating remarkable abilities across various complex tasks.However, their stateless nature limits the capability to retain information across interactions, hindering performance in scenarios requiring historical context recall.To mitigate this, current approaches primarily use explicit memory to allow LLMs to store useful information, which is accessible, readable, and interpretable.Nevertheless, explicit memory lacks the reliable learning mechanisms of implicit memory, which can be optimized end-to-end.To harness the benefits of both, we introduce EM 2 , a novel framework enhancing explicit memory updates via the Expectation-Maximization (EM) algorithm.EM 2 treats memory as a latent variable, ensuring continual learning and improvement during updates.Experimental results on streaming inference tasks demonstrate that EM 2 outperforms existing methods without memory or with static external memory.Our in-depth analysis highlights that EM 2 significantly enhances performance across various backbones and memory strategies, providing a robust solution for advancing LLM memory management and enabling explicit memory to learn and improve similarly to implicit memory. Zhangyue Yin, Qiushi Sun, Qipeng Guo, Zhiyuan Zeng 0004, Qinyuan Cheng, Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 7 |
| 2024 | TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer CapabilitiesabstractMing Zhang, Caishuang Huang, Yilong Wu, Shichun Liu, Huiyuan Zheng, Yurui Dong, Yujiong Shen, Shihan Dou, Jun Zhao, Junjie Ye, Qi Zhang, Tao Gui, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ming Zhang 0030, Caishuang Huang, Yilong Wu, Shichun Liu, Huiyuan Zheng, Yurui Dong 0001, Yujiong Shen, Shihan Dou, Jun Zhao 0019, Junjie Ye 0005, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001 |
EMNLP | 13 |
| 2024 | Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning Through Trap ProblemsabstractHuman cognition exhibits systematic compositionality, the algebraic ability to generate infinite novel combinations from finite learned components, which is the key to understanding and reasoning about complex logic.In this work, we investigate the compositionality of large language models (LLMs) in mathematical reasoning.Specifically, we construct a new dataset MATHTRAP ‡ by introducing carefully designed logical traps into the problem descriptions of MATH and GSM8K.Since problems with logical flaws are quite rare in the real world, these represent "unseen" cases to LLMs.Solving these requires the models to systematically compose (1) the mathematical knowledge involved in the original problems with (2) knowledge related to the introduced traps.Our experiments show that while LLMs possess both components of requisite knowledge, they do not spontaneously combine them to handle these novel cases.We explore several methods to mitigate this deficiency, such as natural language prompts, few-shot demonstrations, and fine-tuning.Additionally, we test the recently released OpenAI o1 model and find that human-like 'slow thinking' helps improve the compositionality of LLMs.Overall, systematic compositionality remains an open challenge for large language models. Jun Zhao 0019, Jingqi Tong, Yurong Mou, Ming Zhang 0030, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 6 |
| 2024 | Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMsabstractXin Zhou, Ping Nie, Yiwen Guo, Haojie Wei, Zhanqiu Zhang, Pasquale Minervini, Ruotian Ma, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Xin Zhou 0012, Ping Nie, Yiwen Guo, Haojie Wei, Zhanqiu Zhang, Pasquale Minervini, Ruotian Ma, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 10 |
| 2024 | A Soft Contrastive Learning-Based Prompt Model for Few-Shot Sentiment AnalysisabstractFew-shot text classification has attracted great interest in both academia and industry due to the lack of labeled data in many fields. Different from general text classification (e.g., topic classification), few-shot sentiment classification is more challenging because the semantic distances among the classes are more subtle. For instance, the semantic distances between the sentiment labels in a positive or negative polarity (e.g., "love" and "joy", "remorse" and "sadness") are close, while the distances are large for the sentiment labels in two opposite polarities (e.g., "love" and "sadness"). To address this problem, we propose a Soft Contrastive learning-based Prompt (SCP) model for few-shot sentiment analysis. First, we design a sentiment-aware chain of thought prompt module to guide the model to predict the sentiment from coarse grain to fine grain via a series of intermediate reasoning steps. Then, we propose a soft contrastive learning algorithm to take the correlation of the labels into account. A series of experiments on several sentiment analysis datasets show the great advantages of SCP by comparing it with SOTA baselines (e.g., ChatGPT). Jie Zhou 0015, Jiabao Zhao, Siyin Wang, Haijun Shan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICASSP | 8 |
| 2024 | Improving Generalization of Alignment with Human Preferences through Group Invariant LearningabstractThe success of AI assistants based on language models (LLMs) hinges crucially on Reinforcement Learning from Human Feedback (RLHF), which enables the generation of responses more aligned with human preferences.
As universal AI assistants, there's a growing expectation for them to perform consistently across various domains.
However, previous work shows that Reinforcement Learning (RL) often exploits shortcuts to attain high rewards and overlooks challenging samples.
This focus on quick reward gains undermines both the stability in training and the model's ability to generalize to new, unseen data.
In this work, we propose a novel approach that can learn a consistent policy via RL across various data groups or domains.
Given the challenges associated with acquiring group annotations, our method automatically classifies data into different groups, deliberately maximizing performance variance.
Then, we optimize the policy to perform well on challenging groups.
Lastly, leveraging the established groups, our approach adaptively adjusts the exploration space, allocating more learning capacity to more challenging data and preventing the model from over-optimizing on simpler data. Experimental results indicate that our approach significantly enhances training stability and model generalization. Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou 0005, Zhiheng Xi, Xiao Wang 0001, Haoran Huang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICLR | 12 |
| 2024 | Efficient and Effective Time-Series Forecasting with Spiking Neural NetworksabstractSpiking neural networks (SNNs), inspired by the spiking behavior of biological neurons, provide a unique pathway for capturing the intricacies of temporal data. However, applying SNNs to time-series forecasting is challenging due to difficulties in effective temporal alignment, complexities in encoding processes, and the absence of standardized guidelines for model selection. In this paper, we propose a framework for SNNs in time-series forecasting tasks, leveraging the efficiency of spiking neurons in processing temporal information. Through a series of experiments, we demonstrate that our proposed SNN-based approaches achieve comparable or superior results to traditional time-series forecasting methods on diverse benchmarks with much less energy consumption. Furthermore, we conduct detailed analysis experiments to assess the SNN’s capacity to capture temporal dependencies within time-series data, offering valuable insights into its nuanced strengths and effectiveness in modeling the intricate dynamics of temporal data. Our study contributes to the expanding field of SNNs and offers a promising alternative for time-series forecasting tasks, presenting a pathway for the development of more biologically inspired and temporally aware forecasting models. Our code is available at https://github.com/microsoft/SeqSNN. Changze Lv, Yansen Wang, Xiaoqing Zheng, Xuanjing Huang 0001, Dongsheng Li 0002 |
ICML | 5 |
| 2024 | Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement LearningabstractIn this paper, we propose R$^3$: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the benefits of process supervision for large language models. The core challenge in applying RL to complex reasoning is to identify a sequence of actions that result in positive rewards and provide appropriate supervision for optimization. Outcome supervision provides sparse rewards for final results without identifying error locations, whereas process supervision offers step-wise rewards but requires extensive manual annotation. R$^3$ overcomes these limitations by learning from correct demonstrations. Specifically, R$^3$ progressively slides the start state of reasoning from a demonstration’s end to its beginning, facilitating easier model exploration at all stages. Thus, R$^3$ establishes a step-wise curriculum, allowing outcome supervision to offer step-level signals and precisely pinpoint errors. Using Llama2-7B, our method surpasses RL baseline on eight reasoning tasks by $4.1$ points on average. Notably, in program-based reasoning, 7B-scale models perform comparably to larger models or closed-source models with our R$^3$. Zhiheng Xi, Wenxiang Chen, Boyang Hong, Senjie Jin, Wei He 0024, Yiwen Ding, Shichun Liu, Junzhe Wang 0001, Honglin Guo, Xiaoran Fan, Yuhao Zhou 0005, Shihan Dou, Xiao Wang 0001, Xinbo Zhang, Peng Sun 0006, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ICML | 21 |
| 2024 | ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented BenchmarksabstractRecent years have witnessed remarkable progress in the development of large vision-language models (LVLMs). Benefiting from the strong language backbones and efficient cross-modal alignment strategies, LVLMs exhibit surprising capabilities to perceive visual signals and perform visually grounded reasoning. However, the capabilities of LVLMs have not been comprehensively and quantitatively evaluated. Most existing multi-modal benchmarks require task-oriented input-output formats, posing great challenges to automatically assess the free-form text output of LVLMs. To effectively leverage the annotations available and reduce the manual efforts required for constructing new benchmarks, we propose to re-formulate existing benchmarks into unified LVLM-compatible formats. Through systematic data collection and reformulation, we present ReForm-Eval benchmark, offering substantial data for evaluating various capabilities of LVLMs. Through extensive experiments and analysis in ReForm-Eval, we demonstrate the comprehensiveness and reliability of ReForm-Eval in assessing various LVLMs. Our benchmark and evaluation framework is now available at https://github.com/FudanDISC/ReForm-Eval Mengfei Du, Qingwen Liu 0002, Binhao Wu, Jiwen Zhang, Chengxing Zhou, Zhihao Fan, Jie Fu 0001, Jingjing Chen 0001, Zhongyu Wei, Xuanjing Huang 0001 |
ACM Multimedia | 12 |
| 2024 | Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and BaselineabstractExisting video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watching the videos. Induced sentiment of viewers is essential for inferring the public response to videos and has broad application in analyzing public societal sentiment, effectiveness of advertising and other areas. The micro videos and the related comments provide a rich application scenario for viewers’ induced sentiment analysis. In light of this, we introduces a novel research task, Multimodal Sentiment Analysis for Comment Response of Video Induced(MSA-CRVI), aims to infer opinions and emotions according to comments response to micro video. Meanwhile, we manually annotate a dataset named Comment Sentiment toward to Micro Video (CSMV) to support this research. It is the largest video multi-modal sentiment dataset in terms of scale and video duration to our knowledge, containing 107, 267 comments and 8, 210 micro videos with a video duration of 68.83 hours. To infer the induced sentiment of comment should leverage the video content, we propose the Video Content-aware Comment Sentiment Analysis (VC-CSA) method as a baseline to address the challenges inherent in this new task. Extensive experiments demonstrate that our method is showing significant improvements over other established baselines. We make the dataset and source code publicly available at https://github.com/IEIT-AGI/MSA-CRVI. Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Guoguang Du 0001, Zhenhua Guo 0003, Yaqian Zhao, Xuanjing Huang 0001, RenGang Li |
NeurIPS | 9 |
| 2024 | Advancing Spiking Neural Networks for Sequential Modeling with Central Pattern GeneratorsabstractSpiking neural networks (SNNs) represent a promising approach to developing artificial neural networks that are both energy-efficient and biologically plausible.
However, applying SNNs to sequential tasks, such as text classification and time-series forecasting, has been hindered by the challenge of creating an effective and hardware-friendly spike-form positional encoding (PE) strategy.
Drawing inspiration from the central pattern generators (CPGs) in the human brain, which produce rhythmic patterned outputs without requiring rhythmic inputs, we propose a novel PE technique for SNNs, termed CPG-PE.
We demonstrate that the commonly used sinusoidal PE is mathematically a specific solution to the membrane potential dynamics of a particular CPG.
Moreover, extensive experiments across various domains, including time-series forecasting, natural language processing, and image classification, show that SNNs with CPG-PE outperform their conventional counterparts.
Additionally, we perform analysis experiments to elucidate the mechanism through which SNNs encode positional information and to explore the function of CPGs in the human brain.
This investigation may offer valuable insights into the fundamental principles of neural computation. Changze Lv, Yansen Wang, Xiaoqing Zheng, Xuanjing Huang 0001, Dongsheng Li 0002 |
NeurIPS | 5 |
| 2024 | Automating Dataset Updates Towards Reliable and Timely Evaluation of Large Language ModelsabstractLarge language models (LLMs) have achieved impressive performance across various natural language benchmarks, prompting a continual need to curate more difficult datasets for larger LLMs, which is costly and time-consuming. In this paper, we propose to automate dataset updating and provide systematical analysis regarding its effectiveness in dealing with benchmark leakage issue, difficulty control, and stability. Thus, once current benchmark has been mastered or leaked, we can update it for timely and reliable evaluation. There are two updating strategies: 1) mimicking strategy to generate similar samples based on original data, preserving stylistic and contextual essence, and 2) extending strategy that further expands existing samples at varying cognitive levels by adapting Bloom’s taxonomy of educational objectives. Extensive experiments on updated MMLU and BIG-Bench demonstrate the stability of the proposed strategies and find that the mimicking strategy can effectively alleviate issues of overestimation from benchmark leakage. In cases where the efficient mimicking strategy fails, our extending strategy still shows promising results. Additionally, by controlling the difficulty, we can better discern the models’ performance and enable fine-grained analysis — neither too difficult nor too easy an exam can fairly judge students’ learning status. To the best of our knowledge, we are the first to automate updating benchmarks for reliable and timely evaluation. Our demo leaderboard can be found at https://yingjiahao14.github.io/Automating-DatasetUpdates/. Jiahao Ying, Yixin Cao 0002, Yushi Bai, Qianru Sun, Wei Tang 0015, Zhaojun Ding, Yizhe Yang, Xuanjing Huang 0001, Shuicheng Yan |
NeurIPS | 9 |
| 2024 | Empowering LLMs for Long-Text Information Extraction in Chinese Legal Documents
Chenchen Shen, Chengwei Ji, Shengbin Yue, Yun Song, Xuanjing Huang 0001, Zhongyu Wei |
NLPCC (1) | 6 |
| 2024 | CausalABSC: Causal Inference for Aspect Debiasing in Aspect-Based Sentiment ClassificationabstractAs the primary subtask of sentiment analysis, aspect-based sentiment classification (ABSC) aims to predict the sentiment polarity for a given aspect. While recent deep neural models for ABSC have shown good performance, their robustness is limited due to their reliance on spurious correlations between aspects and sentiment. Specifically, most existing models tend to assign the most frequent sentiment label to a certain aspect in different sentences, or assume different aspects in the same sentence to have the same sentiment polarity, which is vulnerable due to the aspect biases. In this paper, we propose a causal graph to identify and analyze the causal relationships among treatment variables (e.g., aspect, sentence), intermediate variables (e.g., aspect-aware content), and outcome variables (e.g., sentiment polarity) for ABSC. To address the issue of spurious relationships that mislead sentiment polarity prediction, we introduce a novel causal inference framework called CausalABSC. CausalABSC is model agnostic, allowing integration into existing methods. We conduct extensive experiments on five benchmark datasets, which demonstrate the state-of-the-art performance of CausalABSC and its effectiveness in debiasing. Jie Zhou 0015, Yuanbiao Lin, Qin Chen 0001, Qi Zhang 0001, Xuanjing Huang 0001, Liang He 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | UTC-IE: A Unified Token-pair Classification Architecture for Information ExtractionabstractInformation Extraction (IE) spans several tasks with different output structures, such as named entity recognition, relation extraction and event extraction.Previously, those tasks were solved with different models because of diverse task output structures.Through re-examining IE tasks, we find that all of them can be interpreted as extracting spans and span relations.They can further be decomposed into tokenpair classification tasks by using the start and end token of a span to pinpoint the span, and using the start-to-start and end-to-end token pairs of two spans to determine the relation.Based on the reformulation, we propose a Unified Token-pair Classification architecture for Information Extraction (UTC-IE), where we introduce Plusformer on top of the tokenpair feature matrix.Specifically, it models axis-aware interaction with plus-shaped selfattention and local interaction with Convolutional Neural Network over token pairs.Experiments show that our approach outperforms task-specific and unified models on all tasks in 10 datasets, and achieves better or comparable results on 2 joint IE datasets.Moreover, UTC-IE speeds up over state-of-the-art models on IE tasks significantly in most datasets, which verifies the effectiveness of our architecture.1 * Equal contribution. Hang Yan 0001, Yu Sun 0031, Yunhua Zhou, Xuanjing Huang 0001, Xipeng Qiu |
ACL (1) | 5 |
| 2023 | TableVLM: Multi-modal Pre-training for Table Structure RecognitionabstractTables are widely used in research and business, and are suitable for human consumption, but not easily machine-processable, particularly when tables are present in images.One of the main challenges to extracting data from images of tables is to accurately recognize table structures, especially for complex tables with cross rows and columns.In this study, we propose a novel multi-modal pre-training model for table structure recognition, named TableVLM.With a two-stream multi-modal transformer-based encoder-decoder architecture, TableVLM learns to capture rich table structure-related features by multiple carefullydesigned unsupervised objectives inspired by the notion of masked visual-language modeling.To pre-train this model, we also created a dataset, called ComplexTable, which consists of 1, 000K samples to be released publicly.Experiment results show that the model built on pre-trained TableVLM can improve the performance up to 1.97% in tree-editing-distancescore on ComplexTable. Leiyuan Chen, Chengsong Huang, Xiaoqing Zheng, Jinshu Lin, Xuanjing Huang 0001 |
ACL (1) | 5 |
| 2023 | DiffusionBERT: Improving Generative Masked Language Models with Diffusion ModelsabstractWe present DiffusionBERT, a new generative masked language model based on discrete diffusion models.Diffusion models and many pretrained language models have a shared training objective, i.e., denoising, making it possible to combine the two powerful models and enjoy the best of both worlds.On the one hand, diffusion models offer a promising training strategy that helps improve the generation quality.On the other hand, pre-trained denoising language models (e.g., BERT) can be used as a good initialization that accelerates convergence.We explore training BERT to learn the reverse process of a discrete diffusion process with an absorbing state and elucidate several designs to improve it.First, we propose a new noise schedule for the forward diffusion process that controls the degree of noise added at each step based on the information of each token.Second, we investigate several designs of incorporating the time step into BERT.Experiments on unconditional text generation demonstrate that DiffusionBERT achieves significant improvement over existing diffusion models for text (e.g., D3PM and Diffusion-LM) and previous generative masked language models in terms of perplexity and BLEU score.Promising results in conditional generation tasks show that DiffusionBERT can generate texts of comparable quality and more diverse than a series of established baselines. Zhengfu He, Tianxiang Sun, Qiong Tang, Kuanning Wang, Xuanjing Huang 0001, Xipeng Qiu |
ACL (1) | 5 |
| 2023 | Unifying Cross-Lingual and Cross-Modal Modeling Towards Weakly Supervised Multilingual Vision-Language Pre-trainingabstractMultilingual Vision-Language Pre-training (VLP) is a promising but challenging topic due to the lack of large-scale multilingual imagetext pairs.Existing works address the problem by translating English data into other languages, which is intuitive and the generated data is usually limited in form and scale.In this paper, we explore a more practical and scalable setting: weakly supervised multilingual VLP with only English image-text pairs and multilingual text corpora.We argue that the universal multilingual representation learned from texts allows the cross-modal interaction learned in English to be transferable to other languages.To this end, we propose a framework to effectively unify cross-lingual and cross-modal pre-training.For unified modeling on different data, we design an architecture with flexible modules to learn different interactions.Moreover, two unified tasks are introduced to efficiently guide the unified crosslingual cross-modal learning.Extensive experiments demonstrate that our pre-trained model learns universal multilingual multimodal representations, allowing effective cross-lingual transfer on multimodal tasks.Code and models are available at https://github.com/ FudanDISC/weakly-supervised-mVLP. Zhihao Fan, Jingjing Chen 0001, Qi Zhang 0001, Xuanjing Huang 0001, Zhongyu Wei |
ACL (1) | 5 |
| 2023 | CodeIE: Large Code Generation Models are Better Few-Shot Information ExtractorsabstractPeng Li, Tianxiang Sun, Qiong Tang, Hang Yan, Yuanbin Wu, Xuanjing Huang, Xipeng Qiu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tianxiang Sun, Qiong Tang, Hang Yan 0001, Yuanbin Wu, Xuanjing Huang 0001, Xipeng Qiu |
ACL (1) | 6 |
| 2023 | UPPAM: A Unified Pre-training Architecture for Political Actor Modeling based on LanguageabstractModeling political actors is at the core of quantitative political science.Existing works have incorporated contextual information to better learn the representation of political actors for specific tasks through graph models.However, they are limited to the structure and objective of training settings and can not be generalized to all politicians and other tasks.In this paper, we propose a Unified Pre-training Architecture for Political Actor Modeling based on language (UPPAM).In UPPAM, we aggregate statements to represent political actors and learn the mapping from languages to representation, instead of learning the representation of particular persons.We further design structureaware contrastive learning and behavior-driven contrastive learning tasks, to inject multidimensional information in the political context into the mapping.In this framework, we can profile political actors from different aspects and solve various downstream tasks.Experimental results demonstrate the effectiveness and capability of generalization of our method. * Corresponding author.• No church needs to provide contraception under ObamaCare.• Recovery package must provide state aid, hazard pay.• LGBT rights are in jeopardy from Supreme Court.• I oppose school busing because it fails, not for racism. languages social network behaviorsVoted NO on defining unborn child as eligible for SCHIP. Xinyi Mou, Zhongyu Wei, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2023 | Multitask Pre-training of Modular Prompt for Chinese Few-Shot LearningabstractPrompt tuning is a parameter-efficient approach to adapting pre-trained language models to downstream tasks.Although prompt tuning has been shown to match the performance of full model tuning when training data is sufficient, it tends to struggle in few-shot learning settings.In this paper, we present Multi-task Pre-trained Modular Prompt (MP 2 ) to boost prompt tuning for few-shot learning.MP 2 is a set of combinable prompts pre-trained on 38 Chinese tasks.On downstream tasks, the pre-trained prompts are selectively activated and combined, leading to strong compositional generalization to unseen tasks.To bridge the gap between pre-training and fine-tuning, we formulate upstream and downstream tasks into a unified machine reading comprehension task.Extensive experiments under two learning paradigms, i.e., gradient descent and black-box tuning, show that MP 2 significantly outperforms prompt tuning, full model tuning, and prior prompt pretraining methods in few-shot settings.In addition, we demonstrate that MP 2 can achieve surprisingly fast and strong adaptation to downstream tasks by merely learning 8 parameters to combine the pre-trained modular prompts. Tianxiang Sun, Zhengfu He, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 5 |
| 2023 | Query Structure Modeling for Inductive Logical Reasoning Over Knowledge GraphsabstractSiyuan Wang, Zhongyu Wei, Meng Han, Zhihao Fan, Haijun Shan, Qi Zhang, Xuanjing Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Haijun Shan, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 7 |
| 2023 | Open Set Relation Extraction via Unknown-Aware TrainingabstractJun Zhao, Xin Zhao, WenYu Zhan, Qi Zhang, Tao Gui, Zhongyu Wei, Yun Wen Chen, Xiang Gao, Xuanjing Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jun Zhao 0019, Wenyu Zhan, Qi Zhang 0001, Tao Gui, Zhongyu Wei, Yun Wen Chen, Xiang Gao 0017, Xuanjing Huang 0001 |
ACL (1) | 9 |
| 2023 | Graph Structure Learning via Lottery Hypothesis at Scale
Yuxin Wang 0005, Xiannian Hu, Jiaqing Xie, Zhangyue Yin, Yunhua Zhou, Xipeng Qiu, Xuanjing Huang 0001 |
ACML | 7 |
| 2023 | Hi-ArG: Exploring the Integration of Hierarchical Argumentation Graphs in Language PretrainingabstractJingcong Liang, Rong Ye, Meng Han, Qi Zhang, Ruofei Lai, Xinyu Zhang, Zhao Cao, Xuanjing Huang, Zhongyu Wei. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jingcong Liang, Rong Ye, Qi Zhang 0001, Ruofei Lai, Xinyu Zhang 0019, Zhao Cao, Xuanjing Huang 0001, Zhongyu Wei |
EMNLP | 8 |
| 2023 | Argue with Me Tersely: Towards Sentence-Level Counter-Argument GenerationabstractCounter-argument generation-a captivating area in computational linguistics-seeks to craft statements that offer opposing views.While most research has ventured into paragraph-level generation, sentence-level counter-argument generation beckons with its unique constraints and brevity-focused challenges.Furthermore, the diverse nature of counter-arguments poses challenges for evaluating model performance solely based on ngram-based metrics.In this paper, we present the ArgTersely benchmark for sentence-level counter-argument generation, drawing from a manually annotated dataset from the Change-MyView debate forum 1 .We also propose Arg-LlaMA for generating high-quality counterargument.For better evaluation, we trained a BERT-based evaluator Arg-Judge with human preference data.We conducted comparative experiments involving various baselines such as LlaMA, Alpaca, GPT-3, and others.The results show the competitiveness of our proposed framework and evaluator in counter-argument generation tasks. Rong Ye, Qi Zhang 0001, Ruofei Lai, Xinyu Zhang 0019, Zhao Cao, Xuanjing Huang 0001, Zhongyu Wei |
EMNLP | 8 |
| 2023 | Towards Building More Robust NER datasets: An Empirical Study on NER Dataset Bias from a Dataset Difficulty ViewabstractRecently, many studies have illustrated the robustness problem of Named Entity Recognition (NER) systems: the NER models often rely on superficial entity patterns for predictions, without considering evidence from the context.Consequently, even state-of-theart NER models generalize poorly to outof-domain scenarios when out-of-distribution (OOD) entity patterns are introduced.Previous research attributes the robustness problem to the existence of NER dataset bias, where simpler and regular entity patterns induce shortcut learning.In this work, we bring new insights into this problem by comprehensively investigating the NER dataset bias from a dataset difficulty view.We quantify the entity-context difficulty distribution in existing datasets and explain their relationship with model robustness.Based on our findings, we explore three potential ways to de-bias the NER datasets by altering entity-context distribution, and we validate the feasibility with intensive experiments.Finally, we show that the de-biased datasets can transfer to different models and even benefit existing model-based robustness-improving methods, indicating that building more robust datasets is fundamental for building more robust NER systems. * Equal contribution. † Corresponding authors.Ms. Hall expects Cathay 's profit to grow around 13 % annually this year. Ruotian Ma, Xin Zhou 0012, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2023 | Hallucination Detection for Generative Large Language Models by Bayesian Sequential EstimationabstractLarge Language Models (LLMs) have made remarkable advancements in the field of natural language generation.However, the propensity of LLMs to generate inaccurate or non-factual content, termed "hallucinations", remains a significant challenge.Current hallucination detection methods often necessitate the retrieval of great numbers of relevant evidence, thereby increasing response times.We introduce a unique framework that leverages statistical decision theory and Bayesian sequential analysis to optimize the trade-off between costs and benefits during the hallucination detection process.This approach does not require a predetermined number of observations.Instead, the analysis proceeds in a sequential manner, enabling an expeditious decision towards "belief" or "disbelief" through a stop-or-continue strategy.Extensive experiments reveal that this novel framework surpasses existing methods in both efficiency and precision of hallucination detection.Furthermore, it requires fewer retrieval steps on average, thus decreasing response times 1 . Yuliang Yan, Longtao Huang, Xiaoqing Zheng, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2023 | Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model CommunicationabstractLarge Language Models (LLMs) have recently made significant strides in complex reasoning tasks through the Chain-of-Thought technique.Despite this progress, their reasoning is often constrained by their intrinsic understanding, lacking external insights.To address this, we propose Exchange-of-Thought (EoT), a novel framework that enables cross-model communication during problem-solving.Drawing inspiration from network topology, EoT integrates four unique communication paradigms: Memory, Report, Relay, and Debate.This paper delves into the communication dynamics and volume associated with each paradigm.To counterbalance the risks of incorrect reasoning chains, we implement a robust confidence evaluation mechanism within these communications.Our experiments across diverse complex reasoning tasks demonstrate that EoT significantly surpasses established baselines, underscoring the value of external insights in enhancing LLM performance.Furthermore, we show that EoT achieves these superior results in a cost-effective manner, marking a promising advancement for efficient and collaborative AI problem-solving."Two heads are better than one. Zhangyue Yin, Qiushi Sun, Qipeng Guo, Junqi Dai, Xuanjing Huang 0001, Xipeng Qiu |
EMNLP | 6 |
| 2023 | From Hypergraph Energy Functions to Hypergraph Neural NetworksabstractHypergraphs are a powerful abstraction for representing higher-order interactions between entities of interest. To exploit these relationships in making downstream predictions, a variety of hypergraph neural network architectures have recently been proposed, in large part building upon precursors from the more traditional graph neural network (GNN) literature. Somewhat differently, in this paper we begin by presenting an expressive family of parameterized, hypergraph-regularized energy functions. We then demonstrate how minimizers of these energies effectively serve as node embeddings that, when paired with a parameterized classifier, can be trained end-to-end via a supervised bilevel optimization process. Later, we draw parallels between the implicit architecture of the predictive models emerging from the proposed bilevel hypergraph optimization, and existing GNN architectures in common use. Empirically, we demonstrate state-of-the-art results on various hypergraph node classification benchmarks. Code is available at https://github.com/yxzwang/PhenomNN. Yuxin Wang 0005, Xipeng Qiu, Xuanjing Huang 0001, David P. Wipf |
ICML | 4 |
| 2023 | A benchmark for automatic medical consultation system: frameworks, tasks and datasetsabstractMOTIVATION: In recent years, interest has arisen in using machine learning to improve the efficiency of automatic medical consultation and enhance patient experience. In this article, we propose two frameworks to support automatic medical consultation, namely doctor-patient dialogue understanding and task-oriented interaction. We create a new large medical dialogue dataset with multi-level fine-grained annotations and establish five independent tasks, including named entity recognition, dialogue act classification, symptom label inference, medical report generation and diagnosis-oriented dialogue policy. RESULTS: We report a set of benchmark results for each task, which shows the usability of the dataset and sets a baseline for future studies. AVAILABILITY AND IMPLEMENTATION: Both code and data are available from https://github.com/lemuria-wchen/imcs21. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wei Chen 0088, Hongyi Fang, Qianyuan Yao, Jianye Hao, Qi Zhang 0001, Xuanjing Huang 0001, Jiajie Peng, Zhongyu Wei |
Bioinform. | 8 |
| 2023 | Certified Robustness to Text Adversarial Attacks by Randomized [MASK]abstractAbstract Very recently, few certified defense methods have been developed to provably guarantee the robustness of a text classifier to adversarial synonym substitutions. However, all the existing certified defense methods assume that the defenders have been informed of how the adversaries generate synonyms, which is not a realistic scenario. In this study, we propose a certifiably robust defense method by randomly masking a certain proportion of the words in an input text, in which the above unrealistic assumption is no longer necessary. The proposed method can defend against not only word substitution-based attacks, but also character-level perturbations. We can certify the classifications of over 50% of texts to be robust to any perturbation of five words on AGNEWS, and two words on SST2 dataset. The experimental results show that our randomized smoothing method significantly outperforms recently proposed defense methods across multiple datasets under different attack algorithms. Jiehang Zeng, Jianhan Xu, Xiaoqing Zheng, Xuanjing Huang 0001 |
Comput. Linguistics | 4 |
| 2023 | Multi-modal multi-hop interaction network for dialogue response generation
Jie Zhou 0015, Rui Wang 0005, Yuanbin Wu, Ming Yan 0008, Liang He 0001, Xuanjing Huang 0001 |
Expert Syst. Appl. | 7 |
| 2023 | Improving BERT Fine-Tuning via Self-Ensemble and Self-Distillation
Yige Xu 0001, Xipeng Qiu, Ligao Zhou, Xuanjing Huang 0001 |
J. Comput. Sci. Technol. | 4 |
| 2023 | Chinese Named Entity Recognition Augmented with Lexicon Memory
Yi Zhou 0018, Xiaoqing Zheng, Xuanjing Huang 0001 |
J. Comput. Sci. Technol. | 3 |
| 2022 | CQG: A Simple and Effective Controlled Generation Framework for Multi-hop Question GenerationabstractMulti-hop question generation focuses on generating complex questions that require reasoning over multiple pieces of information of the input passage.Current models with state-of-the-art performance have been able to generate the correct questions corresponding to the answers.However, most models can not ensure the complexity of generated questions, so they may generate shallow questions that can be answered without multi-hop reasoning.To address this challenge, we propose the CQG, which is a simple and effective controlled framework.CQG employs a simple method to generate the multi-hop questions that contain key entities in multi-hop reasoning chains, which ensure the complexity and quality of the questions.In addition, we introduce a novel controlled Transformer-based decoder to guarantee that key entities appear in the questions.Experiment results show that our model greatly improves performance, which also outperforms the state-of-the-art model about 25% by 5 BLEU points on HotpotQA 1 . Zichu Fei, Qi Zhang 0001, Tao Gui, Di Liang, Wei Wu 0014, Xuanjing Huang 0001 |
ACL (1) | 7 |
| 2022 | Flooding-X: Improving BERT's Resistance to Adversarial Attacks via Loss-Restricted Fine-TuningabstractQin Liu, Rui Zheng, Bao Rong, Jingyi Liu, ZhiHua Liu, Zhanzhan Cheng, Liang Qiao, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Qin Liu 0010, Bao Rong, Zhanzhan Cheng, Liang Qiao 0001, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 10 |
| 2022 | MINER: Improving Out-of-Vocabulary Named Entity Recognition from an Information Theoretic PerspectiveabstractXiao Wang, Shihan Dou, Limao Xiong, Yicheng Zou, Qi Zhang, Tao Gui, Liang Qiao, Zhanzhan Cheng, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Xiao Wang 0001, Shihan Dou, Limao Xiong, Yicheng Zou, Qi Zhang 0001, Tao Gui, Liang Qiao 0001, Zhanzhan Cheng, Xuanjing Huang 0001 |
ACL (1) | 9 |
| 2022 | Robust Lottery Tickets for Pre-trained Language ModelsabstractRui Zheng, Bao Rong, Yuhao Zhou, Di Liang, Sirui Wang, Wei Wu, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Bao Rong, Yuhao Zhou 0005, Di Liang, Wei Wu 0014, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 9 |
| 2022 | CoLo: A Contrastive Learning Based Re-ranking Framework for One-Stage SummarizationabstractTraditional training paradigms for extractive and abstractive summarization systems always only use token-level or sentence-level training objectives. However, the output summary is always evaluated from summary-level which leads to the inconsistency in training and evaluation. In this paper, we propose a Contrastive Learning based re-ranking framework for one-stage summarization called CoLo. By modeling a contrastive objective, we show that the summarization model is able to directly generate summaries according to the summary-level score without additional modules and parameters. Extensive experiments demonstrate that CoLo boosts the extractive and abstractive results of one-stage systems on CNN/DailyMail benchmark to 44.58 and 46.33 ROUGE-1 score while preserving the parameter efficiency and inference efficiency. Compared with state-of-the-art multi-stage systems, we save more than 100 GPU training hours and obtaining 3x 8x speed-up ratio during inference while maintaining comparable results. Chenxin An, Ming Zhong 0005, Zhiyong Wu 0003, Xuanjing Huang 0001, Xipeng Qiu |
COLING | 5 |
| 2022 | A Progressive Framework for Role-Aware Rumor ResolutionabstractExisting works on rumor resolution have shown great potential in recognizing word appearance and user participation. However, they ignore the intrinsic propagation mechanisms of rumors and present poor adaptive ability when unprecedented news emerges. To exploit the fine-grained rumor diffusion patterns and generalize rumor resolution methods, we formulate a predecessor task to identify triggering posts, and then exploit their characteristics to facilitate rumor verification. We design a tree-structured annotation interface and extend PHEME dataset with labels on the message level. Data analysis shows that triggers play a critical role in verifying rumors and present similar lingual patterns across irrelevant events. We propose a graph-based model considering the direction and interaction of information flow to implement role-aware rumor resolution. Experimental results demonstrate the effectiveness of our proposed model and progressive scheme. Lei Chen 0082, Guanying Li, Zhongyu Wei, Yang Yang 0009, Baohua Zhou, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 7 |
| 2022 | Decorrelate Irrelevant, Purify Relevant: Overcome Textual Spurious Correlations from a Feature PerspectiveabstractNatural language understanding (NLU) models tend to rely on spurious correlations (i.e., dataset bias) to achieve high performance on in-distribution datasets but poor performance on out-of-distribution ones. Most of the existing debiasing methods often identify and weaken these samples with biased features (i.e., superficial surface features that cause such spurious correlations). However, down-weighting these samples obstructs the model in learning from the non-biased parts of these samples. To tackle this challenge, in this paper, we propose to eliminate spurious correlations in a fine-grained manner from a feature space perspective. Specifically, we introduce Random Fourier Features and weighted re-sampling to decorrelate the dependencies between features to mitigate spurious correlations. After obtaining decorrelated features, we further design a mutual-information-based method to purify them, which forces the model to learn features that are more relevant to tasks. Extensive experiments on two well-studied NLU tasks demonstrate that our method is superior to other comparative approaches. Shihan Dou, Songyang Gao, Junjie Shan, Qi Zhang 0001, Yueming Wu 0001, Xuanjing Huang 0001 |
COLING | 8 |
| 2022 | LFKQG: A Controlled Generation Framework with Local Fine-tuning for Question Generation over Knowledge BasesabstractQuestion generation over knowledge bases (KBQG) aims at generating natural questions about a subgraph, which can be answered by a given answer entity. Existing KBQG models still face two main challenges: (1) Most models often focus on the most relevant part of the answer entity, while neglecting the rest of the subgraph. (2) There are a large number of out-of-vocabulary (OOV) predicates in real-world scenarios, which are hard to adapt for most KBQG models. To address these challenges, we propose LFKQG, a controlled generation framework for Question Generation over Knowledge Bases. (1) LFKQG employs a simple controlled generation method to generate the questions containing the critical entities in the subgraph, ensuring the question is relevant to the whole subgraph. (2) We propose an optimization strategy called local fine-tuning, which can make good use of the rich information hidden in the pre-trained model to improve the ability of the model to adapt the OOV predicates. Extensive experiments show that our method outperforms existing methods significantly on three widely-used benchmark datasets SimpleQuestion, PathQuestions, and WebQuestions. Zichu Fei, Xin Zhou 0012, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 5 |
| 2022 | Improving Abstractive Dialogue Summarization with Speaker-Aware Supervised Contrastive LearningabstractPre-trained models have brought remarkable success on the text summarization task. For dialogue summarization, the subdomain of text summarization, utterances are concatenated to flat text before being processed. As a result, existing summarization systems based on pre-trained models are unable to recognize the unique format of the speaker-utterance pair well in the dialogue. To investigate this issue, we conduct probing tests and manual analysis, and find that the powerful pre-trained model can not identify different speakers well in the conversation, which leads to various factual errors. Moreover, we propose three speaker-aware supervised contrastive learning (SCL) tasks: Token-level SCL, Turn-level SCL, and Global-level SCL. Comprehensive experiments demonstrate that our methods achieve significant performance improvement on two mainstream dialogue summarization datasets. According to detailed human evaluations, pre-trained models equipped with SCL tasks effectively generate summaries with better factual consistency. Zhichao Geng, Ming Zhong 0005, Zhangyue Yin, Xipeng Qiu, Xuanjing Huang 0001 |
COLING | 5 |
| 2022 | A Structure-Aware Argument Encoder for Literature Discourse AnalysisabstractExisting research for argument representation learning mainly treats tokens in the sentence equally and ignores the implied structure information of argumentative context. In this paper, we propose to separate tokens into two groups, namely framing tokens and topic ones, to capture structural information of arguments. In addition, we consider high-level structure by incorporating paragraph-level position information. A novel structure-aware argument encoder is proposed for literature discourse analysis. Experimental results on both a self-constructed corpus and a public corpus show the effectiveness of our model. Resources are available at https://github.com/lemuria-wchen/SAE. Yinzi Li, Wei Chen 0088, Zhongyu Wei, Yujun Huang, Chujun Wang, Siyuan Wang 0025, Qi Zhang 0001, Xuanjing Huang 0001, Libo Wu |
COLING | 8 |
| 2022 | Locate Then Ask: Interpretable Stepwise Reasoning for Multi-hop Question AnsweringabstractMulti-hop reasoning requires aggregating multiple documents to answer a complex question. Existing methods usually decompose the multi-hop question into simpler single-hop questions to solve the problem for illustrating the explainable reasoning process. However, they ignore grounding on the supporting facts of each reasoning step, which tends to generate inaccurate decompositions. In this paper, we propose an interpretable stepwise reasoning framework to incorporate both single-hop supporting sentence identification and single-hop question generation at each intermediate step, and utilize the inference of the current hop for the next until reasoning out the final result. We employ a unified reader model for both intermediate hop reasoning and final hop inference and adopt joint optimization for more accurate and robust multi-hop reasoning. We conduct experiments on two benchmark datasets HotpotQA and 2WikiMultiHopQA. The results show that our method can effectively boost performance and also yields a better interpretable reasoning process without decomposition supervision. Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 5 |
| 2022 | Causal Intervention Improves Implicit Sentiment AnalysisabstractDespite having achieved great success for sentiment analysis, existing neural models struggle with implicit sentiment analysis. It is because they may latch onto spurious correlations (“shortcuts”, e.g., focusing only on explicit sentiment words), resulting in undermining the effectiveness and robustness of the learned model. In this work, we propose a CausaL intervention model for implicit sEntiment ANalysis using instrumental variable (CLEAN). We first review sentiment analysis from a causal perspective and analyze the confounders existing in this task. Then, we introduce instrumental variable to eliminate the confounding causal effects, thus extracting the pure causal effect between sentence and sentiment. We compare the proposed CLEAN with several strong baselines on both the general implicit sentiment analysis and aspect-based implicit sentiment analysis tasks. The results indicate the great advantages of our model and the efficacy of implicit sentiment reasoning. Siyin Wang, Jie Zhou 0015, Changzhi Sun, Junjie Ye 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 7 |
| 2022 | PlugAT: A Plug and Play Module to Defend against Textual Adversarial AttackabstractAdversarial training, which minimizes the loss of adversarially perturbed examples, has received considerable attention. However, these methods require modifying all model parameters and optimizing the model from scratch, which is parameter inefficient and unfriendly to the already deployed models. As an alternative, we propose a pluggable defense module PlugAT, to provide robust predictions by adding a few trainable parameters to the model inputs while keeping the original model frozen. To reduce the potential side effects of using defense modules, we further propose a novel forgetting restricted adversarial training, which filters out bad adversarial examples that impair the performance of original ones. The PlugAT-equipped BERT model substantially improves robustness over several strong baselines on various text classification tasks, whilst training only 9.1% parameters. We observe that defense modules trained under the same model architecture have domain adaptation ability between similar text classification datasets. Rong Bao, Qin Liu 0010, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Rui Xie 0005, Wei Wu 0014 |
COLING | 6 |
| 2022 | Making Parameter-efficient Tuning More Efficient: A Unified Framework for Classification TasksabstractLarge pre-trained language models (PLMs) have demonstrated superior performance in industrial applications. Recent studies have explored parameter-efficient PLM tuning, which only updates a small amount of task-specific parameters while achieving both high efficiency and comparable performance against standard fine-tuning. However, all these methods ignore the inefficiency problem caused by the task-specific output layers, which is inflexible for us to re-use PLMs and introduces non-negligible parameters. In this work, we focus on the text classification task and propose plugin-tuning, a framework that further improves the efficiency of existing parameter-efficient methods with a unified classifier. Specifically, we re-formulate both token and sentence classification tasks into a unified language modeling task, and map label spaces of different tasks into the same vocabulary space. In this way, we can directly re-use the language modeling heads of PLMs, avoiding introducing extra parameters for different tasks. We conduct experiments on six classification benchmarks. The experimental results show that plugin-tuning can achieve comparable performance against fine-tuned PLMs, while further saving around 50% parameters on top of other parameter-efficient methods. Xin Zhou 0012, Ruotian Ma, Yicheng Zou, Xuanting Chen, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Rui Xie 0005, Wei Wu 0014 |
COLING | 7 |
| 2022 | A Multi-Format Transfer Learning Model for Event Argument Extraction via Variational Information BottleneckabstractEvent argument extraction (EAE) aims to extract arguments with given roles from texts, which have been widely studied in natural language processing. Most previous works have achieved good performance in specific EAE datasets with dedicated neural architectures. Whereas, these architectures are usually difficult to adapt to new datasets/scenarios with various annotation schemas or formats. Furthermore, they rely on large-scale labeled data for training, which is unavailable due to the high labelling cost in most cases. In this paper, we propose a multi-format transfer learning model with variational information bottleneck, which makes use of the information especially the common knowledge in existing datasets for EAE in new datasets. Specifically, we introduce a shared-specific prompt framework to learn both format-shared and format-specific knowledge from datasets with different formats. In order to further absorb the common knowledge for EAE and eliminate the irrelevant noise, we integrate variational information bottleneck into our architecture to refine the shared representation. We conduct extensive experiments on three benchmark datasets, and obtain new state-of-the-art performance on EAE. Jie Zhou 0015, Qi Zhang 0001, Qin Chen 0001, Liang He 0001, Xuanjing Huang 0001 |
COLING | 5 |
| 2022 | ProofInfer: Generating Proof via Iterative Hierarchical InferenceabstractProof generation focuses on deductive reasoning: given a hypothesis and a set of theories, including some supporting facts and logical rules expressed in natural language, the model generates a proof tree indicating how to deduce the hypothesis from given theories.Current models with state-of-theart performance employ the stepwise method, linking an individual node to the proof step-bystep.However, these methods actually focus on generating several proof paths rather than a whole tree.To address this problem, we propose ProofInfer, which generates the proof tree via iterative hierarchical inference.At each step, ProofInfer generates the entire layer for proof tree, where all nodes in this layer are generated simultaneously.Since the conventional autoregressive generation architecture cannot simultaneously predict multiple nodes, ProofInfer employs text-to-text paradigm to avoid it.To this end, we propose a divideand-conquer algorithm to encode the proof tree as the plain text recursively without structure information loss.Experimental results show that ProofInfer significantly outperforms the state-of-the-art (SOTA) models on several widely-used datasets.In addition, ProofInfer still performs well with data-limited, achieving comparable performance to the SOTA models with only 40% of the training data. 1 Zichu Fei, Qi Zhang 0001, Xin Zhou 0012, Tao Gui, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2022 | Kernel-Whitening: Overcome Dataset Bias with Isotropic Sentence EmbeddingabstractDataset bias has attracted increasing attention recently for its detrimental effect on the generalization ability of fine-tuned models.The current mainstream solution is designing an additional shallow model to pre-identify biased instances.However, such two-stage methods scale up the computational complexity of training process and obstruct valid feature information while mitigating bias.To address this issue, we utilize the representation normalization method which aims at disentangling the correlations between features of encoded sentences.We find it also promising in eliminating the bias problem by providing isotropic data distribution.We further propose Kernel-Whitening, a Nyström kernel approximation method to achieve more thorough debiasing on nonlinear spurious correlations.Our framework is end-to-end with similar time consumption to fine-tuning.Experiments show that Kernel-Whitening significantly improves the performance of BERT on out-of-distribution datasets while maintaining in-distribution accuracy. Songyang Gao, Shihan Dou, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 4 |
| 2022 | BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text GenerationabstractWARNING: This paper contains examples that are offensive in nature.Automatic evaluation metrics are crucial to the development of generative systems.In recent years, pre-trained language model (PLM) based metrics, such as BERTScore (Zhang et al., 2020), have been commonly adopted in various generation tasks.However, it has been demonstrated that PLMs encode a range of stereotypical societal biases, leading to a concern on the fairness of PLMs as metrics.To that end, this work presents the first systematic study on the social bias in PLM-based metrics.We demonstrate that popular PLM-based metrics exhibit significantly higher social bias than traditional metrics on 6 sensitive attributes, namely race, gender, religion, physical appearance, age, and socioeconomic status.In-depth analysis suggests that choosing paradigms (matching, regression, or generation) of the metric has a greater impact on fairness than choosing PLMs.In addition, we develop debiasing adapters that are injected into PLM layers, mitigating bias in PLM-based metrics while retaining high performance for evaluating text generation. * Equal contribution.Example BERTScore MoverScore BARTScore BLEURT PRISM Tianxiang Sun, Junliang He, Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 4 |
| 2022 | BBTv2: Towards a Gradient-Free Future with Large Language ModelsabstractMost downstream adaptation methods tune all or part of the parameters of pre-trained models (PTMs) through gradient descent, where the tuning cost increases linearly with the growth of the model size.By contrast, gradient-free methods only require the forward computation of the PTM to tune the prompt, retaining the benefits of efficient tuning and deployment.Though, past work on gradient-free tuning often introduces gradient descent to seek a good initialization of prompt and lacks versatility across tasks and PTMs.In this paper, we present BBTv2, an improved version of Black-Box Tuning (Sun et al., 2022b), to drive PTMs for few-shot learning.We prepend continuous prompts to every layer of the PTM and propose a divide-and-conquer gradient-free algorithm to optimize the prompts at different layers alternately.Extensive experiments across various tasks and PTMs show that BBTv2 can achieve comparable performance to full model tuning and state-of-the-art parameter-efficient methods (e.g., Adapter, LoRA, BitFit, etc.) under few-shot settings while maintaining much fewer tunable parameters. Tianxiang Sun, Zhengfu He, Hong Qian, Yunhua Zhou, Xuanjing Huang 0001, Xipeng Qiu |
EMNLP | 5 |
| 2022 | Efficient Adversarial Training with Robust Early-Bird TicketsabstractAdversarial training is one of the most powerful methods to improve the robustness of pretrained language models (PLMs).However, this approach is typically more expensive than traditional fine-tuning because of the necessity to generate adversarial examples via gradient descent.Delving into the optimization process of adversarial training, we find that robust connectivity patterns emerge in the early training phase (typically 0.15 ∼ 0.3 epochs), far before parameters converge.Inspired by this finding, we dig out robust early-bird tickets (i.e., subnetworks) to develop an efficient adversarial training method: (1) searching for robust tickets with structured sparsity in the early stage; (2) fine-tuning robust tickets in the remaining time.To extract the robust tickets as early as possible, we design a ticket convergence metric to automatically terminate the searching process.Experiments show that the proposed efficient adversarial training method can achieve up to 7× ∼ 13× training speedups while maintaining comparable or even better robustness compared to the most competitive state-of-the-art adversarial training methods. Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2022 | Cross-Linguistic Syntactic Difference in Multilingual BERT: How Good is It and How Does It Affect Transfer?abstractMultilingual BERT (mBERT) has demonstrated considerable cross-lingual syntactic ability, whereby it enables effective zero-shot cross-lingual transfer of syntactic knowledge.The transfer is more successful between some languages, but it is not well understood what leads to this variation and whether it fairly reflects difference between languages.In this work, we investigate the distributions of grammatical relations induced from mBERT in the context of 24 typologically different languages.We demonstrate that the distance between the distributions of different languages is highly consistent with the syntactic difference in terms of linguistic formalisms.Such difference learnt via self-supervision plays a crucial role in the zero-shot transfer performance and can be predicted by variation in morphosyntactic properties between languages.These results suggest that mBERT properly encodes languages in a way consistent with linguistic diversity and provide insights into the mechanism of crosslingual transfer. Ningyu Xu, Tao Gui, Ruotian Ma, Qi Zhang 0001, Jingting Ye, Menghan Zhang, Xuanjing Huang 0001 |
EMNLP | 7 |
| 2022 | TextFusion: Privacy-Preserving Pre-trained Model Inference via Token FusionabstractXin Zhou, Jinzhu Lu, Tao Gui, Ruotian Ma, Zichu Fei, Yuran Wang, Yong Ding, Yibo Cheung, Qi Zhang, Xuanjing Huang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xin Zhou 0012, Jinzhu Lu, Tao Gui, Ruotian Ma, Zichu Fei, Yibo Cheung, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 10 |
| 2022 | Black-Box Tuning for Language-Model-as-a-ServiceabstractExtremely large pre-trained language models (PTMs) such as GPT-3 are usually released as a service. It allows users to design task-specific prompts to query the PTMs through some black-box APIs. In such a scenario, which we call Language-Model-as-a-Service (LMaaS), the gradients of PTMs are usually unavailable. Can we optimize the task prompts by only accessing the model inference APIs? This paper proposes the black-box tuning framework to optimize the continuous prompt prepended to the input text via derivative-free optimization. Instead of optimizing in the original high-dimensional prompt space, which is intractable for traditional derivative-free optimization, we perform optimization in a randomly generated subspace due to the low intrinsic dimensionality of large PTMs. The experimental results show that the black-box tuning with RoBERTa on a few labeled samples not only significantly outperforms manual prompt and GPT-3’s in-context learning, but also surpasses the gradient-based counterparts, i.e., prompt tuning and full model tuning. Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang 0001, Xipeng Qiu |
ICML | 4 |
| 2022 | What Dense Graph Do You Need for Self-Attention?abstractTransformers have made progress in miscellaneous tasks, but suffer from quadratic computational and memory complexities. Recent works propose sparse transformers with attention on sparse graphs to reduce complexity and remain strong performance. While effective, the crucial parts of how dense a graph needs to be to perform well are not fully explored. In this paper, we propose Normalized Information Payload (NIP), a graph scoring function measuring information transfer on graph, which provides an analysis tool for trade-offs between performance and complexity. Guided by this theoretical analysis, we present Hypercube Transformer, a sparse transformer that models token interactions in a hypercube and shows comparable or even better results with vanilla transformer while yielding $O(N\log N)$ complexity with sequence length $N$. Experiments on tasks requiring various sequence lengths lay validation for our graph function well. Yuxin Wang 0005, Chu-Tak Lee, Qipeng Guo, Zhangyue Yin, Yunhua Zhou, Xuanjing Huang 0001, Xipeng Qiu |
ICML | 6 |
| 2022 | Constructing Phrase-level Semantic Labels to Form Multi-Grained Supervision for Image-Text RetrievalabstractExisting research for image text retrieval mainly relies on sentence-level supervision to distinguish matched and mismatched sentences for a query image. However, semantic mismatch between an image and sentences usually happens in finer grain, i.e., phrase level. In this paper, we explore to introduce additional phrase-level supervision for the better identification of mismatched units in the text. In practice, multi-grained semantic labels are automatically constructed for a query image in both sentence-level and phrase-level. We construct text scene graphs for the matched sentences and extract entities and triples as the phrase-level labels. In order to integrate both supervision of sentence-level and phrase-level, we propose Semantic Structure Aware Multimodal Transformer (SSAMT) for multi-modal representation learning. Inside the SSAMT, we utilize different kinds of attention mechanisms to enforce interactions of multi-grained semantic units in both sides of vision and language. For the training, we propose multi-scale matching from both global and local perspectives, and penalize mismatched phrases. Experimental results on MS-COCO and Flickr30K show the effectiveness of our approach compared to some state-of-the-art models. Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Haijun Shan, Xuanjing Huang 0001, Jianqing Fan |
ICMR | 6 |
| 2022 | MVPTR: Multi-Level Semantic Alignment for Vision-Language Pre-Training via Multi-Stage LearningabstractPrevious vision-language pre-training models mainly construct multi-modal inputs with tokens and objects (pixels) followed by performing cross-modality interaction between them. We argue that the input of only tokens and object features limits high-level semantic alignment like phrase-to-region grounding. Meanwhile, multi-level alignments are inherently consistent and able to facilitate the representation learning synergistically. Therefore, in this paper, we propose to learn Multi-level semantic alignment for Vision-language Pre-TRaining (MVPTR). In MVPTR, we follow the nested structure of both modalities to introduce concepts as high-level semantics. To ease the learning from multi-modal multi-level inputs, our framework is split into two stages, the first stage focuses on intra-modality multi-level representation learning, the second enforces interactions across modalities via both coarse-grained and fine-grained semantic alignment tasks. In addition to the commonly used image-text matching and masked language model tasks, we introduce a masked concept recovering task in the first stage to enhance the concept representation learning, and two more tasks in the second stage to explicitly encourage multi-level alignments across modalities. Our model achieves state-of-the-art results on several vision and language tasks. Zhihao Fan, Huaixiao Tou, Jingjing Chen 0001, Zhongyu Wei, Xuanjing Huang 0001 |
ACM Multimedia | 6 |
| 2022 | Towards Efficient NLP: A Standard Evaluation and A Strong BaselineabstractXiangyang Liu, Tianxiang Sun, Junliang He, Jiawen Wu, Lingling Wu, Xinyu Zhang, Hao Jiang, Zhao Cao, Xuanjing Huang, Xipeng Qiu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Tianxiang Sun, Junliang He, Jiawen Wu 0002, Lingling Wu, Xinyu Zhang 0019, Hao Jiang 0022, Zhao Cao, Xuanjing Huang 0001, Xipeng Qiu |
NAACL-HLT | 9 |
| 2022 | Template-free Prompt Tuning for Few-shot NERabstractRuotian Ma, Xin Zhou, Tao Gui, Yiding Tan, Linyang Li, Qi Zhang, Xuanjing Huang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ruotian Ma, Xin Zhou 0012, Tao Gui, Yiding Tan, Linyang Li, Qi Zhang 0001, Xuanjing Huang 0001 |
NAACL-HLT | 7 |
| 2022 | CoNT: Contrastive Neural Text GenerationabstractRecently, contrastive learning attracts increasing interests in neural text generation as a new solution to alleviate the exposure bias problem. It introduces a sequence-level training signal which is crucial to generation tasks that always rely on auto-regressive decoding. However, previous methods using contrastive learning in neural text generation usually lead to inferior performance. In this paper, we analyse the underlying reasons and propose a new Contrastive Neural Text generation framework, CoNT. CoNT addresses bottlenecks that prevent contrastive learning from being widely adopted in generation tasks from three aspects -- the construction of contrastive examples, the choice of the contrastive loss, and the strategy in decoding. We validate CoNT on five generation tasks with ten benchmarks, including machine translation, summarization, code comment generation, data-to-text generation and commonsense generation. Experimental results show that CoNT clearly outperforms its baseline on all the ten benchmarks with a convincing margin. Especially, CoNT surpasses previous the most competitive contrastive learning method for text generation, by 1.50 BLEU on machine translation and 1.77 ROUGE-1 on summarization, respectively. It achieves new state-of-the-art on summarization, code comment generation (without external data) and data-to-text generation. Chenxin An, Jiangtao Feng, Kai Lv 0001, Lingpeng Kong, Xipeng Qiu, Xuanjing Huang 0001 |
NeurIPS | 6 |
| 2022 | Regularized Molecular Conformation FieldsabstractPredicting energetically favorable 3-dimensional conformations of organic molecules frommolecular graph plays a fundamental role in computer-aided drug discovery research.However, effectively exploring the high-dimensional conformation space to identify (meta) stable conformers is anything but trivial.In this work, we introduce RMCF, a novel framework to generate a diverse set of low-energy molecular conformations through samplingfrom a regularized molecular conformation field.We develop a data-driven molecular segmentation algorithm to automatically partition each molecule into several structural building blocks to reduce the modeling degrees of freedom.Then, we employ a Markov Random Field to learn the joint probability distribution of fragment configurations and inter-fragment dihedral angles, which enables us to sample from different low-energy regions of a conformation space.Our model constantly outperforms state-of-the-art models for the conformation generation task on the GEOM-Drugs dataset.We attribute the success of RMCF to modeling in a regularized feature space and learning a global fragment configuration distribution for effective sampling.The proposed method could be generalized to deal with larger biomolecular systems. Yi Zhou 0018, Xiaoqing Zheng, Xuanjing Huang 0001, Hao Zhou 0012 |
NeurIPS | 5 |
| 2022 | BART-Reader: Predicting Relations Between Entities via Reading Their Document-Level Context Information
Hang Yan 0001, Yu Sun 0031, Junqi Dai, Xiangkun Hu, Qipeng Guo, Xipeng Qiu, Xuanjing Huang 0001 |
NLPCC (1) | 7 |
| 2022 | Hierarchical reinforcement learning for automatic disease diagnosisabstractMOTIVATION: Disease diagnosis-oriented dialog system models the interactive consultation procedure as the Markov decision process, and reinforcement learning algorithms are used to solve the problem. Existing approaches usually employ a flat policy structure that treat all symptoms and diseases equally for action making. This strategy works well in a simple scenario when the action space is small; however, its efficiency will be challenged in the real environment. Inspired by the offline consultation process, we propose to integrate a hierarchical policy structure of two levels into the dialog system for policy learning. The high-level policy consists of a master model that is responsible for triggering a low-level model, the low-level policy consists of several symptom checkers and a disease classifier. The proposed policy structure is capable to deal with diagnosis problem including large number of diseases and symptoms. RESULTS: Experimental results on three real-world datasets and a synthetic dataset demonstrate that our hierarchical framework achieves higher accuracy and symptom recall in disease diagnosis compared with existing systems. We construct a benchmark including datasets and implementation of existing algorithms to encourage follow-up researches. AVAILABILITY AND IMPLEMENTATION: The code and data are available from https://github.com/FudanDISC/DISCOpen-MedBox-DialoDiagnosis. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kangenbei Liao, Wei Chen 0088, Qianlong Liu, Baolin Peng, Xuanjing Huang 0001, Jiajie Peng, Zhongyu Wei |
Bioinform. | 6 |
| 2022 | Aspect-based sentiment analysis with enhanced aspect-sensitive word embeddings
Yusi Qi, Xiaoqing Zheng, Xuanjing Huang 0001 |
Knowl. Inf. Syst. | 3 |
| 2022 | Sentiment-aware multimodal pre-training for multimodal sentiment analysis
Junjie Ye 0005, Jie Zhou 0015, Rui Wang 0005, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
Knowl. Based Syst. | 8 |
| 2022 | Automatic Math Word Problem Generation With Topic-Expression Co-Attention Mechanism and Reinforcement LearningabstractTaking several topic words and a math expression as input, the aim of math word problem generation is to generate a problem that can be answered by the given expression and related to these topic words. Considerable progress has been achieved by sequence-to-sequence neural network models in many natural language generation tasks, but these models do not effectively consider the characteristics of the math word problem generation task. They may generate problems that are unrelated to the topic words and expressions, and problems that cannot be solved. In this paper, we propose a new model, MWPGen, for automatically generating math word problems. MWPGen has a topic-expression co-attention mechanism to extract relevant information between topic words and expressions. Further, we fine-tune MWPGen with the solving result of the generated problem as the reward for reinforcement learning. MWPGen shows improved performance in popular automatic evaluation metrics and improves the solvability of generated problems. Qinzhuo Wu, Qi Zhang 0001, Xuanjing Huang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Enhancing Scientific Papers Summarization with Citation GraphabstractPrevious work for text summarization in scientific domain mainly focused on the content of the input document, but seldom considering its citation network. However, scientific papers are full of uncommon domain-specific terms, making it almost impossible for the model to understand its true meaning without the help of the relevant research community. In this paper, we redefine the task of scientific papers summarization by utilizing their citation graph and propose a citation graph-based summarization model CGSum which can incorporate the information of both the source paper and its references. In addition, we construct a novel scientific papers summarization dataset Semantic Scholar Network (SSN) which contains 141K research papers in different domains and 661K citation relationships. The entire dataset constitutes a large connected citation graph. Extensive experiments show that our model can achieve competitive performance when compared with the pretrained models even with a simple architecture. The results also indicates the citation graph is crucial to better understand the content of papers and generate high-quality summaries. Chenxin An, Ming Zhong 0005, Yiran Chen 0013, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001 |
AAAI | 6 |
| 2021 | An Unsupervised Sampling Approach for Image-Sentence Matching Using Document-level Structural Information
Zhongyu Wei, Zhihao Fan, Haijun Shan, Xuanjing Huang 0001 |
AAAI | 5 |
| 2021 | Unsupervised Summarization for Chat Logs with Topic-Oriented Ranking and Context-Aware Auto-EncodersabstractAutomatic chat summarization can help people quickly grasp important information from numerous chat messages. Unlike conventional documents, chat logs usually have fragmented and evolving topics. In addition, these logs contain a quantity of elliptical and interrogative sentences, which make the chat summarization highly context dependent. In this work, we propose a novel unsupervised framework called RankAE to perform chat summarization without employing manually labeled data. RankAE consists of a topic-oriented ranking strategy that selects topic utterances according to centrality and diversity simultaneously, as well as a denoising auto-encoder that is carefully designed to generate succinct but context-informative summaries based on the selected utterances. To evaluate the proposed method, we collect a large-scale dataset of chat logs from a customer service environment and build an annotated set only for model evaluation. Experimental results show that RankAE significantly outperforms other unsupervised methods and is able to generate high-quality summaries in terms of relevance and topic coverage. Yicheng Zou, Lujun Zhao, Yangyang Kang, Zhuoren Jiang, Changlong Sun, Qi Zhang 0001, Xuanjing Huang 0001, Xiaozhong Liu 0001 |
AAAI | 8 |
| 2021 | Topic-Oriented Spoken Dialogue Summarization for Customer Service with Saliency-Aware Topic ModelingabstractIn a customer service system, dialogue summarization can boost service efficiency by automatically creating summaries for long spoken dialogues in which customers and agents try to address issues about specific topics. In this work, we focus on topic-oriented dialogue summarization, which generates highly abstractive summaries that preserve the main ideas from dialogues. In spoken dialogues, abundant dialogue noise and common semantics could obscure the underlying informative content, making the general topic modeling approaches difficult to apply. In addition, for customer service, role-specific information matters and is an indispensable part of a summary. To effectively perform topic modeling on dialogues and capture multi-role information, in this work we propose a novel topic-augmented two-stage dialogue summarizer (TDS) jointly with a saliency-aware neural topic model (SATM) for topic-oriented summarization of customer service dialogues. Comprehensive studies on a real-world Chinese customer service dataset demonstrated the superiority of our method against several strong baselines. Yicheng Zou, Lujun Zhao, Yangyang Kang, Minlong Peng, Zhuoren Jiang, Changlong Sun, Qi Zhang 0001, Xuanjing Huang 0001, Xiaozhong Liu 0001 |
AAAI | 9 |
| 2021 | SpanNER: Named Entity Re-/Recognition as Span PredictionabstractJinlan Fu, Xuanjing Huang, Pengfei Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jinlan Fu, Xuanjing Huang 0001, Pengfei Liu 0003 |
ACL/IJCNLP (1) | 2 |
| 2021 | Accelerating BERT Inference for Sequence Labeling via Early-ExitabstractXiaonan Li, Yunfan Shao, Tianxiang Sun, Hang Yan, Xipeng Qiu, Xuanjing Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yunfan Shao, Tianxiang Sun, Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001 |
ACL/IJCNLP (1) | 6 |
| 2021 | SENT: Sentence-level Distant Relation Extraction via Negative TrainingabstractRuotian Ma, Tao Gui, Linyang Li, Qi Zhang, Xuanjing Huang, Yaqian Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruotian Ma, Tao Gui, Linyang Li, Qi Zhang 0001, Xuanjing Huang 0001, Yaqian Zhou 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | Align Voting Behavior with Public Statements for Legislator Representation LearningabstractXinyi Mou, Zhongyu Wei, Lei Chen, Shangyi Ning, Yancheng He, Changjian Jiang, Xuanjing Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinyi Mou, Zhongyu Wei, Lei Chen 0082, Shangyi Ning, Yancheng He, Changjian Jiang, Xuanjing Huang 0001 |
ACL/IJCNLP (1) | 7 |
| 2021 | Math Word Problem Solving with Explicit Numerical ValuesabstractQinzhuo Wu, Qi Zhang, Zhongyu Wei, Xuanjing Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Qinzhuo Wu, Qi Zhang 0001, Zhongyu Wei, Xuanjing Huang 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Defense against Synonym Substitution-based Adversarial Attacks via Dirichlet Neighborhood EnsembleabstractYi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-Wei Chang, Xuanjing Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yi Zhou 0018, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-Wei Chang 0001, Xuanjing Huang 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | Fine-Grained Element Identification in Complaint Text of Internet FraudabstractExisting system dealing with online complaint provides a final decision without explanations. We propose to analyse the complaint text of internet fraud in a fine-grained manner. Considering the complaint text includes multiple clauses with various functions, we propose to identify the role of each clause and classify them into different types of fraud element. We construct a large labeled dataset originated from a real finance service platform. We build an element identification model on top of BERT and propose additional two modules to utilize the context of complaint text for better element label classification, namely, global context encoder and label refiner. Experimental results show the effectiveness of our model. Siyuan Wang 0025, Jingchao Fu, Lei Chen 0082, Zhongyu Wei, Heng Ye, Liaosa Xu, Weiqiang Wang 0002, Xuanjing Huang 0001 |
CIKM | 10 |
| 2021 | TCIC: Theme Concepts Learning Cross Language and Vision for Image CaptioningabstractExisting research for image captioning usually represents an image using a scene graph with low-level facts (objects and relations) and fails to capture the high-level semantics. In this paper, we propose a Theme Concepts extended Image Captioning (TCIC) framework that incorporates theme concepts to represent high-level cross-modality semantics. In practice, we model theme concepts as memory vectors and propose Transformer with Theme Nodes (TTN) to incorporate those vectors for image captioning. Considering that theme concepts can be learned from both images and captions, we propose two settings for their representations learning based on TTN. On the vision side, TTN is configured to take both scene graph based features and theme concepts as input for visual representation learning. On the language side, TTN is configured to take both captions and theme concepts as input for text representation re-construction. Both settings aim to generate target captions with the same transformer-based decoder. During the training, we further align representations of theme concepts learned from images and corresponding captions to enforce the cross-modality learning. Experimental results on MS COCO show the effectiveness of our approach compared to some state-of-the-art models. Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Haijun Shan, Xuanjing Huang 0001 |
IJCAI | 7 |
| 2021 | Mask Attention Networks: Rethinking and Strengthen TransformerabstractZhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, Xuanjing Huang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang 0025, Jian Jiao 0007, Nan Duan 0001, Ruofei Zhang, Xuanjing Huang 0001 |
NAACL-HLT | 9 |
| 2021 | Larger-Context Tagging: When and Why Does It Work?abstractJinlan Fu, Liangjing Feng, Qi Zhang, Xuanjing Huang, Pengfei Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jinlan Fu, Liangjing Feng, Qi Zhang 0001, Xuanjing Huang 0001, Pengfei Liu 0003 |
NAACL-HLT | 4 |
| 2021 | Discrete Argument Representation Learning for Interactive Argument Pair IdentificationabstractLu Ji, Zhongyu Wei, Jing Li, Qi Zhang, Xuanjing Huang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Lu Ji, Zhongyu Wei, Jing Li 0049, Qi Zhang 0001, Xuanjing Huang 0001 |
NAACL-HLT | 5 |
| 2021 | Searching Effective Transformer for Seq2Seq Keyphrase Generation
Yige Xu 0001, Yichao Luo, Zhengyan Li, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001 |
NLPCC (2) | 7 |
| 2021 | Overview of Argumentative Text Understanding for AI Debater Challenge
Liying Cheng, Ruidan He, Yinzi Li, Lidong Bing, Zhongyu Wei, Qin Liu 0010, Chenhui Shen, Shuonan Zhang, Changlong Sun, Luo Si, Changjian Jiang, Xuanjing Huang 0001 |
NLPCC (2) | 13 |
| 2021 | Text information aggregation with centrality attention
Jingjing Gong, Hang Yan 0001, Yining Zheng, Qipeng Guo, Xipeng Qiu, Xuanjing Huang 0001 |
Sci. China Inf. Sci. | 6 |
| 2021 | Information retrieval: a view from the Chinese IR community
Zhumin Chen, Xueqi Cheng 0001, Shoubin Dong, Zhicheng Dou, Jiafeng Guo, Xuanjing Huang 0001, Yanyan Lan, Chenliang Li 0005, Ru Li 0001, Tie-Yan Liu, Yiqun Liu 0001, Jun Ma 0001, Bing Qin 0001, Mingwen Wang 0001, Ji-Rong Wen, Jun Xu 0001, Min Zhang 0006, Peng Zhang 0002, Qi Zhang 0001 |
Frontiers Comput. Sci. | 6 |
| 2021 | Jointly learning bilingual word embeddings and alignments
Zhenqiao Song, Xiaoqing Zheng, Xuanjing Huang 0001 |
Mach. Transl. | 3 |
| 2021 | Generating Responses With a Given Syntactic Pattern in Chinese DialoguesabstractRecently, many efforts have been devoted to generating responses expressing a specific emotion or relating to a given topic in a controlled manner. However, limited attention has been given to generating responses with a specified syntactic pattern, which makes it possible to imitate someone's way of speaking in dialogue. To fulfill this goal, we propose two models to generate syntax-aware responses: a gross-constraint and a specific-constraint model. The former controls the syntactic patterns of generated responses at sentence-level, while the latter works at smaller language units, such as words or phrases, being capable of manipulating the syntactic structures of responses in a more subtle manner. The extensive experimental results on two different datasets show that both the two models not only can generate meaningful responses with a specific and coherent structure but also improve on the diversity of generated responses, with similar gains in readability, relevance, and diversity as measured by human judges. Yi Zhou 0018, Xiaoqing Zheng, Xuanjing Huang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Co-Attention Memory Network for Multimodal Microblog's Hashtag RecommendationabstractHashtags are keywords describing a topic or a theme and are usually chosen by microblogging users. Hence, the hashtags can be used to categorize microblog posts. With the fast development of the social network, the task of recommending suitable hashtags has received considerable attention in recent years. Recently, most neural network methods have treated the task as a multi-class classification problem. In fact, users are constantly introducing new hashtags in a highly dynamic way. Treating the task as a multi-class classification problem with a fixed number of target categories does not allow the method to deal with the new hashtags. To address this problem, the task is reinterpreted as a matching problem and a novel co-attention memory network is proposed to represent the multimodal microblogs and hashtags. We utilize a co-attention mechanism to model the multimodal mircroblogs, and utilize the post history to represent the hashtags. Experimental results on a Twitter-based dataset demonstrated that the proposed method can achieve better performance than the current state-of-the-art methods that treat the task as a multi-class classification problem. Renfeng Ma, Xipeng Qiu, Qi Zhang 0001, Xiangkun Hu, Yu-Gang Jiang 0001, Xuanjing Huang 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2020 | Constructing Multiple Tasks for Augmentation: Improving Neural Image Classification with K-Means FeaturesabstractMulti-task learning (MTL) has received considerable attention, and numerous deep learning applications benefit from MTL with multiple objectives. However, constructing multiple related tasks is difficult, and sometimes only a single task is available for training in a dataset. To tackle this problem, we explored the idea of using unsupervised clustering to construct a variety of auxiliary tasks from unlabeled data or existing labeled data. We found that some of these newly constructed tasks could exhibit semantic meanings corresponding to certain human-specific attributes, but some were non-ideal. In order to effectively reduce the impact of non-ideal auxiliary tasks on the main task, we further proposed a novel meta-learning-based multi-task learning approach, which trained the shared hidden layers on auxiliary tasks, while the meta-optimization objective was to minimize the loss on the main task, ensuring that the optimizing direction led to an improvement on the main task. Experimental results across five image datasets demonstrated that the proposed method significantly outperformed existing single task learning, semi-supervised learning, and some data augmentation methods, including an improvement of more than 9% on the Omniglot dataset. Tao Gui, Lizhi Qing, Qi Zhang 0001, Jiacheng Ye, Hang Yan 0001, Zichu Fei, Xuanjing Huang 0001 |
AAAI | 7 |
| 2020 | Learning Sparse Sharing Architectures for Multiple TasksabstractMost existing deep multi-task learning models are based on parameter sharing, such as hard sharing, hierarchical sharing, and soft sharing. How choosing a suitable sharing mechanism depends on the relations among the tasks, which is not easy since it is difficult to understand the underlying shared factors among these tasks. In this paper, we propose a novel parameter sharing mechanism, named Sparse Sharing. Given multiple tasks, our approach automatically finds a sparse sharing structure. We start with an over-parameterized base network, from which each task extracts a subnetwork. The subnetworks of multiple tasks are partially overlapped and trained in parallel. We show that both hard sharing and hierarchical sharing can be formulated as particular instances of the sparse sharing framework. We conduct extensive experiments on three sequence labeling tasks. Compared with single-task models and three typical multi-task learning baselines, our proposed approach achieves consistent improvement while requiring fewer parameters. Tianxiang Sun, Yunfan Shao, Pengfei Liu 0003, Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001 |
AAAI | 7 |
| 2020 | Storytelling from an Image Stream Using Scene GraphsabstractVisual storytelling aims at generating a story from an image stream. Most existing methods tend to represent images directly with the extracted high-level features, which is not intuitive and difficult to interpret. We argue that translating each image into a graph-based semantic representation, i.e., scene graph, which explicitly encodes the objects and relationships detected within image, would benefit representing and describing images. To this end, we propose a novel graph-based architecture for visual storytelling by modeling the two-level relationships on scene graphs. In particular, on the within-image level, we employ a Graph Convolution Network (GCN) to enrich local fine-grained region representations of objects on scene graphs. To further model the interaction among images, on the cross-images level, a Temporal Convolution Network (TCN) is utilized to refine the region representations along the temporal dimension. Then the relation-aware representations are fed into the Gated Recurrent Unit (GRU) with attention mechanism for story generation. Experiments are conducted on the public visual storytelling dataset. Automatic and human evaluation results indicate that our method achieves state-of-the-art. Zhongyu Wei, Piji Li, Qi Zhang 0001, Xuanjing Huang 0001 |
AAAI | 5 |
| 2020 | FLAT: Chinese NER Using Flat-Lattice TransformerabstractRecently, the character-word lattice structure has been proved to be effective for Chinese named entity recognition (NER) by incorporating the word information.However, since the lattice structure is complex and dynamic, most existing lattice-based models are hard to fully utilize the parallel computation of GPUs and usually have a low inference-speed.In this paper, we propose FLAT: Flat-LAttice Transformer for Chinese NER, which converts the lattice structure into a flat structure consisting of spans.Each span corresponds to a character or latent word and its position in the original lattice.With the power of Transformer and well-designed position encoding, FLAT can fully leverage the lattice information and has an excellent parallelization ability.Experiments on four datasets show FLAT outperforms other lexicon-based models in performance and efficiency. Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001 |
ACL | 4 |
| 2020 | Simplify the Usage of Lexicon in Chinese NERabstractRecently, many works have tried to augment the performance of Chinese named entity recognition (NER) using word lexicons.As a representative, Lattice-LSTM (Zhang and Yang, 2018) has achieved new benchmark results on several public Chinese NER datasets.However, Lattice-LSTM has a complex model architecture.This limits its application in many industrial areas where real-time NER responses are needed.In this work, we propose a simple but effective method for incorporating the word lexicon into the character representations.This method avoids designing a complicated sequence modeling architecture, and for any neural NER model, it requires only subtle adjustment of the character representation layer to introduce the lexicon information.Experimental studies on four benchmark Chinese NER datasets show that our method achieves an inference speed up to 6.15 times faster than those of state-ofthe-art methods, along with a better performance.The experimental results also show that the proposed method can be easily incorporated with pre-trained models like BERT. 1 * Equal contribution. Ruotian Ma, Minlong Peng, Qi Zhang 0001, Zhongyu Wei, Xuanjing Huang 0001 |
ACL | 5 |
| 2020 | Heterogeneous Graph Neural Networks for Extractive Document SummarizationabstractAs a crucial step in extractive document summarization, learning cross-sentence relations has been explored by a plethora of approaches.An intuitive way is to put them in the graphbased neural network, which has a more complex structure for capturing inter-sentence relationships.In this paper, we present a heterogeneous graph-based neural network for extractive summarization (HETERSUMGRAPH), which contains semantic nodes of different granularity levels apart from sentences.These additional nodes act as the intermediary between sentences and enrich the cross-sentence relations.Besides, our graph structure is flexible in natural extension from a singledocument setting to multi-document via introducing document nodes.To our knowledge, we are the first one to introduce different types of nodes into graph-based neural networks for extractive document summarization and perform a comprehensive qualitative analysis to investigate their benefits.The code will be released on Github 1 . Danqing Wang, Pengfei Liu 0003, Yining Zheng, Xipeng Qiu, Xuanjing Huang 0001 |
ACL | 5 |
| 2020 | Evaluating and Enhancing the Robustness of Neural Network-based Dependency Parsing Models with Adversarial ExamplesabstractDespite achieving prominent performance on many important tasks, it has been reported that neural networks are vulnerable to adversarial examples.Previously studies along this line mainly focused on semantic tasks such as sentiment analysis, question answering and reading comprehension.In this study, we show that adversarial examples also exist in dependency parsing: we propose two approaches to study where and how parsers make mistakes by searching over perturbations to existing texts at sentence and phrase levels, and design algorithms to construct such examples in both of the black-box and white-box settings.Our experiments with one of state-of-the-art parsers on the English Penn Treebank (PTB) show that up to 77% of input examples admit adversarial perturbations, and we also show that the robustness of parsing models can be improved by crafting high-quality adversaries and including them in the training stage, while suffering little to no performance drop on the clean input data. Xiaoqing Zheng, Jiehang Zeng, Yi Zhou 0018, Cho-Jui Hsieh, Minhao Cheng, Xuanjing Huang 0001 |
ACL | 6 |
| 2020 | Extractive Summarization as Text MatchingabstractThis paper creates a paradigm shift with regard to the way we build neural extractive summarization systems.Instead of following the commonly used framework of extracting sentences individually and modeling the relationship between sentences, we formulate the extractive summarization task as a semantic text matching problem, in which a source document and candidate summaries will be (extracted from the original text) matched in a semantic space.Notably, this paradigm shift to semantic matching framework is well-grounded in our comprehensive analysis of the inherent gap between sentence-level and summary-level extractors based on the property of the dataset.Besides, even instantiating the framework with a simple form of a matching model, we have driven the state-of-the-art extractive result on CNN/DailyMail to a new level (44.41 in ROUGE-1).Experiments on the other five datasets also show the effectiveness of the matching framework.We believe the power of this matching-based summarization framework has not been fully exploited.To encourage more instantiations in the future, we have released our codes, processed dataset, as well as generated summaries in https://github. com/maszhongming/MatchSum. Ming Zhong 0005, Pengfei Liu 0003, Yiran Chen 0013, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001 |
ACL | 6 |
| 2020 | Modeling Evolution of Message Interaction for Rumor ResolutionabstractPrevious work for rumor resolution concentrates on exploiting time-series characteristics or modeling topology structure separately.However, how local interactive pattern affects global information assemblage has not been explored.In this paper, we attempt to address the problem by learning evolution of message interaction.We model confrontation and reciprocity between message pairs via discrete variational autoencoders which effectively reflects the diversified opinion interactivity.Moreover, we capture the variation of message interaction using a hierarchical framework to better integrate information flow of a rumor cascade.Experiments on PHEME dataset demonstrate our proposed model achieves higher accuracy than existing methods. Lei Chen 0082, Zhongyu Wei, Jing Li 0049, Baohua Zhou, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 6 |
| 2020 | An Enhanced Knowledge Injection Model for Commonsense GenerationabstractCommonsense generation aims at generating plausible everyday scenario description based on a set of provided concepts. Digging the relationship of concepts from scratch is non-trivial, therefore, we retrieve prototypes from external knowledge to assist the understanding of the scenario for better description generation. We integrate two additional modules into the pretrained encoder-decoder model for prototype modeling to enhance the knowledge injection procedure. We conduct experiment on CommonGen benchmark, experimental results show that our method significantly improves the performance on all the metrics. Zhihao Fan, Yeyun Gong, Zhongyu Wei, Siyuan Wang 0025, Yameng Huang, Jian Jiao 0007, Xuanjing Huang 0001, Nan Duan 0001, Ruofei Zhang |
COLING | 7 |
| 2020 | CoLAKE: Contextualized Language and Knowledge EmbeddingabstractWith the emerging branch of incorporating factual knowledge into pre-trained language models such as BERT, most existing models consider shallow, static, and separately pre-trained entity embeddings, which limits the performance gains of these models.Few works explore the potential of deep contextualized knowledge representation when injecting knowledge.In this paper, we propose the Contextualized Language and Knowledge Embedding (CoLAKE), which jointly learns contextualized representation for both language and knowledge with the extended MLM objective.Instead of injecting only entity embeddings, CoLAKE extracts the knowledge context of an entity from large-scale knowledge bases.To handle the heterogeneity of knowledge context and language context, we integrate them in a unified data structure, word-knowledge graph (WK graph).CoLAKE is pre-trained on large-scale WK graphs with the modified Transformer encoder.We conduct experiments on knowledge-driven tasks, knowledge probing tasks, and language understanding tasks.Experimental results show that CoLAKE outperforms previous counterparts on most of the tasks.Besides, CoLAKE achieves surprisingly high performance on our synthetic task called word-knowledge graph completion, which shows the superiority of simultaneously contextualizing language and knowledge representation. 1 Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuanjing Huang 0001, Zheng Zhang 0001 |
COLING | 6 |
| 2020 | Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative Multi-agent CommunicationabstractVisual storytelling aims to generate a narrative paragraph from a sequence of images automatically.Existing approaches construct text description independently for each image and roughly concatenate them as a story, which leads to the problem of generating semantically incoherent content.In this paper, we propose a new way for visual storytelling by introducing a topic description task to detect the global semantic context of an image stream.A story is then constructed with the guidance of the topic description.In order to combine the two generation tasks, we propose a multi-agent communication framework that regards the topic description generator and the story generator as two agents and learn them simultaneously via iterative updating mechanism.We validate our approach on VIST dataset, where quantitative results, ablations, and human evaluation demonstrate our method's good ability in generating stories with higher quality compared to state-of-the-art methods. Zhongyu Wei, Ying Cheng 0005, Piji Li, Haijun Shan, Ji Zhang 0011, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 8 |
| 2020 | RethinkCWS: Is Chinese Word Segmentation a Solved Task?abstractThe performance of the Chinese Word Segmentation (CWS) systems has gradually reached a plateau with the rapid development of deep neural networks, especially the successful use of large pre-trained models.In this paper, we take stock of what we have achieved and rethink what's left in the CWS task.Methodologically, we propose a finegrained evaluation for existing CWS systems, which not only allows us to diagnose the strengths and weaknesses of existing models (under the in-dataset setting), but enables us to quantify the discrepancy between different criterion and alleviate the negative transfer problem when doing multi-criteria learning.Strategically, despite not aiming to propose a novel model in this paper, our comprehensive experiments on eight models and seven datasets, as well as thorough analysis, could search for some promising direction for future research.We make all codes publicly available and release an interface that can quickly evaluate and diagnose user's models: https://github. com/neulab/InterpretEval. Jinlan Fu, Pengfei Liu 0003, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP (1) | 4 |
| 2020 | Uncertainty-Aware Label Refinement for Sequence LabelingabstractConditional random fields (CRF) for label decoding has become ubiquitous in sequence labeling tasks.However, the local label dependencies and inefficient Viterbi decoding have always been a problem to be solved.In this work, we introduce a novel two-stage label decoding framework to model long-term label dependencies, while being much more computationally efficient.A base model first predicts draft labels, and then a novel twostream self-attention model makes refinements on these draft predictions based on longrange label dependencies, which can achieve parallel decoding for a faster prediction.In addition, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with a high probability of being wrong, which can greatly assist in preventing error propagation.The experimental results on three sequence labeling benchmarks demonstrated that the proposed method not only outperformed the CRF-based methods but also greatly accelerated the inference process.* Both authors contributed equally. Tao Gui, Jiacheng Ye, Qi Zhang 0001, Zhengyan Li, Zichu Fei, Yeyun Gong, Xuanjing Huang 0001 |
EMNLP (1) | 7 |
| 2020 | Leveraging Declarative Knowledge in Text and First-Order Logic for Fine-Grained Propaganda DetectionabstractRuize Wang, Duyu Tang, Nan Duan, Wanjun Zhong, Zhongyu Wei, Xuanjing Huang, Daxin Jiang, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Duyu Tang, Nan Duan 0001, Wanjun Zhong, Zhongyu Wei, Xuanjing Huang 0001, Daxin Jiang, Ming Zhou 0001 |
EMNLP (1) | 6 |
| 2020 | PathQG: Neural Question Generation from FactsabstractExisting research for question generation encodes the input text as a sequence of tokens without explicitly modeling fact information.These models tend to generate irrelevant and uninformative questions.In this paper, we explore to incorporate facts in the text for question generation in a comprehensive way.We present a novel task of question generation given a query path in the knowledge graph constructed from the input text.We divide the task into two steps, namely, query representation learning and query-based question generation.We formulate query representation learning as a sequence labeling problem for identifying the involved facts to form a query and employ an RNN-based generator for question generation.We first train the two modules jointly in an end-to-end fashion, and further enforce the interaction between these two modules in a variational framework.We construct the experimental datasets on top of SQuAD and results show that our model outperforms other state-of-the-art approaches, and the performance margin is larger when target questions are complex.Human evaluation also proves that our model is able to generate relevant and informative questions. 1 Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Zengfeng Huang, Weijian Sun, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP (1) | 7 |
| 2020 | A Knowledge-Aware Sequence-to-Tree Network for Math Word Problem SolvingabstractWith the advancements in natural language processing tasks, math word problem solving has received increasing attention.Previous methods have achieved promising results but ignore background common-sense knowledge not directly provided by the problem.In addition, during generation, they focus on local features while neglecting global information.To incorporate external knowledge and global expression information, we propose a novel knowledge-aware sequence-to-tree (KA-S2T) network in which the entities in the problem sequences and their categories are modeled as an entity graph.Based on this entity graph, a graph attention network is used to capture knowledge-aware problem representations.Further, we use a tree-structured decoder with a state aggregation mechanism to capture the long-distance dependency and global expression information.Experimental results on the Math23K dataset revealed that the KA-S2T model can achieve better performance than previously reported best results. Qinzhuo Wu, Qi Zhang 0001, Jinlan Fu, Xuanjing Huang 0001 |
EMNLP (1) | 4 |
| 2020 | Tasty Burgers, Soggy Fries: Probing Aspect Robustness in Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) aims to predict the sentiment towards a specific aspect in the text.However, existing ABSA test sets cannot be used to probe whether a model can distinguish the sentiment of the target aspect from the non-target aspects.To solve this problem, we develop a simple but effective approach to enrich ABSA test sets.Specifically, we generate new examples to disentangle the confounding sentiments of the non-target aspects from the target aspect's sentiment.Based on the SemEval 2014 dataset, we construct the Aspect Robustness Test Set (ARTS) as a comprehensive probe of the aspect robustness of ABSA models.Over 92% data of ARTS show high fluency and desired sentiment on all aspects by human evaluation.Using ARTS, we analyze the robustness of nine ABSA models, and observe, surprisingly, that their accuracy drops by up to 69.73%.We explore several ways to improve aspect robustness, and find that adversarial training can improve models' performance on ARTS by up to 32.85%. 1 Zhijing Jin 0001, Di Jin 0005, Bingning Wang, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP (1) | 6 |
| 2020 | Leveraging Document-Level Label Consistency for Named Entity RecognitionabstractDocument-level label consistency is an effective indicator that different occurrences of a particular token sequence are very likely to have the same entity types. Previous work focused on better context representations and used the CRF for label decoding. However, CRF-based methods are inadequate for modeling document-level label consistency. This work introduces a novel two-stage label refinement approach to handle document-level label consistency, where a key-value memory network is first used to record draft labels predicted by the base model, and then a multi-channel Transformer makes refinements on these draft predictions based on the explicit co-occurrence relationship derived from the memory network. In addition, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with a high probability of being wrong, which can greatly assist in preventing the incorrect refinement of correct draft labels. The experimental results on three named entity recognition benchmarks demonstrated that the proposed method significantly outperformed the state-of-the-art methods. Tao Gui, Jiacheng Ye, Qi Zhang 0001, Yaqian Zhou 0001, Yeyun Gong, Xuanjing Huang 0001 |
IJCAI | 6 |
| 2020 | Learning to Generate Representations for Novel Words: Mimic the OOV Situation in Training
Minlong Peng, Qi Zhang 0001, Qin Liu 0010, Xuanjing Huang 0001 |
NLPCC (1) | 5 |
| 2020 | Recurrent Memory Reasoning Network for Expert Finding in Community Question AnsweringabstractExpert finding is a task designed to enable recommendation of the right person who can provide high-quality answers to a requester's question. Most previous works try to involve a content-based recommendation, which only superficially comprehends the relevance between a requester's question and the expertise of candidate experts by exploring the content or topic similarity between the requester's question and the candidate experts' historical answers. However, if a candidate expert has never answered a question similar to the requester's question, then existing methods have difficulty making a correct recommendation. Therefore, exploring the implicit relevance between a requester's question and a candidate expert's historical records by perception and reasoning should be taken into consideration. In this study, we propose a novel \textslrecurrent memory reasoning network (RMRN) to perform this task. This method focuses on different parts of a question, and accordingly retrieves information from the histories of the candidate expert.Since only a small percentage of historical records are relevant to any requester's question, we introduce a Gumbel-Softmax-based mechanism to select relevant historical records from candidate experts' answering histories. To evaluate the proposed method, we constructed two large-scale datasets drawn from Stack Overflow and Yahoo! Answer. Experimental results on the constructed datasets demonstrate that the proposed method could achieve better performance than existing state-of-the-art methods. Jinlan Fu, Qi Zhang 0001, Qinzhuo Wu, Renfeng Ma, Xuanjing Huang 0001, Yu-Gang Jiang 0001 |
WSDM | 6 |
| 2020 | Chinese Word Segmentation via BiLSTM+Semi-CRF with Relay Node
Nuo Qun, Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001 |
J. Comput. Sci. Technol. | 4 |
| 2020 | A Graph-based Model for Joint Chinese Word Segmentation and Dependency ParsingabstractChinese word segmentation and dependency parsing are two fundamental tasks for Chinese natural language processing. The dependency parsing is defined at the word-level. Therefore word segmentation is the precondition of dependency parsing, which makes dependency parsing suffer from error propagation and unable to directly make use of character-level pre-trained language models (such as BERT). In this paper, we propose a graph-based model to integrate Chinese word segmentation and dependency parsing. Different from previous transition-based joint models, our proposed model is more concise, which results in fewer efforts of feature engineering. Our graph-based joint model achieves better performance than previous joint models and state-of-the-art results in both Chinese word segmentation and dependency parsing. Additionally, when BERT is combined, our model can substantially reduce the performance gap of dependency parsing between joint models and gold-segmented word-based models. Our code is publicly available at https://github.com/fastnlp/JointCwsParser Hang Yan 0001, Xipeng Qiu, Xuanjing Huang 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2019 | Long Short-Term Memory with Dynamic Skip ConnectionsabstractIn recent years, long short-term memory (LSTM) has been successfully used to model sequential data of variable length. However, LSTM can still experience difficulty in capturing long-term dependencies. In this work, we tried to alleviate this problem by introducing a dynamic skip connection, which can learn to directly connect two dependent words. Since there is no dependency information in the training data, we propose a novel reinforcement learning-based method to model the dependency relationship and connect dependent words. The proposed model computes the recurrent transition functions based on the skip connections, which provides a dynamic skipping advantage over RNNs that always tackle entire sentences sequentially. Our experimental results on three natural language processing tasks demonstrate that the proposed method can achieve better performance than existing methods. In the number prediction experiment, the proposed model outperformed LSTM with respect to accuracy by nearly 20%. Tao Gui, Qi Zhang 0001, Lujun Zhao, Yaosong Lin, Minlong Peng, Jingjing Gong, Xuanjing Huang 0001 |
AAAI | 7 |
| 2019 | Contextualized Non-Local Neural Networks for Sequence LearningabstractRecently, a large number of neural mechanisms and models have been proposed for sequence learning, of which selfattention, as exemplified by the Transformer model, and graph neural networks (GNNs) have attracted much attention. In this paper, we propose an approach that combines and draws on the complementary strengths of these two methods. Specifically, we propose contextualized non-local neural networks (CN3), which can both dynamically construct a task-specific structure of a sentence and leverage rich local dependencies within a particular neighbourhood.Experimental results on ten NLP tasks in text classification, semantic matching, and sequence labelling show that our proposed model outperforms competitive baselines and discovers task-specific dependency structures, thus providing better interpretability to users. Pengfei Liu 0003, Shuaichen Chang, Xuanjing Huang 0001, Jackie Chi Kit Cheung |
AAAI | 3 |
| 2019 | Trainable Undersampling for Class-Imbalance LearningabstractUndersampling has been widely used in the class-imbalance learning area. The main deficiency of most existing undersampling methods is that their data sampling strategies are heuristic-based and independent of the used classifier and evaluation metric. Thus, they may discard informative instances for the classifier during the data sampling. In this work, we propose a meta-learning method built on the undersampling to address this issue. The key idea of this method is to parametrize the data sampler and train it to optimize the classification performance over the evaluation metric. We solve the non-differentiable optimization problem for training the data sampler via reinforcement learning. By incorporating evaluation metric optimization into the data sampling process, the proposed method can learn which instance should be discarded for the given classifier and evaluation metric. In addition, as a data level operation, this method can be easily applied to arbitrary evaluation metric and classifier, including non-parametric ones (e.g., C4.5 and KNN). Experimental results on both synthetic and realistic datasets demonstrate the effectiveness of the proposed method. Minlong Peng, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001, Yu-Gang Jiang 0001, Keyu Ding |
AAAI | 5 |
| 2019 | A Multi-Agent Communication Framework for Question-Worthy Phrase Extraction and Question GenerationabstractQuestion generation aims to produce questions automatically given a piece of text as input. Existing research follows a sequence-to-sequence fashion that constructs a single question based on the input. Considering each question usually focuses on a specific fragment of the input, especially in the scenario of reading comprehension, it is reasonable to identify the corresponding focus before constructing the question. In this paper, we propose to identify question-worthy phrases first and generate questions with the assistance of these phrases. We introduce a multi-agent communication framework, taking phrase extraction and question generation as two agents, and learn these two tasks simultaneously via message passing mechanism. The results of experiments show the effectiveness of our framework: we can extract question-worthy phrases, which are able to improve the performance of question generation. Besides, our system is able to extract more than one question worthy phrases and generate multiple questions accordingly. Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Yang Liu 0004, Xuanjing Huang 0001 |
AAAI | 5 |
| 2019 | Style Transformer: Unpaired Text Style Transfer without Disentangled Latent RepresentationabstractDisentangling the content and style in the latent space is prevalent in unpaired text style transfer.However, two major issues exist in most of the current neural models.1) It is difficult to completely strip the style information from the semantics for a sentence.2) The recurrent neural network (RNN) based encoder and decoder, mediated by the latent representation, cannot well deal with the issue of the long-term dependency, resulting in poor preservation of non-stylistic semantic content.In this paper, we propose the Style Transformer, which makes no assumption about the latent representation of source sentence and equips the power of attention mechanism in Transformer to achieve better style transfer and better content preservation.Source code will be available on Github 1 . Jianze Liang, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2019 | Bridging by Word: Image Grounded Vocabulary Construction for Visual CaptioningabstractExisting research for visual captioning usually employs a CNN-RNN architecture that combines a CNN for image encoding with a RNN for caption generation, where the vocabulary is constructed from the entire training dataset as the decoding space.Such approaches typically suffer from the problem of generating N-grams which occur frequently in the training set but are irrelevant to the given image.To tackle this problem, we propose to construct an image-grounded vocabulary that leverages image semantics for more effective caption generation.More concretely, a two-step approach is proposed to construct the vocabulary by incorporating both visual information and relationships among words.Two strategies are then explored to utilize the constructed vocabulary for caption generation.One constrains the generator to select words from the image-grounded vocabulary only and the other integrates the vocabulary information into the RNN cell during the caption generation process.Experimental results on two public datasets show the effectiveness of our framework compared to state-of-the-art models.Our code is available on Github 1 . Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2019 | Distantly Supervised Named Entity Recognition using Positive-Unlabeled LearningabstractIn this work, we explore the way to perform named entity recognition (NER) using only unlabeled data and named entity dictionaries.To this end, we formulate the task as a positive-unlabeled (PU) learning problem and accordingly propose a novel PU learning algorithm to perform the task.We prove that the proposed algorithm can unbiasedly and consistently estimate the task loss as if there is fully labeled data.A key feature of the proposed method is that it does not require the dictionaries to label every entity within a sentence, and it even does not require the dictionaries to label all of the words constituting an entity.This greatly reduces the requirement on the quality of the dictionaries and makes our method generalize well with quite simple dictionaries.Empirical studies on four public NER datasets demonstrate the effectiveness of our proposed method.We have published the source code at https:// github.com/v-mipeng/LexiconNER. Minlong Peng, Qi Zhang 0001, Jinlan Fu, Xuanjing Huang 0001 |
ACL (1) | 5 |
| 2019 | Generating Responses with a Specific Emotion in DialogabstractIt is desirable for dialog systems to have capability to express specific emotions during a conversation, which has a direct, quantifiable impact on improvement of their usability and user satisfaction.After a careful investigation of real-life conversation data, we found that there are at least two ways to express emotions with language.One is to describe emotional states by explicitly using strong emotional words; another is to increase the intensity of the emotional experiences by implicitly combining neutral words in distinct ways.We propose an emotional dialogue system (EmoDS) that can generate the meaningful responses with a coherent structure for a post, and meanwhile express the desired emotion explicitly or implicitly within a unified framework.Experimental results showed EmoDS performed better than the baselines in BLEU, diversity and the quality of emotional expression. Zhenqiao Song, Xiaoqing Zheng, Lu Liu 0009, Mu Xu, Xuanjing Huang 0001 |
ACL (1) | 5 |
| 2019 | Searching for Effective Neural Extractive Summarization: What Works and What's NextabstractThe recent years have seen remarkable success in the use of deep neural networks on text summarization.However, there is no clear understanding of why they perform so well, or how they might be improved.In this paper, we seek to better understand how neural extractive summarization systems could benefit from different types of model architectures, transferable knowledge and learning schemas.Additionally, we find an effective way to improve current frameworks and achieve the state-ofthe-art result on CNN/DailyMail by a large margin based on our observations and analyses.Hopefully, our work could provide more clues for future research on extractive summarization.Source code will be available on Github 1 . Ming Zhong 0005, Pengfei Liu 0003, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 5 |
| 2019 | A Lexicon-Based Graph Neural Network for Chinese NERabstractTao Gui, Yicheng Zou, Qi Zhang, Minlong Peng, Jinlan Fu, Zhongyu Wei, Xuanjing Huang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tao Gui, Yicheng Zou, Qi Zhang 0001, Minlong Peng, Jinlan Fu, Zhongyu Wei, Xuanjing Huang 0001 |
EMNLP/IJCNLP (1) | 7 |
| 2019 | GlossBERT: BERT for Word Sense Disambiguation with Gloss KnowledgeabstractLuyao Huang, Chi Sun, Xipeng Qiu, Xuanjing Huang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Luyao Huang, Chi Sun, Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Asynchronous Deep Interaction Network for Natural Language InferenceabstractDi Liang, Fubao Zhang, Qi Zhang, Xuanjing Huang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Di Liang, Fubao Zhang, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | CNN-Based Chinese NER with Lexicon RethinkingabstractCharacter-level Chinese named entity recognition (NER) that applies long short-term memory (LSTM) to incorporate lexicons has achieved great success. However, this method fails to fully exploit GPU parallelism and candidate lexicons can conflict. In this work, we propose a faster alternative to Chinese NER: a convolutional neural network (CNN)-based method that incorporates lexicons using a rethinking mechanism. The proposed method can model all the characters and potential words that match the sentence in parallel. In addition, the rethinking mechanism can address the word conflict by feeding back the high-level features to refine the networks. Experimental results on four datasets show that the proposed method can achieve better performance than both word-level and character-level baseline methods. In addition, the proposed method performs up to 3.21 times faster than state-of-the-art methods, while realizing better performance. Tao Gui, Ruotian Ma, Qi Zhang 0001, Lujun Zhao, Yu-Gang Jiang 0001, Xuanjing Huang 0001 |
IJCAI | 6 |
| 2019 | Building Personalized Simulator for Interactive SearchabstractInteractive search, where a set of tags is recommended to users together with search results at each turn, is an effective way to guide users to identify their information need. It is a classical sequential decision problem and the reinforcement learning based agent can be introduced as a solution. The training of the agent can be divided into two stages, i.e., offline and online. Existing reinforcement learning based systems tend to perform the offline training in a supervised way based on historical labeled data while the online training is performed via reinforcement learning algorithms based on interactions with real users. The mis-match between online and offline training leads to a cold-start problem for the online usage of the agent. To address this issue, we propose to employ a simulator to mimic the environment for the offline training of the agent. Users' profiles are considered to build a personalized simulator, besides, model-based approach is used to train the simulator and is able to use the data efficiently. Experimental results based on real-world dataset demonstrate the effectiveness of our agent and personalized simulator. Qianlong Liu, Baoliang Cui, Zhongyu Wei, Baolin Peng, Haikuan Huang, Hongbo Deng, Jianye Hao, Xuanjing Huang 0001, Kam-Fai Wong |
IJCAI | 8 |
| 2019 | Learning Task-Specific Representation for Novel Words in Sequence LabelingabstractWord representation is a key component in neural-network-based sequence labeling systems. However, representations of unseen or rare words trained on the end task are usually poor for appreciable performance. This is commonly referred to as the out-of-vocabulary (OOV) problem. In this work, we address the OOV problem in sequence labeling using only training data of the task. To this end, we propose a novel method to predict representations for OOV words from their surface-forms (e.g., character sequence) and contexts. The method is specifically designed to avoid the error propagation problem suffered by existing approaches in the same paradigm. To evaluate its effectiveness, we performed extensive empirical studies on four part-of-speech tagging (POS) tasks and four named entity recognition (NER) tasks. Experimental results show that the proposed method can achieve better or competitive performance on the OOV problem compared with existing state-of-the-art methods. Minlong Peng, Qi Zhang 0001, Tao Gui, Jinlan Fu, Xuanjing Huang 0001 |
IJCAI | 6 |
| 2019 | Model the Long-Term Post History for Hashtag Recommendation
Minlong Peng, Qiyuan Bian, Qi Zhang 0001, Tao Gui, Jinlan Fu, Lanjun Zeng, Xuanjing Huang 0001 |
NLPCC (1) | 7 |
| 2019 | Mention Recommendation in Twitter with Cooperative Multi-Agent Reinforcement LearningabstractIn Twitter-like social networking services, the "@'' symbol can be used with the tweet to mention users whom the user wants to alert regarding the message. An automatic suggestion to the user of a small list of candidate names can improve communication efficiency. Previous work usually used several most recent tweets or randomly select historical tweets to make an inference about this preferred list of names. However, because there are too many historical tweets by users and a wide variety of content types, the use of several tweets cannot guarantee the desired results. In this work, we propose the use of a novel cooperative multi-agent approach to mention recommendation, which incorporates dozens of more historical tweets than earlier approaches. The proposed method can effectively select a small set of historical tweets and cooperatively extract relevant indicator tweets from both the user and mentioned users. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods. Tao Gui, Qi Zhang 0001, Minlong Peng, Yunhua Zhou, Xuanjing Huang 0001 |
SIGIR | 7 |
| 2019 | Adaptive Multi-Attention Network Incorporating Answer Information for Duplicate Question DetectionabstractCommunity-based question answering (CQA), which provides a platform for people with diverse backgrounds to share information and knowledge, has become increasingly popular. With the accumulation of site data, methods to detect duplicate questions in CQA sites have attracted considerable attention. Existing methods typically use only questions to complete the task. However, the paired answers may also provide valuable information. In this paper, we propose an answer information- enhanced adaptive multi-attention network (AMAN) to perform this task. AMAN takes full advantage of the semantic information in the paired answers while alleviating the noise problem caused by adding the answers. To evaluate the proposed method, we use a CQADupStack set and the Quora question-pair dataset expanded with paired answers. Experimental results demonstrate that the proposed model can achieve state-of-the-art performance on the above two data sets. Di Liang, Fubao Zhang, Qi Zhang 0001, Jinlan Fu, Minlong Peng, Tao Gui, Xuanjing Huang 0001 |
SIGIR | 8 |
| 2019 | Hot Topic-Aware Retweet Prediction with Masked Self-attentive ModelabstractSocial media users create millions of microblog entries on various topics each day. Retweet behaviour play a crucial role in spreading topics on social media. Retweet prediction task has received considerable attention in recent years. The majority of existing retweet prediction methods are focus on modeling user preference by utilizing various information, such as user profiles, user post history, user following relationships, etc. Yet, the users exposures towards real-time posting from their followees contribute significantly to making retweet predictions, considering that the users may participate into the hot topics discussed by their followees rather than be limited to their previous interests. To make efficient use of hot topics, we propose a novel masked self-attentive model to perform the retweet prediction task by perceiving the hot topics discussed by the users' followees. We incorporate the posting histories of users with external memory and utilize a hierarchical attention mechanism to construct the users' interests. Hence, our model can be jointly hot-topic aware and user interests aware to make a final prediction. Experimental results on a dataset collected from Twitter demonstrated that the proposed method can achieve better performance than state-of-the-art methods. Renfeng Ma, Xiangkun Hu, Qi Zhang 0001, Xuanjing Huang 0001, Yu-Gang Jiang 0001 |
SIGIR | 4 |
| 2019 | Review Response Generation in E-Commerce Platforms with External Product Informationabstract''User reviews” are becoming an essential component of e-commerce. When buyers write a negative or doubting review, ideally, the sellers need to quickly give a response to minimize the potential impact. When the number of reviews is growing at a frightening speed, there is an urgent need to build a response writing assistant for customer service providers. In order to generate high-quality responses, the algorithm needs to consume and understand the information from both the original review and the target product. The classical sequence-to-sequence (Seq2Seq) methods can hardly satisfy this requirement. In this study, we propose a novel deep neural network model based on the Seq2Seq framework for the review response generation task in e-commerce platforms, which can incorporate product information by a gated multi-source attention mechanism and a copy mechanism. Moreover, we employ a reinforcement learning technique to reduce the exposure bias problem. To evaluate the proposed model, we constructed a large-scale dataset from a popular e-commerce website, which contains product information. Empirical studies on both automatic evaluation metrics and human annotations show that the proposed model can generate informative and diverse responses, significantly outperforming state-of-the-art text generation models. Lujun Zhao, Kaisong Song, Changlong Sun, Qi Zhang 0001, Xuanjing Huang 0001, Xiaozhong Liu 0001 |
WWW | 5 |
| 2019 | Implicit discourse relation detection using concatenated word embeddings and a gated relevance network
Jinlan Fu, Qi Zhang 0001, Jifan Chen, Minlong Peng, Tao Gui, Xipeng Qiu, Xuanjing Huang 0001 |
Sci. China Inf. Sci. | 7 |
| 2019 | Reformulating natural language queries using sequence-to-sequence models
Shunda Pan, Qi Zhang 0001, Yu-Gang Jiang 0001, Xuanjing Huang 0001 |
Sci. China Inf. Sci. | 5 |
| 2019 | Sequence Labeling With Deep Gated Dual Path CNNabstractSequence labeling, such as part-of-speech (POS) tagging, named entity recognition (NER), text chunking, is a classic task in natural language processing. Most existing neural networks models for sequence labeling are based on recurrent neural networks. Recently, convolutional neural networks have been proposed to replace the recurrent components for sequence labeling. However, they are usually shallow compared to deep convolutional networks that achieve start-of-the-art performance in other fields. Due to the vanishing gradient problem, these models usually can not work well when simply increasing the number of layers. In this paper, we propose using deep CNN architecture in sequence labeling, which can capture a large context through stacked convolutions. To reduce the vanishing gradient problem, the proposed method incorporates gated linear units, residual connections, and dense connections. Experimental results on three sequence labeling tasks show that the proposed model can achieve competitive performance to the RNN-based state-of-the-art method while maintaining 2.41 × faster speed, even with up to 10 convolutional layers. Lujun Zhao, Xipeng Qiu, Qi Zhang 0001, Xuanjing Huang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Adaptive Co-attention Network for Named Entity Recognition in TweetsabstractIn this study, we investigate the problem of named entity recognition for tweets. Named entity recognition is an important task in natural language processing and has been carefully studied in recent decades. Previous named entity recognition methods usually only used the textual content when processing tweets. However, many tweets contain not only textual content, but also images. Such visual information is also valuable in the name entity recognition task. To make full use of textual and visual information, this paper proposes a novel method to process tweets that contain multimodal information. We extend a bi-directional long short term memory network with conditional random fields and an adaptive co-attention network to achieve this task. To evaluate the proposed methods, we constructed a large scale labeled dataset that contained multimodal tweets. Experimental results demonstrated that the proposed method could achieve a better performance than the previous methods in most cases. Qi Zhang 0001, Jinlan Fu, Xuanjing Huang 0001 |
AAAI | 4 |
| 2018 | Meta Multi-Task Learning for Sequence ModelingabstractSemantic composition functions have been playing a pivotal role in neural representation learning of text sequences. In spite of their success, most existing models suffer from the underfitting problem: they use the same shared compositional function on all the positions in the sequence, thereby lacking expressive power due to incapacity to capture the richness of compositionality. Besides, the composition functions of different tasks are independent and learned from scratch. In this paper, we propose a new sharing scheme of composition function across multiple tasks. Specifically, we use a shared meta-network to capture the meta-knowledge of semantic composition and generate the parameters of the task-specific semantic composition models. We conduct extensive experiments on two types of tasks, text classification and sequence tagging, which demonstrate the benefits of our approach. Besides, we show that the shared meta-knowledge learned by our proposed model can be regarded as off-the-shelf knowledge and easily transferred to new tasks. Jun-Kun Chen, Xipeng Qiu, Pengfei Liu 0003, Xuanjing Huang 0001 |
AAAI | 4 |
| 2018 | Incorporating Discriminator in Sentence Generation: a Gibbs Sampling MethodabstractGenerating plausible and fluent sentence with desired properties has long been a challenge. Most of the recent works use recurrent neural networks (RNNs) and their variants to predict following words given previous sequence and target label. In this paper, we propose a novel framework to generate constrained sentences via Gibbs Sampling. The candidate sentences are revised and updated iteratively, with sampled new words replacing old ones. Our experiments show the effectiveness of the proposed method to generate plausible and diverse sentences. Jinyue Su, Jiacheng Xu 0001, Xipeng Qiu, Xuanjing Huang 0001 |
AAAI | 4 |
| 2018 | Cross-Domain Sentiment Classification with Target Domain Specific InformationabstractThe task of adopting a model with good performance to a target domain that is different from the source domain used for training has received considerable attention in sentiment analysis.Most existing approaches mainly focus on learning representations that are domain-invariant in both the source and target domains.Few of them pay attention to domain specific information, which should also be informative.In this work, we propose a method to simultaneously extract domain specific and invariant representations and train a classifier on each of the representation, respectively.And we introduce a few target domain labeled data for learning domain-specific information.To effectively utilize the target domain labeled data, we train the domain-invariant representation based classifier with both the source and target domain labeled data and train the domain-specific representation based classifier with only the target domain labeled data.These two classifiers then boost each other in a co-training style.Extensive sentiment analysis experiments demonstrated that the proposed method could achieve better performance than state-of-the-art methods. Minlong Peng, Qi Zhang 0001, Yu-Gang Jiang 0001, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2018 | Incorporating Corporation Relationship via Graph Convolutional Neural Networks for Stock Price PredictionabstractIn this paper, we propose to incorporate information of related corporations of a target company for its stock price prediction. We first construct a graph including all involved corporations based on investment facts from real market and learn a distributed representation for each corporation via node embedding methods applied on the graph. Two approaches are then explored to utilize information of related corporations based on a pipeline model and a joint model via graph convolutional neural networks respectively. Experiments on the data collected from stock market in Mainland China show that the representation learned from our model is able to capture relationships between corporations, and prediction models incorporating related corporations' information are able to make more accurate predictions on stock market. Yingmei Chen, Zhongyu Wei, Xuanjing Huang 0001 |
CIKM | 3 |
| 2018 | Generating Keyword Queries for Natural Language Queries to Alleviate Lexical Chasm ProblemabstractIn recent years, the task of reformulating natural language queries has received considerable attention from both industry and academic communities. Because of the lexical chasm problem between natural language queries and web documents, if we directly use natural language queries as inputs for retrieval, the results are usually unsatisfactory. In this work, we formulated the task as a translation problem to convert natural language queries into keyword queries. Since the nature language queries users input are diverse and multi-faceted, general encoder-decoder models cannot effectively handle low-frequency words and out-of-vocabulary words. We propose a novel encoder-decoder method with two decoders: the pointer decoder firstly extracts query terms directly from the source text via copying mechanism, then the generator decoder generates query terms using two attention modules simultaneously considering the source text and extracted query terms. For evaluation and training, we also proposed a semi-automatic method to construct a large-scale dataset about natural language query-keyword query pairs. Experimental results on this dataset demonstrated that our model could achieve better performance than the previous state-of-the-art methods. Shunda Pan, Qi Zhang 0001, Yu-Gang Jiang 0001, Xuanjing Huang 0001 |
CIKM | 5 |
| 2018 | A Reinforcement Learning Framework for Natural Question Generation using Bi-discriminatorsabstractVisual Question Generation (VQG) aims to ask natural questions about an image automatically. Existing research focus on training model to fit the annotated data set that makes it indifferent from other language generation tasks. We argue that natural questions need to have two specific attributes from the perspectives of content and linguistic respectively, namely, natural and human-written. Inspired by the setting of discriminator in adversarial learning, we propose two discriminators, one for each attribute, to enhance the training. We then use the reinforcement learning framework to incorporate scores from the two discriminators as the reward to guide the training of the question generator. Experimental results on a benchmark VQG dataset show the effectiveness and robustness of our model compared to some state-of-the-art models in terms of both automatic and human evaluation metrics. Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Yang Liu 0004, Xuanjing Huang 0001 |
COLING | 5 |
| 2018 | Information Aggregation via Dynamic Routing for Sequence EncodingabstractWhile much progress has been made in how to encode a text sequence into a sequence of vectors, less attention has been paid to how to aggregate these preceding vectors (outputs of RNN/CNN) into fixed-size encoding vector. Usually, a simple max or average pooling is used, which is a bottom-up and passive way of aggregation and lack of guidance by task information. In this paper, we propose an aggregation mechanism to obtain a fixed-size encoding with a dynamic routing policy. The dynamic routing policy is dynamically deciding that what and how much information need be transferred from each word to the final encoding of the text sequence. Following the work of Capsule Network, we design two dynamic routing policies to aggregate the outputs of RNN/CNN encoding layer into a final encoding vector. Compared to the other aggregation methods, dynamic routing can refine the messages according to the state of final encoding vector. Experimental results on five text classification tasks show that our method outperforms other aggregating models by a significant margin. Related source code is released on our github page. Related source code is released on our github page. Jingjing Gong, Xipeng Qiu, Shaojing Wang, Xuanjing Huang 0001 |
COLING | 4 |
| 2018 | Incorporating Argument-Level Interactions for Persuasion Comments Evaluation using Co-attention ModelabstractIn this paper, we investigate the issue of persuasiveness evaluation for argumentative comments. Most of the existing research explores different text features of reply comments on word level and ignores interactions between participants. In general, viewpoints are usually expressed by multiple arguments and exchanged on argument level. To better model the process of dialogical argumentation, we propose a novel co-attention mechanism based neural network to capture the interactions between participants on argument level. Experimental results on a publicly available dataset show that the proposed model significantly outperforms some state-of-the-art methods for persuasiveness evaluation. Further analysis reveals that attention weights computed in our model are able to extract interactive argument pairs from the original post and the reply. Lu Ji, Zhongyu Wei, Xiangkun Hu, Yang Liu 0004, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 6 |
| 2018 | A Lexicon-Based Supervised Attention Model for Neural Sentiment AnalysisabstractAttention mechanisms have been leveraged for sentiment classification tasks because not all words have the same importance. However, most existing attention models did not take full advantage of sentiment lexicons, which provide rich sentiment information and play a critical role in sentiment analysis. To achieve the above target, in this work, we propose a novel lexicon-based supervised attention model (LBSA), which allows a recurrent neural network to focus on the sentiment content, thus generating sentiment-informative representations. Compared with general attention models, our model has better interpretability and less noise. Experimental results on three large-scale sentiment classification datasets showed that the proposed method outperforms previous methods. Yicheng Zou, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 4 |
| 2018 | Convolutional Interaction Network for Natural Language InferenceabstractAttention-based neural models have achieved great success in natural language inference (NLI).In this paper, we propose the Convolutional Interaction Network (CIN), a general model to capture the interaction between two sentences, which can be an alternative to the attention mechanism for NLI.Specifically, CIN encodes one sentence with the filters dynamically generated based on another sentence.Since the filters may be designed to have various numbers and sizes, CIN can capture more complicated interaction patterns.Experiments on three very large datasets demonstrate CIN's efficacy. Jingjing Gong, Xipeng Qiu, Xinchi Chen, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2018 | Transferring from Formal Newswire Domain with Hypernet for Twitter POS TaggingabstractPart-of-Speech (POS) tagging for Twitter has received considerable attention in recent years.Because most POS tagging methods are based on supervised models, they usually require a large amount of labeled data for training.However, the existing labeled datasets for Twitter are much smaller than those for newswire text.Hence, to help POS tagging for Twitter, most domain adaptation methods try to leverage newswire datasets by learning the shared features between the two domains.However, from a linguistic perspective, Twitter users not only tend to mimic the formal expressions of traditional media, like news, but they also appear to be developing linguistically informal styles.Therefore, POS tagging for the formal Twitter context can be learned together with the newswire dataset, while POS tagging for the informal Twitter context should be learned separately.To achieve this task, in this work, we propose a hypernetworkbased method to generate different parameters to separately model contexts with different expression styles.Experimental results on three different datasets show that our approach achieves better performance than state-of-theart methods in most cases. Tao Gui, Qi Zhang 0001, Jingjing Gong, Minlong Peng, Di Liang, Keyu Ding, Xuanjing Huang 0001 |
EMNLP | 7 |
| 2018 | Automatic Essay Scoring Incorporating Rating Schema via Reinforcement LearningabstractAutomatic essay scoring (AES) is the task of assigning grades to essays without human interference. Existing systems for AES are typically trained to predict the score of each single essay at a time without considering the rating schema. In order to address this issue, we propose a reinforcement learning framework for essay scoring that incorporates quadratic weighted kappa as guidance to optimize the scoring system. Experiment results on benchmark datasets show the effectiveness of our framework. Zhongyu Wei, Yaqian Zhou 0001, Xuanjing Huang 0001 |
EMNLP | 4 |
| 2018 | A Question Type Driven Framework to Diversify Visual Question GenerationabstractVisual question generation aims at asking questions about an image automatically. Existing research works on this topic usually generate a single question for each given image without considering the issue of diversity. In this paper, we propose a question type driven framework to produce multiple questions for a given image with different focuses. In our framework, each question is constructed following the guidance of a sampled question type in a sequence-to-sequence fashion. To diversify the generated questions, a novel conditional variational auto-encoder is introduced to generate multiple questions with a specific question type. Moreover, we design a strategy to conduct the question type distribution learning for each image to select the final questions. Experimental results on three benchmark datasets show that our framework outperforms the state-of-the-art approaches in terms of both relevance and diversity. Zhihao Fan, Zhongyu Wei, Piji Li, Yanyan Lan, Xuanjing Huang 0001 |
IJCAI | 5 |
| 2018 | Toward Diverse Text Generation with Inverse Reinforcement LearningabstractText generation is a crucial task in NLP. Recently, several adversarial generative models have been proposed to improve the exposure bias problem in text generation. Though these models gain great success, they still suffer from the problems of reward sparsity and mode collapse. In order to address these two problems, in this paper, we employ inverse reinforcement learning (IRL) for text generation. Specifically, the IRL framework learns a reward function on training data, and then an optimal policy to maximum the expected total reward. Similar to the adversarial models, the reward and policy function in IRL are optimized alternately. Our method has two advantages: (1) the reward function can produce more dense reward signals. (2) the generation policy, trained by ``entropy regularized'' policy gradient, encourages to generate more diversified texts. Experiment results demonstrate that our proposed method can generate higher quality texts than the previous methods. Xinchi Chen, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 4 |
| 2018 | Mention Recommendation for Multimodal Microblog with Cross-attention Memory NetworkabstractThe users of Twitter-like social media normally use the "@'' sign to select a suitable person to mention. It is a significant role in promoting the user experience and information propagation. To help users easily find the usernames they want to mention, the mention recommendation task has received considerable attention in recent years. Previous methods only incorporated textual information when performing this task. However, many users not only post texts on social media but also the corresponding images. These images can provide additional information that is not included in the text, which could be helpful in improving the accuracy of a mention recommendation. To make full use of textual and visual information, we propose a novel cross-attention memory network to perform the mention recommendation task for multimodal tweets. We incorporate the interests of users with external memory and use the cross-attention mechanism to extract both textual and visual information. Experimental results on a dataset collected from Twitter demonstrated that the proposed method can achieve better performance than state-of-the-art methods that use textual information only. Renfeng Ma, Qi Zhang 0001, Li-Zhen Cui 0001, Xuanjing Huang 0001 |
SIGIR | 5 |
| 2018 | Hashtag recommendation for multimodal microblog posts
Yeyun Gong, Qi Zhang 0001, Xuanjing Huang 0001 |
Neurocomputing | 3 |
| 2017 | Adversarial Multi-Criteria Learning for Chinese Word SegmentationabstractDifferent linguistic perspectives causes many diverse segmentation criteria for Chinese word segmentation (CWS).Most existing methods focus on improve the performance for each single criterion.However, it is interesting to exploit these different criteria and mining their common underlying knowledge.In this paper, we propose adversarial multi-criteria learning for CWS by integrating shared knowledge from multiple heterogeneous segmentation criteria.Experiments on eight corpora with heterogeneous segmentation criteria show that the performance of each corpus obtains a significant improvement, compared to single-criterion learning.Source codes of this paper are available on Github 1 . Xinchi Chen, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2017 | Adversarial Multi-task Learning for Text ClassificationabstractNeural network models have shown their promising opportunities for multi-task learning, which focus on learning the shared layers to extract the common and task-invariant features.However, in most existing approaches, the extracted shared features are prone to be contaminated by task-specific features or the noise brought by other tasks.In this paper, we propose an adversarial multi-task learning framework, alleviating the shared and private latent feature spaces from interfering with each other.We conduct extensive experiments on 16 different text classification tasks, which demonstrates the benefits of our approach.Besides, we show that the shared knowledge learned by our proposed model can be regarded as off-the-shelf knowledge and easily transferred to new tasks.The datasets of all 16 tasks are publicly available at Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 3 |
| 2017 | Part-of-Speech Tagging for Twitter with Adversarial Neural NetworksabstractIn this work, we study the problem of partof-speech tagging for Tweets.In contrast to newswire articles, Tweets are usually informal and contain numerous out-ofvocabulary words.Moreover, there is a lack of large scale labeled datasets for this domain.To tackle these challenges, we propose a novel neural network to make use of out-of-domain labeled data, unlabeled in-domain data, and labeled indomain data.Inspired by adversarial neural networks, the proposed method tries to learn common features through adversarial discriminator.In addition, we hypothesize that domain-specific features of target domain should be preserved in some degree.Hence, the proposed method adopts a sequence-to-sequence autoencoder to perform this task.Experimental results on three different datasets show that our method achieves better performance than state-of-the-art methods. Tao Gui, Qi Zhang 0001, Haoran Huang, Minlong Peng, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2017 | Idiom-Aware Compositional Distributed SemanticsabstractIdioms are peculiar linguistic constructions that impose great challenges for representing the semantics of language, especially in current prevailing end-to-end neural models, which assume that the semantics of a phrase or sentence can be literally composed from its constitutive words.In this paper, we propose an idiomaware distributed semantic model to build representation of sentences on the basis of understanding their contained idioms.Our models are grounded in the literalfirst psycholinguistic hypothesis, which can adaptively learn semantic compositionality of a phrase literally or idiomatically.To better evaluate our models, we also construct an idiom-enriched sentiment classification dataset with considerable scale and abundant peculiarities of idioms.The qualitative and quantitative experimental analyses demonstrate the efficacy of our models.The newly-introduced datasets are publicly available at Pengfei Liu 0003, Kaiyu Qian, Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 4 |
| 2017 | A Mixed Generative-Discriminative Based Hashing MethodabstractHashing methods have proven to be useful for a variety of tasks and have attracted extensive attention in recent years. Various hashing approaches have been proposed to capture similarities between textual, visual, and cross-media information. However, most of the existing works use a bag-of-words methods to represent textual information. Since words with different forms may have similar meaning, semantic level text similarities can not be well processed in these methods. To address these challenges, in this paper, we propose a novel method called semantic cross-media hashing (SCMH), which uses continuous word representations to capture the textual similarity at the semantic level and use a deep belief network (DBN) to construct the correlation between different modalities. To demonstrate the effectiveness of the proposed method, we evaluate the proposed method on three commonly used cross-media data sets are used in this work. Experimental results show that the proposed method achieves significantly better performance than state-of-the-art approaches. Moreover, the efficiency of the proposed method is comparable to or better than that of some other hashing methods. Qi Zhang 0001, Yang Wang 0091, Binbin Deng, Xuanjing Huang 0001 |
ICDE | 5 |
| 2017 | A Feature-Enriched Neural Model for Joint Chinese Word Segmentation and Part-of-Speech TaggingabstractRecently, neural network models for natural language processing tasks have been increasingly focused on for their ability of alleviating the burden of manual feature engineering. However, the previous neural models cannot extract the complicated feature compositions as the traditional methods with discrete features. In this work, we propose a feature-enriched neural model for joint Chinese word segmentation and part-of-speech tagging task. Specifically, to simulate the feature templates of traditional discrete feature based models, we use different filters to model the complex compositional features with convolutional and pooling layer, and then utilize long distance dependency information with recurrent layer. Experimental results on five different datasets show the effectiveness of our proposed model. Xinchi Chen, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 3 |
| 2017 | Mention Recommendation for Twitter with End-to-end Memory NetworkabstractIn this study, we investigated the problem of recommending usernames when people attempt to use the ``@'' sign to mention other people in twitter-like social media. With the extremely rapid development of social networking services, this problem has received considerable attention in recent years. Previous methods have studied the problem from different aspects. Because most of Twitter-like microblogging services limit the length of posts, statistical learning methods may be affected by the problems of word sparseness and synonyms. Although recent progress in neural word embedding methods have advanced the state-of-the-art in many natural language processing tasks, the benefits of word embedding have not been taken into consideration for this problem. In this work, we proposed a novel end-to-end memory network architecture to perform this task. We incorporated the interests of users with external memory. A hierarchical attention mechanism was also applied to better consider the interests of users. The experimental results on a dataset we collected from Twitter demonstrated that the proposed method could outperform state-of-the-art approaches. Haoran Huang, Qi Zhang 0001, Xuanjing Huang 0001 |
IJCAI | 3 |
| 2017 | Dynamic Compositional Neural Networks over Tree StructureabstractTree-structured neural networks have proven to be effective in learning semantic representations by exploitingsyntactic information. In spite of their success, most existing models suffer from the underfitting problem: they recursively use the same shared compositional function throughout the whole compositional process and lack expressive power due to inability to capture the richness of compositionality.In this paper, we address this issue by introducing the dynamic compositional neural networks over tree structure (DC-TreeNN), in which the compositional function is dynamically generated by a meta network.The role of meta-network is to capture the metaknowledge across the different compositional rules and formulate them. Experimental results on two typical tasks show the effectiveness of the proposed models. Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 3 |
| 2017 | Adaptive Semantic Compositionality for Sentence ModellingabstractRepresenting a sentence with a fixed vector has shown its effectiveness in various NLP tasks. Most of the existing methods are based on neural network, which recursively apply different composition functions to a sequence of word vectors thereby obtaining a sentence vector.A hypothesis behind these approaches is that the meaning of any phrase can be composed of the meanings of its constituents.However, many phrases, such as idioms, are apparently non-compositional.To address this problem, we introduce a parameterized compositional switch, which outputs a scalar to adaptively determine whether the meaning of a phrase should be composed of its two constituents.We evaluate our model on five datasets of sentiment classification and demonstrate its efficacy with qualitative and quantitative experimental analysis . Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 3 |
| 2017 | Knowledge Graph Representation with Jointly Structural and Textual EncodingabstractThe objective of knowledge graph embedding is to encode both entities and relations of knowledge graphs into continuous low-dimensional vector spaces. Previously, most works focused on symbolic representation of knowledge graph with structure information, which can not handle new entities or entities with few facts well. In this paper, we propose a novel deep architecture to utilize both structural and textual information of entities. Specifically, we introduce three neural models to encode the valuable information from text description of entity, among which an attentive model can select related information as needed. Then, a gating mechanism is applied to integrate representations of structure and text into a unified architecture. Experiments show that our models outperform baseline and obtain state-of-the-art results on link prediction and triplet classification tasks. Jiacheng Xu 0001, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 4 |
| 2017 | Hashtag Recommendation for Multimodal Microblog Using Co-Attention NetworkabstractIn microblogging services, authors can use hashtags to mark keywords or topics. Many live social media applications (e.g., microblog retrieval, classification) can gain great benefits from these manually labeled tags. However, only a small portion of microblogs contain hashtags inputed by users. Moreover, many microblog posts contain not only textual content but also images. These visual resources also provide valuable information that may not be included in the textual content. So that it can also help to recommend hashtags more accurately. Motivated by the successful use of the attention mechanism, we propose a co-attention network incorporating textual and visual information to recommend hashtags for multimodal tweets. Experimental result on the data collected from Twitter demonstrated that the proposed method can achieve better performance than state-of-the-art methods using textual information only. Qi Zhang 0001, Haoran Huang, Xuanjing Huang 0001, Yeyun Gong |
IJCAI | 4 |
| 2017 | A Learning Error Analysis for Structured Prediction with Approximate InferenceabstractIn this work, we try to understand the differences between exact and approximate inference algorithms in structured prediction. We compare the estimation and approximation error of both underestimate and overestimate models. The result shows that, from the perspective of learning errors, performances of approximate inference could be as good as exact inference. The error analyses also suggest a new margin for existing learning algorithms. Empirical evaluations on text classification, sequential labelling and dependency parsing witness the success of approximate inference and the benefit of the proposed margin. Yuanbin Wu, Man Lan, Shiliang Sun, Qi Zhang 0001, Xuanjing Huang 0001 |
NIPS | 5 |
| 2017 | Hierarchical Dirichlet Processes with Social Influence
Yeyun Gong, Qi Zhang 0001, Xuanjing Huang 0001 |
NLPCC | 4 |
| 2017 | Overview of the NLPCC 2017 Shared Task: Chinese News Headline Categorization
Xipeng Qiu, Jingjing Gong, Xuanjing Huang 0001 |
NLPCC | 3 |
| 2017 | Hyper-Gated Recurrent Neural Networks for Chinese Word Segmentation
Xinchi Chen, Xipeng Qiu, Xuanjing Huang 0001 |
NLPCC | 4 |
| 2017 | Predicting Which Topics You Will Join in the Future on Social MediaabstractEvery day, social media users send millions of microblogs on every imaginable topics. If we could predict which topics a user will join in the future, it would be easy to determine what topics will become popular and what kinds of users a topic may attract. It also can be of great interest for many applications. In this study, we investigate the problem of predicting whether a user will join a topic based on his posting history. We introduce a novel deep convolutional neural network with external neural memory and attention mechanism to perform this problem. User's posting history and topics were modeled with an external neural memory architecture. The convolutional neural network based matching methods were used to construct the relations between users and topics. Final decisions were made based on these matching results. To train and evaluate the proposed method, we collected a large-scale dataset from Twitter. The experimental results demonstrated that the proposed method could perform significantly better than other methods. Comparing to the state-of-the-art deep neural networks, our approach achieves a relative improvement of 18.2\% in F1-score and 28.9\% in [email protected] Haoran Huang, Qi Zhang 0001, Jindou Wu, Xuanjing Huang 0001 |
SIGIR | 4 |
| 2017 | Phrase-based hashtag recommendation for microblog posts
Yeyun Gong, Qi Zhang 0001, Xiaoying Han, Xuanjing Huang 0001 |
Sci. China Inf. Sci. | 4 |
| 2017 | Preface
Hang Li 0001, Xiang Bai, Xuanjing Huang 0001, Changshui Zhang |
J. Comput. Sci. Technol. | 3 |
| 2016 | Discourse Relations Detection via a Mixed Generative-Discriminative FrameworkabstractWord embeddings, which can better capture the fine-grained semantics of words, have proven to be useful for a variety of natural language processing tasks. However, because discourse structures describe the relationships between segments of discourse, word embeddings cannot be directly integrated to perform the task. In this paper, we introduce a mixed generative-discriminative framework, in which we use vector offsets between embeddings of words to represent the semantic relations between text segments and Fisher kernel framework to convert a variable number of vector offsets into a fixed length vector. In order to incorporate the weights of these offsets into the vector, we also propose the Weighted Fisher Vector. Experimental results on two different datasets show that the proposed method without using manually designed features can achieve better performance on recognizing the discourse level relations in most cases. Jifan Chen, Qi Zhang 0001, Pengfei Liu 0003, Xuanjing Huang 0001 |
AAAI | 4 |
| 2016 | Implicit Discourse Relation Detection via a Deep Architecture with Gated Relevance NetworkabstractWord pairs, which are one of the most easily accessible features between two text segments, have been proven to be very useful for detecting the discourse relations held between text segments.However, because of the data sparsity problem, the performance achieved by using word pair features is limited.In this paper, in order to overcome the data sparsity problem, we propose the use of word embeddings to replace the original words.Moreover, we adopt a gated relevance network to capture the semantic interaction between word pairs, and then aggregate those semantic interactions using a pooling layer to select the most informative interactions.Experimental results on Penn Discourse Tree Bank show that the proposed method without using manually designed features can achieve better performance on recognizing the discourse level relations in all of the relations. Jifan Chen, Qi Zhang 0001, Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 5 |
| 2016 | Deep Fusion LSTMs for Text Semantic MatchingabstractRecently, there is rising interest in modelling the interactions of text pair with deep neural networks.In this paper, we propose a model of deep fusion LSTMs (DF-LSTMs) to model the strong interaction of text pair in a recursive matching way.Specifically, DF-LSTMs consist of two interdependent LSTMs, each of which models a sequence under the influence of another.We also use external memory to increase the capacity of LSTMs, thereby possibly capturing more complicated matching patterns.Experiments on two very large datasets demonstrate the efficacy of our proposed architecture.Furthermore, we present an elaborate qualitative analysis of our models, giving an intuitive understanding how our model worked. Pengfei Liu 0003, Xipeng Qiu, Jifan Chen, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2016 | Investigating Language Universal and Specific Properties in Word EmbeddingsabstractRecently, many NLP tasks have benefited from distributed word representation. However, it remains unknown whether embedding models are really immune to the typological diversity of languages, despite the language-independent architecture. Here we investigate three representative models on a large set of language samples by mapping dense embedding to sparse linguistic property space. Experiment results reveal the language universal and specific properties encoded in various word representation. Additionally, strong evidence supports the utility of word form, especially for inflectional languages. Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 3 |
| 2016 | A New Psychometric-inspired Evaluation Metric for Chinese Word SegmentationabstractWord segmentation is a fundamental task for Chinese language processing.However, with the successive improvements, the standard metric is becoming hard to distinguish state-of-the-art word segmentation systems.In this paper, we propose a new psychometric-inspired evaluation metric for Chinese word segmentation, which addresses to balance the very skewed word distribution at different levels of difficulty 1 .The performance on a real evaluation shows that the proposed metric gives more reasonable and distinguishable scores and correlates well with human judgement.In addition, the proposed metric can be easily extended to evaluate other sequence labelling based NLP tasks. Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 3 |
| 2016 | Incorporate Group Information to Enhance Network EmbeddingabstractThe problem of representing large-scale networks with low-dimensional vectors has received considerable attention in recent years. Except the networks that include only vertices and edges, a variety of networks contain information about groups or communities. For example, on Facebook, in addition to users and the follower-followee relations between them, users can also create and join groups. However, previous studies have rarely utilized this valuable information to generate embeddings of vertices. In this paper, we investigate a novel method for learning the network embeddings with valuable group information for large-scale networks. The proposed methods take both the inner structures of the groups and the information across groups into consideration. Experimental results demonstrate that the embeddings generated by the proposed methods significantly outperform state-of-the-art network embedding methods on two different scale real-world network Jifan Chen, Qi Zhang 0001, Xuanjing Huang 0001 |
CIKM | 3 |
| 2016 | Retweet Prediction with Attention-based Deep Neural NetworkabstractOn Twitter-like social media sites, the re-posting statuses or tweets of other users are usually considered to be the key mechanism for spreading information. How to predict whether a tweet will be retweeted by a user has received increasing attention in recent years. Previous methods studied the problem using various linguistic features, personal information of users, and many other manually constructed features to achieve the task. Usually, feature engineering is a laborious task, we require to obtain the external sources and they are difficult or not always available. Recently, deep learning methods have been used in the industry and research community for their ability to learn optimal features automatically and in many tasks, deep learning methods can achieve state-of-the art performance, such as natural language processing, computer vision, image classification and so on. In this work, we proposed a novel attention-based deep neural network to incorporate contextual and social information for this task. We used embeddings to represent the user, the user's attention interests, the author and tweet respectively. To train and evaluate the proposed methods, we also constructed a large dataset collected from Twitter. Experimental results showed that the proposed method could achieve better results than the previous state-of-the-art methods. Qi Zhang 0001, Yeyun Gong, Jindou Wu, Haoran Huang, Xuanjing Huang 0001 |
CIKM | 5 |
| 2016 | Hashtag Recommendation Using End-To-End Memory Networks with Hierarchical AttentionabstractOn microblogging services, people usually use hashtags to mark microblogs, which have a specific theme or content, making them easier for users to find. Hence, how to automatically recommend hashtags for microblogs has received much attention in recent years. Previous deep neural network-based hashtag recommendation approaches converted the task into a multi-class classification problem. However, most of these methods only took the microblog itself into consideration. Motivated by the intuition that the history of users should impact the recommendation procedure, in this work, we extend end-to-end memory networks to perform this task. We incorporate the histories of users into the external memory and introduce a hierarchical attention mechanism to select more appropriate histories. To train and evaluate the proposed method, we also construct a dataset based on microblogs collected from Twitter. Experimental results demonstrate that the proposed methods can significantly outperform state-of-the-art methods. By incorporating the hierarchical attention mechanism, the relative improvement in the proposed method over the state-of-the-art method is around 67.9% in the F1-score. Haoran Huang, Qi Zhang 0001, Yeyun Gong, Xuanjing Huang 0001 |
COLING | 4 |
| 2016 | Attention-Based Convolutional Neural Network for Semantic Relation ExtractionabstractNowadays, neural networks play an important role in the task of relation classification. In this paper, we propose a novel attention-based convolutional neural network architecture for this task. Our model makes full use of word embedding, part-of-speech tag embedding and position embedding information. Word level attention mechanism is able to better determine which parts of the sentence are most influential with respect to the two entities of interest. This architecture enables learning some important features from task-specific labeled data, forgoing the need for external knowledge such as explicit dependency structures. Experiments on the SemEval-2010 Task 8 benchmark dataset show that our model achieves better performances than several state-of-the-art neural network models and can achieve a competitive performance just with minimal feature engineering. Yatian Shen, Xuanjing Huang 0001 |
COLING | 2 |
| 2016 | Deep Multi-Task Learning with Shared Memory for Text Classification
Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 3 |
| 2016 | Modelling Interaction of Sentence Pair with Coupled-LSTMsabstractRecently, there is rising interest in modelling the interactions of two sentences with deep neural networks.However, most of the existing methods encode two sequences with separate encoders, in which a sentence is encoded with little or no information from the other sentence.In this paper, we propose a deep architecture to model the strong interaction of sentence pair with two coupled-LSTMs.Specifically, we introduce two coupled ways to model the interdependences of two LSTMs, coupling the local contextualized interactions of two sentences.We then aggregate these interactions and use a dynamic pooling to select the most informative features.Experiments on two very large datasets demonstrate the efficacy of our proposed architectures. Pengfei Liu 0003, Xipeng Qiu, Yaqian Zhou 0001, Jifan Chen, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2016 | Analyzing Linguistic Knowledge in Sequential Model of SentenceabstractSentence modelling is a fundamental topic in computational linguistics.Recently, deep learning-based sequential models of sentence, such as recurrent neural network, have proved to be effective in dealing with the non-sequential properties of human language.However, little is known about how a recurrent neural network captures linguistic knowledge.Here we propose to correlate the neuron activation pattern of a LSTM language model with rich language features at sequential, lexical and compositional level.Qualitative visualization as well as quantitative analysis under multilingual perspective reveals the effectiveness of gate neurons and indicates that LSTM learns to allow different neurons selectively respond to linguistic knowledge at different levels.Cross-language evidence shows that the model captures different aspects of linguistic properties for different languages due to the variance of syntactic complexity.Additionally, we analyze the influence of modelling strategy on linguistic knowledge encoded implicitly in different sequential models. Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 3 |
| 2016 | Cached Long Short-Term Memory Neural Networks for Document-Level Sentiment ClassificationabstractRecently, neural networks have achieved great success on sentiment classification due to their ability to alleviate feature engineering.However, one of the remaining challenges is to model long texts in document-level sentiment classification under a recurrent architecture because of the deficiency of the memory unit.To address this problem, we present a Cached Long Short-Term Memory neural networks (CLSTM) to capture the overall semantic information in long texts.CLSTM introduces a cache mechanism, which divides memory into several groups with different forgetting rates and thus enables the network to keep sentiment information better within a recurrent unit.The proposed CLSTM outperforms the state-of-the-art models on three publicly available document-level sentiment analysis datasets. Jiacheng Xu 0001, Danlu Chen, Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 4 |
| 2016 | Generating Abbreviations for Chinese Named Entities Using Recurrent Neural Network with Dynamic DictionaryabstractChinese named entities occur frequently in formal and informal environments.Various approaches have been formalized the problem as a sequence labelling task and utilize a character-based methodology, in which character is treated as the basic classification unit.One of the main drawbacks of these methods is that some of the generated abbreviations may not follow the conventional wisdom of Chinese.To address this problem, we propose a novel neural network architecture to perform task.It combines recurrent neural network (RNN) with an architecture determining whether a given sequence of characters can be a word or not.For demonstrating the effectiveness of the proposed method, we evaluate it on Chinese named entity generation and opinion target extraction tasks.Experimental results show that the proposed method can achieve better performance than state-ofthe-art methods. Qi Zhang 0001, Yaqian Zhou 0001, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2016 | Keyphrase Extraction Using Deep Recurrent Neural Networks on TwitterabstractKeyphrases can provide highly condensed and valuable information that allows users to quickly acquire the main ideas.The task of automatically extracting them have received considerable attention in recent decades.Different from previous studies, which are usually focused on automatically extracting keyphrases from documents or articles, in this study, we considered the problem of automatically extracting keyphrases from tweets.Because of the length limitations of Twitter-like sites, the performances of existing methods usually drop sharply.We proposed a novel deep recurrent neural network (RNN) model to combine keywords and context information to perform this problem.To evaluate the proposed method, we also constructed a large-scale dataset collected from Twitter.The experimental results showed that the proposed method performs significantly better than previous methods. Qi Zhang 0001, Yang Wang 0091, Yeyun Gong, Xuanjing Huang 0001 |
EMNLP | 4 |
| 2016 | Recurrent Neural Network for Text Classification with Multi-Task Learning
Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 3 |
| 2016 | Bridging LSTM Architecture and the Neural Dynamics during Reading
Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 3 |
| 2016 | A Mixed Generative-Discriminative Based Hashing MethodabstractHashing methods have proven to be useful for a variety of tasks and have attracted extensive attention in recent years. Various hashing approaches have been proposed to capture similarities between textual, visual, and cross-media information. However, most of the existing works use a bag-of-words methods to represent textual information. Since words with different forms may have similar meaning, semantic level text similarities can not be well processed in these methods. To address these challenges, in this paper, we propose a novel method called semantic cross-media hashing (SCMH), which uses continuous word representations to capture the textual similarity at the semantic level and use a deep belief network (DBN) to construct the correlation between different modalities. To demonstrate the effectiveness of the proposed method, we evaluate the proposed method on three commonly used cross-media data sets are used in this work. Experimental results show that the proposed method achieves significantly better performance than state-of-the-art approaches. Moreover, the efficiency of the proposed method is comparable to or better than that of some other hashing methods. Qi Zhang 0001, Yang Wang 0091, Xuanjing Huang 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | Retweet Behavior Prediction Using Hierarchical Dirichlet ProcessabstractThe task of predicting retweet behavior is an important and essential step for various social network applications, such as business intelligence, popular event prediction, and so on. Due to the increasing requirements, in recent years, the task has attracted extensive attentions. In this work, we propose a novel method using non-parametric statistical models to combine structural, textual, and temporal information together to predict retweet behavior. To evaluate the proposed method, we collect a large number of microblogs and their corresponding social networks from a real microblog service. Experimental results on the constructed dataset demonstrate that the proposed method can achieve better performance than state-of-the-art methods. The relative improvement of the the proposed over the method using only textual information is more than 38.5% in terms of F1-Score. Qi Zhang 0001, Yeyun Gong, Xuanjing Huang 0001 |
AAAI | 4 |
| 2015 | Gated Recursive Neural Network for Chinese Word SegmentationabstractXinchi Chen, Xipeng Qiu, Chenxi Zhu, Xuanjing Huang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Xinchi Chen, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2015 | A Re-ranking Model for Dependency Parser with Recursive Convolutional Neural NetworkabstractChenxi Zhu, Xipeng Qiu, Xinchi Chen, Xuanjing Huang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Xipeng Qiu, Xinchi Chen, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2015 | Who Will You "@"?abstractIn Twitter-like social networking services, people can use the "@" symbol to mention other users in tweets and send them a message or link to their profiles. In recent years, social media services are rapidly growing with thousands of millions of users participating in them every day. When the "@" symbol is entered, there should be an automatic suggestion function which recommends a small list of candidates in order to help users to easily identify and input usernames. In this paper, we present our work on building a recommendation system for the mention function in microblogging services. The recommendation strategy we used takes into consideration not only content of the microblog but also histories of candidate users. To better handle these textual information, we propose a novel method that extends the translation-based model. Experimental results on the dataset we collected from a real world microblogging service demonstrate that the proposed method outperforms state-of-the-art approaches. Yeyun Gong, Qi Zhang 0001, Xuyang Sun, Xuanjing Huang 0001 |
CIKM | 4 |
| 2015 | Long Short-Term Memory Neural Networks for Chinese Word SegmentationabstractCurrently most of state-of-the-art methods for Chinese word segmentation are based on supervised learning, whose features are mostly extracted from a local context.These methods cannot utilize the long distance information which is also crucial for word segmentation.In this paper, we propose a novel neural network model for Chinese word segmentation, which adopts the long short-term memory (LSTM) neural network to keep the previous important information in memory cell and avoids the limit of window size of local context.Experiments on PKU, MSRA and CTB6 benchmark datasets show that our model outperforms the previous neural network models and state-of-the-art methods. Xinchi Chen, Xipeng Qiu, Pengfei Liu 0003, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2015 | Sentence Modeling with Gated Recursive Neural NetworkabstractRecently, neural network based sentence modeling methods have achieved great progress.Among these methods, the recursive neural networks (RecNNs) can effectively model the combination of the words in sentence.However, RecNNs need a given external topological structure, like syntactic tree.In this paper, we propose a gated recursive neural network (GRNN) to model sentences, which employs a full binary tree (FBT) structure to control the combinations in recursive structure.By introducing two kinds of gates, our model can better model the complicated combinations of features.Experiments on three text classification datasets show the effectiveness of our model. Xinchi Chen, Xipeng Qiu, Shiyu Wu, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2015 | Transition-based Dependency Parsing Using Two Heterogeneous Gated Recursive Neural NetworksabstractRecently, neural network based dependency parsing has attracted much interest, which can effectively alleviate the problems of data sparsity and feature engineering by using the dense features.However, it is still a challenge problem to sufficiently model the complicated syntactic and semantic compositions of the dense features in neural network based methods.In this paper, we propose two heterogeneous gated recursive neural networks: tree structured gated recursive neural network (Tree-GRNN) and directed acyclic graph structured gated recursive neural network (DAG-GRNN).Then we integrate them to automatically learn the compositions of the dense features for transition-based dependency parsing.Specifically, Tree-GRNN models the feature combinations for the trees in stack, which already have partial dependency structures.DAG-GRNN models the feature combinations of the nodes whose dependency relations have not been built yet.Experiment results on two prevalent benchmark datasets (PTB3 and CTB5) show the effectiveness of our proposed model. Xinchi Chen, Yaqian Zhou 0001, Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2015 | Hashtag Recommendation Using Dirichlet Process Mixture Models Incorporating Types of HashtagsabstractIn recent years, the task of recommending hashtags for microblogs has been given increasing attention.Various methods have been proposed to study the problem from different aspects.However, most of the recent studies have not considered the differences in the types or uses of hashtags.In this paper, we introduce a novel nonparametric Bayesian method for this task.Based on the Dirichlet Process Mixture Models (DPMM), we incorporate the type of hashtag as a hidden variable.The results of experiments on the data collected from a real world microblogging service demonstrate that the proposed method outperforms stateof-the-art methods that do not consider these aspects.By taking these aspects into consideration, the relative improvement of the proposed method over the state-of-theart methods is around 12.2% in F1-score. Yeyun Gong, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP | 3 |
| 2015 | Multi-Timescale Long Short-Term Memory Neural Network for Modelling Sentences and DocumentsabstractNeural network based methods have obtained great progress on a variety of natural language processing tasks.However, it is still a challenge task to model long texts, such as sentences and documents.In this paper, we propose a multi-timescale long short-term memory (MT-LSTM) neural network to model long texts.MT-LSTM partitions the hidden states of the standard LSTM into several groups.Each group is activated at different time periods.Thus, MT-LSTM can model very long documents as well as short sentences.Experiments on four benchmark datasets show that our model outperforms the other neural models in text classification task. Pengfei Liu 0003, Xipeng Qiu, Xinchi Chen, Shiyu Wu, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2015 | Learning Context-Sensitive Word Embeddings with Neural Tensor Skip-Gram Model
Pengfei Liu 0003, Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 3 |
| 2015 | Convolutional Neural Tensor Network Architecture for Community-Based Question Answering
Xipeng Qiu, Xuanjing Huang 0001 |
IJCAI | 2 |
| 2015 | Overview of the NLPCC 2015 Shared Task: Chinese Word Segmentation and POS Tagging for Micro-blog TextsabstractIn this paper, we give an overview for the shared task at the 4th CCF Conference on Natural Language Processing & Chinese Computing (NLPCC 2015): Chinese word segmentation and part-of-speech (POS) tagging for micro-blog texts. Different with the popular used newswire datasets, the dataset of this shared task consists of the relatively informal micro-texts. The shared task has two sub-tasks: (1) individual Chinese word segmentation and (2) joint Chinese word segmentation and POS Tagging. Each subtask has three tracks to distinguish the systems with different resources. We first introduce the dataset and task, then we characterize the different approaches of the participating systems, report the test results, and provide a overview analysis of these results. An online system is available for open registration and evaluation at http://nlp.fudan.edu.cn/nlpcc2015 . Xipeng Qiu, Liusong Yin, Shiyu Wu, Xuanjing Huang 0001 |
NLPCC | 5 |
| 2015 | Transition-Based Dependency Parsing with Long Distance CollocationsabstractLong distance dependency relation is one of the main challenges for the state-of-the-art transition-based dependency parsing algorithms. In this paper, we propose a method to improve the performance of transition-based parsing with long distance collocations. With these long distance collocations, our method provides an approximate global view of the entire sentence, which is a little bit similar to top-down parsing. To further improve the accuracy of decision, we extend the set of parsing actions with two more fine-grained actions based on the types of arcs. Experimental results show that our method improve the performance of parsing effectively, especially for long sentence. Xipeng Qiu, Xuanjing Huang 0001 |
NLPCC | 3 |
| 2014 | A Generative Model for Identifying Target Companies of Microblogs
Yeyun Gong, Yaqian Zhou 0001, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 5 |
| 2014 | Automatic Corpus Expansion for Chinese Word Segmentation by Exploiting the Redundancy of Web Information
Xipeng Qiu, Chaochao Huang, Xuanjing Huang 0001 |
COLING | 3 |
| 2014 | Time-aware Personalized Hashtag Recommendation on Social Media
Qi Zhang 0001, Yeyun Gong, Xuyang Sun, Xuanjing Huang 0001 |
COLING | 4 |
| 2014 | Continuous word embeddings for detecting local text reuses at the semantic levelabstractText reuse is a common phenomenon in a variety of user-generated content. Along with the quick expansion of social media, reuses of local text are occurring much more frequently than ever before. The task of detecting these local reuses serves as an essential step for many applications. It has attracted extensive attention in recent years. However, semantic level similarities have not received consideration in most previous works. In this paper, we introduce a novel method to efficiently detect local reuses at the semantic level for large scale problems. We propose to use continuous vector representations of words to capture the semantic level similarities between short text segments. In order to handle tens of billions of documents, methods based on information geometry and hashing methods are introduced to aggregate and map text segments presented by word embeddings to binary hash codes. Experimental results demonstrate that the proposed methods achieve significantly better performance than state-of-the-art approaches in all six document collections belonging to four different categories. At some recall levels, the precisions of the proposed method are even 10 times higher than previous methods. Moreover, the efficiency of the proposed method is comparable to or better than that of some other hashing methods. Qi Zhang 0001, Jihua Kang, Xuanjing Huang 0001 |
SIGIR | 4 |
| 2014 | Chinese-English mixed text normalizationabstractAlong with the expansion of globalization, multilingualism has become a popular social phenomenon. More than one language may occur in the context of a single conversation. This phenomenon is also prevalent in China. A huge variety of informal Chinese texts contain English words, especially in emails, social media, and other user generated informal contents. Since most of the existing natural language processing algorithms were designed for processing monolingual information, mixed multilingual texts cannot be well analyzed by them. Hence, it is of critical importance to preprocess the mixed texts before applying other tasks. In this paper, we firstly analyze the phenomena of mixed usage of Chinese and English in Chinese microblogs. Then, we detail the proposed two-stage method for normalizing mixed texts. We propose to use a noisy channel approach to translate in-vocabulary words into Chinese. For better incorporating the historical information of users, we introduce a novel user aware neural network language model. For the out-of-vocabulary words (such as pronunciations, informal expressions and et al.), we propose to use a graph-based unsupervised method to categorize them. Experimental results on a manually annotated microblog dataset demonstrate the effectiveness of the proposed method. We also evaluate three natural language parsers with and without using the proposed method as the preprocessing step. From the results, we can see that the proposed method can significantly benefit other NLP tasks in processing mixed text. Qi Zhang 0001, Huan Chen 0012, Xuanjing Huang 0001 |
WSDM | 3 |
| 2013 | Map search via a factor graph modelabstractMap search has received considerable attention in recent years. With map search, users can specify target locations with textual queries. However, these queries do not always include well-formed addresses or place names. They may contain transpositions, misspellings, fragments and so on. Queries may significantly differ from items stored in the spatial database. In this paper, we propose to connect this task to the semi-structured retrieval problem. A novel factor graph-based semi-structured retrieval framework is introduced to incorporate concept weighting, attribute selection, and word-based similarity metrics together. We randomly sampled a number of queries from logs of a commercial map search engine and manually labeled their categories and relevant results for analysis and evaluation. The results of several experimental comparisons demonstrate that our method outperforms both state-of-the-art semi-structured retrieval methods and some commercial systems in retrieving freeform location queries. Qi Zhang 0001, Jihua Kang, Yeyun Gong, Huan Chen 0012, Yaqian Zhou 0001, Xuanjing Huang 0001 |
CIKM | 6 |
| 2013 | A pattern-based selective recrawling approach for object-level vertical searchabstractTraditional recrawling methods learn navigation patterns in order to crawl related web pages. However, they cannot remove the redundancy found on the web, especially at the object level. To deal with this problem, we propose a new hypertext resource discovery method, called ``selective recrawling'' for object-level vertical search applications. The goal of selective recrawling is to automatically generate URL patterns, then select those pages that have the widest coverage, and least irrelevance and redundancy relative to a pre-defined vertical domain. This method only requires a few seed objects and can select the set of URL patterns that covers the greatest number of objects. The selected set can continue to be used for some time to recrawl web pages and can be renewed periodically. This leads to significant savings in hardware and network resources. Yaqian Zhou 0001, Qi Zhang 0001, Xuanjing Huang 0001, Lide Wu |
CIKM | 3 |
| 2013 | Joint Chinese Word Segmentation and POS Tagging on Heterogeneous Annotated Corpora with Multiple Task LearningabstractChinese word segmentation and part-ofspeech tagging (S&T) are fundamental steps for more advanced Chinese language processing tasks.Recently, it has attracted more and more research interests to exploit heterogeneous annotation corpora for Chinese S&T.In this paper, we propose a unified model for Chinese S&T with heterogeneous annotation corpora.We first automatically construct a loose and uncertain mapping between two representative heterogeneous corpora, Penn Chinese Treebank (CTB) and PKU's People's Daily (PPD).Then we regard the Chinese S&T with heterogeneous corpora as two "related" tasks and train our model on two heterogeneous corpora simultaneously.Experiments show that our method can boost the performances of both of the heterogeneous corpora by using the shared information, and achieves significant improvements over the state-of-the-art methods. Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 3 |
| 2013 | Discourse Level Explanatory Relation Extraction from Product Reviews Using First-Order LogicabstractExplanatory sentences are employed to clarify reasons, details, facts, and so on.High quality online product reviews usually include not only positive or negative opinions, but also a variety of explanations of why these opinions were given.These explanations can help readers get easily comprehensible information of the discussed products and aspects.Moreover, explanatory relations can also benefit sentiment analysis applications.In this work, we focus on the task of identifying subjective text segments and extracting their corresponding explanations from product reviews in discourse level.We propose a novel joint extraction method using firstorder logic to model rich linguistic features and long distance constraints.Experimental results demonstrate the effectiveness of the proposed method. Qi Zhang 0001, Huan Chen 0012, Jihua Kang, Xuanjing Huang 0001 |
EMNLP | 5 |
| 2013 | Learning Topical Translation Model for Microblog Hashtag Suggestion
Zhuoye Ding, Xipeng Qiu, Qi Zhang 0001, Xuanjing Huang 0001 |
IJCAI | 4 |
| 2013 | Chinese Named Entity Abbreviation Generation Using First-Order Logic
Huan Chen 0012, Qi Zhang 0001, Xuanjing Huang 0001 |
IJCNLP | 4 |
| 2013 | Detecting Spammers in Community Question Answering
Zhuoye Ding, Yeyun Gong, Yaqian Zhou 0001, Qi Zhang 0001, Xuanjing Huang 0001 |
IJCNLP | 5 |
| 2013 | Understanding the Semantic Intent of Natural Language Query
Qi Zhang 0001, Xuanjing Huang 0001 |
IJCNLP | 3 |
| 2012 | Discovering logical knowledge for deep question answeringabstractMost open-domain question answering systems achieve better performances with large corpora, such as Web, by taking advantage of information redundancy. However, explicit answers are not always mentioned in the corpus, many answers are implicitly contained and can only be deducted by inference. In this paper, we propose an approach to discover logical knowledge for deep question answering, which automatically extracts knowledge in an unsupervised, domain-independent manner from background texts and reasons out implicit answers for the questions. Firstly, we use semantic role labeling to transform natural language expressions to predicates in first-order logic. Then we use association analysis to uncover the implicit relations among these predicates and build propositions for inference. Since our knowledge is drawn from different sources, we use Markov logic to merge multiple knowledge bases without resolving their inconsistencies. Our experiments show that these propositions can improve the performance of question answering significantly. Xipeng Qiu, Ling Cao, Xuanjing Huang 0001 |
CIKM | 4 |
| 2012 | Selecting expansion terms as a set via integer linear programmingabstractPseudo-relevance feedback via query expansion has been widely studied from various perspectives in the past decades. Its effectiveness in improving retrieval effectiveness has been shown in many tasks. A variety of criteria were proposed to select additional terms for the original queries. However, most of the existing methods weight and select terms individually and do not consider the impact of term-to-term relationship. In this paper, we first examine the influence of combinations of terms through data analysis, which demonstrate the significant effect of term-to-term relationship on retrieval effectiveness. Then, to address this problem, we formalize the query expansion task as an integer linear programming (ILP) problem. The model combines the weights learned from a supervised method for individual terms, and integrates constraints to capture relations between terms. Finally, three standard TREC collections are used to evaluate the proposed method. Experimental results demonstrate that the proposed method can significantly improve the effectiveness of retrieval. Qi Zhang 0001, Xuanjing Huang 0001 |
CIKM | 3 |
| 2012 | Part-of-Speech Tagging for Chinese-English Mixed Texts with Dynamic Features
Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP-CoNLL | 5 |
| 2012 | Learning hash codes for efficient content reuse detectionabstractContent reuse is extremely common in user generated mediums. Reuse detection serves as be the basis for many applications. However, along with the explosion of Internet and continuously growing uses of user generated mediums, the task becomes more critical and difficult. In this paper, we present a novel efficient and scalable approach to detect content reuse. We propose a new signature generation algorithm, which is based on learned hash functions for words. In order to deal with tens of billions of documents, we implement the detection approach on graphical processing units (GPUs). The experimental comparison in this paper involves studies of efficiency and effectiveness of the proposed approach in different types of document collections, including ClueWeb09, Tweets2011, and so on. Experimental results show that the proposed approach can achieve the same detection rates with state-of-the-art systems while uses significantly less execution time than them (from 400X to 1500X speedup). Qi Zhang 0001, Zhuoye Ding, Xuanjing Huang 0001 |
SIGIR | 4 |
| 2012 | Recognizing Inference in Texts with Markov Logic NetworksabstractRecognizing inference in texts (RITE) attracts growing attention of natural language processing (NLP) researchers in recent years. In this article, we propose a novel approach to recognize inference with probabilistic logical reasoning. Our approach is built on Markov logic networks (MLNs) framework, which is a probabilistic extension of first-order logic. We design specific semantic rules based on the surface, syntactic, and semantic representations of texts, and map these rules to logical representations. We also extract information from some knowledge bases as common sense logic rules. Then we utilize MLNs framework to make predictions with combining statistical and logical reasoning. Experiment results shows that our system can achieve better performance than state-of-the-art RITE systems. Xipeng Qiu, Ling Cao, Xuanjing Huang 0001 |
ACM Trans. Asian Lang. Inf. Process. | 4 |
| 2011 | Labelwise Margin Maximization for Sequence Labeling
Xipeng Qiu, Xuanjing Huang 0001 |
CICLing (1) | 3 |
| 2011 | Structural Opinion Mining for Graph-based Sentiment Representation
Yuanbin Wu, Qi Zhang 0001, Xuanjing Huang 0001, Lide Wu |
EMNLP | 3 |
| 2011 | Keyphrase Extraction from Online News Using Binary Integer Programming
Zhuoye Ding, Qi Zhang 0001, Xuanjing Huang 0001 |
IJCNLP | 3 |
| 2011 | Efficient Near-Duplicate Detection for Q&A Forum
Qi Zhang 0001, Xuanjing Huang 0001 |
IJCNLP | 3 |
| 2011 | A Fast Accurate Two-stage Training Algorithm for L1-regularized CRFs with Heuristic Line Search Strategy
Jinlong Zhou, Xipeng Qiu, Xuanjing Huang 0001 |
IJCNLP | 3 |
| 2011 | An Effective Feature Selection Method for Text Categorization
Xipeng Qiu, Jinlong Zhou, Xuanjing Huang 0001 |
PAKDD (1) | 3 |
| 2011 | Opinion Mining with Sentiment GraphabstractOpinion mining became an active research topic in recent years due to its wide range of applications. A number of companies offer opinion mining services. One problem that has not been well studied so far is the representation model. In this paper, we propose a novel sentence level sentiment representation model. By taking the observation that lots of sentences which have complicated opinion relations can not be represented well by slots filling or feature-based model, the novel representation model sentiment graph is described in this paper. A supervised structural learning method is presented and used to construct sentiment graphs from sentences. Experimental results in a manually labeled corpus are given to show the effectiveness of the proposed approach. Qi Zhang 0001, Yuanbin Wu, Xuanjing Huang 0001 |
Web Intelligence | 4 |
| 2010 | Mining Uncertain Sentences with Multiple Instance Learning
Xipeng Qiu, Xuanjing Huang 0001 |
ADMA (1) | 3 |
| 2010 | 2D Trie for Fast Parsing
Xian Qian, Qi Zhang 0001, Xuanjing Huang 0001, Lide Wu |
COLING | 3 |
| 2010 | Joint Training and Decoding Using Virtual Nodes for Cascaded Segmentation and Tagging Tasks
Xian Qian, Qi Zhang 0001, Yaqian Zhou 0001, Xuanjing Huang 0001, Lide Wu |
EMNLP | 4 |
| 2010 | Efficient partial-duplicate detection based on sequence matchingabstractWith the ever-increasing growth of the Internet, numerous copies of documents become serious problem for search engine, opinion mining and many other web applications. Since partial-duplicates only contain a small piece of text taken from other sources and most existing near-duplicate detection approaches focus on document level, partial duplicates can not be dealt with well. In this paper, we propose a novel algorithm to realize the partial-duplicate detection task. Besides the similarities between documents, our proposed algorithm can simultaneously locate the duplicated parts. The main idea is to divide the partial-duplicate detection task into two subtasks: sentence level near-duplicate detection and sequence matching. For evaluation, we compare the proposed method with other approaches on both English and Chinese web collections. Experimental results appear to support that our proposed method is effectively and efficiently to detect both partial-duplicates on large web collections. Qi Zhang 0001, Yue Zhang 0004, Haomin Yu, Xuanjing Huang 0001 |
SIGIR | 4 |
| 2010 | Selective recrawling for object-level vertical searchabstractIn this paper we propose a novel recrawling method based on navigation patterns called Selective Recrawling. The goal of selective recrawling is to automatically select page collections that have large coverage and little redundancy to a pre-defined vertical domain. It only requires several seed objects and can select a set of URL patterns to cover most objects. The selected set can be used to recrawl the web pages for quite a period of time and renewed periodically. Experiments on local event data show that our method can greatly reduce the downloading of web pages while keep the comparative object coverage. Yaqian Zhou 0001, Mengjing Jiang, Qi Zhang 0001, Xuanjing Huang 0001, Lide Wu |
WWW | 4 |
| 2009 | A unified relevance model for opinion retrievalabstractRepresenting the information need is the greatest challenge for opinion retrieval. Typical queries for opinion retrieval are composed of either just content words, or content words with a small number of cue "opinion" words. Both are inadequate for retrieving opinionated documents. In this paper, we develop a general formal framework--the opinion relevance model--to represent an information need for opinion retrieval. We explore a series of methods to automatically identify the most appropriate opinion words for query expansion, including using query independent sentiment resources. We also propose a relevance feedback-based approach to extract opinion words. Both query-independent and query-dependent methods can also be integrated into a more effective mixture relevance model. Finally, opinion retrieval experiments are presented for the Blog06 and COAE08 text collections. The results show that, significant improvements can always be obtained by this opinion relevance model whether sentiment resources are available or not. Xuanjing Huang 0001, W. Bruce Croft |
CIKM | 1 |
| 2009 | Phrase Dependency Parsing for Opinion Mining
Yuanbin Wu, Qi Zhang 0001, Xuanjing Huang 0001, Lide Wu |
EMNLP | 3 |
| 2009 | Sparse higher order conditional random fields for improved sequence labelingabstractIn real sequence labeling tasks, statistics of many higher order features are not sufficient due to the training data sparseness, very few of them are useful. We describe Sparse Higher Order Conditional Random Fields (SHO-CRFs), which are able to handle local features and sparse higher order features together using a novel tractable exact inference algorithm. Our main insight is that states and transitions with same potential functions can be grouped together, and inference is performed on the grouped states and transitions. Though the complexity is not polynomial, SHO-CRFs are still efficient in practice because of the feature sparseness. Experimental results on optical character recognition and Chinese organization name recognition show that with the same higher order feature set, SHO-CRFs significantly outperform previous approaches. Xian Qian, Xiaoqian Jiang, Qi Zhang 0001, Xuanjing Huang 0001, Lide Wu |
ICML | 4 |
| 2009 | Template-independent wrapper for web forumsabstractThis paper presents a novel work on the task of extracting data from Web forums. Millions of users contribute rich information to Web forum everyday, which has become an important resource for manyWeb applications, such as product opinion retrieval, social network analysis, and so on. The novelty of the proposed algorithm is that it can not only extract the pure text but also distinguish between the original post and replies. Experimental results on a large number of real Web forums indicate that the proposed algorithm can correctly ex Qi Zhang 0001, Xuanjing Huang 0001, Lide Wu |
SIGIR | 3 |
| 2009 | Mining product reviews based on shallow dependency parsingabstractThis paper presents a novel method for mining product reviews, where it mines reviews by identifying product features, expressions of opinions and relations between them. By taking advantage of the fact that most of product features are phrases, a concept of shallow dependency parsing is introduced, which extends traditional dependency parsing to phrase level. This concept is then implemented for extracting relation between product features and expressions of opinions. Experimental evaluations show that the mining task can benefit from shallow dependency parsing. Qi Zhang 0001, Yuanbin Wu, Tao Li 0001, Mitsunori Ogihara, Xuanjing Huang 0001 |
SIGIR | 6 |
| 2009 | Using query expansion in graph-based approach for query-focused multi-document summarization
Lide Wu, Xuanjing Huang 0001 |
Inf. Process. Manag. | 3 |
| 2008 | Answering Definition Question: Ranking for Top-kabstractAs an important form of complex questions, definition question attracts much attention from QA researchers. For many of the definition question answering systems, it is a core step to rank the candidate answer sentences, so that the top-k in the ranked list can be extracted. We integrate these evidences as features into a whole framework, and propose a novel method to learning weights of these features to rank the candidate answer sentences. Xipeng Qiu, Xuanjing Huang 0001, Lide Wu |
ECAI | 3 |
| 2006 | Mining the Relation between Sentiment Expression and Target Using Dependency of Words
Zhongchao Fei, Xuanjing Huang 0001, Lide Wu |
PACLIC | 2 |
| 2005 | Answering Definition Questions Using Web Knowledge Bases
Zhushuo Zhang, Yaqian Zhou 0001, Xuanjing Huang 0001, Lide Wu |
IJCNLP | 3 |
| 2004 | A Novel Pattern Learning Method for Open Domain Question Answering
Yongping Du, Xuanjing Huang 0001, Lide Wu |
IJCNLP | 2 |
| 2004 | BBS Based Hot Topic Retrieval Using Back-Propagation Neural Network
Lan You, Yongping Du, Jiayin Ge, Xuanjing Huang 0001, Lide Wu |
IJCNLP | 4 |
| 2004 | Exploring Various Features to Optimize Hot Topic Retrieval on WEB
Lan You, Xuanjing Huang 0001, Lide Wu, Hao Yu 0005, Fumihito Nishino |
ISNN (1) | 2 |