VLDB 2026 Research / reviewers in the wild / expert
Shizhu He
dblp:136/8650
· DBLP profile ↗
68ranked-venue papers
4as first author
42since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 4 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TaREx: Reinforcement Learning for Code-Driven Table Reasoning
Fangyu Lei, Jinxiang Meng, Shizhu He, Jun Zhao 0001, Kang Liu 0001 |
AAAI | 4 |
| 2026 | SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel PruningabstractLong-context inference in large language models (LLMs) is increasingly constrained by the KV cache bottleneck: memory usage grows linearly with sequence length, while attention computation scales quadratically. Existing approaches address this issue by compressing the KV cache along the temporal axis through strategies such as token eviction or merging to reduce memory and computational overhead. However, these methods often neglect fine-grained importance variations across feature dimensions (i.e., the channel axis), thereby limiting their ability to effectively balance efficiency and model accuracy. In reality, we observe that channel saliency varies dramatically across both queries and positions: certain feature channels carry near-zero information for a given query, while others spike in relevance. To address this oversight, we propose SPARK, a training-free plug-and-play method that applies unstructured sparsity by pruning KV at the channel level, while dynamically restoring the pruned entries during attention score computation. Notably, our approach is orthogonal to existing KV compression and quantization techniques, making it compatible for integration with them to achieve further acceleration. By reducing channel-level redundancy, SPARK enables processing of longer sequences within the same memory budget. For sequences of equal length, SPARK not only preserves or improves model accuracy but also reduces KV cache storage by over 30% compared to eviction-based methods. Furthermore, even in an aggressive pruning ratio of 80%, SPARK maintains performance with less degradation than 5% compared to the based eviction method, demonstrating robustness and effectiveness. Our code will be available at \url{https://github.com/AMD-AIG-AIMA/AMD-Spark}. Huanxuan Liao, Yixing Xu, Shizhu He, Xuanwu Yin, Dong Li 0025, Emad Barsoum, Jun Zhao 0001, Kang Liu 0001 |
AAAI | 3 |
| 2026 | Seeing Is Believing: Grounding Long-Video Understanding in Spatio-Temporal Visual EvidenceabstractAlthough Vision Language Models (VLMs) have excelled at image and video understanding, applying them to hour-long videos is held back by two interrelated challenges: exorbitant computational expense and a qualitative breakdown in long-term temporal reasoning. Thus, models tend to generate answers based on speculation instead of solid visual facts, causing both factually incorrect and plausible hallucinations. This problem is compounded by current benchmarks that, by only emphasizing final answers, lack an effective mechanism to check whether reasoning is substantiated by specific visual evidence. This makes it hard to differentiate between true understanding and pretend comprehension, inhibiting targeted model refinement. To address these interrelated challenges of model fragility and evaluation weakness, we adopt a twofold strategy. First, we present EV²-Bench, a large-scale benchmark that breaks new ground by an evaluation paradigm built upon spatio-temporal visual evidence, forcing models to justify answers with checkable hints. Second, we put forward DynamicSelect, an adaptive token compression system that efficiently condenses salient information by a dynamic semantic selector and a hierarchical compression strategy. Comprehensive experiments demonstrate that DynamicSelect significantly outperforms the baselines on EV²-Bench as well as other public benchmarks. Our study offers not only a more effective approach to long-video understanding but also a more stringent evaluation paradigm, indicating the way toward more robust models. Zhaoyang Wei, Guohua Gao, Yanchao Hao, Wenchao Ding 0007, Shizhu He, Xuehui Yu |
AAAI | 8 |
| 2026 | Harmonizing the Past, Present, and Future: A Null-Space Constrained Region-Specific Method for Continual Learning in LLMsabstractJinhui Chen, Shizhu He, Xingchang Yang, Huanxuan Liao, Yequan Wang, Xiangwen Liao, Wenhao Teng, Kang Liu, Jun Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shizhu He, Xingchang Yang, Huanxuan Liao, Yequan Wang, Xiangwen Liao, Wenhao Teng, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 2 |
| 2026 | Spectral Disentanglement: Rank-Aware Task Adaptation for Rehearsal-free Continual Learning in LLMsabstractHuanxuan Liao, Shizhu He, Yupu Hao, Yequan Wang, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Huanxuan Liao, Shizhu He, Yupu Hao, Yequan Wang, Wenhao Teng, Xiangwen Liao, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 2 |
| 2026 | GATE: Graph-based Adaptive Tool Evolution Across Diverse TasksabstractJianwen Luo, Yiming Huang, Jinxiang Meng, Fangyu Lei, Shizhu He, Xiao Liu, Shanshan Jiang, Bin Dong, Jun Zhao, Kang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jinxiang Meng, Fangyu Lei, Shizhu He, Shanshan Jiang 0001, Bin Dong 0003, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 5 |
| 2026 | Shuttle Between Symbolic Instructions and Neural Parameters of Large Language ModelsabstractWangtao Sun, Haotian Xu, Huanxuan Liao, Xuanqing Yu, Zhongtao Jiang, Shizhu He, Jun Zhao, Kang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wangtao Sun, Huanxuan Liao, Xuanqing Yu, Zhongtao Jiang, Shizhu He, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 6 |
| 2025 | Neural-Symbolic Collaborative Distillation: Advancing Small Language Models for Complex Reasoning TasksabstractIn this paper, we propose Neural-Symbolic Collaborative Distillation (NesyCD), a novel knowledge distillation method for learning the complex reasoning abilities of Large Language Models (LLMs, e.g., \textgreater 13B). We argue that complex reasoning tasks are difficult for Small Language Models (SLMs, e.g., $\leq$ 7B), as these tasks demand not only general cognitive abilities but also specialized knowledge, which is often sparse and difficult for these neural-based SLMs to effectively capture. Therefore, NesyCD distills the general capabilities and specialized knowledge in LLMs using different manners.On the one hand, we distill only general abilities from teacher LLMs into the student SLMs of parameterized neural networks. On the other hand, for the specialized abilities and uncommon knowledge of a complex reasoning task, we employ a symbolic knowledge distillation approach to obtain and store the specialized knowledge within a symbolic knowledge base (KB).By decoupling general and specialized capabilities, the proposed NesyCD can achieve superior performance cost-effectively, utilizing smaller models and blending parameterized neural networks with symbolic KB. Moreover, the specialized KB generalizes well and is comprehended and manipulated by humans.Our experiments show that NesyCD significantly boosts SLMs' complex reasoning performance on in-domain (BBH, GSM8K) and out-of-domain (AGIEval, ARC) datasets. Notably, our approach enabled the LLaMA3-8B and Qwen2-7B to surpass GPT-3.5-turbo in performance and come close to matching LLaMA3-70B, despite the latter having nine times more parameters. Huanxuan Liao, Shizhu He, Yuanzhe Zhang, Kang Liu 0001, Jun Zhao 0001 |
AAAI | 2 |
| 2025 | HFF-Tracker: A Hierarchical Fine-grained Fusion Tracker for Referring Multi-Object TrackingabstractReferring Multi-Object Tracking (RMOT) aims to track multiple objects based on a provided language expression. Although prior studies have sought to accomplish this by integrating an textual module into the multi-object tracker, these methods combine text and image features in a basic way, neglecting the importance of text features. In this study, we propose a Hierarchical Fine-grained text-image Fusion tracker, named HFF-Tracker, which can perform fine-grained fusion of pixel-level visual features and text features across various semantic levels. Specifically, we have devised a Hierarchical Multi-Modal Fusion (HMMF) module to merge text and image features at an early stage in a hierarchical and detailed manner. The Text-Guided Decoder (TGD) is designed to provide the query with prior semantic information during the decoding process. Additionally, we have crafted a Text-Guided Prediction Head (TGPH) that utilizes text information to enhance the performance of the prediction head. Furthermore, we have implemented an adaptive Look-Back training strategy to maximize the utilization of valuable labeled data. Extensive experiments on the Refer-KITTI dataset and the Refer-KITTI-V2 dataset demonstrate that our proposed HFF-Tracker outperforms other state-of-the-art methods with remarkable margins. Zeyong Zhao, Yanchao Hao, Qingbin Liu, Dianbo Sui, Shizhu He, Xi Chen 0003 |
AAAI | 7 |
| 2025 | Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language ModelsabstractLarge Language Models (LLMs) offer a transparent brain with accessible parameters that encode extensive knowledge, which can be analyzed, located and transferred.Consequently, a key research challenge is to transcend traditional knowledge transfer paradigms rooted in symbolic language and achieve genuine Parametric Knowledge Transfer (PKT).Significantly, exploring effective methods for transferring knowledge across LLMs of different scales through parameters presents an intriguing and valuable research direction.In this paper, we first demonstrate Alignment in parametric space is the fundamental prerequisite to achieve successful cross-scale PKT.We redefine the previously explored knowledge transfer as Post-Align PKT (PostPKT), which utilizes extracted parameters for LoRA initialization and requires subsequent fine-tune for alignment.Hence, to reduce cost for further fine-tuning, we introduce a novel Pre-Align PKT (PrePKT) paradigm and propose a solution called LaTen (Locate-Then-Align) that aligns the parametric spaces of LLMs across scales only using several training steps without following training.Comprehensive experiments on four benchmarks demonstrate that both PostPKT and PrePKT face challenges in achieving consistently stable transfer.Through in-depth analysis, we identify Neural Incompatibility as the ethological and parametric structural differences between LLMs of varying scales, presenting fundamental challenges to achieving effective PKT.These findings provide fresh insights into the parametric architectures of LLMs and highlight promising directions for future research on efficient PKT.Our code is available at https://github.com/ Trae1ounG/Neural_Incompatibility. Yuqiao Tan, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 2 |
| 2025 | SKIntern: Internalizing Symbolic Knowledge for Distilling Better CoT Capabilities into Small Language ModelsabstractSmall Language Models (SLMs) are attracting attention due to the high computational demands and privacy concerns of Large Language Models (LLMs). Some studies fine-tune SLMs using Chains of Thought (CoT) data distilled from LLMs, aiming to enhance their reasoning ability. Furthermore, Some CoT distillation methods introduce external symbolic knowledge into the generation process to improve the limited knowledge memory, reasoning ability and out-of-domain (OOD) generalization of SLMs. However, the introduction of symbolic knowledge increases computational overhead and introduces potential noise. In this paper, we introduce SKIntern, an innovative approach that empowers SLMs to internalize symbolic knowledge and few-shot examples gradually through a progressive fine-tuning process, guided by a predefined linear decay schedule under curriculum learning. By efficiently internalizing knowledge, SKIntern reduces computational overhead and speeds up the reasoning process by focusing solely on the question during inference. It outperforms state-of-the-art baselines by over 5%, while reducing inference costs (measured in FLOPs) by up to 4\times across a wide range of SLMs in both in-domain (ID) and out-of-domain (OOD) tasks. Our code will be available at https://github.com/Xnhyacinth/SKIntern. Huanxuan Liao, Shizhu He, Yupu Hao, Yuanzhe Zhang, Jun Zhao 0001, Kang Liu 0001 |
COLING | 2 |
| 2025 | Awakening Augmented Generation: Learning to Awaken Internal Knowledge of Large Language Models for Question AnsweringabstractRetrieval-Augmented-Generation and Generation-Augmented-Generation have been proposed to enhance the knowledge required for question answering with Large Language Models (LLMs) by leveraging richer context. However, the former relies on external resources, and both require incorporating explicit documents into the context, which increases execution costs and susceptibility to noise data during inference. Recent works indicate that LLMs model rich knowledge, but it is often not effectively activated and awakened. Inspired by this, we propose a novel knowledge-augmented framework, Awakening-Augmented-Generation (AAG), which mimics the human ability to answer questions using only thinking and recalling to compensate for knowledge gaps, thereby awaking relevant knowledge in LLMs without relying on external resources. AAG consists of two key components for awakening richer context. Explicit awakening fine-tunes a context generator to create a synthetic, compressed document that functions as symbolic context. Implicit awakening utilizes a hypernetwork to generate adapters based on the question and synthetic document, which are inserted into LLMs to serve as parameter context. Experimental results on three datasets demonstrate that AAG exhibits significant advantages in both open-domain and closed-book settings, as well as in out-of-distribution generalization. Our code will be available at https://github.com/Xnhyacinth/IAG. Huanxuan Liao, Shizhu He, Yuanzhe Zhang, Shengping Liu, Kang Liu 0001, Jun Zhao 0001 |
COLING | 2 |
| 2025 | Why and How LLMs Benefit from Knowledge Introspection in Commonsense ReasoningabstractLarge Language Models (LLMs) can improve commonsense reasoning through generating intermediate knowledge.However, the effectiveness of this knowledge introspection is not always guaranteed.This paper first systematically investigates and reveals an introspection paradox: while simple introspection tends to benefit weaker models, it often degrades the performance of stronger ones, particularly on simpler tasks.Our deep analysis indicates that this paradox arises from a complex interplay among model capability, task difficulty and the quality of generated knowledge.Further interpretability analysis reveals the origins of low-quality knowledge generation.To better employ introspected knowledge in LLM, this paper proposes a training-free Adaptive Introspection Strategy that operates in two stages using only the model's internal states: Knowledge Detection, which dynamically identifies and discards potentially low-quality knowledge, and Knowledge Regeneration, which employs attention smoothing to guide the model away from harmful failure modes during knowledge generation.Extensive experiments on five Llama models with different sizes and eight commonsense reasoning benchmarks demonstrate that our approach effectively mitigates the limitations of standard introspection and has consistent performance gains across almost all settings. Chengfeng Zhao, Shizhu He, Shanshan Jiang 0001, Bin Dong 0003, Jun Zhao 0001, Kang Liu 0001 |
EMNLP | 2 |
| 2025 | LLaSA: Large Language and Structured Data AssistantabstractYao Xu, Shizhu He, Jiabei Chen, ZengXiangrong ZengXiangrong, Bingning Wang, Guang Liu, Jun Zhao, Kang Liu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shizhu He, Jiabei Chen, ZengXiangrong ZengXiangrong, Bingning Wang, Jun Zhao 0001, Kang Liu 0001 |
NAACL (Long Papers) | 2 |
| 2024 | ItD: Large Language Models Can Teach Themselves Induction through DeductionabstractAlthough Large Language Models (LLMs) are showing impressive performance on a wide range of Natural Language Processing tasks, researchers have found that they still have limited ability to conduct induction.Recent works mainly adopt "post processes" paradigms to improve the performance of LLMs on induction (e.g., the hypothesis search & refinement methods), but their performance is still constrained by the inherent inductive capability of the LLMs.In this paper, we propose a novel framework, Induction through Deduction (ItD), to enable the LLMs to teach themselves induction through deduction.The ItD framework is composed of two main components: a Deductive Data Generation module to generate induction data and a Naive Bayesian Induction module to optimize the fine-tuning and decoding of LLMs.Our empirical results showcase the effectiveness of ItD on two induction benchmarks, achieving relative performance improvement of 36% and 10% compared with previous state-of-the-art, respectively.Our ablation study verifies the effectiveness of two key modules of ItD.We also verify the effectiveness of ItD across different LLMs and deductors.The data and code of this paper can be found at https://github.com/forangel2014/ItD.(a) Hypothesis Search & Refinement (b) ItD Input …… Induction Result: Output Fine-tune … … Wangtao Sun, Xuanqing Yu, Shizhu He, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 5 |
| 2024 | BP4ER: Bootstrap Prompting for Explicit Reasoning in Medical Dialogue GenerationabstractMedical dialogue generation (MDG) has gained increasing attention due to its substantial practical value. Previous works typically employ a sequence-to-sequence framework to generate medical responses by modeling dialogue context as sequential text with annotated medical entities. While these methods have been successful in generating fluent responses, they fail to provide process explanations of reasoning and require extensive entity annotation. To address these limitations, we propose the method Bootstrap Prompting for Explicit Reasoning in MDG (BP4ER), which explicitly model MDG’s multi-step reasoning process and iteratively enhance this reasoning process. We employ a least-to-most prompting strategy to guide a large language model (LLM) in explicit reasoning, breaking down MDG into simpler sub-questions. These sub-questions build on answers from previous ones. Additionally, we also introduce two distinct bootstrapping techniques for prompting, which autonomously correct errors and facilitate the LLM’s explicit reasoning. This approach eliminates the need for entity annotation and increases the transparency of the MDG process by explicitly generating the intermediate reasoning chain. Experimental results on the two publicly datasets show that BP4ER outperforms state-of-the-art methods across both objective and subjective evaluation. Shizhu He |
LREC/COLING | 3 |
| 2024 | MoDE-CoTD: Chain-of-Thought Distillation for Complex Reasoning Tasks with Mixture of Decoupled LoRA-ExpertsabstractChain-of-thought Distillation (CoTD) aims at distilling Chain-of-thought (CoT) reasoning ability of large language models (LLMs) to much smaller student models. The core of CoTD is using a large teacher model to generate rationales and fine-tune smaller student models. However, current Chain-of-thought Distillation works have the following limitations: 1) Student models are separately distilled from specific reasoning tasks and lack a collaboration mechanism, hindering the enhancement of reasoning performance through collaboration among various reasoning tasks. 2) The parameter update of student models severely harms the CoT reasoning ability on other unseen reasoning tasks not included in the distillation process. In this work, we introduce a novel CoT Distillation method, MoDE-CoTD, which decouples the CoT reasoning abilities out of the student model by distilling multiple LoRA-Experts and freezing the parameters of the student model. Sequentially, LoRA-Experts are combined and adapted to handle both seen and unseen reasoning tasks, enabling collaboration among diverse reasoning tasks to further enhance CoT reasoning performance. Experimental results on 14 datasets (including 4 unseen datasets) demonstrate the strength of MoDE-CoTD, with an average accuracy gain of 6.3% on seen datasets and 7.8% on unseen datasets. Shizhu He, Zhao Yang 0004, Yang jun Jun, Kang Liu 0001, Jun Zhao 0001 |
LREC/COLING | 2 |
| 2024 | Towards Graph-hop Retrieval and Reasoning in Complex Question Answering over Textual DatabaseabstractIn textual question answering (TQA) systems, complex questions often require retrieving multiple textual fact chains with multiple reasoning steps. While existing benchmarks are limited to single-chain or single-hop retrieval scenarios. In this paper, we propose to conduct Graph-Hop —— a novel multi-chains and multi-hops retrieval and reasoning paradigm in complex question answering. We construct a new benchmark called ReasonGraphQA, which provides explicit and fine-grained evidence graphs for complex question to support comprehensive and detailed reasoning. In order to further study how graph-based evidential reasoning can be performed, we explore what form of Graph-Hop works best for generating textual evidence explanations in knowledge reasoning and question answering. We have thoroughly evaluated existing evidence retrieval and reasoning models on the ReasonGraphQA. Experiments highlight Graph-Hop is a promising direction for answering complex questions, but it still has certain limitations. We have further studied mitigation strategies to meet these challenges and discuss future directions. Minjun Zhu, Yixuan Weng, Shizhu He, Kang Liu 0001, Yang jun Jun, Jun Zhao 0001 |
LREC/COLING | 3 |
| 2024 | DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsabstractYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang, Fangyu Lei, Yifan Wei, Shizhu He, Lifu Huang, Xiao Liu, Jun Zhao, Kang Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Fangyu Lei, Yifan Wei 0001, Shizhu He, Lifu Huang, Jun Zhao 0001, Kang Liu 0001 |
EMNLP | 7 |
| 2024 | Does Large Language Model Contain Task-Specific Neurons?abstractLarge language models (LLMs) have demonstrated remarkable capabilities in comprehensively handling various types of natural language processing (NLP) tasks.However, there are significant differences in the knowledge and abilities required for different tasks.Therefore, it is important to understand whether the same LLM processes different tasks in the same way.Are there specific neurons in a LLM for different tasks?Inspired by neuroscience, this paper pioneers the exploration of whether distinct neurons are activated when a LLM handles different tasks.Compared with current research exploring the neurons of language and knowledge, task-specific neurons present a greater challenge due to their abstractness, diversity, and complexity.To address these challenges, this paper proposes a method for task-specific neuron localization based on Causal Gradient Variation with Special Tokens (CGVST).CGVST identifies task-specific neurons by concentrating on the most significant tokens during task processing, thereby eliminating redundant tokens and minimizing interference from non-essential neurons.Compared to traditional neuron localization methods, our approach can more effectively identify task-specific neurons.We conduct experiments across eight different public tasks.Experiments involving the inhibition and amplification of identified neurons demonstrate that our method can accurately locate task-specific neurons. Ran Song 0002, Shizhu He, Shuting Jiang, Yantuan Xian, Shengxiang Gao, Kang Liu 0001, Zhengtao Yu 0001 |
EMNLP | 2 |
| 2024 | Generate-on-Graph: Treat LLM as both Agent and KG for Incomplete Knowledge Graph Question AnsweringabstractTo address the issues of insufficient knowledge and hallucination in Large Language Models (LLMs), numerous studies have explored integrating LLMs with Knowledge Graphs (KGs).However, these methods are typically evaluated on conventional Knowledge Graph Question Answering (KGQA) with complete KGs, where all factual triples required for each question are entirely covered by the given KG.In such cases, LLMs primarily act as an agent to find answer entities within the KG, rather than effectively integrating the internal knowledge of LLMs and external knowledge sources such as KGs.In fact, KGs are often incomplete to cover all the knowledge required to answer questions.To simulate these real-world scenarios and evaluate the ability of LLMs to integrate internal and external knowledge, we propose leveraging LLMs for QA under Incomplete Knowledge Graph (IKGQA), where the provided KG lacks some of the factual triples for each question, and construct corresponding datasets.To handle IKGQA, we propose a training-free method called Generate-on-Graph (GoG), which can generate new factual triples while exploring KGs.Specifically, GoG performs reasoning through a Thinking-Searching-Generating framework, which treats LLM as both Agent and KG in IKGQA.Experimental results on two datasets demonstrate that our GoG outperforms all previous methods. Shizhu He, Jiabei Chen, Zihao Wang 0001, Yangqiu Song, Hanghang Tong, Jun Zhao 0001, Kang Liu 0001 |
EMNLP | 2 |
| 2024 | Unsupervised Learning of Neural Semantic Mappings with the Hungarian Algorithm for Compositional SemanticsabstractNeural semantic parsing maps natural languages (NL) to equivalent formal semantics which are compositional and deduce the sentence meanings by composing smaller parts. To learn a well-defined semantics, semantic parsers must recognize small parts, which are semantic mappings between NL and semantic tokens. Attentions in recent neural models are usually explained as one-on-one semantic mappings. However, attention weights with end-to-end training are shown only weakly correlated with human-labeled mappings. Despite the usefulness, supervised mappings are expensive. We propose the unsupervised Hungarian tweaks on attentions to better model mappings. Experiments have shown our methods is competitive with the supervised approach on performance and mappings recognition, and outperform other baselines. Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
ICASSP | 2 |
| 2024 | Mastering Symbolic Operations: Augmenting Language Models with Compiled Neural NetworksabstractLanguage models' (LMs) proficiency in handling deterministic symbolic reasoning and rule-based tasks remains limited due to their dependency implicit learning on textual data. To endow LMs with genuine rule comprehension abilities, we propose "Neural Comprehension" - a framework that synergistically integrates compiled neural networks (CoNNs) into the standard transformer architecture. CoNNs are neural modules designed to explicitly encode rules through artificially generated attention weights. By incorporating CoNN modules, the Neural Comprehension framework enables LMs to accurately and robustly execute rule-intensive symbolic tasks. Extensive experiments demonstrate the superiority of our approach over existing techniques in terms of length generalization, efficiency, and interpretability for symbolic operations. Furthermore, it can be applied to LMs across different model scales, outperforming tool-calling methods in arithmetic reasoning tasks while maintaining superior inference efficiency. Our work highlights the potential of seamlessly unifying explicit rule learning via CoNNs and implicit pattern learning in LMs, paving the way for true symbolic comprehension capabilities. The code is released at: \url{https://github.com/wengsyx/Neural-Comprehension}. Yixuan Weng, Minjun Zhu, Bin Li 0083, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
ICLR | 5 |
| 2024 | S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language ModelabstractFangyu Lei, Qian Liu, Yiming Huang, Shizhu He, Jun Zhao, Kang Liu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fangyu Lei, Qian Liu 0033, Shizhu He, Jun Zhao 0001, Kang Liu 0001 |
NAACL-HLT | 4 |
| 2024 | From Instance Training to Instruction Learning: Task Adapters Generation from InstructionsabstractLarge language models (LLMs) have acquired the ability to solve general tasks by utilizing instruction finetuning (IFT). However, IFT still relies heavily on instance training of extensive task data, which greatly limits the adaptability of LLMs to real-world scenarios where labeled task instances are scarce and broader task generalization becomes paramount. Contrary to LLMs, humans acquire skills and complete tasks not merely through repeated practice but also by understanding and following instructional guidelines. This paper is dedicated to simulating human learning to address the shortcomings of instance training, focusing on instruction learning to enhance cross-task generalization. Within this context, we introduce Task Adapters Generation from Instructions (TAGI), which automatically constructs the task-specific model in a parameter generation manner based on the given task instructions without retraining for unseen tasks. Specifically, we utilize knowledge distillation to enhance the consistency between TAGI developed through Learning with Instruction and task-specific models developed through Training with Instance, by aligning the labels, output logits, and adapter parameters between them. TAGI is endowed with cross-task generalization capabilities through a two-stage training process that includes hypernetwork pretraining and finetuning. We evaluate TAGI on the Super-Natural Instructions and P3 datasets. The experimental results demonstrate that TAGI can match or even outperform traditional meta-trained models and other hypernetwork models, while significantly reducing computational requirements. Our code will be available at https://github.com/Xnhyacinth/TAGI. Huanxuan Liao, Shizhu He, Yuanzhe Zhang, Yanchao Hao, Shengping Liu, Kang Liu 0001, Jun Zhao 0001 |
NeurIPS | 2 |
| 2024 | Large Language Models With Holistically Thought Could Be Better Doctors
Yixuan Weng, Bin Li 0083, Minjun Zhu, Bin Sun 0001, Shizhu He, Shengping Liu, Kang Liu 0001, Shutao Li 0001, Jun Zhao 0001 |
NLPCC (2) | 6 |
| 2024 | Seq2Set2Seq: A Two-stage Disentangled Method for Reply Keyword Generation in Social MediaabstractSocial media produces large amounts of content every day. How to predict the potential influences of the contents from a social reply feedback perspective is a key issue that has not been explored. Thus, we propose a novel task named reply keyword prediction in social media, which aims to predict the keywords in the potential replies in as many aspects as possible. One prerequisite challenge is that the accessible social media datasets labeling such keywords remain absent. To solve this issue, we propose a new dataset, 1 to study the reply keyword prediction in social media. This task could be seen as a single-turn dialogue keyword prediction for open-domain dialogue system. However, existing methods for dialogue keyword prediction cannot be adopted directly, which has two main drawbacks. First, they do not provide an explicit mechanism to model topic complementarity between keywords which is crucial in social media to controllably model all aspects of replies. Second, the collocations of keywords are not explicitly modeled, which also makes it less controllable to optimize for fine-grained prediction since the context information is much less than that in dialogue. To address these issues, we propose a two-stage disentangled framework, which can optimize the complementarity and collocation explicitly in a disentangled fashion. In the first stage, we use a sequence-to-set paradigm via multi-label prediction and determinantal point processes, to generate a set of keyword seeds satisfying the complementarity. In the second stage, we adopt a set-to-sequence paradigm via seq2seq model with the keyword seeds guidance from the set, to generate the more-fine-grained keywords with collocation. Experiments show that this method can generate not only a more diverse set of keywords but also more relevant and consistent keywords. Furthermore, the keywords obtained based on this method can achieve better reply generation results in the retrieval-based system than others. Jie Liu 0022, Shizhu He, Shun Wu, Kang Liu 0001, Shenping Liu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2024 | Towards Better Quantity Representations for Solving Math Word ProblemsabstractSolving a math word problem requires selecting quantities in it and performing appropriate arithmetic operations to obtain the answer. For deep learning-based methods, it is vital to obtain good quantity representations, i.e., to selectively and emphatically aggregate information in the context of quantities. However, existing works have not paid much attention to this aspect. Many works simply encode quantities as ordinary tokens, or use some implicit or rule-based methods to select information in their context. This leads to poor results when dealing with linguistic variations and confounding quantities. This article proposes a novel method to identify question-related distinguishing features of quantities by contrasting their context with the question and the context of other quantities, thereby enhancing the representation of quantities. Our method not only considers the contrastive relationship between quantities but also considers multiple relationships jointly. Besides, we propose two auxiliary tasks to further guide the representation learning of quantities: (1) predicting whether a quantity is used in the question and (2) predicting the relations (operators) between quantities given the question. Experimental results show that our method outperforms previous methods on SVAMP and ASDiv-A under similar settings, even some newly released strong baselines. Supplementary experiments further confirm that our method indeed improves the performance of quantity selection by improving the representation of both quantities and questions. Runxin Sun, Shizhu He, Jun Zhao 0001, Kang Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | On the Effects of Structural Modeling for Neural Semantic ParsingabstractSemantic parsing aims to map natural language sentences to predefined formal languages, such as logic forms and programming languages, as the semantic annotation.From the theoretic views of linguistic and programming language, structures play an important role in both languages, which had motivated semantic parsers since the task was proposed in the beginning.But in the neural era, semantic parsers treating both natural and formal language as sequences, such as Seq2Seq and LLMs, have got more attentions.On the other side, lots of neural progress have been made for grammar induction, which only focuses on natural languages.Although closely related in the sense of structural modeling, these techniques hadn't been jointly analyzed on the semantic parsing testbeds.To gain the better understanding on structures for semantic parsing, we design a taxonomy of structural modeling methods, and evaluate some representative techniques on semantic parsing, including both compositional and i.i.d.generalizations.In addition to the previous opinion that structures will help in general, we find that (1) structures must be designed for the specific dataset and generalization level, and (2) what really matters is not the structure choice of either source or target side, but the choice combination of both sides.Based on the finding, we further propose a metric that can evaluate the structure choice, which we believe can boost the automation of grammar designs for specific datasets and domains. Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
CoNLL | 2 |
| 2023 | Find Parent then Label Children: A Two-stage Taxonomy Completion Method with Pre-trained Language ModelabstractTaxonomies, which organize domain concepts into hierarchical structures, are crucial for building knowledge systems and downstream applications.As domain knowledge evolves, taxonomies need to be continuously updated to include new concepts.Previous approaches have mainly focused on adding concepts to the leaf nodes of the existing hierarchical tree, which does not fully utilize the taxonomy's knowledge and is unable to update the original taxonomy structure (usually involving nonleaf nodes).In this paper, we propose a twostage method called ATTEMPT for taxonomy completion.Our method inserts new concepts into the correct position by finding a parent node and labeling child nodes.Specifically, by combining local nodes with prompts to generate natural sentences, we take advantage of pre-trained language models for hypernym/hyponymy recognition.Experimental results on two public datasets (including six domains) show that ATTEMPT performs best on both taxonomy completion and extension tasks, surpassing existing methods. Yixuan Weng, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
EACL | 3 |
| 2023 | Learning to Build Reasoning Chains by Reliable Path RetrievalabstractQuestion answering (QA) systems have long pursued the ability to reason over explicit knowledge credibly. Recent work has incorporated knowledge into fine-grained sentences and constructed natural language database (NLDB) task, and conducts complex QA with explicit reasoning chains. Existing models focus on retrieving evidence by combining multiple modules or discretely. However, these models ignore utilizing path information (e.g. sentence order), which is proven to be important for evidence retrievers. In this work, we propose a ReliAble Path-retrieval (RAP) to generate varying length evidence chains iteratively. It comprehensively models reasoning chains and introduces loss from two views. The experimental results show that our model demonstrates state-of-the-art performance on both evidence chain retrieval and question-answering tasks. Additional experiments on sequential supervised and sequential unsupervised retrieval fully indicate the significance of RAP. Minjun Zhu, Yixuan Weng, Shizhu He, Cunguang Wang, Kang Liu 0001, Jun Zhao 0001 |
ICASSP | 3 |
| 2023 | Unsupervised Domain Adaptation on Sentence Matching Through Self-Supervision
Guirong Bai, Qingbin Liu, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
J. Comput. Sci. Technol. | 3 |
| 2023 | Unsupervised Dialogue State Tracking for End-to-End Task-Oriented Dialogue with a Multi-Span Prediction Network
Qingbin Liu, Shizhu He, Cao Liu, Kang Liu 0001, Jun Zhao 0001 |
J. Comput. Sci. Technol. | 2 |
| 2023 | Bidirectional Sentence Ordering with Interactive DecodingabstractSentence ordering aims at restoring orders of shuffled sentences in a paragraph. Previous methods usually predict orders in a single direction, i.e., from head to tail. However, unidirectional prediction inevitably causes error accumulation, which restricts performance. In this article, we propose a bidirectional ordering method, which predicts orders in both head-to-tail and tail-to-head directions at the same time. In our bidirectional ordering method, two directions can interact with each other and help alleviate the error accumulation problem of ordering. Experiments demonstrate that our method can effectively improve performance of previous models. Guirong Bai, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Answering Numerical Reasoning Questions in Table-Text Hybrid Contents with Graph-based Encoder and Tree-based Decoder
Fangyu Lei, Shizhu He, Jun Zhao 0001, Kang Liu 0001 |
COLING | 2 |
| 2022 | Decoupling Mixture-of-Graphs: Unseen Relational Learning for Knowledge Graph Completion by Fusing Ontology and Textual ExpertsabstractKnowledge Graph Embedding (KGE) has been proposed and successfully utilized to knowledge Graph Completion (KGC). But classic KGE paradigm often fail in unseen relation representations. Previous studies mainly utilize the textual descriptions of relations and its neighbor relations to represent unseen relations. In fact, the semantics of a relation can be expressed by three kinds of graphs: factual graph, ontology graph, textual description graph, and they can complement each other. A more common scenario in the real world is that seen and unseen relations appear at the same time. In this setting, the training set (only seen relations) and testing set (both seen and unseen relations) own different distributions. And the train-test inconsistency problem will make KGE methods easiy overfit on seen relations and under-performance on unseen relations. In this paper, we propose decoupling mixture-of-graph experts (DMoG) for unseen relations learning, which could represent the unseen relations in the factual graph by fusing ontology and textual graphs, and decouple fusing space and reasoning space to alleviate overfitting for seen relations. The experiments on two unseen only public datasets and a mixture dataset verify the effectiveness of the proposed method, which improves the state-of-the-art methods by 6.84% in Hits@10 on average. Ran Song 0002, Shizhu He, Suncong Zheng, Shengxiang Gao, Kang Liu 0001, Zhengtao Yu 0001, Jun Zhao 0001 |
COLING | 2 |
| 2022 | Example-guided stylized response generation in zero-shot setting
Guirong Bai, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
Sci. China Inf. Sci. | 2 |
| 2022 | Using Pre-trained Language Model to Enhance Active Learning for Sentence MatchingabstractActive learning is an effective method to substantially alleviate the problem of expensive annotation cost for data-driven models. Recently, pre-trained language models have been demonstrated to be powerful for learning language representations. In this article, we demonstrate that the pre-trained language model can also utilize its learned textual characteristics to enrich criteria of active learning. Specifically, we provide extra textual criteria with the pre-trained language model to measure instances, including noise, coverage, and diversity. With these extra textual criteria, we can select more efficient instances for annotation and obtain better results. We conduct experiments on both English and Chinese sentence matching datasets. The experimental results show that the proposed active learning approach can be enhanced by the pre-trained language model and obtain better performance. Guirong Bai, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Fact-Driven Abstractive Summarization by Utilizing Multi-Granular Multi-Relational KnowledgeabstractAbstractive summarization generates a concise summary to capture the key ideas of the source text. This task underpins important applications like information retrieval, document comprehension, and event tracking. While much progress has been achieved, state-of-the-art summarization approaches often fail to generate high-quality summaries to reproduce factual details accurately. One of the key limitations of existing solutions is that they are primarily concerned about extracting facts from the source text but overlook other crucial factual information, such as the related time, locations, reasons, consequences, purposes, participants and involved parties. Furthermore, the current summarization frameworks are inadequate in modeling the complex semantic relations among facts and the corresponding factual information, leaving much room for improvement. This paper presentsFFSum, a novel summarization framework for exploiting multi-grained factual information to improve text summarization. To this end,FFSumconstructs an individual fine-grained factual graph with multiple relations among facts and the corresponding factual information. It employs a fact-driven graph attention network to integrate multi-granular factual representations at the encoding stage. It then uses a hybrid pointer network to retrieve factual pieces from the graph for the summary generation. We evaluate theFFSumby applying it to two real-world datasets. Experimental results show that theFFSumconsistently outperforms a state-of-the-art approach across evaluation datasets. Qianren Mao, Jianxin Li 0002, Hao Peng 0001, Shizhu He, Philip S. Yu, Zheng Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Domain-Lifelong Learning for Dialogue State Tracking via Knowledge Preservation NetworksabstractDialogue state tracking (DST), which estimates user goals given a dialogue context, is an essential component of task-oriented dialogue systems.Conventional DST models are usually trained offline, which requires a fixed dataset prepared in advance.This paradigm is often impractical in real-world applications since online dialogue systems usually involve continually emerging new data and domains.Therefore, this paper explores Domain-Lifelong Learning for Dialogue State Tracking (DLL-DST), which aims to continually train a DST model on new data to learn incessantly emerging new domains while avoiding catastrophically forgetting old learned domains.To this end, we propose a novel domainlifelong learning method, called Knowledge Preservation Networks (KPN), which consists of multi-prototype enhanced retrospection and multi-strategy knowledge distillation, to solve the problems of expression diversity and combinatorial explosion in the DLL-DST task.Experimental results show that KPN effectively alleviates catastrophic forgetting and outperforms previous state-of-the-art lifelong learning methods by 4.25% and 8.27% of whole joint goal accuracy on the MultiWOZ benchmark and the SGD benchmark, respectively. Qingbin Liu, Cao Liu, Jiansong Chen, Fan Yang 0087, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
EMNLP (1) | 7 |
| 2021 | A Unified Shared-Private Network with Denoising for Dialogue State Tracking
Qingbin Liu, Shizhu He, Kang Liu 0001, Shengping Liu, Jun Zhao 0001 |
J. Comput. Sci. Technol. | 2 |
| 2021 | Heterogeneous Relational Graph Neural Networks with Adaptive Objective for End-to-End Task-Oriented Dialogue
Qingbin Liu, Guirong Bai, Shizhu He, Cao Liu, Kang Liu 0001, Jun Zhao 0001 |
Knowl. Based Syst. | 3 |
| 2020 | Pre-trained Language Model Based Active Learning for Sentence MatchingabstractActive learning is able to significantly reduce the annotation cost for data-driven techniques.However, previous active learning approaches for natural language processing mainly depend on the entropy-based uncertainty criterion, and ignore the characteristics of natural language.In this paper, we propose a pre-trained language model based active learning approach for sentence matching.Differing from previous active learning, it can provide linguistic criteria from the pre-trained language model to measure instances and help select more effective instances for annotation.Experiments demonstrate our approach can achieve greater accuracy with fewer labeled training instances. Guirong Bai, Shizhu He, Kang Liu 0001, Jun Zhao 0001, Zaiqing Nie |
COLING | 2 |
| 2019 | Learning to Align Question and Answer Utterances in Customer Service Conversation with Recurrent Pointer NetworksabstractCustomers ask questions, and customer service staffs answer those questions. It is the basic service manner of customer service (CS). The progress of CS is a typical multi-round conversation. However, there are no explicit corresponding relations among conversational utterances. This paper focuses on obtaining explicit alignments of question and answer utterances in CS. It not only is an important task of dialogue analysis, but also able to obtain lots of valuable train data for learning dialogue systems. In this work, we propose end-to-end models for aligning question (Q) and answer (A) utterances in CS conversation with recurrent pointer networks (RPN). On the one hand, RPN-based alignment models are able to model the conversational contexts and the mutual influence of different Q-A alignments. On the other hand, they are able to address the issue of empty and multiple alignments for some utterances in a unified manner. We construct a dataset from an in-house online CS. The experimental results demonstrate that the proposed models are effective to learn the alignments of question and answer utterances. Shizhu He, Kang Liu 0001, Weiting An |
AAAI | 1 |
| 2019 | Vocabulary Pyramid Network: Multi-Pass Encoding and Decoding with Multi-Level Vocabularies for Response GenerationabstractWe study the task of response generation.Conventional methods employ a fixed vocabulary and one-pass decoding, which not only make them prone to safe and general responses but also lack further refining to the first generated raw sequence.To tackle the above two problems, we present a Vocabulary Pyramid Network (VPN) which is able to incorporate multi-pass encoding and decoding with multi-level vocabularies into response generation.Specifically, the dialogue input and output are represented by multi-level vocabularies which are obtained from hierarchical clustering of raw words.Then, multi-pass encoding and decoding are conducted on the multilevel vocabularies.Since VPN is able to leverage rich encoding and decoding information with multi-level vocabularies, it has the potential to generate better responses.Experiments on English Twitter and Chinese Weibo datasets demonstrate that VPN remarkably outperforms strong baselines. Cao Liu, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 2 |
| 2019 | AdaNSP: Uncertainty-driven Adaptive Decoding in Neural Semantic ParsingabstractNeural semantic parsers utilize the encoderdecoder framework to learn an end-to-end model for semantic parsing that transduces a natural language sentence to the formal semantic representation.To keep the model aware of the underlying grammar in target sequences, many constrained decoders were devised in a multi-stage paradigm, which decode to the sketches or abstract syntax trees first, and then decode to target semantic tokens.We instead to propose an adaptive decoding method to avoid such intermediate representations.The decoder is guided by model uncertainty and automatically uses deeper computations when necessary.Thus it can predict tokens adaptively.Our model outperforms the state-of-the-art neural models and does not need any expertise like predefined grammar or sketches in the meantime. Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 2 |
| 2019 | Incorporating Interlocutor-Aware Context into Response Generation on Multi-Party ChatbotsabstractConventional chatbots focus on two-party response generation, which simplifies the real dialogue scene.In this paper, we strive toward a novel task of Response Generation on Multi-Party Chatbot (RGMPC), where the generated responses heavily rely on the interlocutors' roles (e.g., speaker and addressee) and their utterances.Unfortunately, complex interactions among the interlocutors' roles make it challenging to precisely capture conversational contexts and interlocutors' information.Facing this challenge, we present a response generation model which incorporates Interlocutor-aware Contexts into Recurrent Encoder-Decoder frameworks (ICRED) for RGMPC.Specifically, we employ interactive representations to capture dialogue contexts for different interlocutors.Moreover, we leverage an addressee memory to enhance contextual interlocutor information for the target addressee.Finally, we construct a corpus for RGMPC based on an existing open-access dataset.Automatic and manual evaluations demonstrate that the ICRED remarkably outperforms strong baselines. Cao Liu, Kang Liu 0001, Shizhu He, Zaiqing Nie, Jun Zhao 0001 |
CoNLL | 3 |
| 2019 | Generating Questions for Knowledge Bases via Incorporating Diversified Contexts and Answer-Aware LossabstractCao Liu, Kang Liu, Shizhu He, Zaiqing Nie, Jun Zhao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Cao Liu, Kang Liu 0001, Shizhu He, Zaiqing Nie, Jun Zhao 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Learning the Extraction Order of Multiple Relational Facts in a Sentence with Reinforcement LearningabstractXiangrong Zeng, Shizhu He, Daojian Zeng, Kang Liu, Shengping Liu, Jun Zhao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiangrong Zeng, Shizhu He, Daojian Zeng, Kang Liu 0001, Shengping Liu, Jun Zhao 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Variational Attention for Commonsense Knowledge Aware Conversation Generation
Guirong Bai, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
NLPCC (1) | 2 |
| 2018 | Large Scaled Relation Extraction With Reinforcement LearningabstractSentence relation extraction aims to extract relational facts from sentences, which is an important task in natural language processing field. Previous models rely on the manually labeled supervised dataset. However, the human annotation is costly and limits to the number of relation and data size, which is difficult to scale to large domains. In order to conduct largely scaled relation extraction, we utilize an existing knowledge base to heuristically align with texts, which not rely on human annotation and easy to scale. However, using distant supervised data for relation extraction is facing a new challenge: sentences in the distant supervised dataset are not directly labeled and not all sentences that mentioned an entity pair can represent the relation between them. To solve this problem, we propose a novel model with reinforcement learning. The relation of the entity pair is used as distant supervision and guide the training of relation extractor with the help of reinforcement learning method. We conduct two types of experiments on a publicly released dataset. Experiment results demonstrate the effectiveness of the proposed method compared with baseline models, which achieves 13.36\% improvement. Xiangrong Zeng, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
AAAI | 2 |
| 2018 | Extracting Relational Facts by an End-to-End Neural Model with Copy MechanismabstractThe relational facts in sentences are often complicated.Different relational triplets may have overlaps in a sentence.We divided the sentences into three types according to triplet overlap degree, including Normal, EntityPairOverlap and SingleEn-tiyOverlap. Existing methods mainly focus on Normal class and fail to extract relational triplets precisely.In this paper, we propose an end-to-end model based on sequence-to-sequence learning with copy mechanism, which can jointly extract relational facts from sentences of any of these classes.We adopt two different strategies in decoding process: employing only one united decoder or applying multiple separated decoders.We test our models in two public datasets and our model outperform the baseline method significantly. Xiangrong Zeng, Daojian Zeng, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 3 |
| 2018 | Pattern-revising Enhanced Simple Question Answering over Knowledge BasesabstractQuestion Answering over Knowledge Bases (KB-QA), which automatically answer natural language questions based on the facts contained by a knowledge base, is one of the most important natural language processing (NLP) tasks. Simple questions constitute a large part of questions queried on the web, still being a challenge to QA systems. In this work, we propose to conduct pattern extraction and entity linking first, and put forward pattern revising procedure to mitigate the error propagation problem. In order to learn to rank candidate subject-predicate pairs to enable the relevant facts retrieval given a question, we propose to do joint fact selection enhanced by relation detection. Multi-level encodings and multi-dimension information are leveraged to strengthen the whole procedure. The experimental results demonstrate that our approach sets a new record in this task, outperforming the current state-of-the-art by an absolute large margin. Yanchao Hao, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
COLING | 3 |
| 2018 | Curriculum Learning for Natural Answer GenerationabstractBy reason of being able to obtain natural language responses, natural answers are more favored in real-world Question Answering (QA) systems. Generative models learn to automatically generate natural answers from large-scale question answer pairs (QA-pairs). However, they are suffering from the uncontrollable and uneven quality of QA-pairs crawled from the Internet. To address this problem, we propose a curriculum learning based framework for natural answer generation (CL-NAG), which is able to take full advantage of the valuable learning data from a noisy and uneven-quality corpus. Specifically, we employ two practical measures to automatically measure the quality (complexity) of QA-pairs. Based on the measurements, CL-NAG firstly utilizes simple and low-quality QA-pairs to learn a basic model, and then gradually learns to produce better answers with richer contents and more complete syntaxes based on more complex and higher-quality QA-pairs. In this way, all valuable information in the noisy and uneven-quality corpus could be fully exploited. Experiments demonstrate that CL-NAG outperforms the state-of-the-arts, which increases 6.8% and 8.7% in the accuracy for simple and complex questions, respectively. Cao Liu, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
IJCAI | 2 |
| 2017 | Distant Supervision for Relation Extraction with Sentence-Level Attention and Entity DescriptionsabstractDistant supervision for relation extraction is an efficient method to scale relation extraction to very large corpora which contains thousands of relations. However, the existing approaches have flaws on selecting valid instances and lack of background knowledge about the entities. In this paper, we propose a sentence-level attention model to select the valid instances, which makes full use of the supervision information from knowledge bases. And we extract entity descriptions from Freebase and Wikipedia pages to supplement background knowledge for our task. The background knowledge not only provides more information for predicting relations, but also brings better entity representations for the attention module. We conduct three experiments on a widely used dataset and the experimental results show that our approach outperforms all the baseline systems significantly. Guoliang Ji, Kang Liu 0001, Shizhu He, Jun Zhao 0001 |
AAAI | 3 |
| 2017 | An End-to-End Model for Question Answering over Knowledge Base with Cross-Attention Combining Global KnowledgeabstractWith the rapid growth of knowledge bases (KBs) on the web, how to take full advantage of them becomes increasingly important.Question answering over knowledge base (KB-QA) is one of the promising approaches to access the substantial knowledge.Meanwhile, as the neural networkbased (NN-based) methods develop, NNbased KB-QA has already achieved impressive results.However, previous work did not put more emphasis on question representation, and the question is converted into a fixed vector regardless of its candidate answers.This simple representation strategy is not easy to express the proper information in the question.Hence, we present an end-to-end neural network model to represent the questions and their corresponding scores dynamically according to the various candidate answer aspects via cross-attention mechanism.In addition, we leverage the global knowledge inside the underlying KB, aiming at integrating the rich KB information into the representation of the answers.As a result, it could alleviates the out-of-vocabulary (OOV) problem, which helps the crossattention model to represent the question more precisely.The experimental results on WebQuestions demonstrate the effectiveness of the proposed approach. Yanchao Hao, Yuanzhe Zhang, Kang Liu 0001, Shizhu He, Zhanyi Liu, Hua Wu 0003, Jun Zhao 0001 |
ACL (1) | 4 |
| 2017 | Generating Natural Answers by Incorporating Copying and Retrieving Mechanisms in Sequence-to-Sequence LearningabstractGenerating answer with natural language sentence is very important in real-world question answering systems, which needs to obtain a right answer as well as a coherent natural response.In this paper, we propose an end-to-end question answering system called COREQA in sequence-to-sequence learning, which incorporates copying and retrieving mechanisms to generate natural answers within an encoder-decoder framework.Specifically, in COREQA, the semantic units (words, phrases and entities) in a natural answer are dynamically predicted from the vocabulary, copied from the given question and/or retrieved from the corresponding knowledge base jointly.Our empirical study on both synthetic and realworld datasets demonstrates the efficiency of COREQA, which is able to generate correct, coherent and natural answers for knowledge inquired questions. Shizhu He, Cao Liu, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 1 |
| 2017 | Which is the Effective Way for Gaokao: Information Retrieval or Neural Networks?abstractAs one of the most important test of China, Gaokao is designed to be difficult enough to distinguish the excellent high school students.In this work, we detailed the Gaokao History Multiple Choice Questions(GKHMC) and proposed two different approaches to address them using various resources.One approach is based on entity search technique (IR approach), the other is based on text entailment approach where we specifically employ deep neural networks(NN approach).The result of experiment on our collected real Gaokao questions showed that they are good at different categories of questions, i.e.IR approach performs much better at entity questions(EQs) while NN approach shows its advantage on sentence questions(SQs).Our new method achieves state-of-the-art performance and show that it's indispensable to apply hybrid method when participating in the real-world tests. Shangmin Guo, Xiangrong Zeng, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
EACL (1) | 3 |
| 2016 | Knowledge Graph Completion with Adaptive Sparse Transfer MatrixabstractWe model knowledge graphs for their completion by encoding each entity and relation into a numerical space. All previous work including Trans(E, H, R, and D) ignore the heterogeneity (some relations link many entity pairs and others do not) and the imbalance (the number of head entities and that of tail entities in a relation could be different) of knowledge graphs. In this paper, we propose a novel approach TranSparse to deal with the two issues. In TranSparse, transfer matrices are replaced by adaptive sparse matrices, whose sparse degrees are determined by the number of entities (or entity pairs) linked by relations. In experiments, we design structured and unstructured sparse patterns for transfer matrices and analyze their advantages and disadvantages. We evaluate our approach on triplet classification and link prediction tasks. Experimental results show that TranSparse outperforms Trans(E, H, R, and D) significantly, and achieves state-of-the-art performance. Guoliang Ji, Kang Liu 0001, Shizhu He, Jun Zhao 0001 |
AAAI | 3 |
| 2016 | A Probabilistic Soft Logic Based Approach to Exploiting Latent and Global Information in Event ClassificationabstractGlobal information such as event-event association, and latent local information such as fine-grained entity types, are crucial to event classification. However, existing methods typically focus on sophisticated local features such as part-of-speech tags, either fully or partially ignoring the aforementioned information. By contrast, this paper focuses on fully employing them for event classification. We notice that it is difficult to encode some global information such as event-event association for previous methods. To resolve this problem, we propose a feasible approach which encodes global information in the form of logic using Probabilistic Soft Logic model. Experimental results show that, our proposed approach advances state-of-the-art methods, and achieves the best F1 score to date on the ACE data set. Kang Liu 0001, Shizhu He, Jun Zhao 0001 |
AAAI | 3 |
| 2016 | A Joint Model for Question Answering over Multiple Knowledge BasesabstractAs the amount of knowledge bases (KBs) grows rapidly, the problem of question answering (QA) over multiple KBs has drawn more attention. The most significant distinction between multiple KB-QA and single KB-QA is that the former must consider the alignments between KBs. The pipeline strategy first constructs the alignments independently, and then uses the obtained alignments to construct queries. However, alignment construction is not a trivial task, and the introduced noises would be passed on to query construction. By contrast, we notice that alignment construction and query construction are interactive steps, and jointly considering them would be beneficial. To this end, we present a novel joint model based on integer linear programming (ILP), uniting these two procedures into a uniform framework. The experimental results demonstrate that the proposed approach outperforms state-of-the-art systems, and is able to improve the performance of both alignment construction and query construction. Yuanzhe Zhang, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
AAAI | 2 |
| 2016 | Leveraging FrameNet to Improve Automatic Event DetectionabstractFrames defined in FrameNet (FN) share highly similar structures with events in ACE event extraction program.An event in ACE is composed of an event trigger and a set of arguments.Analogously, a frame in FN is composed of a lexical unit and a set of frame elements, which play similar roles as triggers and arguments of ACE events respectively.Besides having similar structures, many frames in FN actually express certain types of events.The above observations motivate us to explore whether there exists a good mapping from frames to event-types and if it is possible to improve event detection by using FN.In this paper, we propose a global inference approach to detect events in FN.Further, based on the detected results, we analyze possible mappings from frames to event-types.Finally, we improve the performance of event detection and achieve a new state-of-the-art result by using the events automatically detected from FN. Yubo Chen 0001, Shizhu He, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 3 |
| 2016 | Learning to Represent Review with Tensor Decomposition for Spam DetectionabstractReview spam detection is a key task in opinion mining.To accomplish this type of detection, previous work has focused mainly on effectively representing fake and non-fake reviews with discriminative features, which are discovered or elaborately designed by experts or developers.This paper proposes a novel review spam detection method that learns the representation of reviews automatically instead of heavily relying on experts' knowledge in a data-driven manner.More specifically, according to 11 relations (generated automatically from two basic patterns) between reviewers and products, we employ tensor decomposition to learn the embeddings of the reviewers and products in a vector space.We collect relations between any two entities (reviewers and products), which results in much useful and global information.We concatenate the review text, the embeddings of the reviewer and the reviewed product as the representation of a review.Based on such representations, the classifier could identify the opinion spam more precisely.Experimental results on an open Yelp dataset show that our method could effectively enhance the spam detection accuracy compared with the stateof-the-art methods. Xuepeng Wang, Kang Liu 0001, Shizhu He, Jun Zhao 0001 |
EMNLP | 3 |
| 2016 | Employing External Rich Knowledge for Machine Comprehension
Bingning Wang, Shangmin Guo, Kang Liu 0001, Shizhu He, Jun Zhao 0001 |
IJCAI | 4 |
| 2015 | Knowledge Graph Embedding via Dynamic Mapping MatrixabstractGuoliang Ji, Shizhu He, Liheng Xu, Kang Liu, Jun Zhao. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Guoliang Ji, Shizhu He, Liheng Xu, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 2 |
| 2015 | Learning to Represent Knowledge Graphs with Gaussian EmbeddingabstractThe representation of a knowledge graph (KG) in a latent space recently has attracted more and more attention. To this end, some proposed models (e.g., TransE) embed entities and relations of a KG into a "point" vector space by optimizing a global loss function which ensures the scores of positive triplets are higher than negative ones. We notice that these models always regard all entities and relations in a same manner and ignore their (un)certainties. In fact, different entities and relations may contain different certainties, which makes identical certainty insufficient for modeling. Therefore, this paper switches to density-based embedding and propose KG2E for explicitly modeling the certainty of entities and relations, which learn the representations of KGs in the space of multi-dimensional Gaussian distributions. Each entity/relation is represented by a Gaussian distribution, where the mean denotes its position and the covariance (currently with diagonal covariance) can properly represent its certainty. In addition, compared with the symmetric measures used in point-based methods, we employ the KL-divergence for scoring triplets, which is a natural asymmetry function for effectively modeling multiple types of relations. We have conducted extensive experiments on link prediction and triplet classification with multiple benchmark datasets (WordNet and Freebase). Our experimental results demonstrate that our method can effectively model the (un)certainties of entities and relations in a KG, and it significantly outperforms state-of-the-art methods (including TransH and TransR). Shizhu He, Kang Liu 0001, Guoliang Ji, Jun Zhao 0001 |
CIKM | 1 |
| 2014 | Question Answering over Linked Data Using First-order LogicabstractQuestion Answering over Linked Data (QALD) aims to evaluate a question an-swering system over structured data, the key objective of which is to translate questions posed using natural language into structured queries. This technique can help common users to directly ac-cess open-structured knowledge on the Web and, accordingly, has attracted much attention. To this end, we propose a novel method using first-order logic. We formulate the knowledge for resolving the ambiguities in the main three steps of QALD (phrase detection, phrase-to-semantic-item mapping and semantic item grouping) as first-order logic clauses in a Markov Logic Network. All clauses can then produce interacted effects in a unified framework and can jointly resolve all am-biguities. Moreover, our method adopts a pattern-learning strategy for semantic item grouping. In this way, our method can cover more text expressions and answer more questions than previous methods us-ing manually designed patterns. The ex-perimental results using open benchmarks demonstrate the effectiveness of the pro-posed method. 1 Shizhu He, Kang Liu 0001, Yuanzhe Zhang, Liheng Xu, Jun Zhao 0001 |
EMNLP | 1 |
| 2013 | Statistical Machine Translation Improves Question Retrieval in Community Question Answering via Matrix Factorization
Guangyou Zhou, Yang Liu 0021, Shizhu He, Jun Zhao 0001 |
ACL (1) | 4 |