EDBT 2026 Demo / reviewers in the wild / expert
Baotian Hu
dblp:155/1902
· DBLP profile ↗
75ranked-venue papers
3as first author
62since 2021 · last 2026
0000-0001-7490-684XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 57 · 3 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 18 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Value-based Process Verifier via Low-Cost Variance ReductionabstractLarge language models (LLMs) have achieved remarkable success in a wide range of tasks. However, their reasoning capabilities, particularly in complex domains like mathematics, remain a significant challenge. Value-based process verifiers, which estimate the probability of a partial reasoning chain leading to a correct solution, are a promising approach for improving reasoning. Nevertheless, their effectiveness is often hindered by estimation error in their training annotations, a consequence of the limited number of Monte Carlo (MC) samples feasible due to the high cost of LLM inference. In this paper, we identify that the estimation error primarily arises from high variance rather than bias, and the MC estimator is a Minimum Variance Unbiased Estimator (MVUE). To address the problem, we propose the Compound Monte Carlo Sampling (ComMCS) method, which constructs an unbiased estimator by linearly combining the MC estimators from the current and subsequent steps. Theoretically, we show that our method leads to a predictable reduction in variance, while maintaining an unbiased estimation without additional LLM inference cost. We also perform empirical experiments on the MATH-500 and GSM8K benchmarks to demonstrate the effectiveness of our method. Notably, ComMCS outperforms regression-based optimization method by 2.8 points, the non-variance-reduced baseline by 2.2 points on MATH-500 on Best-of-32 sampling experiment. Zetian Sun, Dongfang Li 0002, Baotian Hu, Min Zhang 0005 |
AAAI | 3 |
| 2026 | Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement LearningabstractLarge Language Models (LLMs) face severe challenges in long-context processing, including quadratic computational costs, information forgetting, and the context fragmentation inherent in Retrieval-Augmented Generation (RAG).We introduce LycheeMemory, a cognitively inspired framework that enables efficient long-context inference via chunk-wise compression and selective memory recall, rather than processing all raw tokens.LycheeMemory segments the input into chunks and encodes each into compressed KV-cache-style representations using a Compressor.A Gate then dynamically selects relevant memory blocks, which a Reasoner iteratively processes with an evolving working memory to solve downstream tasks.The Compressor and Reasoner are jointly optimized via end-to-end reinforcement learning, while the Gate is trained separately as a classifier.Experimental results demonstrate that LycheeMemory achieves competitive accuracy (up to 82% in ablation variants) on multi-hop reasoning benchmarks (e.g., RULER-HQA), successfully extrapolates context length from 7K to 1.75M, and provides a favorable accuracy-efficiency trade-off against strong long-context baselines.Notably, compared to MemAgent, LycheeMemory achieves an average 2× reduction in peak GPU memory usage and a 6× speedup during inference. Zhuoen Chen, Dongfang Li 0002, Meishan Zhang, Baotian Hu, Min Zhang 0005 |
ACL (1) | 4 |
| 2026 | Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question AnsweringabstractExisting long-document question answering systems typically process texts as flat sequences or use heuristic chunking, which overlook the discourse structures that naturally guide human comprehension.We present a discourse-aware hierarchical framework that leverages rhetorical structure theory (RST) for long document question answering.Our approach converts discourse trees into sentence-level representations and employs LLM-enhanced node representations to bridge structural and semantic information.The framework involves three key innovations: language-universal discourse parsing for lengthy documents, LLM-based enhancement of discourse relation nodes, and structure-guided hierarchical retrieval.Extensive experiments on four datasets demonstrate consistent improvements over existing approaches through the incorporation of discourse structure, across multiple genres and languages.Moreover, the proposed framework exhibits strong robustness across diverse document types and linguistic settings. Huiyao Chen, Meishan Zhang, Baotian Hu, Min Zhang 0005 |
ACL (1) | 5 |
| 2026 | ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded ExecutionabstractLarge Language Models (LLMs) enhance their problem-solving capability by utilizing external tools.However, in open-world scenarios with massive and evolving tool repositories, existing methods relying on static embedding retrieval or parameter memorization of tools struggle to align user intent with tool semantics or generalize to unseen tools, respectively, leading to suboptimal accuracy of open-world tool retrieval and execution.To address these, we present ToolOmni, a unified agentic framework that enables LLMs for open-world tool use by proactive retrieval and grounded execution within a reasoning loop.First, we construct a cold-start multi-turn interaction dataset to instill foundational agentic capabilities via Supervised Fine-Tuning (SFT).Then, we introduce open-world tool learning based on a Decoupled Multi-Objective GRPO algorithm, which simultaneously optimizes LLMs for both tool retrieval accuracy and execution efficacy in online environments.Extensive experiments demonstrate that ToolOmni achieves state-ofthe-art performance both in retrieval and execution, surpassing strong baselines by a significant margin of +10.8% in end-to-end execution success rate, while exhibiting exceptional robustness and generalization capabilities. Shouzheng Huang, Meishan Zhang, Baotian Hu, Min Zhang 0005 |
ACL (1) | 3 |
| 2026 | UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-ExpertsabstractZhenyu Liu, Yunxin li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Xinyu Chen, Haoyuan Shi, Haolan Chen, Fanbo Meng, Mingjun Zhao, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yunxin Li, Xuanyu Zhang 0006, Qixun Teng, Shenyuan Jiang, Xinyu Chen 0003, Haolan Chen, Mingjun Zhao, Yancheng He, Baotian Hu, Haizhou Li 0001, Min Zhang 0005 |
ACL (1) | 13 |
| 2026 | Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMsabstractZhenyu Liu, Xuanyu Zhang, Yunxin li, Qixun Teng, Shenyuan Jiang, Haolan Chen, Mingjun Zhao, Fanbo Meng, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xuanyu Zhang 0006, Yunxin Li, Qixun Teng, Shenyuan Jiang, Haolan Chen, Mingjun Zhao, Yancheng He, Baotian Hu, Haizhou Li 0001, Min Zhang 0005 |
ACL (1) | 11 |
| 2026 | Structured Episodic Event MemoryabstractCurrent approaches to memory in Large Language Models (LLMs) predominantly rely on static Retrieval-Augmented Generation (RAG), which often results in scattered retrieval and fails to capture the structural dependencies required for complex reasoning.For autonomous agents, these passive and flat architectures lack the cognitive organization necessary to model the dynamic and associative nature of longterm interaction.To address this, we propose Structured Episodic Event Memory (SEEM), a hierarchical framework that synergizes a graph memory layer for relational facts with a dynamic episodic memory layer for narrative progression.Grounded in cognitive frame theory, SEEM transforms interaction streams into structured Episodic Event Frames (EEFs) anchored by precise provenance pointers.Furthermore, we introduce an agentic associative fusion and Reverse Provenance Expansion (RPE) mechanism to reconstruct coherent narrative contexts from fragmented evidence.Experimental results on the LoCoMo and Long-MemEval benchmarks demonstrate that SEEM significantly outperforms baselines, enabling agents to maintain superior narrative coherence and logical consistency. Zhengxuan Lu, Dongfang Li 0002, Yukun Shi, Beilun Wang, Longyue Wang, Baotian Hu |
ACL (1) | 6 |
| 2026 | Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning FrameworkabstractChenyuan Zhang, Qiguang Chen, Xie Chen, Zhuotao Tian, Bowen Xing, Meishan Zhang, Libo Qin, Baotian Hu, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qiguang Chen, Xie Chen 0001, Zhuotao Tian, Meishan Zhang, Libo Qin 0001, Baotian Hu, Min Zhang 0005 |
ACL (1) | 8 |
| 2026 | Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMsabstractRecent advancements in multimodal large language models (MLLMs) have achieved significant multimodal generation capabilities, akin to GPT-4. These models predominantly map visual information into language representation space, leveraging the vast knowledge and powerful text generation abilities of LLMs to produce multimodal instruction-following responses. We could term this method as LLMs for Vision because of its employing LLMs for visual understanding and reasoning, yet observe that these MLLMs neglect the potential of harnessing visual knowledge to enhance the overall capabilities of LLMs, which could be regarded as Vision Enhancing LLMs. In this paper, we propose an approach called MKS2, aimed at enhancing LLMs through empowering Multimodal Knowledge Storage and Sharing in LLMs. Specifically, we introduce Modular Visual Memory (MVM), a component integrated into the internal blocks of LLMs, designed to store open-world visual information efficiently. Additionally, we present a soft Mixture of Multimodal Experts (MoMEs) architecture in LLMs to invoke multimodal knowledge collaboration during text generation. Our comprehensive experiments demonstrate that MKS2 substantially augments the reasoning capabilities of LLMs in contexts necessitating physical or commonsense knowledge. It also delivers competitive results on image-text understanding multimodal benchmarks. The codes will be available at: https://github.com/HITsz-TMG/MKS2-Multimodal-Knowledge-Storage-and-Sharing. Yunxin Li, Baotian Hu, Wei Wang 0335, Xiaochun Cao, Min Zhang 0005 |
IEEE Trans. Image Process. | 3 |
| 2025 | CMT: A Memory Compression Method for Continual Knowledge Learning of Large Language ModelsabstractLarge Language Models (LLMs) need to adapt to the continuous changes in data, tasks, and user preferences. Due to their massive size and the high costs associated with training, LLMs are not suitable for frequent retraining. However, updates are necessary to keep them in sync with rapidly evolving human knowledge. To address these challenges, this paper proposes the Compression Memory Training (CMT) method, an efficient and effective online adaptation framework for LLMs that features robust knowledge retention capabilities. Inspired by human memory mechanisms, CMT compresses and extracts information from new documents to be stored in a memory bank. When answering to queries related to these new documents, the model aggregates these document memories from the memory bank to better answer user questions. The parameters of the LLM itself do not change during training and inference, reducing the risk of catastrophic forgetting. To enhance the encoding, retrieval, and aggregation of memory, we further propose three new general and flexible techniques, including memory-aware objective, self-matching and top-k aggregation. Extensive experiments conducted on three continual learning datasets (i.e., StreamingQA, SQuAD and ArchivalQA) demonstrate that the proposed method improves model adaptability and robustness across multiple base LLMs (e.g., +4.07 EM & +4.19 F1 in StreamingQA with Llama-2-7b). Dongfang Li 0002, Zetian Sun, Xinshuo Hu, Baotian Hu, Min Zhang 0005 |
AAAI | 4 |
| 2025 | VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video ComprehensionabstractXinyu Chen, Yunxin Li, Haoyuan Shi, Baotian Hu, Wenhan Luo, Yaowei Wang, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xinyu Chen 0003, Yunxin Li, Baotian Hu, Wenhan Luo, Yaowei Wang 0001, Min Zhang 0005 |
ACL (1) | 4 |
| 2025 | A Unified Agentic Framework for Evaluating Conditional Image GenerationabstractConditional image generation has gained significant attention for its ability to personalize content. However, the field faces challenges in developing task-agnostic, reliable, and explainable evaluation metrics. This paper introduces CIGEval, a unified agentic framework for comprehensive evaluation of conditional image generation tasks. CIGEval utilizes large multimodal models (LMMs) as its core, integrating a multi-functional toolbox and establishing a fine-grained evaluation framework. Additionally, we synthesize evaluation trajectories for fine-tuning, empowering smaller LMMs to autonomously select appropriate tools and conduct nuanced analyses based on tool outputs. Experiments across seven prominent conditional image generation tasks demonstrate that CIGEval (GPT-4o version) achieves a high correlation of 0.4625 with human assessments, closely matching the inter-annotator correlation of 0.47. Notably, when implemented with 7B open-source LMMs using only 2.3K training trajectories, CIGEval surpasses the previous GPT-4o-based state-of-the-art method. These findings indicate that CIGEval holds great potential for automating evaluation of image generation tasks while maintaining human-level reliability. Jifang Wang, Yangxue, Longyue Wang, Zhenran Xu, Yaowei Wang 0001, Weihua Luo, Kaifu Zhang, Baotian Hu, Min Zhang 0005 |
ACL (1) | 9 |
| 2025 | Advancing Temporal Sensitive Question Answering through Progressive Multi-Step ReflectionabstractRetrieval-augmented generation (RAG) has demonstrated strong potential in enhancing large language models (LLMs) for complex, real-world question answering. However, existing RAG frameworks remain inadequate for temporal scenarios, primarily due to their inability to jointly model temporal constraints in both retrieval and reasoning. On the retrieval side, traditional approaches focus on semantic similarity, often returning outdated or temporally misaligned evidence. On the generation side, these systems frequently produce factually incorrect or hallucinated answers when confronted with incomplete or temporally inconsistent information. Motivated by the observed limitations, we propose ChronoReflect+, a temporal logic-aware RAG framework that incorporates hybrid temporal-aware retrieval and progressive multi-step reflection. Our method iteratively refines both retrieval and reasoning, identifying and bridging information gaps as context accumulates. Extensive experiments demonstrate that ChronoReflect+ significantly outperforms state-of-the-art RAG baselines-improving end-to-end accuracy by 15.2%-particularly on questions involving implicit time expressions and multi-hop reasoning. Erxue Min, Xiang Zhao 0002, Yunxin Li, Jinzhi Liao, Shuaiqiang Wang, Baotian Hu, Dawei Yin 0001 |
CIKM | 8 |
| 2025 | AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip GenerationabstractDespite rapid advancements in video generation models, generating coherent, long-form storytelling videos that span multiple scenes and characters remains challenging. Current methods often rigidly convert pre-generated keyframes into fixed-length clips, resulting in disjointed narratives and pacing issues. Furthermore, the inherent instability of video generation models means that even a single low-quality clip can significantly degrade the entire output animation’s logical coherence and visual continuity. To overcome these obstacles, we introduce AniMaker, a multi-agent framework enabling efficient multi-candidate clip generation and storytelling-aware clip selection, thus creating globally consistent and story-coherent animation solely from text input. The framework is structured around specialized agents, including the Director Agent for storyboard generation, the Photography Agent for video clip generation, the Reviewer Agent for evaluation, and the Post-Production Agent for editing and voiceover, collectively realizing multi-character, multi-scene animation. Central to AniMaker’s approach are two key technical components: MCTS-Gen in Photography Agent, an efficient Monte Carlo Tree Search (MCTS)-inspired strategy that intelligently navigates the candidate space to generate high-potential clips while optimizing resource usage; and AniEval in Reviewer Agent, the first framework specifically designed for multi-shot animation evaluation, which assesses critical aspects such as story-level consistency, action completion, and animation-specific features by considering each clip in the context of its preceding and succeeding clips. Experiments demonstrate that AniMaker achieves superior quality as measured by popular metrics including VBench and our proposed AniEval framework, while significantly improving the efficiency of multi-candidate generation, pushing AI-generated storytelling animation closer to production standards. Code and data for this paper are at https://animaker-dev.github.io/ Yunxin Li, Xinyu Chen 0003, Longyue Wang, Baotian Hu, Min Zhang 0005 |
SIGGRAPH Asia | 5 |
| 2025 | Uni-MoE: Scaling Unified Multimodal LLMs With Mixture of ExpertsabstractRecent advancements in Multimodal Large Language Models (MLLMs) underscore the significance of scalable models and data to boost performance, yet this often incurs substantial computational costs. Although the Mixture of Experts (MoE) architecture has been employed to scale large language or visual-language models efficiently, these efforts typically involve fewer experts and limited modalities. To address this, our work presents the pioneering attempt to develop a unified MLLM with the MoE architecture, named Uni-MoE that can handle a wide array of modalities. Specifically, it features modality-specific encoders with connectors for a unified multimodal representation. We also implement a sparse MoE architecture within the LLMs to enable efficient training and inference through modality-level data parallelism and expert-level model parallelism. To enhance the multi-expert collaboration and generalization, we present a progressive training strategy: 1) Cross-modality alignment using various connectors with different cross-modality data, 2) Training modality-specific experts with cross-modality instruction data to activate experts' preferences, and 3) Tuning the whole Uni-MoE framework utilizing Low-Rank Adaptation (LoRA) on mixed multimodal instruction data. We evaluate the instruction-tuned Uni-MoE on a comprehensive set of multimodal datasets. The extensive experimental results demonstrate Uni-MoE's principal advantage of significantly reducing performance bias in handling mixed multimodal datasets, alongside improved multi-expert collaboration and generalization. Yunxin Li, Shenyuan Jiang, Baotian Hu, Longyue Wang, Wanqi Zhong, Wenhan Luo, Lin Ma 0002, Min Zhang 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | A vision-language model with multi-granular knowledge fusion in medical imaging
Kai Chen 0020, Yunxin Li, Xiwen Zhu, Wentai Zhang 0003, Baotian Hu |
World Wide Web (WWW) | 5 |
| 2024 | Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module OperationabstractLarge language models (LLMs) have been widely used in various applications but are known to suffer from issues related to untruthfulness and toxicity. While parameter-efficient modules (PEMs) have demonstrated their effectiveness in equipping models with new skills, leveraging PEMs for deficiency unlearning remains underexplored. In this work, we propose a PEMs operation approach, namely Extraction-before-Subtraction (Ext-Sub), to enhance the truthfulness and detoxification of LLMs through the integration of ``expert'' PEM and ``anti-expert'' PEM. Remarkably, even anti-expert PEM possess valuable capabilities due to their proficiency in generating fabricated content, which necessitates language modeling and logical narrative competence. Rather than merely negating the parameters, our approach involves extracting and eliminating solely the deficiency capability within anti-expert PEM while preserving the general capabilities. To evaluate the effectiveness of our approach in terms of truthfulness and detoxification, we conduct extensive experiments on LLMs, encompassing additional abilities such as language modelling and mathematical reasoning. Our empirical results demonstrate that our approach effectively improves truthfulness and detoxification, while largely preserving the fundamental abilities of LLMs. Xinshuo Hu, Dongfang Li 0002, Baotian Hu, Min Zhang 0005 |
AAAI | 3 |
| 2024 | Temporal Knowledge Question Answering via Abstract Reasoning InductionabstractIn this study, we address the challenge of enhancing temporal knowledge reasoning in Large Language Models (LLMs).LLMs often struggle with this task, leading to the generation of inaccurate or misleading responses.This issue mainly arises from their limited ability to handle evolving factual knowledge and complex temporal logic.To overcome these limitations, we propose Abstract Reasoning Induction (ARI) framework, which divides temporal reasoning into two distinct phases: Knowledgeagnostic and Knowledge-based.This framework offers factual knowledge support to LLMs while minimizing the incorporation of extraneous noisy data.Concurrently, informed by the principles of constructivism, ARI provides LLMs the capability to engage in proactive, self-directed learning from both correct and incorrect historical reasoning samples.By teaching LLMs to actively construct knowledge and methods, it can significantly boosting their temporal reasoning abilities.Our approach achieves significant improvements, with relative gains of 29.7% and 9.27% on two temporal QA datasets, underscoring its efficacy in advancing temporal reasoning in LLMs.The code can be found at https: //github.com/czy1999/ARI-QA. Dongfang Li 0002, Xiang Zhao 0002, Baotian Hu, Min Zhang 0005 |
ACL (1) | 4 |
| 2024 | Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge AlignmentabstractEvaluating and Rethinking the current landscape of Large Multimodal Models (LMMs), we observe that widely-used visual-language projection approaches (e.g., Q-former or MLP) focus on the alignment of image-text descriptions yet ignore the visual knowledgedimension alignment, i.e., connecting visuals to their relevant knowledge.Visual knowledge plays a significant role in analyzing, inferring, and interpreting information from visuals, helping improve the accuracy of answers to knowledge-based visual questions.In this paper, we mainly explore improving LMMs with visual-language knowledge alignment, especially aimed at challenging knowledge-based visual question answering (VQA).To this end, we present a Cognitive Visual-Language Mapper (CVLM), which contains a pretrained Visual Knowledge Aligner (VKA) and a Finegrained Knowledge Adapter (FKA) used in the multimodal instruction tuning stage.Specifically, we design the VKA based on the interaction between a small language model and a visual encoder, training it on collected imageknowledge pairs to achieve visual knowledge acquisition and projection.FKA is employed to distill the fine-grained visual knowledge of an image and inject it into Large Language Models (LLMs).We conduct extensive experiments on knowledge-based VQA benchmarks and experimental results show that CVLM significantly improves the performance of LMMs on knowledge-based VQA (average gain by 5.0%).Ablation studies also verify the effectiveness of VKA and FKA, respectively.1 Yunxin Li, Xinyu Chen 0003, Baotian Hu, Min Zhang 0005 |
ACL (1) | 3 |
| 2024 | Does the Generator Mind Its Contexts? An Analysis of Generative Model Faithfulness under Context Transferabstracthe present study introduces the knowledge-augmented generator, which is specifically designed to produce information that remains grounded in contextual knowledge, regardless of alterations in the context. Previous research has predominantly focused on examining hallucinations stemming from static input, such as in the domains of summarization or machine translation. However, our investigation delves into the faithfulness of generative question answering in the presence of dynamic knowledge. Our objective is to explore the existence of hallucinations arising from parametric memory when contextual knowledge undergoes changes, while also analyzing the underlying causes for their occurrence. In order to efficiently address this issue, we propose a straightforward yet effective measure for detecting such hallucinations. Intriguingly, our investigation uncovers that all models exhibit a tendency to generate previous answers as hallucinations. To gain deeper insights into the underlying causes of this phenomenon, we conduct a series of experiments that verify the critical role played by context in hallucination, both during training and testing, from various perspectives. Xinshuo Hu, Dongfang Li 0002, Yuxiang Wu, Lifeng Shang, Baotian Hu |
LREC/COLING | 6 |
| 2024 | A Multimodal In-Context Tuning Approach for E-Commerce Product Description GenerationabstractIn this paper, we propose a new setting for generating product descriptions from images, augmented by marketing keywords. It leverages the combined power of visual and textual information to create descriptions that are more tailored to the unique features of products. For this setting, previous methods utilize visual and textual encoders to encode the image and keywords and employ a language model-based decoder to generate the product description. However, the generated description is often inaccurate and generic since same-category products have similar copy-writings, and optimizing the overall framework on large-scale samples makes models concentrate on common words yet ignore the product features. To alleviate the issue, we present a simple and effective Multimodal In-Context Tuning approach, named ModICT, which introduces a similar product sample as the reference and utilizes the in-context learning capability of language models to produce the description. During training, we keep the visual encoder and language model frozen, focusing on optimizing the modules responsible for creating multimodal in-context references and dynamic prompts. This approach preserves the language generation prowess of large language models (LLMs), facilitating a substantial increase in description diversity. To assess the effectiveness of ModICT across various language model scales and types, we collect data from three distinct product categories within the E-commerce domain. Extensive experiments demonstrate that ModICT significantly improves the accuracy (by up to 3.3% on Rouge-L) and diversity (by up to 9.4% on D-5) of generated results compared to conventional methods. Our findings underscore the potential of ModICT as a valuable tool for enhancing the automatic generation of product descriptions in a wide range of applications. Data and code are at https://github.com/HITsz-TMG/Multimodal-In-Context-Tuning Yunxin Li, Baotian Hu, Wenhan Luo, Lin Ma 0002, Min Zhang 0005 |
LREC/COLING | 2 |
| 2024 | Generative Multimodal Entity LinkingabstractMultimodal Entity Linking (MEL) is the task of mapping mentions with multimodal contexts to the referent entities from a knowledge base. Existing MEL methods mainly focus on designing complex multimodal interaction mechanisms and require fine-tuning all model parameters, which can be prohibitively costly and difficult to scale in the era of Large Language Models (LLMs). In this work, we propose GEMEL, a Generative Multimodal Entity Linking framework based on LLMs, which directly generates target entity names. We keep the vision and language model frozen and only train a feature mapper to enable cross-modality interactions. To adapt LLMs to the MEL task, we leverage the in-context learning capability of LLMs by retrieving multimodal instances as demonstrations. Extensive experiments show that, with only ∼0.3% of the model parameters fine-tuned, GEMEL achieves state-of-the-art results on two well-established MEL datasets (7.7% accuracy gains on WikiDiverse and 8.8% accuracy gains on WikiMEL). The performance gain stems from mitigating the popularity bias of LLM predictions and disambiguating less common entities effectively. Further analysis verifies the generality and scalability of GEMEL. Our framework is compatible with any off-the-shelf language model, paving the way towards an efficient and general solution for utilizing LLMs in the MEL task. Our code is available at https://github.com/HITsz-TMG/GEMEL. Senbao Shi, Zhenran Xu, Baotian Hu, Min Zhang 0005 |
LREC/COLING | 3 |
| 2024 | Take Off the Training Wheels! Progressive In-Context Learning for Effective AlignmentabstractRecent studies have explored the working mechanisms of In-Context Learning (ICL).However, they mainly focus on classification and simple generation tasks, limiting their broader application to more complex generation tasks in practice.To address this gap, we investigate the impact of demonstrations on token representations within the practical alignment tasks.We find that the transformer embeds the task function learned from demonstrations into the separator token representation, which plays an important role in the generation of prior response tokens.Once the prior response tokens are determined, the demonstrations become redundant.Motivated by this finding, we propose an efficient Progressive In-Context Alignment (PICA) method consisting of two stages.In the first few-shot stage, the model generates several prior response tokens via standard ICL while concurrently extracting the ICL vector that stores the task function from the separator token representation.In the following zero-shot stage, this ICL vector guides the model to generate responses without further demonstrations.Extensive experiments demonstrate that our PICA not only surpasses vanilla ICL but also achieves comparable performance to other alignment tuning methods.The proposed training-free method reduces the time cost (e.g., 5.45×) with improved alignment performance (e.g., 6.57+).Consequently, our work highlights the application of ICL for alignment and calls for a deeper understanding of ICL for complex generations. Dongfang Li 0002, Xinshuo Hu, Xinping Zhao, Yibin Chen, Baotian Hu, Min Zhang 0005 |
EMNLP | 6 |
| 2024 | SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented GenerationabstractRecent studies in Retrieval-Augmented Generation (RAG) have investigated extracting evidence from retrieved passages to reduce computational costs and enhance the final RAG performance, yet it remains challenging.Existing methods heavily rely on heuristic-based augmentation, encountering several issues: (1) Poor generalization due to hand-crafted context filtering; (2) Semantics deficiency due to rulebased context chunking; (3) Skewed length due to sentence-wise filter learning.To address these issues, we propose a model-based evidence extraction learning framework, SEER, optimizing a vanilla model as an evidence extractor with desired properties through selfaligned learning.Extensive experiments show that our method largely improves the final RAG performance, enhances the faithfulness, helpfulness, and conciseness of the extracted evidence, and reduces the evidence length by 9.25 times.The code will be available at https://github.com/HITsz-TMG/SEER. Xinping Zhao, Dongfang Li 0002, Boren Hu, Yibin Chen, Baotian Hu, Min Zhang 0005 |
EMNLP | 6 |
| 2024 | VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual ContextabstractLarge Multimodal Models (LMMs) have achieved impressive success in visual reasoning, particularly in visual mathematics. However, problem-solving capabilities in graph theory remain less explored for LMMs, despite being a crucial aspect of mathematical reasoning that requires an accurate understanding of graphical structures and multi-step reasoning on visual graphs. To step forward in this direction, we are the first to design a benchmark named VisionGraph, used to explore the capabilities of advanced LMMs in solving multimodal graph theory problems. It encompasses eight complex graph problem tasks, from connectivity to shortest path problems. Subsequently, we present a Description-Program-Reasoning (DPR) chain to enhance the logical accuracy of reasoning processes through graphical structure description generation and algorithm-aware multi-step reasoning. Our extensive study shows that 1) GPT-4V outperforms Gemini Pro in multi-step graph reasoning; 2) All LMMs exhibit inferior perception accuracy for graphical structures, whether in zero/few-shot settings or with supervised fine-tuning (SFT), which further affects problem-solving performance; 3) DPR significantly improves the multi-step graph reasoning capabilities of LMMs and the GPT-4V (DPR) agent achieves SOTA performance. Yunxin Li, Baotian Hu, Wei Wang 0164, Longyue Wang, Min Zhang 0005 |
ICML | 2 |
| 2024 | Enhancing Attributed Graph Networks with Alignment and Uniformity Constraints for Session-based RecommendationabstractSession-based Recommendation (SBR), seeking to predict a user’s next action based on an anonymous session, has drawn increasing attention for its practicability. Most SBR models only rely on the contextual transitions within a short session to learn item representations while neglecting additional valuable knowledge. As such, their model capacity is largely limited by the data sparsity issue caused by short sessions. A few studies have exploited the Modeling of Item Attributes (MIA) to enrich item representations. However, they usually involve specific model designs that can hardly transfer to existing attribute-agnostic SBR models and thus lack universality. In this paper, we propose a model-agnostic framework, named AttrGAU (Attributed Graph Networks with Alignment and Uniformity Constraints), to bring the MIA’s superiority into existing attribute-agnostic models, to improve their accuracy and robustness for recommendation. Specifically, we first build a bipartite attributed graph and design an attribute-aware graph convolution to exploit the rich attribute semantics hidden in the heterogeneous item-attribute relationship. We then decouple existing attribute-agnostic SBR models into the graph neural network and attention readout sub-modules to satisfy the non-intrusive requirement. Lastly, we design two representation constraints, i.e., alignment and uniformity, to optimize distribution discrepancy in representation between the attribute semantics and collaborative semantics. Extensive experiments on three public benchmark datasets demonstrate that the proposed AttrGAU framework can significantly enhance backbone models’ recommendation performance and robustness against data sparsity and data noise issues. Our implementation codes will be available at https://github.com/ItsukiFujii/AttrGAU. Xinping Zhao, Chaochao Chen 0001, Jiajie Su, Yizhao Zhang, Baotian Hu |
ICWS | 5 |
| 2024 | In-Context Learning State Vector with Inner and Momentum OptimizationabstractLarge Language Models (LLMs) have exhibited an impressive ability to perform In-Context Learning (ICL) from only a few examples. Recent works have indicated that the functions learned by ICL can be represented through compressed vectors derived from the transformer. However, the working mechanisms and optimization of these vectors are yet to be thoroughly explored. In this paper, we address this gap by presenting a comprehensive analysis of these compressed vectors, drawing parallels to the parameters trained with gradient descent, and introducing the concept of state vector. Inspired by the works on model soup and momentum-based gradient descent, we propose inner and momentum optimization methods that are applied to refine the state vector progressively as test-time adaptation. Moreover, we simulate state vector aggregation in the multiple example setting, where demonstrations comprising numerous examples are usually too lengthy for regular ICL, and further propose a divide-and-conquer aggregation method to address this challenge. We conduct extensive experiments using Llama-2 and GPT-J in both zero-shot setting and few-shot setting. The experimental results show that our optimization method effectively enhances the state vector and achieves the state-of-the-art performance on diverse tasks. Dongfang Li 0002, Xinshuo Hu, Zetian Sun, Baotian Hu, Min Zhang 0005 |
NeurIPS | 5 |
| 2024 | SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-ReflectionabstractInstruction tuning (IT) is crucial to tailoring large language models (LLMs) towards human-centric interactions. Recent advancements have shown that the careful selection of a small, high-quality subset of IT data can significantly enhance the performance of LLMs. Despite this, common approaches often rely on additional models or data, which increases costs and limits widespread adoption. In this work, we propose a novel approach, termed $\textit{SelectIT}$, that capitalizes on the foundational capabilities of the LLM itself. Specifically, we exploit the intrinsic uncertainty present in LLMs to more effectively select high-quality IT data, without the need for extra resources. Furthermore, we introduce a curated IT dataset, the $\textit{Selective Alpaca}$, created by applying SelectIT to the Alpaca-GPT4 dataset. Empirical results demonstrate that IT using Selective Alpaca leads to substantial model ability enhancement. The robustness of SelectIT has also been corroborated in various foundation models and domain-specific tasks. Our findings suggest that longer and more computationally intensive IT data may serve as superior sources of IT, offering valuable insights for future research in this area. Data, code, and scripts are freely available at https://github.com/Blue-Raincoat/SelectIT. Liangxin Liu, Xuebo Liu 0002, Derek F. Wong, Dongfang Li 0002, Baotian Hu, Min Zhang 0005 |
NeurIPS | 6 |
| 2024 | Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generationabstractdemands substantial human effort and incurs high training Yunxin Li, Baotian Hu, Longyue Wang, Jiashun Zhu, Jinyi Xu, Zhen Zhao 0001, Min Zhang 0005 |
SIGGRAPH Asia | 3 |
| 2024 | BioPRO: Context-Infused Prompt Learning for Biomedical Entity LinkingabstractRecent research tends to address the biomedical entity linking problem in a unified framework solely based on surface form matching between mentions and entities. Specifically, these methods focus on addressing thevarietychallenge of the heterogeneous naming of biomedical concepts. Yet, theambiguitychallenge that the same word under different contexts can be used to refer to distinct concepts is usually ignored. To address this challenge, we propose BioPRO, a two-stage entity linking algorithm to enhance the biomedical entity representations based on context-infused prompt learning. The first stage includes a coarse-grained retrieval from a representation space defined by a bi-encoder that independently embeds the mention and entity's surface forms. Unlike previous one-model-fits-all systems, each candidate is then re-ranked with a fine-grained encoder based on prompt-tuning that sufficiently stimulates knowledge in contextual information of mentions and entities. Furthermore, the trained fine-grained encoder can be utilized to generate deep representations of bio-entities and boost candidate retrieval in the first stage. Extensive experiments show that our model achieves promising performance improvements compared with several state-of-the-art (SOTA) techniques on 4 biomedical corpora. We also observe by cases that the proposed context-infused prompt-tuning strategy is effective in solving both thevarietyandambiguitychallenges in the linking task. Tiantian Zhu 0002, Yang Qin 0001, Ming Feng, Qingcai Chen, Baotian Hu, Yang Xiang 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | LMEye: An Interactive Perception Network for Large Language ModelsabstractCurrent efficient approaches to building Multimodal Large Language Models (MLLMs) mainly incorporate visual information into LLMs with a simple visual mapping network such as a linear projection layer, a multilayer perceptron (MLP), or Q-former from BLIP-2. Such networks project the image feature once and do not consider the interaction between the image and the human inputs. Hence, the obtained visual information without being connected to human intention may be inadequate for LLMs to generate intention-following responses, which we refer to as static visual information. To alleviate this issue, our paper introduces LMEye, a human-like eye with a play-and-plug interactive perception network, designed to enable dynamic interaction between LLMs and external visual information. It can allow the LLM to request the desired visual information aligned with various human instructions, which we term dynamic visual information acquisition. Specifically, LMEye consists of a simple visual mapping network to provide the basic perception of an image for LLMs. It also contains additional modules responsible for acquiring requests from LLMs, performing request-based visual information seeking, and transmitting the resulting interacted visual information to LLMs, respectively. In this way, LLMs act to understand the human query, deliver the corresponding request to the request-based visual information interaction module, and generate the response based on the interleaved multimodal information. We evaluate LMEye through extensive experiments on multimodal benchmarks, demonstrating that it significantly improves zero-shot performances on various multimodal tasks compared to previous methods, with fewer parameters. Moreover, we also verify its effectiveness and scalability on various language models and video understanding, respectively. Yunxin Li, Baotian Hu, Xinyu Chen 0003, Lin Ma 0002, Yong Xu 0001, Min Zhang 0005 |
IEEE Trans. Multim. | 2 |
| 2023 | Improving Biomedical Entity Linking with Cross-Entity InteractionabstractBiomedical entity linking (EL) is the task of linking mentions in a biomedical document to corresponding entities in a knowledge base (KB). The challenge in biomedical EL lies in leveraging mention context to select the most appropriate entity among possible candidates. Although some EL models achieve competitive results by retrieving candidate entities and then exploiting context to re-rank them, these re-ranking models concatenate mention context with one candidate at a time. They lack fine-grained interaction among candidates, and potentially cannot handle ambiguous mentions when facing candidates both with high lexical similarity. We cope with this issue using a re-ranking model based on prompt tuning, which represents mention context and all candidates at once, letting candidates in comparison attend to each other. We also propose a KB-enhanced self-supervised pretraining strategy. Instead of large-scale pretraining on biomedical EL data in previous work, we use masked language modeling with synonyms from KB. Our method achieves state-of-the-art results on 3 biomedical EL datasets: NCBI disease, BC5CDR and COMETA, showing the effectiveness of cross-entity interaction and KB-enhanced pretraining strategy. Code is available at https://github.com/HITsz-TMG/Prompt-BioEL. Zhenran Xu, Baotian Hu |
AAAI | 3 |
| 2023 | A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual CluesabstractConditional inference on joint textual and visual clues is a multi-modal reasoning task that textual clues provide prior permutation or external knowledge, which are complementary with visual content and pivotal to deducing the correct option.Previous methods utilizing pretrained vision-language models (VLMs) have achieved impressive performances, yet they show a lack of multimodal context reasoning capability, especially for text-modal information.To address this issue, we propose a Multi-modal Context Reasoning approach, named ModCR.Compared to VLMs performing reasoning via cross modal semantic alignment, it regards the given textual abstract semantic and objective image information as the pre-context information and embeds them into the language model to perform context reasoning.Different from recent vision-aided language models used in natural language processing, ModCR incorporates the multi-view semantic alignment information between language and vision by introducing the learnable alignment prefix between image and text in the pretrained language model.This makes the language model well-suitable for such multi-modal reasoning scenario on joint textual and visual clues.We conduct extensive experiments on two corresponding data sets and experimental results show significantly improved performance (exact gain by 4.8% on PMR test set) compared to previous strong baselines.Code Yunxin Li, Baotian Hu, Xinyu Chen 0003, Lin Ma 0002, Min Zhang 0005 |
ACL (1) | 2 |
| 2023 | A Neural Divide-and-Conquer Reasoning Framework for Image Retrieval from Linguistically Complex TextabstractPretrained Vision-Language Models (VLMs) have achieved remarkable performance in image retrieval from text.However, their performance drops drastically when confronted with linguistically complex texts that they struggle to comprehend.Inspired by the Divide-and-Conquer (Smith, 1985) algorithm and dualprocess theory (Groves and Thompson, 1970), in this paper, we regard linguistically complex texts as compound proposition texts composed of multiple simple proposition sentences and propose an end-to-end Neural Divide-and-Conquer Reasoning framework, dubbed NDCR.It contains three main components: 1) Divide: a proposition generator divides the compound proposition text into simple proposition sentences and produces their corresponding representations, 2) Conquer: a pretrained VLMsbased visual-linguistic interactor achieves the interaction between decomposed proposition sentences and images, 3) Combine: a neuralsymbolic reasoner combines the above reasoning states to obtain the final solution via a neural logic reasoning approach.According to the dual-process theory, the visual-linguistic interactor and neural-symbolic reasoner could be regarded as analogical reasoning System 1 and logical reasoning System 2. We conduct extensive experiments on a challenging image retrieval from contextual descriptions data set.Experimental results and analyses indicate NDCR significantly improves performance in the complex image-text reasoning problem. Yunxin Li, Baotian Hu, Lin Ma 0002, Min Zhang 0005 |
ACL (1) | 2 |
| 2023 | Revisiting Sparse Retrieval for Few-shot Entity LinkingabstractEntity linking aims to link ambiguous mentions to their corresponding entities in a knowledge base.One of the key challenges comes from insufficient labeled data for specific domains.Although dense retrievers have achieved excellent performance on several benchmarks, their performance decreases significantly when only a limited amount of in-domain labeled data is available.In such few-shot setting, we revisit the sparse retrieval method, and propose an ELECTRA-based keyword extractor to denoise the mention context and construct a better query expression.For training the extractor, we propose a distant supervision method to automatically generate training data based on overlapping tokens between mention contexts and entity descriptions.Experimental results on the ZESHEL dataset demonstrate that the proposed method outperforms state-of-the-art models by a significant margin across all test domains, showing the effectiveness of keywordenhanced sparse retrieval.Code is available at https://github.com/HITsz-TMG/ Zhenran Xu, Baotian Hu, Min Zhang 0005 |
EMNLP | 3 |
| 2023 | HITSZ TMG at ICASSP 2023 SPGC Shared Task: Leveraging Pre-Training and Distillation Method for Title Generation with Limited ResourceabstractIn this paper, we present our proposed method for the shared task of the ICASSP 2023 Signal Processing Grand Challenge (SPGC). We participate in Topic Title Generation (TTG), Track 3 of General Meeting Understanding and Generation (MUG) [1] in SPGC. The primary objective of this task is to generate a title that effectively summarizes the given topic segment. With the constraints of limited model size and external dataset availability, we propose a method as Pre-training - Distillation / Fine-tuning (PDF), which can efficiently leverage the knowledge from large model and corpus. Our method achieves first place during preliminary and final contests in ICASSP2023 MUG Challenge Track 3. Tianxiao Xu, Xinshuo Hu, Zetian Sun, Yu Zhao 0043, Baotian Hu |
ICASSP | 6 |
| 2023 | PPAT: Progressive Graph Pairwise Attention Network for Event Causality IdentificationabstractEvent Causality Identification (ECI) aims to identify the causality between a pair of event mentions in a document, which is composed of sentence-level ECI (SECI) and document-level ECI (DECI). Previous work applies various reasoning models to identify the implicit event causality. However, they indiscriminately reason all event causality in the same way, ignoring that most inter-sentence event causality depends on intra-sentence event causality to infer. In this paper, we propose a progressive graph pairwise attention network (PPAT) to consider the above dependence. PPAT applies a progressive reasoning strategy, as it first predicts the intra-sentence event causality, and then infers the more implicit inter-sentence event causality based on the SECI result. We construct a sentence boundary event relational graph, and PPAT leverages a simple pairwise attention mechanism, which attends to different reasoning chains on the graph. In addition, we propose a causality-guided training strategy for assisting PPAT in learning causality-related representations on every layer. Extensive experiments show that our model achieves state-of-the-art performance on three benchmark datasets (5.5%, 2.2% and 4.5% F1 gains on EventStoryLine, MAVEN-ERE and Causal-TimeBank). Code is available at https://github.com/HITsz-TMG/PPAT. Baotian Hu, Zhenran Xu, Min Zhang 0005 |
IJCAI | 2 |
| 2023 | Enhancing Multi-modal Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-GenerationabstractMulti-modal multi-hop question answering involves answering a question by reasoning over multiple input sources from different modalities. Existing methods often retrieve evidences separately and then use a language model to generate an answer based on the retrieved evidences, and thus do not adequately connect candidates and are unable to model the interdependent relations during retrieval. Moreover, the pipelined approaches of retrieval and generation might result in poor generation performance when retrieval performance is low. To address these issues, we propose a Structured Knowledge and Unified Retrieval-Generation (SKURG) approach. SKURG employs an Entity-centered Fusion Encoder to align sources from different modalities using shared entities. It then uses a unified Retrieval-Generation Decoder to integrate intermediate retrieval results for answer generation and also adaptively determine the number of retrieval steps. Extensive experiments on two representative multi-modal multi-hop QA datasets MultimodalQA and WebQA demonstrate that SKURG outperforms the state-of-the-art models in both source retrieval and answer generation performance with fewer parameters1. Qian Yang 0007, Qian Chen 0003, Wen Wang 0001, Baotian Hu, Min Zhang 0005 |
ACM Multimedia | 4 |
| 2023 | Overview of NLPCC 2023 Shared Task 6: Chinese Few-Shot and Zero-Shot Entity Linking
Zhenran Xu, Zifei Shan, Baotian Hu, Min Zhang 0005 |
NLPCC (3) | 3 |
| 2023 | Hansel: A Chinese Few-Shot and Zero-Shot Entity Linking BenchmarkabstractModern Entity Linking (EL) systems entrench a popularity bias, yet there is no dataset focusing on tail and emerging entities in languages other than English. We present Hansel, a new benchmark in Chinese that fills the vacancy of non-English few-shot and zero-shot EL challenges. The test set of Hansel is human annotated and reviewed, created with a novel method for collecting zero-shot EL datasets. It covers 10K diverse documents in news, social media posts and other web articles, with Wikidata as its target Knowledge Base. We demonstrate that the existing state-of-the-art EL system performs poorly on Hansel ([email protected] of 36.6% on Few-Shot). We then establish a strong baseline that scores a [email protected] of 46.2% on Few-Shot and 76.6% on Zero-Shot on our dataset. We also show that our baseline achieves competitive results on TAC-KBP2015 Chinese Entity Linking task. Datasets and codes are released at https://github.com/HITsz-TMG/Hansel. Zhenran Xu, Zifei Shan, Baotian Hu, Bing Qin 0001 |
WSDM | 4 |
| 2023 | Learning to generate complex question with intent prediction from long passage
Youcheng Pan, Baotian Hu, Shiyue Wang, Xiaolong Wang 0001, Qingcai Chen, Zenglin Xu, Min Zhang 0005 |
Appl. Intell. | 2 |
| 2023 | Fine-grained biomedical knowledge negation detection via contrastive learning
Tiantian Zhu 0002, Yang Xiang 0003, Qingcai Chen, Yang Qin 0001, Baotian Hu, Wentai Zhang 0003 |
Knowl. Based Syst. | 5 |
| 2023 | Fast and Robust Online Handwritten Chinese Character Recognition With Deep Spatial and Contextual Information Fusion NetworkabstractDeep convolutional neuralnetworks have achieved fairly high accuracy for single online handwritten Chinese character recognition (SOLHCCR). However, in real application scenarios, users always write multiple characters to form a complete sentence, and previous contextual information holds significant potential for improving the accuracy, robustness and efficiency of recognition. In this work, we first propose a simple and straightforward model named the vanilla compositional network (VCN) by coupling convolutional neural network with a sequence modeling architecture (i.e., a recurrent neural network or Transformer), which exploits the handwritten character’s previous contextual information. Although VCN performs much better than the previous state-of-the-art SOLHCCR models, it is a two-stage architecture in nature. It suffers from high fragility when confronting with poorly written characters such as sloppy writing, and missing or broken strokes, due to relying heavily on contextual information. To improve the robustness of the OLHCCR model, we further propose a novel deep spatial & contextual information fusion network (DSCIFN). It utilizes an autoregresssive framework pre-trained on a large-scale sentence corpora as the backbone component, and highly integrates the spatial features of handwritten characters and their previous contextual information in a multi-layer fusion module. To verify the effectiveness of models, we reorganize a new form of online Chinese handwritten character with its previous context dataset, named OHCCC. Extensive experimental results demonstrate that DSCIFN achieves state-of-the-art performance and has increased strong robustness compared to VCN and previous SOLHCCR models. The in-depth empirical analysis and case study indicate that DSCIFN can significantly improve the efficiency of handwriting input because it does not need complete strokes to recognize a handwritten Chinese character precisely. Yunxin Li, Qian Yang 0007, Qingcai Chen, Baotian Hu, Xiaolong Wang 0001, Lin Ma 0002 |
IEEE Trans. Multim. | 4 |
| 2022 | Unifying Model Explainability and Robustness for Joint Text Classification and Rationale ExtractionabstractRecent works have shown explainability and robustness are two crucial ingredients of trustworthy and reliable text classification. However, previous works usually address one of two aspects: i) how to extract accurate rationales for explainability while being beneficial to prediction; ii) how to make the predictive model robust to different types of adversarial attacks. Intuitively, a model that produces helpful explanations should be more robust against adversarial attacks, because we cannot trust the model that outputs explanations but changes its prediction under small perturbations. To this end, we propose a joint classification and rationale extraction model named AT-BMC. It includes two key mechanisms: mixed Adversarial Training (AT) is designed to use various perturbations in discrete and embedding space to improve the model’s robustness, and Boundary Match Constraint (BMC) helps to locate rationales more precisely with the guidance of boundary information. Performances on benchmark datasets demonstrate that the proposed AT-BMC outperforms baselines on both classification and rationale extraction by a large margin. Robustness analysis shows that the proposed AT-BMC decreases the attack success rate effectively by up to 69%. The results indicate that there are connections between robust models and better explanations. Dongfang Li 0002, Baotian Hu, Qingcai Chen, Tujie Xu, Jingcong Tao, Yunan Zhang 0003 |
AAAI | 2 |
| 2022 | Prompt-based Text Entailment for Low-Resource Named Entity RecognitionabstractPre-trained Language Models (PLMs) have been applied in NLP tasks and achieve promising results. Nevertheless, the fine-tuning procedure needs labeled data of the target domain, making it difficult to learn in low-resource and non-trivial labeled scenarios. To address these challenges, we propose Prompt-based Text Entailment (PTE) for low-resource named entity recognition, which better leverages knowledge in the PLMs. We first reformulate named entity recognition as the text entailment task. The original sentence with entity type-specific prompts is fed into PLMs to get entailment scores for each candidate. The entity type with the top score is then selected as final label. Then, we inject tagging labels into prompts and treat words as basic units instead of n-gram spans to reduce time complexity in generating candidates by n-grams enumeration. Experimental results demonstrate that the proposed method PTE achieves competitive performance on the CoNLL03 dataset, and better than fine-tuned counterparts on the MIT Movie and Few-NERD dataset in low-resource settings. Dongfang Li 0002, Baotian Hu, Qingcai Chen |
COLING | 2 |
| 2022 | Calibration Meets Explanation: A Simple and Effective Approach for Model Confidence EstimatesabstractCalibration strengthens the trustworthiness of black-box models by producing better accurate confidence estimates on given examples.However, little is known about if model explanations can help confidence calibration.Intuitively, humans look at important features attributions and decide whether the model is trustworthy.Similarly, the explanations can tell us when the model may or may not know.Inspired by this, we propose a method named CME that leverages model explanations to make the model less confident with non-inductive attributions.The idea is that when the model is not highly confident, it is difficult to identify strong indications of any class, and the tokens accordingly do not have high attribution scores for any class and vice versa.We conduct extensive experiments on six datasets with two popular pre-trained language models in the in-domain and out-of-domain settings.The results show that CME improves calibration performance in all settings.The expected calibration errors are further reduced when combined with temperature scaling.Our findings highlight that model explanations can help calibrate posterior estimates. Dongfang Li 0002, Baotian Hu, Qingcai Chen |
EMNLP | 2 |
| 2022 | An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP TasksabstractAccess to external knowledge is essential for many natural language processing tasks, such as question answering and dialogue.Existing methods often rely on a parametric model that stores knowledge in its parameters, or use a retrieval-augmented model that has access to an external knowledge source.Parametric and retrieval-augmented models have complementary strengths in terms of computational efficiency and predictive accuracy.To combine the strength of both approaches, we propose the Efficient Memory-Augmented Transformer (EMAT) -it encodes external knowledge into a key-value memory and exploits the fast maximum inner product search for memory querying.We also introduce pre-training tasks that allow EMAT to encode informative key-value representations, and to learn an implicit strategy to integrate multiple memory slots into the transformer.Experiments on various knowledge-intensive tasks such as question answering and dialogue datasets show that, simply augmenting parametric models (T5-base) using our method produces more accurate results (e.g., 25.8 → 44.3 EM on NQ) while retaining a high throughput (e.g., 1000 queries/s on NQ).Compared to retrievalaugmented models, EMAT runs substantially faster across the board and produces more accurate results on WoW and ELI5. 1 Yuxiang Wu, Yu Zhao 0043, Baotian Hu, Pasquale Minervini, Pontus Stenetorp, Sebastian Riedel 0001 |
EMNLP | 3 |
| 2022 | Multi-Role Event Argument Extraction as Machine Reading Comprehension with Argument Match OptimizationabstractExtracting arguments for the pre-defined roles is a crucial step for event extraction. Recently, there are some insightful works that view it as a machine reading comprehension problem and achieve significant progress. However, most of them need multi-turns to extract the arguments of each role independently, which ignores the relationships among roles in the same event. To alleviate this problem, we propose a novel Multi-Role Argument Extraction method named MRAE which can exploit the relationship of event roles by extracting all arguments for an event simultaneously. To force MRAE to locate more arguments accurately, we propose an argument match optimization loss based on the minimum risk training to exploit sentence-level F1 score. We conduct experiments on the widely used ACE2005 dataset. The experimental results demonstrate that MRAE outperforms the competitor methods by at least +1.2% F1 score on argument extraction, and also shows superiority on data scarce scenarios. Jingcong Tao, Youcheng Pan, Baotian Hu, Weihua Peng, Cuiyun Han, Xiaolong Wang 0001 |
ICASSP | 4 |
| 2022 | Enhancing Entity Representations with Prompt Learning for Biomedical Entity LinkingabstractBiomedical entity linking aims to map mentions in biomedical text to standardized concepts or entities in a curated knowledge base (KB) such as Unified Medical Language System (UMLS). The latest research tends to solve this problem in a unified framework solely based on surface form matching between mentions and entities. Specifically, these methods focus on addressing the variety challenge of the heterogeneous naming of biomedical concepts. Yet, the ambiguity challenge that the same word under different contexts may refer to distinct entities is usually ignored. To address this challenge, we propose a two-stage linking algorithm to enhance the entity representations based on prompt learning. The first stage includes a coarser-grained retrieval from a representation space defined by a bi-encoder that independently embeds the mention and entity’s surface forms. Unlike previous one-model-fits-all systems, each candidate is then re-ranked with a finer-grained encoder based on prompt-tuning that utilizes the contextual information. Extensive experiments show that our model achieves promising performance improvements compared with several state-of-the-art techniques on the largest biomedical public dataset MedMentions and the NCBI disease corpus. We also observe by cases that the proposed prompt-tuning strategy is effective in solving both the variety and ambiguity challenges in the linking task. Tiantian Zhu 0002, Yang Qin 0001, Qingcai Chen, Baotian Hu, Yang Xiang 0003 |
IJCAI | 4 |
| 2022 | Medical Dialogue Response Generation with Pivotal Information RecallingabstractMedical dialogue generation is an important yet challenging task. Most previous works rely on the attention mechanism and large-scale pretrained language models. However, these methods often fail to acquire pivotal information from the long dialogue history to yield an accurate and informative response, due to the fact that the medical entities usually scatters throughout multiple utterances along with the complex relationships between them. To mitigate this problem, we propose a medical response generation model with Pivotal Information Recalling (MedPIR), which is built on two components, i.e., knowledge-aware dialogue graph encoder and recall-enhanced generator. The knowledge-aware dialogue graph encoder constructs a dialogue graph by exploiting the knowledge relationships between entities in the utterances, and encodes it with a graph attention network. Then, the recall-enhanced generator strengthens the usage of these pivotal information by generating a summary of the dialogue before producing the actual response. Experimental results on two large-scale medical dialogue datasets show that MedPIR outperforms the strong baselines in BLEU scores and medical entities F1 measure. Yu Zhao 0043, Yunxin Li, Yuxiang Wu, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001, Min Zhang 0005 |
KDD | 4 |
| 2022 | Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language ExplanationsabstractVisual Entailment with natural language explanations aims to infer the relationship between a text-image pair and generate a sentence to explain the decision-making process. Previous methods rely mainly on a pre-trained vision-language model to perform the relation inference and a language model to generate the corresponding explanation. However, the pre-trained vision-language models mainly build token-level alignment between text and image yet ignore the high-level semantic alignment between the phrases (chunks) and visual contents, which is critical for vision-language reasoning. Moreover, the explanation generator based only on the encoded joint representation does not explicitly consider the critical decision-making points of relation inference. Thus the generated explanations are less faithful to visual-language reasoning. To mitigate these problems, we propose a unified Chunk-aware Alignment and Lexical Constraint based method, dubbed as CALeC. It contains a Chunk-aware Semantic Interactor (arr. CSI), a relation inferrer, and a Lexical Constraint-aware Generator (arr. LeCG). Specifically, CSI exploits the sentence structure inherent in language and various image regions to build chunk-aware semantic alignment. Relation inferrer uses an attention-based reasoning network to incorporate the token-level and chunk-level vision-language representations. LeCG utilizes lexical constraints to expressly incorporate the words or chunks focused by the relation inferrer into explanation generation, improving the faithfulness and informativeness of the explanations. We conduct extensive experiments on three datasets, and experimental results indicate that CALeC significantly outperforms other competitor models on inference accuracy and quality of generated explanations. Qian Yang 0007, Yunxin Li, Baotian Hu, Lin Ma 0002, Min Zhang 0005 |
ACM Multimedia | 3 |
| 2022 | Enhancing Entity Linking with Contextualized Entity Embeddings
Zhenran Xu, Senbao Shi, Baotian Hu |
NLPCC (2) | 4 |
| 2022 | Biomedical relation extraction via knowledge-enhanced reading comprehensionabstractBACKGROUND: In biomedical research, chemical and disease relation extraction from unstructured biomedical literature is an essential task. Effective context understanding and knowledge integration are two main research problems in this task. Most work of relation extraction focuses on classification for entity mention pairs. Inspired by the effectiveness of machine reading comprehension (RC) in the respect of context understanding, solving biomedical relation extraction with the RC framework at both intra-sentential and inter-sentential levels is a new topic worthy to be explored. Except for the unstructured biomedical text, many structured knowledge bases (KBs) provide valuable guidance for biomedical relation extraction. Utilizing knowledge in the RC framework is also worthy to be investigated. We propose a knowledge-enhanced reading comprehension (KRC) framework to leverage reading comprehension and prior knowledge for biomedical relation extraction. First, we generate questions for each relation, which reformulates the relation extraction task to a question answering task. Second, based on the RC framework, we integrate knowledge representation through an efficient knowledge-enhanced attention interaction mechanism to guide the biomedical relation extraction. RESULTS: The proposed model was evaluated on the BioCreative V CDR dataset and CHR dataset. Experiments show that our model achieved a competitive document-level F1 of 71.18% and 93.3%, respectively, compared with other methods. CONCLUSION: Result analysis reveals that open-domain reading comprehension data and knowledge representation can help improve biomedical relation extraction in our proposed KRC framework. Our work can encourage more research on bridging reading comprehension and biomedical relation extraction and promote the biomedical relation extraction. Baotian Hu, Weihua Peng, Qingcai Chen, Buzhou Tang |
BMC Bioinform. | 2 |
| 2021 | Multi-hop Graph Convolutional Network with High-order Chebyshev Approximation for Text ReasoningabstractShuoran Jiang, Qingcai Chen, Xin Liu, Baotian Hu, Lisai Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Shuoran Jiang, Qingcai Chen, Xin Liu 0054, Baotian Hu, Lisai Zhang |
ACL/IJCNLP (1) | 4 |
| 2021 | A Large-Scale Chinese Long-Text Extractive Summarization CorpusabstractRecently, large-scale datasets have vastly facilitated the development in nearly domains of Natural Language Processing. However, lacking large scale Chinese corpus is still a critical bottleneck for further research on deep text summarization methods. In this paper, we publish a large-scale Chinese Long-text Extractive Summarization corpus named CLES. The CLES contains about 104Kpairs, which is originally collected from Sina Weibo1. To verify the quality of the corpus, we also manually tagged the relevance score of 5,000pairs. Our benchmark models on the proposed corpus include conventional deep learning based extractive models and several pre-trained Bert-based algorithms. Their performances are reported and briefly analyzed to facilitate further research on the corpus. We will release this corpus for further research2. Kai Chen 0020, Guanyu Fu, Qingcai Chen, Baotian Hu |
ICASSP | 4 |
| 2021 | Enriching BERT With Knowledge Graph Embedding For Industry Classification
Shiyue Wang, Youcheng Pan, Zhenran Xu, Baotian Hu, Xiaolong Wang 0001 |
ICONIP (6) | 4 |
| 2021 | MSDF: A General Open-Domain Multi-skill Dialog Framework
Yu Zhao 0043, Xinshuo Hu, Yunxin Li, Baotian Hu, Dongfang Li 0002, Sichao Chen, Xiaolong Wang 0001 |
NLPCC (2) | 4 |
| 2021 | Distantly supervised biomedical relation extraction using piecewise attentive convolutional neural network and reinforcement learningabstractOBJECTIVE: There have been various methods to deal with the erroneous training data in distantly supervised relation extraction (RE), however, their performance is still far from satisfaction. We aimed to deal with the insufficient modeling problem on instance-label correlations for predicting biomedical relations using deep learning and reinforcement learning. MATERIALS AND METHODS: In this study, a new computational model called piecewise attentive convolutional neural network and reinforcement learning (PACNN+RL) was proposed to perform RE on distantly supervised data generated from Unified Medical Language System with MEDLINE abstracts and benchmark datasets. In PACNN+RL, PACNN was introduced to encode semantic information of biomedical text, and the RL method with memory backtracking mechanism was leveraged to alleviate the erroneous data issue. Extensive experiments were conducted on 4 biomedical RE tasks. RESULTS: The proposed PACNN+RL model achieved competitive performance on 8 biomedical corpora, outperforming most baseline systems. Specifically, PACNN+RL outperformed all baseline methods with the F1-score of 0.5592 on the may-prevent dataset, 0.6666 on the may-treat dataset, and 0.3838 on the DDI corpus, 2011. For the protein-protein interaction RE task, we obtained new state-of-the-art performance on 4 out of 5 benchmark datasets. CONCLUSIONS: The performance on many distantly supervised biomedical RE tasks was substantially improved, primarily owing to the denoising effect of the proposed model. It is anticipated that PACNN+RL will become a useful tool for large-scale RE and other downstream tasks to facilitate biomedical knowledge acquisition. We also made the demonstration program and source code publicly available at http://112.74.48.115:9000/. Tiantian Zhu 0002, Yang Qin 0001, Yang Xiang 0003, Baotian Hu, Qingcai Chen, Weihua Peng |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Neural data-to-text generation with dynamic content planning
Kai Chen 0020, Fayuan Li, Baotian Hu, Weihua Peng, Qingcai Chen, Hong Yu 0001, Yang Xiang 0003 |
Knowl. Based Syst. | 3 |
| 2021 | Attentive capsule network for click-through rate and conversion rate prediction in online advertising
Dongfang Li 0002, Baotian Hu, Qingcai Chen, Quanchang Qi, Liubin Wang, Haishan Liu |
Knowl. Based Syst. | 2 |
| 2021 | Decomposing word embedding with the capsule network
Xin Liu 0054, Qingcai Chen, Yan Liu 0004, Joanna Siebert, Baotian Hu, Xiangping Wu 0001, Buzhou Tang |
Knowl. Based Syst. | 5 |
| 2021 | LCSegNet: An Efficient Semantic Segmentation Network for Large-Scale Complex Chinese Character RecognitionabstractComplex scene character recognition is a challenging yet important task in machine learning, especially for languages with large character sets, such as Chinese, which is composed of hieroglyphics with large-scale categories and similar glyphs. Recently, state-of-the-art methods based on semantic segmentation have achieved great success in scene parsing and have been applied in scene text recognition. However, because of limitations in terms of memory and computation, they are only applied in the small category recognition tasks, such as tasks involving English alphabets and digits. In this paper, we propose an efficient semantic segmentation model based on label coding (LC), called LCSegNet, to recognize large-scale Chinese characters. First, to reduce the number of labels, we design a new label coding method based on the Wubi Chinese characters code, called Wubi-CRF. In this method, glyphs and structure information of Chinese characters are encoded into 140-bit labels. Second, we employ an efficient semantic segmentation model for pixel-wise prediction and utilize a conditional random field (CRF) module to learn the constraint rules of Wubi-like coding. Finally, experiments are conducted on three benchmarks: a large Chinese text dataset in the wild (CTW), ICDAR2019-ReCTS, and HIT-OR3C dataset. Results show that the proposed method achieves state-of-the-art performances in both complex scene and handwritten character recognition tasks. Xiangping Wu 0001, Qingcai Chen, Yulun Xiao, Xin Liu 0054, Baotian Hu |
IEEE Trans. Multim. | 6 |
| 2020 | MedWriter: Knowledge-Aware Medical Text GenerationabstractTo exploit the domain knowledge to guarantee the correctness of generated text has been a hot topic in recent years, especially for high professional domains such as medical.However, most of recent works only consider the information of unstructured text rather than structured information of the knowledge graph.In this paper, we focus on the medical topic-to-text generation task and adapt a knowledge-aware text generation model to the medical domain, named MedWriter, which not only introduces the specific knowledge from the external MKG but also is capable of learning graph-level representation.We conduct experiments on a medical literature dataset collected from medical journals, each of which has a set of topic words, an abstract of medical literature and a corresponding knowledge graph from CMeKG.Experimental results demonstrate incorporating knowledge graph into generation model can improve the quality of the generated text and has robust superiority over the competitor methods. Youcheng Pan, Qingcai Chen, Weihua Peng, Xiaolong Wang 0001, Baotian Hu, Xin Liu 0054, Wenxiu Zhou |
COLING | 5 |
| 2020 | Towards Medical Machine Reading Comprehension with Structural Knowledge and Plain TextabstractMachine reading comprehension (MRC) has achieved significant progress on the open domain in recent years, mainly due to large-scale pre-trained language models.However, it performs much worse in specific domains such as the medical field due to the lack of extensive training data and professional structural knowledge neglect.As an effort, we first collect a large scale medical multi-choice question dataset (more than 21k instances) for the National Licensed Pharmacist Examination in China.It is a challenging medical examination with a passing rate of less than 14.2% in 2018.Then we propose a novel reading comprehension model KMQA, which can fully exploit the structural medical knowledge (i.e., medical knowledge graph) and the reference medical plain text (i.e., text snippets retrieved from reference books).The experimental results indicate that the KMQA outperforms existing competitive models with a large margin and passes the exam with 61.8% accuracy rate on the test set. Dongfang Li 0002, Baotian Hu, Qingcai Chen, Weihua Peng |
EMNLP (1) | 2 |
| 2020 | Learning to Generate Diverse Questions from KeywordsabstractDiverse text generation has been emerging as an important topic of natural language generation. Traditional studies on question generation mainly investigate how to generate one question based on a given input (one-to-one). In this paper, we focus on a more complex question generation task, i.e., generating a series of questions for each set of keywords (one-to-many). As an effort towards this, we propose a novel neural generative model, which incorporates context information and control signal to produce multiple diverse questions from a given fixed set of keywords. The control signal is designed to increase the diversity of questions by capturing the diverse patterns from the entire dataset. The context information is used to guarantee the generated questions are highly related to the given keywords. To evaluate the effectiveness of the proposed model, we collect a dataset which contains 62835 questions with respect to 12567 sets of keywords.1To the best of our knowledge, it's the first Chinese financial dataset for diverse question generation. The experimental results show that our model outperforms the competitor methods in terms of BLEU and Distinct. The qualitative evaluation indicates that our model is able to generate diverse and meaningful questions. Youcheng Pan, Baotian Hu, Qingcai Chen, Yang Xiang 0003, Xiaolong Wang 0001 |
ICASSP | 2 |
| 2020 | AdaHGNN: Adaptive Hypergraph Neural Networks for Multi-Label Image ClassificationabstractMulti-label image classification is an important and challenging task in computer vision and multimedia fields. Most of the recent works only capture the pair-wise dependencies among multiple labels through statistical co-occurrence information, which cannot model the high-order semantic relations automatically. In this paper, we propose a high-order semantic learning model based on adaptive hypergraph neural networks (AdaHGNN) to boost multi-label classification performance. Firstly, an adaptive hypergraph is constructed by using label embeddings automatically. Secondly, image features are decoupled into feature vectors corresponding to each label, and hypergraph neural networks (HGNN) are employed to correlate these vectors and explore the high-order semantic interactions. In addition, multi-scale learning is used to reduce sensitivity to object size inconsistencies. Experiments are conducted on four benchmarks: MS-COCO, NUS-WIDE, Visual Genome, and Pascal VOC 2007, which cover large, medium, and small-scale categories. State-of-the-art performances are achieved on three of them. Results and analysis demonstrate that the proposed method has the ability to capture high-order semantic dependencies. Xiangping Wu 0001, Qingcai Chen, Yulun Xiao, Baotian Hu |
ACM Multimedia | 5 |
| 2020 | Text-Guided Neural Image InpaintingabstractImage inpainting task requires filling the corrupted image with contents coherent with the context. This research field has achieved promising progress by using neural image inpainting methods. Nevertheless, there is still a critical challenge in guessing the missed content with only the context pixels. The goal of this paper is to fill the semantic information in corrupted images according to the provided descriptive text. Unique from existing text-guided image generation works, the inpainting models are required to compare the semantic content of the given text and the remaining part of the image, then find out the semantic content that should be filled for missing part. To fulfill such a task, we propose a novel inpainting model named Text-Guided Dual Attention Inpainting Network (TDANet). Firstly, a dual multimodal attention mechanism is designed to extract the explicit semantic information about the corrupted regions, which is done by comparing the descriptive text and complementary image areas through reciprocal attention. Secondly, an image-text matching loss is applied to maximize the semantic similarity of the generated image and the text. Experiments are conducted on two open datasets. Results show that the proposed TDANet model reaches new state-of-the-art on both quantitative and qualitative measures. Result analysis suggests that the generated images are consistent with the guidance text, enabling the generation of various results by providing different descriptions. Codes are available at https://github.com/idealwhite/TDANet Lisai Zhang, Qingcai Chen, Baotian Hu, Shuoran Jiang |
ACM Multimedia | 3 |
| 2020 | Stroke Sequence-Dependent Deep Convolutional Neural Network for Online Handwritten Chinese Character RecognitionabstractWe propose a novel model, called stroke sequence-dependent deep convolutional neural network (SSDCNN), which uses the stroke sequence information and eight-directional features of Chinese characters for online handwritten Chinese character recognition (OLHCCR). SSDCNN learns the representation of OLHCCs by incorporating the natural sequence information of the strokes. Furthermore, it naturally incorporates the eight-directional features. First, SSDCNN inputs the stroke sequence and transforms it into stacks of feature maps following the writing order of the strokes. Second, the fixed-length, stroke sequence-dependent representations of OLHCC are derived through convolutional, residual, and max-pooling operations. Third, the stroke sequence-dependent representation is combined with the eight-directional features via a number of fully connected neural network layers. Finally, the Chinese characters are recognized using a softmax classifier. The SSDCNN is trained in two stages: 1) the whole architecture is pretrained using the training data until the performance converges to an acceptable degree. 2) The stroke sequence-dependent representation is combined with the eight-directional features by a fully connected neural network and a softmax layer for further training. The model was experimentally evaluated on the OLHCCR competition tasks of International Conference on Document Analysis and Recognition (ICDAR) 2013. The recognition error was a maximum 58.28% lower in SSDCNN than in a model using the eight-directional features alone (5.13% versus 2.14%). Owing to its high accuracy (97.86%), the proposed SSDCNN reduced the recognition error by approximately 18.0% as compared with that of the winning system in the ICDAR 2013 competition. SSDCNN integrated with an adaptation mechanism, called the SSDCNN+Adapt model, and reached a new state-of-the-art (SOTA) standard with an accuracy of 97.94%. The SSDCNN exploits the stroke sequence information to learn high-quality OLHCC representations. Moreover, the learned representation and the classical eight-directional features complement each other within the SSDCNN architecture. Xin Liu 0054, Baotian Hu, Qingcai Chen, Xiangping Wu 0001, Jinghan You |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Learning to Extract Coherent Summary via Deep Reinforcement LearningabstractCoherence plays a critical role in producing a high-quality summary from a document. In recent years, neural extractive summarization is becoming increasingly attractive. However, most of them ignore the coherence of summaries when extracting sentences. As an effort towards extracting coherent summaries, we propose a neural coherence model to capture the cross-sentence semantic and syntactic coherence patterns. The proposed neural coherence model obviates the need for feature engineering and can be trained in an end-to-end fashion using unlabeled data. Empirical results show that the proposed neural coherence model can efficiently capture the cross-sentence coherence patterns. Using the combined output of the neural coherence model and ROUGE package as the reward, we design a reinforcement learning method to train a proposed neural extractive summarizer which is named Reinforced Neural Extractive Summarization (RNES) model. The RNES model learns to optimize coherence and informative importance of the summary simultaneously. The experimental results show that the proposed RNES outperforms existing baselines and achieves state-of-the-art performance in term of ROUGE on CNN/Daily Mail dataset. The qualitative evaluation indicates that summaries produced by RNES are more coherent and readable. Yuxiang Wu, Baotian Hu |
AAAI | 2 |
| 2018 | Recurrent convolutional neural network for answer selection in community question answering
Xiaoqiang Zhou, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001 |
Neurocomputing | 2 |
| 2016 | A novel word embedding learning model using the dissociation between nouns and verbs
Baotian Hu, Buzhou Tang, Qingcai Chen, Longbiao Kang |
Neurocomputing | 1 |
| 2015 | LCSTS: A Large Scale Chinese Short Text Summarization DatasetabstractAutomatic text summarization is widely regarded as the highly difficult problem, partially because of the lack of large text summarization data set.Due to the great challenge of constructing the large scale summaries for full text, in this paper, we introduce a large corpus of Chinese short text summarization dataset constructed from the Chinese microblogging website Sina Weibo, which is released to the public 1 .This corpus consists of over 2 million real Chinese short texts with short summaries given by the author of each text.We also manually tagged the relevance of 10,666 short summaries with their corresponding short texts.Based on the corpus, we introduce recurrent neural network for the summary generation and achieve promising results, which not only shows the usefulness of the proposed corpus for short text summarization research, but also provides a baseline for further research on this topic. Baotian Hu, Qingcai Chen, Fangze Zhu |
EMNLP | 1 |
| 2015 | An Auto-Encoder for Learning Conversation Representation Using LSTM
Xiaoqiang Zhou, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001 |
ICONIP (1) | 2 |
| 2014 | Convolutional Neural Network Architectures for Matching Natural Language Sentences
Baotian Hu, Zhengdong Lu, Hang Li 0001, Qingcai Chen |
NIPS | 1 |
| 2014 | A Short Texts Matching Method Using Shallow Features and Deep Features
Longbiao Kang, Baotian Hu, Xiangping Wu 0001, Qingcai Chen |
NLPCC | 2 |