EDBT 2026 Demo / reviewers in the wild / expert
Man Lan
dblp:01/800
· DBLP profile ↗
82ranked-venue papers
11as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 9 first-author · 34 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 11 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMsabstractRecent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet, a significant gap persists in their adaptation to real-world multimodal scenarios, most notably, vision-language tasks, due to a heavy focus on single-modal language settings. While efforts to transplant reinforcement learning techniques from NLP to Visual Language Models (VLMs) have emerged, these approaches often remain confined to perception-centric tasks or reduce images to textual summaries, failing to fully exploit visual context and commonsense knowledge, ultimately constraining the generalization of reasoning capabilities across diverse multimodal environments. To address this limitation, we introduce a novel fine-tuning task, Masked Prediction via Context and Commonsense (MPCC), which forces models to integrate visual context and commonsense reasoning by reconstructing semantically meaningful content from occluded images, thereby laying the foundation for generalized reasoning. To systematically evaluate the model’s performance in generalized reasoning, we developed a specialized evaluation benchmark, MPCC-Eval, and employed various fine-tuning strategies to guide reasoning. Among these, we introduced an innovative training method, Reinforcement Fine-Tuning with Prior Sampling, which not only enhances model performance but also improves its generalized reasoning capabilities in out-of-distribution (OOD) and cross-task scenarios. Jiaao Yu 0001, Shenwei Li, Mingjie Han, Yifei Yin, Wenzheng Song, Chenghao Jia, Man Lan |
AAAI | 7 |
| 2026 | CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware RewardsabstractLarge Language Model (LLM) based Chinese Grammatical Error Correction (CGEC) systems face two critical challenges: generalpurpose models lack specialized linguistic priors for subtle grammatical distinctions, and Supervised Fine-Tuning (SFT) with Maximum Likelihood Estimation fails to optimize for precision-focused metrics, leading to systematic over-correction.We propose CSRP, a three-stage framework that progressively builds correction capability through Continual Pretraining (CPT) on 5.9M balanced samples to internalize domain knowledge, Chain-of-Thought SFT with explicit error reasoning for diagnostic transparency, and Group Relative Policy Optimization with a novel Efficiency-Aware Reward that explicitly penalizes unnecessary edits.On the NACGEC benchmark, CSRP achieves state-of-the-art performance with 50.99 F 0.5 and 57.17 precision, substantially outperforming previous best results while effectively mitigating the over-correction bias inherent in MLE-trained models.Our method also advances CSCD spelling correction to 59.61 F1, surpassing GPT-4 by 5.20 points.Comprehensive ablation studies demonstrate that the RL alignment stage contributes a 8% relative gain over the SFT baseline, and that this gain is orthogonal to the contribution of large-scale CPT, validating that explicit optimization for edit efficiency is essential for highquality grammatical error correction.Our code is available at https://github.com/TW-NLP/ ChineseErrorCorrector. Man Lan |
ACL (1) | 3 |
| 2026 | Generative Gamer: Learning Equilibrium Strategy by LLM-driven Dynamic DeductionabstractLarge Language Models (LLMs) have demonstrated remarkable general capabilities, yet they falter in domains requiring deep strategic reasoning.A primary obstacle is the need to navigate a game tree that grows exponentially with search depth, a task for which their generative nature is ill-suited.To address this, we introduce Generative Gamer (GenGamer), a framework that trains LLMs to reason like an expert player.Instead of attempting an exhaustive search, GenGamer learns to generate a compact, pruned reasoning trajectory termed as a Dynamic Deduction.This is achieved by integrating three key strategies: action pruning based on policy confidence, state pruning via value estimation, and branch pruning inspired by alpha-beta principles.Furthermore, to train the model effectively, we propose the Deduction Tree Reward (DTR), a process-oriented mechanism that provides step-by-step feedback on the quality of the reasoning process, rather than relying solely on the final game outcome.Experiments on complex games such as Tic-Tac-Toe and Leduc Poker demonstrate that GenGamer significantly enhances the strategic capabilities of LLMs, enabling them to achieve performance that surpasses current state-of-theart language models. Xinshu Shen, Yupei Ren, Shangqing Zhao, Man Lan |
ACL (1) | 5 |
| 2025 | ReactGPT: Understanding of Chemical Reactions via In-Context TuningabstractThe interdisciplinary field of chemistry and artificial intelligence (AI) is an active area of research aimed at accelerating scientific discovery. Large language Models (LLMs) have shown significant promise in biochemical tasks, especially the molecule caption translation, which aims to align between molecules and natural language texts. However, existing works mainly focus on single molecules, while alignment between chemical reactions and natural language text remains largely unexplored. Additionally, the description of reactions is an essential part in biochemical patents and literature, and research on this aspect not only can help better understand chemical reactions but also promote research on automating chemical synthesis and retrosynthesis. In this work, we propose \textbf{ReactGPT}, a framework aiming to bridge the gap between chemical reaction and text. ReactGPT allows a new task: reaction captioning, by adapting LLMs to learn reaction-text alignment from context examples via In-Context Tuning. Specifically, ReactGPT jointly leverages a Fingerprints-based Reaction Retrieval module, a Domain-Specific Prompt Design module, and a two-stage In-Context Tuning module. We evaluate the effectiveness of ReactGPT on reaction captioning and experimental procedure prediction, both of these tasks can reflect the understanding of chemical reactions. Experimental results show that compared to previous models, ReactGPT exhibits competitive capabilities in resolving chemical reactions and generating high-quality text with correct structure. Zhe Fang, Wenhao Tian, Zhaoguang Long, Changzhi Sun, Yuefeng Chen, Man Lan |
AAAI | 9 |
| 2025 | Towards Comprehensive Argument Analysis in Education: Dataset, Tasks, and MethodabstractArgument mining has garnered increasing attention over the years, with the recent advancement of Large Language Models (LLMs) further propelling this trend. However, current argument relations remain relatively simplistic and foundational, struggling to capture the full scope of argument information. To address this limitation, we propose a systematic framework comprising 14 fine-grained relation types from the perspectives of vertical argument relations and horizontal discourse relations, thereby capturing the intricate interplay between argument components for a thorough understanding of argument structure. On this basis, we conducted extensive experiments on three tasks: argument component prediction, relation prediction, and automated essay grading. Additionally, we explored the impact of writing quality on argument component prediction and relation prediction, as well as the connections between discourse relations and argumentative features. The findings highlight the importance of fine-grained argumentative annotations for argumentative writing assessment and encourage multi-dimensional argument analysis. Yupei Ren, Shangqing Zhao, Man Lan, Xiaopeng Bai |
ACL (1) | 5 |
| 2025 | FinDABench: Benchmarking Financial Data Analysis Ability of Large Language ModelsabstractLarge Language Models (LLMs) have demonstrated impressive capabilities across a wide range of tasks. However, their proficiency and reliability in the specialized domain of financial data analysis, particularly focusing on data-driven thinking, remain uncertain. To bridge this gap, we introduce FinDABench, a comprehensive benchmark designed to evaluate the financial data analysis capabilities of LLMs within this context. The benchmark comprises 15,200 training instances and 8,900 test instances, all meticulously crafted by human experts. FinDABench assesses LLMs across three dimensions: 1) Core Ability, evaluating the models’ ability to perform financial indicator calculation and corporate sentiment risk assessment; 2) Analytical Ability, determining the models’ ability to quickly comprehend textual information and analyze abnormal financial reports; and 3) Technical Ability, examining the models’ use of technical knowledge to address real-world data analysis challenges involving analysis generation and charts visualization from multiple perspectives. We will release FinDABench, and the evaluation scripts at https://github.com/xxx. FinDABench aims to provide a measure for in-depth analysis of LLM abilities and foster the advancement of LLMs in the field of financial data analysis. Shangqing Zhao, Chenghao Jia, Xinlin Zhuang, Zhaoguang Long, Aimin Zhou, Man Lan, Yang Chong |
COLING | 8 |
| 2025 | Semantic Attention and LLM-based Layout Guidance for Text-to-Image GenerationabstractDiffusion models have substantially advanced text-to-image generation, achieving remarkable performance in creating high-quality images from textual prompts. However, they often struggle with accurately generating images representing spatial locations described or implied in the prompts. To address this, we introduce SALT, a training-free method leveraging semantic attention and layout guidance from Large Language Models (LLMs) for text-to-image generation. This method effectively guides both cross-attention and self-attention layers within diffusion models, steering generation toward the direction of high-attention values provided by the layout guidance. During the denoising process of the diffusion model, image features in the latent space are iteratively refined based on the loss function calculated from the desired attention maps. Our approach has been executed on two benchmarks, providing detailed qualitative examples and comprehensive quantitative analyses. Results demonstrate that SALT outperforms existing training-free methods in controlling object layouts and generating attributes.1 Yuxiang Song, Zhaoguang Long, Man Lan, Changzhi Sun, Aimin Zhou, Yuefeng Chen |
ICASSP | 3 |
| 2025 | K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic ReasoningabstractYadong Zhang, Shaoguang Mao, Tao Ge, Xun Wang, Yan Xia, Man Lan, Furu Wei. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shaoguang Mao, Tao Ge 0001, Xun Wang 0012, Yan Xia 0005, Man Lan, Furu Wei |
NAACL (Long Papers) | 6 |
| 2025 | Protein Design with Dynamic Protein VocabularyabstractProtein design is a fundamental challenge in biotechnology, aiming to design novel sequences with specific functions within the vast space of possible proteins. Recent advances in deep generative models have enabled function-based protein design from textual descriptions, yet struggle with structural plausibility. Inspired by classical protein design methods that leverage natural protein structures, we explore whether incorporating fragments from natural proteins can enhance foldability in generative models. Our empirical results show that even random incorporation of fragments improves foldability. Building on this insight, we introduce ProDVa, a novel protein design approach that integrates a text encoder for functional descriptions, a protein language model for designing proteins, and a fragment encoder to dynamically retrieve protein fragments based on textual functional descriptions. Experimental results demonstrate that our approach effectively designs protein sequences that are both functionally aligned and structurally plausible. Compared to state-of-the-art models, ProDVa achieves comparable function alignment using less than 0.04% of the training data, while designing significantly more well-folded proteins, with the proportion of proteins having pLDDT above 70 increasing by 7.38% and those with PAE below 10 increasing by 9.62%. Nuowei Liu, Jiahao Kuang, Changzhi Sun, Man Lan, Yuanbin Wu |
NeurIPS | 6 |
| 2025 | Overview of the NLPCC 2025 Shared Task2: Evaluation of Essay On-Topic Graded Comments(EOTGC)
Haoxiang Dong, Xiayu Sun, Man Lan, Xiaopeng Bai, Lixin Ye |
NLPCC (4) | 3 |
| 2025 | Overview of the NLPCC 2025 Shared Task 3: Comprehensive Argument Analysis for Chinese Argumentative Essay
Zheqin Yin, Yupei Ren, Man Lan, Yuanbin Wu, Aimin Zhou, Xiaopeng Bai |
NLPCC (4) | 3 |
| 2025 | Dongba Machine Translation with Transfer Learning: Leveraging Pre-trained Ancient Chinese ModelsabstractThe Dongba script, a logographic writing system used by the Naxi people in religious activities, faces challenges in translation due to the advanced age of Dongba script experts and the time-consuming nature of manual deciphering. This study focuses on translating the resource-scarce Dongba script into Modern Chinese using a novel approach based on cross-lingual transfer learning from Ancient Chinese. By examining translation patterns from Ancient Chinese to Modern Chinese, we determine the feasibility of transferring knowledge from Ancient Chinese to Dongba script translation. We propose the Dongba Machine Translation Model (DMTM), a pre-trained, low-resource machine translation model that utilizes the linguistic similarities between Ancient Chinese and Dongba script to improve translation quality. The model undergoes pre-training on a large-scale Ancient Chinese corpus and fine-tuning on a small-scale Dongba script corpus, enabling effective knowledge transfer. To address the scarcity of Dongba script translation resources, we present DongBa Corpus 1.0, a fine-grained parallel dataset of Dongba script and Modern Chinese. Experimental results demonstrate that our proposed DMTM achieves a translation score of 50.01% BLEU on the test set. As no prior methods exist for Dongba script translation, we compared various architectures commonly used in low-resource translation tasks, and DMTM exhibited the best performance with a 5.39% improvement over alternative architectures tested. The implementation codes and dataset for our approach are available at https://github.com/Chloe-mxxxxc/DMTM . Xinchen Ma, Man Lan, Wenbo Hu 0008, Yue Lu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | From Coarse to Fine: A Distillation Method for Fine-Grained Emotion-Causal Span Pair Extraction in ConversationabstractWe study the problem of extracting emotions and the causes behind these emotions in conversations. Existing methods either tackle them separately or jointly model them at the coarse-grained level of emotions (fewer emotion categories) and causes (utterance-level causes). In this work, we aim to jointly extract more fine-grained emotions and causes. We construct a fine-grained dataset FG-RECCON, includes 16 fine-grained emotion categories and span-level causes. To further improve the fine-grained extraction performance, we propose to utilize the casual discourse knowledge in a knowledge distillation way. Specifically, the teacher model learns to predict causal connective words between utterances, and then guides the student model in identifying both the fine-grained emotion labels and causal spans. Experimental results demonstrate that our distillation method achieves state-of-the-art performance on both RECCON and FG-RECCON dataset. Xinhao Chen, Changzhi Sun, Man Lan, Aimin Zhou |
AAAI | 4 |
| 2024 | A Lightweight and Effective Multi-View Knowledge Distillation Framework for Text-Image RetrievalabstractLarge-scale dual-stream Vision-Language Pre-training (VLP) models provide an efficient solution for text-image retrieval tasks. Despite this, their performance often falls short of the most current single-stream models, primarily due to limited fine-grained text-image interactions. Recent trends indicate a union of these two types of networks. Some methods adopt a retrieve and rerank strategy, their performance improvements largely hinge on the single-stream encoder during inference. Other approaches utilize knowledge distillation to strengthen either the single-stream encoder or the dual-stream encoder, surpassing their previous capabilities. However, existing distillation techniques typically focus on a single knowledge type, neglecting the richer insights available in the teacher model. To bridge this gap, we introduce a Lightweight and Effective Multi-View Knowledge Distillation approach, named LEMKD, for text-image retrieval. This method effectively utilizes response-based, feature-based and relation-based knowledge, transferring the knowledge from the single-stream encoder to the dual-stream encoder. Our approach is executed on the widely used MS-COCO and Flickr30K datasets. Results demonstrate that LEMKD not only matches the exceptional performance of the most advanced single-stream models but also excels in dual-stream encoder performance amidst the recent integration of single-stream and dual-stream models. Yuxiang Song, Yuxuan Zheng, Shangqing Zhao, Xinlin Zhuang, Zhaoguang Long, Changzhi Sun, Aimin Zhou, Man Lan |
IJCNN | 9 |
| 2024 | Overview of the NLPCC 2024 Shared Task 5: Argument Mining for Chinese Argumentative Essay
Zheqin Yin, Yupei Ren, Man Lan, Yuanbin Wu, Aimin Zhou, Xiaopeng Bai |
NLPCC (5) | 3 |
| 2024 | Overview of the NLPCC 2024 Shared Task: Chinese Essay Discourse Logic Evaluation and Integration
Hongyi Wu, Xinshu Shen, Man Lan, Yuanbin Wu, Xiaopeng Bai, Shaoguang Mao, Tao Ge 0001, Yan Xia 0005 |
NLPCC (5) | 4 |
| 2024 | Bread: A Hybrid Approach for Instruction Data Mining Through Balanced Retrieval and Dynamic Data Sampling
Xinlin Zhuang, Xin Mao 0002, Hongyi Wu, Shangqing Zhao, Yuxiang Song, Chenghao Jia, Man Lan |
NLPCC (2) | 12 |
| 2024 | Self-supervised BGP-graph reasoning enhanced complex KBQA via SPARQL generation
Yan Yang 0008, Peng Gao 0005, Shangqing Zhao, Yuefeng Chen, Man Lan, Aimin Zhou, Liang He 0001 |
Inf. Process. Manag. | 8 |
| 2023 | Connective Prediction for Implicit Discourse Relation Recognition via Knowledge DistillationabstractImplicit discourse relation recognition (IDRR) remains a challenging task in discourse analysis due to the absence of connectives.Most existing methods utilize one-hot labels as the sole optimization target, ignoring the internal association among connectives.Besides, these approaches spend lots of effort on template construction, negatively affecting the generalization capability.To address these problems, we propose a novel Connective Prediction via Knowledge Distillation (CP-KD) approach to instruct large-scale pre-trained language models (PLMs) mining the latent correlations between connectives and discourse relations, which is meaningful for IDRR.Experimental results on the PDTB 2.0/3.0 and CoNLL 2016 datasets show that our method significantly outperforms the state-of-the-art models on coarse-grained and fine-grained discourse relations.Moreover, our approach can be transferred to explicit discourse relation recognition (EDRR) and achieve acceptable performance.Our code is released in https://github.com/cubenlp/CP_KD-for-IDRR. Hongyi Wu, Man Lan, Yuanbin Wu |
ACL (1) | 3 |
| 2023 | A Multi-Task Dataset for Assessing Discourse Coherence in Chinese Essays: Structure, Theme, and Logic AnalysisabstractThis paper introduces the Chinese Essay Discourse Coherence Corpus (CEDCC), a multi-task dataset for assessing discourse coherence.Existing research tends to focus on isolated dimensions of discourse coherence, a gap which the CEDCC addresses by integrating coherence grading, topical continuity, and discourse relations.This approach, alongside detailed annotations, captures the subtleties of real-world texts and stimulates progress in Chinese discourse coherence analysis.Our contributions include the development of the CEDCC, the establishment of baselines for further research, and the demonstration of the impact of coherence on discourse relation recognition and automated essay scoring.The dataset and related codes is available at https: //github.com/cubenlp/CEDCC_corpus. Hongyi Wu, Xinshu Shen, Man Lan, Shaoguang Mao, Xiaopeng Bai, Yuanbin Wu |
EMNLP | 3 |
| 2023 | An Effective and Efficient Time-aware Entity Alignment Framework via Two-aspect Three-view Label PropagationabstractEntity alignment (EA) aims to find the equivalent entity pairs between different knowledge graphs (KGs), which is crucial to promote knowledge fusion. With the wide use of temporal knowledge graphs (TKGs), time-aware EA (TEA) methods appear to enhance EA. Existing TEA models are based on Graph Neural Networks (GNN) and achieve state-of-the-art (SOTA) performance, but it is difficult to transfer them to large-scale TKGs due to the scalability issue of GNN. In this paper, we propose an effective and efficient non-neural EA framework between TKGs, namely LightTEA, which consists of four essential components: (1) Two-aspect Three-view Label Propagation, (2) Sparse Similarity with Temporal Constraints, (3) Sinkhorn Operator, and (4) Temporal Iterative Learning. All of these modules work together to improve the performance of EA while reducing the time consumption of the model. Extensive experiments on public datasets indicate that our proposed model significantly outperforms the SOTA methods for EA between TKGs, and the time consumed by LightTEA is only dozens of seconds at most, no more than 10% of the most efficient TEA method. Xin Mao 0002, Youshao Xiao, Changxu Wu, Man Lan |
IJCAI | 5 |
| 2023 | Document-level Relation Extraction with Entity Interaction and Commonsense KnowledgeabstractDocument-Level Relation Extraction(DLRE) is a more challenging task than sentence-level relation extraction because of the characteristics such as more extended context, more interactions between entities, and the need for common sense to help the relation inference. In this paper, we propose an effective model to address the problems of complex entity interactions and the lack of commonsense knowledge. Specifically, we propose a Transformer-based entity interaction module instead of graph neural networks to model the correlation across entities, thus avoiding the information loss problem triggered by predefined edge-building rules. In addition, the initial word vector from the word embedding layer of a pre-trained language model is injected into entity representation to boost the performance in the extraction of relational facts that need commonsense knowledge. Experiments show that our model obtains competitive performance, especially compared with graph-based methods, which is faster and more effective. The source code, trained checkpoint files, and the commit results will be released to the public. Xinshu Shen, Man Lan |
IJCNN | 4 |
| 2023 | CCC: Chinese Commercial Contracts Dataset for Documents Layout Understanding
Yongnan Jin, Harry Lu, Shangqing Zhao, Man Lan, Yuefeng Chen |
NLPCC (2) | 5 |
| 2023 | Overview of the NLPCC 2023 Shared Task: Chinese Essay Discourse Coherence Evaluation
Hongyi Wu, Xinshu Shen, Man Lan, Xiaopeng Bai, Yuanbin Wu, Aimin Zhou, Shaoguang Mao, Tao Ge 0001, Yan Xia 0005 |
NLPCC (3) | 3 |
| 2022 | Understanding Gender Bias in Knowledge Base EmbeddingsabstractKnowledge base (KB) embeddings have been shown to contain gender biases (Fisher et al., 2020b).In this paper, we study two questions regarding these biases: how to quantify them, and how to trace their origins in KB? Specifically, first, we develop two novel bias measures respectively for a group of person entities and an individual person entity.Evidence of their validity is observed by comparison with real-world census data.Second, we use influence function to inspect the contribution of each triple in KB to the overall group bias.To exemplify the potential applications of our study, we also present two strategies (by adding and removing KB triples) to mitigate gender biases in KB embeddings. Yupei Du, Yuanbin Wu, Man Lan, Yan Yang 0008, Meirong Ma |
ACL (1) | 4 |
| 2022 | An Effective and Efficient Entity Alignment Decoding Algorithm via Third-Order Tensor IsomorphismabstractXin Mao, Meirong Ma, Hao Yuan, Jianchao Zhu, ZongYu Wang, Rui Xie, Wei Wu, Man Lan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Xin Mao 0002, Meirong Ma, Jianchao Zhu, Zongyu Wang, Man Lan |
ACL (1) | 8 |
| 2022 | A Simple Temporal Information Matching Mechanism for Entity Alignment between Temporal Knowledge GraphsabstractEntity alignment (EA) aims to find entities in different knowledge graphs (KGs) that refer to the same object in the real world. Recent studies incorporate temporal information to augment the representations of KGs. The existing methods for EA between temporal KGs (TKGs) utilize a time-aware attention mechanisms to incorporate relational and temporal information into entity embeddings. The approaches outperform the previous methods by using temporal information. However, we believe that it is not necessary to learn the embeddings of temporal information in KGs since most TKGs have uniform temporal representations. Therefore, we propose a simple GNN model combined with a temporal information matching mechanism, which achieves better performance with less time and fewer parameters. Furthermore, since alignment seeds are difficult to label in real-world applications, we also propose a method to generate unsupervised alignment seeds via the temporal information of TKG. Extensive experiments on public datasets indicate that our supervised method significantly outperforms the previous methods and the unsupervised one has competitive performance. Xin Mao 0002, Meirong Ma, Jianchao Zhu, Man Lan |
COLING | 6 |
| 2022 | Few Clean Instances Help Denoising Distant SupervisionabstractExisting distantly supervised relation extractors usually rely on noisy data for both model training and evaluation, which may lead to garbage-in-garbage-out systems. To alleviate the problem, we study whether a small clean dataset could help improve the quality of distantly supervised models. We show that besides getting a more convincing evaluation of models, a small clean dataset also helps us to build more robust denoising models. Specifically, we propose a new criterion for clean instance selection based on influence functions. It collects sample-level evidence for recognizing good instances (which is more informative than loss-level evidence). We also propose a teacher-student mechanism for controlling purity of intermediate results when bootstrapping the clean set. The whole approach is model-agnostic and demonstrates strong performances on both denoising real (NYT) and synthetic noisy datasets. Yufang Liu, Ziyin Huang, Changzhi Sun, Man Lan, Yuanbin Wu, Xiaofeng Mou |
COLING | 5 |
| 2022 | LightEA: A Scalable, Robust, and Interpretable Entity Alignment Framework via Three-view Label PropagationabstractEntity Alignment (EA) aims to find equivalent entity pairs between KGs, which is the core step of bridging and integrating multi-source KGs.In this paper, we argue that existing GNNbased EA methods inherit the inborn defects from their neural network lineage: weak scalability and poor interpretability.Inspired by recent studies, we reinvent the Label Propagation algorithm to effectively run on KGs and propose a non-neural EA framework -LightEA, consisting of three efficient components: (i) Random Orthogonal Label Generation, (ii) Three-view Label Propagation, and (iii) Sparse Sinkhorn Iteration.According to the extensive experiments on public datasets, LightEA has impressive scalability, robustness, and interpretability.With a mere tenth of time consumption, LightEA achieves comparable results to state-of-the-art methods across all datasets and even surpasses them on many. Xin Mao 0002, Yuanbin Wu, Man Lan |
EMNLP | 4 |
| 2022 | Multi-modal chemical information reconstruction from images and texts for exploring the near-drug spaceabstractIdentification of new chemical compounds with desired structural diversity and biological properties plays an essential role in drug discovery, yet the construction of such a potential space with elements of 'near-drug' properties is still a challenging task. In this work, we proposed a multimodal chemical information reconstruction system to automatically process, extract and align heterogeneous information from the text descriptions and structural images of chemical patents. Our key innovation lies in a heterogeneous data generator that produces cross-modality training data in the form of text descriptions and Markush structure images, from which a two-branch model with image- and text-processing units can then learn to both recognize heterogeneous chemical entities and simultaneously capture their correspondence. In particular, we have collected chemical structures from ChEMBL database and chemical patents from the European Patent Office and the US Patent and Trademark Office using keywords 'A61P, compound, structure' in the years from 2010 to 2020, and generated heterogeneous chemical information datasets with 210K structural images and 7818 annotated text snippets. Based on the reconstructed results and substituent replacement rules, structural libraries of a huge number of near-drug compounds can be generated automatically. In quantitative evaluations, our model can correctly reconstruct 97% of the molecular images into structured format and achieve an F1-score around 97-98% in the recognition of chemical entities, which demonstrated the effectiveness of our model in automatic information extraction from chemical patents, and hopefully transforming them to a user-friendly, structured molecular database enriching the near-drug space to realize the intelligent retrieval technology of chemical knowledge. Jie Wang 0146, Zihao Shen, Yichen Liao, Shiliang Li, Gaoqi He, Man Lan, Xuhong Qian, Kai Zhang 0001, Honglin Li 0003 |
Briefings Bioinform. | 7 |
| 2021 | Generating CCG Categories
Yufang Liu, Yuanbin Wu, Man Lan |
AAAI | 4 |
| 2021 | Target-dependent Event Detection: A New Task to Event Extraction from NewsabstractEvent extraction aims to detect events and extract event arguments. However, various events are not only too nuanced and complex to distinguish, but also involve multiple entities in the real-world scenario, especially in the financial field. This brings a great challenge to the current event extraction. To address these problems, previous event-centric methods detect events first and then extract arguments. Due to the diversity and complexity of events, event detection has a low performance, which is unfit for the huge amount of news in the real world. Given that the performance of named entity recognition (NER) is satisfactory, we shift our perspective from event-centric to target-centric view. In this paper, we propose a new task: target-dependent event detection (TDED), which aims to extract target entities and detect their corresponding events. We also propose a semantic and syntactic aware approach to support thousands of target entity extraction first and dozens of event types detection, that can be applied to massive corpora. Experimental results on a real-world Chinese financial dataset demonstrate that our model outperforms previous methods, especially in complex scenarios. Xin Mao 0002, Meirong Ma, Jianchao Zhu, Man Lan |
IEEE BigData | 7 |
| 2021 | Are Negative Samples Necessary in Entity Alignment?: An Approach with High Performance, Scalability and RobustnessabstractEntity alignment (EA) aims to find the equivalent entities in different KGs, which is a crucial step in integrating multiple KGs. However, most existing EA methods have poor scalability and are unable to cope with large-scale datasets. We summarize three issues leading to such high time-space complexity in existing EA methods: (1) Inefficient graph encoders, (2) Dilemma of negative sampling, and (3) "Catastrophic forgetting" in semi-supervised learning. To address these challenges, we propose a novel EA method with three new components to enable high Performance, high Scalability, and high Robustness (PSR): (1) Simplified graph encoder with relational graph sampling, (2) Symmetric negative-free alignment loss, and (3) Incremental semi-supervised learning. Furthermore, we conduct detailed experiments on several public datasets to examine the effectiveness and efficiency of our proposed method. The experimental results show that PSR not only surpasses the previous SOTA in performance but also has impressive scalability and robustness. Xin Mao 0002, Yuanbin Wu, Man Lan |
CIKM | 4 |
| 2021 | From Alignment to Assignment: Frustratingly Simple Unsupervised Entity AlignmentabstractCross-lingual entity alignment (EA) aims to find the equivalent entities between crosslingual KGs (Knowledge Graphs), which is a crucial step for integrating KGs.Recently, many GNN-based EA methods are proposed and show decent performance improvements on several public datasets.However, existing GNN-based EA methods inevitably inherit poor interpretability and low efficiency from neural networks.Motivated by the isomorphic assumption of GNN-based methods, we successfully transform the cross-lingual EA problem into an assignment problem.Based on this re-definition, we propose a frustratingly Simple but Effective Unsupervised entity alignment method (SEU) without neural networks.Extensive experiments have been conducted to show that our proposed unsupervised approach even beats advanced supervised methods across all public datasets while having high efficiency, interpretability, and stability. Xin Mao 0002, Yuanbin Wu, Man Lan |
EMNLP (1) | 4 |
| 2021 | A Dual-Attention Neural Network for Pun Location and Using Pun-Gloss Pairs for Interpretation
Meirong Ma, Jianguo Zhu 0001, Yuanbin Wu, Man Lan |
NLPCC (1) | 6 |
| 2021 | A Unified Information Extraction System Based on Role Recognition and Combination
Man Lan |
NLPCC (2) | 2 |
| 2021 | RoKGDS: A Robust Knowledge Grounded Dialog System
Jun Zhang 0098, Yushi Zhang, Weijie Xu, Jiahao Ying, Yan Yang 0008, Man Lan, Meirong Ma, Jianguo Zhu 0001 |
NLPCC (2) | 7 |
| 2021 | Boosting the Speed of Entity Alignment 10 ×: Dual Attention Matching Network with Normalized Hard Sample MiningabstractSeeking the equivalent entities among multi-source Knowledge Graphs (KGs) is the pivotal step to KGs integration, also known as entity alignment (EA). However, most existing EA methods are inefficient and poor in scalability. A recent summary points out that some of them even require several days to deal with a dataset containing 200,000 nodes (DWY100K). We believe over-complex graph encoder and inefficient negative sampling strategy are the two main reasons. In this paper, we propose a novel KG encoder — Dual Attention Matching Network (Dual-AMN), which not only models both intra-graph and cross-graph information smartly, but also greatly reduces computational complexity. Furthermore, we propose the Normalized Hard Sample Mining Loss to smoothly select hard negative samples with reduced loss shift. The experimental results on widely used public datasets indicate that our method achieves both high accuracy and high efficiency. On DWY100K, the whole running process of our method could be finished in 1,100 seconds, at least 10 × faster than previous work. The performances of our method also outperform previous works across all datasets, where [email protected] and MRR have been improved from 6% to 13%. Xin Mao 0002, Yuanbin Wu, Man Lan |
WWW | 4 |
| 2020 | A Span-based Linearization for Constituent TreesabstractWe propose a novel linearization of a constituent tree, together with a new locally normalized model.For each split point in a sentence, our model computes the normalizer on all spans ending with that split point, and then predicts a tree span from them.Compared with global models, our model is fast and parallelizable.Different from previous local models, our linearization method is tied on the spans directly and considers more local features when performing span prediction, which is more interpretable and effective.Experiments on PTB (95.8 F1) and CTB (92.1 F1) show that our model significantly outperforms existing local models and efficiently achieves competitive results with global models. Yuanbin Wu, Man Lan |
ACL | 3 |
| 2020 | Relational Reflection Entity AlignmentabstractEntity alignment aims to identify equivalent entity pairs from different Knowledge Graphs (KGs), which is essential in integrating multi-source KGs. Recently, with the introduction of GNNs into entity alignment, the architectures of recent models have become more and more complicated. We even find two counter-intuitive phenomena within these methods: (1) The standard linear transformation in GNNs is not working well. (2) Many advanced KG embedding models designed for link prediction task perform poorly in entity alignment. In this paper, we abstract existing entity alignment methods into a unified framework, Shape-Builder & Alignment, which not only successfully explains the above phenomena but also derives two key criteria for an ideal transformation operation. Furthermore, we propose a novel GNNs-based method, Relational Reflection Entity Alignment (RREA). RREA leverages Relational Reflection Transformation to obtain relation specific embeddings for each entity in a more efficient way. The experimental results on real-world datasets show that our model significantly outperforms the state-of-the-art methods, exceeding by 5.8%-10.9% on [email protected] Xin Mao 0002, Yuanbin Wu, Man Lan |
CIKM | 5 |
| 2020 | MRAEA: An Efficient and Robust Entity Alignment Approach for Cross-lingual Knowledge GraphabstractEntity alignment to find equivalent entities in cross-lingual Knowledge Graphs (KGs) plays a vital role in automatically integrating multiple KGs. Existing translation-based entity alignment methods jointly model the cross-lingual knowledge and monolingual knowledge into one unified optimization problem. On the other hand, the Graph Neural Network (GNN) based methods either ignore the node differentiations, or represent relation through entity or triple instances. They all fail to model the meta semantics embedded in relation nor complex relations such as n-to-n and multi-graphs. To tackle these challenges, we propose a novel Meta Relation Aware Entity Alignment (MRAEA) to directly model cross-lingual entity embeddings by attending over the node's incoming and outgoing neighbors and its connected relations' meta semantics. In addition, we also propose a simple and effective bi-directional iterative strategy to add new aligned seeds during training. Our experiments on all three benchmark entity alignment datasets show that our approach consistently outperforms the state-of-the-art methods, exceeding by 15%-58% on [email protected] Through an extensive ablation study, we validate that the proposed meta relation aware representations, relation aware self-attention and bi-directional iterative strategy of new seed selection all make contributions to significant performance improvement. The code is available at https://github.com/MaoXinn/MRAEA. Xin Mao 0002, Man Lan, Yuanbin Wu |
WSDM | 4 |
| 2019 | Graph-based Dependency Parsing with Graph Neural NetworksabstractWe investigate the problem of efficiently incorporating high-order features into neural graph-based dependency parsing.Instead of explicitly extracting high-order features from intermediate parse trees, we develop a more powerful dependency tree node representation which captures high-order information concisely and efficiently.We use graph neural networks (GNNs) to learn the representations and discuss several new configurations of GNN's updating and aggregation functions.Experiments on PTB show that our parser achieves the best UAS and LAS on PTB (96.0%, 94.3%) among systems without using any external resources. Yuanbin Wu, Man Lan |
ACL (1) | 3 |
| 2019 | Joint Type Inference on Entities and Relations via Graph Convolutional NetworksabstractWe develop a new paradigm for the task of joint entity relation extraction.It first identifies entity spans, then performs a joint inference on entity types and relation types.To tackle the joint type inference task, we propose a novel graph convolutional network (GCN) running on an entity-relation bipartite graph.By introducing a binary relation classification task, we are able to utilize the structure of entity-relation bipartite graph in a more efficient and interpretable way.Experiments on ACE05 show that our model outperforms existing joint models in entity performance and is competitive with the state-of-the-art in relation performance. Changzhi Sun, Yeyun Gong, Yuanbin Wu, Ming Gong 0001, Daxin Jiang, Man Lan, Shiliang Sun, Nan Duan 0001 |
ACL (1) | 6 |
| 2019 | Scaling up Open Tagging from Tens to Thousands: Comprehension Empowered Attribute Value Extraction from Product TitleabstractSupplementing product information by extracting attribute values from title is a crucial task in e-Commerce domain.Previous studies treat each attribute only as an entity type and build one set of NER tags (e.g., BIO) for each of them, leading to a scalability issue which unfits to the large sized attribute system in real world e-Commerce.In this work, we propose a novel approach to support value extraction scaling up to thousands of attributes without losing performance: (1) We propose to regard attribute as a query and adopt only one global set of BIO tags for any attributes to reduce the burden of attribute tag or model explosion;(2) We explicitly model the semantic representations for attribute and title, and develop an attention mechanism to capture the interactive semantic relations in-between to enforce our framework to be attribute comprehensive.We conduct extensive experiments in real-life datasets.The results show that our model not only outperforms existing state-of-the-art N-ER tagging models, but also is robust and generates promising results for up to 8, 906 attributes. Xin Mao 0002, Man Lan |
ACL (1) | 5 |
| 2019 | Exploring Human Gender Stereotypes with Word Association TestabstractYupei Du, Yuanbin Wu, Man Lan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yupei Du, Yuanbin Wu, Man Lan |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Handwritten Chinese Text Recognizer Applying Multi-level Multimodal Fusion NetworkabstractHandwritten Chinese text recognition (HCTR) has received extensive attention from the community of pattern recognition in the past decades. Most existing deep learning methods consist of two stages, i.e., training a text recognition network on the base of visual information, followed by incorporating language constrains with various language models. Therefore, the inherent linguistic semantic information is often neglected when designing the recognition network. To tackle this problem, in this work, we propose a novel multi-level multimodal fusion network and properly embed it into an attention-based LSTM so that both the visual information and the linguistic semantic information can be fully leveraged when predicting sequential outputs from the feature vectors. Experimental results on the ICDAR-2013 competition dataset demonstrate a comparable result with the state-of-the-art approaches. Yuhuan Xiu, Hongjian Zhan, Man Lan, Yue Lu 0001 |
ICDAR | 4 |
| 2019 | Residual Connection-Based Multi-step Reasoning via Commonsense Knowledge for Multiple Choice Machine Reading Comprehension
Yixuan Sheng, Man Lan |
ICONIP (3) | 2 |
| 2019 | Hierarchical Intention Enhanced Network for Automatic Dialogue Coherence AssessmentabstractDialogue coherence across multiple turns is still an open challenge. The entity grid model is arguably the most popular approach for coherence modeling. However, it heavily relies on the distribution of entities across adjacent sentences but ignores the emotional context embedded in non-entity text and fails to model long dependencies between speech intentions. These limitations become even more severe when applied to dialogue domain since sentences in dialogue are short, informal and colloquial, thereby, less entities could be extracted and less coherence information could be expressed in these grids. To address the limitations of entity gird methods and incorporate the structure knowledge of dialogue, we propose a new neural network architecture, Hierarchical Intention Enhance Network, to integrate semantic context and speech intention in both utterance and dialogue levels to hierarchically model the global coherence without any entity grids. Our proposed model outperforms the state-of-the-art entity-grid based coherence model on text discrimination task by 17.13% increase in accuracy, confirming the effectiveness of our hierarchical modeling in dialogue context and the crucial importance of intention information in dialogue coherence assessment. Yunxiao Zhou, Man Lan |
IJCNN | 2 |
| 2018 | Inference on Syntactic and Semantic Structures for Machine ComprehensionabstractHidden variable models are important tools for solving open domain machine comprehension tasks and have achieved remarkable accuracy in many question answering benchmark datasets. Existing models impose strong independence assumptions on hidden variables, which leaves the interaction among them unexplored. Here we introduce linguistic structures to help capturing global evidence in hidden variable modeling. In the proposed algorithms, question-answer pairs are scored based on structured inference results on parse trees and semantic frames, which aims to assign hidden variables in a global optimal way. Experiments on the MCTest dataset demonstrate that the proposed models are highly competitive with state-of-the-art machine comprehension systems. Yuanbin Wu, Man Lan |
AAAI | 3 |
| 2018 | A Multi-Task Learning Approach for Improving Product Title Compression with User Search Log DataabstractIt is a challenging and practical research problem to obtain effective compression of lengthy product titles for E-commerce. This is particularly important as more and more users browse mobile E-commerce apps and more merchants make the original product titles redundant and lengthy for Search Engine Optimization. Traditional text summarization approaches often require a large amount of preprocessing costs and do not capture the important issue of conversion rate in E-commerce. This paper proposes a novel multi-task learning approach for improving product title compression with user search log data. In particular, a pointer network-based sequence-to-sequence approach is utilized for title compression with an attentive mechanism as an extractive method and an attentive encoder-decoder approach is utilized for generating user search queries. The encoding parameters (i.e., semantic embedding of original titles) are shared among the two tasks and the attention distributions are jointly optimized. An extensive set of experiments with both human annotated data and online deployment demonstrate the advantage of the proposed research for both compression qualities and online business values. Jingang Wang, Long Qiu, Sheng Li 0017, Jun Lang 0001, Luo Si, Man Lan |
AAAI | 7 |
| 2018 | An Adversarial Joint Learning Model for Low-Resource Language Semantic Textual Similarity
Man Lan, Yuanbin Wu, Jingang Wang, Long Qiu, Sheng Li 0017, Jun Lang 0001, Luo Si |
ECIR | 2 |
| 2018 | Extracting Entities and Relations with Joint Minimum Risk TrainingabstractWe investigate the task of joint entity relation extraction.Unlike prior efforts, we propose a new lightweight joint learning paradigm based on minimum risk training (MRT).Specifically, our algorithm optimizes a global loss function which is flexible and effective to explore interactions between the entity model and the relation model.We implement a strong and simple neural network where the MRT is executed.Experiment results on the benchmark ACE05 and NYT datasets show that our model is able to achieve state-of-the-art joint extraction performances. Changzhi Sun, Yuanbin Wu, Man Lan, Shiliang Sun, Kuang-chih Lee, Kewen Wu 0003 |
EMNLP | 3 |
| 2018 | Memory-Based Model with Multiple Attentions for Multi-turn Response Selection
Xingwu Lu, Man Lan, Yuanbin Wu |
ICONIP (2) | 2 |
| 2018 | Towards a One-stop Solution to Both Aspect Extraction and Sentiment Analysis Tasks with Neural Multi-task LearningabstractPrevious studies usually divided aspect-based sentiment analysis into several subtasks in pipeline, i.e., first aspect term and/or opinion term extraction, then aspect-based sentiment prediction, resulting in error propagation and external resources dependency. To overcome the problems mentioned above, in this work we present a novel one-stop solution on aspect-based sentiment analysis. Specifically, we propose a novel multi-task neural learning framework to jointly tackle aspect extraction and sentiment prediction subtasks at the same time, and leverage attention mechanisms to learn the joint representation of aspect-sentiment relationship. We have conducted extensive comparative experiments on two benchmark datasets from SemEval-2014. The experiment results demonstrate the effectiveness of our proposed solution. Especially, our multi-task model outperforms the state-of-the-art systems on aspect extraction subtask. Feixiang Wang, Man Lan |
IJCNN | 2 |
| 2018 | A Neural Generation-based Conversation Model Using Fine-grained Emotion-guide AttentionabstractHuman emotion interaction is crucial to social communications. However, existing generation-based conversation systems mainly put emphasis on the content of responses in terms of naturalness, diversity and coherence without consideration of the emotion interaction between conversation. In order to reduce the gap between human-generated and computer-generated responses, in this work we present a human-like Emotional Conversation Generation Model, named ECGM, by imitating human conversation. Specifically, ECGM applies an emotion-guide attention which captures and integrates the emotion of the given post into neural response generation. Comparative experiments evaluated by computerised and manual methods show that our proposed model is capable of generating more human-like emotional responses and relevant content as well. Man Lan, Yuanbin Wu |
IJCNN | 2 |
| 2018 | Memory-Based Matching Models for Multi-turn Response Selection in Retrieval-Based Chatbots
Xingwu Lu, Man Lan, Yuanbin Wu |
NLPCC (1) | 2 |
| 2017 | Large-scale Opinion Relation Extraction with Distantly Supervised Neural NetworkabstractWe investigate the task of open domain opinion relation extraction. Different from works on manually labeled corpus, we propose an efficient distantly supervised framework based on pattern matching and neural network classifiers. The patterns are designed to automatically generate training data, and the deep learning model is design to capture various lexical and syntactic features. The result algorithm is fast and scalable on large-scale corpus. We test the system on the Amazon online review dataset. The result shows that our model is able to achieve promising performances without any human annotations. Changzhi Sun, Yuanbin Wu, Man Lan, Shiliang Sun, Qi Zhang 0001 |
EACL (1) | 3 |
| 2017 | Multi-task Attention-based Neural Networks for Implicit Discourse Relationship Representation and IdentificationabstractWe present a novel multi-task attentionbased neural network model to address implicit discourse relationship representation and identification through two types of representation learning, an attentionbased neural network for learning discourse relationship representation with two arguments and a multi-task framework for learning knowledge from annotated and unannotated corpora.The extensive experiments have been performed on two benchmark corpora (i.e., PDTB and CoNLL-2016 datasets).Experimental results show that our proposed model outperforms the state-of-the-art systems on benchmark corpora. Man Lan, Jianxiang Wang, Yuanbin Wu, Zhengyu Niu, Haifeng Wang 0001 |
EMNLP | 1 |
| 2017 | An Effective Gated and Attention-Based Neural Network Model for Fine-Grained Financial Target-Dependent Sentiment Analysis
Mengxiao Jiang, Jianxiang Wang, Man Lan, Yuanbin Wu |
KSEM | 3 |
| 2017 | A Learning Error Analysis for Structured Prediction with Approximate InferenceabstractIn this work, we try to understand the differences between exact and approximate inference algorithms in structured prediction. We compare the estimation and approximation error of both underestimate and overestimate models. The result shows that, from the perspective of learning errors, performances of approximate inference could be as good as exact inference. The error analyses also suggest a new margin for existing learning algorithms. Empirical evaluations on text classification, sequential labelling and dependency parsing witness the success of approximate inference and the benefit of the proposed margin. Yuanbin Wu, Man Lan, Shiliang Sun, Qi Zhang 0001, Xuanjing Huang 0001 |
NIPS | 2 |
| 2017 | Effective Semantic Relationship Classification of Context-Free Chinese Words with Simple Surface and Embedding Features
Yunxiao Zhou, Man Lan, Yuanbin Wu |
NLPCC | 2 |
| 2016 | Building mutually beneficial relationships between question retrieval and answer ranking to improve performance of community question answeringabstractIn community-based question answering (CQA) domain, there are two main tasks, i.e., question retrieval and answer ranking. Previous studies addressed these two tasks in an independent manner or in a sequential fashion without information communication. In this work we propose a novel method to improve the performance of CQA by mutually promoting the two tasks with the help of each other. Specifically, we propose two methods to improve question retrieval task by utilizing the rank of answers or extracting novel features from Q-A pairs respectively. Meanwhile, to improve answer ranking, we also present novel features with the help of similar questions. Experimental results on benchmark dataset showed that this mutually beneficial strategy between question retrieval and answer ranking not only improved the individual performance of these two tasks but also improved the overall performance of CQA through reducing errors propagating from question retrieval to answer ranking. Man Lan, GuoShun Wu, Chunyun Xiao, Yuanbin Wu, Ju Wu |
IJCNN | 1 |
| 2016 | Three Convolutional Neural Network-based models for learning Sentiment Word Vectors towards sentiment analysisabstractWith the development of deep learning, word vectors (i.e., word embeddings) have been extensively explored and applied to many Natural Language Processing tasks (e.g., parsing, Named Entity Recognition, etc). However, the semantic word vectors learned from context have insufficient sentiment information for performing sentiment analysis at different text levels. In this work, we present three Convolutional Neural Network (CNN)-based models to learn sentiment word vectors (SWV), which integrate sentiment information with semantic and syntactic information into word representations in three different strategies. Experimental results on benchmark datasets showed that sentiment word vectors are able to capture both sentiment and semantic information and outperform semantic word vectors for word-level and sentence-level sentiment analysis. Moreover, in combination with traditional NLP features, the sentiment word vectors achieve the best performance so far. Man Lan, Ju Wu |
IJCNN | 1 |
| 2015 | Integrating word embeddings and traditional NLP features to measure textual entailment and semantic relatedness of sentence pairsabstractRecent years the distributed representations of words (i.e., word embeddings) have been shown to be able to significantly improve performance in many natural language processing tasks, such as pos-of-tag tagging, chunking, named entity recognition and sentiment polarity judgement, etc. However, previous tasks only involve a single sentence. In contrast, this paper evaluates the effectiveness of word embeddings in sentence pair classification or regression problems. Specifically, we propose novel simple yet effective features based on word embeddings and extract many traditional linguistic features. Then these features serve as input of a classification/regression algorithm in isolation and in combination. Evaluations are conducted on three sentence pair classification/regression tasks, i.e., textual entailment, cross-lingual textual entailment and semantic relatedness estimation. Experiments on benchmark datasets provided by Semantic Evaluation 2013 and 2014 showed that using word embeddings is able to significantly improve the performance and our results outperform the best achieved results so far. Jiang Zhao, Man Lan, Zhengyu Niu, Yue Lu 0001 |
IJCNN | 2 |
| 2015 | Building a High Performance End-to-End Explicit Discourse Parser for Practical ApplicationabstractTo build practical end-to-end discourse parser, labeling arguments to discourse is the bottleneck to improve performance of whole parser. In consideration of the difference between syntactic and discourse arguments of connectives and the difference between two arguments to discourse in SS and PS cases, we present a method to build two separate argument extractors for two arguments. To evaluate the performance of whole parser, we build an end-to-end explicit discourse parser on PDTB. Experimental results showed that our proposed discourse parser achieved the best performance on explicit discourse so far. Jianxiang Wang, Man Lan |
KSEM | 2 |
| 2014 | Recognizing cross-lingual textual entailment with co-training using similarity and difference viewsabstractCross-lingual textual entailment is a relatively new problem that detects the entailment relationship between two text fragments written in different languages. Previous work adopted machine learning algorithms and similarity measures as features to address this task. In order to overcome the high cost of human annotation and further improve the recognition performance, we present a novel co-training approach to solve this problem. We first use an off-the-shelf machine translation tool to eliminate the language gap between two texts. Then we measure the similarities and differences between two texts and regard them as sufficient and redundant views. We use those two views to conduct the co-training procedure to perform classification. Besides, a new effective Kullback-Leibler (KL) based criterion is proposed to select the results from all possible iterations. Experiments on cross-lingual datasets provided by SemEval 2013 show that our method significantly outperforms the baseline systems and previous work. Jiang Zhao, Man Lan, Zhengyu Niu, Donghong Ji |
IJCNN | 2 |
| 2013 | From Semantic to Emotional Space in Probabilistic Sense Sentiment AnalysisabstractThis paper proposes an effective approach to model the emotional space of words to infer their Sense Sentiment Similarity (SSS). SSS reflects the distance between the words regarding their senses and underlying sentiments. We propose a probabilistic approach that is built on a hidden emotional model in which the basic human emotions are considered as hidden. This leads to predict a vector of emotions for each sense of the words, and then to infer the sense sentiment similarity. The effectiveness of the proposed approach is investigated in two Natural Language Processing tasks: Indirect yes/no Question Answer Pairs Inference and Sentiment Orientation Prediction. Mitra Mohtarami, Man Lan, Chew Lim Tan |
AAAI | 2 |
| 2013 | Leveraging Synthetic Discourse Data via Multi-task Learning for Implicit Discourse Relation Recognition
Man Lan, Zhengyu Niu |
ACL (1) | 1 |
| 2013 | Probabilistic Sense Sentiment Similarity through Hidden Emotions
Mitra Mohtarami, Man Lan, Chew Lim Tan |
ACL (1) | 2 |
| 2012 | Sense Sentiment Similarity: An AnalysisabstractThis paper describes an emotion-based approach to acquire sentiment similarity of word pairs with respect to their senses. Sentiment similarity indicates the similarity between two words from their underlying sentiments. Our approach is built on a model which maps from senses of words to vectors of twelve basic emotions. The emotional vectors are used to measure the sentiment similarity of word pairs. We show the utility of measuring sentiment similarity in two main natural language processing tasks, namely, indirect yes/no question answer pairs (IQAP) Inference and sentiment orientation (SO) prediction. Extensive experiments demonstrate that our approach can effectively capture the sentiment similarity of word pairs and utilize this information to address the above mentioned tasks. Mitra Mohtarami, Hadi Amiri, Man Lan, Thanh Phu Tran, Chew Lim Tan |
AAAI | 3 |
| 2012 | Connective prediction using machine learning for implicit discourse relation classificationabstractImplicit discourse relation classification is a challenge task due to missing discourse connective. Some work directly adopted machine learning algorithms and linguistically informed features to address this task. However, one interesting solution is to automatically predict implicit discourse connective. In this paper, we present a novel two-step machine learning-based approach to implicit discourse relation classification. We first use machine learning method to automatically predict the discourse connective that can best express the implicit discourse relation. Then the predicted implicit discourse connective is used to classify the implicit discourse relation. Experiments on Penn Discourse Treebank 2.0 (PDTB) and Biomedical Discourse Relation Bank (BioDRB) show that our method performs better than the baseline system and previous work. Man Lan, Yue Lu 0001, Zhengyu Niu, Chew Lim Tan |
IJCNN | 2 |
| 2011 | Predicting the uncertainty of sentiment adjectives in indirect answersabstractOpinion question answering (QA) requires automatic and correct interpretation of an answer relative to its question. However, the ambiguity that often exists in the question-answer pairs causes complexity in interpreting the answers. This paper aims to infer yes/no answers from indirect yes/no question-answer pairs (IQAPs) that are ambiguous due to the presence of ambiguous sentiment adjectives. We propose a method to measure the uncertainty of the answer in an IQAP relative to its question. In particular, to infer the yes or no response from an IQAP, our method employs antonyms, synonyms, word sense disambiguation as well as the semantic association between the sentiment adjectives that appear in the IQAP. Extensive experiments demonstrate the effectiveness of our method over the baseline. Mitra Mohtarami, Hadi Amiri, Man Lan, Chew Lim Tan |
CIKM | 3 |
| 2010 | The Effects of Discourse Connectives Prediction on Implicit Discourse Relation Recognition
Zhi-Min Zhou, Man Lan, Zhengyu Niu, Jian Su 0002 |
SIGDIAL Conference | 2 |
| 2010 | Empirical Investigations into Full-Text Protein Interaction Article Categorization Task (ACT) in the BioCreative II.5 ChallengeabstractThe selection of protein interaction documents is one important application for biology research and has a direct impact on the quality of downstream BioNLP applications, i.e., information extraction and retrieval, summarization, QA, etc. The BioCreative II.5 Challenge Article Categorization task (ACT) involves doing a binary text classification to determine whether a given structured full-text article contains protein interaction information. This may be the first attempt at classification of full-text protein interaction documents in wide community. In this paper, we compare and evaluate the effectiveness of different section types in full-text articles for text classification. Moreover, in practice, the less number of true-positive samples results in unstable performance and unreliable classifier trained on it. Previous research on learning with skewed class distributions has altered the class distribution using oversampling and downsampling. We also investigate the skewed protein interaction classification and analyze the effect of various issues related to the choice of external sources, oversampling training sets, classifiers, etc. We report on the various factors above to show that 1) a full-text biomedical article contains a wealth of scientific information important to users that may not be completely represented by abstracts and/or keywords, which improves the accuracy performance of classification and 2) reinforcing true-positive samples significantly increases the accuracy and stability performance of classification. Man Lan, Jian Su 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2009 | CpG-discover: A machine learning approach for CpG islands identification from human DNA sequenceabstractCpG islands (CGIs) play a fundamental role in genome analysis as genomic markers and tumor markers. Identification of potential CGIs has contributed not only to the prediction of promoters of most house-keeping genes and many tissue-specific genes but also to the understanding of the epigenetic causes of cancer. The most current methods for identifying CGIs suffered from various limitations and involved a lot of human intervention for search purpose. In this paper, we implement a HMM-based CGIs identification system, namely CpG-Discover. Experiments have been conducted on the EMBL human DNA database and in comparison with other widely-used tools. The controlled experimental results indicate that our system is a promising tool and has the capability of locating CGIs accurately. In addition, our system has significant differences from other tools in that it avoids the disadvantages of using sliding windows and it reduces the large amount of human intervention needed to search for or to combine potential CGIs (such as, the thresholds of initial density or distance seed). Therefore, given annotated training data set, our system has the adaptability to find other specific nucleotides sequences in DNA. Man Lan, Ying Zuo, Chew Lim Tan, Jian Su 0002 |
IJCNN | 1 |
| 2009 | Feature generation and representations for protein-protein interaction classification
Man Lan, Chew Lim Tan, Jian Su 0002 |
J. Biomed. Informatics | 1 |
| 2009 | Supervised and Traditional Term Weighting Methods for Automatic Text CategorizationabstractIn vector space model (VSM), text representation is the task of transforming the content of a textual document into a vector in the term space so that the document could be recognized and classified by a computer or a classifier. Different terms (i.e. words, phrases, or any other indexing units used to identify the contents of a text) have different importance in a text. The term weighting methods assign appropriate weights to the terms to improve the performance of text categorization. In this study, we investigate several widely-used unsupervised (traditional) and supervised term weighting methods on benchmark data collections in combination with SVM and kappa NN algorithms. In consideration of the distribution of relevant documents in the collection, we propose a new simple supervised term weighting method, i.e. tf.rf, to improve the terms' discriminating power for text categorization task. From the controlled experimental results, these supervised term weighting methods have mixed performance. Specifically, our proposed supervised term weighting method, tf.rf, has a consistently better performance than other term weighting methods while other supervised term weighting methods based on information theory or statistical metric perform the worst in all experiments. On the other hand, the popularly used tf.idf method has not shown a uniformly good performance in terms of different data sets. Man Lan, Chew Lim Tan, Jian Su 0002, Yue Lu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Adaptive EEG signal classification using stochastic approximation methodsabstractClassification of time-varying electrophysiological signals is an important problem in the development of brain-computer interfaces (BCIs). Designing adaptive classifiers is a potential way to address this task. In this paper, Bayesian classifiers with Gaussian mixture models (GMMs) are adopted as the decision rule to classify electroencephalogram (EEG) signals. The stochastic approximation method (SAM) is used as the specific gradient descent method for updating the parameters of mean values and covariance matrices in the distribution of GMMs, where the parameters are simultaneously updated in a batch mode. Experimental results using data from a BCI show that the stochastic approximation method is effective for EEG classification tasks. Shiliang Sun, Man Lan, Yue Lu 0001 |
ICASSP | 2 |
| 2007 | Text Representations for Text Categorization: A Case Study in Biomedical DomainabstractIn vector space model (VSM), textual documents are represented as vectors in the term space. Therefore, there are two issues in this representation, i.e. (1) what should a term be and (2) how to weight a term. This paper examined ways to represent text from the above two aspects to improve the performance of text categorization. Different representations have been evaluated using SVM on three biomedical corpora. The controlled experiments showed that the straightforward usage of named entities as terms in VSM does not show performance improvements over the bag-of-words representation. On the other hand, the term weighting method slightly improved the performance. However, to further improve the performance of text categorization, more advanced techniques and more effective usages of natural language processing for text representations appear needed. Man Lan, Chew Lim Tan, Jian Su 0002, Hwee-Boon Low |
IJCNN | 1 |
| 2006 | Proposing a New Term Weighting Scheme for Text Categorization
Man Lan, Chew Lim Tan, Hwee-Boon Low |
AAAI | 1 |
| 2005 | A comparative study on term weighting schemes for text categorizationabstractThe term weighting scheme, which is used to convert documents into vectors in the term spaces, is a vital step in automatic text categorization. The previous studies showed that term weighting schemes dominate the performance rather than the kernel functions of SVMs for the text categorization task. In this paper, we conducted experiments to compare various term weighting schemes with SVM on two widely-used benchmark data sets. We also presented a new term weighting scheme tf.rf for text categorization. The cross-scheme comparison was performed by using McNemar's tests. The controlled experimental results showed that the newly proposed tf.rf scheme is significantly better than other term weighting schemes. Compared with schemes related with tf factor alone, the idf factor does not improve or even decrease the term's discriminating power for text categorization. The binary and tf.chi representations significantly underperform the other term weighting schemes. Man Lan, Sam Yuan Sung, Hwee-Boon Low, Chew Lim Tan |
IJCNN | 1 |
| 2004 | Initialization of cluster refinement algorithms: a review and comparative studyabstractVarious iterative refinement clustering methods are dependent on the initial state of the model and are capable of obtaining one of their local optima only. Since the task of identifying the global optimization is NP-hard, the study of the initialization method towards a sub-optimization is of great value. This paper reviews the various cluster initialization methods in the literature by categorizing them into three major families, namely random sampling methods, distance optimization methods, and density estimation methods. In addition, using a set of quantitative measures, we assess their performance on a number of synthetic and real-life data sets. Our controlled benchmark identifies two distance optimization methods, namely SCS and KKZ, as complements of the k-means learning characteristics towards a better cluster separation in the output solution. Man Lan, Chew Lim Tan, Sam Yuan Sung, Hwee-Boon Low |
IJCNN | 2 |