Lili Mou

dblp:127/0779 · DBLP profile ↗
← Back
70ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0001-7753-4295ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 63 · 7 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
abstract
Yihong Dong, Xue Jiang, Yongding Tao, Huanyu Liu, Kechi Zhang, Lili Mou, Rongyu Cao, Yingwei MA, Jue Chen, Binhua Li, Zhi Jin, Fei Huang, Yongbin Li, Ge Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yihong Dong, Yongding Tao, Huanyu Liu 0001, Kechi Zhang, Lili Mou, Rongyu Cao, Yingwei Ma, Jue Chen 0003, Binhua Li, Zhi Jin 0001, Fei Huang 0002, Yongbin Li 0001, Ge Li 0001
ACL (1)6
2025 Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency Parsing
abstract
We address unsupervised dependency parsing by building an ensemble of diverse existing models through post hoc aggregation of their output dependency parse structures. We observe that these ensembles often suffer from low robustness against weak ensemble components due to error accumulation. To tackle this problem, we propose an efficient ensemble-selection approach that considers error diversity and avoids error accumulation. Results demonstrate that our approach outperforms each individual model as well as previous ensemble techniques. Additionally, our experiments show that the proposed ensemble-selection method significantly enhances the performance and robustness of our ensemble, surpassing previously proposed strategies, which have not accounted for error diversity.
Behzad Shayegh, Hobie H.-B. Lee, Xiaodan Zhu 0001, Jackie Chi Kit Cheung, Lili Mou
AAAI5
2025 EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine Translation
abstract
The ability of zero-shot translation emerges when we train a multilingual model with certain translation directions; the model can then directly translate in unseen directions. Alternatively, zero-shot translation can be accomplished by pivoting through a third language (e.g., English). In our work, we observe that both direct and pivot translations are noisy and achieve less satisfactory performance. We propose EBBS, an ensemble method with a novel bi-level beam search algorithm, where each ensemble component explores its own prediction step by step at the lower level but all components are synchronized by a "soft voting" mechanism at the upper level. Results on two popular multilingual translation datasets show that EBBS consistently outperforms direct and pivot translations, as well as existing ensemble techniques. Further, we can distill the ensemble's knowledge back to the multilingual model to improve inference efficiency; profoundly, our EBBS-distilled model can even outperform EBBS as it learns from the ensemble knowledge.
Yuqiao Wen, Behzad Shayegh, Chenyang Huang 0001, Yanshuai Cao, Lili Mou
AAAI5
2024 Tree-Averaging Algorithms for Ensemble-Based Unsupervised Discontinuous Constituency Parsing
abstract
We address unsupervised discontinuous constituency parsing, where we observe a high variance in the performance of the only previous model in the literature.We propose to build an ensemble of different runs of the existing discontinuous parser by averaging the predicted trees, to stabilize and boost performance.To begin with, we provide comprehensive computational complexity analysis (in terms of P and NP-complete) for tree averaging under different setups of binarity and continuity.We then develop an efficient exact algorithm to tackle the task, which runs in a reasonable time for all samples in our experiments.Results on three datasets show our method outperforms all baselines in all metrics; we also provide in-depth analyses of our approach.1
Behzad Shayegh, Yuqiao Wen, Lili Mou
ACL (1)3
2024 A Dual-View Approach to Classifying Radiology Reports by Co-Training
abstract
Radiology report analysis provides valuable information that can aid with public health initiatives, and has been attracting increasing attention from the research community. In this work, we present a novel insight that the structure of a radiology report (namely, the Findings and Impression sections) offers different views of a radiology scan. Based on this intuition, we further propose a co-training approach, where two machine learning models are built upon the Findings and Impression sections, respectively, and use each other’s information to boost performance with massive unlabeled data in a semi-supervised manner. We conducted experiments in a public health surveillance study, and results show that our co-training approach is able to improve performance using the dual views and surpass competing supervised and semi-supervised methods.
Yutong Han, Lili Mou
LREC/COLING3
2024 LLMR: Knowledge Distillation with a Large Language Model-Induced Reward
abstract
Large language models have become increasingly popular and demonstrated remarkable performance in various natural language processing (NLP) tasks. However, these models are typically computationally expensive and difficult to be deployed in resource-constrained environments. In this paper, we propose LLMR, a novel knowledge distillation (KD) method based on a reward function induced from large language models. We conducted experiments on multiple datasets in the dialogue generation and summarization tasks. Empirical results demonstrate that our LLMR approach consistently outperforms traditional KD methods in different tasks and datasets.
Dongheng Li, Yongchang Hao, Lili Mou
LREC/COLING3
2024 Claim-Centric and Sentiment Guided Graph Attention Network for Rumour Detection
abstract
Automatic rumour detection has gained attention due to the influence of social media on individuals and its pervasiveness. In this work, we construct a representation that takes into account the claim in the source tweet, considering both the propagation graph and the accompanying text alongside tweet sentiment. This is achieved through the implementation of a hierarchical attention mechanism, which not only captures the embedding of documents from individual word vectors but also combines these document representations as nodes within the propagation graph. Furthermore, to address potential overfitting concerns, we employ generative models to augment the existing datasets. This involves rephrasing the claims initially made in the source tweet, thereby creating a more diverse and robust dataset. In addition, we augment the dataset with sentiment labels to improve the performance of the rumour detection task. This holistic and refined approach yields a significant enhancement in the performance of our model across three distinct datasets designed for rumour detection. Quantitative and qualitative analysis proves the effectiveness of our methodology, surpassing the achievements of prior methodologies.
Sajad Ramezani, Mauajama Firdaus, Lili Mou
LREC/COLING3
2024 Ensemble Distillation for Unsupervised Constituency Parsing
abstract
We investigate the unsupervised constituency parsing task, which organizes words and phrases of a sentence into a hierarchical structure without using linguistically annotated data. We observe that existing unsupervised parsers capture different aspects of parsing structures, which can be leveraged to enhance unsupervised parsing performance. To this end, we propose a notion of "tree averaging," based on which we further propose a novel ensemble method for unsupervised parsing. To improve inference efficiency, we further distill the ensemble knowledge into a student model; such an ensemble-then-distill process is an effective approach to mitigate the over-smoothing problem existing in common multi-teacher distilling methods. Experiments show that our method surpasses all previous approaches, consistently demonstrating its effectiveness and robustness across various runs, with different ensemble components, and under domain-shift conditions.
Behzad Shayegh, Yanshuai Cao, Xiaodan Zhu 0001, Jackie Chi Kit Cheung, Lili Mou
ICLR5
2024 Zero-Shot Continuous Prompt Transfer: Generalizing Task Semantics Across Language Models
abstract
Prompt tuning in natural language processing (NLP) has become an increasingly popular method for adapting large language models to specific tasks. However, the transferability of these prompts, especially continuous prompts, between different models remains a challenge. In this work, we propose a zero-shot continuous prompt transfer method, where source prompts are encoded into relative space and the corresponding target prompts are searched for transferring to target models. Experimental results confirm the effectiveness of our method, showing that 'task semantics' in continuous prompts can be generalized across various language models. Moreover, we find that combining 'task semantics' from multiple source models can further enhance the performance of transfer.
Zijun Wu 0002, Yongkang Wu, Lili Mou
ICLR3
2024 Flora: Low-Rank Adapters Are Secretly Gradient Compressors
abstract
Despite large neural networks demonstrating remarkable abilities to complete different tasks, they require excessive memory usage to store the optimization states for training. To alleviate this, the low-rank adaptation (LoRA) is proposed to reduce the optimization states by training fewer parameters. However, LoRA restricts overall weight update matrices to be low-rank, limiting the model performance. In this work, we investigate the dynamics of LoRA and identify that it can be approximated by a random projection. Based on this observation, we propose Flora, which is able to achieve high-rank updates by resampling the projection matrices while enjoying the sublinear space complexity of optimization states. We conduct experiments across different tasks and model architectures to verify the effectiveness of our approach.
Yongchang Hao, Yanshuai Cao, Lili Mou
ICML3
2023 f-Divergence Minimization for Sequence-Level Knowledge Distillation
abstract
Knowledge distillation (KD) is the process of transferring knowledge from a large model to a small one.It has gained increasing attention in the natural language processing community, driven by the demands of compressing evergrowing language models.In this work, we propose an f -DISTILL framework, which formulates sequence-level knowledge distillation as minimizing a generalized f -divergence function.We propose four distilling variants under our framework and show that existing SeqKD and ENGINE approaches are approximations of our f -DISTILL methods.We further derive step-wise decomposition for our f -DISTILL, reducing intractable sequence-level divergence to word-level losses that can be computed in a tractable manner.Experiments across four datasets show that our methods outperform existing KD approaches, and that our symmetric distilling losses can better force the student to learn from the teacher distribution.1
Yuqiao Wen, Zichao Li 0001, Wenyu Du, Lili Mou
ACL (1)4
2023 An Equal-Size Hard EM Algorithm for Diverse Dialogue Generation
Yuqiao Wen, Yongchang Hao, Yanshuai Cao, Lili Mou
ICLR4
2023 Weakly Supervised Explainable Phrasal Reasoning with Neural Fuzzy Logic
Zijun Wu 0002, Zi Xuan Zhang, Atharva Naik, Zhijian Mei, Mauajama Firdaus, Lili Mou
ICLR6
2022 Non-autoregressive Translation with Layer-Wise Prediction and Deep Supervision
abstract
How do we perform efficient inference while retaining high translation quality? Existing neural machine translation models, such as Transformer, achieve high performance, but they decode words one by one, which is inefficient. Recent non-autoregressive translation models speed up the inference, but their quality is still inferior. In this work, we propose DSLP, a highly efficient and high-performance model for machine translation. The key insight is to train a non-autoregressive Transformer with Deep Supervision and feed additional Layer-wise Predictions. We conducted extensive experiments on four translation tasks (both directions of WMT'14 EN-DE and WMT'16 EN-RO). Results show that our approach consistently improves the BLEU scores compared with respective base models. Specifically, our best variant outperforms the autoregressive model on three translation tasks, while being 14.8 times more efficient in inference.
Chenyang Huang 0001, Hao Zhou 0012, Osmar R. Zaïane, Lili Mou, Lei Li 0005
AAAI4
2022 Generalized Equivariance and Preferential Labeling for GNN Node Classification
abstract
Existing graph neural networks (GNNs) largely rely on node embeddings, which represent a node as a vector by its identity, type, or content. However, graphs with unattributed nodes widely exist in real-world applications (e.g., anonymized social networks). Previous GNNs either assign random labels to nodes (which introduces artefacts to the GNN) or assign one embedding to all nodes (which fails to explicitly distinguish one node from another). Further, when these GNNs are applied to unattributed node classification problems, they have an undesired equivariance property, which are fundamentally unable to address the data with multiple possible outputs. In this paper, we analyze the limitation of existing approaches to node classification problems. Inspired by our analysis, we propose a generalized equivariance property and a Preferential Labeling technique that satisfies the desired property asymptotically. Experimental results show that we achieve high performance in several unattributed node classification tasks.
Zeyu Sun 0004, Wenjie Zhang 0007, Lili Mou, Qihao Zhu, Yingfei Xiong 0001, Lu Zhang 0023
AAAI3
2022 Search and Learn: Improving Semantic Coverage for Data-to-Text Generation
abstract
Data-to-text generation systems aim to generate text descriptions based on input data (often represented in the tabular form). A typical system uses huge training samples for learning the correspondence between tables and texts. However, large training sets are expensive to obtain, limiting the applicability of these approaches in real-world scenarios. In this work, we focus on few-shot data-to-text generation. We observe that, while fine-tuned pretrained language models may generate plausible sentences, they suffer from the low semantic coverage problem in the few-shot setting. In other words, important input slots tend to be missing in the generated text. To this end, we propose a search-and-learning approach that leverages pretrained language models but inserts the missing slots to improve the semantic coverage. We further finetune our system based on the search results to smooth out the search noise, yielding better-quality text and improving inference efficiency to a large extent. Experiments show that our model achieves high performance on E2E and WikiBio datasets. Especially, we cover 98.35% of input slots on E2E, largely alleviating the low coverage problem.
Shailza Jolly, Zi Xuan Zhang, Andreas Dengel 0001, Lili Mou
AAAI4
2022 Learning Non-Autoregressive Models from Search for Unsupervised Sentence Summarization
abstract
Text summarization aims to generate a short summary for an input text.In this work, we propose a Non-Autoregressive Unsupervised Summarization (NAUS) approach, which does not require parallel data for training.Our NAUS first performs edit-based search towards a heuristically defined score, and generates a summary as pseudo-groundtruth.Then, we train an encoder-only non-autoregressive Transformer based on the search result.We also propose a dynamic programming approach for length-control decoding, which is important for the summarization task.Experiments on two datasets show that NAUS achieves state-of-the-art performance for unsupervised summarization, yet largely improving inference efficiency.Further, our algorithm is able to perform explicit length-transfer summary generation. 1
Puyuan Liu, Chenyang Huang 0001, Lili Mou
ACL (1)3
2022 An Empirical Study on the Overlapping Problem of Open-Domain Dialogue Datasets
abstract
Open-domain dialogue systems aim to converse with humans through text, and dialogue research has heavily relied on benchmark datasets. In this work, we observe the overlapping problem in DailyDialog and OpenSubtitles, two popular open-domain dialogue benchmark datasets. Our systematic analysis then shows that such overlapping can be exploited to obtain fake state-of-the-art performance. Finally, we address this issue by cleaning these datasets and setting up a proper data processing procedure for future research.
Yuqiao Wen, Guoqing Luo, Lili Mou
LREC3
2022 Document-Level Relation Extraction with Sentences Importance Estimation and Focusing
abstract
Document-level relation extraction (DocRE) aims to determine the relation between two entities from a document of multiple sentences.Recent studies typically represent the entire document by sequence-or graph-based models to predict the relations of all entity pairs.However, we find that such a model is not robust and exhibits bizarre behaviors: it predicts correctly when an entire test document is fed as input, but errs when non-evidence sentences are removed.To this end, we propose a Sentence Importance Estimation and Focusing (SIEF) framework for DocRE, where we design a sentence importance score and a sentence focusing loss, encouraging DocRE models to focus on evidence sentences.Experimental results on two domains show that our SIEF not only improves overall performance, but also makes DocRE models more robust.Moreover, SIEF is a general framework, shown to be effective when combined with a variety of base DocRE models. 1
Kehai Chen, Lili Mou, Tiejun Zhao
NAACL-HLT3
2022 Teacher Forcing Recovers Reward Functions for Text Generation
abstract
Reinforcement learning (RL) has been widely used in text generation to alleviate the exposure bias issue or to utilize non-parallel datasets. The reward function plays an important role in making RL training successful. However, previous reward functions are typically task-specific and sparse, restricting the use of RL. In our work, we propose a task-agnostic approach that derives a step-wise reward function directly from a model trained with teacher forcing. We additionally propose a simple modification to stabilize the RL training on non-parallel datasets with our induced reward function. Empirical results show that our method outperforms self-training and reward regression methods on several text generation tasks, confirming the effectiveness of our reward function.
Yongchang Hao, Lili Mou
NeurIPS3
2022 A Character-Level Length-Control Algorithm for Non-Autoregressive Sentence Summarization
abstract
Sentence summarization aims at compressing a long sentence into a short one that keeps the main gist, and has extensive real-world applications such as headline generation. In previous work, researchers have developed various approaches to improve the ROUGE score, which is the main evaluation metric for summarization, whereas controlling the summary length has not drawn much attention. In our work, we address a new problem of explicit character-level length control for summarization, and propose a dynamic programming algorithm based on the Connectionist Temporal Classification (CTC) model. Results show that our approach not only achieves higher ROUGE scores but also yields more complete sentences.
Puyuan Liu, Lili Mou
NeurIPS3
2022 MBCT: Tree-Based Feature-Aware Binning for Individual Uncertainty Calibration
abstract
Most machine learning classifiers only concern classification accuracy, while certain applications (such as medical diagnosis, meteorological forecasting, and computation advertising) require the model to predict the true probability, known as a calibrated estimate. In previous work, researchers have developed several calibration methods to post-process the outputs of a predictor to obtain calibrated values, such as binning and scaling methods. Compared with scaling, binning methods are shown to have distribution-free theoretical guarantees, which motivates us to prefer binning methods for calibration. However, we notice that existing binning methods have several drawbacks: (a) the binning scheme only considers the original prediction values, thus limiting the calibration performance; and (b) the binning approach is non-individual, mapping multiple samples in a bin to the same value, and thus is not suitable for order-sensitive applications. In this paper, we propose a feature-aware binning framework, called Multiple Boosting Calibration Trees (MBCT), along with a multi-view calibration loss to tackle the above issues. Our MBCT optimizes the binning scheme by the tree structures of features, and adopts a linear function in a tree node to achieve individual calibration. Our MBCT is non-monotonic, and has the potential to improve order accuracy, due to its learnable binning scheme and the individual calibration. We conduct comprehensive experiments on three datasets in different fields. Results show that our method outperforms all competing models in terms of both calibration error and order accuracy. We also conduct simulation experiments, justifying that the proposed multi-view calibration loss is a better metric in modeling calibration error. In addition, our approach is deployed in a real-world online advertising platform; an A/B test over two weeks further demonstrates the effectiveness and great business value of our approach.
Siguang Huang, Yunli Wang, Lili Mou, Huayue Zhang, Han Zhu 0001, Chuan Yu 0002, Bo Zheng 0007
WWW3
2021 Simulated Annealing for Emotional Dialogue Systems
abstract
Explicitly modeling emotions in dialogue generation has important applications, such as building empathetic personal companions. In this study, we consider the task of expressing a specific emotion for dialogue generation. Previous approaches take the emotion as a training signal, which may be ignored during inference. Here, we propose a search-based emotional dialogue system by simulated annealing (SA). Specifically, we first define a scoring function that combines contextual coherence and emotional correctness. Then, SA iteratively edits a general response, and search for a generation with a high score. In this way, we enforce the presence of the desired emotion. We evaluate our system on the NLPCC2017 dataset. The proposed method shows about 12% improvements in emotion accuracy compared with the previous state-of-the-art method, without hurting the generation quality (measured by BLEU).
Chengzhang Dong, Chenyang Huang 0001, Osmar R. Zaïane, Lili Mou
CIKM4
2021 Seq2Emo: A Sequence to Multi-Label Emotion Classification Model
abstract
Chenyang Huang, Amine Trabelsi, Xuebin Qin, Nawshad Farruque, Lili Mou, Osmar Zaïane. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Chenyang Huang 0001, Amine Trabelsi, Xuebin Qin, Nawshad Farruque, Lili Mou, Osmar R. Zaïane
NAACL-HLT5
2021 Simulated annealing for optimization of graphs and sequences
Xianggen Liu, Pengyong Li, Fandong Meng, Hao Zhou 0012, Huasong Zhong, Jie Zhou 0016, Lili Mou, Sen Song
Neurocomputing7
2020 TreeGen: A Tree-Based Transformer Architecture for Code Generation
abstract
A code generation system generates programming language code based on an input natural language description. State-of-the-art approaches rely on neural networks for code generation. However, these code generators suffer from two problems. One is the long dependency problem, where a code element often depends on another far-away code element. A variable reference, for example, depends on its definition, which may appear quite a few lines before. The other problem is structure modeling, as programs contain rich structural information. In this paper, we propose a novel tree-based neural architecture, TreeGen, for code generation. TreeGen uses the attention mechanism of Transformers to alleviate the long-dependency problem, and introduces a novel AST reader (encoder) to incorporate grammar rules and AST structures into the network. We evaluated TreeGen on a Python benchmark, HearthStone, and two semantic parsing benchmarks, ATIS and GEO. TreeGen outperformed the previous state-of-the-art approach by 4.5 percentage points on HearthStone, and achieved the best accuracy among neural network-based approaches on ATIS (89.1%) and GEO (89.6%). We also conducted an ablation test to better understand each component of our model.
Zeyu Sun 0004, Qihao Zhu, Yingfei Xiong 0001, Yican Sun, Lili Mou, Lu Zhang 0023
AAAI5
2020 Iterative Edit-Based Unsupervised Sentence Simplification
abstract
We present a novel iterative, edit-based approach to unsupervised sentence simplification.Our model is guided by a scoring function involving fluency, simplicity, and meaning preservation.Then, we iteratively perform word and phrase-level edits on the complex sentence.Compared with previous approaches, our model does not require a parallel training set, but is more controllable and interpretable.Experiments on Newsela and WikiLarge datasets show that our approach is nearly as effective as state-of-the-art supervised approaches. 1
Dhruv Kumar 0005, Lili Mou, Lukasz Golab, Olga Vechtomova
ACL2
2020 Unsupervised Paraphrasing by Simulated Annealing
abstract
We propose UPSA, a novel approach that accomplishes Unsupervised Paraphrasing by Simulated Annealing.We model paraphrase generation as an optimization problem and propose a sophisticated objective function, involving semantic similarity, expression diversity, and language fluency of paraphrases.UPSA searches the sentence space towards this objective by performing a sequence of local edits.We evaluate our approach on various datasets, namely, Quora, Wikianswers, MSCOCO, and Twitter.Extensive results show that UPSA achieves the state-of-the-art performance compared with previous unsupervised methods in terms of both automatic and human evaluations.Further, our approach outperforms most existing domain-adapted supervised models, showing the generalizability of UPSA. 1
Xianggen Liu, Lili Mou, Fandong Meng, Hao Zhou 0012, Jie Zhou 0016, Sen Song
ACL2
2020 Discrete Optimization for Unsupervised Sentence Summarization with Word-Level Extraction
abstract
Automatic sentence summarization produces a shorter version of a sentence, while preserving its most important information.A good summary is characterized by language fluency and high information overlap with the source sentence.We model these two aspects in an unsupervised objective function, consisting of language modeling and semantic similarity metrics.We search for a high-scoring summary by discrete optimization.Our proposed method achieves a new state-of-the art for unsupervised sentence summarization according to ROUGE scores.Additionally, we demonstrate that the commonly reported ROUGE F1 metric is sensitive to summary length.Since this is unwillingly exploited in recent work, we emphasize that future evaluation should explicitly group summarization systems by output length brackets.1
Raphael Schumann, Lili Mou, Olga Vechtomova, Katja Markert
ACL2
2020 Adversarial Learning on the Latent Space for Diverse Dialog Generation
abstract
Generating relevant responses in a dialog is challenging, and requires not only proper modeling of context in the conversation, but also being able to generate fluent sentences during inference.In this paper, we propose a two-step framework based on generative adversarial nets for generating conditioned responses.Our model first learns a meaningful representation of sentences by autoencoding, and then learns to map an input query to the response representation, which is in turn decoded as a response sentence.Both quantitative and qualitative evaluations show that our model generates more fluent, relevant, and diverse responses than existing state-of-the-art methods. 1
Kashif Khan, Gaurav Sahu, Vikash Balasubramanian, Lili Mou, Olga Vechtomova
COLING4
2020 Formality Style Transfer with Shared Latent Space
abstract
Conventional approaches for formality style transfer borrow models from neural machine translation, which typically requires massive parallel data for training. However, the dataset for formality style transfer is considerably smaller than translation corpora. Moreover, we observe that informal and formal sentences closely resemble each other, which is different from the translation task where two languages have different vocabularies and grammars. In this paper, we present a new approach, Sequence-to-Sequence with Shared Latent Space (S2S-SLS), for formality style transfer, where we propose two auxiliary losses and adopt joint training of bi-directional transfer and auto-encoding. Experimental results show that S2S-SLS (with either RNN or Transformer architectures) consistently outperforms baselines in various settings, especially when we have limited data.
Yunli Wang, Yu Wu 0012, Lili Mou, Zhoujun Li 0001, Wen-Han Chao
COLING3
2020 Improving Word Sense Disambiguation with Translations
abstract
It has been conjectured that multilingual information can help monolingual word sense disambiguation (WSD).However, existing WSD systems rarely consider multilingual information, and no effective method has been proposed for improving WSD by generating translations.In this paper, we present a novel approach that improves the performance of a base WSD system using machine translation.Since our approach is language independent, we perform WSD experiments on several languages.The results demonstrate that our methods can consistently improve the performance of WSD systems, and obtain state-ofthe-art results in both English and multilingual WSD.To facilitate the use of lexical translation information, we also propose BABALIGN, an precise bitext alignment algorithm which is guided by multilingual lexical correspondences from BabelNet.
Yixing Luan, Bradley Hauer, Lili Mou, Grzegorz Kondrak
EMNLP (1)3
2020 Progressive Memory Banks for Incremental Domain Adaptation
Nabiha Asghar, Lili Mou, Kira A. Selby, Kevin D. Pantasdo, Pascal Poupart, Xin Jiang 0002
ICLR2
2020 Unsupervised Text Generation by Learning from Search
abstract
In this work, we propose TGLS, a novel framework for unsupervised Text Generation by Learning from Search. We start by applying a strong search algorithm (in particular, simulated annealing) towards a heuristically defined objective that (roughly) estimates the quality of sentences. Then, a conditional generative model learns from the search results, and meanwhile smooth out the noise of search. The alternation between search and learning can be repeated for performance bootstrapping. We demonstrate the effectiveness of TGLS on two real-world natural language generation tasks, unsupervised paraphrasing and text formalization. Our model significantly outperforms unsupervised baseline methods in both tasks. Especially, it achieves comparable performance to strong supervised methods for paraphrase generation.
Jingjing Li 0007, Zichao Li 0001, Lili Mou, Xin Jiang 0002, Michael R. Lyu, Irwin King
NeurIPS3
2020 Finding decision jumps in text classification
Xianggen Liu, Lili Mou, Haotian Cui, Zhengdong Lu, Sen Song
Neurocomputing2
2019 CGMH: Constrained Sentence Generation by Metropolis-Hastings Sampling
abstract
In real-world applications of natural language generation, there are often constraints on the target sentences in addition to fluency and naturalness requirements. Existing language generation techniques are usually based on recurrent neural networks (RNNs). However, it is non-trivial to impose constraints on RNNs while maintaining generation quality, since RNNs generate sentences sequentially (or with beam search) from the first word to the last. In this paper, we propose CGMH, a novel approach using Metropolis-Hastings sampling for constrained sentence generation. CGMH allows complicated constraints such as the occurrence of multiple keywords in the target sentences, which cannot be handled in traditional RNN-based approaches. Moreover, CGMH works in the inference stage, and does not require parallel corpora for training. We evaluate our method on a variety of tasks, including keywords-to-sentence generation, unsupervised sentence paraphrasing, and unsupervised sentence error correction. CGMH achieves high performance compared with previous supervised methods for sentence generation. Our code is released at https://github.com/NingMiao/CGMH
Ning Miao, Hao Zhou 0012, Lili Mou, Rui Yan 0001, Lei Li 0005
AAAI3
2019 A Grammar-Based Structural CNN Decoder for Code Generation
abstract
Code generation maps a program description to executable source code in a programming language. Existing approaches mainly rely on a recurrent neural network (RNN) as the decoder. However, we find that a program contains significantly more tokens than a natural language sentence, and thus it may be inappropriate for RNN to capture such a long sequence. In this paper, we propose a grammar-based structural convolutional neural network (CNN) for code generation. Our model generates a program by predicting the grammar rules of the programming language; we design several CNN modules, including the tree-based convolution and pre-order convolution, whose information is further aggregated by dedicated attentive pooling layers. Experimental results on the HearthStone benchmark dataset show that our CNN code generator significantly outperforms the previous state-of-the-art method by 5 percentage points; additional experiments on several semantic parsing tasks demonstrate the robustness of our model. We also conduct in-depth ablation test to better understand each component of our model.
Zeyu Sun 0004, Qihao Zhu, Lili Mou, Yingfei Xiong 0001, Ge Li 0001, Lu Zhang 0023
AAAI3
2019 Generating Sentences from Disentangled Syntactic and Semantic Spaces
abstract
Variational auto-encoders (VAEs) are widely used in natural language generation due to the regularization of the latent space.However, generating sentences from the continuous latent space does not explicitly model the syntactic information.In this paper, we propose to generate sentences from disentangled syntactic and semantic spaces.Our proposed method explicitly models syntactic information in the VAE's latent space by using the linearized tree sequence, leading to better performance of language generation.Additionally, the advantage of sampling in the disentangled syntactic and semantic latent spaces enables us to perform novel applications, such as the unsupervised paraphrase generation and syntaxtransfer generation.Experimental results show that our proposed model achieves similar or better performance in various tasks, compared with state-of-the-art related work.
Hao Zhou 0012, Shujian Huang, Lei Li 0005, Lili Mou, Olga Vechtomova, Xinyu Dai, Jiajun Chen 0001
ACL (1)5
2019 Disentangled Representation Learning for Non-Parallel Text Style Transfer
abstract
This paper tackles the problem of disentangling the latent representations of style and content in language models.We propose a simple yet effective approach, which incorporates auxiliary multi-task and adversarial objectives, for style prediction and bag-of-words prediction, respectively.We show, both qualitatively and quantitatively, that the style and content are indeed disentangled in the latent space.This disentangled latent representation learning can be applied to style transfer on non-parallel corpora.We achieve high performance in terms of transfer accuracy, content preservation, and language fluency, in comparison to various previous approaches. 1
Vineet John, Lili Mou, Hareesh Bahuleyan, Olga Vechtomova
ACL (1)2
2019 An Imitation Learning Approach to Unsupervised Parsing
abstract
Recently, there has been an increasing interest in unsupervised parsers that optimize semantically oriented objectives, typically using reinforcement learning.Unfortunately, the learned trees often do not match actual syntax trees well.Shen et al. (2018) propose a structured attention mechanism for language modeling (PRPN), which induces better syntactic structures but relies on ad hoc heuristics.Also, their model lacks interpretability as it is not grounded in parsing actions.In our work, we propose an imitation learning approach to unsupervised parsing, where we transfer the syntactic knowledge induced by the PRPN to a Tree-LSTM model with discrete parsing actions.Its policy is then refined by Gumbel-Softmax training towards a semantically oriented objective.We evaluate our approach on the All Natural Language Inference dataset and show that it achieves a new state of the art in terms of parsing F -score, outperforming our base models, including the PRPN. 1
Bowen Li 0002, Lili Mou, Frank Keller
ACL (1)2
2019 Harnessing Pre-Trained Neural Networks with Rules for Formality Style Transfer
abstract
Yunli Wang, Yu Wu, Lili Mou, Zhoujun Li, Wenhan Chao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yunli Wang, Yu Wu 0012, Lili Mou, Zhoujun Li 0001, Wen-Han Chao
EMNLP/IJCNLP (1)3
2019 Why Do Neural Dialog Systems Generate Short and Meaningless Replies? a Comparison between Dialog and Translation
abstract
This paper addresses the question: In neural dialog systems, why do sequence-to-sequence (Seq2Seq) neural networks generate short and meaningless replies for open-domain response generation? We conjecture that in a dialog system, due to the randomness of spoken language, there may be multiple equally plausible replies for one utterance, causing the deficiency of a Seq2Seq model. To evaluate our conjecture, we propose a systematic way to mimic the dialog scenario in machine translation systems with both real datasets and toy datasets generated elaborately. Experimental results show that we manage to reproduce the phenomenon of generating short and meaningless sentences in the translation setting.
Bolin Wei, Lili Mou, Hao Zhou 0012, Pascal Poupart, Ge Li 0001, Zhi Jin 0001
ICASSP3
2018 Towards Neural Speaker Modeling in Multi-Party Conversation: The Task, Dataset, and Models
abstract
In this paper, we address the problem of speaker classification in multi-party conversation, and collect massive data to facilitate research in this direction. We further investigate temporal-based and content-based models of speakers, and propose several hybrids of them. Experiments show that speaker classification is feasible, and that hybrid models outperform each single component.
Lili Mou, Zhi Jin 0001
AAAI2
2018 Order-Planning Neural Text Generation From Structured Data
abstract
Generating texts from structured data (e.g., a table) is important for various natural language processing tasks such as question answering and dialog systems. In recent studies, researchers use neural language models and encoder-decoder frameworks for table-to-text generation. However, these neural network-based approaches typically do not model the order of content during text generation. When a human writes a summary based on a given table, he or she would probably consider the content order before wording. In this paper, we propose an order-planning text generation model, where order information is explicitly captured by link-based attention. Then a self-adaptive gate combines the link-based attention with traditional content-based attention. We conducted experiments on the WikiBio dataset and achieve higher performance than previous methods in terms of BLEU, ROUGE, and NIST scores; we also performed ablation tests to analyze each component of our model.
Lei Sha, Lili Mou, Tianyu Liu 0001, Pascal Poupart, Sujian Li, Baobao Chang, Zhifang Sui
AAAI2
2018 RUBER: An Unsupervised Method for Automatic Evaluation of Open-Domain Dialog Systems
abstract
Open-domain human-computer conversation has been attracting increasing attention over the past few years. However, there does not exist a standard automatic evaluation metric for open-domain dialog systems; researchers usually resort to human annotation for model evaluation, which is time- and labor-intensive. In this paper, we propose RUBER, a Referenced metric and Unreferenced metric Blended Evaluation Routine, which evaluates a reply by taking into consideration both a groundtruth reply and a query (previous user-issued utterance). Our metric is learnable, but its training does not require labels of human satisfaction. Hence, RUBER is flexible and extensible to different datasets and languages. Experiments on both retrieval and generative dialog systems show that RUBER has a high correlation with human annotation, and that RUBER has fair transferability over different datasets.
Chongyang Tao, Lili Mou, Dongyan Zhao 0001, Rui Yan 0001
AAAI2
2018 Variational Attention for Sequence-to-Sequence Models
abstract
The variational encoder-decoder (VED) encodes source information as a set of random variables using a neural network, which in turn is decoded into target data using another neural network. In natural language processing, sequence-to-sequence (Seq2Seq) models typically serve as encoder-decoder networks. When combined with a traditional (deterministic) attention mechanism, the variational latent space may be bypassed by the attention model, and thus becomes ineffective. In this paper, we propose a variational attention mechanism for VED, where the attention vector is also modeled as Gaussian distributed random variables. Results on two experiments show that, without loss of quality, our proposed method alleviates the bypassing phenomenon as it increases the diversity of generated sentences.
Hareesh Bahuleyan, Lili Mou, Olga Vechtomova, Pascal Poupart
COLING2
2018 Affective Neural Response Generation
Nabiha Asghar, Pascal Poupart, Jesse Hoey, Xin Jiang 0002, Lili Mou
ECIR5
2018 Jumper: Learning When to Make Classification Decision in Reading
abstract
In early years, text classification is typically accomplished by feature-based classifiers; recently, neural networks, as powerful classifiers, make it possible to work with raw input as the text stands. In this paper, we propose a novel framework, Jumper, inspired by the cognitive process of text reading, that models text classification as a sequential decision process. Basically, Jumper is a neural system that can scan a piece of text sequentially and make classification decision at the time it chooses. Both the classification and when to make the classification are part of the decision process which are controlled by the policy net and trained with reinforcement learning to maximize the overall classification accuracy. Experimental results show that a properly trained Jumper has the following properties: (1) It can make decisions whenever the evidence is enough, therefore reducing the total text reading by 30~40% and often finding the key rationale of prediction. (2) It can achieve classification accuracy better or comparable to state-of-the-art model in several benchmark and industrial datasets.
Xianggen Liu, Lili Mou, Haotian Cui, Zhengdong Lu, Sen Song
IJCAI2
2018 Towards Neural Speaker Modeling in Multi-Party Conversation: The Task, Dataset, and Models
Lili Mou, Zhi Jin 0001
LREC2
2018 Modeling Past and Future for Neural Machine Translation
abstract
Existing neural machine translation systems do not explicitly model what has been translated and what has not during the decoding phase. To address this problem, we propose a novel mechanism that separates the source information into two parts: translated Past contents and untranslated Future contents, which are modeled by two additional recurrent layers. The Past and Future contents are fed to both the attention model and the decoder states, which provides Neural Machine Translation (NMT) systems with the knowledge of translated and untranslated contents. Experimental results show that the proposed approach significantly improves the performance in Chinese-English, German-English, and English-German translation tasks. Specifically, the proposed model outperforms the conventional coverage model in terms of both the translation quality and the alignment error rate.
Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Lili Mou, Xinyu Dai, Jiajun Chen 0001, Zhaopeng Tu
Trans. Assoc. Comput. Linguistics4
2017 Hierarchical RNN with Static Sentence-Level Attention for Text-Based Speaker Change Detection
abstract
Speaker change detection (SCD) is an important task in dialog modeling. Our paper addresses the problem of text-based SCD, which differs from existing audio-based studies and is useful in various scenarios, for example, processing dialog transcripts where speaker identities are missing (e.g., OpenSubtitle), and enhancing audio SCD with textual information. We formulate text-based SCD as a matching problem of utterances before and after a certain decision point; we propose a hierarchical recurrent neural network (RNN) with static sentence-level attention. Experimental results show that neural networks consistently achieve better performance than feature-based approaches, and that our attention-based model significantly outperforms non-attention neural networks.
Lili Mou, Zhi Jin 0001
CIKM2
2017 Coupling Distributed and Symbolic Execution for Natural Language Queries
abstract
Building neural networks to query a knowledge base (a table) with natural language is an emerging research topic in deep learning. An executor for table querying typically requires multiple steps of execution because queries may have complicated structures. In previous studies, researchers have developed either fully distributed executors or symbolic executors for table querying. A distributed executor can be trained in an end-to-end fashion, but is weak in terms of execution efficiency and explicit interpretability. A symbolic executor is efficient in execution, but is very difficult to train especially at initial stages. In this paper, we propose to couple distributed and symbolic execution for natural language queries, where the symbolic executor is pretrained with the distributed executor’s intermediate execution results in a step-by-step fashion. Experiments show that our approach significantly outperforms both distributed and symbolic executors, exhibiting high accuracy, high learning efficiency, high execution efficiency, and high interpretability.
Lili Mou, Zhengdong Lu, Hang Li 0001, Zhi Jin 0001
ICML1
2016 Convolutional Neural Networks over Tree Structures for Programming Language Processing
abstract
Programming language processing (similar to natural language processing) is a hot research topic in the field of software engineering; it has also aroused growing interest in the artificial intelligence community. However, different from a natural language sentence, a program contains rich, explicit, and complicated structural information. Hence, traditional NLP models may be inappropriate for programs. In this paper, we propose a novel tree-based convolutional neural network (TBCNN) for programming language processing, in which a convolution kernel is designed over programs' abstract syntax trees to capture structural information. TBCNN is a generic architecture for programming language processing; our experiments show its effectiveness in two different program analysis tasks: classifying programs according to functionality, and detecting code snippets of certain patterns. TBCNN outperforms baseline methods, including several neural models for NLP.
Lili Mou, Ge Li 0001, Lu Zhang 0023, Tao Wang 0080, Zhi Jin 0001
AAAI1
2016 Compressing Neural Language Models by Sparse Word Representations
abstract
Neural networks are among the state-ofthe-art techniques for language modeling.Existing neural language models typically map discrete words to distributed, dense vector representations.After information processing of the preceding context words by hidden layers, an output layer estimates the probability of the next word.Such approaches are time-and memory-intensive because of the large numbers of parameters for word embeddings and the output layer.In this paper, we propose to compress neural language models by sparse word representations.In the experiments, the number of parameters in our model increases very slowly with the growth of the vocabulary size, which is almost imperceptible.Moreover, our approach not only reduces the parameter space to a large extent, but also improves the performance in terms of the perplexity measure. 1
Yunchuan Chen, Lili Mou, Yan Xu 0013, Ge Li 0001, Zhi Jin 0001
ACL (1)2
2016 Distilling Word Embeddings: An Encoding Approach
abstract
Distilling knowledge from a well-trained cumbersome network to a small one has recently become a new research topic, as lightweight neural networks with high performance are particularly in need in various resource-restricted systems. This paper addresses the problem of distilling word embeddings for NLP tasks. We propose an encoding approach to distill task-specific knowledge from a set of high-dimensional embeddings, so that we can reduce model complexity by a large margin as well as retain high accuracy, achieving a good compromise between efficiency and performance. Experiments reveal the phenomenon that distilling knowledge from cumbersome embeddings is better than directly training neural networks with small embeddings.
Lili Mou, Ran Jia, Yan Xu 0013, Ge Li 0001, Lu Zhang 0023, Zhi Jin 0001
CIKM1
2016 Sequence to Backward and Forward Sequences: A Content-Introducing Approach to Generative Short-Text Conversation
abstract
Using neural networks to generate replies in human-computer dialogue systems is attracting increasing attention over the past few years. However, the performance is not satisfactory: the neural network tends to generate safe, universally relevant replies which carry little meaning. In this paper, we propose a content-introducing approach to neural network-based generative dialogue systems. We first use pointwise mutual information (PMI) to predict a noun as a keyword, reflecting the main gist of the reply. We then propose seq2BF, a “sequence to backward and forward sequences” model, which generates a reply containing the given keyword. Experimental results show that our approach significantly outperforms traditional sequence-to-sequence models in terms of human evaluation and the entropy measure, and that the predicted keyword can appear at an appropriate position in the reply.
Lili Mou, Yiping Song, Rui Yan 0001, Ge Li 0001, Lu Zhang 0023, Zhi Jin 0001
COLING1
2016 Improved relation classification by deep recurrent neural networks with data augmentation
abstract
Nowadays, neural networks play an important role in the task of relation classification. By designing different neural architectures, researchers have improved the performance to a large extent in comparison with traditional methods. However, existing neural networks for relation classification are usually of shallow architectures (e.g., one-layer convolutional neural networks or recurrent networks). They may fail to explore the potential representation space in different abstraction levels. In this paper, we propose deep recurrent neural networks (DRNNs) for relation classification to tackle this challenge. Further, we propose a data augmentation method by leveraging the directionality of relations. We evaluated our DRNNs on the SemEval-2010 Task 8, and achieve an F1-score of 86.1%, outperforming previous state-of-the-art recorded results.
Yan Xu 0013, Ran Jia, Lili Mou, Ge Li 0001, Yunchuan Chen, Yangyang Lu, Zhi Jin 0001
COLING3
2016 How Transferable are Neural Networks in NLP Applications?
abstract
Transfer learning is aimed to make use of valuable knowledge in a source domain to help model performance in a target domain.It is particularly important to neural networks, which are very likely to be overfitting.In some fields like image processing, many studies have shown the effectiveness of neural network-based transfer learning.For neural NLP, however, existing studies have only casually applied transfer learning, and conclusions are inconsistent.In this paper, we conduct systematic case studies and provide an illuminating picture on the transferability of neural networks in NLP. 1
Lili Mou, Rui Yan 0001, Ge Li 0001, Yan Xu 0013, Lu Zhang 0023, Zhi Jin 0001
EMNLP1
2016 StalemateBreaker: A Proactive Content-Introducing Approach to Automatic Human-Computer Conversation
Xiang Li 0065, Lili Mou, Rui Yan 0001, Ming Zhang 0004
IJCAI2
2016 Dialogue Session Segmentation by Embedding-Enhanced TextTiling
abstract
In human-computer conversation systems, the context of a userissued utterance is particularly important because it provides useful background information of the conversation.However, it is unwise to track all previous utterances in the current session as not all of them are equally important.In this paper, we address the problem of session segmentation.We propose an embedding-enhanced TextTiling approach, inspired by the observation that conversation utterances are highly noisy, and that word embeddings provide a robust way of capturing semantics.Experimental results show that our approach achieves better performance than the TextTiling, MMD approaches.
Yiping Song, Lili Mou, Rui Yan 0001, Zinan Zhu, Xiaohua Hu 0001, Ming Zhang 0004
INTERSPEECH2
2016 Context-Aware Tree-Based Convolutional Neural Networks for Natural Language Inference
Lili Mou, Ge Li 0001, Zhi Jin 0001
KSEM2
2015 Discriminative Neural Sentence Modeling by Tree-Based Convolution
abstract
This paper proposes a tree-based convolutional neural network (TBCNN) for discriminative sentence modeling.Our model leverages either constituency trees or dependency trees of sentences.The tree-based convolution process extracts sentences structural features, which are then aggregated by max pooling.Such architecture allows short propagation paths between the output layer and underlying feature detectors, enabling effective structural feature learning and extraction.We evaluate our models on two tasks: sentiment analysis and question classification.In both experiments, TBCNN outperforms previous state-of-the-art results, including existing neural networks and dedicated feature/rule engineering.We also make efforts to visualize the tree-based convolution process, shedding light on how our models work.
Lili Mou, Hao Peng 0017, Ge Li 0001, Yan Xu 0013, Lu Zhang 0023, Zhi Jin 0001
EMNLP1
2015 A Comparative Study on Regularization Strategies for Embedding-based Neural Networks
abstract
This paper aims to compare different regularization strategies to address a common phenomenon, severe overfitting, in embedding-based neural networks for NLP.We chose two widely studied neural models and tasks as our testbed.We tried several frequently applied or newly proposed regularization strategies, including penalizing weights (embeddings excluded), penalizing embeddings, reembedding words, and dropout.We also emphasized on incremental hyperparameter tuning, and combining different regularizations.The results provide a picture on tuning hyperparameters for neural NLP models.
Hao Peng 0017, Lili Mou, Ge Li 0001, Yunchuan Chen, Yangyang Lu, Zhi Jin 0001
EMNLP2
2015 Classifying Relations via Long Short Term Memory Networks along Shortest Dependency Paths
abstract
Relation classification is an important research arena in the field of natural language processing (NLP).In this paper, we present SDP-LSTM, a novel neural network to classify the relation of two entities in a sentence.Our neural architecture leverages the shortest dependency path (SDP) between two entities; multichannel recurrent neural networks, with long short term memory (LSTM) units, pick up heterogeneous information along the SDP.Our proposed model has several distinct features: (1) The shortest dependency paths retain most relevant information (to relation classification), while eliminating irrelevant words in the sentence.(2) The multichannel LSTM networks allow effective information integration from heterogeneous sources over the dependency paths.(3) A customized dropout strategy regularizes the neural network to alleviate overfitting.We test our model on the SemEval 2010 relation classification task, and achieve an F 1 -score of 83.7%, higher than competing methods in the literature.
Yan Xu 0013, Lili Mou, Ge Li 0001, Yunchuan Chen, Hao Peng 0017, Zhi Jin 0001
EMNLP2
2015 Building Program Vector Representations for Deep Learning
abstract
Deep learning has made significant breakthroughs in various fields of artificial intelligence. However, it is still virtually impossible to use deep learning to analyze programs since deep architectures cannot be trained effectively with pure back propagation. In this pioneering paper, we propose the “coding criterion” to build program vector representations, which are the premise of deep learning for program analysis. We evaluate the learned vector representations both qualitatively and quantitatively. We conclude, based on the experiments, the coding criterion is successful in building program representations. To evaluate whether deep learning is beneficial for program analysis, we feed the representations to deep neural networks, and achieve higher accuracy in the program classification task than “shallow” methods. This result confirms the feasibility of deep learning to analyze programs.
Hao Peng 0017, Lili Mou, Ge Li 0001, Yuxuan Liu 0004, Lu Zhang 0023, Zhi Jin 0001
KSEM2
2014 Verification Based on Hyponymy Hierarchical Characteristics for Web-Based Hyponymy Discovery
Lili Mou, Ge Li 0001, Zhi Jin 0001, Lu Zhang 0023
KSEM1
2014 Learning Non-Taxonomic Relations on Demand for Ontology Extension
abstract
Learning non-taxonomic relations becomes an important research topic in ontology extension. Most of the existing learning approaches are mainly based on expert crafted corpora. These approaches are normally domain-specific and the corpora acquisition is laborious and costly. On the other hand, based on the static corpora, it is not able to meet personalized needs of semantic relations discovery for various taxonomies. In this paper, we propose a novel approach for learning non-taxonomic relations on demand. For any supplied taxonomy, it can focus on the segment of the taxonomy and collect information dynamically about the taxonomic concepts by using Wikipedia as a learning source. Based on the newly generated corpus, non-taxonomic relations are acquired through three steps: a) semantic relatedness detection; b) relations extraction between concepts; and c) relations generalization within a hierarchy. The proposed approach is evaluated on three different predefined taxonomies and the experimental results show that it is effective in capturing non-taxonomic relations as needed and has good potential for the ontology extension on demand.
Yan Xu 0013, Ge Li 0001, Lili Mou, Yangyang Lu
Int. J. Softw. Eng. Knowl. Eng.3
2013 Domain Hyponymy Hierarchy Discovery by Iterative Web Searching and Inferable Semantics Based Concept Selecting
abstract
The hyponymy hierarchy is an essential part of domain knowledge, which is widely used in many applications. With the development of the Internet, the World Wide Web is now an invaluable resource of hyponymy discovering. However, acquiring domain hyponymy hierarchy from the web is still a low efficient work, because the hyponymy acquiring process is often disturbed by numerous irrelevant terms. In this paper, we propose a new iterative domain hyponymy hierarchy discovering method, where irrelevant terms can be eliminated automatically by inferable semantic information. Our approach is evaluated by the experiments in two programming-related domains. The results show that our approach works well.
Lili Mou, Ge Li 0001, Zhi Jin 0001
COMPSAC1
2013 MCT: a tool for commenting programs by multimedia comments
abstract
Program comments have always been the key to understanding code. However, typical text comments can easily become verbose or evasive. Thus sometimes code reviewers find an audio or video code narration quite helpful. In this paper, we present our tool, called MCT (Multimedia Commenting Tool), which is an integrated development environment-based tool that enables programmers to easily explain their code by voice, video and mouse movement in the form of comments. With this tool, programmers can replay the audio or video when they feel like. A demonstration video can be accessed at: http://www.youtube.com/watch?v=tHEHqZme4VE.
Yiyang Hao, Ge Li 0001, Lili Mou, Lu Zhang 0023, Zhi Jin 0001
ICSE3
2012 Discovering Domain Concepts and Hyponymy Relations by Text Relevance Classifying Based Iterative Web Searching
abstract
Domain concepts and taxonomic relationships are an essential part of a domain ontology. They are used in a number of applications, including natural language processing, information retrieval, knowledge management and so on. Nowadays, with the continuous permeation of various kinds of Internet knowledge applications, numerous new concepts are emerged and released on to the Internet. So, the Internet has become an invaluable source of new concepts for almost every possible domain of knowledge. In order to ensure the domain ontologies keep pace with fast changing knowledge, we proposed an web searching based concepts and taxonomic relationships discovering approach. By our approach, the potential concepts on the Internet, which are taxonomically related with the give seeds concepts, can be discovered autonomously and iteratively. In this paper, the approach and a corresponding application in Chinese web pages are reported in detail. The experiments show that, our approach can catch the related domain concepts precisely, meanwhile, can reject irrelevant concepts and figure out the domain knowledge border definitely.
Lili Mou, Ge Li 0001, Zhi Jin 0001, Yangyang Lu, Yiyang Hao
APSEC1