VLDB 2026 Research / reviewers in the wild / expert
Fuli Luo
dblp:220/4216
· DBLP profile ↗
29ranked-venue papers
9as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree SearchabstractLean is an advanced proof assistant designed to facilitate formal theorem proving by providing a variety of interactive feedback. In this paper, we explore methodologies to leverage proof assistant feedback to augment the capabilities of large language models in constructing formal proofs. First, we deploy online reinforcement learning using Lean verification outcomes as the reward signal to improve the proof completion policy. This straightforward approach shows great promise in enhancing the model's alignment with the formal verification system. In addition, we propose RMaxTS, a variant of Monte-Carlo tree search that employs an intrinsic-reward-driven exploration strategy to generate diverse proof paths. The tree structure is organized to represent the transitions of intermediate tactic states, extracted from the compilation messages given by Lean's tactic mode. The intrinsic reward is constructed to incentivize the discovery of novel tactic states, which helps to to mitigate the sparse-reward problem inherent in proof search. These techniques lead to a more efficient planning scheme for formal proof generation, achieving new state-of-the-art results on both miniF2F and ProofNet benchmarks. Huajian Xin, Z. Z. Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Qiushi Du, Qihao Zhu, Dejian Yang, Zhibin Gou, Z. F. Wu, Fuli Luo, Chong Ruan |
ICLR | 17 |
| 2024 | DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsabstractDamai Dai, Chengqi Deng, Chenggang Zhao, R.x. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y.k. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, Wenfeng Liang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, Wenfeng Liang |
ACL (1) | 14 |
| 2024 | Comparative Analysis of Two Angle Normalization Approaches for SAR Backscatter: Simulation and Satellite Observation-Based Evaluation in Soil Moisture RetrievalabstractLocal incidence angle (LIA) normalization is an important method to improve the accuracy of active microwave remote sensing-based soil moisture retrieval in mountainous areas. In this study, the differences between two commonly used synthetic aperture radar (SAR) backscatter LIA normalization methods, cosine-based and liner-based, were compared using simulated and Sentinel-1 SAR data. The influence of two backscatter normalization methods on soil moisture retrieval in the dual-temporal dual-channel (DTDC) algorithm is analyzed. The results show that the difference between the normalized backscatter by the two methods is less than 0.3 dB in most cases. Despite the simplicity of the methods, both angle normalization techniques can rectify variations in backscattering induced by the LIA effect and improve the accuracy of soil moisture retrieval. Dong Fan, Fuli Luo, Jiliu Hu, Junxuan Liu, Bo-Hui Tang |
IGARSS | 3 |
| 2024 | Quantification of Correlation Between Geometric Characteristics for The Hexagonal Discrete Global Grid SystemsabstractAlthough the discrete global grid systems can serve as a robust digital Earth framework, the geometric complexity of grid cells increase with the level of divisions due to the inherent topology of the spherical surface. Research on the correlation of geometric characteristics is essential for describing these complexities. Recently, researches have paid more attention to global triangular and diamond grid systems, but have not yet investigated the correlation relationships among multiple geometric features of global hexagonal grid systems. This constrains significantly the application expansion in high-precision spatial data processing and analysis, Earth system modeling, and complex phenomena simulation. In this study, we selected seven indicators from the Goodchild Criteria to quantify the geometric features, calculated and identified features’ correlations at various resolutions, and constructed a correlation model of multiple quantitative indicators. Results with the ISEA4H grid model as the subject indicate a strong correlation among intrinsic geometric features, while topological features exhibit notably high independence. Our study reveals the correlation among multiple geometric features of hexagonal discrete global grid systems, laying a foundation for their comprehensive quality evaluation system. Fuli Luo |
IGARSS | 1 |
| 2023 | Knowledgeable Salient Span Mask for Enhancing Language Models as Knowledge Base
Cunxiang Wang, Fuli Luo, Yanyang Li, Runxin Xu, Fei Huang 0002, Yue Zhang 0004 |
NLPCC (2) | 2 |
| 2023 | Making Pre-trained Language Models End-to-end Few-shot Learners with Contrastive Prompt TuningabstractPre-trained Language Models (PLMs) have achieved remarkable performance for various language understanding tasks in IR systems, which require the fine-tuning process based on labeled training data. For low-resource scenarios, prompt-based learning for PLMs exploits prompts as task guidance and turns downstream tasks into masked language problems for effective few-shot fine-tuning. In most existing approaches, the high performance of prompt-based learning heavily relies on handcrafted prompts and verbalizers, which may limit the application of such approaches in real-world scenarios. To solve this issue, we present CP-Tuning, an end-to-end Contrastive Prompt Tuning framework for fine-tuning PLMs without any manual engineering of task-specific prompts and verbalizers. It is integrated with the task-invariant continuous prompt encoding technique with fully trainable prompt parameters. We further propose the pair-wise cost-sensitive contrastive learning procedure to optimize the model in order to achieve verbalizer-free class mapping and enhance the task-invariance of prompts. It explicitly learns to distinguish different classes and makes the decision boundary smoother by assigning different costs to easy and hard cases. Experiments over a variety of language understanding tasks and different PLMs show that CP-Tuning outperforms state-of-the-art methods. Ziyun Xu, Chengyu Wang 0001, Minghui Qiu, Fuli Luo, Runxin Xu, Songfang Huang, Jun Huang 0007 |
WSDM | 4 |
| 2022 | From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model CompressionabstractPre-trained Language Models (PLMs) have achieved great success in various Natural Language Processing (NLP) tasks under the pre-training and fine-tuning paradigm. With large quantities of parameters, PLMs are computation-intensive and resource-hungry. Hence, model pruning has been introduced to compress large-scale PLMs. However, most prior approaches only consider task-specific knowledge towards downstream tasks, but ignore the essential task-agnostic knowledge during pruning, which may cause catastrophic forgetting problem and lead to poor generalization ability. To maintain both task-agnostic and task-specific knowledge in our pruned model, we propose ContrAstive Pruning (CAP) under the paradigm of pre-training and fine-tuning. It is designed as a general framework, compatible with both structured and unstructured pruning. Unified in contrastive learn- ing, CAP enables the pruned model to learn from the pre-trained model for task-agnostic knowledge, and fine-tuned model for task-specific knowledge. Besides, to better retain the performance of the pruned model, the snapshots (i.e., the intermediate models at each pruning iteration) also serve as effective supervisions for pruning. Our extensive experiments show that adopting CAP consistently yields significant improvements, especially in extremely high sparsity scenarios. With only 3% model parameters reserved (i.e., 97% sparsity), CAP successfully achieves 99.2% and 96.3% of the original BERT performance in QQP and MNLI tasks. In addition, our probing experiments demonstrate that the model pruned by CAP tends to achieve better generalization ability. Runxin Xu, Fuli Luo, Chengyu Wang 0001, Baobao Chang, Jun Huang 0007, Songfang Huang, Fei Huang 0002 |
AAAI | 2 |
| 2022 | Probing Structured Pruning on Multilingual Pre-trained Models: Settings, Algorithms, and EfficiencyabstractStructured pruning has been extensively studied on monolingual pre-trained language models and is yet to be fully evaluated on their multilingual counterparts.This work investigates three aspects of structured pruning on multilingual pre-trained language models: settings, algorithms, and efficiency.Experiments on nine downstream tasks show several counterintuitive phenomena: for settings, individually pruning for each language does not induce a better result; for algorithms, the simplest method performs the best; for efficiency, a fast model does not imply that it is also small.To facilitate the comparison on all sparsity levels, we present Dynamic Sparsification, a simple approach that allows training the model once and adapting to different model sizes at inference.We hope this work fills the gap in the study of structured pruning on multilingual pre-trained models and sheds light on future research. Yanyang Li, Fuli Luo, Runxin Xu, Songfang Huang, Fei Huang 0002, Liwei Wang 0009 |
ACL (1) | 2 |
| 2022 | Parameter-Efficient Sparsity for Large Language Models Fine-TuningabstractWith the dramatically increased number of parameters in language models, sparsity methods have received ever-increasing research focus to compress and accelerate the models. While most research focuses on how to accurately retain appropriate weights while maintaining the performance of the compressed model, there are challenges in the computational overhead and memory footprint of sparse training when compressing large-scale language models. To address this problem, we propose a Parameter-efficient Sparse Training (PST) method to reduce the number of trainable parameters during sparse-aware training in downstream tasks. Specifically, we first combine the data-free and data-driven criteria to efficiently and accurately measure the importance of weights. Then we investigate the intrinsic redundancy of data-driven weight importance and derive two obvious characteristics i.e. low-rankness and structuredness. Based on that, two groups of small matrices are introduced to compute the data-driven importance of weights, instead of using the original large importance score matrix, which therefore makes the sparse training resource-efficient and parameter-efficient. Experiments with diverse networks (i.e. BERT, RoBERTa and GPT-2) on dozens of datasets demonstrate PST performs on par or better than previous sparsity methods, despite only training a small number of parameters. For instance, compared with previous sparsity methods, our PST only requires 1.5% trainable parameters to achieve comparable performance on BERT. Fuli Luo, Chuanqi Tan, Songfang Huang |
IJCAI | 2 |
| 2021 | VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and GenerationabstractFuli Luo, Wei Wang, Jiahao Liu, Yijia Liu, Bin Bi, Songfang Huang, Fei Huang, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fuli Luo, Wei Wang 0225, Bin Bi, Songfang Huang, Fei Huang 0002, Luo Si |
ACL/IJCNLP (1) | 1 |
| 2021 | Rethinking Denoised Auto-Encoding in Language Pre-TrainingabstractPre-trained self-supervised models such as BERT have achieved striking success in learning sequence representations, especially for natural language processing.These models typically corrupt the given sequences with certain types of noise, such as masking, shuffling, or substitution, and then try to recover the original input.However, such pre-training approaches are prone to learning representations that are covariant with the noise, leading to the discrepancy between the pre-training and finetuning stage.To remedy this, we present Con-trAstive Pre-Training (CAPT) to learn noise invariant sequence representations.The proposed CAPT encourages the consistency between representations of the original sequence and its corrupted version via unsupervised instance-wise training signals.In this way, it not only alleviates the pretrain-finetune discrepancy induced by the noise of pre-training, but also aids the pre-trained model in better capturing global semantics of the input via more effective sentence-level supervision.Different from most prior work that focuses on a particular modality, comprehensive empirical evidence on 11 natural language understanding and cross-modal tasks illustrates that CAPT is applicable for both language and vision-language tasks, and obtains surprisingly consistent improvement, including 0.6% absolute gain on GLUE benchmarks and 0.8% absolute increment on NLVR 2 . * Equal Contribution.Models Noise types BERT (Devlin et al., 2019) Mask tokens SpanBERT (Joshi et al., 2019) Mask spans RoBERTa (Liu et al., 2019) Mask token XLNet (Yang et al., 2019) Shuffle token ELECTRA (Clark et al., 2019) Replace tokens StructBERT (Wang et al., 2019b) Mask + Shuffle tokens BART (Lewis et al., 2019) Mask + Shuffle + Replace.UNITER (Chen et al., 2019) Mask tokens/regions LXMERT (Tan and Bansal, 2019) Mask tokens/regions Fuli Luo, Xuancheng Ren, Xu Sun 0001, Songfang Huang, Fei Huang 0002 |
EMNLP (1) | 1 |
| 2021 | Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningabstractRecent pretrained language models extend from millions to billions of parameters.Thus the need to fine-tune an extremely large pretrained model with a limited training corpus arises in various downstream tasks.In this paper, we propose a straightforward yet effective fine-tuning technique, CHILD-TUNING, which updates a subset of parameters (called child network) of large pretrained models via strategically masking out the gradients of the non-child network during the backward process.Experiments on various downstream tasks in GLUE benchmark show that CHILD-TUNING consistently outperforms the vanilla fine-tuning by 1.5 ∼ 8.6 average score among four different pretrained models, and surpasses the prior fine-tuning techniques by 0.6 ∼ 1.3 points.Furthermore, empirical results on domain transfer and task transfer show that CHILD-TUNING can obtain better generalization performance by large margins. Runxin Xu, Fuli Luo, Chuanqi Tan, Baobao Chang, Songfang Huang, Fei Huang 0002 |
EMNLP (1) | 2 |
| 2019 | WSD-GAN: Word Sense Disambiguation Using Generative Adversarial NetworksabstractWord Sense Disambiguation (WSD), as a tough task in Natural Language Processing (NLP), aims to identify the correct sense of an ambiguous word in a given context. There are two mainstreams in WSD. Supervised methods mainly utilize labeled context to train a classifier which generates the right probability distribution of word senses. Meanwhile knowledge-based (unsupervised) methods which focus on glosses (word sense definitions) always calculate the similarity of context-gloss pair as score to find out the right word sense. In this paper, we propose a generative adversarial framework WSD-GAN which combines two mainstream methods in WSD. The generative model, based on supervised methods, tries to generate a probability distribution over the word senses. Meanwhile the discriminative model, based on knowledge-based methods, focuses on predicting the relevancy of the context-gloss pairs and identifies the correct pairs over the others. Furthermore, in order to optimize both two models, we leverage policy gradient to enhance the performances of the two models mutually. Our experimental results show that WSD-GAN achieves competitive results on several English all-words WSD datasets. Fuli Luo, Yutong Tan, Wenxin Zeng, Zhifang Sui |
AAAI | 2 |
| 2019 | Hierarchical Encoder with Auxiliary Supervision for Neural Table-to-Text Generation: Learning Better Representation for TablesabstractGenerating natural language descriptions for the structured tables which consist of multiple attribute-value tuples is a convenient way to help people to understand the tables. Most neural table-to-text models are based on the encoder-decoder framework. However, it is hard for a vanilla encoder to learn the accurate semantic representation of a complex table. The challenges are two-fold: firstly, the table-to-text datasets often contain large number of attributes across different domains, thus it is hard for the encoder to incorporate these heterogeneous resources. Secondly, the single encoder also has difficulties in modeling the complex attribute-value structure of the tables. To this end, we first propose a two-level hierarchical encoder with coarse-to-fine attention to handle the attribute-value structure of the tables. Furthermore, to capture the accurate semantic representations of the tables, we propose 3 joint tasks apart from the prime encoder-decoder learning, namely auxiliary sequence labeling task, text autoencoder and multi-labeling classification, as the auxiliary supervisions for the table encoder. We test our models on the widely used dataset WIKIBIO which contains Wikipedia infoboxes and related descriptions. The dataset contains complex tables as well as large number of attributes across different domains. We achieve the state-of-the-art performance on both automatic and human evaluation metrics. Tianyu Liu 0001, Fuli Luo, Qiaolin Xia, Shuming Ma, Baobao Chang, Zhifang Sui |
AAAI | 2 |
| 2019 | Towards Comprehensive Description Generation from Factual Attribute-value TablesabstractThe comprehensive descriptions for factual attribute-value tables, which should be accurate, informative and loyal, can be very helpful for end users to understand the structured data in this form.However previous neural generators might suffer from key attributes missing, less informative and groundless information problems, which impede the generation of high-quality comprehensive descriptions for tables.To relieve these problems, we first propose force attention (FA) method to encourage the generator to pay more attention to the uncovered attributes to avoid potential key attributes missing.Furthermore, we propose reinforcement learning for information richness to generate more informative as well as more loyal descriptions for tables.In our experiments, we utilize the widely used WIKIBIO dataset as a benchmark.Additionally we create WB-filter based on WIKIBIO to test our model in the simulated user-oriented scenarios, in which the generated descriptions should accord with particular user interests.Experimental results show that our model outperforms the state-of-the-art baselines on both automatic and human evaluation. Tianyu Liu 0001, Fuli Luo, Wei Wu 0044, Baobao Chang, Zhifang Sui |
ACL (1) | 2 |
| 2019 | Learning to Control the Fine-grained Sentiment for Story Ending GenerationabstractAutomatic story ending generation is an interesting and challenging task in natural language generation.Previous studies are mainly limited to generate coherent, reasonable and diversified story endings, and few works focus on controlling the sentiment of story endings.This paper focuses on generating a story ending which meets the given fine-grained sentiment intensity.There are two major challenges to this task.First is the lack of story corpus which has fine-grained sentiment labels.Second is the difficulty of explicitly controlling sentiment intensity when generating endings.Therefore, we propose a generic and novel framework which consists of a sentiment analyzer and a sentimental generator, respectively addressing the two challenges.The sentiment analyzer adopts a series of methods to acquire sentiment intensities of the story dataset.The sentimental generator introduces the sentiment intensity into decoder via a Gaussian Kernel Layer to control the sentiment of the output.To the best of our knowledge, this is the first endeavor to control the fine-grained sentiment for story ending generation without manually annotating sentiment labels.Experiments show that our proposed framework can generate story endings which are not only more coherent and fluent but also able to meet the given sentiment intensity better. 1 Fuli Luo, Damai Dai, Tianyu Liu 0001, Baobao Chang, Zhifang Sui, Xu Sun 0001 |
ACL (1) | 1 |
| 2019 | Towards Fine-grained Text Sentiment TransferabstractIn this paper, we focus on the task of finegrained text sentiment transfer (FGST).This task aims to revise an input sequence to satisfy a given sentiment intensity, while preserving the original semantic content.Different from conventional sentiment transfer task that only reverses the sentiment polarity (positive/negative) of text, the FTST task requires more nuanced and fine-grained control of sentiment.To remedy this, we propose a novel Seq2SentiSeq model.Specifically, the numeric sentiment intensity value is incorporated into the decoder via a Gaussian kernel layer to finely control the sentiment intensity of the output.Moreover, to tackle the problem of lacking parallel data, we propose a cycle reinforcement learning algorithm to guide the model training.In this framework, the elaborately designed rewards can balance both sentiment transformation and content preservation, while not requiring any ground truth output.Experimental results show that our approach can outperform existing methods by a large margin in both automatic evaluation and human evaluation.Our code and data, including outputs of all baselines and our model are available at https://github.com/luofuli/ Fine-grained-Sentiment-Transfer. 1 Fuli Luo, Peng Li 0030, Jie Zhou 0016, Yutong Tan, Baobao Chang, Zhifang Sui, Xu Sun 0001 |
ACL (1) | 1 |
| 2019 | A Hierarchical Reinforced Sequence Operation Method for Unsupervised Text Style TransferabstractUnsupervised text style transfer aims to alter text styles while preserving the content, without aligned data for supervision.Existing seq2seq methods face three challenges: 1) the transfer is weakly interpretable, 2) generated outputs struggle in content preservation, and 3) the trade-off between content and style is intractable.To address these challenges, we propose a hierarchical reinforced sequence operation method, named Point-Then-Operate (PTO), which consists of a high-level agent that proposes operation positions and a lowlevel agent that alters the sentence.We provide comprehensive training objectives to control the fluency, style, and content of the outputs and a mask-based inference algorithm that allows for multi-step revision based on the single-step trained agents.Experimental results on two text style transfer datasets show that our method significantly outperforms recent methods and effectively addresses the aforementioned challenges. 1 Chen Wu 0005, Xuancheng Ren, Fuli Luo, Xu Sun 0001 |
ACL (1) | 3 |
| 2019 | MAAM: A Morphology-Aware Alignment Model for Unsupervised Bilingual Lexicon InductionabstractThe task of unsupervised bilingual lexicon induction (UBLI) aims to induce word translations from monolingual corpora in two languages.Previous work has shown that morphological variation is an intractable challenge for the UBLI task, where the induced translation in failure case is usually morphologically related to the correct translation.To tackle this challenge, we propose a morphology-aware alignment model for the UBLI task.The proposed model aims to alleviate the adverse effect of morphological variation by introducing grammatical information learned by the pre-trained denoising language model.Results show that our approach can substantially outperform several state-of-the-art unsupervised systems, and even achieves competitive performance compared to supervised methods. Fuli Luo, Tianyu Liu 0001, Xu Sun 0001 |
ACL (1) | 2 |
| 2019 | Enhancing Topic-to-Essay Generation with External Commonsense KnowledgeabstractAutomatic topic-to-essay generation is a challenging task since it requires generating novel, diverse, and topic-consistent paragraph-level text with a set of topics as input.Previous work tends to perform essay generation based solely on the given topics while ignoring massive commonsense knowledge.However, this commonsense knowledge provides additional background information, which can help to generate essays that are more novel and diverse.Towards filling this gap, we propose to integrate commonsense from the external knowledge base into the generator through dynamic memory mechanism.Besides, the adversarial training based on a multi-label discriminator is employed to further improve topic-consistency.We also develop a series of automatic evaluation metrics to comprehensively assess the quality of the generated essay.Experiments show that with external commonsense knowledge and adversarial training, the generated essays are more novel, diverse, and topic-consistent than existing methods in terms of both automatic and human evaluation. Lei Li 0039, Fuli Luo, Tianyu Liu 0001, Xu Sun 0001 |
ACL (1) | 3 |
| 2019 | A Deep Reinforced Sequence-to-Set Model for Multi-Label ClassificationabstractMulti-label classification (MLC) aims to predict a set of labels for a given instance.Based on a pre-defined label order, the sequence-tosequence (Seq2Seq) model trained via maximum likelihood estimation method has been successfully applied to the MLC task and shows powerful ability to capture high-order correlations between labels.However, the output labels are essentially an unordered set rather than an ordered sequence.This inconsistency tends to result in some intractable problems, e.g., sensitivity to the label order.To remedy this, we propose a simple but effective sequence-to-set model.The proposed model is trained via reinforcement learning, where reward feedback is designed to be independent of the label order.In this way, we can reduce the dependence of the model on the label order, as well as capture high-order correlations between labels.Extensive experiments show that our approach can substantially outperform competitive baselines, as well as effectively reduce the sensitivity to the label order. 1 Fuli Luo, Shuming Ma, Junyang Lin, Xu Sun 0001 |
ACL (1) | 2 |
| 2019 | Cross-Modal Commentator: Automatic Machine Commenting Based on Cross-Modal InformationabstractAutomatic commenting of online articles can provide additional opinions and facts to the reader, which improves user experience and engagement on social media platforms.Previous work focuses on automatic commenting based solely on textual content.However, in real-scenarios, online articles usually contain multiple modal contents.For instance, graphic news contains plenty of images in addition to text.Contents other than text are also vital because they are not only more attractive to the reader but also may provide critical information.To remedy this, we propose a new task: cross-model automatic commenting (CMAC), which aims to make comments by integrating multiple modal contents.We construct a largescale dataset for this task and explore several representative methods.Going a step further, an effective co-attention model is presented to capture the dependency between textual and visual information.Evaluation results show that our proposed model can achieve better performance than competitive baselines.1 Zhihan Zhang 0001, Fuli Luo, Lei Li 0039, Chengyang Huang, Xu Sun 0001 |
ACL (1) | 3 |
| 2019 | Pun-GAN: Generative Adversarial Network for Pun GenerationabstractFuli Luo, Shunyao Li, Pengcheng Yang, Lei Li, Baobao Chang, Zhifang Sui, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Fuli Luo, Shunyao Li, Lei Li 0039, Baobao Chang, Zhifang Sui, Xu Sun 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2019 | A Dual Reinforcement Learning Framework for Unsupervised Text Style TransferabstractUnsupervised text style transfer aims to transfer the underlying style of text but keep its main content unchanged without parallel data. Most existing methods typically follow two steps: first separating the content from the original style, and then fusing the content with the desired style. However, the separation in the first step is challenging because the content and style interact in subtle ways in natural language. Therefore, in this paper, we propose a dual reinforcement learning framework to directly transfer the style of the text via a one-step mapping model, without any separation of content and style. Specifically, we consider the learning of the source-to-target and target-to-source mappings as a dual task, and two rewards are designed based on such a dual structure to reflect the style accuracy and content preservation, respectively. In this way, the two one-step mapping models can be trained via reinforcement learning, without any use of parallel data. Automatic evaluations show that our model outperforms the state-of-the-art systems by a large margin, especially with more than 10 BLEU points improvement averaged on two benchmark datasets. Human evaluations also validate the effectiveness of our model in terms of style accuracy, content preservation and fluency. Our code and data, including outputs of all baselines and our model are available at https://github.com/luofuli/DualRL. Fuli Luo, Peng Li 0030, Jie Zhou 0016, Baobao Chang, Xu Sun 0001, Zhifang Sui |
IJCAI | 1 |
| 2019 | Knowledgeable Storyteller: A Commonsense-Driven Generative Model for Visual StorytellingabstractThe visual storytelling (VST) task aims at generating a reasonable and coherent paragraph-level story with the image stream as input. Different from caption that is a direct and literal description of image content, the story in the VST task tends to contain plenty of imaginary concepts that do not appear in the image. This requires the AI agent to reason and associate with the imaginary concepts based on implicit commonsense knowledge to generate a reasonable story describing the image stream. Therefore, in this work, we present a commonsense-driven generative model, which aims to introduce crucial commonsense from the external knowledge base for visual storytelling. Our approach first extracts a set of candidate knowledge graphs from the knowledge base. Then, an elaborately designed vision-aware directional encoding schema is adopted to effectively integrate the most informative commonsense. Besides, we strive to maximize the semantic similarity within the output during decoding to enhance the coherence of the generated text. Results show that our approach can outperform the state-of-the-art systems by a large margin, which achieves a 29\% relative improvement of CIDEr score. With additional commonsense and semantic-relevance based objective, the generated stories are more diverse and coherent. Fuli Luo, Lei Li 0039, Zhiyi Yin, Xiaodong He 0001, Xu Sun 0001 |
IJCAI | 2 |
| 2019 | Learning Unsupervised Word Mapping via Maximum Mean Discrepancy
Fuli Luo, Shuangzhi Wu, Jingjing Xu 0001, Dongdong Zhang 0001 |
NLPCC (1) | 2 |
| 2019 | Cross-language document summarization via extraction and ranking of multiple summaries
Xiaojun Wan 0001, Fuli Luo, Songfang Huang, Jin-ge Yao |
Knowl. Inf. Syst. | 2 |
| 2018 | Incorporating Glosses into Neural Word Sense DisambiguationabstractWord Sense Disambiguation (WSD) aims to identify the correct meaning of polysemous words in the particular context.Lexical resources like WordNet which are proved to be of great help for WSD in the knowledge-based methods.However, previous neural networks for WSD always rely on massive labeled data (context), ignoring lexical resources like glosses (sense definitions).In this paper, we integrate the context and glosses of the target word into a unified framework in order to make full use of both labeled data and lexical knowledge.Therefore, we propose GAS: a gloss-augmented WSD neural network which jointly encodes the context and glosses of the target word.GAS models the semantic relationship between the context and the gloss in an improved memory network framework, which breaks the barriers of the previous supervised methods and knowledge-based methods.We further extend the original gloss of word sense via its semantic relations in WordNet to enrich the gloss information.The experimental results show that our model outperforms the state-of-theart systems on several English all-words WSD datasets. Fuli Luo, Tianyu Liu 0001, Qiaolin Xia, Baobao Chang, Zhifang Sui |
ACL (1) | 1 |
| 2018 | Leveraging Gloss Knowledge in Neural Word Sense Disambiguation by Hierarchical Co-AttentionabstractThe goal of Word Sense Disambiguation (WSD) is to identify the correct meaning of a word in the particular context.Traditional supervised methods only use labeled data (context), while missing rich lexical knowledge such as the gloss which defines the meaning of a word sense.Recent studies have shown that incorporating glosses into neural networks for WSD has made significant improvement.However, the previous models usually build the context representation and gloss representation separately.In this paper, we find that the learning for the context and gloss representation can benefit from each other.Gloss can help to highlight the important words in the context, thus building a better context representation.Context can also help to locate the key words in the gloss of the correct word sense.Therefore, we introduce a co-attention mechanism to generate co-dependent representations for the context and gloss.Furthermore, in order to capture both word-level and sentence-level information, we extend the attention mechanism in a hierarchical fashion.Experimental results show that our model achieves the state-of-the-art results on several standard English all-words WSD test datasets. Fuli Luo, Tianyu Liu 0001, Zexue He, Qiaolin Xia, Zhifang Sui, Baobao Chang |
EMNLP | 1 |