EDBT 2026 Demo / reviewers in the wild / expert
Dayiheng Liu
dblp:189/4488
· DBLP profile ↗
56ranked-venue papers
10as first author
40since 2021 · last 2026
0000-0002-8755-8941ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 10 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Controllable LLM Reasoning via Sparse Autoencoder-Based SteeringabstractYi Fang, Wenjie Wang, Mingfeng Xue, Boyi Deng, Fengli Xu, Dayiheng Liu, Fuli Feng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yi Fang 0010, Wenjie Wang 0007, Mingfeng Xue, Boyi Deng, Fengli Xu, Dayiheng Liu, Fuli Feng |
ACL (1) | 6 |
| 2026 | MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning EvaluationabstractXiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoyuan Li 0001, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang 0007, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu |
ACL (1) | 9 |
| 2026 | PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal PracticeabstractYuzhen Shi, Huanghai Liu, Yiran HU, Song Gaojie, Xu Xinran, Yubo Ma, Tianyi Tang, Li Zhang, Qingjing Chen, Feng Di, Wenbo Lv, Weiheng Wu, Kexin Yang, Sen Yang, Wei Wang, Rongyao Shi, Qiu Yuanyang, Yuemeng Qi, Zhang Jingwen, Sui Xiaoyu, Yifan Chen, Zhang Yi, An Yang, Bowen Yu, Dayiheng Liu, Junyang Lin, Weixing Shen, Bing Zhao, Charles L. A. Clarke, HU Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song, Xinran Xu, Yubo Ma, Qingjing Chen, Di Feng, Wenbo Lv, Weiheng Wu, Kexin Yang 0002, Wei Wang 0225, Rongyao Shi, Yuanyang Qiu, Yuemeng Qi, Xiaoyu Sui, Yi Zhang 0101, An Yang, Bowen Yu 0002, Dayiheng Liu, Junyang Lin, Weixing Shen, Charles L. A. Clarke, Hu Wei |
ACL (1) | 25 |
| 2025 | Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert ModelsabstractZihan Qiu, Zeyu Huang, Bo Zheng, Kaiyue Wen, Zekun Wang, Rui Men, Ivan Titov, Dayiheng Liu, Jingren Zhou, Junyang Lin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zihan Qiu, Bo Zheng 0007, Kaiyue Wen, Rui Men, Ivan Titov 0001, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin |
ACL (1) | 8 |
| 2025 | ProcessBench: Identifying Process Errors in Mathematical ReasoningabstractChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, Junyang Lin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chujie Zheng, Zhenru Zhang, Runji Lin, Keming Lu, Bowen Yu 0002, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin |
ACL (1) | 7 |
| 2025 | START: Self-taught Reasoner with ToolsabstractChengpeng Li, Mingfeng Xue, Zhenru Zhang, Jiaxi Yang, Beichen Zhang, Bowen Yu, Binyuan Hui, Junyang Lin, Xiang Wang, Dayiheng Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chengpeng Li 0001, Mingfeng Xue, Zhenru Zhang, Jiaxi Yang 0004, Bowen Yu 0002, Binyuan Hui, Junyang Lin, Xiang Wang 0010, Dayiheng Liu |
EMNLP | 10 |
| 2025 | P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMsabstractYidan Zhang, Yu Wan, Boyi Deng, Baosong Yang, Hao-Ran Wei, Fei Huang, Bowen Yu, Dayiheng Liu, Junyang Lin, Fei Huang, Jingren Zhou. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yidan Zhang 0004, Yu Wan 0004, Boyi Deng, Baosong Yang, Fei Huang 0002, Bowen Yu 0002, Dayiheng Liu, Junyang Lin, Fei Huang 0005, Jingren Zhou 0001 |
EMNLP | 8 |
| 2025 | NOVA-63: Native Omni-lingual Versatile Assessments of 63 DisciplinesabstractThe multilingual capabilities of large language models (LLMs) have attracted considerable attention over the past decade. Assessing the accuracy with which LLMs provide answers in multilingual contexts is essential for determining their level of multilingual proficiency. Nevertheless, existing multilingual benchmarks generally reveal severe drawbacks, such as overly translated content (translationese), the absence of difficulty control, constrained diversity, and disciplinary imbalance, making the benchmarking process unreliable and showing low convincingness. To alleviate those shortcomings, we introduce NOVA-63 (Native Omni-lingual Versatile Assessments of 63 Disciplines), a comprehensive, difficult multilingual benchmark featuring 93,536 questions sourced from native speakers across 14 languages and 63 academic disciplines. Leveraging a robust pipeline that integrates LLM-assisted formatting, expert quality verification, and multi-level difficulty screening, NOVA-63 is balanced on disciplines with consistent difficulty standards while maintaining authentic linguistic elements. Extensive experimentation with current LLMs has shown significant insights into cross-lingual consistency among language families, and exposed notable disparities in models’ capabilities across various disciplines. This work provides valuable benchmarking data for the future development of multilingual models. Furthermore, our findings underscore the importance of moving beyond overall scores and instead conducting fine-grained analyses of model performance. Kexin Yang 0002, Yu Wan 0004, Muyang Ye, Baosong Yang, Junyang Lin, Dayiheng Liu |
EMNLP | 8 |
| 2025 | DataMan: Data Manager for Pre-training Large Language ModelsabstractThe performance emergence of large language models (LLMs) driven by data scaling laws makes the selection of pre-training data increasingly important.
However, existing methods rely on limited heuristics and human intuition, lacking comprehensive and clear guidelines.
To address this, we are inspired by *``reverse thinking''* -- prompting LLMs to self-identify which criteria benefit its performance.
As its pre-training capabilities are related to perplexity (PPL), we derive 14 quality criteria from the causes of text perplexity anomalies and introduce 15 common application domains to support domain mixing.
In this paper, we train a **Data** **Man**ager (**DataMan**) to learn quality ratings and domain recognition from pointwise rating, and use it to annotate a 447B token pre-training corpus with 14 quality ratings and domain type.
Our experiments validate our approach, using DataMan to select 30B tokens to train a 1.3B-parameter language model, demonstrating significant improvements in in-context learning (ICL), perplexity, and instruction-following ability over the state-of-the-art baseline.
The best-performing model, based on the *Overall Score l=5* surpasses a model trained with 50% more data using uniform sampling.
We continue pre-training with high-rated, domain-specific data annotated by DataMan to enhance domain-specific ICL performance and thus verify DataMan's domain mixing ability.
Our findings emphasize the importance of quality ranking, the complementary nature of quality criteria, and their low correlation with perplexity, analyzing misalignment between PPL and ICL performance.
We also thoroughly analyzed our pre-training dataset, examining its composition, the distribution of quality ratings, and the original document sources. Ru Peng, Kexin Yang 0002, Yawen Zeng, Junyang Lin, Dayiheng Liu, Junbo Zhao 0002 |
ICLR | 5 |
| 2025 | Parallel Scaling Law for Language ModelsabstractIt is commonly believed that scaling language models should commit a significant space or time cost, by increasing the parameters (parameter scaling) or output tokens (inference-time scaling). We introduce another and more inference-efficient scaling paradigm: increasing the model's parallel computation during both training and inference time. We apply $P$ diverse and learnable transformations to the input, execute forward passes of the model in parallel, and dynamically aggregate the $P$ outputs. This method, namely parallel scaling (ParScale), scales parallel computation by reusing existing parameters and can be applied to any model structure, optimization procedure, data, or task. We theoretically propose a new scaling law and validate it through large-scale pre-training, which shows that a model with $P$ parallel streams is similar to scaling the parameters by $\mathcal O(\log P)$ while showing superior inference efficiency. For example, ParScale can use up to 22$\times$ less memory increase and 6$\times$ less latency increase compared to parameter scaling that achieves the same performance improvement. It can also recycle an off-the-shelf pre-trained model into a parallelly scaled one by post-training on a small amount of tokens, further reducing the training budget. The new scaling law we discovered potentially facilitates the deployment of more powerful models in low-resource scenarios, and provides an alternative perspective for the role of computation in machine learning. Our code and 67 trained model checkpoints are publicly available at https://github.com/QwenLM/ParScale and https://huggingface.co/ParScale. Mouxiang Chen, Binyuan Hui, Zeyu Cui, Jiaxi Yang 0004, Dayiheng Liu, Jianling Sun, Junyang Lin, Zhongxin Liu 0002 |
NeurIPS | 5 |
| 2025 | Chain of Execution Supervision Promotes General Reasoning in Large Language ModelsabstractBuilding robust and general reasoning ability is a central goal in the development of large language models (LLMs). Recent efforts increasingly turn to code as a rich training source, given its inherent logical structure and diverse reasoning paradigms—such as divide-and-conquer, topological ordering, and enumeration. However, reasoning in code is often expressed implicitly and entangled with syntactic or implementation noise, making direct training on raw code suboptimal. To address this, we introduce TraceMind, a large-scale corpus of 2.6 million samples that transforms code execution into explicit, step-by-step chain-of-thought style rationales, which we call Chain of Execution (CoE).
The corpus spans domains including mathematics, classical algorithms and algorithmic competition, and is enriched with variable-tracing questions and code rewritings to enhance logical granularity and code diversity.
We evaluate Tracepile using three training setups—continue-pretraining, instruction tuning after pretraining, and two-stage finetuning. Experiments across four base models (LLaMA 3, LLaMA 3.1, Qwen-2.5, and Qwen-2.5 Coder) and 20 benchmarks covering math, code, logic, and algorithms demonstrate consistent improvements. Notably, Tracepile boosts LLaMA3-8B by 9.2\% on average across nine math datasets and delivers clear gains on LiveCodeBench, CRUX, and Zebra Logic under two-stage finetuning. Keqin Bao, Junyang Lin, Dayiheng Liu |
NeurIPS | 5 |
| 2025 | Teaching Language Models to Reason with ToolsabstractLarge reasoning models (LRMs) like OpenAI-o1 have shown impressive capabilities in natural language reasoning. However, these models frequently demonstrate inefficiencies or inaccuracies when tackling complex mathematical operations. While integrating computational tools such as Code Interpreters (CIs) offers a promising solution, it introduces a critical challenge: a conflict between the model's internal, probabilistic reasoning and the external, deterministic knowledge provided by the CI, which often leads models to unproductive deliberation. To overcome this, we introduce CoRT (Code-Optimized Reasoning Training), a post-training framework designed to teach LRMs to effectively utilize CIs. We propose **Hint-Engineering**, a new data synthesis strategy that strategically injects diverse hints at optimal points within reasoning paths. This approach generates high-quality, code-integrated reasoning data specifically tailored to optimize LRM-CI interaction. Using this method, we have synthesized 30 high-quality samples to post-train models ranging from 1.5B to 32B parameters through supervised fine-tuning. CoRT further refines the multi-round interleaving of external CI usage and internal thinking by employing rejection sampling and reinforcement learning. Our experimental evaluations demonstrate CoRT's effectiveness, yielding absolute improvements of 4\% and 8\% on DeepSeek-R1-Distill-Qwen-32B and DeepSeek-R1-Distill-Qwen-1.5B, respectively, across five challenging mathematical reasoning datasets. Moreover, CoRT significantly enhances efficiency, reducing token usage by approximately 30\% for the 32B model and 50\% for the 1.5B model compared to pure natural language reasoning baselines. The models and code are available at: [this url](https://github.com/ChengpengLi1003/CoRT). Chengpeng Li 0001, Zhengyang Tang, Ziniu Li, Mingfeng Xue, Keqin Bao, Tian Ding, Ruoyu Sun 0001, Benyou Wang, Xiang Wang 0010, Junyang Lin, Dayiheng Liu |
NeurIPS | 11 |
| 2025 | Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeabstractGating mechanisms have been widely utilized, from early models like LSTMs and Highway Networks to recent state space models, linear attention, and also softmax attention.
Yet, existing literature rarely examines the specific effects of gating.
In this work, we conduct comprehensive experiments to systematically investigate gating-augmented softmax attention variants.
Specifically, we perform a comprehensive comparison over 30 variants of 15B Mixture-of-Experts (MoE) models and 1.7B dense models trained on a 3.5 trillion token dataset.
Our central finding is that a simple modification—applying a head-specific sigmoid gate after the Scaled Dot-Product Attention (SDPA)—consistently improves performance.
This modification also enhances training stability, tolerates larger learning rates, and improves scaling properties.
By comparing various gating positions and computational variants, we attribute this effectiveness to two key factors: (1) introducing non-linearity upon the low-rank mapping in the softmax attention, and (2) applying query-dependent sparse gating scores to modulate the SDPA output.
Notably, we find this sparse gating mechanism mitigates `massive activation`, `attention sink` and enhances long-context extrapolation performance.
We also release related codes (https://github.com/qiuzh20/gated_attention}) and models (https://huggingface.co/QwQZh/gated_attention) to facilitate future research.
Furthermore, the most effective SDPA output gating is used in the Qwen3-Next models (https://huggingface.co/collections/Qwen/qwen3-next). Zihan Qiu, Bo Zheng 0007, Kaiyue Wen, Rui Men, Suozhi Huang, Dayiheng Liu, Jingren Zhou 0001, Junyang Lin |
NeurIPS | 11 |
| 2024 | Talk Funny! A Large-Scale Humor Response Dataset with Chain-of-Humor InterpretationabstractHumor is a crucial part of human communication. Understanding humor and generating humorous responses in dialogue can provide natural and empathic human-computer interactions. However, most existing pre-trained language models (PLMs) perform unsatisfactorily in humor generation. On one hand, the serious shortage of humor corpus and datasets pose challenges for constructing models that can understand and generate humorous expressions. On the other hand, humor generation relies on rich knowledge and commonsense, which is often tacit and unspoken. In this paper, we construct the largest Chinese Explainable Humor Response Dataset to date with chain-of-humor and humor mind map annotations, which can be used to comprehensively evaluate as well as improve the humorous response ability of PLMs. We further design humor-related auxiliary tasks to further enhance PLMs' humorous response performance. Extensive evaluations demonstrate that our proposed dataset and auxiliary tasks effectively help PLMs to generate humorous responses, laying the groundwork for future humor research. Yuyan Chen, Panjun Liu, Dayiheng Liu, Qinghao Guan, Mengfei Guo, Haiming Peng, Bang Liu 0003, Zhixu Li, Yanghua Xiao |
AAAI | 4 |
| 2024 | How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data CompositionabstractGuanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li, Mingfeng Xue, Dayiheng Liu, Wei Wang, Zheng Yuan, Chang Zhou, Jingren Zhou. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Guanting Dong 0001, Hongyi Yuan, Keming Lu, Chengpeng Li 0001, Mingfeng Xue, Dayiheng Liu, Wei Wang 0225, Zheng Yuan 0002, Chang Zhou 0005, Jingren Zhou 0001 |
ACL (1) | 6 |
| 2024 | MoNMT: Modularly Leveraging Monolingual and Bilingual Knowledge for Neural Machine TranslationabstractThe effective use of monolingual and bilingual knowledge represents a critical challenge within the neural machine translation (NMT) community. In this paper, we propose a modular strategy that facilitates the cooperation of these two types of knowledge in translation tasks, while avoiding the issue of catastrophic forgetting and exhibiting superior model generalization and robustness. Our model is comprised of three functionally independent modules: an encoding module, a decoding module, and a transferring module. The former two acquire large-scale monolingual knowledge via self-supervised learning, while the latter is trained on parallel data and responsible for transferring latent features between the encoding and decoding modules. Extensive experiments in multi-domain translation tasks indicate our model yields remarkable performance, with up to 7 BLEU improvements in out-of-domain tests over the conventional pretrain-and-finetune approach. Our codes are available at https://github.com/NLP2CT/MoNMT. Jianhui Pang, Baosong Yang, Derek F. Wong, Dayiheng Liu, Xiangpeng Wei, Lidia S. Chao |
LREC/COLING | 4 |
| 2024 | Knowledge Enhanced Pre-training for Cross-lingual Dense RetrievalabstractIn recent years, multilingual pre-trained language models (mPLMs) have achieved significant progress in cross-lingual dense retrieval. However, most mPLMs neglect the importance of knowledge. Knowledge always conveys similar semantic concepts in a language-agnostic manner, while query-passage pairs in cross-lingual retrieval also share common factual information. Motivated by this observation, we introduce KEPT, a novel mPLM that effectively leverages knowledge to learn language-agnostic semantic representations. To achieve this, we construct a multilingual knowledge base using hyperlinks and cross-language page alignment data annotated by Wiki. From this knowledge base, we mine intra- and cross-language pairs by extracting symmetrically linked segments and multilingual entity descriptions. Subsequently, we adopt contrastive learning with the mined pairs to pre-train KEPT. We evaluate KEPT on three widely-used benchmarks, considering both zero-shot cross-lingual transfer and supervised multilingual fine-tuning scenarios. Extensive experimental results demonstrate that KEPT achieves strong multilingual and cross-lingual retrieval performance with significant improvements over existing mPLMs. Hang Zhang 0029, Yeyun Gong, Dayiheng Liu, Shunyu Zhang, Xingwei He 0003, Jiancheng Lv 0001, Jian Guo 0016 |
LREC/COLING | 3 |
| 2024 | An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train ModelabstractRecent studies applied Parameter Efficient Fine-Tuning techniques (PEFTs) to efficiently narrow the performance gap between pre-training and downstream. There are two important factors for various PEFTs, namely, the accessible data size and fine-tunable parameter size. A natural expectation for PEFTs is that the performance of various PEFTs is positively related to the data size and fine-tunable parameter size. However, according to the evaluation of five PEFTs on two downstream vision-language (VL) tasks, we find that such an intuition holds only if the downstream data and task are not consistent with pre-training. For downstream fine-tuning consistent with pre-training, data size no longer affects the performance, while the influence of fine-tunable parameter size is not monotonous. We believe such an observation could guide the choice of training strategy for various PEFTs. Mouxing Yang, Yunfan Li 0003, Dayiheng Liu, Xingzhang Ren, Xi Peng 0001, Jiancheng Lv 0001 |
ICME | 4 |
| 2024 | Rethinking the Exploitation of Monolingual Data for Low-Resource Neural Machine TranslationabstractAbstract The utilization of monolingual data has been shown to be a promising strategy for addressing low-resource machine translation problems. Previous studies have demonstrated the effectiveness of techniques such as back-translation and self-supervised objectives, including masked language modeling, causal language modeling, and denoise autoencoding, in improving the performance of machine translation models. However, the manner in which these methods contribute to the success of machine translation tasks and how they can be effectively combined remains an under-researched area. In this study, we carry out a systematic investigation of the effects of these techniques on linguistic properties through the use of probing tasks, including source language comprehension, bilingual word alignment, and translation fluency. We further evaluate the impact of pre-training, back-translation, and multi-task learning on bitexts of varying sizes. Our findings inform the design of more effective pipelines for leveraging monolingual data in extremely low-resource and low-resource machine translation tasks. Experiment results show consistent performance gains in seven translation directions, which provide further support for our conclusions and understanding of the role of monolingual data in machine translation. Jianhui Pang, Baosong Yang, Derek F. Wong, Yu Wan 0004, Dayiheng Liu, Lidia S. Chao |
Comput. Linguistics | 5 |
| 2023 | Tailor: A Soft-Prompt-Based Approach to Attribute-Based Controlled Text GenerationabstractKexin Yang, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Mingfeng Xue, Boxing Chen, Jun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kexin Yang 0002, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Mingfeng Xue, Boxing Chen |
ACL (1) | 2 |
| 2023 | Bridging the Domain Gaps in Context Representations for k-Nearest Neighbor Neural Machine TranslationabstractZhiwei Cao, Baosong Yang, Huan Lin, Suhang Wu, Xiangpeng Wei, Dayiheng Liu, Jun Xie, Min Zhang, Jinsong Su. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Baosong Yang, Suhang Wu, Xiangpeng Wei, Dayiheng Liu, Min Zhang 0005, Jinsong Su |
ACL (1) | 6 |
| 2023 | Fantastic Expressions and Where to Find Them: Chinese Simile Generation with Multiple ConstraintsabstractKexin Yang, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Xiangpeng Wei, Zhengyuan Liu, Jun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kexin Yang 0002, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Xiangpeng Wei, Zhengyuan Liu |
ACL (1) | 2 |
| 2023 | Hallucination Detection: Robustly Discerning Reliable Answers in Large Language ModelsabstractLarge language models (LLMs) have gained widespread adoption in various natural language processing tasks, including question answering and dialogue systems. However, a major drawback of LLMs is the issue of hallucination, where they generate unfaithful or inconsistent content that deviates from the input source, leading to severe consequences. In this paper, we propose a robust discriminator named RelD to effectively detect hallucination in LLMs' generated answers. RelD is trained on the constructed RelQA, a bilingual question-answering dialogue dataset along with answers generated by LLMs and a comprehensive set of metrics. Our experimental results demonstrate that the proposed RelD successfully detects hallucination in the answers generated by diverse LLMs. Moreover, it performs well in distinguishing hallucination in LLMs' generated answers from both in-distribution and out-of-distribution datasets. Additionally, we also conduct a thorough analysis of the types of hallucinations that occur and present valuable insights. This research significantly contributes to the detection of reliable answers generated by LLMs and holds noteworthy implications for mitigating hallucination in the future work. Yuyan Chen, Qiang Fu 0015, Zhihao Wen, Ge Fan, Dayiheng Liu, Dongmei Zhang 0001, Zhixu Li, Yanghua Xiao |
CIKM | 6 |
| 2023 | Unifying Discrete and Continuous Representations for Unsupervised Paraphrase GenerationabstractMingfeng Xue, Dayiheng Liu, Wenqiang Lei, Jie Fu, Jian Lan, Mei Li, Baosong Yang, Jun Xie, Yidan Zhang, Dezhong Peng, Jiancheng Lv. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Mingfeng Xue, Dayiheng Liu, Wenqiang Lei, Jie Fu 0001, Baosong Yang, Yidan Zhang 0004, Dezhong Peng, Jiancheng Lv 0001 |
EMNLP | 2 |
| 2023 | EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation LearningabstractExpressing universal semantics common to all languages is helpful to understand the meanings of complex and culture-specific sentences. The research theme underlying this scenario focuses on learning universal representations across languages with the usage of massive parallel corpora. However, due to the sparsity and scarcity of parallel data, there is still a big challenge in learning authentic ``universals'' for any two languages. In this paper, we propose Emma-X: an EM-like Multilingual pre-training Algorithm, to learn Cross-lingual universals with the aid of excessive multilingual non-parallel data. Emma-X unifies the cross-lingual representation learning task and an extra semantic relation prediction task within an EM framework. Both the extra semantic classifier and the cross-lingual sentence encoder approximate the semantic relation of two sentences, and supervise each other until convergence. To evaluate Emma-X, we conduct experiments on xrete, a newly introduced benchmark containing 12 widely studied cross-lingual tasks that fully depend on sentence-level representations. Results reveal that Emma-X achieves state-of-the-art performance. Further geometric analysis of the built representation space with three requirements demonstrates the superiority of Emma-X over advanced models. Ping Guo 0002, Xiangpeng Wei, Yue Hu 0002, Baosong Yang, Dayiheng Liu, Fei Huang 0002 |
NeurIPS | 5 |
| 2022 | KGR4: Retrieval, Retrospect, Refine and Rethink for Commonsense GenerationabstractGenerative commonsense reasoning requires machines to generate sentences describing an everyday scenario given several concepts, which has attracted much attention recently. However, existing models cannot perform as well as humans, since sentences they produce are often implausible and grammatically incorrect. In this paper, inspired by the process of humans creating sentences, we propose a novel Knowledge-enhanced Commonsense Generation framework, termed KGR4, consisting of four stages: Retrieval, Retrospect, Refine, Rethink. Under this framework, we first perform retrieval to search for relevant sentences from external corpus as the prototypes. Then, we train the generator that either edits or copies these prototypes to generate candidate sentences, of which potential errors will be fixed by an autoencoder-based refiner. Finally, we select the output sentence from candidate sentences produced by generators with different hyper-parameters. Experimental results and in-depth analysis on the CommonGen benchmark strongly demonstrate the effectiveness of our framework. Particularly, KGR4 obtains 33.56 SPICE in the official leaderboard, outperforming the previously-reported best result by 2.49 SPICE and achieving state-of-the-art performance. We release the code at https://github.com/DeepLearnXMU/KGR-4. Xin Liu 0066, Dayiheng Liu, Baosong Yang, Haibo Zhang 0013, Junwei Ding, Wenqing Yao, Weihua Luo, Jinsong Su |
AAAI | 2 |
| 2022 | Frequency-Aware Contrastive Learning for Neural Machine TranslationabstractLow-frequency word prediction remains a challenge in modern neural machine translation (NMT) systems. Recent adaptive training methods promote the output of infrequent words by emphasizing their weights in the overall training objectives. Despite the improved recall of low-frequency words, their prediction precision is unexpectedly hindered by the adaptive objectives. Inspired by the observation that low-frequency words form a more compact embedding space, we tackle this challenge from a representation learning perspective. Specifically, we propose a frequency-aware token-level contrastive learning method, in which the hidden state of each decoding step is pushed away from the counterparts of other target words, in a soft contrastive way based on the corresponding word frequencies. We conduct experiments on widely used NIST Chinese-English and WMT14 English-German translation tasks. Empirical results show that our proposed methods can not only significantly improve the translation quality but also enhance lexical diversity and optimize word representation space. Further investigation reveals that, comparing with related adaptive training strategies, the superiority of our method on low-frequency word prediction lies in the robustness of token-level recall across different frequencies without sacrificing precision. Tong Zhang 0001, Wei Ye 0004, Baosong Yang, Long Zhang 0012, Xingzhang Ren, Dayiheng Liu, Jinan Sun, Shikun Zhang, Haibo Zhang 0013 |
AAAI | 6 |
| 2022 | UniTE: Unified Translation EvaluationabstractYu Wan, Dayiheng Liu, Baosong Yang, Haibo Zhang, Boxing Chen, Derek Wong, Lidia Chao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yu Wan 0004, Dayiheng Liu, Baosong Yang, Haibo Zhang 0013, Boxing Chen, Derek F. Wong, Lidia S. Chao |
ACL (1) | 2 |
| 2022 | Competency-Aware Neural Machine Translation: Can Machine Translation Know its Own Translation Quality?abstractNeural machine translation (NMT) is often criticized for failures that happen without awareness.The lack of competency awareness makes NMT untrustworthy.This is in sharp contrast to human translators who give feedback or conduct further investigations whenever they are in doubt about predictions.To fill this gap, we propose a novel competency-aware NMT by extending conventional NMT with a selfestimator, offering abilities to translate a source sentence and estimate its competency.The selfestimator encodes the information of the decoding procedure and then examines whether it can reconstruct the original semantics of the source sentence.Experimental results on four translation tasks demonstrate that the proposed method not only carries out translation tasks intact but also delivers outstanding performance on quality estimation.Without depending on any reference or annotated data typically required by state-of-the-art metric and quality estimation methods, our model yields an even higher correlation with human quality judgments than a variety of aforementioned methods, such as BLEURT, COMET, and BERTScore.Quantitative and qualitative analyses show better robustness of competency awareness in our model.1 Pei Zhang 0011, Baosong Yang, Dayiheng Liu, Kai Fan 0002, Luo Si |
EMNLP | 4 |
| 2022 | Should We Rely on Entity Mentions for Relation Extraction? Debiasing Relation Extraction with Counterfactual AnalysisabstractYiwei Wang, Muhao Chen, Wenxuan Zhou, Yujun Cai, Yuxuan Liang, Dayiheng Liu, Baosong Yang, Juncheng Liu, Bryan Hooi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yiwei Wang 0001, Muhao Chen 0001, Wenxuan Zhou 0002, Yujun Cai, Yuxuan Liang 0002, Dayiheng Liu, Baosong Yang, Bryan Hooi |
NAACL-HLT | 6 |
| 2022 | Effective Approaches to Neural Query Language IdentificationabstractAbstract Query language identification (Q-LID) plays a crucial role in a cross-lingual search engine. There exist two main challenges in Q-LID: (1) insufficient contextual information in queries for disambiguation; and (2) the lack of query-style training examples for low-resource languages. In this article, we propose a neural Q-LID model by alleviating the above problems from both model architecture and data augmentation perspectives. Concretely, we build our model upon the advanced Transformer model. In order to enhance the discrimination of queries, a variety of external features (e.g., character, word, as well as script) are fed into the model and fused by a multi-scale attention mechanism. Moreover, to remedy the low resource challenge in this task, a novel machine translation–based strategy is proposed to automatically generate synthetic query-style data for low-resource languages. We contribute the first Q-LID test set called QID-21, which consists of search queries in 21 languages. Experimental results reveal that our model yields better classification accuracy than strong baselines and existing LID systems on both query and traditional LID tasks.1 Xingzhang Ren, Baosong Yang, Dayiheng Liu, Haibo Zhang 0013, Xiaoyu Lv |
Comput. Linguistics | 3 |
| 2022 | Prediction, selection, and generation: a knowledge-driven conversation system
Dayiheng Liu, Chanjuan Li, Jiancheng Lv 0001 |
Neural Comput. Appl. | 2 |
| 2022 | CoupGAN: Chinese couplet generation via encoder-decoder model and adversarial training under global control
Qian Qu, Jiancheng Lv 0001, Dayiheng Liu, Kexin Yang 0002 |
Soft Comput. | 3 |
| 2021 | Towards User-Driven Neural Machine TranslationabstractHuan Lin, Liang Yao, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Degen Huang, Jinsong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Baosong Yang, Dayiheng Liu, Haibo Zhang 0013, Weihua Luo, Degen Huang, Jinsong Su |
ACL/IJCNLP (1) | 4 |
| 2021 | Bridging Subword Gaps in Pretrain-Finetune Paradigm for Natural Language GenerationabstractXin Liu, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Min Zhang, Haiying Zhang, Jinsong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xin Liu 0066, Baosong Yang, Dayiheng Liu, Haibo Zhang 0013, Weihua Luo, Min Zhang 0005, Jinsong Su |
ACL/IJCNLP (1) | 3 |
| 2021 | POS-Constrained Parallel Decoding for Non-autoregressive GenerationabstractKexin Yang, Wenqiang Lei, Dayiheng Liu, Weizhen Qi, Jiancheng Lv. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kexin Yang 0002, Wenqiang Lei, Dayiheng Liu, Weizhen Qi, Jiancheng Lv 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale PretrainingabstractIn this paper, we propose BANG, a new pretraining model to Bridge the gap between Autoregressive (AR) and Non-autoregressive (NAR) Generation. AR and NAR generation can be uniformly regarded as to what extent previous tokens can be attended, and BANG bridges AR and NAR generation through designing a novel model structure for large-scale pre-training. A pretrained BANG model can simultaneously support AR, NAR, and semi-NAR generation to meet different requirements. Experiments on question generation (SQuAD 1.1), summarization (XSum), and dialogue generation (PersonaChat) show that BANG improves NAR and semi-NAR performance significantly as well as attaining comparable performance with strong AR pretrained models. Compared with the semi-NAR strong baselines, BANG achieves absolute improvements of 14.01 and 5.24 in the overall scores of SQuAD 1.1 and XSum, respectively. In addition, BANG achieves absolute improvements of 10.73, 6.39, and 5.90 in the overall scores of SQuAD, XSUM, and PersonaChat compared with the NAR strong baselines, respectively. Our code will be made publicly available. Weizhen Qi, Yeyun Gong, Jian Jiao 0007, Weizhu Chen, Dayiheng Liu, Kewen Tang, Houqiang Li, Jiusheng Chen, Ruofei Zhang, Ming Zhou 0001, Nan Duan 0001 |
ICML | 6 |
| 2021 | AnchiBERT: A Pre-Trained Model for Ancient Chinese Language Understanding and GenerationabstractAncient Chinese is the essence of Chinese culture. There are several natural language processing tasks of ancient Chinese domain, such as ancient-modern Chinese translation, poem generation, and couplet generation. Previous studies usually use the supervised models which deeply rely on parallel data. However, it is difficult to obtain large-scale parallel data of ancient Chinese. In order to make full use of the more easily available monolingual ancient Chinese corpora, we release An-chiBERT, a pre-trained language model based on the architecture of BERT, which is trained on large-scale ancient Chinese corpora. We evaluate AnchiBERT on both language understanding and generation tasks, including poem classification, ancient-modern Chinese translation, poem generation, and couplet generation. The experimental results show that AnchiBERT outperforms BERT as well as the non-pretrained models and achieves state-of - the-art results in all cases. Huishuang Tian, Kexin Yang 0002, Dayiheng Liu, Jiancheng Lv 0001 |
IJCNN | 3 |
| 2021 | Mask Attention Networks: Rethinking and Strengthen TransformerabstractZhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, Xuanjing Huang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang 0025, Jian Jiao 0007, Nan Duan 0001, Ruofei Zhang, Xuanjing Huang 0001 |
NAACL-HLT | 3 |
| 2021 | An automatic evaluation metric for Ancient-Modern Chinese translation
Kexin Yang 0002, Dayiheng Liu, Qian Qu, Yongsheng Sang, Jiancheng Lv 0001 |
Neural Comput. Appl. | 2 |
| 2020 | Revision in Continuous Space: Unsupervised Text Style Transfer without Adversarial LearningabstractTypical methods for unsupervised text style transfer often rely on two key ingredients: 1) seeking the explicit disentanglement of the content and the attributes, and 2) troublesome adversarial learning. In this paper, we show that neither of these components is indispensable. We propose a new framework that utilizes the gradients to revise the sentence in a continuous space during inference to achieve text style transfer. Our method consists of three key components: a variational auto-encoder (VAE), some attribute predictors (one for each attribute), and a content predictor. The VAE and the two types of predictors enable us to perform gradient-based optimization in the continuous space, which is mapped from sentences in a discrete space, to find the representation of a target sentence with the desired attributes and preserved content. Moreover, the proposed method naturally has the ability to simultaneously manipulate multiple fine-grained attributes, such as sentence length and the presence of specific words, when performing text style transfer tasks. Compared with previous adversarial learning based methods, the proposed method is more interpretable, controllable and easier to train. Extensive experimental studies on three popular text style transfer tasks show that the proposed method significantly outperforms five state-of-the-art methods. Dayiheng Liu, Jie Fu 0001, Yidan Zhang 0004, Christopher Joseph Pal, Jiancheng Lv 0001 |
AAAI | 1 |
| 2020 | Deep Poetry: A Chinese Classical Poetry Generation SystemabstractIn this work, we demonstrate a Chinese classical poetry generation system called Deep Poetry. Existing systems for Chinese classical poetry generation are mostly template-based and very few of them can accept multi-modal input. Unlike previous systems, Deep Poetry uses neural networks that are trained on over 200 thousand poems and 3 million ancient Chinese prose. Our system can accept plain text, images or artistic conceptions as inputs to generate Chinese classical poetry. More importantly, users are allowed to participate in the process of writing poetry by our system. For the user's convenience, we deploy the system at the WeChat applet platform, users can use the system on the mobile device whenever and wherever possible. Yusen Liu 0001, Dayiheng Liu, Jiancheng Lv 0001 |
AAAI | 2 |
| 2020 | RikiNet: Reading Wikipedia Pages for Natural Question AnsweringabstractReading long documents to answer opendomain questions remains challenging in natural language understanding.In this paper, we introduce a new model, called RikiNet, which reads Wikipedia pages for natural question answering.RikiNet contains a dynamic paragraph dual-attention reader and a multi-level cascaded answer predictor.The reader dynamically represents the document and question by utilizing a set of complementary attention mechanisms.The representations are then fed into the predictor to obtain the span of the short answer, the paragraph of the long answer, and the answer type in a cascaded manner.On the Natural Questions (NQ) dataset, a single RikiNet achieves 74.3 F1 and 57.9 F1 on longanswer and short-answer tasks.To our best knowledge, it is the first single model that outperforms the single human performance.Furthermore, an ensemble RikiNet obtains 76.1 F1 and 61.3 F1 on long-answer and shortanswer tasks, achieving the best performance on the official NQ leaderboard 1 . Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Jiusheng Chen, Daxin Jiang, Jiancheng Lv 0001, Nan Duan 0001 |
ACL | 1 |
| 2020 | Herb-Know: Knowledge Enhanced Prescription Generation for Traditional Chinese MedicineabstractPrescription generation of traditional Chinese medicine (TCM) is a meaningful and challenging problem. Previous researches mainly model the relationship between symptoms and herbal prescription directly. However, TCM practitioners often take herb effects into consideration when prescribing. Few works focus on fusing the external knowledge of herbs. In this paper, we explore how to generate a prescription with the knowledge of herb effects under the given symptoms. We propose Herb-Know, a sequence to sequence (seq2seq) model with pointer network, where the prescription is conditioned over two inputs (symptoms and pre-selected herb candidates). To the best of our knowledge, this is the first attempt to generate a prescription with a knowledge enhanced seq2seq model. The experimental results demonstrate that our method can make use of knowledge to generate informative and reasonable herbs, which outperforms other baseline models. Chanjuan Li, Dayiheng Liu, Kexin Yang 0002, Jiancheng Lv 0001 |
BIBM | 2 |
| 2020 | Tell Me How to Ask Again: Question Data Augmentation with Controllable Rewriting in Continuous SpaceabstractIn this paper, we propose a novel data augmentation method, referred to as Controllable Rewriting based Question Data Augmentation (CRQDA), for machine reading comprehension (MRC), question generation, and question-answering natural language inference tasks.We treat the question data augmentation task as a constrained question rewriting problem to generate context-relevant, high-quality, and diverse question data samples.CRQDA utilizes a Transformer autoencoder to map the original discrete question into a continuous embedding space.It then uses a pre-trained MRC model to revise the question representation iteratively with gradientbased optimization.Finally, the revised question representations are mapped back into the discrete space, which serve as additional question data.Comprehensive experiments on SQuAD 2.0, SQuAD 1.1 question generation, and QNLI tasks demonstrate the effectiveness of CRQDA 1 . Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Jiusheng Chen, Jiancheng Lv 0001, Nan Duan 0001, Ming Zhou 0001 |
EMNLP (1) | 1 |
| 2020 | Diverse, Controllable, and Keyphrase-Aware: A Corpus and Method for News Multi-Headline GenerationabstractNews headline generation aims to produce a short sentence to attract readers to read the news.One news article often contains multiple keyphrases that are of interest to different users, which can naturally have multiple reasonable headlines.However, most existing methods focus on the single headline generation.In this paper, we propose generating multiple headlines with keyphrases of user interests, whose main idea is to generate multiple keyphrases of interest to users for the news first, and then generate multiple keyphrase-relevant headlines.We propose a multi-source Transformer decoder, which takes three sources as inputs: (a) keyphrase, (b) keyphrase-filtered article, and (c) original article to generate keyphrase-relevant, highquality, and diverse headlines.Furthermore, we propose a simple and effective method to mine the keyphrases of interest in the news article and build a first large-scale keyphraseaware news headline corpus, which contains over 180K aligned triples of news article, headline, keyphrase .Extensive experimental comparisons on the real-world dataset show that the proposed method achieves state-of-theart results in terms of quality and diversity 1 . Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Daxin Jiang, Jiancheng Lv 0001, Nan Duan 0001 |
EMNLP (1) | 1 |
| 2020 | Exploration on the Generation of Chinese Palindrome Poetry
Zhichen Lai 0001, Dayiheng Liu, Jiancheng Lv 0001, Yongsheng Sang |
ICONIP (1) | 3 |
| 2020 | Generating Chinese Poetry from Images via Concrete and Abstract InformationabstractIn recent years, the automatic generation of classical Chinese poetry has made great progress. Besides focusing on improving the quality of the generated poetry, there is a new topic about generating poetry from an image. However, the existing methods for this topic still have the problem of topic drift and semantic inconsistency, and the image-poem pairs dataset is hard to be built when training these models. In this paper, we extract and integrate the Concrete and Abstract information from images to address those issues. We proposed an infilling-based Chinese poetry generation model which can infill the Concrete keywords into each line of poems in an explicit way, and an abstract information embedding to integrate the Abstract information into generated poems. In addition, we use non-parallel data during training and construct separate image datasets and poem datasets to train the different components in our framework. Both automatic and human evaluation results show that our approach can generate poems which have better consistency with images without losing the quality. Yusen Liu 0001, Dayiheng Liu, Jiancheng Lv 0001, Yongsheng Sang |
IJCNN | 2 |
| 2020 | μ-Forcing: Training Variational Recurrent Autoencoders for Text GenerationabstractIt has been previously observed that training Variational Recurrent Autoencoders (VRAE) for text generation suffers from serious uninformative latent variables problems. The model would collapse into a plain language model that totally ignores the latent variables and can only generate repeating and dull samples. In this article, we explore the reason behind this issue and propose an effective regularizer-based approach to address it. The proposed method directly injects extra constraints on the posteriors of latent variables into the learning process of VRAE, which can flexibly and stably control the tradeoff between the Kullback-Leibler (KL) term and the reconstruction term, making the model learn dense and meaningful latent representations. The experimental results show that the proposed method outperforms several strong baselines and can make the model learn interpretable latent variables and generate diverse meaningful sentences. Furthermore, the proposed method can perform well without using other strategies, such as KL annealing. Dayiheng Liu, Yuanyuan Chen 0006, Jiancheng Lv 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2020 | Ancient-Modern Chinese Translation with a New Large Training DatasetabstractAncient Chinese brings the wisdom and spirit culture of the Chinese nation. Automatic translation from ancient Chinese to modern Chinese helps to inherit and carry forward the quintessence of the ancients. However, the lack of large-scale parallel corpus limits the study of machine translation in ancient–modern Chinese. In this article, we propose an ancient–modern Chinese clause alignment approach based on the characteristics of these two languages. This method combines both lexical-based information and statistical-based information, which achieves 94.2 F1-score on our manual annotation Test set. We use this method to create a new large-scale ancient–modern Chinese parallel corpus that contains 1.24M bilingual pairs. To our best knowledge, this is the first large high-quality ancient–modern Chinese dataset. Furthermore, we analyzed and compared the performance of the SMT and various NMT models on this dataset and provided a strong baseline for this task. Dayiheng Liu, Kexin Yang 0002, Qian Qu, Jiancheng Lv 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2019 | TIGS: An Inference Algorithm for Text Infilling with Gradient SearchabstractText infilling is defined as a task for filling in the missing part of a sentence or paragraph, which is suitable for many real-world natural language generation scenarios.However, given a well-trained sequential generative model, generating missing symbols conditioned on the context is challenging for existing greedy approximate inference algorithms.In this paper, we propose an iterative inference algorithm based on gradient search, which is the first inference algorithm that can be broadly applied to any neural sequence generative models for text infilling tasks.We compare the proposed method with strong baselines on three text infilling tasks with various mask ratios and different mask strategies.The results show that our proposed method is effective and efficient for fill-in-the-blank tasks, consistently outperforming all baselines.1 Dayiheng Liu, Jie Fu 0001, Pengfei Liu 0003, Jiancheng Lv 0001 |
ACL (1) | 1 |
| 2019 | Deep learning-based automatic downbeat tracking: a brief review
Bijue Jia, Jiancheng Lv 0001, Dayiheng Liu |
Multim. Syst. | 3 |
| 2019 | BFGAN: Backward and Forward Generative Adversarial Networks for Lexically Constrained Sentence GenerationabstractIncorporating prior knowledge like lexical constraints into the model's output to generate meaningful and coherent sentences has many applications in dialogue system, machine translation, image captioning, etc. However, existing auto-regressive models incrementally generate sentences from left to right via beam search, which makes it difficult to directly introduce lexical constraints into the generated sentences. In this paper, we propose a new algorithmic framework, dubbed BFGAN, to address this challenge. Specifically, we employ a backward generator and a forward generator to generate lexically constrained sentences together, and use a discriminator to guide the joint training of two generators by assigning them reward signals. Due to the difficulty of BFGAN training, we propose several training techniques to make the training process more stable and efficient. Our extensive experiments on three large-scale datasets with human evaluation demonstrate that BFGAN has significant improvements over previous methods. Dayiheng Liu, Jie Fu 0001, Qian Qu, Jiancheng Lv 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | A Multi-Modal Chinese Poetry Generation ModelabstractRecent studies in sequence-to-sequence learning demonstrate that RNN encoder-decoder structure can successfully generate Chinese poetry. However, existing methods can only generate poetry with a given first line or user's intent theme. In this paper, we proposed a three-stage multi-modal Chinese poetry generation approach. Given a picture, the first line, the title and the other lines of the poem are successively generated in three stages. According to the characteristics of Chinese poems, we propose a hierarchy-attention seq2seq model which can effectively capture character, phrase, and sentence information between contexts and improve the symmetry delivered in poems. In addition, the Latent Dirichlet allocation (LDA) model is utilized for title generation and improve the relevance of the whole poem and the title. Compared with strong baseline, the experimental results demonstrate the effectiveness of our approach, using machine evaluations as well as human judgments. Dayiheng Liu, Quan Guo, Wubo Li, Jiancheng Lv 0001 |
IJCNN | 1 |
| 2018 | Method to Improve the Performance of Restricted Boltzmann Machines
Jing Yin, Qingyu Mao, Dayiheng Liu, Jiancheng Lv 0001 |
ISNN | 3 |
| 2016 | A neural words encoding modelabstractThis paper proposes a neural network model and learning algorithm that can be applied to encode words. The model realizes the function of words encoding and decoding which can be applied to text encryption/decryption and word-based compression. The model is based on Deep Belief Networks (DBNs) and it differs from traditional DBNs in that it is asymmetric structured and the output of it is a binary vector. With pre-training of multi-layer Restricted Boltzmann Machines (RBMs) and fine-tuning to reconstruct word set, the output of code layer can be used as a kind of representation code of words. We can change the number of neurons of code layer to control the length of representation code for different applications. This paper reports on experiments using English words of American National Corpus to train a neural words encoding model which can be used to encode/decode English words, realizing text encryption and data compression. Dayiheng Liu, Jiancheng Lv 0001, Xiaofeng Qi, Jiangshu Wei |
IJCNN | 1 |