Yiming Cui 0001

dblp:130/6308-1 · DBLP profile ↗
← Back
28ranked-venue papers
13as first author
9since 2021 · last 2025
0000-0002-2452-375XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 12 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
16 papers
Question answering and dialogue systems · 28% Language models and text generation · 19% Efficient and distributed learning · 11%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
2.052022
Teaching Machines to Read, Answer and Explain · IEEE ACM Trans. Audio Speech Lang. Process. 2022
A Span-Extraction Dataset for Chinese Machine Reading Comprehension · EMNLP/IJCNLP (1) 2019
Cross-Lingual Machine Reading Comprehension · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation › natural language understanding › question answering
multiple-choice question answering
1.022022
Teaching Machines to Read, Answer and Explain · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Convolutional Spatial Attention Model for Reading Comprehension with Multiple-Choice Questions · AAAI 2019
Computer vision › Vision and language › vision-language model › multimodal large language model › chart understanding
chart-to-code generation
0.912025
Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code Generation · EMNLP 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
chart understanding
0.912025
Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code Generation · EMNLP 2025
Natural language and speech › Language models and text generation
pre-trained language model
0.832023
Pre-Training With Whole Word Masking for Chinese BERT · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Gradient-based Intra-attention Pruning on Pre-trained Language Models · ACL (1) 2023
Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting · EMNLP (1) 2020
Machine learning › Learning paradigms
lifelong learning
0.812024
Self-Evolving GPT: A Lifelong Autonomous Experiential Learner · ACL (1) 2024
Natural language and speech › Question answering and dialogue systems › dialogue modeling
dialogue context modeling
0.712023
A Static and Dynamic Attention Framework for Multi Turn Dialogue Generation · ACM Trans. Inf. Syst. 2023
Machine learning › Efficient and distributed learning
model compression
0.712023
Gradient-based Intra-attention Pruning on Pre-trained Language Models · ACL (1) 2023
Natural language and speech › Question answering and dialogue systems › dialogue generation
multi-turn dialogue generation
0.712023
A Static and Dynamic Attention Framework for Multi Turn Dialogue Generation · ACM Trans. Inf. Syst. 2023
Machine learning › Efficient and distributed learning › attention efficiency
self-attention pruning
0.712023
Gradient-based Intra-attention Pruning on Pre-trained Language Models · ACL (1) 2023
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning
0.712023
Gradient-based Intra-attention Pruning on Pre-trained Language Models · ACL (1) 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › explanation generation
answer explanation
0.612022
Teaching Machines to Read, Answer and Explain · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Natural language and speech › Language models and text generation › natural language understanding › question answering
explainable question answering
0.612022
Teaching Machines to Read, Answer and Explain · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Natural language and speech › Language models and text generation
masked language modeling
0.512021
Pre-Training With Whole Word Masking for Chinese BERT · IEEE ACM Trans. Audio Speech Lang. Process. 2021
Machine learning › Learning paradigms › continual learning › catastrophic forgetting
catastrophic forgetting mitigation
0.412020
Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting · EMNLP (1) 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.412020
Discriminative Sentence Modeling for Story Ending Prediction · AAAI 2020
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.412020
Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting · EMNLP (1) 2020
Machine learning › Graph learning › graph neural network
graph neural network for NLP
0.412020
Is Graph Structure Necessary for Multi-hop Question Answering? · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.412020
Is Graph Structure Necessary for Multi-hop Question Answering? · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis
narrative understanding
0.412020
Discriminative Sentence Modeling for Story Ending Prediction · AAAI 2020
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.412020
Conversational Word Embedding for Retrieval-Based Dialog System · ACL 2020
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder
0.412019
Exploiting Persona Information for Diverse Generation of Conversational Responses · IJCAI 2019
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation
0.412019
Exploiting Persona Information for Diverse Generation of Conversational Responses · IJCAI 2019
Natural language and speech › Question answering and dialogue systems › personalized dialogue
persona-grounded dialogue
0.412019
Exploiting Persona Information for Diverse Generation of Conversational Responses · IJCAI 2019
Machine learning › Deep learning architectures and training › attention mechanism › visual attention
spatial attention
0.412019
Convolutional Spatial Attention Model for Reading Comprehension with Multiple-Choice Questions · AAAI 2019
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
cloze-style reading comprehension
0.312017
Attention-over-Attention Neural Networks for Reading Comprehension · ACL (1) 2017
Natural language and speech › Information extraction and text analysis
coreference resolution
0.312017
Generating and Exploiting Large-scale Pseudo Training Data for Zero Pronoun Resolution · ACL (1) 2017
Natural language and speech › Information extraction and text analysis › coreference resolution
zero pronoun resolution
0.312017
Generating and Exploiting Large-scale Pseudo Training Data for Zero Pronoun Resolution · ACL (1) 2017
Machine learning › Trustworthy machine learning
interpretability
0.212022
Teaching Machines to Read, Answer and Explain · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Machine learning › Trustworthy machine learning › interpretability › explainable AI
self-interpretable models
0.212022
Teaching Machines to Read, Answer and Explain · IEEE ACM Trans. Audio Speech Lang. Process. 2022

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 1.7data synthesis · 1.7experiential learning · 0.8static attention · 0.7knowledge distillation · 0.7hierarchical recurrent encoder-decoder · 0.7gradient separation · 0.7dynamic attention · 0.7unsupervised learning · 0.6recursive dynamic gating · 0.6
YearPublicationVenuePosition
2025 Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code Generation
abstract
Chart2Code has recently received significant attention in the multimodal community due to its potential to reduce the burden of visualization and promote a more detailed understanding of charts.However, existing Chart2Coderelated training datasets suffer from at least one of the following issues: (1) limited scale, (2) limited type coverage, and ( 3) inadequate complexity.To address these challenges, we seek more diverse sources that better align with real-world user distributions and propose dual data synthesis pipelines: (1) Synthesize based on online plotting code.(2) Synthesize based on the chart images in the academic paper.We create a large-scale Chart2Code training dataset Chart2Code53, including 53 chart types, 130K Chart-code pairs based on the pipeline.Experimental results demonstrate that even with few parameters, the model finetuned on Chart2Code53 achieves state-ofthe-art performance on multiple Chart2Code benchmarks within open-source models 1 .
Tianhao Niu, Yiming Cui 0001, Baoxin Wang, Xiao Xu 0005, Qingfu Zhu, Dayong Wu, Shijin Wang 0001, Wanxiang Che
EMNLP2
2025 You Might Not Need Attention Diagonals
abstract
Pre-trained language models, such as GPT, BERT, have revolutionized natural language processing tasks across various fields. However, the current multi-head self-attention mechanisms in these models exhibit an “over self-confidence” issue, which has been underexplored in prior research, causing the model to attend heavily to itself rather than other tokens. In this study, we propose a simple yet efficient solution: discarding diagonal elements in the attention matrix, allowing the model to focus more on other tokens. Our experiments reveal that the proposed approach not only consistently improves upon vanilla attention in transformer models for diverse natural language understanding tasks, particularly for smaller models in resource-limited conditions, but also exhibits faster convergence in training speed. This effectiveness generalizes well across different languages, model types, and various natural language understanding tasks, while requiring almost no additional computation. Our findings challenge previous assumptions about multi-head self-attention and suggest a promising direction for developing more effective pre-trained language models.
Yiming Cui 0001, Shijin Wang 0001
IEEE Signal Process. Lett.1
2024 Self-Evolving GPT: A Lifelong Autonomous Experiential Learner
abstract
Jinglong Gao, Xiao Ding, Yiming Cui, Jianbai Zhao, Hepeng Wang, Ting Liu, Bing Qin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jinglong Gao, Yiming Cui 0001, Jianbai Zhao, Hepeng Wang, Ting Liu 0001, Bing Qin 0001
ACL (1)3
2023 Gradient-based Intra-attention Pruning on Pre-trained Language Models
abstract
Pre-trained language models achieve superior performance but are computationally expensive.Techniques such as pruning and knowledge distillation have been developed to reduce their sizes and latencies.In this work, we propose a structured pruning method GRAIN (Gradientbased Intra-attention pruning), which performs task-specific pruning with knowledge distillation and yields highly effective models.Different from common approaches that prune each attention head as a whole, GRAIN inspects and prunes intra-attention structures, which greatly expands the structure search space and enables more flexible models.We also propose a gradient separation strategy that reduces the interference of distillation on pruning for a better combination of the two approaches.Experiments on GLUE, SQuAD, and CoNLL 2003 show that GRAIN notably outperforms other methods, especially in the high sparsity regime, and achieves 6 ∼ 7× speedups while maintaining 93% ∼ 99% performance.Under extreme compression where only 3% transformer weights remain, the pruned model is still competitive compared to larger models. 1
Ziqing Yang 0001, Yiming Cui 0001, Shijin Wang 0001
ACL (1)2
2023 A Static and Dynamic Attention Framework for Multi Turn Dialogue Generation
abstract
Recently, research on open domain dialogue systems have attracted extensive interests of academic and industrial researchers. The goal of an open domain dialogue system is to imitate humans in conversations. Previous works on single turn conversation generation have greatly promoted the research of open domain dialogue systems. However, understanding multiple single turn conversations is not equal to the understanding of multi turn dialogue due to the coherent and context dependent properties of human dialogue. Therefore, in open domain multi turn dialogue generation, it is essential to modeling the contextual semantics of the dialogue history rather than only according to the last utterance. Previous research had verified the effectiveness of the hierarchical recurrent encoder-decoder framework on open domain multi turn dialogue generation. However, using an RNN-based model to hierarchically encoding the utterances to obtain the representation of dialogue history still face the problem of a vanishing gradient. To address this issue, in this article, we proposed a static and dynamic attention-based approach to model the dialogue history and then generate open domain multi turn dialogue responses. Experimental results on the Ubuntu and Opensubtitles datasets verify the effectiveness of the proposed static and dynamic attention-based approach on automatic and human evaluation metrics in various experimental settings. Meanwhile, we also empirically verify the performance of combining the static and dynamic attentions on open domain multi turn dialogue generation.
Weinan Zhang 0003, Yiming Cui 0001, Yifa Wang, Qingfu Zhu, Lingzhi Li 0003, Ting Liu 0001
ACM Trans. Inf. Syst.2
2022 CINO: A Chinese Minority Pre-trained Language Model
abstract
Multilingual pre-trained language models have shown impressive performance on cross-lingual tasks. It greatly facilitates the applications of natural language processing on low-resource languages. However, there are still some languages that the current multilingual models do not perform well on. In this paper, we propose CINO (Chinese Minority Pre-trained Language Model), a multilingual pre-trained language model for Chinese minority languages. It covers Standard Chinese, Yue Chinese, and six other ethnic minority languages. To evaluate the cross-lingual ability of the multilingual model on ethnic minority languages, we collect documents from Wikipedia and news websites, and construct two text classification datasets, WCM (Wiki-Chinese-Minority) and CMNews (Chinese-Minority-News). We show that CINO notably outperforms the baselines on various classification tasks. The CINO model and the datasets are publicly available at http://cino.hfl-rc.com.
Ziqing Yang 0001, Zihang Xu, Yiming Cui 0001, Baoxin Wang, Dayong Wu, Zhigang Chen 0003
COLING3
2022 Interactive Gated Decoder for Machine Reading Comprehension
abstract
Owing to the availability of various large-scale Machine Reading Comprehension ( MRC ) datasets, building an effective model to extract passage spans for question answering has been well studied in previous works. However, in reality, there are some questions that cannot be answered through the passage information, which brings more challenges to this task. In this article, we propose an Interactive Gated Decoder ( IG Decoder ), which focuses on modeling the interactions between the answer span prediction and no-answer prediction with a gating mechanism. We also propose a simple but effective approach for automatically generating pseudo training data, which aims to enrich the training data of the unanswerable questions. Experimental results on popular benchmark SQuAD 2.0 and NewsQA show that the proposed approaches yield consistent improvements over traditional BERT-large and strong ALBERT-xxlarge baseline systems. We also provide detailed ablations of the proposed method and error analysis on hard samples, which could be helpful in future research.
Yiming Cui 0001, Wanxiang Che, Ziqing Yang 0001, Ting Liu 0001, Bing Qin 0001, Shijin Wang 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2022 Teaching Machines to Read, Answer and Explain
abstract
With various Pre-trained Language Models (PLMs) blooming, Machine Reading Comprehension (MRC) systems have embraced significant improvements on various benchmarks and even surpassed human performances. However, most existing works only focus on the accuracy of the answer predictions and neglect the importance of the explanations for the prediction, which is a big obstacle when utilizing these models in real-life applications to convince humans. This paper proposes a novel unsupervised self-explainable framework, called Recursive Dynamic Gating (RDG), for the machine reading comprehension task. The main idea is that the proposed system tries to use less passage information and achieves similar results to the system that uses the whole passage, while the filtered passage is used as text explanations. We carried out experiments on three multiple-choice MRC datasets (including English and Chinese) and found that the proposed system can not only achieve better performance in answer prediction but also provide informative explanations compared to the attention mechanism.
Yiming Cui 0001, Ting Liu 0001, Wanxiang Che, Zhigang Chen 0003, Shijin Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Pre-Training With Whole Word Masking for Chinese BERT
abstract
Bidirectional Encoder Representations from Transformers (BERT) has shown marvelous improvements across various NLP tasks, and its consecutive variants have been proposed to further improve the performance of the pre-trained language models. In this paper, we aim to first introduce the whole word masking (wwm) strategy for Chinese BERT, along with a series of Chinese pre-trained language models. Then we also propose a simple but effective model called MacBERT, which improves upon RoBERTa in several ways. Especially, we propose a new masking strategy called MLM as correction (Mac). To demonstrate the effectiveness of these models, we create a series of Chinese pre-trained language models as our baselines, including BERT, RoBERTa, ELECTRA, RBT, etc. We carried out extensive experiments on ten Chinese NLP tasks to evaluate the created Chinese pre-trained language models as well as the proposed MacBERT. Experimental results show that MacBERT could achieve state-of-the-art performances on many NLP tasks, and we also ablate details with several findings that may help future research. We open-source our pre-trained language models for further facilitating our research community.
Yiming Cui 0001, Wanxiang Che, Ting Liu 0001, Bing Qin 0001, Ziqing Yang 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Discriminative Sentence Modeling for Story Ending Prediction
abstract
Story Ending Prediction is a task that needs to select an appropriate ending for the given story, which requires the machine to understand the story and sometimes needs commonsense knowledge. To tackle this task, we propose a new neural network called Diff-Net for better modeling the differences of each ending in this task. The proposed model could discriminate two endings in three semantic levels: contextual representation, story-aware representation, and discriminative representation. Experimental results on the Story Cloze Test dataset show that the proposed model siginificantly outperforms various systems by a large margin, and detailed ablation studies are given for better understanding our model. We also carefully examine the traditional and BERT-based models on both SCT v1.0 and v1.5 with interesting findings that may potentially help future studies.
Yiming Cui 0001, Wanxiang Che, Weinan Zhang 0003, Ting Liu 0001, Shijin Wang 0001
AAAI1
2020 Conversational Word Embedding for Retrieval-Based Dialog System
abstract
Human conversations contain many types of information, e.g., knowledge, common sense, and language habits.In this paper, we propose a conversational word embedding method named PR-Embedding, which utilizes the conversation pairs post, reply 1 to learn word embedding.Different from previous works, PR-Embedding uses the vectors from two different semantic spaces to represent the words in post and reply.To catch the information among the pair, we first introduce the word alignment model from statistical machine translation to generate the cross-sentence window, then train the embedding on word-level and sentence-level.We evaluate the method on single-turn and multi-turn response selection tasks for retrieval-based dialog systems.The experiment results show that PR-Embedding can improve the quality of the selected response.2
Yiming Cui 0001, Ting Liu 0001, Shijin Wang 0001
ACL2
2020 A Sentence Cloze Dataset for Chinese Machine Reading Comprehension
abstract
Owing to the continuous efforts by the Chinese NLP community, more and more Chinese machine reading comprehension datasets become available.To add diversity in this area, in this paper, we propose a new task called Sentence Cloze-style Machine Reading Comprehension (SC-MRC).The proposed task aims to fill the right candidate sentence into the passage that has several blanks.We built a Chinese dataset called CMRC 2019 to evaluate the difficulty of the SC-MRC task.Moreover, to add more difficulties, we also made fake candidates that are similar to the correct ones, which requires the machine to judge their correctness in the context.The proposed dataset contains over 100K blanks (questions) within over 10K passages, which was originated from Chinese narrative stories.To evaluate the dataset, we implement several baseline systems based on the pre-trained models, and the results show that the stateof-the-art model still underperforms human performance by a large margin.We release the dataset and baseline system to further facilitate our community.
Yiming Cui 0001, Ting Liu 0001, Ziqing Yang 0001, Zhipeng Chen 0001, Wanxiang Che, Shijin Wang 0001
COLING1
2020 CharBERT: Character-aware Pre-trained Language Model
abstract
Most pre-trained language models (PLMs) construct word representations at subword level with Byte-Pair Encoding (BPE) or its variations, by which OOV (out-of-vocab) words are almost avoidable.However, those methods split a word into subword units and make the representation incomplete and fragile.In this paper, we propose a character-aware pre-trained language model named CharBERT improving on the previous methods (such as BERT, RoBERTa) to tackle these problems.We first construct the contextual word embedding for each token from the sequential character representations, then fuse the representations of characters and the subword representations by a novel heterogeneous interaction module.We also propose a new pre-training task named NLM (Noisy LM) for unsupervised character representation learning.We evaluate our method on question answering, sequence labeling, and text classification tasks, both on the original datasets and adversarial misspelling test sets.The experimental results show that our method can significantly improve the performance and robustness of PLMs simultaneously.Pretrained models, evaluation sets, and code are available at https
Yiming Cui 0001, Chenglei Si, Ting Liu 0001, Shijin Wang 0001
COLING2
2020 CLUE: A Chinese Language Understanding Evaluation Benchmark
abstract
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, Zhenzhong Lan. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Liang Xu 0011, Hai Hu 0001, Xuanwei Zhang, Chenjie Cao, Yudong Li 0001, Yechen Xu, Kai Sun 0006, Dian Yu 0001, Cong Yu 0010, Yin Tian, Qianqian Dong, Weitang Liu, Yiming Cui 0001, Rongzhao Wang, Weijian Xie, Yina Patterson, Zuoyu Tian, Shaoweihua Liu, Zhe Zhao 0006, Qipeng Zhao, Cong Yue, Zhengliang Yang, Kyle Richardson 0001, Zhen-Zhong Lan
COLING15
2020 Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting
abstract
Deep pretrained language models have achieved great success in the way of pretraining first and then fine-tuning.But such a sequential transfer learning paradigm often confronts the catastrophic forgetting problem and leads to sub-optimal performance.To fine-tune with less forgetting, we propose a recall and learn mechanism, which adopts the idea of multi-task learning and jointly learns pretraining tasks and downstream tasks.Specifically, we propose a Pretraining Simulation mechanism to recall the knowledge from pretraining tasks without data, and an Objective Shifting mechanism to focus the learning on downstream tasks gradually.Experiments show that our method achieves state-of-the-art performance on the GLUE benchmark.Our method also enables BERT-base to achieve better performance than directly fine-tuning of BERT-large.Further, we provide the open-source RECADAM optimizer, which integrates the proposed mechanisms into Adam optimizer, to facility the NLP community.
Sanyuan Chen, Yutai Hou, Yiming Cui 0001, Wanxiang Che, Ting Liu 0001, Xiangzhan Yu
EMNLP (1)3
2020 Is Graph Structure Necessary for Multi-hop Question Answering?
abstract
Recently, attempting to model texts as graph structure and introducing graph neural networks to deal with it has become a trend in many NLP research areas.In this paper, we investigate whether the graph structure is necessary for multi-hop question answering.Our analysis is centered on HotpotQA.We construct a strong baseline model to establish that, with the proper use of pre-trained models, graph structure may not be necessary for multi-hop question answering.We point out that both graph structure and adjacency matrix are task-related prior knowledge, and graphattention can be considered as a special case of self-attention.Experiments and visualized analysis demonstrate that graph-attention or the entire graph structure can be replaced by self-attention or Transformers.
Yiming Cui 0001, Ting Liu 0001, Shijin Wang 0001
EMNLP (1)2
2019 Convolutional Spatial Attention Model for Reading Comprehension with Multiple-Choice Questions
abstract
Machine Reading Comprehension (MRC) with multiplechoice questions requires the machine to read given passage and select the correct answer among several candidates. In this paper, we propose a novel approach called Convolutional Spatial Attention (CSA) model which can better handle the MRC with multiple-choice questions. The proposed model could fully extract the mutual information among the passage, question, and the candidates, to form the enriched representations. Furthermore, to merge various attention results, we propose to use convolutional operation to dynamically summarize the attention values within the different size of regions. Experimental results show that the proposed model could give substantial improvements over various state-of- the-art systems on both RACE and SemEval-2018 Task11 datasets.
Zhipeng Chen 0001, Yiming Cui 0001, Shijin Wang 0001
AAAI2
2019 TripleNet: Triple Attention Network for Multi-Turn Response Selection in Retrieval-Based Chatbots
abstract
We consider the importance of different utterances in the context for selecting the response usually depends on the current query. In this paper, we propose the model TripleNet to fully model the task with the triple instead of in previous works. The heart of TripleNet is a novel attention mechanism named triple attention to model the relationships within the triple at four levels. The new mechanism updates the representation for each element based on the attention with the other two concurrently and symmetrically. We match the triple centered on the response from char to context level for prediction. Experimental results on two large-scale multi-turn response selection datasets show that the proposed model can significantly outperform the state-of-the-art methods. TripleNet source code is available at https://github.com/wtma/TripleNet
Yiming Cui 0001, Su He, Weinan Zhang 0003, Ting Liu 0001, Shijin Wang 0001
CoNLL2
2019 Cross-Lingual Machine Reading Comprehension
abstract
Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Shijin Wang, Guoping Hu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yiming Cui 0001, Wanxiang Che, Ting Liu 0001, Bing Qin 0001, Shijin Wang 0001
EMNLP/IJCNLP (1)1
2019 A Span-Extraction Dataset for Chinese Machine Reading Comprehension
abstract
Yiming Cui, Ting Liu, Wanxiang Che, Li Xiao, Zhipeng Chen, Wentao Ma, Shijin Wang, Guoping Hu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yiming Cui 0001, Ting Liu 0001, Wanxiang Che, Zhipeng Chen 0001, Shijin Wang 0001
EMNLP/IJCNLP (1)1
2019 Exploiting Persona Information for Diverse Generation of Conversational Responses
abstract
In human conversations, due to their personalities in mind, people can easily carry out and maintain the conversations. Giving conversational context with persona information to a chatbot, how to exploit the information to generate diverse and sustainable conversations is still a non-trivial task. Previous work on persona-based conversational models successfully make use of predefined persona information and have shown great promise in delivering more realistic responses. And they all learn with the assumption that given a source input, there is only one target response. However, in human conversations, there are massive appropriate responses to a given input message. In this paper, we propose a memory-augmented architecture to exploit persona information from context and incorporate a conditional variational autoencoder model together to generate diverse and sustainable conversations. We evaluate the proposed model on a benchmark persona-chat dataset. Both automatic and human evaluations show that our model can deliver more diverse and more engaging persona-based responses than baseline approaches.
Haoyu Song 0002, Weinan Zhang 0003, Yiming Cui 0001, Ting Liu 0001
IJCAI3
2018 Context-Sensitive Generation of Open-Domain Conversational Responses
abstract
Despite the success of existing works on single-turn conversation generation, taking the coherence in consideration, human conversing is actually a context-sensitive process. Inspired by the existing studies, this paper proposed the static and dynamic attention based approaches for context-sensitive generation of open-domain conversational responses. Experimental results on two public datasets show that the proposed static attention based approach outperforms all the baselines on automatic and human evaluation.
Weinan Zhang 0003, Yiming Cui 0001, Yifa Wang, Qingfu Zhu, Lingzhi Li 0003, Lianqiang Zhou, Ting Liu 0001
COLING2
2018 Dataset for the First Evaluation on Chinese Machine Reading Comprehension
Yiming Cui 0001, Ting Liu 0001, Zhipeng Chen 0001, Shijin Wang 0001
LREC1
2017 Attention-over-Attention Neural Networks for Reading Comprehension
abstract
Cloze-style queries are representative problems in reading comprehension. Over the past few months, we have seen much progress that utilizing neural network approach to solve Cloze-style questions. In this paper, we present a novel model called attention-over-attention reader for the Cloze-style reading comprehension task. Our model aims to place another attention mechanism over the document-level attention, and induces "attended attention" for final predictions. Unlike the previous works, our neural network model requires less pre-defined hyper-parameters and uses an elegant architecture for modeling. Experimental results show that the proposed attention-over-attention model significantly outperforms various state-of-the-art systems by a large margin in public datasets, such as CNN and Children's Book Test datasets.
Yiming Cui 0001, Zhipeng Chen 0001, Si Wei, Shijin Wang 0001, Ting Liu 0001
ACL (1)1
2017 Generating and Exploiting Large-scale Pseudo Training Data for Zero Pronoun Resolution
abstract
Most existing approaches for zero pronoun resolution are heavily relying on annotated data, which is often released by shared task organizers.Therefore, the lack of annotated data becomes a major obstacle in the progress of zero pronoun resolution task.Also, it is expensive to spend manpower on labeling the data for better performance.To alleviate the problem above, in this paper, we propose a simple but novel approach to automatically generate large-scale pseudo training data for zero pronoun resolution.Furthermore, we successfully transfer the cloze-style reading comprehension neural network model into zero pronoun resolution task and propose a two-step training mechanism to overcome the gap between the pseudo training data and the real one.Experimental results show that the proposed approach significantly outperforms the state-of-the-art systems with an absolute improvements of 3.1% F-score on OntoNotes 5.0 data.
Ting Liu 0001, Yiming Cui 0001, Qingyu Yin, Weinan Zhang 0003, Shijin Wang 0001
ACL (1)2
2016 Consensus Attention-based Neural Networks for Chinese Reading Comprehension
abstract
Reading comprehension has embraced a booming in recent NLP research. Several institutes have released the Cloze-style reading comprehension data, and these have greatly accelerated the research of machine comprehension. In this work, we firstly present Chinese reading comprehension datasets, which consist of People Daily news dataset and Children’s Fairy Tale (CFT) dataset. Also, we propose a consensus attention-based neural network architecture to tackle the Cloze-style reading comprehension problem, which aims to induce a consensus attention over every words in the query. Experimental results show that the proposed neural network significantly outperforms the state-of-the-art baselines in several public datasets. Furthermore, we setup a baseline for Chinese reading comprehension task, and hopefully this would speed up the process for future research.
Yiming Cui 0001, Ting Liu 0001, Zhipeng Chen 0001, Shijin Wang 0001
COLING1
2016 LSTM Neural Reordering Feature for Statistical Machine Translation
abstract
Artificial neural networks are powerful models, which have been widely applied into many aspects of machine translation, such as language modeling and translation modeling.Though notable improvements have been made in these areas, the reordering problem still remains a challenge in statistical machine translations.In this paper, we present a novel neural reordering model that directly models word pairs and their alignment.Further by utilizing LSTM recurrent neural networks, much longer context could be learned for reordering prediction.Experimental results on NIST OpenMT12 Arabic-English and Chinese-English 1000-best rescoring task show that our LSTM neural reordering feature is robust, and achieves significant improvements over various baseline systems.
Yiming Cui 0001, Shijin Wang 0001
HLT-NAACL1
2013 Phrase Table Combination Deficiency Analyses in Pivot-Based SMT
Yiming Cui 0001, Conghui Zhu, Tiejun Zhao, Dequan Zheng
NLDB1