EDBT 2026 Demo / reviewers in the wild / expert
Jingjing Xu 0001
dblp:25/624-1
· DBLP profile ↗
34ranked-venue papers
10as first author
15since 2021 · last 2025
0000-0002-5763-8369ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 10 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Long Chain-of-Thought Fine-tuning via Understanding-to-Reasoning TransitionabstractChenxin An, Zhihui Xie, Xiaonan Li, Ming Zhong, Shansan Gong, Lei Li, Jun Zhang, Jingjing Xu, Lingpeng Kong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chenxin An, Zhihui Xie 0002, Ming Zhong 0005, Shansan Gong, Lei Li 0039, Jun Zhang 0003, Jingjing Xu 0001, Lingpeng Kong |
EMNLP | 8 |
| 2025 | Let the Code LLM Edit Itself When You Edit the CodeabstractIn this work, we investigate a typical scenario in code generation where a developer edits existing code in real time and requests a code assistant, e.g., a large language model, to re-predict the next token or next line on the fly. Naively, the LLM needs to re-encode the entire KV cache to provide an accurate prediction. However, this process is computationally expensive, especially when the sequence length is long. Simply encoding the edited subsequence and integrating it to the original KV cache meets the temporal confusion problem, leading to significantly worse performance. We address this efficiency and accuracy trade-off by introducing $\underline{\textbf{P}\text{ositional}\ \textbf{I}\text{ntegrity}\ \textbf{E}\text{ncoding}}$ (PIE). Building upon the rotary positional encoding, PIE first removes the rotary matrices in the Key cache that introduce temporal confusion and then reapplies the correct rotary matrices. This process ensures that positional relationships between tokens are correct and requires only a single round of matrix multiplication. We validate the effectiveness of PIE through extensive experiments on the RepoBench-C-8k dataset, utilizing DeepSeek-Coder models with 1.3B, 6.7B, and 33B parameters. Our evaluation includes three real-world coding tasks: code insertion, code deletion, and multi-place code editing. Results demonstrate that PIE reduces computational overhead by over 85% compared to the standard full recomputation approach across all model sizes and tasks while well approximating the model performance. Zhenyu He 0012, Jun Zhang 0003, Shengjie Luo, Jingjing Xu 0001, Zhi Zhang 0005, Di He 0001 |
ICLR | 4 |
| 2025 | Why Does the Effective Context Length of LLMs Fall Short?abstractAdvancements in distributed training and efficient attention mechanisms have significantly expanded the context window sizes of large language models (LLMs). However, recent work reveals that the effective context lengths of open-source LLMs often fall short, typically not exceeding half of their training lengths. In this work, we attribute this limitation to the left-skewed frequency distribution of relative positions formed in LLMs pretraining and post-training stages, which impedes their ability to effectively gather distant information.
To address this challenge, we introduce Shifted Rotray Position Embedding (STRING). STRING shifts well-trained positions to overwrite the original ineffective positions during inference, enhancing performance within their existing training lengths.
Experimental results show that without additional training, STRING dramatically improves the performance of the latest large-scale models, such as Llama3.1 70B and Qwen2 72B, by over 10 points on popular long-context benchmarks RULER and InfiniteBench, establishing new state-of-the-art results for open-source LLMs. Compared to commercial models, Llama 3.1 70B with STRING even achieves better performance than GPT-4-128K and clearly surpasses Claude 2 and Kimi-chat. Chenxin An, Jun Zhang 0003, Ming Zhong 0005, Lei Li 0039, Shansan Gong, Yao Luo, Jingjing Xu 0001, Lingpeng Kong |
ICLR | 7 |
| 2025 | Teaching Language Models to Critique via Reinforcement LearningabstractTeaching large language models (LLMs) to critique and refine their outputs is crucial for building systems that can iteratively improve, yet it is fundamentally limited by the ability to provide *accurate judgments* and *actionable suggestions*. In this work, we study LLM critics for code generation and propose $\texttt{CTRL}$, a framework for $\texttt{C}$ritic $\texttt{T}$raining via $\texttt{R}$einforcement $\texttt{L}$earning, which trains a critic model to generate feedback that maximizes correction performance for a fixed generator model without human supervision. Our results demonstrate that critics trained with $\texttt{CTRL}$ significantly enhance pass rates and mitigate compounding errors across both base and stronger generator models.
Furthermore, we show that these critic models act as accurate generative reward models and enable test-time scaling through iterative critique-revision, achieving up to 106.1\% relative improvements across challenging code generation benchmarks. Zhihui Xie 0002, Liyu Chen, Weichao Mao, Jingjing Xu 0001, Lingpeng Kong |
ICML | 5 |
| 2024 | A Survey on In-context LearningabstractQingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, Zhifang Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Qingxiu Dong, Lei Li 0039, Damai Dai, Jingyuan Ma, Rui Li 0094, Heming Xia, Jingjing Xu 0001, Zhiyong Wu 0011, Baobao Chang, Xu Sun 0001, Lei Li 0005, Zhifang Sui |
EMNLP | 8 |
| 2024 | Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length ExtrapolationabstractIn this work, we leverage the intrinsic segmentation of language sequences and design a new positional encoding method called Bilevel Positional Encoding (BiPE). For each position, our BiPE blends an intra-segment encoding and an inter-segment encoding. The intra-segment encoding identifies the locations within a segment and helps the model capture the semantic information therein via absolute positional encoding. The inter-segment encoding specifies the segment index, models the relationships between segments, and aims to improve extrapolation capabilities via relative positional encoding. Theoretical analysis shows this disentanglement of positional information makes learning more effective. The empirical results also show that our BiPE has superior length extrapolation capabilities across a wide range of tasks in diverse text modalities. Zhenyu He 0012, Guhao Feng, Shengjie Luo, Liwei Wang 0001, Jingjing Xu 0001, Zhi Zhang 0005, Hongxia Yang, Di He 0001 |
ICML | 6 |
| 2023 | INK: Injecting kNN Knowledge in Nearest Neighbor Machine TranslationabstractNeural machine translation has achieved promising results on many translation tasks.However, previous studies have shown that neural models induce a non-smooth representation space, which harms its generalization results.Recently, kNN-MT has provided an effective paradigm to smooth the prediction based on neighbor representations during inference.Despite promising results, kNN-MT usually requires large inference overhead.We propose an effective training framework INK to directly smooth the representation space via adjusting representations of kNN neighbors with a small number of new parameters.The new parameters are then used to refresh the whole representation datastore to get new kNN knowledge asynchronously.This loop keeps running until convergence.Experiments on four benchmark datasets show that INK achieves average gains of 1.99 COMET and 1.0 BLEU, outperforming the state-of-the-art kNN-MT system with 0.02× memory space and 1.9× inference speedup 1 . Jingjing Xu 0001, Shujian Huang, Lingpeng Kong, Jiajun Chen 0001 |
ACL (1) | 2 |
| 2023 | Can Language Models Understand Physical Concepts?abstractLanguage models (LMs) gradually become general-purpose interfaces in the interactive and embodied world, where the understanding of physical concepts is an essential prerequisite.However, it is unclear whether LMs can understand physical concepts in the human world.To investigate this, we design a benchmark VEC that covers the tasks of (i) Visual concepts, such as the shape and material of objects, and (ii) Embodied Concepts, learned from the interaction with the world such as the temperature of objects.Our zero (few)-shot prompting results show that the understanding of certain visual concepts emerges as scaling up LMs, but there are still basic concepts to which the scaling law does not apply.For example, OPT-175B performs close to humans with a zero-shot accuracy of 85% on the material concept, yet behaves like random guessing on the mass concept.Instead, vision-augmented LMs such as CLIP and BLIP achieve a human-level understanding of embodied concepts.Analysis indicates that the rich semantics in visual representation can serve as a valuable source of embodied knowledge.Inspired by this, we propose a distillation method to transfer embodied knowledge from VLMs to LMs, achieving performance gain comparable with that by scaling up parameters of LMs 134×. 1 o 1 : This is a photo of the water.o 2 : This is a photo of a frying oil.Attribute: This is a photo of a cold object. Lei Li 0039, Jingjing Xu 0001, Qingxiu Dong, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049 |
EMNLP | 2 |
| 2023 | Can We Edit Factual Knowledge by In-Context Learning?abstractPrevious studies have shown that large language models (LLMs) like GPTs store massive factual knowledge in their parameters.However, the stored knowledge could be false or outdated.Traditional knowledge editing methods refine LLMs via fine-tuning on texts containing specific knowledge.However, with the increasing scales of LLMs, these gradient-based approaches bring large computation costs.The trend of model-as-a-service also makes it impossible to modify knowledge in black-box LLMs.Inspired by in-context learning (ICL), a new paradigm based on demonstration contexts without parameter updating, we explore whether ICL can edit factual knowledge.To answer this question, we give a comprehensive empirical study of ICL strategies.Experiments show that in-context knowledge editing (IKE), without any gradient and parameter updating, achieves a competitive success rate compared to gradient-based methods on GPT-J (6B) but with much fewer side effects, including less over-editing on similar but unrelated facts and less knowledge forgetting on previously stored knowledge.We also apply the method to larger LMs with tens or hundreds of parameters like OPT-175B, which shows the scalability of our method.The code is available at https://github.com/pkunlp-icler/IKE. Lei Li 0039, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu 0011, Jingjing Xu 0001, Baobao Chang |
EMNLP | 6 |
| 2023 | Statistical Knowledge Assessment for Large Language ModelsabstractGiven varying prompts regarding a factoid question, can a large language model (LLM) reliably generate factually correct answers? Existing LLMs may generate distinct responses for different prompts. In this paper, we study the problem of quantifying knowledge contained in an LLM regarding a given set of facts. We propose KaRR, a statistical approach to assess factual knowledge for LLMs. The main idea is to estimate the ratio of LLM generating text corresponding to the answer entity given diverse prompts of the subject and the querying relation, versus it generating by random chances. Our assessment suite contains a comprehensive set of 994,123 entities and 600 relations, with 1,395,905 text aliases. We use our method to evaluate 20 LLMs of various sizes, including LLaMA, Alpaca, OPT, etc. Experiments show that our results have a strong correlation (0.43 Kendall's $\tau$) with the results of human assessment on LLMs. Our results reveal that the knowledge in LLMs with the same backbone architecture adheres to the scaling law, while tuning on instruction-following data sometimes compromises the model's capability to generate factually correct text reliably. Qingxiu Dong, Jingjing Xu 0001, Lingpeng Kong, Zhifang Sui, Lei Li 0005 |
NeurIPS | 2 |
| 2022 | Contextual Representation Learning beyond Masked Language ModelingabstractHow do masked language models (MLMs) such as BERT learn contextual representations?In this work, we analyze the learning dynamics of MLMs.We find that MLMs adopt sampled embeddings as anchors to estimate and inject contextual semantics to representations, which limits the efficiency and effectiveness of MLMs.To address these issues, we propose TACO, a simple yet effective representation learning approach to directly model global semantics.TACO extracts and aligns contextual semantics hidden in contextualized representations to encourage models to attend global semantics when generating contextualized representations.Experiments on the GLUE benchmark show that TACO achieves up to 5x speedup and up to 1.2 points average improvement over existing MLMs.The code is available at https:// github.com/FUZHIYI/TACO. Zhiyi Fu, Wangchunshu Zhou, Jingjing Xu 0001, Hao Zhou 0012, Lei Li 0005 |
ACL (1) | 3 |
| 2022 | switch-GLAT: Multilingual Parallel Machine Translation Via Code-Switch Decoder
Zhenqiao Song, Hao Zhou 0012, Lihua Qian, Jingjing Xu 0001, Shanbo Cheng, Mingxuan Wang, Lei Li 0005 |
ICLR | 4 |
| 2021 | Vocabulary Learning via Optimal Transport for Neural Machine TranslationabstractJingjing Xu, Hao Zhou, Chun Gan, Zaixiang Zheng, Lei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jingjing Xu 0001, Hao Zhou 0012, Chun Gan, Zaixiang Zheng, Lei Li 0005 |
ACL/IJCNLP (1) | 1 |
| 2021 | KNAS: Green Neural Architecture SearchabstractMany existing neural architecture search (NAS) solutions rely on downstream training for architecture evaluation, which takes enormous computations. Considering that these computations bring a large carbon footprint, this paper aims to explore a green (namely environmental-friendly) NAS solution that evaluates architectures without training. Intuitively, gradients, induced by the architecture itself, directly decide the convergence and generalization results. It motivates us to propose the gradient kernel hypothesis: Gradients can be used as a coarse-grained proxy of downstream training to evaluate random-initialized networks. To support the hypothesis, we conduct a theoretical analysis and find a practical gradient kernel that has good correlations with training loss and validation performance. According to this hypothesis, we propose a new kernel based architecture search approach KNAS. Experiments show that KNAS achieves competitive results with orders of magnitude faster than “train-then-test” paradigms on image classification tasks. Furthermore, the extremely low search cost enables its wide applications. The searched network also outperforms strong baseline RoBERTA-large on two text classification tasks. Jingjing Xu 0001, Junyang Lin, Rundong Gao, Xu Sun 0001, Hongxia Yang |
ICML | 1 |
| 2021 | Duplex Sequence-to-Sequence Learning for Reversible Machine TranslationabstractSequence-to-sequence learning naturally has two directions. How to effectively utilize supervision signals from both directions? Existing approaches either require two separate models, or a multitask-learned model but with inferior performance. In this paper, we propose REDER (Reversible Duplex Transformer), a parameter-efficient model and apply it to machine translation. Either end of REDER can simultaneously input and output a distinct language. Thus REDER enables {\em reversible machine translation} by simply flipping the input and output ends. Experiments verify that REDER achieves the first success of reversible machine translation, which helps outperform its multitask-trained baselines by up to 1.3 BLEU. Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Jiajun Chen 0001, Jingjing Xu 0001, Lei Li 0005 |
NeurIPS | 5 |
| 2020 | Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question AnsweringabstractCommonsense question answering aims to answer questions which require background knowledge that is not explicitly expressed in the question. The key challenge is how to obtain evidence from external knowledge and make predictions based on the evidence. Recent studies either learn to generate evidence from human-annotated evidence which is expensive to collect, or extract evidence from either structured or unstructured knowledge bases which fails to take advantages of both sources simultaneously. In this work, we propose to automatically extract evidence from heterogeneous knowledge sources, and answer questions based on the extracted evidence. Specifically, we extract evidence from both structured knowledge base (i.e. ConceptNet) and Wikipedia plain texts. We construct graphs for both sources to obtain the relational structures of evidence. Based on these graphs, we propose a graph-based approach consisting of a graph-based contextual word representation learning module and a graph-based inference module. The first module utilizes graph structural information to re-define the distance between words for learning better contextual word representations. The second module adopts graph convolutional network to encode neighbor information into the representations of nodes, and aggregates evidence with graph attention mechanism for predicting the final answer. Experimental results on CommonsenseQA dataset illustrate that our graph-based approach over both knowledge sources brings improvement over strong baselines. Our approach achieves the state-of-the-art accuracy (75.3%) on the CommonsenseQA dataset. Shangwen Lv, Daya Guo, Jingjing Xu 0001, Duyu Tang, Nan Duan 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Songlin Hu 0001 |
AAAI | 3 |
| 2020 | Reasoning Over Semantic-Level Graph for Fact CheckingabstractFact checking is a challenging task because verifying the truthfulness of a claim requires reasoning about multiple retrievable evidence.In this work, we present a method suitable for reasoning about the semantic-level structure of evidence.Unlike most previous works, which typically represent evidence sentences with either string concatenation or fusing the features of isolated evidence sentences, our approach operates on rich semantic structures of evidence obtained by semantic role labeling.We propose two mechanisms to exploit the structure of evidence while leveraging the advances of pre-trained models like BERT, GPT or XLNet.Specifically, using XLNet as the backbone, we first utilize the graph structure to re-define the relative distances of words, with the intuition that semantically related words should have short distances.Then, we adopt graph convolutional network and graph attention network to propagate and aggregate information from neighboring nodes on the graph.We evaluate our system on FEVER, a benchmark dataset for fact checking, and find that rich structural information is helpful and both our graph-based mechanisms improve the accuracy.Our model is the state-of-the-art system in terms of both official evaluation metrics, namely claim verification accuracy and FEVER score. Wanjun Zhong, Jingjing Xu 0001, Duyu Tang, Zenan Xu, Nan Duan 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001 |
ACL | 2 |
| 2020 | Modeling the Stock Relation with Graph Network for Overnight Stock Movement PredictionabstractStock movement prediction is a hot topic in the Fintech area. Previous works usually predict the price movement in a daily basis, although the market impact of news can be absorbed much shorter, and the exact time is hard to estimate. In this work, we propose a more practical objective to predict the overnight stock movement between the previous close price and the open price. As no trading operation occurs after market close, the market impact of overnight news will be reflected by the overnight movement. One big obstacle for such task is the lacking of data, in this work we collect and publish the overnight stock price movement dataset of Reuters Financial News. Another challenge is that the stocks in the market are not independent, which is omitted by previous works. To make use of the connection among stocks, we propose a LSTM Relational Graph Convolutional Network (LSTM-RGCN) model, which models the connection among stocks with their correlation matrix. Extensive experiment results show that our model outperforms the baseline models. Further analysis shows that the introduction of the graph enables our model to predict the movement of stocks that are not directly associated with news as well as the whole market, which is not available in most previous methods. Wei Li 0101, Ruihan Bao, Keiko Harimoto, Deli Chen, Jingjing Xu 0001, Qi Su 0001 |
IJCAI | 5 |
| 2020 | Training Simplification and Model Simplification for Deep Learning : A Minimal Effort Back Propagation MethodabstractWe propose a simple yet effective technique to simplify the training and the resulting model of neural networks. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-k elements (in terms of magnitude) are kept. As a result, only k rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction in the computational cost. Based on the sparsified gradients, we further simplify the model by eliminating the rows or columns that are seldom updated, which will reduce the computational cost both in the training and decoding, and potentially accelerate decoding in real-world applications. Surprisingly, experimental results demonstrate that most of the time we only need to update fewer than 5 percent of the weights at each back propagation pass. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given. The model simplification results show that we could adaptively simplify the model which could often be reduced by around 9x, without any loss on accuracy or even with improved accuracy. Xu Sun 0001, Xuancheng Ren, Shuming Ma, Bingzhen Wei, Wei Li 0101, Jingjing Xu 0001, Houfeng Wang, Yi Zhang 0050 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2019 | Coherent Comments Generation for Chinese Articles with a Graph-to-Sequence ModelabstractAutomatic article commenting is helpful in encouraging user engagement and interaction on online news platforms.However, the news documents are usually too long for traditional encoder-decoder based models, which often results in general and irrelevant comments.In this paper, we propose to generate comments with a graph-to-sequence model that models the input news as a topic interaction graph.By organizing the article into graph structure, our model can better understand the internal structure of the article and the connection between topics, which makes it better able to understand the story.We collect and release a large scale news-comment corpus from a popular Chinese online news platform Tencent Kuaibao. 1 Extensive experiment results show that our model can generate much more coherent and informative comments compared with several strong baseline models.2 Wei Li 0101, Jingjing Xu 0001, Yancheng He, Shengli Yan, Yunfang Wu, Xu Sun 0001 |
ACL (1) | 2 |
| 2019 | Asking Clarification Questions in Knowledge-Based Question AnsweringabstractJingjing Xu, Yuechen Wang, Duyu Tang, Nan Duan, Pengcheng Yang, Qi Zeng, Ming Zhou, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jingjing Xu 0001, Yuechen Wang, Duyu Tang, Nan Duan 0001, Qi Zeng 0001, Ming Zhou 0001, Xu Sun 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2019 | LexicalAT: Lexical-Based Adversarial Reinforcement Training for Robust Sentiment ClassificationabstractJingjing Xu, Liang Zhao, Hanqi Yan, Qi Zeng, Yun Liang, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jingjing Xu 0001, Hanqi Yan, Qi Zeng 0001, Yun Liang 0001, Xu Sun 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Specificity-Driven Cascading Approach for Unsupervised Sentiment ModificationabstractPengcheng Yang, Junyang Lin, Jingjing Xu, Jun Xie, Qi Su, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Junyang Lin, Jingjing Xu 0001, Qi Su 0001, Xu Sun 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Understanding and Improving Layer NormalizationabstractLayer normalization (LayerNorm) is a technique to normalize the distributions of intermediate layers. It enables smoother gradients, faster training, and better generalization accuracy. However, it is still unclear where the effectiveness stems from. In this paper, our main contribution is to take a step further in understanding LayerNorm. Many of previous studies believe that the success of LayerNorm comes from forward normalization. Unlike them, we find that the derivatives of the mean and variance are more important than forward normalization by re-centering and re-scaling backward gradients. Furthermore, we find that the parameters of LayerNorm, including the bias and gain, increase the risk of over-fitting and do not work in most cases. Experiments show that a simple version of LayerNorm (LayerNorm-simple) without the bias and gain outperforms LayerNorm on four datasets. It obtains the state-of-the-art performance on En-Vi machine translation. To address the over-fitting problem, we propose a new normalization method, Adaptive Normalization (AdaNorm), by replacing the bias and gain with a new transformation function. Experiments show that AdaNorm demonstrates better results than LayerNorm on seven out of eight datasets. Jingjing Xu 0001, Xu Sun 0001, Zhiyuan Zhang 0001, Guangxiang Zhao, Junyang Lin |
NeurIPS | 1 |
| 2019 | Evaluating Semantic Rationality of a Sentence: A Sememe-Word-Matching Neural Network Based on HowNet
Jingjing Xu 0001, Xuancheng Ren |
NLPCC (1) | 2 |
| 2019 | Knowledge-Aware Conversational Semantic Parsing over Web Tables
Duyu Tang, Jingjing Xu 0001, Nan Duan 0001, Bing Qin 0001, Ting Liu 0001, Ming Zhou 0001 |
NLPCC (1) | 3 |
| 2019 | Learning Unsupervised Word Mapping via Maximum Mean Discrepancy
Fuli Luo, Shuangzhi Wu, Jingjing Xu 0001, Dongdong Zhang 0001 |
NLPCC (1) | 4 |
| 2018 | Unpaired Sentiment-to-Sentiment Translation: A Cycled Reinforcement Learning ApproachabstractJingjing Xu, Xu Sun, Qi Zeng, Xiaodong Zhang, Xuancheng Ren, Houfeng Wang, Wenjie Li. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Jingjing Xu 0001, Xu Sun 0001, Qi Zeng 0001, Xiaodong Zhang 0022, Xuancheng Ren, Houfeng Wang, Wenjie Li 0002 |
ACL (1) | 1 |
| 2018 | An Auto-Encoder Matching Model for Learning Utterance-Level Semantic Dependency in Dialogue GenerationabstractGenerating semantically coherent responses is still a major challenge in dialogue generation.Different from conventional text generation tasks, the mapping between inputs and responses in conversations is more complicated, which highly demands the understanding of utterance-level semantic dependency, a relation between the whole meanings of inputs and outputs.To address this problem, we propose an Auto-Encoder Matching (AEM) model to learn such dependency.The model contains two auto-encoders and one mapping module.The auto-encoders learn the semantic representations of inputs and responses, and the mapping module learns to connect the utterance-level representations.Experimental results from automatic and human evaluations demonstrate that our model is capable of generating responses of high coherence and fluency compared to baseline models. 1 Liangchen Luo, Jingjing Xu 0001, Junyang Lin, Qi Zeng 0001, Xu Sun 0001 |
EMNLP | 2 |
| 2018 | Diversity-Promoting GAN: A Cross-Entropy Based Generative Adversarial Network for Diversified Text GenerationabstractExisting text generation methods tend to produce repeated and "boring" expressions. To tackle this problem, we propose a new text generation model, called Diversity-Promoting Generative Adversarial Network (DP-GAN).The proposed model assigns low reward for repeatedly generated text and high reward for "novel" and fluent text, encouraging the generator to produce diverse and informative text.Moreover, we propose a novel languagemodel based discriminator, which can better distinguish novel text from repeated text without the saturation problem compared with existing classifier-based discriminators.The experimental results on review generation and dialogue generation tasks demonstrate that our model can generate substantially more diverse and informative text than existing baselines.1 Jingjing Xu 0001, Xuancheng Ren, Junyang Lin, Xu Sun 0001 |
EMNLP | 1 |
| 2018 | A Skeleton-Based Model for Promoting Coherence Among Sentences in Narrative Story GenerationabstractNarrative story generation is a challenging problem because it demands the generated sentences with tight semantic connections, which has not been well studied by most existing generative models.To address this problem, we propose a skeleton-based model to promote the coherence of generated stories.Different from traditional models that generate a complete sentence at a stroke, the proposed model first generates the most critical phrases, called skeleton, and then expands the skeleton to a complete and fluent sentence.The skeleton is not manually defined, but learned by a reinforcement learning method.Compared to the state-of-the-art models, our skeleton-based model can generate significantly more coherent text according to human evaluation and automatic evaluation.The G-score is improved by 20.1% in human evaluation. 1 Jingjing Xu 0001, Xuancheng Ren, Yi Zhang 0050, Qi Zeng 0001, Xiaoyan Cai, Xu Sun 0001 |
EMNLP | 1 |
| 2018 | Learning Sentiment Memories for Sentiment Modification without Parallel DataabstractThe task of sentiment modification requires reversing the sentiment of the input and preserving the sentiment-independent content.However, aligned sentences with the same content but different sentiments are usually unavailable.Due to the lack of such parallel data, it is hard to extract sentiment independent content and reverse the sentiment in an unsupervised way.Previous work usually can not reconcile sentiment transformation and content preservation.In this paper, motivated by the fact the non-emotional context (e.g., "staff") provides strong cues for the occurrence of emotional words (e.g., "friendly"), we propose a novel method that automatically extracts appropriate sentiment information from the learned sentiment memories according to the specific context.Experiments show that our method substantially improves the content preservation degree and achieves the state-of-the-art performance.1 Yi Zhang 0050, Jingjing Xu 0001, Xu Sun 0001 |
EMNLP | 2 |
| 2018 | Cross-Domain and Semisupervised Named Entity Recognition in Chinese Social Media: A Unified ModelabstractNamed entity recognition (NER) in Chinese social media is an important, but challenging task because Chinese social media language is informal and noisy. Most previous methods on NER focus on in-domain supervised learning, which is limited by scarce annotated data in social media. In this paper, we present that sufficient corpora in formal domains and massive unannotated text can be combined to improve the NER performance in social media. We propose a unified model which can learn from out-of-domain corpora and in-domain unannotated text. The unified model is composed of two parts. One is for cross-domain learning and the other is for semisupervised learning. Cross-domain learning can learn out-of-domain information based on domain similarity. Semisupervised learning can learn in-domain unannotated information by self-training. Experimental results show that our unified model yields a 9.57% improvement over strong baselines and achieves the state-of-the-art performance. Jingjing Xu 0001, Hangfeng He 0001, Xu Sun 0001, Xuancheng Ren, Sujian Li |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Transfer Deep Learning for Low-Resource Chinese Word Segmentation with a Novel Neural Network
Jingjing Xu 0001, Shuming Ma, Yi Zhang 0050, Bingzhen Wei, Xiaoyan Cai, Xu Sun 0001 |
NLPCC | 1 |