EDBT 2026 Demo / reviewers in the wild / expert
Xuancheng Ren
dblp:202/2250
· DBLP profile ↗
39ranked-venue papers
2as first author
16since 2021 · last 2025
0000-0002-6994-2114ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsabstractRecent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities.
However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54% and 52.55% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning. Bofei Gao, Feifan Song 0001, Zhe Yang 0013, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li 0039, Chenghao Ma, Liang Chen 0024, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang 0009, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu 0001, Baobao Chang |
ICLR | 18 |
| 2025 | Rethinking Natural Language Generation with Layer-Wise Multi-View DecodingabstractIn natural language generation, language models, particularly those based on decoder-only architectures as in popular Large Language Models (LLMs), have demonstrated impressive performance across a wide range of tasks. However, encoder-decoder architectures remain highly effective for tasks involving non-text data, such as images and time-series data. The decoder relies on the attention mechanism to efficiently extract information from the encoder. While it is common practice to draw information from only the last encoder layer, this might lead to insufficient training of the encoder layer stack due to the hierarchy bypassing problem. In this work, we propose layer-wise multi-view decoding for improved encoder-decoder language models, where for each decoder layer, together with the representations from the last encoder layer, which serve as a global view, those from other encoder layers are supplemented for a stereoscopic view of the source inputs. Systematic experiments and analyses show that we successfully address the hierarchy bypassing problem, require almost negligible parameter increase, and improve the performance of sequence learning with deep representations on diverse tasks, i.e., machine translation, abstractive summarization, image captioning, video captioning, medical report generation, and paraphrase generation. In particular, our approach achieves new state-of-the-art results on benchmark datasets, including a low-resource machine translation dataset and low-resource medical report generation datasets. Xuancheng Ren, Guangxiang Zhao, Chenyu You, Sherry Ma, Xian Wu 0001, Wei Fan 0001, Xu Sun 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2023 | An Adaptive Learning Method for Solving the Extreme Learning Rate Problem of Transformer
Jianbang Ding, Xuancheng Ren, Ruixuan Luo |
NLPCC (1) | 2 |
| 2022 | Well-Classified Examples Are Underestimated in Classification with Deep Neural NetworksabstractThe conventional wisdom behind learning deep classification models is to focus on bad-classified examples and ignore well-classified examples that are far from the decision boundary. For instance, when training with cross-entropy loss, examples with higher likelihoods (i.e., well-classified examples) contribute smaller gradients in back-propagation. However, we theoretically show that this common practice hinders representation learning, energy optimization, and margin growth. To counteract this deficiency, we propose to reward well-classified examples with additive bonuses to revive their contribution to the learning process. This counterexample theoretically addresses these three issues. We empirically support this claim by directly verifying the theoretical results or significant performance improvement with our counterexample on diverse tasks, including image classification, graph classification, and machine translation. Furthermore, this paper shows that we can deal with complex scenarios, such as imbalanced classification, OOD detection, and applications under adversarial attacks because our idea can solve these three issues. Code is available at https://github.com/lancopku/well-classified-examples-are-underestimated. Guangxiang Zhao, Wenkai Yang, Xuancheng Ren, Lei Li 0039, Yunfang Wu, Xu Sun 0001 |
AAAI | 3 |
| 2022 | Rethinking the Promotion Brought by Contrastive Learning to Semi-Supervised Node ClassificationabstractGraph Contrastive Learning (GCL) has proven highly effective in promoting the performance of Semi-Supervised Node Classification (SSNC). However, existing GCL methods are generally transferred from other fields like CV or NLP, whose underlying working mechanism remains underexplored. In this work, we first deeply probe the working mechanism of GCL in SSNC, and find that the promotion brought by GCL is severely unevenly distributed: the improvement mainly comes from subgraphs with less annotated information, which is fundamentally different from contrastive learning in other fields. However, existing GCL methods generally ignore this uneven distribution of annotated information and apply GCL evenly to the whole graph. To remedy this issue and further improve GCL in SSNC, we propose the Topology InFormation gain-Aware Graph Contrastive Learning (TIFA-GCL) framework that considers the annotated information distribution across graph in GCL. Extensive experiments on six benchmark graph datasets, including the enormous OGB-Products graph, show that TIFA-GCL can bring a larger improvement than existing GCL methods in both transductive and inductive settings. Further experiments demonstrate the generalizability and interpretability of TIFA-GCL. Deli Chen, Yankai Lin 0001, Lei Li 0039, Xuancheng Ren, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001 |
IJCAI | 4 |
| 2022 | DiMBERT: Learning Vision-Language Grounded Representations with Disentangled Multimodal-AttentionabstractVision-and-language (V-L) tasks require the system to understand both vision content and natural language, thus learning fine-grained joint representations of vision and language (a.k.a. V-L representations) is of paramount importance. Recently, various pre-trained V-L models are proposed to learn V-L representations and achieve improved results in many tasks. However, the mainstream models process both vision and language inputs with the same set of attention matrices. As a result, the generated V-L representations are entangled in one common latent space . To tackle this problem, we propose DiMBERT (short for Di sentangled M ultimodal-Attention BERT ), which is a novel framework that applies separated attention spaces for vision and language, and the representations of multi-modalities can thus be disentangled explicitly. To enhance the correlation between vision and language in disentangled spaces, we introduce the visual concepts to DiMBERT which represent visual information in textual format. In this manner, visual concepts help to bridge the gap between the two modalities. We pre-train DiMBERT on a large amount of image–sentence pairs on two tasks: bidirectional language modeling and sequence-to-sequence language modeling. After pre-train, DiMBERT is further fine-tuned for the downstream tasks. Experiments show that DiMBERT sets new state-of-the-art performance on three tasks (over four datasets), including both generation tasks (image captioning and visual storytelling) and classification tasks (referring expressions). The proposed DiM (short for Di sentangled M ultimodal-Attention) module can be easily incorporated into existing pre-trained V-L models to boost their performance, up to a 5% increase on the representative task. Finally, we conduct a systematic analysis and demonstrate the effectiveness of our DiM and the introduced visual concepts. Xian Wu 0001, Shen Ge, Xuancheng Ren, Wei Fan 0001, Xu Sun 0001, Yuexian Zou |
ACM Trans. Knowl. Discov. Data | 4 |
| 2021 | Exploring the Vulnerability of Deep Neural Networks: A Study of Parameter CorruptionabstractWe argue that the vulnerability of model parameters is of crucial value to the study of model robustness and generalization but little research has been devoted to understanding this matter. In this work, we propose an indicator to measure the robustness of neural network parameters by exploiting their vulnerability via parameter corruption. The proposed indicator describes the maximum loss variation in the non-trivial worst-case scenario under parameter corruption. For practical purposes, we give a gradient-based estimation, which is far more effective than random corruption trials that can hardly induce the worst accuracy degradation. Equipped with theoretical support and empirical validation, we are able to systematically investigate the robustness of different model parameters and reveal vulnerability of deep neural networks that has been rarely paid attention to before. Moreover, we can enhance the models accordingly with the proposed adversarial corruption-resistant training, which not only improves the parameter robustness but also translates into accuracy elevation. Xu Sun 0001, Zhiyuan Zhang 0001, Xuancheng Ren, Ruixuan Luo, Liangyou Li |
AAAI | 3 |
| 2021 | Collaborative Group LearningabstractCollaborative learning has successfully applied knowledge transfer to guide a pool of small student networks towards robust local minima. However, previous approaches typically struggle with drastically aggravated student homogenization when the number of students rises. In this paper, we propose Collaborative Group Learning, an efficient framework that aims to diversify the feature representation and conduct an effective regularization. Intuitively, similar to the human group study mechanism, we induce students to learn and exchange different parts of course knowledge as collaborative groups. First, each student is established by randomly routing on a modular neural network, which facilitates flexible knowledge communication between students due to random levels of representation sharing and branching. Second, to resist the student homogenization, students first compose diverse feature sets by exploiting the inductive bias from sub-sets of training data, and then aggregate and distill different complementary knowledge by imitating a random sub-group of students at each time step. Overall, the above mechanisms are beneficial for maximizing the student population to further improve the model generalization without sacrificing computational efficiency. Empirical evaluations on both image and text tasks indicate that our method significantly outperforms various state-of-the-art collaborative approaches whilst enhancing computational efficiency. Shaoxiong Feng, Hongshen Chen, Xuancheng Ren, Zhuoye Ding, Kan Li 0001, Xu Sun 0001 |
AAAI | 3 |
| 2021 | Multi-View Feature Representation for Dialogue Generation with Bidirectional DistillationabstractNeural dialogue models suffer from low-quality responses when interacted in practice, demonstrating difficulty in generalization beyond training data. Recently, knowledge distillation has been used to successfully regularize the student by transferring knowledge from the teacher. However, the teacher and the student are trained on the same dataset and tend to learn similar feature representations, whereas the most general knowledge should be found through differences. The finding of general knowledge is further hindered by the unidirectional distillation, as the student should obey the teacher and may discard some knowledge that is truly general but refuted by the teacher. To this end, we propose a novel training framework, where the learning of general knowledge is more in line with the idea of reaching consensus, i.e., finding common knowledge that is beneficial to different yet all datasets through diversified learning partners. Concretely, the training task is divided into a group of subtasks with the same number of students. Each student assigned to one subtask not only is optimized on the allocated subtask but also imitates multi-view feature representation aggregated from other students (i.e., student peers), which induces students to capture common knowledge among different subtasks and alleviates the over-fitting of students on the allocated subtasks. To further enhance generalization, we extend the unidirectional distillation to the bidirectional distillation that encourages the student and its student peers to co-evolve by exchanging complementary knowledge with each other. Empirical results and analysis demonstrate that our training framework effectively improves the model generalization without sacrificing training efficiency. Shaoxiong Feng, Xuancheng Ren, Kan Li 0001, Xu Sun 0001 |
AAAI | 2 |
| 2021 | Towards Semantics-Enhanced Pre-Training: Can Lexicon Definitions Help Learning Sentence Meanings?abstractSelf-supervised pre-training techniques, albeit relying on large amounts of text, have enabled rapid growth in learning language representations for natural language understanding. However, as radically empirical models on sentences, they are subject to the input data distribution, inevitably incorporating data bias and reporting bias, which may lead to inaccurate understanding of sentences. To address this problem, we propose to adopt a human learner's approach: when we cannot make sense of a word in a sentence, we often consult the dictionary for specific meanings; but can the same work for empirical models? In this work, we try to inform the pre-trained masked language models of word meanings for semantics-enhanced pre-training. To achieve a contrastive and holistic view of word meanings, a definition pair of two related words is presented to the masked language model such that the model can better associate a word with its crucial semantic features. Both intrinsic and extrinsic evaluations validate the proposed approach on semantics-orientated tasks, with an almost negligible increase of training data. Xuancheng Ren, Xu Sun 0001, Houfeng Wang, Qun Liu 0001 |
AAAI | 1 |
| 2021 | Rethinking Denoised Auto-Encoding in Language Pre-TrainingabstractPre-trained self-supervised models such as BERT have achieved striking success in learning sequence representations, especially for natural language processing.These models typically corrupt the given sequences with certain types of noise, such as masking, shuffling, or substitution, and then try to recover the original input.However, such pre-training approaches are prone to learning representations that are covariant with the noise, leading to the discrepancy between the pre-training and finetuning stage.To remedy this, we present Con-trAstive Pre-Training (CAPT) to learn noise invariant sequence representations.The proposed CAPT encourages the consistency between representations of the original sequence and its corrupted version via unsupervised instance-wise training signals.In this way, it not only alleviates the pretrain-finetune discrepancy induced by the noise of pre-training, but also aids the pre-trained model in better capturing global semantics of the input via more effective sentence-level supervision.Different from most prior work that focuses on a particular modality, comprehensive empirical evidence on 11 natural language understanding and cross-modal tasks illustrates that CAPT is applicable for both language and vision-language tasks, and obtains surprisingly consistent improvement, including 0.6% absolute gain on GLUE benchmarks and 0.8% absolute increment on NLVR 2 . * Equal Contribution.Models Noise types BERT (Devlin et al., 2019) Mask tokens SpanBERT (Joshi et al., 2019) Mask spans RoBERTa (Liu et al., 2019) Mask token XLNet (Yang et al., 2019) Shuffle token ELECTRA (Clark et al., 2019) Replace tokens StructBERT (Wang et al., 2019b) Mask + Shuffle tokens BART (Lewis et al., 2019) Mask + Shuffle + Replace.UNITER (Chen et al., 2019) Mask tokens/regions LXMERT (Tan and Bansal, 2019) Mask tokens/regions Fuli Luo, Xuancheng Ren, Xu Sun 0001, Songfang Huang, Fei Huang 0002 |
EMNLP (1) | 4 |
| 2021 | A Global Past-Future Early Exit Method for Accelerating Inference of Pre-trained Language ModelsabstractKaiyuan Liao, Yi Zhang, Xuancheng Ren, Qi Su, Xu Sun, Bin He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Kaiyuan Liao, Yi Zhang 0050, Xuancheng Ren, Qi Su 0001, Xu Sun 0001 |
NAACL-HLT | 3 |
| 2021 | Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP ModelsabstractWenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, Bin He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Wenkai Yang, Lei Li 0039, Zhiyuan Zhang 0001, Xuancheng Ren, Xu Sun 0001 |
NAACL-HLT | 4 |
| 2021 | Neural Network Surgery: Injecting Data Patterns into Pre-trained Models with Minimal Instance-wise Side EffectsabstractZhiyuan Zhang, Xuancheng Ren, Qi Su, Xu Sun, Bin He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhiyuan Zhang 0001, Xuancheng Ren, Qi Su 0001, Xu Sun 0001 |
NAACL-HLT | 2 |
| 2021 | Topology-Imbalance Learning for Semi-Supervised Node ClassificationabstractThe class imbalance problem, as an important issue in learning node representations, has drawn increasing attention from the community. Although the imbalance considered by existing studies roots from the unequal quantity of labeled examples in different classes (quantity imbalance), we argue that graph data expose a unique source of imbalance from the asymmetric topological properties of the labeled nodes, i.e., labeled nodes are not equal in terms of their structural role in the graph (topology imbalance). In this work, we first probe the previously unknown topology-imbalance issue, including its characteristics, causes, and threats to semisupervised node classification learning. We then provide a unified view to jointly analyzing the quantity- and topology- imbalance issues by considering the node influence shift phenomenon with the Label Propagation algorithm. In light of our analysis, we devise an influence conflict detection–based metric Totoro to measure the degree of graph topology imbalance and propose a model-agnostic method ReNode to address the topology-imbalance issue by re-weighting the influence of labeled nodes adaptively based on their relative positions to class boundaries. Systematic experiments demonstrate the effectiveness and generalizability of our method in relieving topology-imbalance issue and promoting semi-supervised node classification. The further analysis unveils varied sensitivity of different graph neural networks (GNNs) to topology imbalance, which may serve as a new perspective in evaluating GNN architectures. Deli Chen, Yankai Lin 0001, Guangxiang Zhao, Xuancheng Ren, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001 |
NeurIPS | 4 |
| 2021 | Adversarial parameter defense by multi-step risk minimization
Zhiyuan Zhang 0001, Ruixuan Luo, Xuancheng Ren, Qi Su 0001, Liangyou Li, Xu Sun 0001 |
Neural Networks | 3 |
| 2020 | Rethinking Skip Connection with Layer NormalizationabstractSkip connection is a widely-used technique to improve the performance and the convergence of deep neural networks, which is believed to relieve the difficulty in optimization due to non-linearity by propagating a linear component through the neural network layers. However, from another point of view, it can also be seen as a modulating mechanism between the input and the output, with the input scaled by a pre-defined value one. In this work, we investigate how the scale factors in the effectiveness of the skip connection and reveal that a trivial adjustment of the scale will lead to spurious gradient exploding or vanishing in line with the deepness of the models, which could by addressed by normalization, in particular, layer normalization, which induces consistent improvements over the plain skip connection. Inspired by the findings, we further propose to adaptively adjust the scale of the input by recursively applying skip connection with layer normalization, which promotes the performance substantially and generalizes well across diverse tasks including both machine translation and image classification datasets. Xuancheng Ren, Zhiyuan Zhang 0001, Xu Sun 0001, Yuexian Zou |
COLING | 2 |
| 2020 | Regularizing Dialogue Generation by Imitating Implicit ScenariosabstractHuman dialogues are scenario-based and appropriate responses generally relate to the latent context knowledge entailed by the specific scenario.To enable responses that are more meaningful and context-specific, we propose to improve generative dialogue systems from the scenario perspective, where both dialogue history and future conversation are taken into account to implicitly reconstruct the scenario knowledge.More importantly, the conversation scenarios are further internalized using imitation learning framework, where the conventional dialogue model that has no access to future conversations is effectively regularized by transferring the scenario knowledge contained in hierarchical supervising signals from the scenario-based dialogue model, so that the future conversation is not required in actual inference.Extensive evaluations show that our approach significantly outperforms state-of-theart baselines on diversity and relevance, and expresses scenario-specific knowledge. Shaoxiong Feng, Xuancheng Ren, Hongshen Chen, Bin Sun 0004, Kan Li 0001, Xu Sun 0001 |
EMNLP (1) | 2 |
| 2020 | Prophet Attention: Predicting Attention with Future AttentionabstractRecently, attention based models have been used extensively in many sequence-to-sequence learning systems. Especially for image captioning, the attention based models are expected to ground correct image regions with proper generated words. However, for each time step in the decoding process, the attention based models usually use the hidden state of the current input to attend to the image regions. Under this setting, these attention models have a deviated focus'' problem that they calculate the attention weights based on previous words instead of the one to be generated, impairing the performance of both grounding and captioning. In this paper, we propose the Prophet Attention, similar to the form of self-supervision. In the training stage, this module utilizes the future information to calculate theideal'' attention weights towards image regions. These calculated ideal'' weights are further used to regularize thedeviated'' attention. In this manner, image regions are grounded with the correct words. The proposed Prophet Attention can be easily incorporated into existing image captioning models to improve their performance of both grounding and captioning. The experiments on the Flickr30k Entities and the MSCOCO datasets show that the proposed Prophet Attention consistently outperforms baselines in both automatic metrics and human evaluations. It is worth noticing that we set new state-of-the-arts on the two benchmark datasets and achieve the 1st place on the leaderboard of the online MSCOCO benchmark in terms of the default ranking score, i.e., CIDEr-c40. Xuancheng Ren, Xian Wu 0001, Shen Ge, Wei Fan 0001, Yuexian Zou, Xu Sun 0001 |
NeurIPS | 2 |
| 2020 | Memorized sparse backpropagation
Zhiyuan Zhang 0001, Xuancheng Ren, Qi Su 0001, Xu Sun 0001 |
Neurocomputing | 3 |
| 2020 | Training Simplification and Model Simplification for Deep Learning : A Minimal Effort Back Propagation MethodabstractWe propose a simple yet effective technique to simplify the training and the resulting model of neural networks. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-k elements (in terms of magnitude) are kept. As a result, only k rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction in the computational cost. Based on the sparsified gradients, we further simplify the model by eliminating the rows or columns that are seldom updated, which will reduce the computational cost both in the training and decoding, and potentially accelerate decoding in real-world applications. Surprisingly, experimental results demonstrate that most of the time we only need to update fewer than 5 percent of the weights at each back propagation pass. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given. The model simplification results show that we could adaptively simplify the model which could often be reduced by around 9x, without any loss on accuracy or even with improved accuracy. Xu Sun 0001, Xuancheng Ren, Shuming Ma, Bingzhen Wei, Wei Li 0101, Jingjing Xu 0001, Houfeng Wang, Yi Zhang 0050 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2019 | A Hierarchical Reinforced Sequence Operation Method for Unsupervised Text Style TransferabstractUnsupervised text style transfer aims to alter text styles while preserving the content, without aligned data for supervision.Existing seq2seq methods face three challenges: 1) the transfer is weakly interpretable, 2) generated outputs struggle in content preservation, and 3) the trade-off between content and style is intractable.To address these challenges, we propose a hierarchical reinforced sequence operation method, named Point-Then-Operate (PTO), which consists of a high-level agent that proposes operation positions and a lowlevel agent that alters the sentence.We provide comprehensive training objectives to control the fluency, style, and content of the outputs and a mask-based inference algorithm that allows for multi-step revision based on the single-step trained agents.Experimental results on two text style transfer datasets show that our method significantly outperforms recent methods and effectively addresses the aforementioned challenges. 1 Chen Wu 0005, Xuancheng Ren, Fuli Luo, Xu Sun 0001 |
ACL (1) | 2 |
| 2019 | Exploring and Distilling Cross-Modal Information for Image CaptioningabstractRecently, attention-based encoder-decoder models have been used extensively in image captioning. Yet there is still great difficulty for the current methods to achieve deep image understanding. In this work, we argue that such understanding requires visual attention to correlated image regions and semantic attention to coherent attributes of interest. To perform effective attention, we explore image captioning from a cross-modal perspective and propose the Global-and-Local Information Exploring-and-Distilling approach that explores and distills the source information in vision and language. It globally provides the aspect vector, a spatial and relational representation of images based on caption contexts, through the extraction of salient region groupings and attribute collocations, and locally extracts the fine-grained regions and attributes in reference to the aspect vector for word selection. Our fully-attentive model achieves a CIDEr score of 129.3 in offline COCO evaluation with remarkable efficiency in terms of accuracy, speed, and parameter budget. Xuancheng Ren, Yuanxin Liu, Kai Lei, Xu Sun 0001 |
IJCAI | 2 |
| 2019 | Aligning Visual Regions and Textual Concepts for Semantic-Grounded Image RepresentationsabstractIn vision-and-language grounding problems, fine-grained representations of the image are considered to be of paramount importance. Most of the current systems incorporate visual features and textual concepts as a sketch of an image. However, plainly inferred representations are usually undesirable in that they are composed of separate components, the relations of which are elusive. In this work, we aim at representing an image with a set of integrated visual regions and corresponding textual concepts, reflecting certain semantics. To this end, we build the Mutual Iterative Attention (MIA) module, which integrates correlated visual features and textual concepts, respectively, by aligning the two modalities. We evaluate the proposed approach on two representative vision-and-language grounding tasks, i.e., image captioning and visual question answering. In both tasks, the semantic-grounded image representations consistently boost the performance of the baseline models under all metrics across the board. The results demonstrate that our approach is effective and generalizes well to a wide range of models for image-related applications. (The code is available at \url{https://github.com/fenglinliu98/MIA) Yuanxin Liu, Xuancheng Ren, Xiaodong He 0001, Xu Sun 0001 |
NeurIPS | 3 |
| 2019 | Evaluating Semantic Rationality of a Sentence: A Sememe-Word-Matching Neural Network Based on HowNet
Jingjing Xu 0001, Xuancheng Ren |
NLPCC (1) | 3 |
| 2019 | Towards easier and faster sequence labeling for natural language processing: A search-based probabilistic online learning framework (SAPO)
Xu Sun 0001, Shuming Ma, Yi Zhang 0050, Xuancheng Ren |
Inf. Sci. | 4 |
| 2019 | Regularizing Output Distribution of Abstractive Chinese Social Media Text Summarization for Improved Semantic ConsistencyabstractAbstractive text summarization is a highly difficult problem, and the sequence-to-sequence model has shown success in improving the performance on the task. However, the generated summaries are often inconsistent with the source content in semantics. In such cases, when generating summaries, the model selects semantically unrelated words with respect to the source content as the most probable output. The problem can be attributed to heuristically constructed training data, where summaries can be unrelated to the source content, thus containing semantically unrelated words and spurious word correspondence. In this article, we propose a regularization approach for the sequence-to-sequence model and make use of what the model has learned to regularize the learning objective to alleviate the effect of the problem. In addition, we propose a practical human evaluation method to address the problem that the existing automatic evaluation method does not evaluate the semantic consistency with the source content properly. Experimental results demonstrate the effectiveness of the proposed approach, which outperforms almost all the existing models. Especially, the proposed approach improves the semantic consistency by 4% in terms of human evaluation. Bingzhen Wei, Xuancheng Ren, Yi Zhang 0050, Xiaoyan Cai, Qi Su 0001, Xu Sun 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2018 | Unpaired Sentiment-to-Sentiment Translation: A Cycled Reinforcement Learning ApproachabstractJingjing Xu, Xu Sun, Qi Zeng, Xiaodong Zhang, Xuancheng Ren, Houfeng Wang, Wenjie Li. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Jingjing Xu 0001, Xu Sun 0001, Qi Zeng 0001, Xiaodong Zhang 0022, Xuancheng Ren, Houfeng Wang, Wenjie Li 0002 |
ACL (1) | 5 |
| 2018 | Deconvolution-Based Global Decoding for Neural Machine TranslationabstractA great proportion of sequence-to-sequence (Seq2Seq) models for Neural Machine Translation (NMT) adopt Recurrent Neural Network (RNN) to generate translation word by word following a sequential order. As the studies of linguistics have proved that language is not linear word sequence but sequence of complex structure, translation at each step should be conditioned on the whole target-side context. To tackle the problem, we propose a new NMT model that decodes the sequence with the guidance of its structural prediction of the context of the target sequence. Our model generates translation based on the structural prediction of the target-side context so that the translation can be freed from the bind of sequential order. Experimental results demonstrate that our model is more competitive compared with the state-of-the-art methods, and the analysis reflects that our model is also robust to translating sentences of different lengths and it also reduces repetition with the instruction from the target-side context for decoding. Junyang Lin, Xu Sun 0001, Xuancheng Ren, Shuming Ma, Jinsong Su, Qi Su 0001 |
COLING | 3 |
| 2018 | Does Higher Order LSTM Have Better Accuracy for Segmenting and Labeling Sequence Data?abstractExisting neural models usually predict the tag of the current token independent of the neighboring tags. The popular LSTM-CRF model considers the tag dependencies between every two consecutive tags. However, it is hard for existing neural models to take longer distance dependencies between tags into consideration. The scalability is mainly limited by the complex model structures and the cost of dynamic programming during training. In our work, we first design a new model called “high order LSTM” to predict multiple tags for the current token which contains not only the current tag but also the previous several tags. We call the number of tags in one prediction as “order”. Then we propose a new method called Multi-Order BiLSTM (MO-BiLSTM) which combines low order and high order LSTMs together. MO-BiLSTM keeps the scalability to high order models with a pruning technique. We evaluate MO-BiLSTM on all-phrase chunking and NER datasets. Experiment results show that MO-BiLSTM achieves the state-of-the-art result in chunking and highly competitive results in two NER datasets. Yi Zhang 0050, Xu Sun 0001, Shuming Ma, Yang Yang 0125, Xuancheng Ren |
COLING | 5 |
| 2018 | Learning When to Concentrate or Divert Attention: Self-Adaptive Attention Temperature for Neural Machine TranslationabstractMost of the Neural Machine Translation (NMT) models are based on the sequence-tosequence (Seq2Seq) model with an encoderdecoder framework equipped with the attention mechanism.However, the conventional attention mechanism treats the decoding at each time step equally with the same matrix, which is problematic since the softness of the attention for different types of words (e.g.content words and function words) should differ.Therefore, we propose a new model with a mechanism called Self-Adaptive Control of Temperature (SACT) to control the softness of attention by means of an attention temperature.Experimental results on the Chinese-English translation and English-Vietnamese translation demonstrate that our model outperforms the baseline models, and the analysis and the case study show that our model can attend to the most relevant elements in the source-side contexts and generate the translation of high quality. Junyang Lin, Xu Sun 0001, Xuancheng Ren, Muyu Li, Qi Su 0001 |
EMNLP | 3 |
| 2018 | simNet: Stepwise Image-Topic Merging Network for Generating Detailed and Comprehensive Image CaptionsabstractThe encode-decoder framework has shown recent success in image captioning.Visual attention, which is good at detailedness, and semantic attention, which is good at comprehensiveness, have been separately proposed to ground the caption on the image.In this paper, we propose the Stepwise Image-Topic Merging Network (simNet) that makes use of the two kinds of attention at the same time.At each time step when generating the caption, the decoder adaptively merges the attentive information in the extracted topics and the image according to the generated context, so that the visual information and the semantic information can be effectively combined.The proposed approach is evaluated on two benchmark datasets and reaches the state-of-the-art performances.1 Xuancheng Ren, Yuanxin Liu, Houfeng Wang, Xu Sun 0001 |
EMNLP | 2 |
| 2018 | Diversity-Promoting GAN: A Cross-Entropy Based Generative Adversarial Network for Diversified Text GenerationabstractExisting text generation methods tend to produce repeated and "boring" expressions. To tackle this problem, we propose a new text generation model, called Diversity-Promoting Generative Adversarial Network (DP-GAN).The proposed model assigns low reward for repeatedly generated text and high reward for "novel" and fluent text, encouraging the generator to produce diverse and informative text.Moreover, we propose a novel languagemodel based discriminator, which can better distinguish novel text from repeated text without the saturation problem compared with existing classifier-based discriminators.The experimental results on review generation and dialogue generation tasks demonstrate that our model can generate substantially more diverse and informative text than existing baselines.1 Jingjing Xu 0001, Xuancheng Ren, Junyang Lin, Xu Sun 0001 |
EMNLP | 2 |
| 2018 | A Skeleton-Based Model for Promoting Coherence Among Sentences in Narrative Story GenerationabstractNarrative story generation is a challenging problem because it demands the generated sentences with tight semantic connections, which has not been well studied by most existing generative models.To address this problem, we propose a skeleton-based model to promote the coherence of generated stories.Different from traditional models that generate a complete sentence at a stroke, the proposed model first generates the most critical phrases, called skeleton, and then expands the skeleton to a complete and fluent sentence.The skeleton is not manually defined, but learned by a reinforcement learning method.Compared to the state-of-the-art models, our skeleton-based model can generate significantly more coherent text according to human evaluation and automatic evaluation.The G-score is improved by 20.1% in human evaluation. 1 Jingjing Xu 0001, Xuancheng Ren, Yi Zhang 0050, Qi Zeng 0001, Xiaoyan Cai, Xu Sun 0001 |
EMNLP | 2 |
| 2018 | A Hierarchical End-to-End Model for Jointly Improving Text Summarization and Sentiment ClassificationabstractText summarization and sentiment classification both aim to capture the main ideas of the text but at different levels. Text summarization is to describe the text within a few sentences, while sentiment classification can be regarded as a special type of summarization which ``summarizes'' the text into a even more abstract fashion, i.e., a sentiment class. Based on this idea, we propose a hierarchical end-to-end model for joint learning of text summarization and sentiment classification, where the sentiment classification label is treated as the further ``summarization'' of the text summarization output. Hence, the sentiment classification layer is put upon the text summarization layer, and a hierarchical structure is derived. Experimental results on Amazon online reviews datasets show that our model achieves better performance than the strong baseline systems on both abstractive summarization and sentiment classification. Shuming Ma, Xu Sun 0001, Junyang Lin, Xuancheng Ren |
IJCAI | 4 |
| 2018 | Building an Ellipsis-aware Chinese Dependency Treebank for Web Text
Xuancheng Ren, Xu Sun 0001, Ji Wen, Bingzhen Wei, Weidong Zhan, Zhiyuan Zhang 0001 |
LREC | 1 |
| 2018 | Query and Output: Generating Words by Querying Distributed Word Representations for Paraphrase GenerationabstractShuming Ma, Xu Sun, Wei Li, Sujian Li, Wenjie Li, Xuancheng Ren. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Shuming Ma, Xu Sun 0001, Wei Li 0101, Sujian Li, Wenjie Li 0002, Xuancheng Ren |
NAACL-HLT | 6 |
| 2018 | Cross-Domain and Semisupervised Named Entity Recognition in Chinese Social Media: A Unified ModelabstractNamed entity recognition (NER) in Chinese social media is an important, but challenging task because Chinese social media language is informal and noisy. Most previous methods on NER focus on in-domain supervised learning, which is limited by scarce annotated data in social media. In this paper, we present that sufficient corpora in formal domains and massive unannotated text can be combined to improve the NER performance in social media. We propose a unified model which can learn from out-of-domain corpora and in-domain unannotated text. The unified model is composed of two parts. One is for cross-domain learning and the other is for semisupervised learning. Cross-domain learning can learn out-of-domain information based on domain similarity. Semisupervised learning can learn in-domain unannotated information by self-training. Experimental results show that our unified model yields a 9.57% improvement over strong baselines and achieves the state-of-the-art performance. Jingjing Xu 0001, Hangfeng He 0001, Xu Sun 0001, Xuancheng Ren, Sujian Li |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | meProp: Sparsified Back Propagation for Accelerated Deep Learning with Reduced OverfittingabstractWe propose a simple yet effective technique for neural network learning. The forward propagation is computed as usual. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-$k$ elements (in terms of magnitude) are kept. As a result, only $k$ rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction ($k$ divided by the vector dimension) in the computational cost. Surprisingly, experimental results demonstrate that we can update only 1–4\% of the weights at each back propagation pass. This does not result in a larger number of training iterations. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given. Xu Sun 0001, Xuancheng Ren, Shuming Ma, Houfeng Wang |
ICML | 2 |