Xu Sun 0001

dblp:37/1971-1 · DBLP profile ↗
← Back
14ranked-venue papers in the field
9as first author
4since 2021 · last 2025
0000-0001-8241-9320ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 7 (3 first)Database Systems & Data Management · 3 (3 first)Information Retrieval & Web Search · 3 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
YearPublicationVenuePosition
2025 Rethinking Natural Language Generation with Layer-Wise Multi-View Decoding
abstract
In natural language generation, language models, particularly those based on decoder-only architectures as in popular Large Language Models (LLMs), have demonstrated impressive performance across a wide range of tasks. However, encoder-decoder architectures remain highly effective for tasks involving non-text data, such as images and time-series data. The decoder relies on the attention mechanism to efficiently extract information from the encoder. While it is common practice to draw information from only the last encoder layer, this might lead to insufficient training of the encoder layer stack due to the hierarchy bypassing problem. In this work, we propose layer-wise multi-view decoding for improved encoder-decoder language models, where for each decoder layer, together with the representations from the last encoder layer, which serve as a global view, those from other encoder layers are supplemented for a stereoscopic view of the source inputs. Systematic experiments and analyses show that we successfully address the hierarchy bypassing problem, require almost negligible parameter increase, and improve the performance of sequence learning with deep representations on diverse tasks, i.e., machine translation, abstractive summarization, image captioning, video captioning, medical report generation, and paraphrase generation. In particular, our approach achieves new state-of-the-art results on benchmark datasets, including a low-resource machine translation dataset and low-resource medical report generation datasets.
Xuancheng Ren, Guangxiang Zhao, Chenyu You, Sherry Ma, Xian Wu 0001, Wei Fan 0001, Xu Sun 0001
ACM Trans. Knowl. Discov. Data8
2022 Stock Trading Volume Prediction with Dual-Process Meta-Learning
Wei Li 0089, Zhiyuan Zhang 0001, Ruihan Bao, Keiko Harimoto, Xu Sun 0001
ECML/PKDD (6)6
2022 Distributional Correlation-Aware Knowledge Distillation for Stock Trading Volume Prediction
Lei Li 0039, Zhiyuan Zhang 0001, Ruihan Bao, Keiko Harimoto, Xu Sun 0001
ECML/PKDD (6)5
2022 DiMBERT: Learning Vision-Language Grounded Representations with Disentangled Multimodal-Attention
abstract
Vision-and-language (V-L) tasks require the system to understand both vision content and natural language, thus learning fine-grained joint representations of vision and language (a.k.a. V-L representations) is of paramount importance. Recently, various pre-trained V-L models are proposed to learn V-L representations and achieve improved results in many tasks. However, the mainstream models process both vision and language inputs with the same set of attention matrices. As a result, the generated V-L representations are entangled in one common latent space . To tackle this problem, we propose DiMBERT (short for Di sentangled M ultimodal-Attention BERT ), which is a novel framework that applies separated attention spaces for vision and language, and the representations of multi-modalities can thus be disentangled explicitly. To enhance the correlation between vision and language in disentangled spaces, we introduce the visual concepts to DiMBERT which represent visual information in textual format. In this manner, visual concepts help to bridge the gap between the two modalities. We pre-train DiMBERT on a large amount of image–sentence pairs on two tasks: bidirectional language modeling and sequence-to-sequence language modeling. After pre-train, DiMBERT is further fine-tuned for the downstream tasks. Experiments show that DiMBERT sets new state-of-the-art performance on three tasks (over four datasets), including both generation tasks (image captioning and visual storytelling) and classification tasks (referring expressions). The proposed DiM (short for Di sentangled M ultimodal-Attention) module can be easily incorporated into existing pre-trained V-L models to boost their performance, up to a 5% increase on the representative task. Finally, we conduct a systematic analysis and demonstrate the effectiveness of our DiM and the introduced visual concepts.
Xian Wu 0001, Shen Ge, Xuancheng Ren, Wei Fan 0001, Xu Sun 0001, Yuexian Zou
ACM Trans. Knowl. Discov. Data6
2020 Training Simplification and Model Simplification for Deep Learning : A Minimal Effort Back Propagation Method
abstract
We propose a simple yet effective technique to simplify the training and the resulting model of neural networks. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-k elements (in terms of magnitude) are kept. As a result, only k rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction in the computational cost. Based on the sparsified gradients, we further simplify the model by eliminating the rows or columns that are seldom updated, which will reduce the computational cost both in the training and decoding, and potentially accelerate decoding in real-world applications. Surprisingly, experimental results demonstrate that most of the time we only need to update fewer than 5 percent of the weights at each back propagation pass. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given. The model simplification results show that we could adaptively simplify the model which could often be reduced by around 9x, without any loss on accuracy or even with improved accuracy.
Xu Sun 0001, Xuancheng Ren, Shuming Ma, Bingzhen Wei, Wei Li 0101, Jingjing Xu 0001, Houfeng Wang, Yi Zhang 0050
IEEE Trans. Knowl. Data Eng.1
2019 Towards easier and faster sequence labeling for natural language processing: A search-based probabilistic online learning framework (SAPO)
Xu Sun 0001, Shuming Ma, Yi Zhang 0050, Xuancheng Ren
Inf. Sci.1
2013 A unified graph model for personalized query-oriented reference paper recommendation
abstract
With the tremendous amount of research publications, it has become increasingly important to provide a researcher with a rapid and accurate recommendation of a list of reference papers about a research field or topic. In this paper, we propose a unified graph model that can easily incorporate various types of useful information (e.g., content, authorship, citation and collaboration networks etc.) for efficient recommendation. The proposed model not only allows to thoroughly explore how these types of information can be better combined, but also makes personalized query-oriented reference paper recommendation possible, which as far as we know is a new issue that has not been explicitly addressed in the past. The experiments have demonstrated the clear advantages of personalized recommendation over non-personalized recommendation.
Fanqi Meng, Dehong Gao, Wenjie Li 0002, Xu Sun 0001, Yuexian Hou
CIKM4
2013 Probabilistic Chinese word segmentation with non-local information and stochastic training
Xu Sun 0001, Takuya Matsuzaki, Yoshimasa Tsuruoka, Jun'ichi Tsujii
Inf. Process. Manag.1
2013 Large-Scale Personalized Human Activity Recognition Using Online Multitask Learning
abstract
Personalized activity recognition usually has the problem of highly biased activity patterns among different tasks/persons. Traditional methods face problems on dealing with those conflicted activity patterns. We try to effectively model the activity patterns among different persons via casting this personalized activity recognition problem as a multitask learning issue. We propose a novel online multitask learning method for large-scale personalized activity recognition. In contrast with existing work of multitask learning that assumes fixed task relationships, our method can automatically discover task relationships from real-world data. Convergence analysis shows reasonable convergence properties of the proposed method. Experiments on two different activity data sets demonstrate that the proposed method significantly outperforms existing methods in activity recognition.
Xu Sun 0001, Hisashi Kashima, Naonori Ueda
IEEE Trans. Knowl. Data Eng.1
2013 Latent Structured Perceptrons for Large-Scale Learning with Hidden Information
abstract
Many real-world data mining problems contain hidden information (e.g., unobservable latent dependencies). We propose a perceptron-style method, latent structured perceptron, for fast discriminative learning of structured classification with hidden information. We also give theoretical analysis and demonstrate good convergence properties of the proposed method. Our method extends the perceptron algorithm for the learning task with hidden information, which can be hardly captured by traditional models. It relies on Viterbi decoding over latent variables, combined with simple additive updates. We perform experiments on one synthetic data set and two real-world structured classification tasks. Compared to conventional nonlatent models (e.g., conditional random fields, structured perceptrons), our method is more accurate on real-world tasks. Compared to existing heavy probabilistic models of latent variables (e.g., latent conditional random fields), our method lowers the training cost significantly (almost one order magnitude faster) yet with comparable or even superior classification accuracy. In addition, experiments demonstrate that the proposed method has good scalability on large-scale problems.
Xu Sun 0001, Takuya Matsuzaki, Wenjie Li 0002
IEEE Trans. Knowl. Data Eng.1
2012 Fast multi-task learning for query spelling correction
abstract
In this paper, we explore the use of a novel online multi-task learning framework for the task of search query spelling correction. In our procedure, correction candidates are initially generated by a ranker-based system and then re-ranked by our multi-task learning algorithm. With the proposed multi-task learning method, we are able to effectively transfer information from different and highly biased training datasets, for improving spelling correction on all datasets. Our experiments are conducted on three query spelling correction datasets including the well-known TREC benchmark dataset. The experimental results demonstrate that our proposed method considerably outperforms the existing baseline systems in terms of accuracy. Importantly, the proposed method is about one order of magnitude faster than baseline systems in terms of training speed. Compared to the commonly used online learning methods which typically require more than (e.g.,) 60 training passes, our proposed method is able to closely reach the empirical optimum in about 5 passes.
Xu Sun 0001, Anshumali Shrivastava, Ping Li 0001
CIKM1
2011 A New Multi-task Learning Method for Personalized Activity Recognition
abstract
Personalized activity recognition usually faces the problem of data sparseness. We aim at improving accuracy of personalized activity recognition by incorporating the information from other persons. We propose a new online multi-task learning method for personalized activity recognition. The proposed online multi-task learning method automatically learns the ``transfer-factors" (similarities) among different tasks (i.e., among different persons in our case). Experiments demonstrate that the proposed method significantly outperforms existing methods. The novelty of this paper is twofold: (1) A new multi-task learning framework, which can naturally learn similarities among tasks, (2) To our knowledge, this is the first study of large-scale personalized activity recognition.
Xu Sun 0001, Hisashi Kashima, Ryota Tomioka, Naonori Ueda, Ping Li 0001
ICDM1
2011 Large Scale Real-Life Action Recognition Using Conditional Random Fields with Stochastic Training
Xu Sun 0001, Hisashi Kashima, Ryota Tomioka, Naonori Ueda
PAKDD (2)1
2010 Averaged Stochastic Gradient Descent with Feedback: An Accurate, Robust, and Fast Training Method
abstract
On large datasets, the popular training approach has been stochastic gradient descent (SGD). This paper proposes a modification of SGD, called averaged SGD with feedback (ASF), that significantly improves the performance (robustness, accuracy, and training speed) over the traditional SGD. The proposal is based on three simple ideas: averaging the weight vectors across SGD iterations, feeding the averaged weights back into the SGD update process, and deciding when to perform the feedback (linearly slowing down feedback). Theoretically, we demonstrate the reasonable convergence properties of the ASF. Empirically, the ASF outperforms several strong baselines in terms of accuracy, robustness over the noise, and the training speed. To our knowledge, this is the first study of ``feedback'' in stochastic gradient learning. Although we choose latent conditional models for verifying the ASF in this paper, the ASF is a general purpose technique just like SGD, and can be directly applied to other models.
Xu Sun 0001, Hisashi Kashima, Takuya Matsuzaki, Naonori Ueda
ICDM1