VLDB 2026 Research / reviewers in the wild / expert
Tingyu Xia
dblp:266/4705
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language ModelsabstractOne core capability of large language models~(LLMs) is to follow natural language instructions. However, the issue of automatically constructing high-quality training data to enhance the complex instruction-following abilities of LLMs without manual annotation remains unresolved. In this paper, we introduce AutoIF, the first scalable and reliable method for automatically generating instruction-following training data. AutoIF transforms the validation of instruction-following data quality into code verification, requiring LLMs to generate instructions, the corresponding code to verify the correctness of the instruction responses, and unit test samples to cross-validate the code's correctness. Then, execution feedback-based rejection sampling can generate data for Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) training. AutoIF achieves significant improvements across three training algorithms, SFT, Offline DPO, and Online DPO, when applied to the advanced open-source LLMs, Qwen2 and LLaMA3, in self-alignment and strong-to-weak distillation settings. Using two widely-used and three challenging general instruction-following benchmarks, we demonstrate that AutoIF significantly improves LLM performance across a wide range of natural instruction constraints. Notably, AutoIF is the first to surpass 90\% accuracy in IFEval’s loose instruction accuracy, without compromising general, math and coding capabilities. Further analysis of quality, scaling, combination, and data efficiency highlights AutoIF's strong generalization and alignment potential. Our code are available at https://github.com/QwenLM/AutoIF Guanting Dong 0001, Keming Lu, Chengpeng Li 0001, Tingyu Xia, Bowen Yu 0002, Chang Zhou 0005, Jingren Zhou 0001 |
ICLR | 4 |
| 2025 | A survey of RWKV
Tingyu Xia, Yi Chang 0001, Yuan Wu 0002 |
Neurocomputing | 2 |
| 2025 | Selective fine-tuning for large language models via matrix nuclear norm
Tingyu Xia, Yahan Li, Yuan Wu 0002, Yi Chang 0001 |
Inf. Process. Manag. | 1 |
| 2024 | Enhancing inter-sentence attention for Semantic Textual Similarity
Ying Zhao 0034, Tingyu Xia, Yunqi Jiang, Yuan Tian 0016 |
Inf. Process. Manag. | 2 |
| 2024 | DAW-GAN: a generative adversarial network based on the dynamic adaptive weight for image super-resolution
Tingyu Xia, Xin Yang 0002, Yitian Zhu |
Multim. Tools Appl. | 1 |
| 2023 | Few-shot relation classification using clustering-based prototype modification
Mingtong Wen, Tingyu Xia, Bowen Liao, Yuan Tian 0016 |
Knowl. Based Syst. | 2 |
| 2022 | FastClass: A Time-Efficient Approach to Weakly-Supervised Text ClassificationabstractWeakly-supervised text classification aims to train a classifier using only class descriptions and unlabeled data.Recent research shows that keyword-driven methods can achieve state-ofthe-art performance on various tasks.However, these methods not only rely on carefullycrafted class descriptions to obtain classspecific keywords but also require substantial amount of unlabeled data and takes a long time to train.This paper proposes FastClass, an efficient weakly-supervised classification approach.It uses dense text representation to retrieve class-relevant documents from external unlabeled corpus and selects an optimal subset to train a classifier.Compared to keyworddriven methods, our approach is less reliant on initial class descriptions as it no longer needs to expand each class description into a set of class-specific keywords.Experiments on a wide range of classification tasks show that the proposed approach frequently outperforms keyword-driven models in terms of classification accuracy and often enjoys orders-ofmagnitude faster training speed. Tingyu Xia, Yue Wang 0035, Yuan Tian 0016, Yi Chang 0001 |
EMNLP | 1 |
| 2021 | Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching TasksabstractWe study the problem of incorporating prior knowledge into a deep Transformer-based model, i.e., Bidirectional Encoder Representations from Transformers (BERT), to enhance its performance on semantic textual matching tasks. By probing and analyzing what BERT has already known when solving this task, we obtain better understanding of what task-specific knowledge BERT needs the most and where it is most needed. The analysis further motivates us to take a different approach than most existing works. Instead of using prior knowledge to create a new training task for fine-tuning BERT, we directly inject knowledge into BERT’s multi-head attention mechanism. This leads us to a simple yet effective approach that enjoys fast training stage as it saves the model from training on additional data or tasks other than the main task. Extensive experiments demonstrate that the proposed knowledge-enhanced BERT is able to consistently improve semantic textual matching performance over the original BERT model, and the performance benefit is most salient when training data is scarce. Tingyu Xia, Yue Wang 0035, Yuan Tian 0016, Yi Chang 0001 |
WWW | 1 |
| 2020 | Unsupervised Nonlinear Feature Selection from High-Dimensional Signed NetworksabstractWith the rapid development of social media services in recent years, relational data are explosively growing. The signed network, which consists of a mixture of positive and negative links, is an effective way to represent the friendly and hostile relations among nodes, which can represent users or items. Because the features associated with a node of a signed network are usually incomplete, noisy, unlabeled, and high-dimensional, feature selection is an important procedure to eliminate irrelevant features. However, existing network-based feature selection methods are linear methods, which means they can only select features that having the linear dependency on the output values. Moreover, in many social data, most nodes are unlabeled; therefore, selecting features in an unsupervised manner is generally preferred. To this end, in this paper, we propose a nonlinear unsupervised feature selection method for signed networks, called SignedLasso. This method can select a small number of important features with nonlinear associations between inputs and output from a high-dimensional data. More specifically, we formulate unsupervised feature selection as a nonlinear feature selection problem with the Hilbert-Schmidt Independence Criterion Lasso (HSIC Lasso), which can find a small number of features in a nonlinear manner. Then, we propose the use of a deep learning-based node embedding to represent node similarity without label information and incorporate the node embedding into the HSIC Lasso. Through experiments on two real world datasets, we show that the proposed algorithm is superior to existing linear unsupervised feature selection methods. Tingyu Xia, Huiyan Sun, Makoto Yamada, Yi Chang 0001 |
AAAI | 2 |