VLDB 2026 Research / reviewers in the wild / expert
Yi Xu 0003
dblp:14/5580-3
· DBLP profile ↗
19ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0004-2076-5374ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IBCB: Efficient Inverse Batched Contextual Bandit for Behavioral Evolution HistoryabstractTraditional imitation learning focuses on modeling the behavioral mechanisms of experts, which requires a large amount of interaction history generated by some fixed expert. However, in many streaming applications, such as streaming recommender systems, online decision-makers typically engage in online learning during the decision-making process, meaning that the interaction history generated by online decision-makers includes their behavioral evolution from novice expert to experienced expert. This poses a new challenge for existing imitation learning approaches that can only utilize data from experienced experts. To address this issue, this paper proposes an inverse batched contextual bandit (IBCB) framework that can efficiently perform estimations of environment reward parameters and learned policy based on the expert's behavioral evolution history. Specifically, IBCB formulates the inverse problem into a simple quadratic programming problem by utilizing the behavioral evolution history of the batched contextual bandit with inaccessible rewards, and it can be extended to fairness-aware expert limitation. We demonstrate that IBCB is a unified framework for both deterministic and randomized bandit policies. The experimental results indicate that IBCB outperforms several existing imitation learning algorithms on synthetic and real-world data and significantly reduces running time. Additionally, empirical analyses reveal that IBCB exhibits better imitation ability for fairness-aware experts, out-of-distribution generalization and is highly effective in learning the bandit policy from the interaction history of novice experts. The code is publicly available. Yi Xu 0003, Weiran Shen, Jun Xu 0001, Xiao Zhang 0034, Ji-Rong Wen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | NeRF-MIR: Toward High-Quality Restoration of Masked Images With Neural Radiance FieldsabstractNeural radiance fields (NeRFs) have demonstrated remarkable performance in novel view synthesis. However, there is much room for restoring 3-D scenes based on NeRF from corrupted images, which are common in natural scene captures and can significantly impact the effectiveness of NeRF. This article introduces NeRF-MIR, a novel neural rendering approach specifically proposed for the restoration of masked images, demonstrating the potential of NeRF in this domain. Recognizing that randomly emitting rays to pixels in NeRF may not effectively learn intricate image textures, we propose a patch-based entropy for ray emitting (PERE) strategy to distribute emitted rays properly. This enables NeRF-MIR to fuse comprehensive information from images of different views. In addition, we introduce a progressively iterative restoration (PIRE) mechanism to restore the masked regions in a self-training process. Furthermore, we design a dynamically weighted loss function that automatically recalibrates the loss weights for masked regions. As existing datasets do not support NeRF-based masked image restoration, we construct three masked datasets to simulate corrupted scenarios. Extensive experiments on real data and constructed datasets demonstrate the superiority of NeRF-MIR over its counterparts in masked image restoration. Xianliang Huang, Zhizhou Zhong, Yi Xu 0003, Jihong Guan, Shuigeng Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | MAPS: Motivation-Aware Personalized Search via LLM-Driven Consultation AlignmentabstractWeicong Qin, Yi Xu, Weijie Yu, Chenglei Shen, Ming He, Jianping Fan, Xiao Zhang, Jun Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weicong Qin, Yi Xu 0003, Weijie Yu 0003, Chenglei Shen, Jianping Fan 0001, Xiao Zhang 0034, Jun Xu 0001 |
ACL (1) | 2 |
| 2025 | Similarity = Value? Consultation Value-Assessment and Alignment for Personalized SearchabstractPersonalized search systems in e-commerce platforms increasingly involve user interactions with AI assistants, where users consult about products, usage scenarios, and more.Leveraging consultation to personalize search services is trending.Existing methods typically rely on semantic similarity to align historical consultations with current queries due to the absence of 'value' labels, but we observe that semantic similarity alone often fails to capture the true value of consultation for personalization.To address this, we propose a consultation value assessment framework that evaluates historical consultations from three novel perspectives: (1) Scenario Scope Value, (2) Posterior Action Value, and (3) Time Decay Value.Based on this, we introduce VAPS, a value-aware personalized search model that selectively incorporates high-value consultations through a consultation-user action interaction module and an explicit objective that aligns consultations with user actions.Experiments on both public and commercial datasets show that VAPS consistently outperforms baselines in both retrieval and ranking tasks.Codes are available at https://github.com/E-qin/VAPS. Weicong Qin, Yi Xu 0003, Weijie Yu 0003, Teng Shi, Chenglei Shen, Jianping Fan 0001, Xiao Zhang 0034, Jun Xu 0001 |
EMNLP | 2 |
| 2025 | MoRE: A Mixture of Reflectors Framework for Large Language Model-Based Sequential Recommendation
Weicong Qin, Yi Xu 0003, Weijie Yu 0003, Chenglei Shen, Xiao Zhang 0034, Jianping Fan 0001, Jun Xu 0001 |
RecSys | 2 |
| 2025 | HiREN: Towards higher supervision quality for better scene text image super-resolution
Minyi Zhao, Yi Xu 0003, Bingjia Li, Jihong Guan, Shuigeng Zhou |
Neurocomputing | 2 |
| 2023 | Mining and Applying Composition Knowledge of Dance Moves for Style-Concentrated Dance GenerationabstractChoreography refers to creation of dance motions according to both music and dance knowledge, where the created dances should be style-specific and consistent. However, most of the existing methods generate dances using the given music as the only reference, lacking the stylized dancing knowledge, namely, the flag motion patterns contained in different styles. Without the stylized prior knowledge, these approaches are not promising to generate controllable style or diverse moves for each dance style, nor new dances complying with stylized knowledge. To address this issue, we propose a novel music-to-dance generation framework guided by style embedding, considering both input music and stylized dancing knowledge. These style embeddings are learnt representations of style-consistent kinematic abstraction of reference dance videos, which can act as controllable factors to impose style constraints on dance generation in a latent manner. Hence, we can make the style embedding fit into any given style while allowing the flexibility to generate new compatible dance moves by modifying the style embedding according to the learnt representations of a certain style. We are the first to achieve knowledge-driven style control in dance generation tasks. To support this study, we build a large multi-style music-to-dance dataset referred to as I-Dance. The qualitative and quantitative evaluations demonstrate the advantage of the proposed framework, as well as the ability to synthesize diverse moves under a dance style directed by style embedding. Xinjian Zhang, Su Yang 0001, Yi Xu 0003, Weishan Zhang, Longwen Gao |
AAAI | 3 |
| 2022 | EPiDA: An Easy Plug-in Data Augmentation Framework for High Performance Text ClassificationabstractMinyi Zhao, Lu Zhang, Yi Xu, Jiandong Ding, Jihong Guan, Shuigeng Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Minyi Zhao, Lu Zhang 0060, Yi Xu 0003, Jiandong Ding, Jihong Guan, Shuigeng Zhou |
NAACL-HLT | 3 |
| 2021 | GIF Thumbnails: Attract More Clicks to Your VideosabstractWith the rapid increase of mobile devices and online media, more and more people prefer posting/viewing videos online. Generally, these videos are presented on video streaming sites with image thumbnails and text titles. While facing huge amounts of videos, a viewer clicks through a certain video with high probability because of its eye-catching thumbnail. However, current video thumbnails are created manually, which is time-consuming and quality-unguaranteed. And static image thumbnails contain very limited information of the corresponding videos, which prevents users from successfully clicking what they really want to view. In this paper, we address a novel problem, namely GIF thumbnail generation, which aims to automatically generate GIF thumbnails for videos and consequently boost their Click-Through-Rate (CTR). Here, a GIF thumbnail is an animated GIF file consisting of multiple segments from the video, containing more information of the target video than a static image thumbnail. To support this study, we build the first GIF thumbnails benchmark dataset that consists of 1070 videos covering a total duration of 69.1 hours, and 5394 corresponding manually-annotated GIFs. To solve this problem, we propose a learning-based automatic GIF thumbnail generation model, which is called Generative Variational Dual-Encoder (GEVADEN). As not relying on any user interaction information (e.g. time-sync comments and real-time view counts), this model is applicable to newly-uploaded/rarely-viewed videos. Experiments on our built dataset show that GEVADEN significantly outperforms several baselines, including video-summarization and highlight-detection based ones. Furthermore, we develop a pilot application of the proposed model on an online video platform with 9814 videos covering 1231 hours, which shows that our model achieves a 37.5% CTR improvement over traditional image thumbnails. This further validates the effectiveness of the proposed model and the promising application prospect of GIF thumbnails. Yi Xu 0003, Fan Bai 0001, Yingxuan Shi, Qiuyu Chen, Longwen Gao, Kai Tian 0001, Shuigeng Zhou, Huyang Sun |
AAAI | 1 |
| 2021 | Weakly-supervised Text Classification Based on Keyword GraphabstractWeakly-supervised text classification has received much attention in recent years for it can alleviate the heavy burden of annotating massive data.Among them, keyword-driven methods are the mainstream where user-provided keywords are exploited to generate pseudolabels for unlabeled texts.However, existing methods treat keywords independently, thus ignore the correlation among them, which should be useful if properly exploited.In this paper, we propose a novel framework called ClassKG to explore keyword-keyword correlation on keyword graph by GNN.Our framework is an iterative process.In each iteration, we first construct a keyword graph, so the task of assigning pseudo labels is transformed to annotating keyword subgraphs.To improve the annotation quality, we introduce a self-supervised task to pretrain a subgraph annotator, and then finetune it.With the pseudo labels generated by the subgraph annotator, we then train a text classifier to classify the unlabeled texts.Finally, we re-extract keywords from the classified texts.Extensive experiments on both long-text and short-text datasets show that our method substantially outperforms the existing ones. Lu Zhang 0060, Jiandong Ding, Yi Xu 0003, Yingyao Liu, Shuigeng Zhou |
EMNLP (1) | 3 |
| 2021 | Recursive Fusion and Deformable Spatiotemporal Attention for Video Compression Artifact ReductionabstractA number of deep learning based algorithms have been proposed to recover high-quality videos from low-quality compressed ones. Among them, some restore the missing details of each frame via exploring the spatiotemporal information of neighboring frames. However, these methods usually suffer from a narrow temporal scope, thus may miss some useful details from some frames outside the neighboring ones. In this paper, to boost artifact removal, on the one hand, we propose a Recursive Fusion (RF) module to model the temporal dependency within a long temporal range. Specifically, RF utilizes both the current reference frames and the preceding hidden state to conduct better spatiotemporal compensation. On the other hand, we design an efficient and effective Deformable Spatiotemporal Attention (DSTA) module such that the model can pay more effort on restoring the artifact-rich areas like the boundary area of a moving object. Extensive experiments show that our method outperforms the existing ones on the MFQE 2.0 dataset in terms of both fidelity and perceptual effect. Code is available at https://github.com/zhaominyiz/RFDA-PyTorch. Minyi Zhao, Yi Xu 0003, Shuigeng Zhou |
ACM Multimedia | 2 |
| 2021 | DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled SamplesabstractThe scarcity of labeled data is a critical obstacle to deep learning. Semi-supervised learning (SSL) provides a promising way to leverage unlabeled data by pseudo labels. However, when the size of labeled data is very small (say a few labeled samples per class), SSL performs poorly and unstably, possibly due to the low quality of learned pseudo labels. In this paper, we propose a new SSL method called DP-SSL that adopts an innovative data programming (DP) scheme to generate probabilistic labels for unlabeled data. Different from existing DP methods that rely on human experts to provide initial labeling functions (LFs), we develop a multiple-choice learning~(MCL) based approach to automatically generate LFs from scratch in SSL style. With the noisy labels produced by the LFs, we design a label model to resolve the conflict and overlap among the noisy labels, and finally infer probabilistic labels for unlabeled samples. Extensive experiments on four standard SSL benchmarks show that DP-SSL can provide reliable labels for unlabeled data and achieve better classification performance on test sets than existing SSL methods, especially when only a small number of labeled samples are available. Concretely, for CIFAR-10 with only 40 labeled samples, DP-SSL achieves 93.82% annotation accuracy on unlabeled data and 93.46% classification accuracy on test data, which are higher than the SOTA results. Yi Xu 0003, Jiandong Ding, Lu Zhang 0060, Shuigeng Zhou |
NeurIPS | 1 |
| 2020 | Network as Regularization for Training Deep Neural Networks: Framework, Model and PerformanceabstractDespite powerful representation ability, deep neural networks (DNNs) are prone to over-fitting, because of over-parametrization. Existing works have explored various regularization techniques to tackle the over-fitting problem. Some of them employed soft targets rather than one-hot labels to guide network training (e.g. label smoothing in classification tasks), which are called target-based regularization approaches in this paper. To alleviate the over-fitting problem, here we propose a new and general regularization framework that introduces an auxiliary network to dynamically incorporate guided semantic disturbance to the labels. We call it Network as Regularization (NaR in short). During training, the disturbance is constructed by a convex combination of the predictions of the target network and the auxiliary network. These two networks are initialized separately. And the auxiliary network is trained independently from the target network, while providing instance-level and class-level semantic information to the latter progressively. We conduct extensive experiments to validate the effectiveness of the proposed method. Experimental results show that NaR outperforms many state-of-the-art target-based regularization methods, and other regularization approaches (e.g. mixup) can also benefit from combining with NaR. Kai Tian 0001, Yi Xu 0003, Jihong Guan, Shuigeng Zhou |
AAAI | 2 |
| 2020 | Adaptive Fractional Dilated Convolution Network for Image Aesthetics AssessmentabstractTo leverage deep learning for image aesthetics assessment, one critical but unsolved issue is how to seamlessly incorporate the information of image aspect ratios to learn more robust models. In this paper, an adaptive fractional dilated convolution (AFDC), which is aspect-ratio-embedded, composition-preserving and parameter-free, is developed to tackle this issue natively in convolutional kernel level. Specifically, the fractional dilated kernel is adaptively constructed according to the image aspect ratios, where the interpolation of nearest two integer dilated kernels are used to cope with the misalignment of fractional sampling. Moreover, we provide a concise formulation for mini-batch training and utilize a grouping strategy to reduce computational overhead. As a result, it can be easily implemented by common deep learning libraries and plugged into popular CNN architectures in a computation-efficient manner. Our experimental results demonstrate that our proposed method achieves state-of-the-art performance on image aesthetics assessment over the AVA dataset. Qiuyu Chen, Wei Zhang 0016, Yi Xu 0003, Yu Zheng 0006, Jianping Fan 0001 |
CVPR | 5 |
| 2020 | Optimal Trade Execution Based on Deep Deterministic Policy Gradient
Zekun Ye, Weijie Deng, Shuigeng Zhou, Yi Xu 0003, Jihong Guan |
DASFAA (1) | 4 |
| 2019 | Versatile Multiple Choice Learning and Its Application to Vision ComputingabstractMost existing ensemble methods aim to train the underlying embedded models independently and simply aggregate their final outputs via averaging or weighted voting. As many prediction tasks contain uncertainty, most of these ensemble methods just reduce variance of the predictions without considering the collaborations among the ensembles. Different from these ensemble methods, multiple choice learning (MCL) methods exploit the cooperation among all the embedded models to generate multiple diverse hypotheses. In this paper, a new MCL method, called vMCL (the abbreviation of versatile Multiple Choice Learning), is developed to extend the application scenarios of MCL methods by ensembling deep neural networks. Our vMCL method keeps the advantage of existing MCL methods while overcoming their major drawback, thus achieves better performance. The novelty of our vMCL lies in three aspects: (1) a choice network is designed to learn the confidence level of each specialist which can provide the best prediction base on multiple hypotheses; (2) a hinge loss is introduced to alleviate the overconfidence issue in MCL settings; (3) Easy to be implemented and can be trained in an end-to-end manner, which is a very attractive feature for many real-world applications. Experiments on image classification and image segmentation task show that vMCL outperforms the existing state-of-the-art MCL methods. Kai Tian 0001, Yi Xu 0003, Shuigeng Zhou, Jihong Guan |
CVPR | 2 |
| 2019 | Non-Local ConvLSTM for Video Compression Artifact ReductionabstractVideo compression artifact reduction aims to recover high-quality videos from low-quality compressed videos. Most existing approaches use a single neighboring frame or a pair of neighboring frames (preceding and/or following the target frame) for this task. Furthermore, as frames of high quality overall may contain low-quality patches, and high-quality patches may exist in frames of low quality overall, current methods focusing on nearby peak-quality frames (PQFs) may miss high-quality details in low-quality frames. To remedy these shortcomings, in this paper we propose a novel end-to-end deep neural network called non-local ConvLSTM (NL-ConvLSTM in short) that exploits multiple consecutive frames. An approximate non-local strategy is introduced in NL-ConvLSTM to capture global motion patterns and trace the spatiotemporal dependency in a video sequence. This approximate strategy makes the non-local module work in a fast and low space-cost way. Our method uses the preceding and following frames of the target frame to generate a residual, from which a higher quality frame is reconstructed. Experiments on two datasets show that NL-ConvLSTM outperforms the existing methods. Yi Xu 0003, Longwen Gao, Kai Tian 0001, Shuigeng Zhou, Huyang Sun |
ICCV | 1 |
| 2019 | Simple Is Better: A Global Semantic Consistency Based End-to-End Framework for Effective Zero-Shot Learning
Shuigeng Zhou, Yi Xu 0003, Jihong Guan, Jun Huan |
PRICAI (1) | 4 |
| 2018 | Effective community division based on improved spectral clustering
Yi Xu 0003, Zhi Zhuang, Weimin Li 0001, Xiaokang Zhou |
Neurocomputing | 1 |