VLDB 2026 Research / reviewers in the wild / expert
Yao-Xiang Ding 0001
dblp:186/8301-1 · also Yaoxiang Ding 0001
· DBLP profile ↗
17ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0001-8580-1103ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Harnessing the Power of Reinforcement Learning for Language-Model-Based Information Retriever via Query-Document Co-Augmentation
Jingming Liu, Yao-Xiang Ding 0001, Hui Su, Kun Zhou 0001 |
PAKDD (3) | 4 |
| 2026 | Expressive Head Avatar Modeling From Monocular Video of Neutral ExpressionabstractWe study the reconstruction of high-quality 3D head avatars. Our goal is to reduce the reliance on dense capture data for most existing approaches, which limits their practicality. Recent advances address this using single or few input images by either training a prior model or fine-tuning multi-view diffusion models to generate pseudo training points. While they fall short in producing multi-view-consistent, high-fidelity results aligned with the input data. This motivates us to explore a more practical and user-friendly input setting. Modern smartphones such as Apple's Face ID already guide users to slowly rotate their heads in front of a single camera, enabling the capture of facial data across varying viewpoints with minimal effort. This simple and intuitive scanning motion has become a widely accepted user habit and provides sufficient geometric information-highlighting a natural opportunity for 3D head avatar creation from monocular videos of neutral expression. Under this settinng, we introduce R$^{2}$2 Avatar, a lightweight and user-friendly framework for generating expressive 3D head avatars, which adopts a Reconstruction-by-Restoration strategy avoiding large-scale model pretraining while achieving high-quality animatable avatars. Specifically, a geometry-guided warping module first synthesizes coarse expression variaHons from the neutral input. Then, a restoration module refines the warped results by recovering highfrequency facial details, including mouth interior, with the help of a data-driven 2D animation prior. These restored images serve as supervision targets to optimize the final avatar. Experiments demonstrate that our method produces realistic avatars with improved expression diversity and view consistency compared to baseline approaches. Yao-Xiang Ding 0001, Kun Zhou 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Enhancing Identity-Deformation Disentanglement in StyleGAN for One-Shot Face Video Re-EnactmentabstractThe task of one-shot face video re-enactment aims at generating target video of faces with the same identity of one source frame and facial deformation of the driving video. To achieve high quality generation, it is essential to precisely disentangle identity-related and identity-independent characteristics, meanwhile build expressive features keeping high-frequency facial details, which still remain unaddressed for existing approaches. To deal with these two challenges, we propose a two-stage generation model based on StyleGAN, whose key novel techniques lie in better disentangling identity and deformation codes in the latent space through an identity-based modeling and manipulating intermediate StyleGAN features at the second stage for augmenting facial details of the generating targets. To further improve identity consistency, a data augmentation method is introduced during training for enhancing the key features affecting identity such as hair and wrinkles. Extensive experimental results demonstrate the superiority of our approach compared to state-of-the-art methods. Yao-Xiang Ding 0001, Kun Zhou 0001 |
AAAI | 2 |
| 2025 | Achieving Nearly-Optimal Regret and Sample Complexity in Dueling Bandits with Applications in Online RecommendationsabstractWe focus on the dueling bandits problem, which has recently drawn significant attention due to its wide-ranging applications in online recommendation systems and the alignment of large language models (LLMs), considers an online preference learning scenario where the learner iteratively selects arms based on pairwise comparison feedback to infer user preferences. Two primary objectives are typically considered in dueling bandits: Regret Minimization (RM), which aims to improve the overall quality of selected arms over time, and Best Arm Identification (BAI), which seeks to efficiently identify the best item with minimal user feedback. For instance, RM is exemplified by the objective of consistently providing high-quality items, while BAI reduces the required human feedback by minimizing the number of necessary comparisons. Conventional research treats RM and BAI as two conflicting objectives, optimizing one at the expense of the other. In this paper, we propose a novel framework that demonstrates the near-consistency of RM and BAI in dueling bandits by reducing the BAI in dueling bandits into a sequential noisy identification problem. Based on our formulation, we propose a black-box reduction technique that transforms any RM algorithm into a BAI algorithm, and prove that such reduction with optimal RM algorithm achieves optimal sample complexity and nearly-optimal cumulative weak regret simultaneously. Our proposed algorithm acheives a nearly-optimal BAI sample complexity and attains a cumulative weak regret that is order-wise equivalent to the best-known result simultaneously. Experiments on both synthetic benchmarks and real-world online recommendation tasks validate the effectiveness of the proposed method, providing empirical evidences for our theoretical findings. Lanjihong Ma, Yao-Xiang Ding 0001, Zhen-Yu Zhang, Zhi-Hua Zhou |
KDD (1) | 2 |
| 2025 | Generating by Understanding: Neural Visual Generation with Logical Symbol GroundingsabstractMaking neural visual generative models controllable by logical reasoning systems is promising for improving faithfulness, transparency, and generalizability. We propose the Abductive visual Generation (AbdGen) approach to build such logic-integrated models. A vector-quantized symbol grounding mechanism and the corresponding disentanglement training method are introduced to enhance the controllability of logical symbols over generation. Furthermore, we propose two logical abduction methods to make our approach require few labeled training data and support the induction of latent logical generative rules from data. We experimentally show that our approach can be utilized to integrate various neural generative models with logical reasoning systems, by both learning from scratch or utilizing pre-trained models directly. The code is released at https://github.com/future-item/AbdGen. Yifei Peng, Zijie Zha, Zhexu Luo, Wang-Zhou Dai, Zhong Ren 0001, Yao-Xiang Ding 0001, Kun Zhou 0001 |
KDD (2) | 7 |
| 2025 | Appearance as reliable evidence: Reconciling appearance and generative priors for monocular motion estimation
Zipei Chen, Zhong Ren 0001, Yao-Xiang Ding 0001, Kun Zhou 0001 |
Comput. Graph. | 4 |
| 2025 | Learning Objective Adaptation by Correlation-Based Model ReuseabstractIn open-environment machine learning (open ML), the learning objectives can vary according to specific real-world requirements. Models tailored for initial objectives may not be appropriate for the varied objectives. Retraining models from scratch for every single objective can be computationally intensive. Therefore, it is desirable to reuse models trained on the original objectives to help learn under the varied objectives. To this end, it is essential to characterize the objective correlations to better reuse the models. Previous works only consider the relative importance between pairs of previous and varied objectives, also known as previous-varied objectives correlations, ignoring correlations among the original objectives themselves. In this article, we demonstrate the importance of cross-original objective correlations. We propose a novel approach that employs the optimal transport technique to model correlations across all previous and varied objectives and then facilitates model reuse by utilizing learned transportation discrepancies to incorporate model reusabilities. Our empirical results show that our approach significantly outperforms existing benchmarks and well captures the underlying objective structure, validating the importance of accurate objective correlation modeling for learning with varied objectives. Lanjihong Ma, Yao-Xiang Ding 0001, Peng Zhao 0006, Zhi-Hua Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Learning Only When It Matters: Cost-Aware Long-Tailed ClassificationabstractMost current long-tailed classification approaches assume the cost-agnostic scenario, where the training distribution of classes is long-tailed while the testing distribution of classes is balanced. Meanwhile, the misclassification costs of all instances are the same. On the other hand, in many real-world applications, it is more proper to assume that the training and testing distributions of classes are the same, while the misclassification cost of tail-class instances is varied. In this work, we model such a scenario as cost-aware long-tailed classification, in which the identification of high-cost tail instances and focusing learning on them thereafter is essential. In consequence, we propose the learning strategy of augmenting new instances based on adaptive region partition in the feature space. We conduct theoretical analysis to show that under the assumption that the feature-space distance and the misclassification cost are correlated, the identification of high-cost tail instances can be realized by building region partitions with a low variance of risk within each region. The resulting AugARP approach could significantly outperform baseline approaches on both benchmark datasets and real-world product sales datasets. Yu-Cheng He, Yao-Xiang Ding 0001, Han-Jia Ye, Zhi-Hua Zhou |
AAAI | 2 |
| 2024 | Handling Varied Objectives by Online Decision MakingabstractConventional machine learning typically assume a fixed learning objective throughout the learning process.However, for real-world tasks in open and dynamic environments, objectives can change frequently.For example, in autonomous driving, a car has several default modes, but a user's concern for speed and fuel consumption varies depending on road conditions and personal needs.We formulate this problem as learning with varied objectives (LVO), where the goal is to optimize a dynamic weighted combination of multiple sub-objectives by sequentially selecting actions that incur different losses on these sub-objectives.We propose the VaRons algorithm, which estimates the action-wise performance on each sub-objective and adaptively selects decisions according to the dynamic requirements on different sub-objectives.Further, we extend our approach to cases involving contextual representations and propose the Con-VaRons algorithm, assuming parameterized linear structure that links contextual features to the main objective.Both the VaRons and ConVaRons are provably minimax optimal with respect to the time horizon 𝑇 , with ConVaRons showing better dependency with the number of sub-objectives 𝐾.Experiments on dynamic classifier and real-world cluster service allocation tasks validate the effectiveness of our methods and support our theoretical findings. Lanjihong Ma, Zhen-Yu Zhang, Yao-Xiang Ding 0001, Zhi-Hua Zhou |
KDD | 3 |
| 2024 | Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
Bohong Chen 0004, Yao-Xiang Ding 0001, Tianjia Shao, Kun Zhou 0001 |
ACM Multimedia | 3 |
| 2024 | CPoser: An Optimization-after-Parsing Approach for Text-to-Pose Generation Using Large Language ModelsabstractText-to-pose generation is challenging due to the complexity of natural language and human posture semantics. Utilizing large language models (LLMs) for text-to-pose generation is appealing due to their strong capabilities in text understanding and reasoning. However, as LLMs are designed for general-purpose language processing and not specifically trained for pose generation, it remains nontrivial to generate precise articulation targets for the full body using LLMs directly. To this end, we propose CPoser, a novel approach to harness the power of LLMs for text-to-pose generation, featuring a prompt parsing stage and a pose optimization stage. The parsing stage utilizes LLMs to turn text prompts into pose intermediate representations (Pose-IRs) through a set of predefined structured queries. These Pose-IRs explicitly describe specific pose conditions, such as squatting depth and knee bending angle, naturally forming an objective function that a target pose should satisfy. The optimization stage solves for expressive poses and hand gestures based on the Pose-IR objective function via robust optimization in a quantized pose prior space. The results are further refined to enhance naturalness and incorporate facial expressions. Experiments show that our approach effectively understands diverse text prompts for pose generation, surpassing existing text-to-pose methods. Bohong Chen 0004, Zhong Ren 0001, Yao-Xiang Ding 0001, Libin Liu 0002, Tianjia Shao, Kun Zhou 0001 |
ACM Trans. Graph. | 4 |
| 2023 | Seeing Differently, Acting Similarly: Heterogeneously Observable Imitation Learning
Xin-Qiang Cai, Yao-Xiang Ding 0001, Zi-Xuan Chen, Yuan Jiang 0001, Masashi Sugiyama, Zhi-Hua Zhou |
ICLR | 2 |
| 2023 | Model Spider: Learning to Rank Pre-Trained Models EfficientlyabstractFiguring out which Pre-Trained Model (PTM) from a model zoo fits the target task is essential to take advantage of plentiful model resources. With the availability of numerous heterogeneous PTMs from diverse fields, efficiently selecting the most suitable one is challenging due to the time-consuming costs of carrying out forward or backward passes over all PTMs. In this paper, we propose Model Spider, which tokenizes both PTMs and tasks by summarizing their characteristics into vectors to enable efficient PTM selection. By leveraging the approximated performance of PTMs on a separate set of training tasks, Model Spider learns to construct representation and measure the fitness score between a model-task pair via their representation. The ability to rank relevant PTMs higher than others generalizes to new tasks. With the top-ranked PTM candidates, we further learn to enrich task repr. with their PTM-specific semantics to re-rank the PTMs for better selection. Model Spider balances efficiency and selection ability, making PTM selection like a spider preying on a web. Model Spider exhibits promising performance across diverse model zoos, including visual models and Large Language Models (LLMs). Code is available at https://github.com/zhangyikaii/Model-Spider. Yi-Kai Zhang, Ting-Ji Huang, Yao-Xiang Ding 0001, De-Chuan Zhan, Han-Jia Ye |
NeurIPS | 3 |
| 2022 | Pre-Trained Model Reusability Evaluation for Small-Data Transfer LearningabstractWe study {\it model reusability evaluation} (MRE) for source pre-trained models: evaluating their transfer learning performance to new target tasks. In special, we focus on the setting under which the target training datasets are small, making it difficult to produce reliable MRE scores using them. Under this situation, we propose {\it synergistic learning} for building the task-model metric, which can be realized by collecting a set of pre-trained models and asking a group of data providers to participate. We provide theoretical guarantees to show that the learned task-model metric distances can serve as trustworthy MRE scores, and propose synergistic learning algorithms and models for general learning tasks. Experiments show that the MRE models learned by synergistic learning can generate significantly more reliable MRE scores than existing approaches for small-data transfer learning. Yao-Xiang Ding 0001, Xi-Zhu Wu, Kun Zhou 0001, Zhi-Hua Zhou |
NeurIPS | 1 |
| 2020 | Boosting-Based Reliable Model ReuseabstractWe study the following model reuse problem: a learner needs to select a subset of models from a model pool to classify an unlabeled dataset without accessing the raw training data of the models. Under this situation, it is challenging to properly estimate the reusability of the models in the pool. In this work, we consider the model reuse protocol under which the learner receives specifications of the models, including reusability indicators to verify the models’ prediction accuracy on any unlabeled instances. We propose MoreBoost, a simple yet powerful boosting algorithm to achieve effective model reuse under the idealized assumption that the reusability indicators are noise-free. When the reusability indicators are noisy, we strengthen MoreBoost with an active rectification mechanism, allowing the learner to query ground-truth indicator values from the model providers actively. The resulted MoreBoost.AR algorithm is guaranteed to significantly reduce the prediction error caused by the indicator noise. We also conduct experiments on both synthetic and benchmark datasets to verify the performance of the proposed approaches. Yao-Xiang Ding 0001, Zhi-Hua Zhou |
ACML | 1 |
| 2018 | Preference Based Adaptation for Learning ObjectivesabstractIn many real-world learning tasks, it is hard to directly optimize the true performance measures, meanwhile choosing the right surrogate objectives is also difficult. Under this situation, it is desirable to incorporate an optimization of objective process into the learning loop based on weak modeling of the relationship between the true measure and the objective. In this work, we discuss the task of objective adaptation, in which the learner iteratively adapts the learning objective to the underlying true objective based on the preference feedback from an oracle. We show that when the objective can be linearly parameterized, this preference based learning problem can be solved by utilizing the dueling bandit model. A novel sampling based algorithm DL^2M is proposed to learn the optimal parameter, which enjoys strong theoretical guarantees and efficient empirical performance. To avoid learning a hypothesis from scratch after each objective function update, a boosting based hypothesis adaptation approach is proposed to efficiently adapt any pre-learned element hypothesis to the current objective. We apply the overall approach to multi-label learning, and show that the proposed approach achieves significant performance under various multi-label performance measures. Yao-Xiang Ding 0001, Zhi-Hua Zhou |
NeurIPS | 1 |
| 2018 | Crowdsourcing with unsure option
Yao-Xiang Ding 0001, Zhi-Hua Zhou |
Mach. Learn. | 1 |