Qinwei Yang

dblp:354/9136 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 81% Transfer learning and domain adaptation · 12% Probabilistic and Bayesian machine learning · 8%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
policy learning
1.622025
Optimal Policy Adaptation Under Covariate Shift · IJCAI 2025
Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards · NeurIPS 2024
Machine learning › Reinforcement learning › offline reinforcement learning
offline policy learning
1.012026
Offline Policy Enhancement and Transfer by Combining Experimental and External One-Sided Treatment Data · KDD (1) 2026
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy transfer
1.012026
Offline Policy Enhancement and Transfer by Combining Experimental and External One-Sided Treatment Data · KDD (1) 2026
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift
0.912025
Optimal Policy Adaptation Under Covariate Shift · IJCAI 2025
Machine learning › Reinforcement learning › off-policy evaluation
doubly robust estimation
0.912025
Optimal Policy Adaptation Under Covariate Shift · IJCAI 2025
Machine learning › Reinforcement learning
semiparametric efficient estimation
0.912025
Optimal Policy Adaptation Under Covariate Shift · IJCAI 2025
Machine learning › Reinforcement learning
multi-objective reinforcement learning
0.812024
Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.312026
Offline Policy Enhancement and Transfer by Combining Experimental and External One-Sided Treatment Data · KDD (1) 2026
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
treatment effect estimation
0.312026
Offline Policy Enhancement and Transfer by Combining Experimental and External One-Sided Treatment Data · KDD (1) 2026
Mathematical optimization
multi-objective optimization
0.212024
Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

ε-constraint · 1.5linear weighting · 1.5decomposition-based policy learning · 1.5importance weighting · 1.0doubly robust estimation · 1.0sensitivity analysis · 0.9semiparametric efficiency bound · 0.9efficient influence function · 0.9
YearPublicationVenuePosition
2026 Offline Policy Enhancement and Transfer by Combining Experimental and External One-Sided Treatment Data
abstract
In this article, we investigate a novel setting for offline policy enhancement and transfer that involves two distinct datasets: an experimental dataset and an external one-sided treatment dataset. The experimental dataset, though unconfounded, is constrained by its small sample. Consequently, methods based solely on the experimental dataset may suffer from low accuracy and limited generalizability. In contrast, the external one-sided treatment dataset typically has a larger sample size but includes observations from only one treatment arm (e.g., all units belong to the control group, with no units receiving the treatment). Based only on the external one-sided treatment dataset, it cannot identify the policy reward. By combining the two datasets, we propose a principled framework to accomplish two key tasks: (1) policy enhancement: improving the accuracy of offline policy evaluation and learning in the experimental dataset by leveraging the external one-sided treatment dataset; (2) policy transfer : enabling offline policy evaluation and learning in the external one-sided treatment dataset by utilizing information from the experimental dataset, and thus making the learned policies applicable to a broader range of data distributions. Extensive experiments demonstrate that our proposed methods not only estimate rewards more accurately but also learn policies that closely approximate the theoretically optimal policy. The code is available at https://anonymous.4open.science/r/Offline-Policy-Enhancement-and-Transfer-E0D1 for double-blind review.
Qinwei Yang, Zhiyu Hao, Peng Wu 0012
KDD (1)2
2026 Counterfactual harm-aware learning of individualized decisions
Qinwei Yang, Jile Chaoge
Neurocomputing1
2025 Optimal Policy Adaptation Under Covariate Shift
abstract
Transfer learning of prediction models has been extensively studied, while the corresponding policy learning approaches are rarely discussed. In this paper, we propose principled approaches for learning the optimal policy in the target domain by leveraging two datasets: one with full information from the source domain and the other from the target domain with only covariates. First, in the setting of covariate shift, we formulate the problem from a perspective of causality and present the identifiability assumptions for the reward induced by a given policy. Then, we derive the efficient influence function and the semiparametric efficiency bound for the reward. Based on this, we construct a doubly robust and semiparametric efficient estimator for the reward and then learn the optimal policy by optimizing the estimated reward. Moreover, we theoretically analyze the bias and the generalization error bound for the learned policy. Furthermore, in the presence of both covariate and concept shifts, we propose a novel sensitivity analysis method to evaluate the robustness of the proposed policy learning approach. Extensive experiments demonstrate that the approach not only estimates the reward more accurately but also yields a policy that closely approximates the theoretically optimal policy.
Qinwei Yang, Zhaoqing Tian, Ruocheng Guo, Peng Wu 0012
IJCAI2
2025 Adaptive Data-Borrowing for Improving Treatment Effect Estimation using External Controls
abstract
Randomized controlled trials (RCTs) often exhibit limited inferential efficiency in estimating treatment effects due to small sample sizes. In recent years, the combination of external controls has gained increasing attention as a means of improving the efficiency of RCTs. However, external controls are not always comparable to RCTs, and direct borrowing without careful evaluation can introduce substantial bias and reduce the efficiency of treatment effect estimation. In this paper, we propose a novel influence-based adaptive sample borrowing approach that effectively quantifies the "comparability'' of each sample in the external controls using influence function theory. Given a selected set of borrowed external controls, we further derive a semiparametric efficient estimator under an exchangeability assumption. Recognizing that the exchangeability assumption may not hold for all possible borrowing sets, we conduct a detailed analysis of the asymptotic bias and variance of the proposed estimator under violations of exchangeability. Building on this bias-variance trade-off, we further develop a data-driven approach to select the optimal subset of external controls for borrowing. Extensive simulations and real-world applications demonstrate that the proposed approach significantly enhances treatment effect estimation efficiency in RCTs, outperforming existing approaches.
Qinwei Yang
NeurIPS1
2025 vClos: Network contention aware scheduling for distributed machine learning tasks in multi-tenant GPU clusters
Xinchi Han, Shizhen Zhao, Yongxi Lv, Peirui Cao, Qinwei Yang, Yunzhuo Liu, Shengkai Lin, Bo Jiang 0003, Ximeng Liu, Yong Cui 0001, Chenghu Zhou, Xinbing Wang
Comput. Networks6
2024 LubeRDMA: A Fail-safe Mechanism of RDMA
abstract
Recent years have witnessed a wide adoption of Remote Direct Memory Access (RDMA) to accelerate distributed systems. As the scale of distributed applications keeps increasing, network failures become more prominent. Although some link/switch failures can be circumvented by in-network rerouting, failures like NIC failure are still fatal in RDMA networks and may cause the entire system to fail.
Shengkai Lin, Qinwei Yang, Zengyin Yang, Shizhen Zhao
APNet2
2024 Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards
abstract
Learning the optimal policy to balance multiple short-term and long-term rewards has extensive applications across various domains. Yet, there is a noticeable scarcity of research addressing policy learning strategies in this context. In this paper, we aim to learn the optimal policy capable of effectively balancing multiple short-term and long-term rewards, especially in scenarios where the long-term outcomes are often missing due to data collection challenges over extended periods. Towards this goal, the conventional linear weighting method, which aggregates multiple rewards into a single surrogate reward through weighted summation, can only achieve sub-optimal policies when multiple rewards are related. Motivated by this, we propose a novel decomposition-based policy learning (DPPL) method that converts the whole problem into subproblems. The DPPL method is capable of obtaining optimal policies even when multiple rewards are interrelated. Nevertheless, the DPPL method requires a set of preference vectors specified in advance, posing challenges in practical applications where selecting suitable preferences is non-trivial. To mitigate this, we further theoretically transform the optimization problem in DPPL into an $\varepsilon$-constraint problem, where $\varepsilon$ represents the minimum acceptable levels of other rewards while maximizing one reward. This transformation provides intuitive into the selection of preference vectors. Extensive experiments are conducted on the proposed method and the results validate the effectiveness of the method.
Qinwei Yang, Yan Zeng 0002, Ruocheng Guo, Yang Liu 0018, Peng Wu 0012
NeurIPS1