EDBT 2026 Demo / reviewers in the wild / expert
Qinwei Yang
dblp:354/9136
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 81% Transfer learning and domain adaptation · 12% Probabilistic and Bayesian machine learning · 8% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
policy learning |
1.6 | 2 | 2025 | Optimal Policy Adaptation Under Covariate Shift · IJCAI 2025 Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards · NeurIPS 2024 |
Machine learning › Reinforcement learning › offline reinforcement learning
offline policy learning |
1.0 | 1 | 2026 | Offline Policy Enhancement and Transfer by Combining Experimental and External One-Sided Treatment Data · KDD (1) 2026 |
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
policy transfer |
1.0 | 1 | 2026 | Offline Policy Enhancement and Transfer by Combining Experimental and External One-Sided Treatment Data · KDD (1) 2026 |
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift |
0.9 | 1 | 2025 | Optimal Policy Adaptation Under Covariate Shift · IJCAI 2025 |
Machine learning › Reinforcement learning › off-policy evaluation
doubly robust estimation |
0.9 | 1 | 2025 | Optimal Policy Adaptation Under Covariate Shift · IJCAI 2025 |
Machine learning › Reinforcement learning
semiparametric efficient estimation |
0.9 | 1 | 2025 | Optimal Policy Adaptation Under Covariate Shift · IJCAI 2025 |
Machine learning › Reinforcement learning
multi-objective reinforcement learning |
0.8 | 1 | 2024 | Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.3 | 1 | 2026 | Offline Policy Enhancement and Transfer by Combining Experimental and External One-Sided Treatment Data · KDD (1) 2026 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
treatment effect estimation |
0.3 | 1 | 2026 | Offline Policy Enhancement and Transfer by Combining Experimental and External One-Sided Treatment Data · KDD (1) 2026 |
Mathematical optimization
multi-objective optimization |
0.2 | 1 | 2024 | Learning the Optimal Policy for Balancing Short-Term and Long-Term Rewards · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
ε-constraint · 1.5linear weighting · 1.5decomposition-based policy learning · 1.5importance weighting · 1.0doubly robust estimation · 1.0sensitivity analysis · 0.9semiparametric efficiency bound · 0.9efficient influence function · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Offline Policy Enhancement and Transfer by Combining Experimental and External One-Sided Treatment DataabstractIn this article, we investigate a novel setting for offline policy enhancement and transfer that involves two distinct datasets: an experimental dataset and an external one-sided treatment dataset. The experimental dataset, though unconfounded, is constrained by its small sample. Consequently, methods based solely on the experimental dataset may suffer from low accuracy and limited generalizability. In contrast, the external one-sided treatment dataset typically has a larger sample size but includes observations from only one treatment arm (e.g., all units belong to the control group, with no units receiving the treatment). Based only on the external one-sided treatment dataset, it cannot identify the policy reward. By combining the two datasets, we propose a principled framework to accomplish two key tasks: (1) policy enhancement: improving the accuracy of offline policy evaluation and learning in the experimental dataset by leveraging the external one-sided treatment dataset; (2) policy transfer : enabling offline policy evaluation and learning in the external one-sided treatment dataset by utilizing information from the experimental dataset, and thus making the learned policies applicable to a broader range of data distributions. Extensive experiments demonstrate that our proposed methods not only estimate rewards more accurately but also learn policies that closely approximate the theoretically optimal policy. The code is available at https://anonymous.4open.science/r/Offline-Policy-Enhancement-and-Transfer-E0D1 for double-blind review. Qinwei Yang, Zhiyu Hao, Peng Wu 0012 |
KDD (1) | 2 |
| 2026 | Counterfactual harm-aware learning of individualized decisions
Qinwei Yang, Jile Chaoge |
Neurocomputing | 1 |
| 2025 | Optimal Policy Adaptation Under Covariate ShiftabstractTransfer learning of prediction models has been extensively studied, while the corresponding policy learning approaches are rarely discussed. In this paper, we propose principled approaches for learning the optimal policy in the target domain by leveraging two datasets: one with full information from the source domain and the other from the target domain with only covariates. First, in the setting of covariate shift, we formulate the problem from a perspective of causality and present the identifiability assumptions for the reward induced by a given policy. Then, we derive the efficient influence function and the semiparametric efficiency bound for the reward. Based on this, we construct a doubly robust and semiparametric efficient estimator for the reward and then learn the optimal policy by optimizing the estimated reward. Moreover, we theoretically analyze the bias and the generalization error bound for the learned policy. Furthermore, in the presence of both covariate and concept shifts, we propose a novel sensitivity analysis method to evaluate the robustness of the proposed policy learning approach. Extensive experiments demonstrate that the approach not only estimates the reward more accurately but also yields a policy that closely approximates the theoretically optimal policy. Qinwei Yang, Zhaoqing Tian, Ruocheng Guo, Peng Wu 0012 |
IJCAI | 2 |
| 2025 | Adaptive Data-Borrowing for Improving Treatment Effect Estimation using External ControlsabstractRandomized controlled trials (RCTs) often exhibit limited inferential efficiency in estimating treatment effects due to small sample sizes. In recent years, the combination of external controls has gained increasing attention as a means of improving the efficiency of RCTs. However, external controls are not always comparable to RCTs, and direct borrowing without careful evaluation can introduce substantial bias and reduce the efficiency of treatment effect estimation. In this paper, we propose a novel influence-based adaptive sample borrowing approach that effectively quantifies the "comparability'' of each sample in the external controls using influence function theory. Given a selected set of borrowed external controls, we further derive a semiparametric efficient estimator under an exchangeability assumption. Recognizing that the exchangeability assumption may not hold for all possible borrowing sets, we conduct a detailed analysis of the asymptotic bias and variance of the proposed estimator under violations of exchangeability. Building on this bias-variance trade-off, we further develop a data-driven approach to select the optimal subset of external controls for borrowing. Extensive simulations and real-world applications demonstrate that the proposed approach significantly enhances treatment effect estimation efficiency in RCTs, outperforming existing approaches. Qinwei Yang |
NeurIPS | 1 |
| 2025 | vClos: Network contention aware scheduling for distributed machine learning tasks in multi-tenant GPU clusters
Xinchi Han, Shizhen Zhao, Yongxi Lv, Peirui Cao, Qinwei Yang, Yunzhuo Liu, Shengkai Lin, Bo Jiang 0003, Ximeng Liu, Yong Cui 0001, Chenghu Zhou, Xinbing Wang |
Comput. Networks | 6 |
| 2024 | LubeRDMA: A Fail-safe Mechanism of RDMAabstractRecent years have witnessed a wide adoption of Remote Direct Memory Access (RDMA) to accelerate distributed systems. As the scale of distributed applications keeps increasing, network failures become more prominent. Although some link/switch failures can be circumvented by in-network rerouting, failures like NIC failure are still fatal in RDMA networks and may cause the entire system to fail. Shengkai Lin, Qinwei Yang, Zengyin Yang, Shizhen Zhao |
APNet | 2 |
| 2024 | Learning the Optimal Policy for Balancing Short-Term and Long-Term RewardsabstractLearning the optimal policy to balance multiple short-term and long-term rewards has extensive applications across various domains. Yet, there is a noticeable scarcity of research addressing policy learning strategies in this context. In this paper, we aim to learn the optimal policy capable of effectively balancing multiple short-term and long-term rewards, especially in scenarios where the long-term outcomes are often missing due to data collection challenges over extended periods. Towards this goal, the conventional linear weighting method, which aggregates multiple rewards into a single surrogate reward through weighted summation, can only achieve sub-optimal policies when multiple rewards are related. Motivated by this, we propose a novel decomposition-based policy learning (DPPL) method that converts the whole problem into subproblems. The DPPL method is capable of obtaining optimal policies even when multiple rewards are interrelated. Nevertheless, the DPPL method requires a set of preference vectors specified in advance, posing challenges in practical applications where selecting suitable preferences is non-trivial. To mitigate this, we further theoretically transform the optimization problem in DPPL into an $\varepsilon$-constraint problem, where $\varepsilon$ represents the minimum acceptable levels of other rewards while maximizing one reward. This transformation provides intuitive into the selection of preference vectors. Extensive experiments are conducted on the proposed method and the results validate the effectiveness of the method. Qinwei Yang, Yan Zeng 0002, Ruocheng Guo, Yang Liu 0018, Peng Wu 0012 |
NeurIPS | 1 |