EDBT 2026 Demo / reviewers in the wild / expert
Dengwei Zhao
dblp:323/9550
· DBLP profile ↗
8ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0003-4764-2759ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Planning, search and constraint satisfaction · 79% Reinforcement learning · 21% | |
| Human-computer interaction and pervasive computing
1 paper |
Games and playful interaction · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search |
2.3 | 3 | 2025 | KeeA*: Epistemic Exploratory A* Search via Knowledge Calibration · NeurIPS 2025 SeeA*: Efficient Exploration-Enhanced A* Search by Selective Sampling · NeurIPS 2024 Generalized Weighted Path Consistency for Mastering Atari Games · NeurIPS 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search › best-first search
a* search |
1.6 | 2 | 2025 | KeeA*: Epistemic Exploratory A* Search via Knowledge Calibration · NeurIPS 2025 SeeA*: Efficient Exploration-Enhanced A* Search by Selective Sampling · NeurIPS 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
1.1 | 3 | 2025 | Generalized Weighted Path Consistency for Mastering Atari Games · NeurIPS 2023 KeeA*: Epistemic Exploratory A* Search via Knowledge Calibration · NeurIPS 2025 SeeA*: Efficient Exploration-Enhanced A* Search by Selective Sampling · NeurIPS 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search
neural-guided search |
0.7 | 1 | 2023 | Generalized Weighted Path Consistency for Mastering Atari Games · NeurIPS 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › local consistency
path consistency |
0.7 | 1 | 2023 | Generalized Weighted Path Consistency for Mastering Atari Games · NeurIPS 2023 |
Machine learning › Reinforcement learning › deep reinforcement learning
alphazero |
0.6 | 1 | 2022 | Efficient Learning for AlphaZero via Path Consistency · ICML 2022 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.6 | 1 | 2022 | Efficient Learning for AlphaZero via Path Consistency · ICML 2022 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
self-play |
0.6 | 1 | 2022 | Efficient Learning for AlphaZero via Path Consistency · ICML 2022 |
Games and playful interaction
board games |
0.2 | 1 | 2022 | Efficient Learning for AlphaZero via Path Consistency · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
monte carlo tree search · 2.6neural network · 1.5path consistency · 1.1cluster sampling · 0.9selective sampling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KeeA*: Epistemic Exploratory A* Search via Knowledge CalibrationabstractIn recent years, neural network-guided heuristic search algorithms, such as Monte-Carlo tree search and A$^\*$ search, have achieved significant advancements across diverse practical applications. Due to the challenges stemming from high state-space complexity, sparse training datasets, and incomplete environmental modeling, heuristic estimations manifest uncontrolled inherent biases towards the actual expected evaluations, thereby compromising the decision-making quality of search algorithms. Sampling exploration enhanced A$^\*$ (SeeA$^\*$) was proposed to improve the efficiency of A$^\*$ search by constructing an dynamic candidate subset through random sampling, from which the expanded node was selected. However, uniform sampling strategy utilized by SeeA$^\*$ facilitates exploration exclusively through the injection of randomness, which completely neglects the heuristic knowledge relevant to open nodes. Moreover, the theoretical support of cluster sampling remains ambiguous. Despite the existence of potential biases, heuristic estimations still encapsulate certain valuable information. In this paper, epistemic exploratory A$^\*$ search (KeeA$^\*$) is proposed to integrate heuristic knowledge for calibrating the sampling process. We first theoretically demonstrate that SeeA$^\*$ with cluster sampling outperforms uniform sampling due to the distribution-aware selection with higher variance. Building on this insight, cluster scouting and path-aware sampling are introduced in KeeA$^\*$ to further exploit heuristic knowledge to increase the sampling mean and variance, respectively, thereby generating higher-quality extreme candidates and enhancing overall decision-making performance. Finally, empirical results on retrosynthetic planning and logic synthesis demonstrate superior performance of KeeA$^*$ compared to state-of-the-art heuristic search algorithms. Dengwei Zhao, Shikui Tu, Yanan Sun 0003, Lei Xu 0001 |
NeurIPS | 1 |
| 2025 | Multi-Objective Structure-Based Drug Design Using Causal DiscoveryabstractStructure-based drug design (SBDD) is a critical subtask in the drug discovery process, with deep generative models playing a pivotal role. Inherently, drug design is a multi-objective task given the fact that a promising drug candidate must satisfy multiple properties. However, existing SBDD methods either focus solely on the binding affinity between molecules and target proteins while neglecting other crucial properties, or they assume that objective properties are independent of each other. Yet there are often potential relationships among properties, which can be conflicting-improving one property may lead to the deterioration of another. The lack of consideration for these relationships in current methods makes it unfeasible to generate molecules that simultaneously meet multiple objectives. To address the above issues, a multi-objective SBDD algorithm is proposed based on the diffusion model to optimize binding affinity and other drug properties simultaneously. Multiple expert networks are trained in parallel to predict properties for molecules in intermediate states and transmit gradients, and a causal graph is constructed through the causal discovery algorithm to unveil the underlying relationships among target properties. During the entire generation process, the joint distribution of target properties is decomposed in a reasonable manner according to the casual graph, and then the gradients of each property are applied to guide the optimizing direction of generation. Experimental results indicate that our model effectively optimizes multiple objectives simultaneously, generating molecules with greater drug potential compared to baseline models. Jingyuan Zhou, Dengwei Zhao, Shikui Tu, Lei Xu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | Novelty Encouraged Beam Clustering Search for Multi-Objective De Novo Diverse Drug DesignabstractThe generation of drug-like, high-quality molecules from scratch within the expansive chemical space is a significant challenge in drug discovery. In previous research, value-based reinforcement learning algorithms have been utilized to optimize multiple desired properties simultaneously. Randomness is injected into the decision-making process through ε-Greedy or stochastic sampling to enable the generation of a diverse ensemble of molecules, which usually encounters a trade-off between the optimality and diversity of these generated molecules. Moreover, novelty has not been explicitly addressed as an optimization objective, and the distinctiveness of generated molecules from the reference molecules is not guaranteed. In this paper, novelty-encouraged beam clustering (NeBC) search algorithm is proposed for de novo drug design. A clustering strategy is integrated with heuristic value-guided beam search to strike a balance between the optimality and diversity of the generated molecules. An intrinsic reward, which is measured by the disagreement of a group of experts trained on reference molecules, is proposed to encourage novelty explicitly. Experimental results demonstrate that NeBC search not only achieves a balanced trade-off between optimality and diversity but also effectively enhances the novelty of the generated molecules. The source code is publicly accessible on https://github.com/CMACH508/NeBC. Dengwei Zhao, Shikui Tu, Lei Xu 0001 |
BIBM | 1 |
| 2024 | SeeA*: Efficient Exploration-Enhanced A* Search by Selective SamplingabstractMonte-Carlo tree search (MCTS) and reinforcement learning contributed crucially to the success of AlphaGo and AlphaZero, and A$^*$ is a tree search algorithm among the most well-known ones in the classical AI literature. MCTS and A$^*$ both perform heuristic search and are mutually beneficial. Efforts have been made to the renaissance of A$^*$ from three possible aspects, two of which have been confirmed by studies in recent years, while the third is about the OPEN list that consists of open nodes of A$^*$ search, but still lacks deep investigation. This paper aims at the third, i.e., developing the Sampling-exploration enhanced A$^*$ (SeeA$^*$) search by constructing a dynamic subset of OPEN through a selective sampling process, such that the node with the best heuristic value in this subset instead of in the OPEN is expanded. Nodes with the best heuristic values in OPEN are most probably picked into this subset, but sometimes may not be included, which enables SeeA$^*$ to explore other promising branches. Three sampling techniques are presented for comparative investigations. Moreover, under the assumption about the distribution of prediction errors, we have theoretically shown the superior efficiency of SeeA$^*$ over A$^*$ search, particularly when the accuracy of the guiding heuristic function is insufficient. Experimental results on retrosynthetic planning in organic chemistry, logic synthesis in integrated circuit design, and the classical Sokoban game empirically demonstrate the efficiency of SeeA$^*$, in comparison with the state-of-the-art heuristic search algorithms. Dengwei Zhao, Shikui Tu, Lei Xu 0001 |
NeurIPS | 1 |
| 2024 | De Novo Drug Design by Multi-Objective Path Consistency Learning With Beam A* SearchabstractGenerating high-quality and drug-like molecules from scratch within the expansive chemical space presents a significant challenge in the field of drug discovery. In prior research, value-based reinforcement learning algorithms have been employed to generate molecules with multiple desired properties iteratively. The immediate reward was defined as the evaluation of intermediate-state molecules at each step, and the learning objective would be maximizing the expected cumulative evaluation scores for all molecules along the generative path. However, this definition of the reward was misleading, as in reality, the optimization target should be the evaluation score of only the final generated molecule. Furthermore, in previous works, randomness was introduced into the decision-making process, enabling the generation of diverse molecules but no longer pursuing the maximum future rewards. In this paper, immediate reward is defined as the improvement achieved through the modification of the molecule to maximize the evaluation score of the final generated molecule exclusively. Originating from the A search, path consistency (PC), i.e., values on one optimal path should be identical, is employed as the objective function in the update of the value estimator to train a multi-objective de novo drug designer. By incorporating the value into the decision-making process of beam search, the DrugBA algorithm is proposed to enable the large-scale generation of molecules that exhibit both high quality and diversity. Experimental results demonstrate a substantial enhancement over the state-of-the-art algorithm QADD in multiple molecular properties of the generated molecules. Dengwei Zhao, Jingyuan Zhou, Shikui Tu, Lei Xu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2023 | DeepTH: Chip Placement with Deep Reinforcement Learning Using a Three-Head Policy NetworkabstractModern very-large-scale integrated (VLSI) circuit placement with huge state space is a critical task for achieving layouts with high performance. Recently, reinforcement learning (RL) algorithms have made a promising breakthrough to dramatically save design time than human effort. However, the previous RL-based works either require a large dataset of chip placements for pre-training or produce illegal final placement solutions. In this paper, DeepTH, a three-head policy gradient placer, is proposed to learn from scratch without the need of pre-training, and generate superior chip floorplans. Graph neural network is initially adopted to extract the features from nodes and nets of chips for estimating the policy and value. To efficiently improve the quality of floorplans, a reconstruction head is employed in the RL network to recover the visual representation of the current placement, by enriching the extracted features of placement embedding. Besides, the reconstruction error is used as a bonus during training to encourage exploration while alleviating the sparse reward problem. Furthermore, the expert knowledge of floorplanning preference is embedded into the decision process to narrow down the potential action space. Experiment results on the ISPD 2005 benchmark have shown that our method achieves 19.02% HPWL improvement than the analytic placer DREAMPlace and 19.89% improvement at least than the state-of-the-art RL algorithms. Dengwei Zhao, Shuai Yuan 0016, Yanan Sun 0003, Shikui Tu, Lei Xu 0001 |
DATE | 1 |
| 2023 | Generalized Weighted Path Consistency for Mastering Atari GamesabstractReinforcement learning with the help of neural-guided search consumes huge computational resources to achieve remarkable performance. Path consistency (PC), i.e., $f$ values on one optimal path should be identical, was previously imposed on MCTS by PCZero to improve the learning efficiency of AlphaZero. Not only PCZero still lacks a theoretical support but also considers merely board games. In this paper, PCZero is generalized into GW-PCZero for real applications with non-zero immediate reward. A weighting mechanism is introduced to reduce the variance caused by scouting's uncertainty on the $f$ value estimation. For the first time, it is theoretically proved that neural-guided MCTS is guaranteed to find the optimal solution under the constraint of PC. Experiments are conducted on the Atari $100$k benchmark with $26$ games and GW-PCZero achieves $198\%$ mean human performance, higher than the state-of-the-art EfficientZero's $194\\%$, while consuming only $25\\%$ of the computational resources consumed by EfficientZero. Dengwei Zhao, Shikui Tu, Lei Xu 0001 |
NeurIPS | 1 |
| 2022 | Efficient Learning for AlphaZero via Path ConsistencyabstractIn recent years, deep reinforcement learning have made great breakthroughs on board games. Still, most of the works require huge computational resources for a large scale of environmental interactions or self-play for the games. This paper aims at building powerful models under a limited amount of self-plays which can be utilized by a human throughout the lifetime. We proposes a learning algorithm built on AlphaZero, with its path searching regularised by a path consistency (PC) optimality, i.e., values on one optimal search path should be identical. Thus, the algorithm is shortly named PCZero. In implementation, historical trajectory and scouted search paths by MCTS makes a good balance between exploration and exploitation, which enhances the generalization ability effectively. PCZero obtains $94.1%$ winning rate against the champion of Hex Computer Olympiad in 2015 on $13\times 13$ Hex, much higher than $84.3%$ by AlphaZero. The models consume only $900K$ self-play games, about the amount humans can study in a lifetime. The improvements by PCZero have been also generalized to Othello and Gomoku. Experiments also demonstrate the efficiency of PCZero under offline learning setting. Dengwei Zhao, Shikui Tu, Lei Xu 0001 |
ICML | 1 |