EDBT 2026 Demo / reviewers in the wild / expert
Qingyuan Wu
dblp:150/9871
· DBLP profile ↗
12ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 93% Planning, search and constraint satisfaction · 7% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
delayed reinforcement learning |
2.4 | 3 | 2025 | Directly Forecasting Belief for Reinforcement Learning with Delays · ICML 2025 Variational Delayed Policy Optimization · NeurIPS 2024 Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays · ICML 2024 |
Machine learning › Reinforcement learning › value-based reinforcement learning
value iteration network |
1.6 | 2 | 2025 | Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning · ICML 2025 Highway Value Iteration Networks · ICML 2024 |
Machine learning › Reinforcement learning › partially observable reinforcement learning
belief state estimation |
0.9 | 1 | 2025 | Directly Forecasting Belief for Reinforcement Learning with Delays · ICML 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
long-term planning |
0.9 | 1 | 2025 | Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning · ICML 2025 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.9 | 1 | 2025 | Directly Forecasting Belief for Reinforcement Learning with Delays · ICML 2025 |
Machine learning › Reinforcement learning
sample efficiency |
0.8 | 1 | 2024 | Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays · ICML 2024 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.8 | 1 | 2024 | Variational Delayed Policy Optimization · NeurIPS 2024 |
Machine learning › Reinforcement learning
state augmentation |
0.8 | 1 | 2024 | Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays · ICML 2024 |
Machine learning › Reinforcement learning
value function approximation |
0.8 | 1 | 2024 | Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays · ICML 2024 |
Machine learning › Reinforcement learning › dynamic programming
value iteration |
0.8 | 1 | 2024 | Highway Value Iteration Networks · ICML 2024 |
Machine learning › Reinforcement learning › markov decision process › Bayesian MDP
Hidden parameter MDPs |
0.3 | 1 | 2025 | Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning · ICML 2025 |
Machine learning › Reinforcement learning › model-based reinforcement learning
model-based planning |
0.3 | 1 | 2025 | Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
skip connections · 1.6transformer · 0.9multi-step bootstrapping · 0.9dynamic transition kernel · 0.9adaptive highway loss · 0.9policy improvement · 0.8highway value iteration · 0.8bootstrapping · 0.8backpropagation · 0.8auxiliary task learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term PlanningabstractThe Value Iteration Network (VIN) is an end-to-end differentiable neural network architecture for planning. It exhibits strong generalization to unseen domains by incorporating a differentiable planning module that operates on a latent Markov Decision Process (MDP). However, VINs struggle to scale to long-term and large-scale planning tasks, such as navigating a $100\times 100$ maze---a task that typically requires thousands of planning steps to solve. We observe that this deficiency is due to two issues: the representation capacity of the latent MDP and the planning module's depth. We address these by augmenting the latent MDP with a dynamic transition kernel, dramatically improving its representational capacity, and, to mitigate the vanishing gradient problem, introduce an "adaptive highway loss" that constructs skip connections to improve gradient flow. We evaluate our method on 2D/3D maze navigation environments, continuous control, and the real-world Lunar rover navigation task. We find that our new method, named Dynamic Transition VIN (DT-VIN), scales to 5000 layers and solves challenging versions of the above tasks. Altogether, we believe that DT-VIN represents a concrete step forward in performing long-term large-scale planning in complex environments. Yuhui Wang 0004, Qingyuan Wu, Dylan R. Ashley, Francesco Faccio, Weida Li, Chao Huang 0015, Jürgen Schmidhuber |
ICML | 2 |
| 2025 | Directly Forecasting Belief for Reinforcement Learning with DelaysabstractReinforcement learning (RL) with delays is challenging as sensory perceptions lag behind the actual events: the RL agent needs to estimate the real state of its environment based on past observations. State-of-the-art (SOTA) methods typically employ recursive, step-by-step forecasting of states. This can cause the accumulation of compounding errors. To tackle this problem, our novel belief estimation method, named Directly Forecasting Belief Transformer (DFBT), directly forecasts states from observations without incrementally estimating intermediate states step-by-step. We theoretically demonstrate that DFBT greatly reduces compounding errors of existing recursively forecasting methods, yielding stronger performance guarantees. In experiments with D4RL offline datasets, DFBT reduces compounding errors with remarkable prediction accuracy. DFBT’s capability to forecast state sequences also facilitates multi-step bootstrapping, thus greatly improving learning efficiency. On the MuJoCo benchmark, our DFBT-based method substantially outperforms SOTA baselines. Code is available at https://github.com/QingyuanWuNothing/DFBT. Qingyuan Wu, Yuhui Wang 0004, Simon Sinong Zhan, Yixuan Wang 0001, Chung-Wei Lin, Chen Lv 0001, Qi Zhu 0002, Jürgen Schmidhuber, Chao Huang 0015 |
ICML | 1 |
| 2025 | Underwater Salient Object Detection via Dual-Stage Self-Paced Learning and Depth EmphasisabstractSalient object detection of underwater scenes (USOD) poses greater challenges than that of traditional terrestrial scenes due to the presence of diverse and complex underwater image degradation. Current deep learning-based USOD methods generally treat all samples equally while failing to account for the varying difficulty levels of different training samples, thus leading to a limited performance. To tackle this challenge, this paper introduces a novel deep USOD method which benefits from iterative Dual-stage Self-paced Learning (DSPL) and Salient Object Depth Emphasis (SODE). Specifically, a DSPL strategy, which enforces the network to only focus on simpler samples in the first stage and then shifts attention to more challenging samples in the second stage, is devised to imitate the learning process of humans. The whole network is iteratively trained with the DSPL strategy and thus gradually adapted to various underwater scenes with different difficulty levels. Additionally, the proposed method involves an SODE module, which adaptively enhances depth information to effectively locate salient objects, addressing the issue of unreliable depth data caused by underwater image quality degradation. Experimental results on two benchmark datasets demonstrate the superior performance of the proposed method against state-of-the-art methods. The source code of our method will be made available athttps://github.com/NIT-JJH/SPDE. Jianhui Jin, Qiuping Jiang, Qingyuan Wu, Binwei Xu, Runmin Cong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Highway Value Iteration NetworksabstractValue iteration networks (VINs) enable end-to-end learning for planning tasks by employing a differentiable "planning module" that approximates the value iteration algorithm. However, long-term planning remains a challenge because training very deep VINs is difficult. To address this problem, we embed highway value iteration—a recent algorithm designed to facilitate long-term credit assignment—into the structure of VINs. This improvement augments the "planning module" of the VIN with three additional components: 1) an "aggregate gate," which constructs skip connections to improve information flow across many layers; 2) an "exploration module," crafted to increase the diversity of information and gradient flow in spatial dimensions; 3) a "filter gate" designed to ensure safe exploration. The resulting novel highway VIN can be trained effectively with hundreds of layers using standard backpropagation. In long-term planning tasks requiring hundreds of planning steps, deep highway VINs outperform both traditional VINs and several advanced, very deep NNs. Yuhui Wang 0004, Weida Li, Francesco Faccio, Qingyuan Wu, Jürgen Schmidhuber |
ICML | 4 |
| 2024 | Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short DelaysabstractReinforcement learning (RL) is challenging in the common case of delays between events and their sensory perceptions. State-of-the-art (SOTA) state augmentation techniques either suffer from state space explosion or performance degeneration in stochastic environments. To address these challenges, we present a novel *Auxiliary-Delayed Reinforcement Learning (AD-RL)* method that leverages auxiliary tasks involving short delays to accelerate RL with long delays, without compromising performance in stochastic environments. Specifically, AD-RL learns a value function for short delays and uses bootstrapping and policy improvement techniques to adjust it for long delays. We theoretically show that this can greatly reduce the sample complexity. On deterministic and stochastic benchmarks, our method significantly outperforms the SOTAs in both sample efficiency and policy performance. Code is available at https://github.com/QingyuanWuNothing/AD-RL. Qingyuan Wu, Simon Sinong Zhan, Yixuan Wang 0001, Yuhui Wang 0004, Chung-Wei Lin, Chen Lv 0001, Qi Zhu 0002, Jürgen Schmidhuber, Chao Huang 0015 |
ICML | 1 |
| 2024 | Variational Delayed Policy OptimizationabstractIn environments with delayed observation, state augmentation by including actions within the delay window is adopted to retrieve Markovian property to enable reinforcement learning (RL). Whereas, state-of-the-art (SOTA) RL techniques with Temporal-Difference (TD) learning frameworks commonly suffer from learning inefficiency, due to the significant expansion of the augmented state space with the delay. To improve the learning efficiency without sacrificing performance, this work novelly introduces Variational Delayed Policy Optimization (VDPO), reforming delayed RL as a variational inference problem. This problem is further modelled as a two-step iterative optimization problem, where the first step is TD learning in the delay-free environment with a small state space, and the second step is behaviour cloning which can be addressed much more efficiently than TD learning. We not only provide a theoretical analysis of VDPO in terms of sample complexity and performance, but also empirically demonstrate that VDPO can achieve consistent performance with SOTA methods, with a significant enhancement of sample efficiency (approximately 50\% less amount of samples) in the MuJoCo benchmark. Qingyuan Wu, Simon Sinong Zhan, Yixuan Wang 0001, Yuhui Wang 0004, Chung-Wei Lin, Chen Lv 0001, Qi Zhu 0002, Chao Huang 0015 |
NeurIPS | 1 |
| 2023 | Topic Driven Adaptive Network for cross-domain sentiment classification
Yicheng Zhu, Yiqiao Qiu, Qingyuan Wu, Fu Lee Wang, Yanghui Rao |
Inf. Process. Manag. | 3 |
| 2022 | Kuramoto Model-Based Analysis Reveals Oxytocin Effects on Brain Network DynamicsabstractThe oxytocin effects on large-scale brain networks such as Default Mode Network (DMN) and Frontoparietal Network (FPN) have been largely studied using fMRI data. However, these studies are mainly based on the statistical correlation or Bayesian causality inference, lacking interpretability at the physical and neuroscience level. Here, we propose a physics-based framework of the Kuramoto model to investigate oxytocin effects on the phase dynamic neural coupling in DMN and FPN. Testing on fMRI data of 59 participants administrated with either oxytocin or placebo, we demonstrate that oxytocin changes the topology of brain communities in DMN and FPN, leading to higher synchronization in the FPN and lower synchronization in the DMN, as well as a higher variance of the coupling strength within the DMN and more flexible coupling patterns at group level. These results together indicate that oxytocin may increase the ability to overcome the corresponding internal oscillation dispersion and support the flexibility in neural synchrony in various social contexts, providing new evidence for explaining the oxytocin modulated social behaviors. Our proposed Kuramoto model-based framework can be a potential tool in network neuroscience and offers physical and neural insights into phase dynamics of the brain. Shuhan Zheng, Zhichao Liang, Youzhi Qu, Qingyuan Wu, Haiyan Wu, Quanying Liu |
Int. J. Neural Syst. | 4 |
| 2017 | A multi-relational term scheme for first story detection
Yanghui Rao, Qing Li 0001, Qingyuan Wu, Haoran Xie 0001, Fu Lee Wang, Tao Wang 0036 |
Neurocomputing | 3 |
| 2015 | Identifying Context Familiarity for Incidental Word Learning Task Recommendations
Haoran Xie 0001, Di Zou, Fu Lee Wang, Tak-Lam Wong, Qingyuan Wu |
ICCE | 5 |
| 2015 | Differential co-expression and regulation analyses reveal different mechanisms underlying major depressive disorder and subsyndromal symptomatic depressionabstractBACKGROUND: Recent depression research has revealed a growing awareness of how to best classify depression into depressive subtypes. Appropriately subtyping depression can lead to identification of subtypes that are more responsive to current pharmacological treatment and aid in separating out depressed patients in which current antidepressants are not particularly effective. Differential co-expression analysis (DCEA) and differential regulation analysis (DRA) were applied to compare the transcriptomic profiles of peripheral blood lymphocytes from patients with two depressive subtypes: major depressive disorder (MDD) and subsyndromal symptomatic depression (SSD). RESULTS: Six differentially regulated genes (DRGs) (FOSL1, SRF, JUN, TFAP4, SOX9, and HLF) and 16 transcription factor-to-target differentially co-expressed gene links or pairs (TF2target DCLs) appear to be the key differential factors in MDD; in contrast, one DRG (PATZ1) and eight TF2target DCLs appear to be the key differential factors in SSD. There was no overlap between the MDD target genes and SSD target genes. Venlafaxine (Efexor™, Effexor™) appears to have a significant effect on the gene expression profile of MDD patients but no significant effect on the gene expression profile of SSD patients. CONCLUSION: DCEA and DRA revealed no apparent similarities between the differential regulatory processes underlying MDD and SSD. This bioinformatic analysis may provide novel insights that can support future antidepressant R&D efforts. Qingyuan Wu, Weihua Shao, Jun Mu, Deyu Yang, Yongtao Yang |
BMC Bioinform. | 4 |
| 2014 | Affective topic model for social emotion detection
Yanghui Rao, Qing Li 0001, Wenyin Liu, Qingyuan Wu, Xiaojun Quan |
Neural Networks | 4 |