EDBT 2026 Demo / reviewers in the wild / expert
Ye-Wen Wang
dblp:367/2496
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0004-4147-5341ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 35% Trustworthy machine learning · 32% Efficient and distributed learning · 16% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
offline reinforcement learning |
1.6 | 2 | 2025 | Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL · ICML 2025 Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL · NeurIPS 2024 |
Machine learning › Reinforcement learning › regularization for reinforcement learning
value regularization |
0.9 | 1 | 2025 | Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL · ICML 2025 |
Machine learning › Efficient and distributed learning
active learning |
0.8 | 1 | 2024 | Bidirectional Uncertainty-Based Active Learning for Open-Set Annotation · ECCV (28) 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › exponential family
dirichlet distribution |
0.8 | 1 | 2024 | Dirichlet-Based Prediction Calibration for Learning with Noisy Labels · AAAI 2024 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.8 | 1 | 2024 | Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.8 | 1 | 2024 | Dirichlet-Based Prediction Calibration for Learning with Noisy Labels · AAAI 2024 |
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning |
0.8 | 1 | 2024 | Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › active learning › active learning for classification
open-set active learning |
0.8 | 1 | 2024 | Bidirectional Uncertainty-Based Active Learning for Open-Set Annotation · ECCV (28) 2024 |
Machine learning › Trustworthy machine learning › calibration
prediction calibration |
0.8 | 1 | 2024 | Dirichlet-Based Prediction Calibration for Learning with Noisy Labels · AAAI 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 1 | 2024 | Dirichlet-Based Prediction Calibration for Learning with Noisy Labels · AAAI 2024 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.2 | 1 | 2024 | Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › uncertainty estimation › neural network uncertainty
evidential deep learning |
0.2 | 1 | 2024 | Dirichlet-Based Prediction Calibration for Learning with Noisy Labels · AAAI 2024 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.2 | 1 | 2024 | Bidirectional Uncertainty-Based Active Learning for Open-Set Annotation · ECCV (28) 2024 |
Methods — techniques the papers use, named apart from their topics
policy constraint · 0.9conservative q-learning · 0.9behavior cloning · 0.9optimistic critic reconstruction · 0.8evidence deep learning · 0.8dirichlet distribution · 0.8constrained fine-tuning · 0.8calibrated softmax · 0.8bidirectional uncertainty · 0.8active learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RLabstractOffline reinforcement learning (RL) aims to learn an effective policy from a static dataset.
To alleviate extrapolation errors, existing studies often uniformly regularize the value function or policy updates across all states.
However, due to substantial variations in data quality, the fixed regularization strength often leads to a dilemma:
Weak regularization strength fails to address extrapolation errors and value overestimation, while strong regularization strength shifts policy learning toward behavior cloning, impeding potential performance enabled by Bellman updates.
To address this issue, we propose the selective state-adaptive regularization method for offline RL. Specifically, we introduce state-adaptive regularization coefficients to trust state-level Bellman-driven results, while selectively applying regularization on high-quality actions, aiming to avoid performance degradation caused by tight constraints on low-quality actions.
By establishing a connection between the representative value regularization method, CQL, and explicit policy constraint methods, we effectively extend selective state-adaptive regularization to these two mainstream offline RL approaches.
Extensive experiments demonstrate that the proposed method significantly outperforms the state-of-the-art approaches in both offline and offline-to-online settings on the D4RL benchmark. The implementation is available at https://github.com/QinwenLuo/SSAR. Qin-Wen Luo, Ming-Kun Xie, Ye-Wen Wang, Sheng-Jun Huang |
ICML | 3 |
| 2024 | Dirichlet-Based Prediction Calibration for Learning with Noisy LabelsabstractLearning with noisy labels can significantly hinder the generalization performance of deep neural networks (DNNs). Existing approaches address this issue through loss correction or example selection methods. However, these methods often rely on the model's predictions obtained from the softmax function, which can be over-confident and unreliable. In this study, we identify the translation invariance of the softmax function as the underlying cause of this problem and propose the \textit{Dirichlet-based Prediction Calibration} (DPC) method as a solution. Our method introduces a calibrated softmax function that breaks the translation invariance by incorporating a suitable constant in the exponent term, enabling more reliable model predictions. To ensure stable model training, we leverage a Dirichlet distribution to assign probabilities to predicted labels and introduce a novel evidence deep learning (EDL) loss. The proposed loss function encourages positive and sufficiently large logits for the given label, while penalizing negative and small logits for other labels, leading to more distinct logits and facilitating better example selection based on a large-margin criterion. Through extensive experiments on diverse benchmark datasets, we demonstrate that DPC achieves state-of-the-art performance. The code is available at https://github.com/chenchenzong/DPC. Chen-Chen Zong, Ye-Wen Wang, Ming-Kun Xie, Sheng-Jun Huang |
AAAI | 2 |
| 2024 | Bidirectional Uncertainty-Based Active Learning for Open-Set Annotation
Chen-Chen Zong, Ye-Wen Wang, Kun-Peng Ning, Haibo Ye, Sheng-Jun Huang |
ECCV (28) | 2 |
| 2024 | Dirichlet-Based Coarse-to-Fine Example Selection For Open-Set AnnotationabstractActive learning (AL) has achieved great success by selecting the most valuable examples from unlabeled data. However, they usually deteriorate in real scenarios where open-set noise gets involved, which is studied as open-set annotation (OSA). In this paper, we owe the deterioration to the unreliable predictions arising from softmax-based translation invariance and propose a Dirichlet-based Coarse-to-Fine Example Selection (DCFS) strategy accordingly. Our method introduces simplex-based evidential deep learning (EDL) to break translation invariance and distinguish known and unknown classes by considering evidence-based data and distribution uncertainty simultaneously. Furthermore, hard known-class examples are identified by model discrepancy generated from two classifier heads, where we amplify and alleviate the model discrepancy respectively for unknown and known classes. Finally, we combine the discrepancy with uncertainties to form a two-stage strategy, selecting the most informative examples from known classes. Extensive experiments on various openness ratio datasets demonstrate that DCFS achieves state-of-art performance. Ye-Wen Wang, Chen-Chen Zong, Ming-Kun Xie, Sheng-Jun Huang |
ICME | 1 |
| 2024 | Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RLabstractOffline-to-online (O2O) reinforcement learning (RL) provides an effective means of leveraging an offline pre-trained policy as initialization to improve performance rapidly with limited online interactions. Recent studies often design fine-tuning strategies for a specific offline RL method and cannot perform general O2O learning from any offline method. To deal with this problem, we disclose that there are evaluation and improvement mismatches between the offline dataset and the online environment, which hinders the direct application of pre-trained policies to online fine-tuning. In this paper, we propose to handle these two mismatches simultaneously, which aims to achieve general O2O learning from any offline method to any online method. Before online fine-tuning, we re-evaluate the pessimistic critic trained on the offline dataset in an optimistic way and then calibrate the misaligned critic with the reliable offline actor to avoid erroneous update. After obtaining an optimistic and and aligned critic, we perform constrained fine-tuning to combat distribution shift during online learning. We show empirically that the proposed method can achieve stable and efficient performance improvement on multiple simulated tasks when compared to the state-of-the-art methods. Qin-Wen Luo, Ming-Kun Xie, Ye-Wen Wang, Sheng-Jun Huang |
NeurIPS | 3 |