VLDB 2026 Research / reviewers in the wild / expert
Tai-Wei Chang
dblp:151/8520
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Trustworthy machine learning · 36% Reinforcement learning · 36% Probabilistic and Bayesian machine learning · 16% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.6 | 2 | 2025 | RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution · EMNLP 2025 Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning · EMNLP 2024 |
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution · EMNLP 2025 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal model |
0.9 | 1 | 2025 | Learning Causal Transition Matrix for Instance-dependent Label Noise · AAAI 2025 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
instance-dependent label noise |
0.9 | 1 | 2025 | Learning Causal Transition Matrix for Instance-dependent Label Noise · AAAI 2025 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.9 | 1 | 2025 | Learning Causal Transition Matrix for Instance-dependent Label Noise · AAAI 2025 |
Machine learning › Reinforcement learning › reward design › reward shaping
reward redistribution |
0.9 | 1 | 2025 | RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution · EMNLP 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning · EMNLP 2024 |
Recommender systems
causal recommendation |
0.6 | 1 | 2022 | ESCM2: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate Estimation · SIGIR 2022 |
Recommender systems
conversion rate prediction |
0.6 | 1 | 2022 | ESCM2: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate Estimation · SIGIR 2022 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.3 | 1 | 2025 | Learning Causal Transition Matrix for Instance-dependent Label Noise · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
transition matrix estimation · 0.9reward model · 0.9causal graph · 0.9mirror descent · 0.8dynamic weighted sum · 0.8RLHF · 0.8RLAIF · 0.8multi-task learning · 0.6counterfactual risk minimization · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Causal Transition Matrix for Instance-dependent Label NoiseabstractNoisy labels are both inevitable and problematic in machine learning methods, as they negatively impact models' generalization ability by causing overfitting. In the context of learning with noise, the transition matrix plays a crucial role in the design of statistically consistent algorithms. However, the transition matrix is often considered unidentifiable. One strand of methods typically addresses this problem by assuming that the transition matrix is instance-independent; that is, the probability of mislabeling a particular instance is not influenced by its characteristics or attributes. This assumption is clearly invalid in complex real-world scenarios. To better understand the transition relationship and relax this assumption, we propose to study the data generation process of noisy labels from a causal perspective. We discover that an unobservable latent variable can affect either the instance itself, the label annotation procedure, or both, which complicates the identification of the transition matrix. To address various scenarios, we have unified these observations within a new causal graph. In this graph, the input instance is divided into a noise-resistant component and a noise-sensitive component based on whether they are affected by the latent variable. These two components contribute to identifying the “causal transition matrix”, which approximates the true transition matrix with theoretical guarantee. In line with this, we have designed a novel training framework that explicitly models this causal relationship and, as a result, achieves a more accurate model for inferring the clean label. Jiahui Li 0003, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011 |
AAAI | 2 |
| 2025 | RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward RedistributionabstractReinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences.Typically, a reward model is trained or supplied to act as a proxy for humans in evaluating generated responses during the reinforcement training phase.However, current reward models operate as sequence-to-one models, allocating a single, sparse, and delayed reward to an entire output sequence.This approach may overlook the significant contributions of individual tokens toward the desired outcome.To this end, we propose a more finegrained, token-level guidance approach for RL training.Specifically, we introduce RED, a novel REward reDistribition method that evaluates and assigns specific credit to each token using an off-the-shelf reward model.Utilizing these fine-grained rewards enhances the model's understanding of language nuances, leading to more precise performance improvements.Notably, our method does not require modifying the reward model or introducing additional training steps, thereby incurring minimal computational costs.Experimental results across diverse datasets and tasks demonstrate the superiority of our approach. Jiahui Li 0003, Lin Li 0065, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011, Cheng Yang 0002 |
EMNLP | 3 |
| 2024 | Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement LearningabstractReinforcement learning from human feedback (RLHF) and AI-generated feedback (RLAIF) have become prominent techniques that significantly enhance the functionality of pre-trained language models (LMs).These methods harness feedback, sourced either from humans or AI, as direct rewards or to shape reward models that steer LM optimization.Nonetheless, the effective integration of rewards from diverse sources presents a significant challenge due to their disparate characteristics.To address this, recent research has developed algorithms incorporating strategies such as weighting, ranking, and constraining to handle this complexity.Despite these innovations, a bias toward disproportionately high rewards can still skew the reinforcement learning process and negatively impact LM performance.This paper explores a methodology for reward composition that enables simultaneous improvements in LMs across multiple dimensions.Inspired by fairness theory, we introduce a training algorithm that aims to reduce Disparity and enhance Stability among various rewards.Our method treats the aggregate reward as a dynamic weighted sum of individual rewards, with alternating updates to the weights and model parameters.For efficient and straightforward implementation, we employ an estimation technique rooted in the mirror descent method for weight updates, eliminating the need for gradient computations.The empirical results under various types of rewards across a wide range of scenarios demonstrate the effectiveness of our method. Jiahui Li 0003, Fengda Zhang, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011 |
EMNLP | 4 |
| 2022 | ESCM2: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate EstimationabstractAccurate estimation of post-click conversion rate is critical for building recommender systems, which has long been confronted with sample selection bias and data sparsity issues. Methods in the Entire Space Multi-task Model (ESMM) family leverage the sequential pattern of user actions, \ie $impression\rightarrow click \rightarrow conversion$ to address data sparsity issue. However, they still fail to ensure the unbiasedness of CVR estimates. In this paper, we theoretically demonstrate that ESMM suffers from the following two problems: (1) Inherent Estimation Bias (IEB) for CVR estimation, where the CVR estimate is inherently higher than the ground truth; (2) Potential Independence Priority (PIP) for CTCVR estimation, where ESMM might overlook the causality from click to conversion. To this end, we devise a principled approach named Entire Space Counterfactual Multi-task Modelling (ESCM$^2$), which employs a counterfactual risk miminizer as a regularizer in ESMM to address both IEB and PIP issues simultaneously. Extensive experiments on offline datasets and online environments demonstrate that our proposed ESCM$^2$ can largely mitigate the inherent IEB and PIP issues and achieve better performance than baseline models. Hao Wang 0049, Tai-Wei Chang, Tianqiao Liu, Jianmin Huang, Zhichao Chen 0001, Ruopeng Li |
SIGIR | 2 |
| 2014 | Interpretation of Chinese Discourse Connectives for Explicit Discourse Relation Recognition
Hen-Hsen Huang, Tai-Wei Chang, Huan-Yuan Chen, Hsin-Hsi Chen |
COLING | 2 |