Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tai-Wei Chang

dblp:151/8520 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 36% Reinforcement learning · 36% Probabilistic and Bayesian machine learning · 16%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.622025
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution · EMNLP 2025
Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning · EMNLP 2024
Natural language and speech › Language models and text generation
alignment
0.912025
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution · EMNLP 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal model
0.912025
Learning Causal Transition Matrix for Instance-dependent Label Noise · AAAI 2025
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
instance-dependent label noise
0.912025
Learning Causal Transition Matrix for Instance-dependent Label Noise · AAAI 2025
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.912025
Learning Causal Transition Matrix for Instance-dependent Label Noise · AAAI 2025
Machine learning › Reinforcement learning › reward design › reward shaping
reward redistribution
0.912025
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution · EMNLP 2025
Machine learning › Trustworthy machine learning
fairness
0.812024
Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning · EMNLP 2024
Recommender systems
causal recommendation
0.612022
ESCM2: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate Estimation · SIGIR 2022
Recommender systems
conversion rate prediction
0.612022
ESCM2: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate Estimation · SIGIR 2022
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.312025
Learning Causal Transition Matrix for Instance-dependent Label Noise · AAAI 2025

Methods — techniques the papers use, named apart from their topics

transition matrix estimation · 0.9reward model · 0.9causal graph · 0.9mirror descent · 0.8dynamic weighted sum · 0.8RLHF · 0.8RLAIF · 0.8multi-task learning · 0.6counterfactual risk minimization · 0.6
YearPublicationVenuePosition
2025 Learning Causal Transition Matrix for Instance-dependent Label Noise
abstract
Noisy labels are both inevitable and problematic in machine learning methods, as they negatively impact models' generalization ability by causing overfitting. In the context of learning with noise, the transition matrix plays a crucial role in the design of statistically consistent algorithms. However, the transition matrix is often considered unidentifiable. One strand of methods typically addresses this problem by assuming that the transition matrix is instance-independent; that is, the probability of mislabeling a particular instance is not influenced by its characteristics or attributes. This assumption is clearly invalid in complex real-world scenarios. To better understand the transition relationship and relax this assumption, we propose to study the data generation process of noisy labels from a causal perspective. We discover that an unobservable latent variable can affect either the instance itself, the label annotation procedure, or both, which complicates the identification of the transition matrix. To address various scenarios, we have unified these observations within a new causal graph. In this graph, the input instance is divided into a noise-resistant component and a noise-sensitive component based on whether they are affected by the latent variable. These two components contribute to identifying the “causal transition matrix”, which approximates the true transition matrix with theoretical guarantee. In line with this, we have designed a novel training framework that explicitly models this causal relationship and, as a result, achieves a more accurate model for inferring the clean label.
Jiahui Li 0003, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011
AAAI2
2025 RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
abstract
Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences.Typically, a reward model is trained or supplied to act as a proxy for humans in evaluating generated responses during the reinforcement training phase.However, current reward models operate as sequence-to-one models, allocating a single, sparse, and delayed reward to an entire output sequence.This approach may overlook the significant contributions of individual tokens toward the desired outcome.To this end, we propose a more finegrained, token-level guidance approach for RL training.Specifically, we introduce RED, a novel REward reDistribition method that evaluates and assigns specific credit to each token using an off-the-shelf reward model.Utilizing these fine-grained rewards enhances the model's understanding of language nuances, leading to more precise performance improvements.Notably, our method does not require modifying the reward model or introducing additional training steps, thereby incurring minimal computational costs.Experimental results across diverse datasets and tasks demonstrate the superiority of our approach.
Jiahui Li 0003, Lin Li 0065, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011, Cheng Yang 0002
EMNLP3
2024 Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning
abstract
Reinforcement learning from human feedback (RLHF) and AI-generated feedback (RLAIF) have become prominent techniques that significantly enhance the functionality of pre-trained language models (LMs).These methods harness feedback, sourced either from humans or AI, as direct rewards or to shape reward models that steer LM optimization.Nonetheless, the effective integration of rewards from diverse sources presents a significant challenge due to their disparate characteristics.To address this, recent research has developed algorithms incorporating strategies such as weighting, ranking, and constraining to handle this complexity.Despite these innovations, a bias toward disproportionately high rewards can still skew the reinforcement learning process and negatively impact LM performance.This paper explores a methodology for reward composition that enables simultaneous improvements in LMs across multiple dimensions.Inspired by fairness theory, we introduce a training algorithm that aims to reduce Disparity and enhance Stability among various rewards.Our method treats the aggregate reward as a dynamic weighted sum of individual rewards, with alternating updates to the weights and model parameters.For efficient and straightforward implementation, we employ an estimation technique rooted in the mirror descent method for weight updates, eliminating the need for gradient computations.The empirical results under various types of rewards across a wide range of scenarios demonstrate the effectiveness of our method.
Jiahui Li 0003, Fengda Zhang, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011
EMNLP4
2022 ESCM2: Entire Space Counterfactual Multi-Task Model for Post-Click Conversion Rate Estimation
abstract
Accurate estimation of post-click conversion rate is critical for building recommender systems, which has long been confronted with sample selection bias and data sparsity issues. Methods in the Entire Space Multi-task Model (ESMM) family leverage the sequential pattern of user actions, \ie $impression\rightarrow click \rightarrow conversion$ to address data sparsity issue. However, they still fail to ensure the unbiasedness of CVR estimates. In this paper, we theoretically demonstrate that ESMM suffers from the following two problems: (1) Inherent Estimation Bias (IEB) for CVR estimation, where the CVR estimate is inherently higher than the ground truth; (2) Potential Independence Priority (PIP) for CTCVR estimation, where ESMM might overlook the causality from click to conversion. To this end, we devise a principled approach named Entire Space Counterfactual Multi-task Modelling (ESCM$^2$), which employs a counterfactual risk miminizer as a regularizer in ESMM to address both IEB and PIP issues simultaneously. Extensive experiments on offline datasets and online environments demonstrate that our proposed ESCM$^2$ can largely mitigate the inherent IEB and PIP issues and achieve better performance than baseline models.
Hao Wang 0049, Tai-Wei Chang, Tianqiao Liu, Jianmin Huang, Zhichao Chen 0001, Ruopeng Li
SIGIR2
2014 Interpretation of Chinese Discourse Connectives for Explicit Discourse Relation Recognition
Hen-Hsen Huang, Tai-Wei Chang, Huan-Yuan Chen, Hsin-Hsi Chen
COLING2