VLDB 2026 Research / reviewers in the wild / expert
Fengda Zhang
dblp:255/0136
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0001-5280-413XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Generative modeling · 34% Trustworthy machine learning · 23% Reinforcement learning · 12% | |
| Computer networks
1 paper |
Edge and fog computing · 100% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.6 | 3 | 2025 | D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025 Decoding Correlation-Induced Misalignment in the Stable Diffusion Workflow for Text-to-Image Generation · ICCV 2025 Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards · CVPR 2025 |
Machine learning › Trustworthy machine learning
fairness |
2.2 | 3 | 2024 | Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning · EMNLP 2024 Distributionally Generative Augmentation for Fair Facial Attribute Classification · CVPR 2024 Fairness-aware Contrastive Learning with Partially Annotated Sensitive Attributes · ICLR 2023 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.0 | 2 | 2025 | Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning · EMNLP 2024 D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular Data · ICML 2025 |
Machine learning › Generative modeling
score-based model |
0.9 | 1 | 2025 | Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular Data · ICML 2025 |
Machine learning › Generative modeling › diffusion model › latent diffusion model
stable diffusion |
0.9 | 1 | 2025 | Decoding Correlation-Induced Misalignment in the Stable Diffusion Workflow for Text-to-Image Generation · ICCV 2025 |
Computer vision › Vision and language › cross-modal alignment › image-text alignment
text-to-image alignment |
0.9 | 1 | 2025 | D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | Decoding Correlation-Induced Misalignment in the Stable Diffusion Workflow for Text-to-Image Generation · ICCV 2025 |
Computer vision › Face, body and person analysis › facial attribute analysis
facial attribute recognition |
0.8 | 1 | 2024 | Distributionally Generative Augmentation for Fair Facial Attribute Classification · CVPR 2024 |
Machine learning › Trustworthy machine learning › fairness › fair classification
fair facial attribute classification |
0.8 | 1 | 2024 | Distributionally Generative Augmentation for Fair Facial Attribute Classification · CVPR 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.7 | 1 | 2023 | Fairness-aware Contrastive Learning with Partially Annotated Sensitive Attributes · ICLR 2023 |
Machine learning › Efficient and distributed learning › distributed training › edge training
device-cloud collaborative learning |
0.7 | 1 | 2023 | Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023 |
Machine learning › Efficient and distributed learning
distributed training |
0.7 | 1 | 2023 | Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023 |
Edge and fog computing
edge-cloud collaboration |
0.7 | 1 | 2023 | Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023 |
Edge and fog computing › distributed learning
edge-cloud collaborative learning |
0.7 | 1 | 2023 | Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023 |
Machine learning › Graph learning
graph neural network |
0.2 | 1 | 2023 | Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023 |
Machine learning › Transfer learning and domain adaptation
pre-trained models |
0.2 | 1 | 2023 | Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AI · IEEE Trans. Knowl. Data Eng. 2023 |
Methods — techniques the papers use, named apart from their topics
score-based model · 0.9policy optimization · 0.9mask-guided self-attention fusion · 0.9direct preference optimization · 0.9density estimation · 0.9branch-based sampling · 0.9backward progressive training · 0.9generative augmentation · 0.8dynamic weighted sum · 0.8diffusion model · 0.8knowledge distillation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse RewardsabstractDiffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and corresponding text prompts. To tackle this issue, reinforcement learning (RL) has been considered for diffusion model fine-tuning. Yet, RL’s effectiveness is limited by the challenge of sparse reward, where feedback is only available at the end of the generation process. This makes it difficult to identify which actions during the de-noising process contribute positively to the final generated image, potentially leading to ineffective or unnecessary de-noising policies. To this end, this paper presents a novel RL-based framework that addresses the sparse reward problem when training diffusion models. Our framework, named B2-DiffuRL, employs two strategies: Backward progressive training and Branch-based sampling. For one thing, backward progressive training focuses initially on the final timesteps of denoising process and gradually extends the training interval to earlier timesteps, easing the learning difficulty from sparse rewards. For another, we perform branch-based sampling for each training interval. By comparing the samples within the same branch, we can identify how much the policies of the current training interval contribute to the final image, which helps to learn effective policies instead of unnecessary ones. B2-DiffuRL is compatible with existing optimization algorithms. Extensive experiments demonstrate the effectiveness of B2-DiffuRL in improving prompt-image alignment and maintaining diversity in generated images. The code for this work is available1. Zijing Hu, Fengda Zhang, Long Chen 0016, Kun Kuang 0001, Jiahui Li 0003, Kaifeng Gao, Jun Xiao 0001, Xin Wang 0019, Wenwu Zhu 0001 |
CVPR | 2 |
| 2025 | Decoding Correlation-Induced Misalignment in the Stable Diffusion Workflow for Text-to-Image Generation
Yunze Tong, Fengda Zhang, Didi Zhu, Jun Xiao 0001, Kun Kuang 0001 |
ICCV | 2 |
| 2025 | D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent SamplesabstractThe practical applications of diffusion models have been limited by the misalignment between generated images and corresponding text prompts. Recent studies have introduced direct preference optimization (DPO) to enhance the alignment of these models. However, the effectiveness of DPO is constrained by the issue of visual inconsistency, where the significant visual disparity between well-aligned and poorly-aligned images prevents diffusion models from identifying which factors contribute positively to alignment during fine-tuning. To address this issue, this paper introduces D-Fusion, a method to construct DPO-trainable visually consistent samples. On one hand, by performing mask-guided self-attention fusion, the resulting images are not only well-aligned, but also visually consistent with given poorly-aligned images. On the other hand, D-Fusion can retain the denoising trajectories of the resulting images, which are essential for DPO training. Extensive experiments demonstrate the effectiveness of D-Fusion in improving prompt-image alignment when applied to different reinforcement learning algorithms. Zijing Hu, Fengda Zhang, Kun Kuang 0001 |
ICML | 2 |
| 2025 | Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular DataabstractMachine learning models often perform well on tabular data by optimizing average prediction accuracy. However, they may underperform on specific subsets due to inherent biases and spurious correlations in the training data, such as associations with non-causal features like demographic information. These biases lead to critical robustness issues as models may inherit or amplify them, resulting in poor performance where such misleading correlations do not hold. Existing mitigation methods have significant limitations: some require prior group labels, which are often unavailable, while others focus solely on the conditional distribution $P(Y|X)$, upweighting misclassified samples without effectively balancing the overall data distribution $P(X)$. To address these shortcomings, we propose a latent score-based reweighting framework. It leverages score-based models to capture the joint data distribution $P(X, Y)$ without relying on additional prior information. By estimating sample density through the similarity of score vectors with neighboring data points, our method identifies underrepresented regions and upweights samples accordingly. This approach directly tackles inherent data imbalances, enhancing robustness by ensuring a more uniform dataset representation. Experiments on various tabular datasets under distribution shifts demonstrate that our method effectively improves performance on imbalanced data. Yunze Tong, Fengda Zhang, Kaifeng Gao, Pengfei Lyu, Jun Xiao 0001, Kun Kuang 0001 |
ICML | 2 |
| 2024 | Distributionally Generative Augmentation for Fair Facial Attribute ClassificationabstractFacial Attribute Classification (FAC) holds substantial promise in widespread applications. However, FAC models trained by traditional methodologies can be unfair by exhibiting accuracy inconsistencies across varied data sub-populations. This unfairness is largely attributed to bias in data, where some spurious attributes (e.g., Male) statistically correlate with the target attribute (e.g., Smiling). Most of existing fairness-aware methods rely on the labels of spurious attributes, which may be unavailable in practice. This work proposes a novel, generation-based two-stage framework to train a fair FAC model on biased data without additional annotation. Initially, we identify the potential spurious attributes based on generative models. Notably, it enhances interpretability by explicitly showing the spurious attributes in image space. Following this, for each image, we first edit the spurious attributes with a random degree sampled from a uniform distribution, while keeping target attribute unchanged. Then we train a fair FAC model by fostering model invariance to these augmentation. Extensive experiments on three common datasets demonstrate the effectiveness of our method in promoting fairness in FAC without compromising accuracy. Codes are in https://github.com/heqianpei/DiGA. Fengda Zhang, Qianpei He, Kun Kuang 0001, Long Chen 0016, Chao Wu 0001, Jun Xiao 0001, Hanwang Zhang |
CVPR | 1 |
| 2024 | Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement LearningabstractReinforcement learning from human feedback (RLHF) and AI-generated feedback (RLAIF) have become prominent techniques that significantly enhance the functionality of pre-trained language models (LMs).These methods harness feedback, sourced either from humans or AI, as direct rewards or to shape reward models that steer LM optimization.Nonetheless, the effective integration of rewards from diverse sources presents a significant challenge due to their disparate characteristics.To address this, recent research has developed algorithms incorporating strategies such as weighting, ranking, and constraining to handle this complexity.Despite these innovations, a bias toward disproportionately high rewards can still skew the reinforcement learning process and negatively impact LM performance.This paper explores a methodology for reward composition that enables simultaneous improvements in LMs across multiple dimensions.Inspired by fairness theory, we introduce a training algorithm that aims to reduce Disparity and enhance Stability among various rewards.Our method treats the aggregate reward as a dynamic weighted sum of individual rewards, with alternating updates to the weights and model parameters.For efficient and straightforward implementation, we employ an estimation technique rooted in the mirror descent method for weight updates, eliminating the need for gradient computations.The empirical results under various types of rewards across a wide range of scenarios demonstrate the effectiveness of our method. Jiahui Li 0003, Fengda Zhang, Tai-Wei Chang, Kun Kuang 0001, Long Chen 0016, Jun Zhou 0011 |
EMNLP | 3 |
| 2023 | Fairness-aware Contrastive Learning with Partially Annotated Sensitive Attributes
Fengda Zhang, Kun Kuang 0001, Long Chen 0016, Chao Wu 0001, Jun Xiao 0001 |
ICLR | 1 |
| 2023 | Federated mutual learning: a collaborative machine learning method for heterogeneous data, models, and objectivesabstractFederated learning (FL) is a novel technique in deep learning that enables clients to collaboratively train a shared model while retaining their decentralized data. However, researchers working on FL face several unique challenges, especially in the context of heterogeneity. Heterogeneity in data distributions, computational capabilities, and scenarios among clients necessitates the development of customized models and objectives in FL. Unfortunately, existing works such as FedAvg may not effectively accommodate the specific needs of each client. To address the challenges arising from heterogeneity in FL, we provide an overview of the heterogeneities in data, model, and objective (DMO). Furthermore, we propose a novel framework called federated mutual learning (FML), which enables each client to train a personalized model that accounts for the data heterogeneity (DH). A “meme model” serves as an intermediary between the personalized and global models to address model heterogeneity (MH). We introduce a knowledge distillation technique called deep mutual learning (DML) to transfer knowledge between these two models on local data. To overcome objective heterogeneity (OH), we design a shared global model that includes only certain parts, and the personalized model is task-specific and enhanced through mutual learning with the meme model. We evaluate the performance of FML in addressing DMO heterogeneities through experiments and compare it with other commonly used FL methods in similar scenarios. The results demonstrate that FML outperforms other methods and effectively addresses the DMO challenges encountered in the FL setting. Tao Shen 0002, Jie Zhang 0081, Xinkang Jia, Fengda Zhang, Zheqi Lv, Kun Kuang 0001, Chao Wu 0001, Fei Wu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2023 | Federated unsupervised representation learningabstractTo leverage the enormous amount of unlabeled data on distributed edge devices, we formulate a new problem in federated learning called federated unsupervised representation learning (FURL) to learn a common representation model without supervision while preserving data privacy. FURL poses two new challenges: (1) data distribution shift (non-independent and identically distributed, non-IID) among clients would make local models focus on different categories, leading to the inconsistency of representation spaces; (2) without unified information among the clients in FURL, the representations across clients would be misaligned. To address these challenges, we propose the federated contrastive averaging with dictionary and alignment (FedCA) algorithm. FedCA is composed of two key modules: a dictionary module to aggregate the representations of samples from each client which can be shared with all clients for consistency of representation space and an alignment module to align the representation of each client on a base model trained on public data. We adopt the contrastive approach for local model training. Through extensive experiments with three evaluation protocols in IID and non-IID settings, we demonstrate that FedCA outperforms all baselines with significant margins. Fengda Zhang, Kun Kuang 0001, Long Chen 0016, Zhaoyang You, Tao Shen 0002, Jun Xiao 0001, Yin Zhang 0006, Chao Wu 0001, Fei Wu 0001, Yueting Zhuang |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2023 | Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AIabstractInfluenced by the great success of deep learning via cloud computing and the rapid development of edge chips, research in artificial intelligence (AI) has shifted to both of the computing paradigms, i.e., cloud computing and edge computing. In recent years, we have witnessed significant progress in developing more advanced AI models on cloud servers that surpass traditional deep learning models owing to model innovations (e.g., Transformers, Pretrained families), explosion of training data and soaring computing capabilities. However, edge computing, especially edge and cloud collaborative computing, are still in its infancy to announce their success due to the resource-constrained IoT scenarios with very limited algorithms deployed. In this survey, we conduct a systematic review for both cloud and edge AI. Specifically, we are the first to set up the collaborative learning mechanism for cloud and edge modeling with a thorough review of the architectures that enable such mechanism. We also discuss potentials and practical experiences of some on-going advanced edge AI topics including pretraining models, graph neural networks and reinforcement learning. Finally, we discuss the promising directions and challenges in this field. Jiangchao Yao, Shengyu Zhang 0001, Feng Wang 0072, Jianwei Zhang 0012, Yunfei Chu, Luo Ji, Kunyang Jia, Tao Shen 0002, Anpeng Wu, Fengda Zhang, Kun Kuang 0001, Chao Wu 0001, Fei Wu 0001, Jingren Zhou 0001, Hongxia Yang |
IEEE Trans. Knowl. Data Eng. | 12 |