VLDB 2026 Research / reviewers in the wild / expert
Zijing Hu
dblp:12/1065
· DBLP profile ↗
3ranked-venue papers
2as first author
2since 2021 · last 2025
0009-0007-6167-3996ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 36% Reinforcement learning · 23% Vision and language · 23% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.7 | 2 | 2025 | D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025 Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards · CVPR 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025 |
Computer vision › Vision and language › cross-modal alignment › image-text alignment
text-to-image alignment |
0.9 | 1 | 2025 | D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.3 | 1 | 2025 | D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
policy optimization · 0.9mask-guided self-attention fusion · 0.9direct preference optimization · 0.9branch-based sampling · 0.9backward progressive training · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse RewardsabstractDiffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and corresponding text prompts. To tackle this issue, reinforcement learning (RL) has been considered for diffusion model fine-tuning. Yet, RL’s effectiveness is limited by the challenge of sparse reward, where feedback is only available at the end of the generation process. This makes it difficult to identify which actions during the de-noising process contribute positively to the final generated image, potentially leading to ineffective or unnecessary de-noising policies. To this end, this paper presents a novel RL-based framework that addresses the sparse reward problem when training diffusion models. Our framework, named B2-DiffuRL, employs two strategies: Backward progressive training and Branch-based sampling. For one thing, backward progressive training focuses initially on the final timesteps of denoising process and gradually extends the training interval to earlier timesteps, easing the learning difficulty from sparse rewards. For another, we perform branch-based sampling for each training interval. By comparing the samples within the same branch, we can identify how much the policies of the current training interval contribute to the final image, which helps to learn effective policies instead of unnecessary ones. B2-DiffuRL is compatible with existing optimization algorithms. Extensive experiments demonstrate the effectiveness of B2-DiffuRL in improving prompt-image alignment and maintaining diversity in generated images. The code for this work is available1. Zijing Hu, Fengda Zhang, Long Chen 0016, Kun Kuang 0001, Jiahui Li 0003, Kaifeng Gao, Jun Xiao 0001, Xin Wang 0019, Wenwu Zhu 0001 |
CVPR | 1 |
| 2025 | D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent SamplesabstractThe practical applications of diffusion models have been limited by the misalignment between generated images and corresponding text prompts. Recent studies have introduced direct preference optimization (DPO) to enhance the alignment of these models. However, the effectiveness of DPO is constrained by the issue of visual inconsistency, where the significant visual disparity between well-aligned and poorly-aligned images prevents diffusion models from identifying which factors contribute positively to alignment during fine-tuning. To address this issue, this paper introduces D-Fusion, a method to construct DPO-trainable visually consistent samples. On one hand, by performing mask-guided self-attention fusion, the resulting images are not only well-aligned, but also visually consistent with given poorly-aligned images. On the other hand, D-Fusion can retain the denoising trajectories of the resulting images, which are essential for DPO training. Extensive experiments demonstrate the effectiveness of D-Fusion in improving prompt-image alignment when applied to different reinforcement learning algorithms. Zijing Hu, Fengda Zhang, Kun Kuang 0001 |
ICML | 1 |
| 2006 | DSEC: A Data Stream Engine Based Clinical Information System
Hongyan Li 0002, Zijing Hu, Jianlong Gao, Shiwei Tang, Xinbiao Zhou |
APWeb | 3 |