Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zijing Hu

dblp:12/1065 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
2since 2021 · last 2025
0009-0007-6167-3996ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 36% Reinforcement learning · 23% Vision and language · 23%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.722025
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025
Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards · CVPR 2025
Natural language and speech › Language models and text generation
preference optimization
0.912025
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025
Computer vision › Vision and language › cross-modal alignment › image-text alignment
text-to-image alignment
0.912025
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.312025
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples · ICML 2025

Methods — techniques the papers use, named apart from their topics

policy optimization · 0.9mask-guided self-attention fusion · 0.9direct preference optimization · 0.9branch-based sampling · 0.9backward progressive training · 0.9
YearPublicationVenuePosition
2025 Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards
abstract
Diffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and corresponding text prompts. To tackle this issue, reinforcement learning (RL) has been considered for diffusion model fine-tuning. Yet, RL’s effectiveness is limited by the challenge of sparse reward, where feedback is only available at the end of the generation process. This makes it difficult to identify which actions during the de-noising process contribute positively to the final generated image, potentially leading to ineffective or unnecessary de-noising policies. To this end, this paper presents a novel RL-based framework that addresses the sparse reward problem when training diffusion models. Our framework, named B2-DiffuRL, employs two strategies: Backward progressive training and Branch-based sampling. For one thing, backward progressive training focuses initially on the final timesteps of denoising process and gradually extends the training interval to earlier timesteps, easing the learning difficulty from sparse rewards. For another, we perform branch-based sampling for each training interval. By comparing the samples within the same branch, we can identify how much the policies of the current training interval contribute to the final image, which helps to learn effective policies instead of unnecessary ones. B2-DiffuRL is compatible with existing optimization algorithms. Extensive experiments demonstrate the effectiveness of B2-DiffuRL in improving prompt-image alignment and maintaining diversity in generated images. The code for this work is available1.
Zijing Hu, Fengda Zhang, Long Chen 0016, Kun Kuang 0001, Jiahui Li 0003, Kaifeng Gao, Jun Xiao 0001, Xin Wang 0019, Wenwu Zhu 0001
CVPR1
2025 D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples
abstract
The practical applications of diffusion models have been limited by the misalignment between generated images and corresponding text prompts. Recent studies have introduced direct preference optimization (DPO) to enhance the alignment of these models. However, the effectiveness of DPO is constrained by the issue of visual inconsistency, where the significant visual disparity between well-aligned and poorly-aligned images prevents diffusion models from identifying which factors contribute positively to alignment during fine-tuning. To address this issue, this paper introduces D-Fusion, a method to construct DPO-trainable visually consistent samples. On one hand, by performing mask-guided self-attention fusion, the resulting images are not only well-aligned, but also visually consistent with given poorly-aligned images. On the other hand, D-Fusion can retain the denoising trajectories of the resulting images, which are essential for DPO training. Extensive experiments demonstrate the effectiveness of D-Fusion in improving prompt-image alignment when applied to different reinforcement learning algorithms.
Zijing Hu, Fengda Zhang, Kun Kuang 0001
ICML1
2006 DSEC: A Data Stream Engine Based Clinical Information System
Hongyan Li 0002, Zijing Hu, Jianlong Gao, Shiwei Tang, Xinbiao Zhou
APWeb3